tiny-geometry-v5-44ce3de-a34456a8-32a80ab5
completed declared
Hypothesis
Applying the v4 deterministic per-step coverage treatment to one-icon as well will make both predeclared cases pass while preserving diverse recovery, exact checkpoint continuation, recognizable renders, and byte-identical artifacts.
Why run it
Final controlled check for Gate E's one-icon overfit, tiny-diverse recovery, resume, and rendering exit evidence before beginning corruption-family comparisons.
Measured behaviour
Table view
| optimizer step | train accuracy |
|---|---|
| 1 | 0.0018 |
| 10 | 0.2672 |
| 20 | 0.5858 |
| 30 | 0.7855 |
| 40 | 0.8848 |
| 50 | 0.9161 |
| 60 | 0.9583 |
| 70 | 0.9645 |
| 80 | 0.9675 |
| 90 | 0.9626 |
| 100 | 0.9877 |
| 110 | 0.9755 |
| 120 | 0.9835 |
| 130 | 0.9810 |
| 140 | 0.9828 |
| 150 | 0.9896 |
| 160 | 0.9920 |
| 1 | 0.0020 |
| 10 | 0.0804 |
| 20 | 0.3095 |
| 30 | 0.5623 |
| 40 | 0.7269 |
| 50 | 0.7879 |
| 60 | 0.8441 |
| 70 | 0.8789 |
| 80 | 0.9031 |
| 90 | 0.9262 |
| 100 | 0.9294 |
| 110 | 0.9299 |
| 120 | 0.9503 |
| 130 | 0.9456 |
| 140 | 0.9608 |
| 150 | 0.9639 |
| 160 | 0.9682 |
| 170 | 0.9671 |
| 180 | 0.9658 |
| 190 | 0.9678 |
| 200 | 0.9793 |
| 210 | 0.9699 |
| 220 | 0.9680 |
| 230 | 0.9795 |
| 240 | 0.9760 |
| 250 | 0.9811 |
| 260 | 0.9791 |
| 270 | 0.9782 |
| 280 | 0.9782 |
| 290 | 0.9854 |
| 300 | 0.9837 |
| 310 | 0.9863 |
| 320 | 0.9852 |
Measured result
Verbatim from the run's summary.json.
e975f17c15497260…1c909c13042da5ba…d7972e2a714662fd…1c909c13042da5ba…7029c413d5bdce0c…8f30cf4c834873bd…9b12879ba3a18236…0d02b532b1dbe573…8e09bccf8d3c1b90…a34456a8e5087700…32a80ab576a6d5a4…2805101702954e39…8eafdf8ce49f84c2…Visual output
Written result
Completed reproducibly with a mixed per-run result and positive cumulative Gate E exit.
One-icon held-out accuracy rises to 99.26% (98.37% changed, 99.72% retained), and diverse-four remains at 98.56%. The fixed one-icon probe is 97.73%, below its legacy 99% memorization threshold, so v5's literal both-criteria hypothesis is false. That probe is an unseen batch under online resampling, not the batch v1 overfit.
Across the controlled sequence, Gate E is complete: v1 records 100% one-icon overfit; v5 records strong one-icon held-out recovery; v4/v5 record strong diverse recovery and recognizable renders; v2-v5 record exact checkpoint continuation and artifact identity.
State transitions
- planned2026-08-30T06:13:13.901491Z
- completed2026-08-30T06:19:09.284898Zone_icon_and_diverse_heldout_recovery_are_strong_but_v5_fixed_probe_misses_legacy_memorization_threshold_gate_e_closes_from_controlled_sequence
Run record
Verbatim from runs/tiny-geometry-v5-44ce3de-a34456a8-32a80ab5/run.yaml, the record committed before launch.
e3b0c44298fc1c14…a34456a8e5087700…32a80ab576a6d5a4…