tiny-geometry-v4-ede1009-7ccbc64e-32a80ab5
completed declared
Hypothesis
Deterministic per-step corruption resampling for diverse-four will raise held-out changed-token accuracy to at least 0.75 and retained-token accuracy to at least 0.90 without increasing model size, batch size, corruption probability, or optimizer steps.
Why run it
Tests whether v3 is limited primarily by seeing only four fixed corruptions per icon, before spending evidence budget on more capacity or longer training.
Measured behaviour
Table view
| optimizer step | train accuracy |
|---|---|
| 1 | 0.0012 |
| 10 | 0.2892 |
| 20 | 0.6366 |
| 30 | 0.8676 |
| 40 | 0.9694 |
| 50 | 0.9975 |
| 60 | 1.0000 |
| 70 | 1.0000 |
| 80 | 1.0000 |
| 90 | 1.0000 |
| 100 | 1.0000 |
| 110 | 1.0000 |
| 120 | 1.0000 |
| 130 | 1.0000 |
| 140 | 1.0000 |
| 150 | 1.0000 |
| 160 | 1.0000 |
| 1 | 0.0020 |
| 10 | 0.0804 |
| 20 | 0.3095 |
| 30 | 0.5623 |
| 40 | 0.7269 |
| 50 | 0.7879 |
| 60 | 0.8441 |
| 70 | 0.8789 |
| 80 | 0.9031 |
| 90 | 0.9262 |
| 100 | 0.9294 |
| 110 | 0.9299 |
| 120 | 0.9503 |
| 130 | 0.9456 |
| 140 | 0.9608 |
| 150 | 0.9639 |
| 160 | 0.9682 |
| 170 | 0.9671 |
| 180 | 0.9658 |
| 190 | 0.9678 |
| 200 | 0.9793 |
| 210 | 0.9699 |
| 220 | 0.9680 |
| 230 | 0.9795 |
| 240 | 0.9760 |
| 250 | 0.9811 |
| 260 | 0.9791 |
| 270 | 0.9782 |
| 280 | 0.9782 |
| 290 | 0.9854 |
| 300 | 0.9837 |
| 310 | 0.9863 |
| 320 | 0.9852 |
Measured result
Verbatim from the run's summary.json.
e975f17c15497260…1c909c13042da5ba…d7972e2a714662fd…1c909c13042da5ba…8e0a9878092058a0…8f30cf4c834873bd…9b12879ba3a18236…0d02b532b1dbe573…8e09bccf8d3c1b90…7ccbc64e51a0f009…32a80ab576a6d5a4…e11c8598e51d6913…8eafdf8ce49f84c2…Visual output
Written result
Completed reproducibly; the diverse corruption-coverage hypothesis passed.
At unchanged batch size, model, probability, and step count, deterministic per-step resampling raised diverse-four held-out accuracy from 73.82% to 98.56%. Changed-token accuracy rose from 58.17% to 96.65%; retained-token accuracy rose from 82.52% to 99.63%. Checkpoint continuation and the full artifact rerun are exact. The paired renders are recognizable at both target sizes.
The unchanged static one-icon control remains below its added held-out threshold, so the generic combined gate flag remains false. Apply the same treatment there before closing Gate E.
State transitions
- planned2026-08-30T06:05:24.460858Z
- completed2026-08-30T06:11:11.007462Zper_step_corruption_coverage_recovers_diverse_heldout_geometry_without_more_model_batch_or_steps
Run record
Verbatim from runs/tiny-geometry-v4-ede1009-7ccbc64e-32a80ab5/run.yaml, the record committed before launch.
e3b0c44298fc1c14…7ccbc64e51a0f009…32a80ab576a6d5a4…