tiny-geometry-v3-93ed354-ff42f4dd-32a80ab5
completed —
Hypothesis
The diverse-four aggregate held-out accuracy materially overstates genuine denoising: accuracy on actually changed coordinate tokens will trail aggregate accuracy by at least 0.15, while all v2 learning outputs remain otherwise unchanged.
Why run it
Determines whether the next Gate E experiment should target corrupted-token recovery rather than model capacity, additional optimization steps, or topology.
Measured behaviour
Table view
| optimizer step | train accuracy |
|---|---|
| 1 | 0.0012 |
| 10 | 0.2892 |
| 20 | 0.6366 |
| 30 | 0.8676 |
| 40 | 0.9694 |
| 50 | 0.9975 |
| 60 | 1.0000 |
| 70 | 1.0000 |
| 80 | 1.0000 |
| 90 | 1.0000 |
| 100 | 1.0000 |
| 110 | 1.0000 |
| 120 | 1.0000 |
| 130 | 1.0000 |
| 140 | 1.0000 |
| 150 | 1.0000 |
| 160 | 1.0000 |
| 1 | 0.0028 |
| 10 | 0.1141 |
| 20 | 0.4179 |
| 30 | 0.7195 |
| 40 | 0.8920 |
| 50 | 0.9649 |
| 60 | 0.9948 |
| 70 | 0.9993 |
| 80 | 1.0000 |
| 90 | 1.0000 |
| 100 | 1.0000 |
| 110 | 1.0000 |
| 120 | 1.0000 |
| 130 | 1.0000 |
| 140 | 1.0000 |
| 150 | 1.0000 |
| 160 | 1.0000 |
| 170 | 1.0000 |
| 180 | 1.0000 |
| 190 | 1.0000 |
| 200 | 1.0000 |
| 210 | 1.0000 |
| 220 | 1.0000 |
| 230 | 1.0000 |
| 240 | 1.0000 |
| 250 | 1.0000 |
| 260 | 1.0000 |
| 270 | 1.0000 |
| 280 | 1.0000 |
| 290 | 1.0000 |
| 300 | 1.0000 |
| 310 | 1.0000 |
| 320 | 1.0000 |
Measured result
Verbatim from the run's summary.json.
52aee590c65f81f5…240bd2ca34fd1b1a…bd0b0814a19c7402…240bd2ca34fd1b1a…8e0a9878092058a0…8f30cf4c834873bd…9b12879ba3a18236…0d02b532b1dbe573…578ec1d402fdb88f…ff42f4dd0b8915f0…32a80ab576a6d5a4…c4e0d9a8c7873466…d4c36c75381689fb…Visual output
Written result
Completed reproducibly; the predeclared diagnostic passed.
For diverse-four, actually-changed-token accuracy is 58.17%, 15.65 points below the 73.82% aggregate. Retained-token accuracy is also limited at 82.52%, so the model often overwrites coordinates already correct in x_t. The one-icon case is much stronger: 88.59% changed and 96.11% retained accuracy. A complete rerun reproduced all artifacts.
This supports testing broader deterministic corruption coverage at the same batch size, model, probability, and optimizer-step count before increasing capacity or steps.
State transitions
- planned2026-08-30T05:55:54.629673Z
- completed2026-08-30T06:01:12.713902Zaggregate_accuracy_materially_overstates_changed_token_recovery_and_diverse_model_also_overwrites_retained_coordinates
Run record
Verbatim from runs/tiny-geometry-v3-93ed354-ff42f4dd-32a80ab5/run.yaml, the record committed before launch.
e3b0c44298fc1c14…ff42f4dd0b8915f0…32a80ab576a6d5a4…