openmoji-g1-corruption-sweep-71a080a-4levels-9b9b1699
completed falsified
Hypothesis
Two mechanisms are in tension and this run adjudicates between them. The regime hypothesis says the fixed 0.35 corruption probability destroys information no model can recover, so one fixed checkpoint evaluated at lower corruption should close a much larger fraction of the render gap between `x_t` and `x_0`. The damage hypothesis says the opposite: held-out retained-token accuracy is only 0.4128, meaning the model overwrites about 59% of the fields that were already correct, and at low corruption there are far more correct fields to damage, so the recovered fraction could fall to zero or go negative and the model could make its input measurably worse. The prediction registered here is the regime hypothesis, because it is the one the render probe motivated and the one the project's next experiment depends on.
Why run it
v2 eliminated data volume and v3 eliminated model capacity, leaving the corruption schedule as the only standing candidate. Before spending training runs on it, this establishes whether the model has any regime in which it is useful at all. It needs no training: one fixed checkpoint, four evaluation levels, CPU only. If recovery improves sharply at lower corruption, a training-schedule experiment is justified. If the model damages its input instead, the binding problem is retained-token preservation and no corruption schedule will fix it.
Scope limits
One checkpoint, one seed, 32 held-out icons per level, one corruption draw per icon. The model trained only at 0.35, so evaluating it at 0.05 is off-distribution by construction; that is the point, but it means a poor result at low corruption does not by itself condemn a model trained there.
Written result
v2 eliminated data volume as the Gate G constraint and v3 eliminated model capacity, both on matched predeclared comparisons. The corruption schedule was the only candidate left, and the render probe had suggested why: at probability 0.35 the input x_t is already visually destroyed, so perhaps no model can recover information that is gone.
This run tests that before spending any training on it. One fixed checkpoint — v2's loss-selected step 840 — evaluated and rendered at four corruption levels, with 0.35 as the control because that is what it trained at. Identical icons, sizes and corruption seeds at every level; only the Bernoulli rate differs. No training.
The statistic is the per-icon recovery fraction, (MAE(x_t) - MAE(x_hat_0)) / MAE(x_t) — the share of the render gap the model closes. Absolute error is not comparable across levels, because x_t itself moves closer to x_0 as corruption falls.
Result
| corruption p | median x_t error | median x_hat_0 error | mean recovery fraction | 95% CI | icons helped |
|---|---|---|---|---|---|
| 0.05 | 0.04973 | 0.13853 | -4.6517 | -7.672 .. -1.632 | 0 of 32 |
| 0.10 | 0.08116 | 0.13914 | -1.2024 | -2.118 .. -0.287 | 4 of 32 |
| 0.20 | 0.13864 | 0.13717 | -0.1315 | -0.285 .. +0.022 | 14 of 32 |
| 0.35 (control) | 0.17606 | 0.14045 | +0.1800 | +0.103 .. +0.257 | 26 of 32 |
At 18 px the picture is the same: -4.07, -1.12, -0.12, +0.19.
Predeclared criteria
| criterion | threshold | observed | outcome |
|---|---|---|---|
| recovery_improves_at_lower_corruption | recovery at 0.10 >= 2x recovery at 0.35 | -1.2024 against +0.3600 | falsified |
| monotone_in_corruption | non-increasing across 0.05, 0.10, 0.20, 0.35 | -4.65, -1.20, -0.13, +0.18 — strictly increasing | falsified |
| a_useful_regime_exists | some level exceeds 0.50 recovery | best is +0.18 | falsified |
| reproducibility | identical artifacts per level | identical | pass |
The regime hypothesis is falsified and the competing damage hypothesis — registered in the same record before launch — is confirmed. At probability 0.05 the model makes its input 2.8x worse, and it helps on zero of thirty-two icons.
Why: the output barely depends on the input
The decisive number is not in the table above. It is that x_hat_0 error hardly moves at all while x_t error moves by a factor of 3.5:
| corruption p | mean x_t error | mean x_hat_0 error |
|---|---|---|
| 0.05 | 0.05098 | 0.13822 |
| 0.10 | 0.08739 | 0.13820 |
| 0.20 | 0.13592 | 0.14728 |
| 0.35 | 0.17958 | 0.14651 |
Per icon, the prediction at p=0.05 correlates with the prediction at p=0.35 at Pearson r = 0.91, with a mean absolute difference of 0.019 against a mean error of 0.14. The model emits nearly the same reconstruction for a given icon whether 5% or 35% of its geometry has been replaced.
It is not denoising. It has learned a group- and subgroup-conditioned prior over icon geometry and emits approximately that, largely ignoring x_t.
The architecture makes this easy to believe. GeometryDenoiser takes the corrupted tokens plus group and subgroup embeddings and nothing else: there is no noise-level or timestep input and no mask marking which fields were replaced. Trained at a single fixed probability of 0.35, it never needed to behave differently at any other level, and nothing in its input tells it how much to trust what it is given.
This corrects how the earlier Gate G numbers should be read
The measurements stand; their interpretation does not.
- The +18% recovery at p=0.35 in the first render probe is not the model recovering geometry. It is a fixed-quality output happening to beat a badly corrupted baseline. The same output scored against a lightly corrupted baseline loses badly.
- Held-out retained-token accuracy of 0.4128 was never a secondary weakness. It was this finding, visible in token space the whole time: the model overwrites correct fields because it is not reading them. The standing 0.90 bar has been measuring the real problem since v1.
- The held-out loss floor near 7.35 that v2 and a 3.53x larger v3 both converge to is consistent with both models learning the same prior. Capacity would not help, and measurably did not.
- v1's 18x-above-untrained result and v2's 25% loss reduction remain real learning. What was learned is a prior over icons, not a conditional denoiser.
What this redirects the project toward
Not a corruption schedule. The falsification meaning was registered before launch and it holds: the binding problem is that the model cannot leave correct fields alone, so the next experiment is an objective or architecture that lets it.
The concrete, cheap candidates, in the order their cost suggests:
- Tell the model the noise level. A timestep or corruption-level embedding, and training across a range of levels rather than one fixed point. Without it the model cannot modulate how much to trust its input, which is exactly the failure here.
- Make copying easy. Predict an edit mask, or a residual against
x_t, so that "leave this field alone" is the default rather than something the model must reconstruct token by token. - Condition on which fields were corrupted, as a diagnostic upper bound. That is not a realistic sampling-time signal, but it cleanly separates "cannot identify the corrupted fields" from "cannot predict their values", and that is worth knowing before designing around either.
The first is the smallest change that addresses the mechanism observed here, and it is the one this project should run next.
Scope
One checkpoint, one seed, 32 held-out icons per level, one corruption draw per icon. The model trained only at 0.35, so evaluating it at 0.05 is off-distribution by construction — that is the point of the sweep, but it means a poor result at low corruption does not by itself condemn a model that was trained across levels. Establishing that is exactly what candidate 1 above would do.
State transitions
- planned2026-09-20T19:00:00Z
- completed2026-09-20T19:15:00Zmodel_output_is_near_independent_of_its_input_it_learned_a_prior_not_a_conditional_denoiser
Run record
Verbatim from runs/openmoji-g1-corruption-sweep-71a080a-4levels-9b9b1699/run.yaml, the record committed before launch.
e791493769907ac4…9b9b1699677a6f97…