MojiDiff

← experiments

openmoji-g1-corruption-sweep-71a080a-4levels-9b9b1699

completed falsified

Hypothesis

Two mechanisms are in tension and this run adjudicates between them. The regime hypothesis says the fixed 0.35 corruption probability destroys information no model can recover, so one fixed checkpoint evaluated at lower corruption should close a much larger fraction of the render gap between `x_t` and `x_0`. The damage hypothesis says the opposite: held-out retained-token accuracy is only 0.4128, meaning the model overwrites about 59% of the fields that were already correct, and at low corruption there are far more correct fields to damage, so the recovered fraction could fall to zero or go negative and the model could make its input measurably worse. The prediction registered here is the regime hypothesis, because it is the one the render probe motivated and the one the project's next experiment depends on.

Why run it

v2 eliminated data volume and v3 eliminated model capacity, leaving the corruption schedule as the only standing candidate. Before spending training runs on it, this establishes whether the model has any regime in which it is useful at all. It needs no training: one fixed checkpoint, four evaluation levels, CPU only. If recovery improves sharply at lower corruption, a training-schedule experiment is justified. If the model damages its input instead, the binding problem is retained-token preservation and no corruption schedule will fix it.

Scope limits

One checkpoint, one seed, 32 held-out icons per level, one corruption draw per icon. The model trained only at 0.35, so evaluating it at 0.05 is off-distribution by construction; that is the point, but it means a poor result at low corruption does not by itself condemn a model trained there.

Written result

v2 eliminated data volume as the Gate G constraint and v3 eliminated model capacity, both on matched predeclared comparisons. The corruption schedule was the only candidate left, and the render probe had suggested why: at probability 0.35 the input x_t is already visually destroyed, so perhaps no model can recover information that is gone.

This run tests that before spending any training on it. One fixed checkpoint — v2's loss-selected step 840 — evaluated and rendered at four corruption levels, with 0.35 as the control because that is what it trained at. Identical icons, sizes and corruption seeds at every level; only the Bernoulli rate differs. No training.

The statistic is the per-icon recovery fraction, (MAE(x_t) - MAE(x_hat_0)) / MAE(x_t) — the share of the render gap the model closes. Absolute error is not comparable across levels, because x_t itself moves closer to x_0 as corruption falls.

Result

corruption pmedian x_t errormedian x_hat_0 errormean recovery fraction95% CIicons helped
0.050.049730.13853-4.6517-7.672 .. -1.6320 of 32
0.100.081160.13914-1.2024-2.118 .. -0.2874 of 32
0.200.138640.13717-0.1315-0.285 .. +0.02214 of 32
0.35 (control)0.176060.14045+0.1800+0.103 .. +0.25726 of 32

At 18 px the picture is the same: -4.07, -1.12, -0.12, +0.19.

Predeclared criteria

criterionthresholdobservedoutcome
recovery_improves_at_lower_corruptionrecovery at 0.10 >= 2x recovery at 0.35-1.2024 against +0.3600falsified
monotone_in_corruptionnon-increasing across 0.05, 0.10, 0.20, 0.35-4.65, -1.20, -0.13, +0.18 — strictly increasingfalsified
a_useful_regime_existssome level exceeds 0.50 recoverybest is +0.18falsified
reproducibilityidentical artifacts per levelidenticalpass

The regime hypothesis is falsified and the competing damage hypothesis — registered in the same record before launch — is confirmed. At probability 0.05 the model makes its input 2.8x worse, and it helps on zero of thirty-two icons.

Why: the output barely depends on the input

The decisive number is not in the table above. It is that x_hat_0 error hardly moves at all while x_t error moves by a factor of 3.5:

corruption pmean x_t errormean x_hat_0 error
0.050.050980.13822
0.100.087390.13820
0.200.135920.14728
0.350.179580.14651

Per icon, the prediction at p=0.05 correlates with the prediction at p=0.35 at Pearson r = 0.91, with a mean absolute difference of 0.019 against a mean error of 0.14. The model emits nearly the same reconstruction for a given icon whether 5% or 35% of its geometry has been replaced.

It is not denoising. It has learned a group- and subgroup-conditioned prior over icon geometry and emits approximately that, largely ignoring x_t.

The architecture makes this easy to believe. GeometryDenoiser takes the corrupted tokens plus group and subgroup embeddings and nothing else: there is no noise-level or timestep input and no mask marking which fields were replaced. Trained at a single fixed probability of 0.35, it never needed to behave differently at any other level, and nothing in its input tells it how much to trust what it is given.

This corrects how the earlier Gate G numbers should be read

The measurements stand; their interpretation does not.

What this redirects the project toward

Not a corruption schedule. The falsification meaning was registered before launch and it holds: the binding problem is that the model cannot leave correct fields alone, so the next experiment is an objective or architecture that lets it.

The concrete, cheap candidates, in the order their cost suggests:

  1. Tell the model the noise level. A timestep or corruption-level embedding, and training across a range of levels rather than one fixed point. Without it the model cannot modulate how much to trust its input, which is exactly the failure here.
  2. Make copying easy. Predict an edit mask, or a residual against x_t, so that "leave this field alone" is the default rather than something the model must reconstruct token by token.
  3. Condition on which fields were corrupted, as a diagnostic upper bound. That is not a realistic sampling-time signal, but it cleanly separates "cannot identify the corrupted fields" from "cannot predict their values", and that is worth knowing before designing around either.

The first is the smallest change that addresses the mechanism observed here, and it is the one this project should run next.

Scope

One checkpoint, one seed, 32 held-out icons per level, one corruption draw per icon. The model trained only at 0.35, so evaluating it at 0.05 is off-distribution by construction — that is the point of the sweep, but it means a poor result at low corruption does not by itself condemn a model that was trained across levels. Establishing that is exactly what candidate 1 above would do.

State transitions

  1. planned2026-09-20T19:00:00Z
  2. completed2026-09-20T19:15:00Zmodel_output_is_near_independent_of_its_input_it_learned_a_prior_not_a_conditional_denoiser

Run record

Verbatim from runs/openmoji-g1-corruption-sweep-71a080a-4levels-9b9b1699/run.yaml, the record committed before launch.

schema_version
1
run_id
openmoji-g1-corruption-sweep-71a080a-4levels-9b9b1699
state
completed
parent_run
openmoji-g1-train-v2-datascale-a50b2e0-c474c94d-9b9b1699
hypothesis
Two mechanisms are in tension and this run adjudicates between them. The regime hypothesis says the fixed 0.35 corruption probability destroys information no model can recover, so one fixed checkpoint evaluated at lower corruption should close a much larger fraction of the render gap between `x_t` and `x_0`. The damage hypothesis says the opposite: held-out retained-token accuracy is only 0.4128, meaning the model overwrites about 59% of the fields that were already correct, and at low corruption there are far more correct fields to damage, so the recovered fraction could fall to zero or go negative and the model could make its input measurably worse. The prediction registered here is the regime hypothesis, because it is the one the render probe motivated and the one the project's next experiment depends on.
expected_information_gain
v2 eliminated data volume and v3 eliminated model capacity, leaving the corruption schedule as the only standing candidate. Before spending training runs on it, this establishes whether the model has any regime in which it is useful at all. It needs no training: one fixed checkpoint, four evaluation levels, CPU only. If recovery improves sharply at lower corruption, a training-schedule experiment is justified. If the model damages its input instead, the binding problem is retained-token preservation and no corruption schedule will fix it.
scope_limits
One checkpoint, one seed, 32 held-out icons per level, one corruption draw per icon. The model trained only at 0.35, so evaluating it at 0.05 is off-distribution by construction; that is the point, but it means a poor result at low corruption does not by itself condemn a model trained there.
method
Evaluate and render the v2 loss-selected step-840 checkpoint at corruption probabilities 0.05, 0.10, 0.20 and 0.35, with 0.35 as the control since it is the trained level. Identical icons, sizes and corruption seeds at every level, so only the Bernoulli rate differs. The reported statistic is the per-icon recovery fraction, `(MAE(x_t) - MAE(x_hat_0)) / MAE(x_t)`, the share of the render gap the model closes. Absolute render error is not comparable across levels because `x_t` itself gets closer to `x_0` as corruption falls.
code
git_commit
71a080a
execution_mode
native-local-cpu
config
paths
configs/learning/openmoji-g1-corruption-sweep-p005.yaml, configs/learning/openmoji-g1-corruption-sweep-p010.yaml, configs/learning/openmoji-g1-corruption-sweep-p020.yaml, configs/learning/openmoji-g1-corruption-sweep-p035.yaml
inputs
checkpoint_sha256
e791493769907ac4…
checkpoint_step
840
trained_corruption_probability
0.35
hybrid_sha256
9b9b1699677a6f97…
execution
icons_per_level
32
levels
0.05, 0.1, 0.2, 0.35
render_sizes
72, 18
network
none
predeclared_criteria
note
Declared before launch. The 12-icon probe at 0.35 measured a median recovery of about 16% of the render gap; the per-icon mean recovery fraction at 0.35 is not yet known and is measured here as the control.
primary
id
recovery_improves_at_lower_corruption
statement
At 72 px, the mean per-icon recovery fraction at probability 0.10 is at least twice the mean recovery fraction at the control probability 0.35.
id
monotone_in_corruption
statement
At 72 px, the mean recovery fraction is non-increasing across the levels in the order 0.05, 0.10, 0.20, 0.35.
id
a_useful_regime_exists
statement
At some level the mean recovery fraction exceeds 0.50, meaning the model closes more than half the render gap between the corrupted input and the clean program.
id
reproducibility
statement
Every level reproduces identical artifacts on a second invocation.
falsification_meaning
A negative or near-zero recovery fraction at low corruption supports the damage hypothesis: the binding problem is that the model overwrites correct fields, which the standing 0.90 retained-preservation bar has been measuring all along at 0.41. In that case the next experiment is not a corruption schedule but an objective or architecture that can leave correct fields alone - for example predicting an edit mask, or conditioning on which fields were corrupted. That redirection would be the most valuable outcome available here and must not be tuned away.
reported_not_gated
absolute median RGBA MAE for `x_t` and `x_hat_0` at every level, the 18 px figures as a consistency check, whether any level produces visually recognizable output
outputs
local_metadata
runs/openmoji-g1-corruption-sweep-71a080a-4levels-9b9b1699
report_roots
reports/learning/openmoji-g1-corruption-sweep-p{005,010,020,035}
planned_at
2026-09-20 19:00:00+00:00
completed_at
2026-09-20 19:15:00+00:00
result
statistic
per-icon recovery fraction (MAE(x_t) - MAE(x_hat_0)) / MAE(x_t), 72 px
levels
0.05
median_x_t
0.04973
median_x_hat_0
0.13853
mean_recovery
-4.6517
ci95
-7.6718, -1.6316
helped
0
of
32
0.10
median_x_t
0.08116
median_x_hat_0
0.13914
mean_recovery
-1.2024
ci95
-2.1178, -0.2869
helped
4
of
32
0.20
median_x_t
0.13864
median_x_hat_0
0.13717
mean_recovery
-0.1315
ci95
-0.2845, 0.0216
helped
14
of
32
0.35
median_x_t
0.17606
median_x_hat_0
0.14045
mean_recovery
0.18
ci95
0.103, 0.257
helped
26
of
32
input_independence
mean_x_t_error_by_level
0.05
0.05098
0.10
0.08739
0.20
0.13592
0.35
0.17958
mean_x_hat_0_error_by_level
0.05
0.13822
0.10
0.1382
0.20
0.14728
0.35
0.14651
per_icon_pearson_r_p005_vs_p035
0.9122
mean_abs_prediction_difference_p005_vs_p035
0.01902
statement
Input render error varies by a factor of 3.5 across the sweep while prediction error moves about 6%, and per-icon predictions at 0.05 and 0.35 correlate at r = 0.91. The model emits nearly the same reconstruction regardless of how corrupted its input is.
identical_rerun
true
predeclared_outcome
overall
falsified
recovery_improves_at_lower_corruption
passed
false
observed
-1.2024
threshold
0.36
monotone_in_corruption
passed
false
observed
-4.6517, -1.2024, -0.1315, 0.18
note
strictly increasing in p
the opposite of the prediction
None
a_useful_regime_exists
passed
false
observed
0.18
threshold
0.5
reproducibility
passed
true
competing_hypothesis_confirmed
damage
conclusion
The regime hypothesis is false and the damage hypothesis registered alongside it is confirmed. At probability 0.05 the model makes its input 2.8x worse and helps on zero of thirty-two icons. The mechanism is that the prediction barely depends on the input: the model has learned a group- and subgroup-conditioned prior over icon geometry and emits approximately that. GeometryDenoiser has no noise-level or timestep input and no mask marking corrupted fields, and it trained at a single fixed probability, so nothing tells it how much to trust what it is given.
reinterprets_earlier_runs
The measurements stand, the interpretation does not. The +18% recovery at 0.35 is a fixed-quality output beating a badly corrupted baseline, not recovered geometry. Retained-token accuracy of 0.4128 was this same finding visible in token space since v1. The shared loss floor near 7.35 across a 3.53x capacity change is consistent with both models learning the same prior. v1 and v2 learned a prior over icons, not a conditional denoiser.
redirects_to
Not a corruption schedule. Give the model a noise-level or timestep embedding and train across a range of levels; make copying easy through an edit mask or a residual against x_t; and, as a diagnostic upper bound only, condition on which fields were corrupted to separate identification failure from prediction failure.
artifact_durability
versioned
Four contact sheets, summaries, and per-icon render metrics under reports/learning/, with metrics and summaries also copied into this run directory. Reproducible from the durable v2 checkpoint on CPU in about six minutes.