MojiDiff

← experiments

openmoji-g1-v16-corruption-sweep-a0cd620-4levels-9b9b1699

completed partially-falsified

Hypothesis

Every result in this gate is at corruption probability 0.35, where the input is already scribble - so "no render is a recognizable icon" may be a statement about the regime rather than the model. The same sweep over the v2 checkpoint found it made a lightly corrupted input 2.8x WORSE and helped zero of 32 icons. v16 should at minimum never damage its input at any level, and at the lightest corruption it should visibly improve it. Whether it ever produces something recognizable is the question this run exists to answer honestly, either way.

Why run it

It is the cheapest test bearing directly on the project's primary objective, needs no training, and its answer decides whether the next effort goes into the corruption schedule or into the model. A negative answer at every level is as informative as a positive one and must be reported as plainly.

Scope limits

v16 trained only at 0.35, so 0.05 through 0.20 are off-distribution by construction - that is deliberate and is the point, but it means a poor result at low corruption does not condemn a model trained across levels. Twelve icons, one corruption draw each, single seed. Recognisability is a judgement and is reported as one, not gated.

Written result

Every result in this gate has been at corruption probability 0.35, where the input is already scribble, and every report has carried the same line: no render in this project is a recognizable icon. This run asks whether that was a statement about the model or about the regime.

It was about the regime.

Result

v16 rendered at four corruption levels, each gated at the threshold derived for that level on the withheld calibration icons:

pthresholdmedian x_tmedian x_hat_0mean recovery95% CIhelped
0.050.8700.037300.02235+0.0308[−0.629, +0.690]9/12
0.100.8150.082470.04206+0.3902[+0.227, +0.554]11/12
0.200.7600.125680.10014+0.2469[+0.089, +0.405]10/12
0.350.6800.169550.13373+0.2511[+0.141, +0.361]11/12

The same sweep over the v2 checkpoint gave −4.6517, −1.2024, −0.1315, +0.1800 and helped zero of 32 icons at p=0.05.

criterionobservedoutcome
never_damagesCI excludes zero at 0.10, 0.20, 0.35 but not 0.05falsified
improves_a_lightly_corrupted_inputmedian 0.02235 against 0.03730, a 40% improvementpass
reproducibilityidentical artifacts at every levelpass

The p=0.05 failure is one icon and a fragile statistic

Nine of twelve icons improve, two are untouched, and exactly one worsens: 26A0, the warning triangle, by 0.05018 absolute. Because its x_t error was only 0.01390, the relative statistic turns that into −3.61 and drags a mean of twelve to +0.0308 from a median of +0.3496.

Switching to the absolute difference does not rescue it either — mean +0.01070 with a 95% interval of [−0.00204, +0.02344]. So this is not only a bad choice of statistic; at twelve icons one bad case is genuinely enough to prevent significance. The criterion is falsified as written and the honest summary is that the model improves nine of twelve lightly corrupted icons by a median 35% and damages one.

What the renders show

At p=0.05 the predictions are recognizable icons, and visibly repaired: the hedgehog's stray diagonal slash removed, the first-aid kit's diagonal gone, "WC" cleaned to near-perfect, the vampire's face restored from under a scribble. 26A0 is the visible failure — a near-perfect input triangle comes out broken.

At p=0.10 the repairs are larger and still land. 1F3CB, the weightlifter, goes from heavy scribble to a clean, recognizable figure; 1F199 comes back as a clean "UP!" badge; the vampire and the WC sign are both recognizable again. This is the level with the best recovery of the four, +0.3902.

At 0.20 and 0.35 neither input nor output is recognizable, which is what every earlier report described — correctly, for those levels.

What this changes

The standing line that "no render in this project is a recognizable icon" was true and is now obsolete. It described corruption probability 0.35, which has been the only level this gate ever trained or reported at, and at which a third of the geometry is simply gone. At 0.05 and 0.10 the same checkpoint produces recognizable output and performs real repairs.

The honest framing is narrow and should stay narrow. This is a denoiser that removes a handful of stray strokes from an icon that is mostly intact — not a generator, and not a model that reconstructs destroyed geometry. But it is the first result in this project where the output is something a person would recognise, and it means the ceiling was set by the corruption schedule rather than by the representation, the model or the objective.

Scope

v16 trained only at 0.35, so 0.05 through 0.20 are off-distribution by construction — which makes the result stronger, not weaker, since the model was never shown these levels. Twelve icons, one corruption draw each, single seed, no training. Recognisability is a judgement and is reported as one.

The obvious next run follows: train across a range of corruption levels and report at each, rather than training and reporting at a single destructive point.

State transitions

  1. planned2026-09-21T06:30:00Z
  2. completed2026-09-21T07:00:00Zproduces_recognizable_icons_at_low_corruption_the_regime_was_the_ceiling_not_the_model

Run record

Verbatim from runs/openmoji-g1-v16-corruption-sweep-a0cd620-4levels-9b9b1699/run.yaml, the record committed before launch.

schema_version
1
run_id
openmoji-g1-v16-corruption-sweep-a0cd620-4levels-9b9b1699
state
completed
parent_run
openmoji-g1-geometric-gate-v16-85ec724-aaa03592-9b9b1699
supersedes_for_v16
openmoji-g1-corruption-sweep-71a080a-4levels-9b9b1699
hypothesis
Every result in this gate is at corruption probability 0.35, where the input is already scribble - so "no render is a recognizable icon" may be a statement about the regime rather than the model. The same sweep over the v2 checkpoint found it made a lightly corrupted input 2.8x WORSE and helped zero of 32 icons. v16 should at minimum never damage its input at any level, and at the lightest corruption it should visibly improve it. Whether it ever produces something recognizable is the question this run exists to answer honestly, either way.
expected_information_gain
It is the cheapest test bearing directly on the project's primary objective, needs no training, and its answer decides whether the next effort goes into the corruption schedule or into the model. A negative answer at every level is as informative as a positive one and must be reported as plainly.
scope_limits
v16 trained only at 0.35, so 0.05 through 0.20 are off-distribution by construction - that is deliberate and is the point, but it means a poor result at low corruption does not condemn a model trained across levels. Twelve icons, one corruption draw each, single seed. Recognisability is a judgement and is reported as one, not gated.
method
Render x_0, x_t and x_hat_0 at probabilities 0.05, 0.10, 0.20 and 0.35, each gated at the threshold derived for that level on the withheld calibration icons - 0.8700, 0.8150, 0.7600 and 0.6800 respectively. Identical icons and corruption seeds across levels.
code
git_commit
a0cd6202d4e3b67af188f17d8fd94722951bc384
execution_mode
native-local-cpu
trains_nothing
true
baselines
identity
recovery_at_every_level
0.0
v2_checkpoint_same_sweep
0.05
-4.6517
0.10
-1.2024
0.20
-0.1315
0.35
0.18
note
uniform corruption
helped 0 of 32 at 0.05
None
v16_at_0.35
paired_render_recovery_72px
0.2511
predeclared_criteria
primary
id
never_damages
statement
At every level, mean paired per-icon render recovery at 72 px is positive with its 95% interval excluding zero. The v2 checkpoint failed this badly at three of four levels; a repaired model should not.
id
improves_a_lightly_corrupted_input
statement
At probability 0.05, median x_hat_0 RGBA MAE at 72 px is below the median x_t error at that level. Improving a nearly clean input is strictly harder than improving a destroyed one and is where v2 failed worst.
id
reproducibility
statement
Every level reproduces identical artifacts on a second invocation.
reported_not_gated
whether any render is recognizable, stated plainly by inspection either way, median x_t and x_hat_0 error at each level, and icons helped
falsification_meaning
If the model still damages lightly corrupted inputs, the repairs did not generalise off the training corruption level and the corruption schedule becomes the next training factor. If it never damages but nothing is recognizable at any level, then the limit is the model rather than the regime, and PROJECT_PLAN.md section 12 branch 4 - Gate I, the cached autoregressive baseline on the same codec - is the remaining route to a positive generation result.
outputs
local_metadata
runs/openmoji-g1-v16-corruption-sweep-a0cd620-4levels-9b9b1699
report_roots
reports/learning/openmoji-g1-v16-sweep-p{005,010,020,035}
planned_at
2026-09-21 06:30:00+00:00
completed_at
2026-09-21 07:00:00+00:00
result
identical_rerun
true
levels
0.05
threshold
0.87
median_x_t
0.0373
median_x_hat_0
0.02235
mean_recovery
0.0308
median_recovery
0.3496
ci95
-0.6289, 0.6904
helped
9
of
12
0.10
threshold
0.815
median_x_t
0.08247
median_x_hat_0
0.04206
mean_recovery
0.3902
ci95
0.2269, 0.5536
helped
11
of
12
0.20
threshold
0.76
median_x_t
0.12568
median_x_hat_0
0.10014
mean_recovery
0.2469
ci95
0.0888, 0.4049
helped
10
of
12
0.35
threshold
0.68
median_x_t
0.16955
median_x_hat_0
0.13373
mean_recovery
0.2511
ci95
0.1408, 0.3614
helped
11
of
12
v2_same_sweep
0.05
-4.6517
0.10
-1.2024
0.20
-0.1315
0.35
0.18
helped_at_0.05
0 of 32
predeclared_outcome
overall
partially-falsified
never_damages
passed
false
detail
The interval excludes zero at 0.10, 0.20 and 0.35 but not at 0.05, where nine of twelve icons improve, two are untouched and exactly one worsens - 26A0, by 0.05018 absolute. Its x_t error was only 0.01390, so the relative statistic turns that into -3.61 and drags a mean of twelve to +0.0308 from a median of +0.3496. The absolute difference does not rescue it either, mean +0.01070 with interval [-0.00204, +0.02344], so at twelve icons one bad case is genuinely enough to prevent significance. Falsified as written.
improves_a_lightly_corrupted_input
passed
true
median_x_hat_0
0.02235
median_x_t
0.0373
reproducibility
passed
true
headline
At probability 0.05 and 0.10 the predictions are RECOGNIZABLE ICONS and visibly repaired - the hedgehog's stray slash removed, the first-aid kit's diagonal gone, "WC" cleaned, the vampire's face restored, and at 0.10 the weightlifter 1F3CB going from heavy scribble to a clean figure. This is the first result in this project where the output is something a person would recognise.
what_it_changes
The standing line that no render in this project is a recognizable icon was true and is now obsolete. It described probability 0.35, the only level this gate ever trained or reported at, where a third of the geometry is simply gone. The ceiling was set by the corruption schedule, not by the representation, the model or the objective.
honest_framing
This is a denoiser that removes a handful of stray strokes from an icon that is mostly intact - not a generator, and not a model that reconstructs destroyed geometry. 26A0 shows it can still damage a near-perfect input. v16 trained only at 0.35, so these levels are off-distribution, which makes the result stronger rather than weaker.
next
Train across a range of corruption levels and report at each, rather than training and reporting at a single destructive point.