MojiDiff

← experiments

openmoji-g1-per-level-comparison-19dd31e-2models-9b9b1699

completed falsified

Hypothesis

Training across corruption levels was retested and restored v4's conclusion that it changes nothing - but that retest measured only at probability 0.35, the one level at which the sweep has since shown nothing can help, because a third of the geometry is simply gone. The sweep also showed that v16, trained at 0.35 alone, produces recognizable icons at 0.05 and 0.10 despite those levels being off-distribution for it. A model that samples its corruption level over 0.05 to 0.50 and is told the level should beat it at exactly those low levels, where v16 is extrapolating and it is not.

Why run it

It is the direct test of the correction the sweep implies: that this gate's entire history reported a single destructive level and that doing so hid where the model works. If the range-trained model wins at low corruption, the project has a better operating point and a reason to train that way. If it does not, then v16's extrapolation is genuinely as good as training for the regime, which would be surprising and worth knowing.

Scope limits

Fixed-topology and geometry-only, single seed. No new training: both checkpoints exist, so this is an evaluation study. 32 icons rather than the 12 used previously, because the last sweep was prevented from a significant result at 0.05 by one bad case. Recognisability is a judgement and is reported as one.

Written result

The sweep showed v16 — trained at probability 0.35 alone — producing recognizable icons at 0.05 and 0.10, levels it had never seen. The obvious inference was that a model which samples its corruption level over 0.05–0.50 and is told the level would do better there. It does not.

Both checkpoints already existed, so this is an evaluation study: 32 icons rather than the previous 12, both models at four levels on the same icons with the same corruption seeds, each gated at the threshold derived for that level on its own withheld icons, and reported per level. The statistic is the paired per-icon absolute improvement, because at low corruption the relative form divides by a denominator going to zero — which is what prevented significance last time.

Result

pmodelmean absolute improvement95% CIhelped
0.05v16+0.01797[+0.0098, +0.0261]26/32
0.05range+0.01671[+0.0114, +0.0220]29/32
0.10v16+0.02987[+0.0226, +0.0372]30/32
0.10range+0.02899[+0.0221, +0.0359]31/32
0.20v16+0.03684[+0.0244, +0.0493]28/32
0.20range+0.03360[+0.0251, +0.0421]30/32
0.35v16+0.03846[+0.0245, +0.0524]27/32
0.35range+0.02446[+0.0161, +0.0328]28/32

Paired difference, range minus v16:

pdifference95% CIrange better on
0.05−0.00125[−0.0092, +0.0067]17/32
0.10−0.00087[−0.0068, +0.0050]14/32
0.20−0.00324[−0.0138, +0.0073]17/32
0.35−0.01400[−0.0264, −0.0016]13/32
criterionobservedoutcome
range_model_wins_at_low_corruption−0.00125, CI spans zerofalsified
range_model_does_not_lose_at_the_trained_level−0.01400, CI excludes zerofalsified
both_models_help_at_every_levelevery CI excludes zeropass
reproducibilityidentical artifacts, all eight runspass

Training across levels is indistinguishable from single-level training at 0.05, 0.10 and 0.20, and measurably worse at 0.35, the level it was supposed to trade away. It buys nothing and costs something.

The calibration figures said so first

This run's record, written before it ran, noted that the range-trained model saves fewer view units per field on its calibration icons at every level — 0.4134 against 0.4636 at 0.05, and 1.9435 against 3.0133 at 0.35 — and flagged that as arguing against the hypothesis. It was right, and it cost nothing to check: the calibration estimate is a usable predictor of the render outcome, which is worth knowing for future runs.

What 32 icons settles

The previous sweep could not clear zero at p=0.05: nine of twelve icons improved but one bad case dragged the interval across zero. At 32 icons both models improve lightly corrupted inputs significantly — v16 on 26 of 32, the range model on 29 of 32, both intervals excluding zero. The earlier failure was sample size, and the corrected rule about using absolute differences at low corruption did its job.

So the sweep's headline result is now properly supported rather than suggestive: this model repairs lightly corrupted icons, significantly, at every level tested.

What is left

The predeclared falsification meaning applies. v16's extrapolation from a single level is as good as training across levels, so the training schedule is not the remaining lever. Together with the earlier retests — data volume matters, capacity does not, noise conditioning does not — every training-side factor this gate has identified is now settled.

That leaves Gate I, the cached autoregressive baseline on the same codec, which is PROJECT_PLAN.md section 12 branch 4. It is the only untried route to a generation result rather than a denoising one, and nothing in the project blocks it.

Scope

No training; both checkpoints predate this run. 32 icons, one corruption draw per icon per level, single seed, fixed topology and geometry only. This compares two denoisers; neither is a generator.

State transitions

  1. planned2026-09-21T07:20:00Z
  2. completed2026-09-21T07:45:00Ztraining_across_corruption_levels_buys_nothing_and_costs_performance_at_the_trained_level

Run record

Verbatim from runs/openmoji-g1-per-level-comparison-19dd31e-2models-9b9b1699/run.yaml, the record committed before launch.

schema_version
1
run_id
openmoji-g1-per-level-comparison-19dd31e-2models-9b9b1699
state
completed
parent_runs
openmoji-g1-geometric-gate-v16-85ec724-aaa03592-9b9b1699, openmoji-g1-retest-eliminations-b949fc2-3arms-9b9b1699
hypothesis
Training across corruption levels was retested and restored v4's conclusion that it changes nothing - but that retest measured only at probability 0.35, the one level at which the sweep has since shown nothing can help, because a third of the geometry is simply gone. The sweep also showed that v16, trained at 0.35 alone, produces recognizable icons at 0.05 and 0.10 despite those levels being off-distribution for it. A model that samples its corruption level over 0.05 to 0.50 and is told the level should beat it at exactly those low levels, where v16 is extrapolating and it is not.
expected_information_gain
It is the direct test of the correction the sweep implies: that this gate's entire history reported a single destructive level and that doing so hid where the model works. If the range-trained model wins at low corruption, the project has a better operating point and a reason to train that way. If it does not, then v16's extrapolation is genuinely as good as training for the regime, which would be surprising and worth knowing.
scope_limits
Fixed-topology and geometry-only, single seed. No new training: both checkpoints exist, so this is an evaluation study. 32 icons rather than the 12 used previously, because the last sweep was prevented from a significant result at 0.05 by one bad case. Recognisability is a judgement and is reported as one.
method
Both models rendered at 0.05, 0.10, 0.20 and 0.35 on the same 32 held-out icons with the same corruption seeds, each gated at the threshold derived for that level on its own withheld calibration icons. Reported per level, never collapsed.
models
v16
trained_at
fixed 0.35
thresholds
0.05
0.87
0.10
0.815
0.20
0.76
0.35
0.68
calibration_view_units_saved
0.05
0.4636
0.10
0.9464
0.20
1.9726
0.35
3.0133
range
trained_at
sampled 0.05-0.50 with 16 noise-level features
thresholds
0.05
0.535
0.10
0.58
0.20
0.615
0.35
0.76
calibration_view_units_saved
0.05
0.4134
0.10
0.8185
0.20
1.537
0.35
1.9435
note
It gates far lower at low corruption and so edits more there, but saves fewer view units per field on calibration at every level - which argues against the hypothesis before the run and is recorded here rather than after.
code
git_commit
19dd31e
execution_mode
native-local-cpu
trains_nothing
true
baselines
v16_at_12_icons
0.05
0.0308
0.10
0.3902
0.20
0.2469
0.35
0.2511
identity
0.0
predeclared_criteria
statistic
Paired per-icon difference in absolute RGBA MAE, x_t minus x_hat_0, at 72 px. The absolute form rather than the relative one, because at low corruption the relative statistic divides by a denominator that goes to zero - which is what prevented a significant result at 0.05 last time.
primary
id
range_model_wins_at_low_corruption
statement
At probability 0.05, the range-trained model's mean paired absolute improvement exceeds v16's, and the 95% interval on the per-icon difference BETWEEN the two models excludes zero. This is the hypothesis.
id
range_model_does_not_lose_at_the_trained_level
statement
At probability 0.35, the range-trained model's mean paired absolute improvement is not below v16's by more than the interval on their difference. Buying low-level performance with a loss at 0.35 would be a trade, not a win, and must be visible.
id
both_models_help_at_every_level
statement
For each model and level, mean paired absolute improvement is positive with its 95% interval excluding zero. At 32 icons a single bad case should no longer prevent this.
id
reproducibility
statement
Every one of the eight render runs reproduces identical artifacts.
reported_not_gated
per-level medians, icons helped, and the relative statistic alongside the absolute, whether renders are recognizable at each level, for each model, stated plainly
falsification_meaning
If the range-trained model does not win at low corruption, then v16's extrapolation from a single level is as good as training across levels, the calibration figures above predicted the outcome correctly, and the remaining lever is not the training schedule. That would leave Gate I - the cached autoregressive baseline on the same codec - as the one untried route to a generation result rather than a denoising one.
outputs
local_metadata
runs/openmoji-g1-per-level-comparison-19dd31e-2models-9b9b1699
report_roots
reports/learning/openmoji-g1-levels-{v16,range}-p{005,010,020,035}
planned_at
2026-09-21 07:20:00+00:00
completed_at
2026-09-21 07:45:00+00:00
result
identical_rerun
true
icons
32
statistic
paired per-icon absolute improvement in RGBA MAE at 72 px
mean_absolute_improvement
v16
0.05
0.01797
0.10
0.02987
0.20
0.03684
0.35
0.03846
range
0.05
0.01671
0.10
0.02899
0.20
0.0336
0.35
0.02446
helped_of_32
v16
0.05
26
0.10
30
0.20
28
0.35
27
range
0.05
29
0.10
31
0.20
30
0.35
28
paired_difference_range_minus_v16
0.05
mean
-0.00125
ci95
-0.00921, 0.0067
0.10
mean
-0.00087
ci95
-0.00675, 0.005
0.20
mean
-0.00324
ci95
-0.01378, 0.0073
0.35
mean
-0.014
ci95
-0.02644, -0.00156
predeclared_outcome
overall
falsified
range_model_wins_at_low_corruption
passed
false
observed
-0.00125
ci95
-0.00921, 0.0067
range_model_does_not_lose_at_the_trained_level
passed
false
observed
-0.014
ci95
-0.02644, -0.00156
both_models_help_at_every_level
passed
true
reproducibility
passed
true
conclusion
Training across corruption levels is indistinguishable from single-level training at 0.05, 0.10 and 0.20, and measurably worse at 0.35 - the level it was supposed to trade away. It buys nothing and costs something. v16's extrapolation from a single level is as good as training for the regime.
prediction_recorded_before_the_run
This run's record noted beforehand that the range-trained model saves fewer view units per field on calibration at every level, 0.4134 against 0.4636 at 0.05 and 1.9435 against 3.0133 at 0.35, and flagged that as arguing against the hypothesis. It was right, so the calibration estimate is a usable cheap predictor of the render outcome.
resolves_the_previous_sample_size_limit
The 12-icon sweep could not clear zero at p=0.05 because one bad case dragged the interval across it. At 32 icons both models improve lightly corrupted inputs significantly - v16 on 26 of 32, the range model on 29 of 32, both intervals excluding zero. The sweep's headline is now properly supported rather than suggestive.
what_is_left
Every training-side factor this gate identified is now settled: data volume matters, capacity does not, noise conditioning does not, and the corruption schedule does not. That leaves Gate I, the cached autoregressive baseline on the same codec - PROJECT_PLAN.md section 12 branch 4 - as the only untried route to a generation result rather than a denoising one.