MojiDiff

← experiments

openmoji-g1-selection-scalar-full-c4e6a8c-powered-9b9b1699

completed passed

Hypothesis

No directional hypothesis. The twelve-icon comparison falsified the only prediction there were grounds for - that the strictly more accurate checkpoint renders better - and left the two scalars statistically indistinguishable. This run re-measures the same paired quantity on the complete 128-icon held-out draw, which shrinks the confidence interval by about 3.3x, to find out whether the scalars can be separated at all or whether the effect is simply small.

Why run it

Checkpoint selection is a knob that every later Gate G run turns. Leaving the rule undecided leaves post-hoc freedom in every one of them. This settles it or bounds it at a cost of CPU render time and no GPU.

Scope limits

One training trajectory, one seed, one corruption probability of 0.35, one corruption draw per icon. Conditional on argmax decoding; a sampler that draws from the distribution would reopen the calibration question. Both checkpoints render as scribble, so this bounds a selection rule, not output quality.

Written result

The twelve-icon comparison falsified the only directional prediction there were grounds for — that the strictly more accurate checkpoint renders better — but left the two candidate scalars statistically indistinguishable. This run re-measures the same paired quantity on the complete 128-icon held-out draw, with the decision rule declared before the numbers were read.

Result

Control: x_t render error is identical across the two arms for all 128 icons at both sizes, so the two checkpoints saw exactly the same inputs.

72 px18 px
median RGBA MAE, loss-selected step 8400.1412150.144372
median RGBA MAE, accuracy-selected step 1,3200.1345770.138735
mean paired difference (positive favours loss)-0.001638-0.001392
95% CI on that mean-0.004867 .. +0.001590-0.004982 .. +0.002198
CI half-width0.0032280.003590
icons where loss-selected is better65 of 12861 of 128
two-sided sign testp = 0.930p = 0.659

Predeclared criteria

criterionthresholdobservedoutcome
control_inputs_identicalidentical for all 128identicalpass
separation_or_boundCI excludes zero, or half-width < 0.01half-width 0.003228pass, by the bound
reproducibilityidentical artifacts on rerunidenticalpass

The confidence interval does not exclude zero and the sign test is a coin flip at both sizes, so the two scalars are not separable. But the interval is now tight: any true difference is below about 0.0032 RGBA MAE at 72 px, against a median render error of about 0.14. The choice of selection scalar moves render quality by at most 2.3% of the error that is already there.

The declared decision rule, applied

The rule committed before the numbers were read says that if the effect is bounded below 0.01, held-out loss governs selection by default — on the separate ground that it is the standard early-stopping signal and the rule v1 and v2 already used — and that the record must state explicitly that the choice is immaterial for render quality at this stage.

Held-out loss is the project's checkpoint-selection scalar. It is chosen by convention, not by evidence of superiority, and the record says so. v2's changed-token criterion, evaluated at the loss-selected checkpoint, remains correctly recorded as falsified: reading it at the final step would not have been justified by render quality, because render quality does not distinguish the two checkpoints at all.

The twelve-icon medians were an artifact, confirmed

At twelve icons the medians favoured the loss arm by 6.5% (0.142433 against 0.151687). At 128 icons they favour the accuracy arm by 4.7% (0.141215 against 0.134577) — the opposite direction — while the paired difference stays near zero in both samples. That is exactly the instability the twelve-icon run's design defect predicted, and it is a useful reminder: on paired data, compare the pairs.

Scope

One training trajectory, one seed, one corruption probability of 0.35, one corruption draw per icon. The conclusion is conditional on argmax decoding, which discards the distribution beyond the winning token; a sampler that draws from the distribution would reopen the calibration question and this result would not transfer. Both checkpoints render as scribble, so this bounds the effect of a selection rule, not output quality.

State transitions

  1. running2026-09-20T18:02:00Z
  2. completed2026-09-20T18:20:00Zselection_scalars_are_not_separable_effect_bounded_below_0.0032_rgba_mae_held_out_loss_adopted_by_convention

Run record

Verbatim from runs/openmoji-g1-selection-scalar-full-c4e6a8c-powered-9b9b1699/run.yaml, the record committed before launch.

schema_version
1
run_id
openmoji-g1-selection-scalar-full-c4e6a8c-powered-9b9b1699
state
completed
parent_run
openmoji-g1-selection-scalar-c4e6a8c-3a34e64b-9b9b1699
hypothesis
No directional hypothesis. The twelve-icon comparison falsified the only prediction there were grounds for - that the strictly more accurate checkpoint renders better - and left the two scalars statistically indistinguishable. This run re-measures the same paired quantity on the complete 128-icon held-out draw, which shrinks the confidence interval by about 3.3x, to find out whether the scalars can be separated at all or whether the effect is simply small.
expected_information_gain
Checkpoint selection is a knob that every later Gate G run turns. Leaving the rule undecided leaves post-hoc freedom in every one of them. This settles it or bounds it at a cost of CPU render time and no GPU.
scope_limits
One training trajectory, one seed, one corruption probability of 0.35, one corruption draw per icon. Conditional on argmax decoding; a sampler that draws from the distribution would reopen the calibration question. Both checkpoints render as scribble, so this bounds a selection rule, not output quality.
disclosure
The two render jobs were launched before these criteria were written. The results had not been produced or inspected when the criteria below were committed, which was verified by checking that neither report directory existed. The probe is a deterministic re-measurement with no tunable parameter beyond the icon count, but the ordering is disclosed rather than presented as a clean predeclaration.
code
git_commit
c4e6a8ccc01b1807f77706b4ff977fd1cbe16418
execution_mode
native-local-cpu
config
loss_arm
configs/learning/openmoji-g1-v2-renders-full.yaml
accuracy_arm
configs/learning/openmoji-g1-v2-final-step-renders-full.yaml
arms
loss_selected
checkpoint_step
840
checkpoint_sha256
e791493769907ac4…
accuracy_selected
checkpoint_step
1320
checkpoint_sha256
359b11a1020ea07b…
execution
icons
128
render_sizes
72, 18
statistic
Mean per-icon difference in RGBA MAE of `x_hat_0` against `x_0`, accuracy arm minus loss arm, so a positive value favours the loss-selected checkpoint. This replaces the twelve-icon run's independent medians, which were the wrong test for paired data.
network
none
predeclared_criteria
primary
id
control_inputs_identical
statement
`x_t` render error is identical across the two arms for all 128 icons at both sizes. A failure voids the comparison.
id
separation_or_bound
statement
At 72 px, either the 95% confidence interval on the mean paired difference excludes zero, or its half-width is below 0.01 RGBA MAE. If neither holds the question is not settled at this sample size either.
id
reproducibility
statement
Both probes reproduce identical artifacts on a second invocation.
decision_rule
note
Declared before the numbers are read, so the conclusion cannot be chosen after seeing them.
if_interval_excludes_zero
The favoured scalar governs checkpoint selection from now on, and is recorded in state/CURRENT.md as the project's selection rule.
if_effect_bounded_below_0.01
Held-out loss governs selection by default, on the separate ground that it is the standard early-stopping signal and the rule v1 and v2 already used, and the record states explicitly that the choice is immaterial for render quality at this stage.
if_neither
The question stays open, and no later run may cite whichever reading favours its result. Any run reporting a recovery number must report both readings.
reported_not_gated
the 18 px interval, as a consistency check on the 72 px decision, per-icon differences and their spread, absolute render error for both arms
outputs
local_metadata
runs/openmoji-g1-selection-scalar-full-c4e6a8c-powered-9b9b1699
report_roots
reports/learning/openmoji-g1-v2-renders-full, reports/learning/openmoji-g1-v2-final-step-renders-full
planned_at
2026-09-20 18:02:00+00:00
completed_at
2026-09-20 18:20:00+00:00
result
icons
128
control_inputs_identical
true
px72
median_rgba_mae_loss_selected
0.141215
median_rgba_mae_accuracy_selected
0.134577
mean_paired_difference
-0.001638
ci95
-0.004867, 0.00159
ci95_half_width
0.003228
excludes_zero
false
loss_selected_better_on
65
of
128
sign_test_p
0.93
px18
median_rgba_mae_loss_selected
0.144372
median_rgba_mae_accuracy_selected
0.138735
mean_paired_difference
-0.001392
ci95
-0.004982, 0.002198
ci95_half_width
0.00359
excludes_zero
false
loss_selected_better_on
61
of
128
sign_test_p
0.659
identical_rerun
true
predeclared_outcome
overall
passed
control_inputs_identical
passed
true
separation_or_bound
passed
true
branch
bound
observed_ci_half_width_72px
0.003228
threshold
0.01
note
The interval does not exclude zero and the sign test is a coin flip at both sizes, so the scalars are not separable. The bound branch of the criterion is met: any true difference is below about 0.0032 RGBA MAE against a median render error near 0.14, at most 2.3% of the error already present.
reproducibility
passed
true
decision_applied
rule
if_effect_bounded_below_0.01
outcome
Held-out loss is the project's checkpoint-selection scalar from now on, chosen as the standard early-stopping signal and the rule v1 and v2 already used, not on evidence of superiority. The record states that the choice is immaterial for render quality at this stage.
corroborates_predecessor_design_defect
The twelve-icon medians favoured the loss arm by 6.5%; the 128-icon medians favour the accuracy arm by 4.7%, the opposite direction, while the paired difference stays near zero in both samples. The median was an unstable statistic on paired data, exactly as the predecessor's defect note recorded.
artifact_durability
durable_root
/home/dev/.cache/openmoji-g1-selection-scalar-full-c4e6a8c-powered-9b9b1699
versioned
Both 128-icon contact sheets, summaries, and per-icon render metrics under reports/learning/, about 8.3 MB, duplicated on the durable sink.