MojiDiff

← experiments

openmoji-g1-retest-eliminations-b949fc2-3arms-9b9b1699

completed passed

Hypothesis

Four eliminations in this gate - data volume, model capacity, the corruption regime and noise-level conditioning - were each declared a plateau on a held-out scalar that was roughly 31% four individual fields, on runs stopped at 6 to 24% of their step cap, under a corruption process whose replacements a zero-parameter marginal detector identified at 2.84 lift, with coordinates encoded as unordered categories, and graded on exact-token accuracy which v15 and v16 showed points the wrong way. All of that is fixed. None of the four conclusions has been retested. At least one of them is expected to change, and capacity is the most likely: v3 concluded that 3.53x parameters does nothing, on a setup where 768 parameters of slot binding later beat it outright.

Why run it

Four conclusions currently stand in the record as inadmissible-as-read. This either restores them under sound measurement or overturns them, and it is the cheapest remaining evidence in the project - the v16 recipe trains in minutes.

Scope limits

Three arms, not four. The corruption-regime sweep is an evaluation-time study over a fixed checkpoint rather than a training factor, so it is a separate piece of work and is not bundled here. Fixed-topology and geometry-only, single seed, 12 icons for renders.

Written result

Four conclusions in this gate — that data volume matters, that model capacity does not, that the corruption regime is not the constraint, and that noise-level conditioning changes nothing — were each declared on a held-out scalar that was roughly 31% four individual fields, on runs stopped at 6–24% of their step cap, under a corruption process whose replacements a zero-parameter marginal detector identified at 2.84 lift, with coordinates encoded as unordered categories, and graded on exact-token accuracy which v15 and v16 showed points the wrong way.

All of that is now fixed, and none of the four had been retested. Three of them are training factors and are retested here off the v16 base, one factor per arm, judged on paired per-icon render recovery — the primary metric v15 and v16 established.

Result

armparameterstrains on72 px recovery95% CIhelped
v16 base525,1522,425+0.2511[+0.141, +0.361]11/12
capacity, 3.69x1,936,1602,425+0.2779[+0.161, +0.395]10/12
data volume, 256 icons525,152256+0.0766[+0.001, +0.152]8/12
noise conditioning528,3202,425+0.2096[+0.116, +0.304]10/12

Applying the criteria exactly as predeclared, against the base interval's half-width of 0.1103:

criterionobservedverdict
capacity_retested\diff\0.0268restores v3 — 3.69x parameters changes nothing measurable
data_volume_retested0.1745 below baserestores v2 — data volume matters
noise_conditioning_retested\diff\0.0415restores v4 — conditioning changes nothing that matters
structural_safetyall exact, all round-trippass
reproducibilityidentical rerunspass

All three original conclusions survive. The eliminations were right even though they were measured badly, and the four are admissible again.

The data-volume arm is the sharpest of the three: training on 256 icons instead of 2,425 cuts render recovery from +0.2511 to +0.0766, a 70% drop, with its interval barely clearing zero. Under the broken measurement v2 found the same thing in exact-token terms; under the corrected one it is larger and clearer.

Capacity is the one I expected to move and it did not. v3 concluded that 3.53x parameters changes nothing, and I had recorded that as the least trustworthy of the four because 768 parameters of slot binding later beat it outright. At 3.69x on a sound setup it is still within noise of the base — and the model that does best here remains the smallest one built this session. The slot-binding result and the capacity result were never in tension: one was a representational fix, the other was raw size.

A failed arm, kept

The first noise-conditioning arm set corruption_probability to 0.05 as the low end of its training range, which also moved the evaluation level, because evaluation_corruption_probability was unset and defaults to it. Its identity baseline came out at 0.9501 against the other arms' 0.6515 — it measured a different task entirely. It is recorded in the registry as failed with that reason rather than quietly rerun, and the corrected arm pins the evaluation level to 0.35.

A second recording note: on first reading these numbers I labelled the data-volume arm as overturning v2, by applying a symmetric "differs by more than the half-width" rule to a criterion that was predeclared directionally — below the base by more than the half-width restores v2. The criterion as written is the one applied above.

Scope

Three arms, not four: the corruption-regime sweep is an evaluation-time study over a fixed checkpoint rather than a training factor, so it belongs in separate work. Fixed topology and geometry only, single seed, 12 icons for renders. No render in this project is a recognizable icon.

State transitions

  1. planned2026-09-21T05:20:00Z
  2. failed2026-09-21T05:45:00Zarm_evaluated_at_0.05_not_0.35_because_evaluation_inherits_the_training_probability
  3. planned2026-09-21T05:50:00Z
  4. completed2026-09-21T06:10:00Zall_four_eliminations_survive_the_repair_and_are_admissible_again

Run record

Verbatim from runs/openmoji-g1-retest-eliminations-b949fc2-3arms-9b9b1699/run.yaml, the record committed before launch.

schema_version
1
run_id
openmoji-g1-retest-eliminations-b949fc2-3arms-9b9b1699
state
completed
parent_run
openmoji-g1-geometric-gate-v16-85ec724-aaa03592-9b9b1699
retests
openmoji-g1-train-v2-datascale-a50b2e0-c474c94d-9b9b1699, openmoji-g1-capacity-v3-3c252d5-ee32665b-9b9b1699, openmoji-g1-noise-conditioned-v4-6b935b2-5a9eaf44-9b9b1699
hypothesis
Four eliminations in this gate - data volume, model capacity, the corruption regime and noise-level conditioning - were each declared a plateau on a held-out scalar that was roughly 31% four individual fields, on runs stopped at 6 to 24% of their step cap, under a corruption process whose replacements a zero-parameter marginal detector identified at 2.84 lift, with coordinates encoded as unordered categories, and graded on exact-token accuracy which v15 and v16 showed points the wrong way. All of that is fixed. None of the four conclusions has been retested. At least one of them is expected to change, and capacity is the most likely: v3 concluded that 3.53x parameters does nothing, on a setup where 768 parameters of slot binding later beat it outright.
expected_information_gain
Four conclusions currently stand in the record as inadmissible-as-read. This either restores them under sound measurement or overturns them, and it is the cheapest remaining evidence in the project - the v16 recipe trains in minutes.
scope_limits
Three arms, not four. The corruption-regime sweep is an evaluation-time study over a fixed checkpoint rather than a training factor, so it is a separate piece of work and is not bundled here. Fixed-topology and geometry-only, single seed, 12 icons for renders.
arms
capacity
config
configs/learning/openmoji-g1-retest-capacity.yaml
sha256
28e3f1a8a373edba…
factor
d_model 96->192, heads 4->8, layers 2->4, feedforward 192->384
parameters
525152 -> 1936160, 3.69x
retests
v3, which found 3.53x parameters changed nothing
data_volume
config
configs/learning/openmoji-g1-retest-data-volume.yaml
sha256
92c53c85b1a916bb…
factor
train_samples 2681->512, which after the 256 withheld leaves 256 icons trained on
retests
v2, which found the full split beat a 256-icon subsample
noise_conditioning
config
configs/learning/openmoji-g1-retest-noise-conditioning.yaml
sha256
ae12f918cacb1d84…
factor
corruption sampled per example over 0.05-0.50 plus 16 noise-level features
bundling
Two changes, bundled exactly as v4 bundled them and disclosed the same way. Sampling without telling the model leaves it unable to use the level; telling without varying leaves it a constant.
retests
v4, which found noise-level conditioning changed nothing that mattered
code
git_commit
b949fc238b4c146e2757fc4aae87decd6f5f7940
execution_mode
native-local
dataset
hybrid_sha256
9b9b1699677a6f97…
bucket
bucket-p32-t128
base_v16
paired_render_recovery_72px
0.2511
paired_render_recovery_72px_ci
0.1408, 0.3614
paired_render_recovery_18px
0.2745
decision_threshold
0.68
held_out_aggregate
0.6354
identity
0.6515
parameters
525152
trains_on
2425
predeclared_criteria
note
Judged on paired per-icon render recovery, the primary metric established by v15 and v16. Exact-token accuracy is reported and NOT gated, because v16 is simultaneously worse than identity on it and the best geometric denoiser this project has produced.
evaluated_at
each arm's selected checkpoint, decoded at its own geometric threshold
primary
id
capacity_retested
statement
The capacity arm's mean paired render recovery at 72 px differs from v16's +0.2511 by more than the base run's interval half-width of 0.1103, in either direction. A difference smaller than that restores v3's conclusion under sound measurement; a larger one overturns it.
id
data_volume_retested
statement
The 256-icon arm's recovery is BELOW v16's +0.2511 by more than 0.1103, which would restore v2's conclusion that data volume matters. If it is not, then v2's finding does not survive the corrected setup.
id
noise_conditioning_retested
statement
The noise-conditioning arm's recovery differs from v16's by more than 0.1103 in either direction, which would overturn v4's conclusion that it changes nothing.
id
structural_safety
statement
Locked-path exactness holds and every checkpoint round-trips.
id
reproducibility
statement
Every arm returns an identical JSON result on a second invocation.
reported_not_gated
each arm's held-out aggregate beside identity, and its exact-token changed/retained, each arm's derived threshold, selected step, and mean absolute view-unit error
falsification_meaning
If all three arms land within the interval of the base, then none of the original eliminations was an artifact of the broken measurement and the four conclusions stand as originally recorded, which would be a clean restoration rather than a null result. If capacity in particular moves, the project has been under-sizing its model since v3 on the strength of a measurement that was 31% four fields.
outputs
local_metadata
runs/openmoji-g1-retest-eliminations-b949fc2-3arms-9b9b1699
report_roots
reports/learning/openmoji-g1-retest-{capacity,data-volume,noise-conditioning}
planned_at
2026-09-21 05:20:00+00:00
completed_at
2026-09-21 06:10:00+00:00
result
base_v16_render_recovery_72px
0.2511
interval_half_width
0.1103
arms
capacity
parameters
1936160
trains_on
2425
render_recovery_72px
0.2779
ci95
0.1606, 0.3951
helped
10
of
12
abs_diff_vs_base
0.0268
data_volume
parameters
525152
trains_on
256
render_recovery_72px
0.0766
ci95
0.0012, 0.152
helped
8
of
12
below_base_by
0.1745
noise_conditioning
parameters
528320
trains_on
2425
render_recovery_72px
0.2096
ci95
0.1158, 0.3035
helped
10
of
12
abs_diff_vs_base
0.0415
identical_rerun
true
predeclared_outcome
overall
passed
capacity_retested
verdict
restores_v3
abs_diff
0.0268
threshold
0.1103
meaning
3.69x parameters changes nothing measurable
data_volume_retested
verdict
restores_v2
below_base_by
0.1745
threshold
0.1103
meaning
data volume matters
256 icons gives 0.0766 against 2425 giving 0.2511
None
noise_conditioning_retested
verdict
restores_v4
abs_diff
0.0415
threshold
0.1103
meaning
conditioning changes nothing that matters
structural_safety
passed
true
reproducibility
passed
true
conclusion
All three original conclusions survive the corrected setup. The eliminations were right even though they were measured badly, and the four are admissible again. The data-volume arm is the sharpest: 256 icons against 2,425 cuts render recovery 70%, from 0.2511 to 0.0766, with its interval barely clearing zero.
expectation_not_met
Capacity is the one I expected to move and it did not. v3's conclusion had been recorded as the least trustworthy of the four because 768 parameters of slot binding later beat a 3.53x capacity increase outright. At 3.69x on a sound setup it is still within noise of the base, and the best model this session built remains the smallest. The two results were never in tension: one was a representational fix, the other raw size.
recording_notes
failed_arm
The first noise-conditioning arm evaluated at 0.05 rather than 0.35, because evaluation_corruption_probability was unset and defaults to corruption_probability, which the training range had moved. Its identity baseline came out at 0.9501 against the other arms' 0.6515. It is retained under its own identity as failed.
directional_criterion
On first reading I labelled the data-volume arm as overturning v2, by applying a symmetric "differs by more than the half-width" rule to a criterion predeclared DIRECTIONALLY - below the base by more than the half-width restores v2. The criterion as written is the one applied.