MojiDiff

← experiments

openmoji-g1-train-v2-renders-eff0038-9573bc74-9b9b1699

completed —

Hypothesis

A read-only probe, not a learning claim, so it carries no predeclared pass/fail criteria. The question is descriptive: at the v2 selected checkpoint's 0.2886 held-out aggregate and 0.0573 changed-token accuracy, what does the model's predicted clean state actually look like, and how much does it improve the render over the raw corrupted state it was given?

Why run it

Token accuracy is a poor proxy for whether a vector icon reads correctly, because a handful of wrong coordinates can destroy a shape that most-tokens-correct would call a success. Rendering the triple decides whether the Gate G scalars are overselling or underselling the visual result, and whether the corruption probability is even in a regime where the visual task is achievable.

Scope limits

Twelve held-out icons from the pilot's own 128-icon validation draw, one corruption draw each, at the single trained corruption probability of 0.35. Descriptive only. It trains nothing, changes no learning artifact, and is not a benchmark.

Measured result

Verbatim from the run's summary.json.

aggregate
x_hat_0-18
count
12.0
median_rgba_mae
0.15277021302348098
x_hat_0-72
count
12.0
median_rgba_mae
0.14243325995521666
x_t-18
count
12.0
median_rgba_mae
0.18287490922294844
x_t-72
count
12.0
median_rgba_mae
0.16954562303316387
checkpoint_sha256
e791493769907ac4…
checkpoint_step
840
config_sha256
9573bc740e90bf3a…
corruption
factorized_role_uniform_geometry, the pilot's held-out draw
corruption_probability
0.35
icons
12
pilot_config
configs/learning/openmoji-g1-dominant-bucket-train-v2-data-scale.yaml
pilot_config_sha256
c474c94db45be13f…
render_metrics_sha256
703c68a6fbe33043…
render_sizes
72, 18
rendered_rows
color/svg/1F994.svg, color/svg/E30A.svg, color/svg/1F6BE.svg, color/svg/1F3CA-1F3FB-200D-2642-FE0F.svg, color/svg/1F3CB-1F3FC-200D-2642-FE0F.svg, color/svg/1F93D-1F3FC.svg … and 6 more
schema_version
1
scope
read-only render probe over a trained pilot checkpoint; trains nothing
sheet_sha256
2e79da3ca8f92e3f…
study_version
openmoji-g1-train-v2-renders

Visual output

trajectory
trajectory.png

Written result

The v2 data-scale run reported 0.2886 held-out aggregate token accuracy and 0.0573 changed-token accuracy at its selected step. Neither number says whether the model's predicted clean state is a recognizable icon with some fields wrong or is rubble. This probe renders the triple for twelve held-out icons drawn from the pilot's own 128-icon validation set, using the pilot's own held-out corruption generator, so the pictures show exactly the inputs those scalars were computed on.

It is descriptive and carries no predeclared criteria. It trains nothing.

Result

state and sizemedian RGBA MAE against x_0
x_t 72 px0.169546
x_hat_0 72 px0.142433
x_t 18 px0.182875
x_hat_0 18 px0.152770

The prediction reduces median RGBA error against the clean render by 16.0% at 72 px and 16.5% at 18 px relative to the corrupted input it was given. A second complete invocation reproduced every artifact hash.

!Held-out trajectory contact sheet

Reading the sheet

None of the twelve predictions is a recognizable icon. Both x_t and x_hat_0 read as scribble at both sizes.

What does survive is instructive. Fills, styles, and topology are not corrupted by this process - only geometry fields are - so large flat colour regions come through intact: the warning triangle 26A0 keeps its yellow field, the tumbler 1F943 its orange, the 1F199 badge its green. The stroke geometry is destroyed in every case. On a few icons the prediction is visibly less tangled than the corrupted input; on others the two are indistinguishable.

What this changes

Two things, and the second matters more.

First, the Gate G scalars oversell the visual result substantially. A 16% reduction in median render error is not a recovered icon, and 0.2886 aggregate token accuracy should not be read as one. Any future claim about Gate G recovery needs a render beside it.

Second, at corruption probability 0.35 the input x_t is already visually destroyed. The model is being trained and evaluated at a corruption level where the visual task may not be achievable at any model size, because a third of the geometry is simply gone. The project has held this probability fixed since the Gate F fixtures, where it was chosen for four-icon experiments, and has never varied it at corpus scale. That makes the corruption schedule a first-class candidate factor for the next experiment rather than a secondary one, and it argues for training across a range of corruption levels rather than at a single fixed point.

This is twelve icons at one corruption draw and one probability. It is a description of one checkpoint, not a benchmark.

State transitions

  1. completed2026-09-20T17:40:00Zscalars_oversell_the_visual_result_no_prediction_is_a_recognizable_icon

Run record

Verbatim from runs/openmoji-g1-train-v2-renders-eff0038-9573bc74-9b9b1699/run.yaml, the record committed before launch.

schema_version
1
run_id
openmoji-g1-train-v2-renders-eff0038-9573bc74-9b9b1699
state
completed
parent_run
openmoji-g1-train-v2-datascale-a50b2e0-c474c94d-9b9b1699
recorded_retroactively
true
hypothesis
A read-only probe, not a learning claim, so it carries no predeclared pass/fail criteria. The question is descriptive: at the v2 selected checkpoint's 0.2886 held-out aggregate and 0.0573 changed-token accuracy, what does the model's predicted clean state actually look like, and how much does it improve the render over the raw corrupted state it was given?
expected_information_gain
Token accuracy is a poor proxy for whether a vector icon reads correctly, because a handful of wrong coordinates can destroy a shape that most-tokens-correct would call a success. Rendering the triple decides whether the Gate G scalars are overselling or underselling the visual result, and whether the corruption probability is even in a regime where the visual task is achievable.
scope_limits
Twelve held-out icons from the pilot's own 128-icon validation draw, one corruption draw each, at the single trained corruption probability of 0.35. Descriptive only. It trains nothing, changes no learning artifact, and is not a benchmark.
code
git_commit
eff003833b6a4cb0a95a2ea1dbc2b0bd946c0f23
execution_mode
native-local-cpu
config
path
configs/learning/openmoji-g1-train-v2-renders.yaml
sha256
9573bc740e90bf3a…
inputs
checkpoint
/home/dev/.cache/openmoji-g1-train-v2-datascale-a50b2e0-c474c94d-9b9b1699/checkpoint/checkpoint.zip
checkpoint_sha256
e791493769907ac4…
checkpoint_step
840
hybrid_sha256
9b9b1699677a6f97…
corruption
factorized_role_uniform_geometry, the pilot's own held-out draw
corruption_probability
0.35
execution
icons
12
render_sizes
72, 18
renderer
isolated bounded subprocess, 20 s timeout
network
none
result
median_rgba_mae_vs_x_0
x_t_72
0.16954562303316387
x_hat_0_72
0.14243325995521666
x_t_18
0.18287490922294844
x_hat_0_18
0.15277021302348098
improvement_over_corrupted_input_72
0.16
visual_assessment
None of the twelve predictions is a recognizable icon. Both x_t and x_hat_0 read as scribble. Large flat colour regions survive because fills, styles, and topology are not corrupted, so a few icons keep a recognisable palette - the warning triangle stays yellow, the tumbler stays orange - but the stroke geometry is destroyed in every case. The prediction is visibly less tangled than the corrupted input on some icons and indistinguishable on others.
identical_rerun
true
render_metrics_sha256
703c68a6fbe33043…
sheet_sha256
2e79da3ca8f92e3f…
conclusion
The Gate G scalars oversell the visual result substantially. A 16% reduction in median RGBA error against the corrupted input is not a recovered icon, and no reading of 0.2886 aggregate token accuracy should have been taken as one. It also shows that at 0.35 the corrupted input is already visually destroyed, so the training distribution sits at a corruption level where the visual task may be unachievable for any model of this size. That makes the corruption schedule a first-class candidate factor rather than a secondary one.
outputs
local_metadata
runs/openmoji-g1-train-v2-renders-eff0038-9573bc74-9b9b1699
report_root
reports/learning/openmoji-g1-train-v2-renders
completed_at
2026-09-20 17:40:00+00:00