MojiDiff

← experiments

openmoji-g1-train-ec2436b-f2a06ca5-9b9b1699

completed falsified

Hypothesis

With the selected packed representation, factorized role-uniform geometry corruption, and group/subgroup conditioning, a 577,552-parameter denoiser trained on a 256-icon family-disjoint subsample of the dominant P32/T128 bucket recovers corrupted geometry fields on 128 held-out icons far above its own untrained same-input control, and preserves already-correct fields reliably.

Why run it

This is the first Gate G run whose outcome is a learning claim rather than pipeline liveness. It moves the icon count from the 4-icon Gate F fixtures to 256 diverse icons across 10 groups and 75 subgroups, and it decides whether a larger full-split run is justified.

Scope limits

Fixed-topology and geometry-only. Not unconditional generation, not topology generation, not a style or AR comparison. A single seed, so small differences are not interpretable across seeds.

Measured behaviour

Held-out loss
lossoptimizer step8101214160200400600
Table view
optimizer stepheld-out loss
015.3908
6010.7155
1209.9711
1809.7997
2409.7932
30010.1371
36010.2798
42010.5504
48010.7512
54011.1266
60011.1485
Held-out token accuracy
aggregatechanged fieldsretained fields
accuracyoptimizer step00.10.20.30.40200400600
Table view
optimizer stepaggregatechanged fieldsretained fields
00.00370.00300.0041
600.06840.02190.0934
1200.14530.03460.2048
1800.18930.04200.2685
2400.20900.04390.2977
3000.21860.04580.3114
3600.22160.04650.3157
4200.22690.04990.3219
4800.23110.05300.3267
5400.23080.05590.3247
6000.23110.05520.3256
Training token accuracy (per batch)
accuracyoptimizer step00.20.40.6200400600
Table view
optimizer steptrain accuracy
10.0048
20.0030
30.0041
40.0032
50.0047
60.0075
70.0058
80.0079
90.0095
100.0094
110.0070
120.0082
130.0108
140.0116
150.0130
160.0155
170.0238
180.0183
190.0191
200.0214
210.0185
220.0191
230.0181
240.0227
250.0270
260.0270
270.0211
280.0222
290.0297
300.0307
310.0325
320.0355
330.0525
340.0411
350.0451
360.0456
370.0412
380.0407
390.0343
400.0494
410.0566
420.0523
430.0396
440.0424
450.0599
460.0569
470.0605
480.0614
490.0811
500.0752
510.0685
520.0810
530.0718
540.0628
550.0571
560.0802
570.1001
580.0736
590.0712
600.0831
610.0889
620.0960
630.1002
640.1067
650.1208
660.1161
670.1009
680.1324
690.0982
700.0854
710.0844
720.1146
730.1317
740.1059
750.1076
760.1154
770.1236
780.1427
790.1463
800.1431
810.1505
820.1669
830.1358
840.1820
850.1250
860.1213
870.1196
880.1514
890.1848
900.1458
910.1423
920.1538
930.1729
940.1923
950.1924
960.1793
970.1877
980.2074
990.1775
1000.2432
1010.1639
1020.1434
1030.1506
1040.1944
1050.2309
1060.1869
1070.1743
1080.1821
1090.2035
1100.2371
1110.2260
1120.2160
1130.2209
1140.2563
1150.2098
1160.2855
1170.1856
1180.1783
1190.1877
1200.2221
1210.2789
1220.2109
1230.2091
1240.2310
1250.2439
1260.2809
1270.2601
1280.2569
1290.2737
1300.2886
1310.2391
1320.3364
1330.2103
1340.2027
1350.2267
1360.2431
1370.3059
1380.2346
1390.2376
1400.2591
1410.2657
1420.3132
1430.2806
1440.2871
1450.2938
1460.3231
1470.2570
1480.3779
1490.2335
1500.2280
1510.2489
1520.2699
1530.3426
1540.2609
1550.2573
1560.2823
1570.2929
1580.3461
1590.3087
1600.3041
1610.3222
1620.3640
1630.2873
1640.3927
1650.2544
1660.2594
1670.2655
1680.2918
1690.3568
1700.2778
1710.2823
1720.3070
1730.3017
1740.3584
1750.3245
1760.3279
1770.3548
1780.3793
1790.3011
1800.4093
1810.2690
1820.2813
1830.2803
1840.3092
1850.3799
1860.2965
1870.3213
1880.3217
1890.3223
1900.3792
1910.3312
1920.3355
1930.3527
1940.4069
1950.3074
1960.4367
1970.2897
1980.3026
1990.2961
2000.3162
2010.3952
2020.3044
2030.3229
2040.3403
2050.3299
2060.4012
2070.3489
2080.3648
2090.3905
2100.3955
2110.3186
2120.4448
2130.2972
2140.3064
2150.3081
2160.3316
2170.4137
2180.3259
2190.3433
2200.3395
2210.3461
2220.4082
2230.3687
2240.3834
2250.4073
2260.4308
2270.3434
2280.4618
2290.3160
2300.3267
2310.3266
2320.3352
2330.4232
2340.3261
2350.3350
2360.3629
2370.3602
2380.4183
2390.3743
2400.3918
2410.4135
2420.4505
2430.3461
2440.4723
2450.3215
2460.3285
2470.3294
2480.3455
2490.4332
2500.3438
2510.3621
2520.3812
2530.3580
2540.4326
2550.3920
2560.3997
2570.4312
2580.4489
2590.3540
2600.5025
2610.3239
2620.3466
2630.3591
2640.3504
2650.4413
2660.3473
2670.3744
2680.3766
2690.3560
2700.4349
2710.3903
2720.4137
2730.4379
2740.4434
2750.3725
2760.5030
2770.3436
2780.3629
2790.3497
2800.3660
2810.4520
2820.3542
2830.3751
2840.4027
2850.3825
2860.4399
2870.4204
2880.4130
2890.4368
2900.4593
2910.3810
2920.5019
2930.3523
2940.3674
2950.3690
2960.3614
2970.4642
2980.3569
2990.3741
3000.3987
3010.3890
3020.4553
3030.4076
3040.4280
3050.4467
3060.4657
3070.3735
3080.5139
3090.3556
3100.3772
3110.3678
3120.3885
3130.4765
3140.3688
3150.4094
3160.4181
3170.3907
3180.4598
3190.4245
3200.4326
3210.4553
3220.4759
3230.4006
3240.5247
3250.3718
3260.3785
3270.3817
3280.3975
3290.4642
3300.3815
3310.3864
3320.4272
3330.3993
3340.4747
3350.4076
3360.4370
3370.4655
3380.4860
3390.3821
3400.5239
3410.3763
3420.4046
3430.3742
3440.3852
3450.4830
3460.3713
3470.4117
3480.4179
3490.4086
3500.4793
3510.4193
3520.4331
3530.4609
3540.4856
3550.3969
3560.5275
3570.3783
3580.3882
3590.3896
3600.4013
3610.4816
3620.3853
3630.4103
3640.4327
3650.4081
3660.4787
3670.4322
3680.4433
3690.4818
3700.5014
3710.4069
3720.5365
3730.3834
3740.4161
3750.3926
3760.4090
3770.4912
3780.3794
3790.4305
3800.4380
3810.4135
3820.4691
3830.4423
3840.4406
3850.4837
3860.4990
3870.4109
3880.5421
3890.3893
3900.4008
3910.3963
3920.4152
3930.4981
3940.3973
3950.4212
3960.4354
3970.4211
3980.4771
3990.4417
4000.4556
4010.4727
4020.5144
4030.4107
4040.5573
4050.3966
4060.4068
4070.3954
4080.4103
4090.5005
4100.3851
4110.4286
4120.4485
4130.4299
4140.4864
4150.4537
4160.4454
4170.4799
4180.5080
4190.4103
4200.5654
4210.3822
4220.4239
4230.4188
4240.4176
4250.4935
4260.3996
4270.4240
4280.4561
4290.4310
4300.4835
4310.4551
4320.4580
4330.5040
4340.5134
4350.4197
4360.5598
4370.3886
4380.4319
4390.4191
4400.4319
4410.5084
4420.4124
4430.4356
4440.4544
4450.4316
4460.4987
4470.4530
4480.4512
4490.4853
4500.5176
4510.4176
4520.5558
4530.4032
4540.4274
4550.4184
4560.4211
4570.5112
4580.4075
4590.4339
4600.4589
4610.4347
4620.5066
4630.4544
4640.4624
4650.5024
4660.5331
4670.4240
4680.5554
4690.4030
4700.4299
4710.4222
4720.4328
4730.5068
4740.4044
4750.4416
4760.4814
4770.4293
4780.5006
4790.4608
4800.4627
4810.5013
4820.5281
4830.4351
4840.5826
4850.4148
4860.4347
4870.4234
4880.4332
4890.5188
4900.4111
4910.4513
4920.4692
4930.4341
4940.5155
4950.4561
4960.4607
4970.5292
4980.5249
4990.4298
5000.5791
5010.4114
5020.4335
5030.4226
5040.4246
5050.5333
5060.4294
5070.4495
5080.4637
5090.4316
5100.5019
5110.4705
5120.4633
5130.5088
5140.5463
5150.4360
5160.5738
5170.4185
5180.4498
5190.4297
5200.4341
5210.5263
5220.4284
5230.4527
5240.4711
5250.4477
5260.5113
5270.4650
5280.4848
5290.5228
5300.5417
5310.4467
5320.5861
5330.4134
5340.4455
5350.4364
5360.4346
5370.5242
5380.4380
5390.4624
5400.4681
5410.4440
5420.5165
5430.4719
5440.4768
5450.5174
5460.5345
5470.4431
5480.5858
5490.4181
5500.4515
5510.4321
5520.4509
5530.5452
5540.4240
5550.4578
5560.4850
5570.4439
5580.5173
5590.4783
5600.4839
5610.5257
5620.5397
5630.4459
5640.5919
5650.4181
5660.4606
5670.4286
5680.4498
5690.5249
5700.4339
5710.4701
5720.4861
5730.4530
5740.5041
5750.4728
5760.4867
5770.5158
5780.5453
5790.4457
5800.5987
5810.4250
5820.4598
5830.4378
5840.4511
5850.5401
5860.4345
5870.4657
5880.4882
5890.4462
5900.5205
5910.4747
5920.4916
5930.5142
5940.5481
5950.4500
5960.5944
5970.4264
5980.4618
5990.4332
6000.4550

Measured result

Verbatim from the run's summary.json.

bucket
bucket-p32-t128
bucket_icons
3359
checkpoint_bytes
7075311
checkpoint_round_trip
true
checkpoint_sha256
ac696b81b3c9cd6f…
config_sha256
f2a06ca55cd18a36…
corruption
factorized_role_uniform_geometry
cuda_version
12.9
deterministic_algorithms
true
device
cuda
eval_every
60
final_train
loss
2.7220990657806396
step
600
train_token_accuracy
0.45504627589246366
group_vocabulary_size
12
locked_path_exact
true
metrics_sha256
b5216607e3631be5…
model_parameters
577552
schema_version
1
scope
dominant-bucket fixed-topology geometry pilot; not unconditional generation
selected_train_rows
color/svg/1F3F4-E0064-E0065-E0062-E0065-E007F.svg, color/svg/1F468-1F3FC-200D-1F9BC.svg, color/svg/1F561.svg, color/svg/1F469-1F3FE-200D-1F9BC.svg, color/svg/1F469-1F3FD-200D-2764-FE0F-200D-1F48B-200D-1F469-1F3FE.svg, color/svg/1F6B5-1F3FB-200D-2642-FE0F.svg … and 250 more
selected_validation_rows
color/svg/1F994.svg, color/svg/E30A.svg, color/svg/1F6BE.svg, color/svg/1F3CA-1F3FB-200D-2642-FE0F.svg, color/svg/1F3CB-1F3FC-200D-2642-FE0F.svg, color/svg/1F93D-1F3FC.svg … and 122 more
steps
600
study_version
openmoji-g1-dominant-bucket-train-v1
subgroup_vocabulary_size
118
torch_version
2.8.0a0+5228986c39.nv25.06
validation
accuracy
0.23112877424515096
changed_accuracy
0.05522964658749464
changed_total
13978
loss
11.148523330688477
retained_accuracy
0.32558586246638493
retained_total
26030
validation_trace_sha256
06dfa07b62c2c59d…
validation_untrained
accuracy
0.003699260147970406
changed_accuracy
0.003004721705537273
changed_total
13978
loss
15.390789985656738
retained_accuracy
0.004072224356511717
retained_total
26030

Written result

Status: staged; training run not yet started.

First Gate G run whose outcome is a learning claim rather than pipeline liveness. It trains the selected 577,552-parameter geometry denoiser for 600 bounded steps on a 256-icon family-disjoint subsample of the exact P32/T128 bucket and evaluates on 128 held-out icons from the disjoint validation split, tracing held-out recovery every 60 steps against an untrained step-0 control on the identical corruption draw.

The predeclared pass/fail criteria are recorded in run.yaml before launch. They are deliberately falsifiable: the retained-token bar in particular is a real bet, because a 60-step CPU preflight reached only 0.0934 retained accuracy.

Result

Completed on the owned RTX 4080 in about 22 seconds including stage verification. Two complete invocations produced identical checkpoint and summary digests.

stepheld-out aggregatechangedretainedheld-out loss
0 (untrained)0.00370.00300.004115.3908
600.06840.02190.093410.7155
1200.14530.03460.20489.9711
1800.18930.04200.26859.7997
2400.20900.04390.29779.7932
3000.21860.04580.311410.1371
3600.22160.04650.315710.2798
4200.22690.04990.321910.5504
4800.23110.05300.326710.7512
5400.23080.05590.324711.1266
6000.23110.05520.325611.1485

Held-out totals are 13,978 changed and 26,030 retained fields.

Predeclared criteria

criterionthresholdobservedoutcome
recovery above control>= 10x18.38xpass
retained preservation>= 0.900.3256fail
monotone learning>= 8/109/10pass
structural safetylocked + round tripboth truepass
reproducibilityidentical rerunidenticalpass

Overall: falsified. The run is retained as a negative result rather than retuned.

Interpretation

The representation is clearly learnable: held-out changed-token recovery reaches 18.4x its own untrained same-input control, and the improvement is monotone. That is a real result at 256 diverse icons across 10 groups and 75 subgroups, far beyond the 4-icon Gate F fixtures.

But the run also falsifies the retained-preservation hypothesis at this data scale, and the trace shows why. Held-out loss reaches its minimum of 9.7932 at step 240 and then rises steadily to 11.1485 while training token accuracy climbs to 0.4550 and training loss falls to 2.7221. Held-out aggregate accuracy is flat at about 0.231 from step 240 onward. This is overfitting to 256 icons, not an optimization failure.

One nuance worth recording rather than smoothing over: held-out changed accuracy keeps creeping up (0.0439 to 0.0552) across the same interval in which held-out loss worsens. The monotone criterion therefore passed while the model was already generalizing worse overall, so that criterion is weaker evidence than it looks in isolation. Aggregate held-out loss is the more honest scalar here.

The 0.90 retained bar was a deliberate bet and it lost by a wide margin. Nothing about this run suggests the bar itself was wrong for a usable denoiser; it suggests 256 icons is too little data for this model to reach it.

Next experiment this justifies

More data, not more steps. The dominant bucket has a 2,681-icon family-disjoint train split, ten times what this run used, and loading it costs about three minutes of CPU at the measured 60 ms per icon. The controlled next run should change only the train-split size, hold the model, seed, corruption, batch, and learning rate fixed, and use an early-stopping or best-checkpoint policy keyed on held-out loss rather than a fixed 600-step budget.

State transitions

  1. planned2026-09-20T11:25:37Z
  2. staged2026-09-20T11:26:52Z
  3. completed2026-09-20T11:28:49Zlearns_far_above_untrained_control_but_overfits_256_icons_and_misses_retained_preservation_bar

Run record

Verbatim from runs/openmoji-g1-train-ec2436b-f2a06ca5-9b9b1699/run.yaml, the record committed before launch.

schema_version
1
run_id
openmoji-g1-train-ec2436b-f2a06ca5-9b9b1699
state
completed
parent_run
openmoji-g1-gpu-7ba1aa4-0bafd5c-9b9b1699
hypothesis
With the selected packed representation, factorized role-uniform geometry corruption, and group/subgroup conditioning, a 577,552-parameter denoiser trained on a 256-icon family-disjoint subsample of the dominant P32/T128 bucket recovers corrupted geometry fields on 128 held-out icons far above its own untrained same-input control, and preserves already-correct fields reliably.
expected_information_gain
This is the first Gate G run whose outcome is a learning claim rather than pipeline liveness. It moves the icon count from the 4-icon Gate F fixtures to 256 diverse icons across 10 groups and 75 subgroups, and it decides whether a larger full-split run is justified.
scope_limits
Fixed-topology and geometry-only. Not unconditional generation, not topology generation, not a style or AR comparison. A single seed, so small differences are not interpretable across seeds.
code
git_commit
ec2436b006be076a4858868ff77e5bb5e172d523
snapshot_mode
git-archive-plus-hash-pinned-384-svg-training-fixture
archive_sha256
efe3f4093ba77bc2…
tree_sha256
1f97ba11d1031ee4…
config
path
configs/learning/openmoji-g1-dominant-bucket-train-v1.yaml
sha256
f2a06ca55cd18a36…
dataset
hybrid_sha256
9b9b1699677a6f97…
assignments_sha256
e0cdb2a3cc8f00df…
bucket
bucket-p32-t128
bucket_icons
3359
train_samples
256
validation_samples
128
staged_raw_svg_count
384
family_disjoint_train_validation
true
train_groups
10
train_subgroups
75
worker
alias
owned-gpu
gpu
NVIDIA GeForce RTX 4080
driver
595.71.05
image
mojidiff/owned-gpu-smoke:cd3250e
image_digest
sha256:f42d0acc2a768b6929a982dbb1afcfb342b382751ea0e974ba00052e5dff38ca
torch
2.8.0a0+5228986c39.nv25.06
cuda
12.9
execution
smoke_id
openmoji-g1-train-v1
seed_set
3101
deterministic_algorithms
true
cublas_workspace_config
:4096:8
optimizer_steps
600
batch_size
16
learning_rate
0.001
corruption_probability
0.35
eval_every
60
resource_cap
max_steps
20000
max_storage_gb
50
network
none
checkpoint_policy
retain checkpoint and compact evidence in the persistent artifact volume
predeclared_criteria
note
Declared before launch. A 60-step CPU preflight was run first to size memory and wall time; it reported an untrained held-out changed-token accuracy of 0.003005 and 0.021892 after 60 steps. That preflight informed these thresholds and is disclosed rather than hidden. The full 600-step outcome was not observed beforehand.
primary
id
held_out_recovery_above_control
statement
Held-out changed-token accuracy is at least 10x the untrained same-input control evaluated on the identical corruption draw.
id
retained_preservation
statement
Held-out retained-token accuracy is at least 0.90.
id
monotone_learning
statement
Held-out changed-token accuracy increases in at least 8 of the 10 recorded trace intervals.
id
structural_safety
statement
Locked-path exactness holds and the canonical checkpoint round-trips.
id
reproducibility
statement
A second complete invocation returns an identical JSON result.
reported_not_gated
absolute held-out changed-token accuracy, held-out aggregate accuracy and loss, the complete 11-point validation trace
chance_level
Random legal-token choice is about 0.0035 for endpoints over 289 bins and 0.0024 for controls over 417 bins, so the untrained control near 0.003 is at chance.
outputs
local_metadata
runs/openmoji-g1-train-ec2436b-f2a06ca5-9b9b1699
remote_workspace
/home/dev/workspace/openmoji-g1-train-ec2436b-f2a06ca5-9b9b1699
durable_artifacts
/home/dev/.cache/openmoji-g1-train-ec2436b-f2a06ca5-9b9b1699
planned_at
2026-09-20 11:25:37+00:00
staged_at
2026-09-20 11:26:52+00:00
completed_at
2026-09-20 11:28:49+00:00
result
device
cuda
wall_seconds_including_stage_verification
22
final_train_loss
2.7220990657806396
final_train_token_accuracy
0.45504627589246366
held_out_untrained_changed_accuracy
0.003004721705537273
held_out_changed_accuracy
0.05522964658749464
held_out_retained_accuracy
0.32558586246638493
held_out_aggregate_accuracy
0.2311
held_out_changed_total
13978
held_out_retained_total
26030
held_out_loss_minimum_step
240
held_out_loss_minimum
9.7932
held_out_loss_final
11.1485
checkpoint_sha256
ac696b81b3c9cd6f…
summary_sha256
b892acbf28078405…
identical_rerun
true
artifact_bytes
7140279
predeclared_outcome
overall
falsified
held_out_recovery_above_control
passed
true
observed_ratio
18.38
threshold
10
retained_preservation
passed
false
observed
0.32558586246638493
threshold
0.9
monotone_learning
passed
true
observed_increasing_intervals
9
threshold
8
structural_safety
passed
true
reproducibility
passed
true