MojiDiff

← experiments

openmoji-g1-train-v2-datascale-a50b2e0-c474c94d-9b9b1699

completed partially-falsified

Hypothesis

The v1 falsification was a data-scale failure, not an optimization or plumbing failure. Training the identical 577,552-parameter denoiser on the full 2,681-icon family-disjoint train split of the dominant P32/T128 bucket, instead of a 256-icon subsample, generalizes materially better on the identical 128-icon held-out draw: a lower held-out loss minimum, a narrower train/held-out gap, and higher changed-token recovery at the selected checkpoint.

Why run it

v1 established that the representation is learnable at real icon diversity but that 256 icons are too few to generalize. This run decides whether 10.5x more data is sufficient at this model size, or whether the next factor is model capacity, augmentation, or the corruption schedule. It is the cheapest test that separates those, because it changes only the data.

Scope limits

Fixed-topology and geometry-only. Not unconditional generation, not topology generation, not a style or AR comparison. A single seed, so small differences are not interpretable across seeds. Group/subgroup conditioning and exact path locks are unchanged from v1.

Measured behaviour

Held-out loss
lossoptimizer step57.51012.51517.505001000selected
Table view
optimizer stepheld-out loss
015.3908
6010.5741
1209.3253
1808.6219
2408.2993
3008.0193
3607.6764
4207.6276
4807.5239
5407.5360
6007.4659
6607.4575
7207.4163
7807.3897
8407.3333
9007.3702
9607.5195
10207.4133
10807.5393
11407.5728
12007.6322
12607.6349
13207.5275
Held-out token accuracy
aggregatechanged fieldsretained fields
accuracyoptimizer step00.20.40.605001000selected
Table view
optimizer stepaggregatechanged fieldsretained fields
00.00370.00300.0041
600.07000.02280.0954
1200.14750.03280.2091
1800.20020.03830.2871
2400.22530.04080.3243
3000.24660.04200.3564
3600.25660.04450.3706
4200.26270.04690.3786
4800.26920.04840.3878
5400.27350.04930.3940
6000.27850.05190.4002
6600.28230.05320.4054
7200.28240.05320.4055
7800.28440.05420.4080
8400.28860.05730.4128
9000.28770.05590.4122
9600.29020.05840.4146
10200.29200.06000.4167
10800.29120.06150.4146
11400.29500.06170.4203
12000.29580.06370.4204
12600.29610.06460.4205
13200.29990.06620.4254
Training token accuracy (per batch)
accuracyoptimizer step00.20.40.65001000selected
Table view
optimizer steptrain accuracy
10.0048
20.0030
30.0041
40.0032
50.0047
60.0075
70.0058
80.0079
90.0095
100.0094
110.0070
120.0082
130.0108
140.0116
150.0130
160.0099
170.0103
180.0174
190.0141
200.0162
210.0178
220.0188
230.0286
240.0155
250.0183
260.0321
270.0216
280.0354
290.0332
300.0318
310.0297
320.0298
330.0194
340.0511
350.0386
360.0437
370.0288
380.0474
390.0427
400.0276
410.0565
420.0628
430.0341
440.0533
450.0516
460.0360
470.0521
480.0686
490.0677
500.0442
510.0711
520.0752
530.0510
540.0538
550.0798
560.0395
570.0851
580.0856
590.0876
600.0793
610.0694
620.0800
630.1154
640.0771
650.0930
660.0740
670.0718
680.1153
690.0853
700.0919
710.1346
720.1337
730.1241
740.1374
750.0978
760.0948
770.0895
780.1333
790.0964
800.1091
810.1181
820.1144
830.1479
840.1129
850.1459
860.1490
870.1260
880.1226
890.1376
900.1519
910.1672
920.1527
930.0963
940.1137
950.1227
960.1483
970.1368
980.1170
990.1049
1000.1817
1010.2118
1020.1545
1030.1529
1040.1797
1050.1804
1060.1377
1070.1536
1080.1738
1090.2224
1100.1484
1110.2099
1120.1568
1130.1334
1140.1676
1150.1468
1160.1550
1170.1751
1180.1611
1190.2383
1200.1449
1210.1413
1220.1993
1230.1822
1240.1181
1250.2077
1260.1398
1270.1585
1280.2737
1290.2315
1300.3062
1310.1176
1320.2193
1330.2028
1340.1959
1350.1785
1360.2779
1370.1867
1380.2772
1390.1730
1400.2237
1410.2323
1420.2325
1430.2361
1440.1837
1450.1979
1460.2133
1470.2515
1480.2492
1490.1939
1500.2634
1510.2224
1520.2197
1530.2620
1540.1874
1550.2366
1560.1940
1570.1763
1580.2295
1590.2211
1600.2036
1610.1924
1620.2371
1630.2143
1640.2364
1650.2564
1660.1625
1670.2136
1680.2702
1690.1804
1700.2574
1710.2716
1720.2756
1730.1926
1740.1836
1750.2797
1760.2362
1770.2927
1780.2173
1790.1852
1800.3242
1810.2478
1820.2985
1830.2988
1840.2265
1850.2714
1860.2554
1870.3537
1880.2747
1890.2373
1900.2625
1910.2829
1920.2778
1930.2929
1940.2578
1950.2861
1960.3002
1970.2764
1980.2768
1990.2951
2000.2767
2010.2098
2020.3484
2030.2957
2040.3058
2050.2782
2060.2776
2070.2664
2080.2151
2090.3318
2100.2847
2110.2081
2120.3535
2130.2389
2140.2451
2150.2405
2160.3215
2170.2646
2180.1920
2190.3315
2200.2434
2210.2710
2220.2842
2230.2896
2240.2932
2250.2324
2260.3043
2270.3167
2280.3020
2290.2998
2300.3288
2310.2908
2320.2914
2330.3107
2340.2987
2350.2081
2360.3689
2370.2463
2380.3004
2390.4222
2400.2879
2410.3669
2420.2982
2430.2866
2440.2738
2450.2546
2460.3053
2470.3052
2480.2741
2490.2228
2500.3631
2510.3138
2520.3224
2530.3542
2540.2789
2550.2887
2560.3129
2570.3384
2580.3086
2590.4179
2600.3088
2610.2272
2620.2753
2630.3288
2640.2832
2650.3018
2660.2898
2670.3000
2680.2851
2690.3496
2700.3325
2710.2977
2720.3732
2730.2982
2740.2752
2750.2769
2760.3746
2770.3512
2780.2629
2790.4323
2800.2651
2810.2676
2820.2999
2830.3302
2840.2786
2850.2947
2860.2952
2870.3681
2880.2827
2890.3176
2900.3422
2910.2500
2920.3014
2930.2581
2940.2842
2950.3168
2960.3465
2970.3925
2980.3679
2990.2767
3000.3331
3010.3516
3020.2730
3030.3439
3040.3503
3050.3454
3060.2997
3070.2823
3080.3531
3090.3260
3100.4149
3110.2714
3120.2967
3130.3053
3140.3526
3150.3497
3160.3364
3170.2640
3180.3967
3190.3270
3200.2981
3210.3505
3220.2418
3230.3550
3240.2553
3250.2800
3260.3281
3270.3351
3280.2602
3290.3005
3300.3044
3310.3144
3320.3513
3330.3235
3340.2451
3350.3758
3360.3047
3370.3417
3380.3000
3390.4273
3400.2688
3410.2705
3420.2814
3430.3649
3440.3677
3450.3256
3460.2824
3470.3226
3480.3274
3490.3837
3500.3801
3510.3611
3520.3394
3530.3228
3540.4492
3550.4096
3560.3074
3570.3331
3580.3657
3590.2974
3600.3852
3610.3648
3620.2832
3630.3811
3640.3838
3650.3687
3660.3411
3670.3666
3680.2938
3690.4103
3700.3714
3710.4094
3720.2854
3730.3775
3740.2836
3750.3520
3760.3328
3770.4026
3780.2718
3790.3779
3800.3724
3810.2834
3820.3116
3830.3347
3840.3781
3850.2744
3860.3409
3870.3218
3880.2792
3890.3403
3900.3832
3910.2820
3920.3641
3930.3107
3940.4123
3950.3454
3960.3710
3970.3259
3980.4129
3990.3022
4000.4215
4010.3817
4020.2922
4030.3753
4040.3510
4050.3592
4060.4078
4070.4376
4080.3616
4090.4025
4100.3053
4110.3289
4120.3208
4130.3710
4140.3136
4150.3065
4160.3357
4170.3487
4180.3991
4190.3470
4200.4145
4210.4108
4220.3545
4230.3347
4240.3466
4250.3917
4260.4643
4270.4124
4280.2874
4290.2641
4300.3511
4310.3371
4320.3995
4330.2933
4340.3070
4350.3766
4360.3820
4370.4034
4380.3140
4390.3993
4400.4019
4410.3346
4420.3238
4430.3867
4440.4166
4450.3572
4460.3955
4470.3419
4480.3182
4490.3352
4500.3317
4510.3192
4520.3426
4530.3330
4540.3538
4550.3796
4560.3118
4570.3936
4580.3492
4590.3021
4600.3626
4610.3075
4620.2995
4630.3997
4640.4122
4650.4563
4660.3054
4670.3614
4680.3893
4690.3538
4700.3424
4710.4480
4720.3377
4730.4224
4740.3102
4750.3658
4760.3737
4770.4175
4780.4061
4790.2842
4800.3653
4810.3348
4820.3733
4830.3927
4840.3685
4850.3610
4860.3856
4870.3685
4880.3659
4890.3266
4900.3072
4910.3762
4920.2919
4930.3535
4940.3229
4950.3515
4960.2714
4970.3813
4980.3335
4990.3757
5000.4047
5010.2942
5020.3317
5030.4300
5040.3004
5050.3797
5060.3881
5070.3991
5080.2906
5090.2767
5100.4104
5110.3253
5120.4308
5130.3227
5140.2788
5150.4779
5160.3464
5170.4330
5180.3945
5190.3747
5200.3843
5210.3613
5220.4881
5230.3757
5240.3383
5250.3397
5260.4232
5270.3728
5280.4080
5290.3449
5300.3706
5310.4424
5320.3815
5330.4173
5340.4058
5350.3982
5360.2845
5370.4364
5380.3865
5390.4053
5400.3630
5410.3901
5420.3593
5430.2912
5440.4369
5450.4026
5460.3258
5470.4309
5480.2981
5490.3909
5500.3116
5510.4033
5520.3818
5530.2646
5540.4285
5550.3062
5560.3246
5570.3953
5580.3858
5590.3296
5600.3368
5610.4106
5620.4148
5630.3205
5640.4151
5650.3989
5660.3854
5670.3853
5680.4064
5690.3843
5700.2771
5710.4662
5720.3350
5730.3872
5740.5077
5750.3776
5760.4269
5770.3935
5780.3367
5790.3924
5800.3196
5810.4227
5820.3735
5830.3537
5840.3095
5850.4364
5860.3972
5870.4095
5880.4534
5890.3760
5900.3614
5910.3809
5920.3751
5930.4205
5940.4974
5950.4038
5960.2908
5970.3084
5980.4123
5990.3536
6000.3581
6010.3690
6020.2848
6030.4302
6040.4655
6050.4088
6060.3824
6070.4069
6080.4267
6090.3410
6100.3365
6110.4420
6120.4438
6130.2946
6140.5055
6150.3404
6160.3264
6170.3778
6180.3969
6190.3614
6200.3624
6210.3286
6220.4371
6230.3263
6240.3839
6250.3620
6260.3780
6270.3604
6280.3178
6290.3443
6300.3299
6310.4399
6320.4401
6330.4718
6340.2872
6350.4398
6360.4076
6370.3247
6380.3899
6390.4365
6400.3409
6410.4211
6420.3305
6430.3926
6440.4226
6450.4504
6460.3714
6470.3246
6480.3559
6490.4073
6500.3885
6510.4283
6520.3249
6530.4737
6540.3793
6550.3505
6560.4234
6570.3036
6580.4041
6590.3059
6600.3380
6610.3681
6620.3912
6630.3236
6640.3425
6650.3588
6660.3714
6670.4165
6680.3665
6690.2968
6700.4371
6710.3602
6720.4164
6730.3261
6740.5218
6750.3560
6760.3272
6770.3200
6780.4183
6790.4237
6800.3987
6810.3030
6820.3767
6830.3906
6840.4406
6850.4536
6860.4284
6870.3640
6880.3959
6890.5025
6900.4604
6910.3488
6920.3775
6930.4100
6940.3696
6950.4136
6960.3885
6970.3856
6980.4196
6990.4345
7000.3836
7010.4175
7020.4045
7030.3664
7040.4489
7050.3753
7060.4448
7070.3721
7080.4170
7090.3818
7100.3710
7110.3384
7120.4800
7130.3304
7140.4199
7150.4284
7160.3175
7170.3743
7180.3517
7190.4498
7200.3482
7210.3621
7220.4082
7230.2795
7240.3729
7250.4050
7260.4024
7270.3814
7280.3714
7290.4459
7300.3912
7310.4274
7320.3800
7330.4661
7340.3053
7350.4653
7360.4187
7370.3598
7380.3746
7390.4287
7400.3560
7410.4491
7420.5208
7430.3630
7440.4996
7450.3279
7460.3737
7470.4091
7480.3950
7490.3876
7500.3563
7510.3737
7520.4060
7530.4378
7540.3652
7550.4468
7560.4527
7570.4514
7580.3493
7590.4212
7600.4263
7610.4856
7620.4927
7630.3705
7640.2931
7650.3887
7660.3819
7670.4004
7680.3710
7690.3519
7700.4032
7710.4357
7720.4181
7730.4132
7740.4592
7750.4107
7760.3965
7770.3881
7780.4207
7790.4440
7800.4225
7810.3899
7820.3983
7830.3744
7840.4043
7850.3235
7860.4124
7870.3481
7880.4140
7890.3727
7900.4037
7910.3585
7920.3925
7930.4015
7940.3485
7950.3803
7960.3485
7970.3444
7980.4259
7990.4292
8000.4744
8010.3990
8020.3753
8030.4341
8040.3889
8050.3544
8060.4886
8070.3636
8080.4328
8090.3676
8100.3586
8110.4657
8120.4385
8130.4625
8140.3059
8150.3982
8160.3659
8170.4084
8180.4000
8190.4019
8200.4117
8210.4266
8220.4159
8230.4160
8240.3457
8250.3167
8260.4367
8270.3343
8280.3918
8290.3537
8300.3654
8310.3235
8320.3995
8330.3297
8340.4071
8350.4698
8360.3319
8370.3210
8380.4789
8390.3514
8400.4154
8410.4367
8420.4367
8430.3201
8440.3053
8450.4505
8460.3741
8470.4605
8480.3510
8490.3205
8500.4657
8510.4031
8520.4661
8530.4890
8540.4006
8550.4149
8560.4197
8570.5192
8580.3827
8590.4107
8600.3532
8610.4536
8620.3762
8630.4396
8640.4200
8650.3777
8660.4776
8670.4114
8680.4583
8690.3913
8700.4800
8710.3228
8720.4700
8730.4365
8740.4505
8750.4028
8760.4373
8770.3814
8780.3123
8790.4770
8800.4437
8810.3392
8820.4714
8830.3627
8840.3834
8850.3294
8860.4324
8870.4289
8880.3168
8890.4186
8900.3486
8910.3572
8920.3919
8930.4853
8940.2959
8950.4177
8960.4037
8970.4749
8980.3983
8990.4812
9000.3689
9010.4548
9020.4162
9030.4571
9040.4232
9050.3177
9060.4611
9070.3963
9080.4224
9090.4915
9100.4728
9110.4067
9120.4513
9130.3533
9140.4077
9150.3468
9160.4477
9170.4084
9180.3236
9190.3988
9200.4329
9210.4162
9220.4264
9230.4709
9240.4756
9250.3768
9260.4020
9270.4288
9280.4543
9290.5002
9300.4588
9310.2978
9320.3281
9330.4504
9340.4034
9350.3962
9360.3967
9370.3139
9380.4534
9390.4771
9400.4705
9410.3595
9420.4874
9430.4311
9440.3692
9450.3858
9460.4800
9470.4663
9480.2949
9490.5457
9500.3808
9510.3491
9520.3845
9530.4078
9540.3819
9550.4024
9560.3725
9570.4491
9580.3610
9590.4238
9600.4005
9610.3967
9620.3485
9630.4140
9640.3279
9650.3678
9660.4826
9670.4628
9680.4805
9690.2980
9700.4532
9710.3787
9720.4008
9730.3985
9740.5162
9750.3388
9760.4253
9770.3445
9780.4233
9790.4434
9800.4791
9810.4160
9820.3461
9830.3810
9840.4230
9850.4174
9860.4669
9870.3389
9880.4490
9890.4704
9900.3156
9910.4805
9920.3323
9930.4047
9940.3659
9950.3438
9960.4043
9970.4007
9980.3505
9990.3716
10000.3763
10010.3713
10020.4555
10030.4169
10040.3102
10050.4450
10060.4012
10070.4012
10080.3774
10090.5004
10100.4415
10110.3218
10120.3596
10130.4242
10140.4221
10150.4451
10160.3549
10170.3744
10180.4766
10190.4227
10200.4994
10210.4525
10220.4100
10230.4187
10240.5037
10250.5113
10260.3644
10270.3896
10280.4456
10290.4152
10300.4255
10310.4296
10320.4125
10330.4028
10340.4745
10350.4145
10360.4668
10370.4218
10380.4300
10390.3407
10400.4920
10410.4305
10420.4311
10430.4572
10440.4025
10450.4004
10460.3587
10470.4807
10480.3796
10490.4609
10500.4285
10510.3272
10520.4402
10530.3394
10540.4258
10550.4129
10560.3613
10570.4239
10580.3106
10590.4041
10600.4250
10610.4244
10620.3969
10630.3924
10640.4491
10650.4437
10660.4646
10670.4197
10680.4432
10690.3836
10700.4821
10710.4319
10720.3917
10730.3637
10740.4869
10750.4099
10760.4509
10770.5244
10780.3817
10790.5784
10800.3470
10810.3808
10820.4229
10830.4057
10840.4167
10850.3772
10860.3996
10870.4025
10880.5073
10890.4188
10900.4506
10910.4815
10920.4595
10930.3332
10940.4629
10950.4270
10960.4518
10970.5622
10980.4065
10990.3224
11000.3667
11010.4728
11020.3865
11030.4030
11040.3830
11050.4164
11060.4470
11070.4154
11080.4608
11090.4714
11100.4230
11110.4262
11120.3719
11130.4239
11140.4891
11150.4673
11160.3775
11170.4718
11180.3973
11190.3984
11200.3891
11210.4465
11220.3550
11230.4314
11240.4069
11250.4368
11260.3886
11270.3991
11280.4398
11290.3809
11300.3799
11310.3858
11320.3282
11330.4066
11340.4461
11350.5142
11360.4339
11370.3866
11380.3779
11390.4539
11400.3603
11410.5271
11420.3677
11430.4645
11440.4133
11450.3173
11460.4897
11470.4551
11480.4823
11490.3397
11500.4373
11510.3689
11520.4179
11530.4310
11540.4423
11550.3737
11560.4669
11570.4308
11580.3834
11590.3909
11600.3435
11610.4462
11620.3380
11630.3853
11640.4452
11650.3783
11660.3451
11670.4342
11680.3422
11690.4341
11700.4814
11710.3686
11720.3349
11730.5018
11740.3448
11750.4298
11760.4360
11770.4945
11780.3284
11790.3364
11800.4221
11810.4252
11820.4798
11830.3646
11840.3487
11850.4723
11860.4006
11870.5254
11880.5053
11890.4182
11900.4335
11910.3849
11920.5590
11930.4755
11940.3534
11950.4085
11960.4860
11970.4063
11980.4692
11990.4338
12000.3415
12010.5482
12020.4179
12030.4730
12040.4200
12050.4843
12060.3431
12070.4927
12080.4584
12090.4747
12100.3632
12110.4869
12120.3791
12130.3737
12140.4859
12150.4813
12160.3798
12170.4563
12180.3876
12190.3883
12200.3478
12210.4363
12220.4847
12230.3430
12240.4287
12250.4031
12260.3381
12270.4356
12280.5005
12290.3240
12300.4272
12310.4083
12320.4843
12330.4074
12340.5220
12350.3420
12360.4952
12370.3950
12380.5042
12390.4721
12400.3157
12410.4949
12420.3806
12430.4544
12440.5128
12450.4691
12460.4407
12470.4608
12480.3762
12490.4462
12500.3923
12510.4588
12520.3823
12530.3824
12540.3933
12550.4472
12560.4154
12570.4552
12580.4690
12590.5047
12600.4103
12610.3996
12620.4505
12630.4754
12640.5156
12650.4981
12660.3310
12670.3516
12680.4138
12690.4628
12700.4096
12710.3998
12720.3246
12730.4753
12740.5042
12750.4440
12760.4160
12770.4732
12780.4684
12790.3983
12800.3972
12810.4723
12820.4969
12830.3225
12840.5698
12850.3709
12860.3803
12870.3951
12880.4540
12890.3716
12900.4111
12910.3931
12920.4931
12930.3611
12940.3701
12950.5061
12960.3882
12970.3419
12980.4252
12990.3491
13000.3579
13010.4713
13020.4593
13030.5576
13040.3178
13050.4339
13060.4376
13070.3918
13080.4170
13090.4930
13100.4136
13110.4460
13120.3608
13130.4266
13140.4789
13150.5171
13160.4435
13170.3619
13180.3976
13190.3913
13200.4534

Measured result

Verbatim from the run's summary.json.

bucket
bucket-p32-t128
bucket_icons
3359
checkpoint_bytes
7075311
checkpoint_round_trip
true
checkpoint_sha256
e791493769907ac4…
config_sha256
c474c94db45be13f…
corruption
factorized_role_uniform_geometry
cuda_version
13.4
deterministic_algorithms
true
device
cuda
eval_every
60
final_train
loss
4.393980503082275
step
1320
train_token_accuracy
0.453393135725429
group_vocabulary_size
12
locked_path_exact
true
metrics_sha256
bc974202b24f9542…
model_parameters
577552
schema_version
1
scope
dominant-bucket fixed-topology geometry pilot; not unconditional generation
selected_train_rows
color/svg/1F3F4-E0064-E0065-E0062-E0065-E007F.svg, color/svg/1F468-1F3FC-200D-1F9BC.svg, color/svg/1F561.svg, color/svg/1F469-1F3FE-200D-1F9BC.svg, color/svg/1F469-1F3FD-200D-2764-FE0F-200D-1F48B-200D-1F469-1F3FE.svg, color/svg/1F6B5-1F3FB-200D-2642-FE0F.svg … and 2675 more
selected_validation_rows
color/svg/1F994.svg, color/svg/E30A.svg, color/svg/1F6BE.svg, color/svg/1F3CA-1F3FB-200D-2642-FE0F.svg, color/svg/1F3CB-1F3FC-200D-2642-FE0F.svg, color/svg/1F93D-1F3FC.svg … and 122 more
selection
completed_steps
1320
evals_without_improvement
8
min_delta
0.0
objective
held_out_loss
patience_evals
8
selected_held_out_loss
7.3332648277282715
selected_step
840
stopped_early
true
steps
6300
study_version
openmoji-g1-dominant-bucket-train-v2-data-scale
subgroup_vocabulary_size
118
torch_version
2.14.0a0+4fdf77b940.nv26.08
validation
accuracy
0.28861727654469105
changed_accuracy
0.057304335384175134
changed_total
13978
loss
7.3332648277282715
retained_accuracy
0.41283134844410296
retained_total
26030
validation_final_step
accuracy
0.29994001199760045
changed_accuracy
0.06624695950779796
changed_total
13978
loss
7.527534484863281
retained_accuracy
0.425432193622743
retained_total
26030
validation_trace_sha256
3d59e8786842a1ec…
validation_untrained
accuracy
0.003699260147970406
changed_accuracy
0.003004721705537273
changed_total
13978
loss
15.390789985656738
retained_accuracy
0.004072224356511717
retained_total
26030

Written result

The v1 falsification left one question: was the failure a matter of data scale? v1 trained the selected 577,552-parameter geometry denoiser on a 256-icon family-disjoint subsample and overfitted it - held-out loss bottomed out at step 240 and then rose to 11.1485 while training accuracy climbed to 0.4550.

This run changes the train split to the complete 2,681-icon family-disjoint split and nothing else about the learning problem: model, seed 3101, corruption probability 0.35, batch size 16, learning rate 0.001, the 128-icon validation draw, the held-out corruption seeds, and the evaluation cadence are all unchanged. The fixed step budget becomes a 6,300-step cap plus held-out early stopping, because a fixed budget is precisely what made v1 report an overfitted model. That stopping rule is a second change, so v1 is read at its own held-out loss minimum, step 240, rather than at its reported step 600. Both arms are therefore compared under the same selection rule.

The predeclared criteria were committed at a50b2e0 before launch.

Result

Completed natively on the owned RTX 4080 in 75.6 seconds, peak 290.1 MiB CUDA memory. Early stopping ended the run at step 1,320 after eight evaluations without improvement; the selected checkpoint is step 840. A second complete invocation returned an identical JSON result.

Held-out trace

stepheld-out aggregatechangedretainedheld-out loss
0 (untrained)0.00370.00300.004115.3908
600.07000.02280.095410.5741
1200.14750.03280.20919.3253
1800.20020.03830.28718.6219
2400.22530.04080.32438.2993
3000.24660.04200.35648.0193
3600.25660.04450.37067.6764
4200.26270.04690.37867.6276
4800.26920.04840.38787.5239
5400.27350.04930.39407.5360
6000.27850.05190.40027.4659
6600.28230.05320.40547.4575
7200.28240.05320.40557.4163
7800.28440.05420.40807.3897
8400.28860.05730.41287.3333
9000.28770.05590.41227.3702
9600.29020.05840.41467.5195
10200.29200.06000.41677.4133
10800.29120.06150.41467.5393
11400.29500.06170.42037.5728
12000.29580.06370.42047.6322
12600.29610.06460.42057.6349
13200.29990.06620.42547.5275

Held-out totals are 13,978 changed and 26,030 retained fields, identical to v1 because the validation draw and corruption seeds are unchanged. The bold row is the selected checkpoint.

Against the v1 baseline, both read at their own held-out optimum

measurev1 @ step 240v2 @ step 840change
held-out loss9.79327.3333-25.1%
held-out aggregate accuracy0.20910.2886+38.0%
held-out changed-token accuracy0.04390.0573+30.7%
held-out retained-token accuracy0.29780.4128+38.6%
train accuracy, mean of 10 steps to the selected step0.36540.3828+4.8%
train minus held-out aggregate gap0.15630.0942-39.7%

Training accuracy is nearly unchanged while every held-out measure improves, so the generalization gap narrows by 39.7%. That is the signature of a data-scale effect rather than an optimization one.

Predeclared criteria

criterionthresholdobservedoutcome
data_scale_lowers_held_out_loss< 9.79327.3333pass
generalization_gap_narrows< 0.18310.0942pass
held_out_changed_recovery_improves>= 0.065782 (1.5x)0.0573 (1.307x)falsified
later_held_out_optimum> step 240step 840pass
structural_safetylocked-path exact, checkpoint round-tripsboth truepass
reproducibilityidentical rerun JSONidenticalpass

Standing Gate G bar, declared as a bar rather than as a prediction of this run:

criterionthresholdobservedoutcome
retained_preservation>= 0.900.4128falsified, as expected

Overall: partially falsified. Ten times the data buys a large, consistent generalization improvement, but not the predicted 1.5x in changed-token recovery.

The finding this run actually produced

The predeclared 1.5x bar is missed at the selected checkpoint and met at the last trained step. Changed-token accuracy at step 1,320 is 0.0662, which is 1.511x the baseline. The two scalars disagree about when to stop:

The criterion was evaluated exactly as predeclared, at the selected checkpoint, so it is recorded as falsified. Reading it at the final step instead would be choosing the stopping rule after seeing which one passed.

The disagreement itself is the result worth carrying forward. v1 saw a weaker version of it and concluded that held-out loss was the more honest scalar. v2 shows that acting on that conclusion has a measurable cost: loss-based selection gives up 15.6% of the relative changed-token recovery available at the cap. Cross-entropy punishes confident errors while accuracy counts only the argmax, so a model can keep getting more answers right while becoming worse calibrated. Which scalar should govern selection is now an open question in its own right, and it should be settled deliberately - on a criterion declared before the next run - rather than by whichever reading is convenient.

What this does and does not establish

It establishes that at this model size the 256-icon result was data-limited, and that the full split substantially closes the generalization gap. It does not establish that data scale alone reaches useful recovery: at 0.0573 changed-token accuracy the model recovers about one corrupted geometry field in seventeen. Retained-token preservation improves from 0.2978 to 0.4128 but remains far from the 0.90 bar.

The bottleneck has moved. Ten times the data no longer produces a proportional gain in changed-field recovery, which points at model capacity, the corruption schedule, or the single-shot prediction objective rather than at data volume. This remains fixed-topology, geometry-only, single-seed work; it is not evidence about unconditional generation, topology, or style.

State transitions

  1. planned2026-09-20T17:05:00Z
  2. running2026-09-20T17:08:00Z
  3. completed2026-09-20T17:12:40Zfull_split_closes_the_generalization_gap_but_misses_changed_token_bar_and_loss_and_accuracy_disagree_on_stopping

Run record

Verbatim from runs/openmoji-g1-train-v2-datascale-a50b2e0-c474c94d-9b9b1699/run.yaml, the record committed before launch.

schema_version
1
run_id
openmoji-g1-train-v2-datascale-a50b2e0-c474c94d-9b9b1699
state
completed
parent_run
openmoji-g1-train-v1-native-c9bf1b9-f2a06ca5-9b9b1699
hypothesis
The v1 falsification was a data-scale failure, not an optimization or plumbing failure. Training the identical 577,552-parameter denoiser on the full 2,681-icon family-disjoint train split of the dominant P32/T128 bucket, instead of a 256-icon subsample, generalizes materially better on the identical 128-icon held-out draw: a lower held-out loss minimum, a narrower train/held-out gap, and higher changed-token recovery at the selected checkpoint.
expected_information_gain
v1 established that the representation is learnable at real icon diversity but that 256 icons are too few to generalize. This run decides whether 10.5x more data is sufficient at this model size, or whether the next factor is model capacity, augmentation, or the corruption schedule. It is the cheapest test that separates those, because it changes only the data.
scope_limits
Fixed-topology and geometry-only. Not unconditional generation, not topology generation, not a style or AR comparison. A single seed, so small differences are not interpretable across seeds. Group/subgroup conditioning and exact path locks are unchanged from v1.
changed_factors
primary
train_samples 256 -> 2681 (the complete family-disjoint train split)
secondary
The fixed 600-step budget becomes a 6,300-step cap plus held-out-loss early stopping with patience 8 evaluations. This is a stopping rule, not a change to the learning problem. It is disclosed as a second change and controlled for by reading the v1 baseline at its own held-out loss minimum, step 240, rather than at its reported step 600. Both arms are therefore compared under the same selection rule.
held_fixed
model (d_model 96, 4 heads, 2 layers, ff 192, 577,552 parameters), seed 3101, corruption probability 0.35, batch size 16, learning rate 0.001, eval_every 60, validation_samples 128 drawn by the identical hash order, codec, bucket, and the held-out corruption seeds.
step_budget_rationale
6,300 steps x 16 / 2,681 icons is 37.6 epochs, matched to v1's 600 x 16 / 256 = 37.5 epochs, so the cap is not an unmatched advantage. Early stopping is expected to end the run before the cap.
code
git_commit
a50b2e0b515d89758190107b7bbddd8e5a6681f5
execution_mode
native-local
config
path
configs/learning/openmoji-g1-dominant-bucket-train-v2-data-scale.yaml
sha256
c474c94db45be13f…
dataset
hybrid_sha256
9b9b1699677a6f97…
assignments_sha256
e0cdb2a3cc8f00df…
bucket
bucket-p32-t128
bucket_icons
3359
train_samples
2681
validation_samples
128
family_disjoint_train_validation
true
worker
alias
gpubox-4080
gpu
NVIDIA GeForce RTX 4080
driver
595.71.05
torch
2.14.0a0+4fdf77b940.nv26.08
cuda
13.4
execution
seed_set
3101
deterministic_algorithms
true
cublas_workspace_config
:4096:8
optimizer_step_cap
6300
early_stopping
objective
held_out_loss
patience_evals
8
min_delta
0.0
batch_size
16
learning_rate
0.001
corruption_probability
0.35
eval_every
60
resource_cap
max_steps
20000
max_storage_gb
50
network
none
baseline
run
openmoji-g1-train-v1-native-c9bf1b9-f2a06ca5-9b9b1699
read_at
held-out loss minimum, step 240, from its committed validation.jsonl
held_out_loss
9.793182373046875
held_out_aggregate_accuracy
0.20908318336332735
held_out_changed_accuracy
0.043854628702246386
held_out_retained_accuracy
0.29781021897810217
train_token_accuracy_at_step
0.3921783625730994
predeclared_criteria
note
Declared before launch. A disclosed 60-step sizing preflight on the full 2,681-icon split was run first to size wall time and memory. It reported 61.7 s wall including data loading, 290.1 MiB peak CUDA memory, train token accuracy 0.0793, held-out loss 10.5741, and held-out changed-token accuracy 0.0228 at step 60. The corresponding v1 step-60 numbers are held-out loss 10.7155 and changed 0.0219, so the preflight distinguishes nothing about the outcome and is disclosed rather than hidden. The full run was not observed beforehand.
evaluated_at
the checkpoint selected by held-out loss, not the last trained step
primary
id
data_scale_lowers_held_out_loss
statement
The minimum held-out loss is below 9.7932, the v1 baseline minimum on the identical validation draw and corruption seeds.
id
generalization_gap_narrows
statement
At the selected step, train token accuracy minus held-out aggregate accuracy is below 0.1831, the v1 value at its own selected step. Train token accuracy is a 16-sample per-step figure, so the mean over the ten training steps ending at the selected step is used for both arms.
id
held_out_changed_recovery_improves
statement
Held-out changed-token accuracy is at least 1.5x the v1 baseline value of 0.043855, that is at least 0.065782.
id
later_held_out_optimum
statement
The selected step is later than v1's step 240, so more data supports longer useful training rather than merely reaching the same optimum sooner.
id
structural_safety
statement
Locked-path exactness holds and the canonical checkpoint round-trips at the selected step.
id
reproducibility
statement
A second complete invocation returns an identical JSON result.
standing_gate_not_a_prediction
id
retained_preservation
statement
Held-out retained-token accuracy is at least 0.90.
note
Carried forward from v1, where it was falsified at 0.3256. It is recorded so the Gate G bar is not quietly dropped, but this run is not predicted to meet it: 10.5x data at an unchanged 577,552 parameters is not expected to close a 0.33-to-0.90 gap. If it is met, that is a stronger result than the hypothesis.
reported_not_gated
the complete held-out trace and the step at which loss turns upward, if it does, held-out aggregate accuracy and retained-token accuracy at the selected step, the last-trained-step held-out numbers, for the overfitting comparison, wall time, peak CUDA memory, and whether early stopping or the cap ended the run
chance_level
Random legal-token choice is about 0.0035 for endpoints over 289 bins and 0.0024 for controls over 417 bins.
outputs
local_metadata
runs/openmoji-g1-train-v2-datascale-a50b2e0-c474c94d-9b9b1699
report_root
reports/learning/openmoji-g1-dominant-bucket-train-v2-data-scale
durable_artifacts
/home/dev/.cache/openmoji-g1-train-v2-datascale-a50b2e0-c474c94d-9b9b1699
planned_at
2026-09-20 17:05:00+00:00
completed_at
2026-09-20 17:12:40+00:00
result
device
cuda
wall_seconds
75.6
peak_cuda_mib
290.1
ended_by
early_stopping
completed_steps
1320
selected_step
840
selected_held_out_loss
7.3332648277282715
selected_held_out_aggregate_accuracy
0.28861727654469105
selected_held_out_changed_accuracy
0.057304335384175134
selected_held_out_retained_accuracy
0.41283134844410296
final_step_held_out_loss
7.527534484863281
final_step_held_out_changed_accuracy
0.06624695950779796
train_token_accuracy_mean10_at_selected_step
0.3828216374269006
generalization_gap_at_selected_step
0.09420436088220956
checkpoint_sha256
e791493769907ac4…
checkpoint_bytes
7075311
identical_rerun
true
predeclared_outcome
overall
partially-falsified
data_scale_lowers_held_out_loss
passed
true
observed
7.3332648277282715
threshold
9.793182373046875
generalization_gap_narrows
passed
true
observed
0.09420436088220956
threshold
0.1831
held_out_changed_recovery_improves
passed
false
observed
0.057304335384175134
observed_ratio
1.3067
threshold
0.065782
threshold_ratio
1.5
note
Met at the last trained step, where changed accuracy is 0.06624695950779796, a 1.511x ratio. It is recorded as falsified because the criterion was predeclared at the selected checkpoint, and choosing the other reading after seeing the numbers would be selecting the stopping rule to pass the test.
later_held_out_optimum
passed
true
observed
840
threshold
240
structural_safety
passed
true
reproducibility
passed
true
retained_preservation
passed
false
observed
0.41283134844410296
threshold
0.9
note
Standing Gate G bar, not a prediction of this run. Improved from v1's 0.3256.
conclusion
Ten times the data lowers held-out loss 25.1% and narrows the train/held-out gap 39.7% with training accuracy nearly unchanged, so the v1 failure was data-limited. It does not deliver the predicted 1.5x changed-token recovery, so data volume is no longer the binding constraint. Held-out loss and held-out accuracy disagree about when to stop, and that disagreement is now its own open question.
artifact_durability
durable_root
/home/dev/.cache/openmoji-g1-train-v2-datascale-a50b2e0-c474c94d-9b9b1699
checkpoint_bytes
7075311
copied_to_repo
summary.json, metrics.jsonl, validation.jsonl