MojiDiff

← experiments

openmoji-g1-capacity-v3-3c252d5-ee32665b-9b9b1699

completed falsified

Hypothesis

The bottleneck v2 exposed is model capacity. At 2,040,976 parameters, 3.53x v2's 577,552, the same denoiser on the same 2,681-icon split recovers materially more corrupted geometry and renders measurably closer to `x_0` than v2's selected checkpoint.

Why run it

v2 showed that 10.5x the data narrows the train/held-out gap 39.7% but lifts changed-token recovery only 1.307x, and leaves that gap at 0.0942 - the model is close to fitting what it can express. If capacity is the constraint, this is the cheapest test that shows it. If it is not, the remaining candidate is the corruption schedule, and that is a different experiment.

Scope limits

Fixed-topology and geometry-only, single seed, single corruption probability of 0.35. Not unconditional generation, not topology generation, not an AR comparison. Capacity is varied by width and depth together, which is how a transformer's capacity is conventionally expressed; head dimension stays at 24 and the feedforward ratio at 2x d_model, so no other shape parameter moves.

Measured behaviour

Held-out loss
lossoptimizer step57.51012.51517.50200400600800selected
Table view
optimizer stepheld-out loss
015.7289
609.4942
1208.1581
1807.6501
2407.5938
3007.5106
3607.3790
4207.6706
4807.5067
5407.5109
6007.6677
6607.5621
7207.6622
7807.8262
8407.6919
Held-out token accuracy
aggregatechanged fieldsretained fields
accuracyoptimizer step00.20.40.60200400600800selected
Table view
optimizer stepaggregatechanged fieldsretained fields
00.00350.00340.0036
600.16080.03410.2288
1200.24850.04320.3587
1800.27520.04740.3975
2400.28720.05440.4123
3000.29820.06070.4257
3600.30360.06390.4323
4200.30550.06800.4330
4800.30540.06800.4330
5400.31350.07150.4434
6000.31140.07550.4381
6600.31450.07710.4420
7200.31640.07720.4449
7800.31390.08230.4382
8400.31710.08310.4427
Training token accuracy (per batch)
accuracyoptimizer step00.20.40.60.8200400600800selected
Table view
optimizer steptrain accuracy
10.0037
20.0032
30.0063
40.0051
50.0041
60.0053
70.0077
80.0134
90.0188
100.0150
110.0151
120.0162
130.0221
140.0256
150.0250
160.0242
170.0300
180.0449
190.0458
200.0558
210.0400
220.0496
230.0697
240.0445
250.0658
260.0862
270.0645
280.0990
290.0910
300.0859
310.0908
320.0997
330.0545
340.1537
350.1182
360.1246
370.0891
380.1238
390.1201
400.0915
410.1686
420.1451
430.0789
440.1554
450.1427
460.1073
470.1418
480.1624
490.1752
500.1264
510.1801
520.1698
530.1361
540.1709
550.2176
560.1152
570.1979
580.1936
590.2333
600.2102
610.2010
620.1945
630.2769
640.2040
650.2515
660.2403
670.1712
680.2705
690.1962
700.2360
710.2840
720.2780
730.2641
740.2756
750.2093
760.2451
770.2145
780.2701
790.2149
800.2335
810.2521
820.2719
830.2919
840.2532
850.3247
860.3129
870.2800
880.2520
890.3027
900.3020
910.3277
920.3025
930.2049
940.2388
950.2462
960.3042
970.2816
980.2418
990.2218
1000.3345
1010.3576
1020.3190
1030.2889
1040.3291
1050.3392
1060.2736
1070.2999
1080.3371
1090.3766
1100.2691
1110.3684
1120.2824
1130.2540
1140.3018
1150.3160
1160.2605
1170.3006
1180.2775
1190.3629
1200.2771
1210.2590
1220.3737
1230.3036
1240.2359
1250.3455
1260.2482
1270.2645
1280.4025
1290.3605
1300.4400
1310.2067
1320.3287
1330.3389
1340.3273
1350.2947
1360.4002
1370.2989
1380.3956
1390.2578
1400.3369
1410.3646
1420.3534
1430.3570
1440.2963
1450.2999
1460.3299
1470.3508
1480.3685
1490.3015
1500.3910
1510.3417
1520.3117
1530.3640
1540.2998
1550.3468
1560.3177
1570.2755
1580.3295
1590.3385
1600.3035
1610.2934
1620.3317
1630.3106
1640.3555
1650.3835
1660.2503
1670.3151
1680.3908
1690.2965
1700.3767
1710.4167
1720.3882
1730.2933
1740.2783
1750.3979
1760.3439
1770.4214
1780.3041
1790.2861
1800.4453
1810.3571
1820.4332
1830.4214
1840.3475
1850.3912
1860.3800
1870.4837
1880.3762
1890.3443
1900.3676
1910.3880
1920.3781
1930.3913
1940.3427
1950.4090
1960.4061
1970.3807
1980.3951
1990.4136
2000.3852
2010.2876
2020.4524
2030.3877
2040.4111
2050.3915
2060.3645
2070.3634
2080.2999
2090.4407
2100.4163
2110.3192
2120.4703
2130.3099
2140.3488
2150.3180
2160.4324
2170.3626
2180.2749
2190.4397
2200.3219
2210.3592
2220.3922
2230.3956
2240.3623
2250.2986
2260.4033
2270.4238
2280.3874
2290.4119
2300.4071
2310.3677
2320.3877
2330.4119
2340.3992
2350.2808
2360.4866
2370.3289
2380.4038
2390.5159
2400.3718
2410.4662
2420.3904
2430.3793
2440.3608
2450.3662
2460.3854
2470.3868
2480.3646
2490.3044
2500.4739
2510.4213
2520.4184
2530.4828
2540.3483
2550.3751
2560.4185
2570.4338
2580.3911
2590.5248
2600.4165
2610.2986
2620.3301
2630.4484
2640.3717
2650.3950
2660.3830
2670.4038
2680.3804
2690.4403
2700.4358
2710.3788
2720.4740
2730.3982
2740.3729
2750.3600
2760.4817
2770.4458
2780.3197
2790.5376
2800.3405
2810.3596
2820.3812
2830.4381
2840.3668
2850.3755
2860.3680
2870.4438
2880.3784
2890.4125
2900.4399
2910.3390
2920.3885
2930.3333
2940.3605
2950.3927
2960.4104
2970.4900
2980.4578
2990.3596
3000.3989
3010.4527
3020.3552
3030.4370
3040.4397
3050.4274
3060.3626
3070.3466
3080.4455
3090.4165
3100.5210
3110.3505
3120.3717
3130.3674
3140.4452
3150.4275
3160.4131
3170.3434
3180.5086
3190.4086
3200.3885
3210.4264
3220.3282
3230.4495
3240.3253
3250.3635
3260.4133
3270.4217
3280.3486
3290.3825
3300.3755
3310.3882
3320.4424
3330.4218
3340.3074
3350.4699
3360.3756
3370.4342
3380.3859
3390.5265
3400.3511
3410.3368
3420.3570
3430.4566
3440.4681
3450.4200
3460.3564
3470.4082
3480.4189
3490.4910
3500.4747
3510.4685
3520.4345
3530.4168
3540.5462
3550.5055
3560.3814
3570.4206
3580.4460
3590.3836
3600.4767
3610.4495
3620.3473
3630.4974
3640.4771
3650.4583
3660.4335
3670.4557
3680.3739
3690.4975
3700.4606
3710.4924
3720.3668
3730.4945
3740.3465
3750.4345
3760.4214
3770.5024
3780.3534
3790.4785
3800.4620
3810.3704
3820.3942
3830.4280
3840.4784
3850.3641
3860.4240
3870.4153
3880.3530
3890.4485
3900.4797
3910.3513
3920.4390
3930.3834
3940.5055
3950.4467
3960.4696
3970.4110
3980.4933
3990.3750
4000.5342
4010.4904
4020.3648
4030.4702
4040.4400
4050.4528
4060.4943
4070.5281
4080.4494
4090.4910
4100.3947
4110.4134
4120.4192
4130.4591
4140.3875
4150.3947
4160.4073
4170.4371
4180.4723
4190.4316
4200.5193
4210.5331
4220.4348
4230.4215
4240.4225
4250.4673
4260.5611
4270.5053
4280.3555
4290.3168
4300.4310
4310.4380
4320.4987
4330.3755
4340.3922
4350.4540
4360.4768
4370.5131
4380.4015
4390.4842
4400.4864
4410.4078
4420.4116
4430.4732
4440.5035
4450.4300
4460.4864
4470.4217
4480.3927
4490.4207
4500.4194
4510.4188
4520.4211
4530.4102
4540.4370
4550.4611
4560.3951
4570.4868
4580.4179
4590.3900
4600.4484
4610.3834
4620.3468
4630.4711
4640.4971
4650.5492
4660.3721
4670.4312
4680.4680
4690.4317
4700.4086
4710.5169
4720.4234
4730.4969
4740.3707
4750.4497
4760.4683
4770.5020
4780.4856
4790.3419
4800.4414
4810.4113
4820.4393
4830.4799
4840.4446
4850.4504
4860.4725
4870.4398
4880.4446
4890.4020
4900.3920
4910.4773
4920.3636
4930.4386
4940.4097
4950.4312
4960.3358
4970.4652
4980.4041
4990.4626
5000.5011
5010.3816
5020.3872
5030.5114
5040.3880
5050.4722
5060.4916
5070.4951
5080.3537
5090.3436
5100.4988
5110.4126
5120.5384
5130.3976
5140.3645
5150.5738
5160.4307
5170.5517
5180.4943
5190.4774
5200.4693
5210.4559
5220.5820
5230.4464
5240.4178
5250.4125
5260.5207
5270.4733
5280.4797
5290.4167
5300.4543
5310.5403
5320.4666
5330.5034
5340.4976
5350.4710
5360.3435
5370.5408
5380.4834
5390.4858
5400.4597
5410.4754
5420.4319
5430.3661
5440.5344
5450.5035
5460.4161
5470.5250
5480.3516
5490.4962
5500.3720
5510.5000
5520.4643
5530.3475
5540.5096
5550.3589
5560.3970
5570.4958
5580.4772
5590.4099
5600.4071
5610.5002
5620.5016
5630.4052
5640.5180
5650.4759
5660.4654
5670.4800
5680.5128
5690.4928
5700.3371
5710.5598
5720.4215
5730.4811
5740.5968
5750.4568
5760.5187
5770.4881
5780.4030
5790.4819
5800.4009
5810.5134
5820.4477
5830.4415
5840.3855
5850.5300
5860.4797
5870.5057
5880.5694
5890.4497
5900.4544
5910.4505
5920.4647
5930.5018
5940.5915
5950.4970
5960.3564
5970.3748
5980.5192
5990.4562
6000.4604
6010.4583
6020.3607
6030.5175
6040.5620
6050.5099
6060.4757
6070.4811
6080.5245
6090.4224
6100.4356
6110.5547
6120.5381
6130.3631
6140.6044
6150.4190
6160.4010
6170.4574
6180.5010
6190.4524
6200.4380
6210.4038
6220.5213
6230.4150
6240.4749
6250.4599
6260.4881
6270.4258
6280.3966
6290.4070
6300.3795
6310.5142
6320.5173
6330.5532
6340.3684
6350.5145
6360.4921
6370.3976
6380.4788
6390.5241
6400.4127
6410.4901
6420.3913
6430.4767
6440.5078
6450.5392
6460.4628
6470.4017
6480.4276
6490.4950
6500.4634
6510.5171
6520.4072
6530.5601
6540.4688
6550.4238
6560.4785
6570.3868
6580.4905
6590.4019
6600.4052
6610.4562
6620.4741
6630.3975
6640.4217
6650.4410
6660.4450
6670.5217
6680.4623
6690.3580
6700.5195
6710.4273
6720.5134
6730.4354
6740.6189
6750.4252
6760.3846
6770.3929
6780.5151
6790.5417
6800.4925
6810.3900
6820.4842
6830.4654
6840.5634
6850.5532
6860.5174
6870.4725
6880.4920
6890.6052
6900.5702
6910.4147
6920.4535
6930.4840
6940.4549
6950.4996
6960.4721
6970.4505
6980.5154
6990.5315
7000.4537
7010.5122
7020.4809
7030.4354
7040.5352
7050.4696
7060.5375
7070.4331
7080.5309
7090.4538
7100.4579
7110.4320
7120.5781
7130.4004
7140.5269
7150.5125
7160.3810
7170.4713
7180.4335
7190.5313
7200.4415
7210.4301
7220.4989
7230.3569
7240.4709
7250.5057
7260.4918
7270.4707
7280.4630
7290.5352
7300.4982
7310.5251
7320.4624
7330.5419
7340.3976
7350.5752
7360.5421
7370.4555
7380.4659
7390.5264
7400.4539
7410.5429
7420.6212
7430.4332
7440.5843
7450.4079
7460.4626
7470.5047
7480.4912
7490.4667
7500.4406
7510.4523
7520.4934
7530.5376
7540.4261
7550.5548
7560.5580
7570.5339
7580.4289
7590.5172
7600.5052
7610.5712
7620.5828
7630.4507
7640.3651
7650.4766
7660.4866
7670.5085
7680.4697
7690.4347
7700.4811
7710.5243
7720.5184
7730.5010
7740.5419
7750.4917
7760.4898
7770.5056
7780.5110
7790.5421
7800.5127
7810.4814
7820.4571
7830.4584
7840.4932
7850.4088
7860.5097
7870.4308
7880.5218
7890.4475
7900.4841
7910.4551
7920.4793
7930.4982
7940.4580
7950.4608
7960.4342
7970.3883
7980.5017
7990.5129
8000.5613
8010.4758
8020.4565
8030.5052
8040.4845
8050.4280
8060.5919
8070.4469
8080.5021
8090.4343
8100.4142
8110.5665
8120.5356
8130.5517
8140.3778
8150.4887
8160.4532
8170.4860
8180.4795
8190.4893
8200.4935
8210.5213
8220.4939
8230.5032
8240.4246
8250.4066
8260.5338
8270.4214
8280.4819
8290.4615
8300.4421
8310.3944
8320.4921
8330.3974
8340.4979
8350.5638
8360.4217
8370.3922
8380.5725
8390.4482
8400.5154

Measured result

Verbatim from the run's summary.json.

bucket
bucket-p32-t128
bucket_icons
3359
checkpoint_bytes
24704555
checkpoint_round_trip
true
checkpoint_sha256
93344f19969eb06e…
config_sha256
ee32665be9e3fda5…
corruption
factorized_role_uniform_geometry
cuda_version
13.4
deterministic_algorithms
true
device
cuda
eval_every
60
final_train
loss
3.004683017730713
step
840
train_token_accuracy
0.5154013015184382
group_vocabulary_size
12
locked_path_exact
true
metrics_sha256
ae7febe80c0f1f33…
model_parameters
2040976
schema_version
1
scope
dominant-bucket fixed-topology geometry pilot; not unconditional generation
selected_train_rows
color/svg/1F3F4-E0064-E0065-E0062-E0065-E007F.svg, color/svg/1F468-1F3FC-200D-1F9BC.svg, color/svg/1F561.svg, color/svg/1F469-1F3FE-200D-1F9BC.svg, color/svg/1F469-1F3FD-200D-2764-FE0F-200D-1F48B-200D-1F469-1F3FE.svg, color/svg/1F6B5-1F3FB-200D-2642-FE0F.svg … and 2675 more
selected_validation_rows
color/svg/1F994.svg, color/svg/E30A.svg, color/svg/1F6BE.svg, color/svg/1F3CA-1F3FB-200D-2642-FE0F.svg, color/svg/1F3CB-1F3FC-200D-2642-FE0F.svg, color/svg/1F93D-1F3FC.svg … and 122 more
selection
completed_steps
840
evals_without_improvement
8
min_delta
0.0
objective
held_out_loss
patience_evals
8
selected_held_out_loss
7.379046440124512
selected_step
360
stopped_early
true
steps
6300
study_version
openmoji-g1-capacity-v3
subgroup_vocabulary_size
118
torch_version
2.14.0a0+4fdf77b940.nv26.08
validation
accuracy
0.3035642871425715
changed_accuracy
0.06388610673916154
changed_total
13978
loss
7.379046440124512
retained_accuracy
0.43227045716480983
retained_total
26030
validation_final_step
accuracy
0.3170865826834633
changed_accuracy
0.08313063385319788
changed_total
13978
loss
7.691944599151611
retained_accuracy
0.44271993853246255
retained_total
26030
validation_trace_sha256
5e08252553b9f432…
validation_untrained
accuracy
0.0034993001399720057
changed_accuracy
0.0033624266704821862
changed_total
13978
loss
15.728937149047852
retained_accuracy
0.0035728006146753745
retained_total
26030

Written result

v2 established that data volume is no longer the binding constraint — 10.5x the data narrowed the train/held-out gap 39.7% but lifted changed-token recovery only 1.307x, leaving that gap at 0.0942 and suggesting the 577,552-parameter denoiser was close to fitting what it could express. This run tests that suggestion directly.

Capacity is the single factor. Width and depth scale together, which is how a transformer's capacity is conventionally expressed, with head dimension held at 24 and the feedforward ratio at 2x d_model so no other shape parameter moves: 577,552 -> 2,040,976 parameters, 3.53x. Split, seed, corruption probability, batch size, learning rate, step cap, evaluation cadence, and the held-out-loss selection policy are all unchanged.

The selection scalar was declared before launch, which matters here: as in v2, held-out accuracy keeps climbing well past the loss minimum, and at the final step changed-token recovery reaches 0.0831 — 1.45x the baseline, which would have passed the criterion. The settled rule says read at the loss-selected checkpoint, and it was settled before this run existed.

Result

Completed natively on the RTX 4080 in 72.1 seconds at 342.5 MiB peak CUDA memory. Early stopping ended it at step 840; the selected checkpoint is step 360. A second complete invocation returned an identical JSON result.

Held-out trace

stepaggregatechangedretainedheld-out loss
0 (untrained)0.00350.00340.003615.7289
600.16080.03410.22889.4942
1200.24850.04320.35878.1581
1800.27520.04740.39757.6501
2400.28720.05440.41237.5938
3000.29820.06070.42577.5106
3600.30360.06390.43237.3790
4200.30550.06800.43307.6706
4800.30540.06800.43307.5067
5400.31350.07150.44347.5109
6000.31140.07550.43817.6677
6600.31450.07710.44207.5621
7200.31640.07720.44497.6622
7800.31390.08230.43827.8262
8400.31710.08310.44277.6919

Against v2, both read at their own held-out loss minimum

v2, 577,552 paramsv3, 2,040,976 params
selected step8403602.3x sooner
held-out loss7.33337.3790+0.6% worse
aggregate accuracy0.28860.30361.052x
changed-token accuracy0.05730.06391.115x
retained-token accuracy0.41280.43231.047x
median render RGBA MAE, 72 px, 128 icons0.1412150.136983
mean paired render difference, 72 px—+0.00185395% CI -0.0026 .. +0.0063

Predeclared criteria

criterionthresholdobservedoutcome
capacity_lowers_held_out_loss< 7.33337.3790falsified
changed_recovery_improves>= 0.071630 (1.25x)0.063886 (1.115x)falsified
renders_measurably_closerCI on paired difference excludes zero-0.0026 .. +0.0063falsified
structural_safetylocked-path exact, checkpoint round-tripsboth truepass
reproducibilityidentical rerun JSONidenticalpass

Standing Gate G bar, not a prediction of this run:

criterionthresholdobservedoutcome
retained_preservation>= 0.900.4323falsified, as expected

Overall: falsified. Capacity is not the binding constraint.

What the shape of the failure says

3.53x the parameters did not reach a better held-out loss. It reached essentially the same floor — 7.3790 against 7.3333, within 0.6% — and reached it in 360 steps instead of 840, then overfitted from there. That is the signature of a problem that is limited by the task rather than by the model: more capacity buys faster fitting of the same ceiling, not a lower one.

The render check agrees and is the more important of the two, because it is powered. The selection-scalar comparison measured the paired-difference interval's half-width at 0.0032 on this exact 128-icon draw, so an effect above roughly 0.0064 would have been detected. The observed effect is +0.001853 with an interval spanning zero. A 3.53x model does not render measurably closer to x_0.

The accuracy-versus-loss divergence is also sharper here than in v2: changed-token recovery climbs from 0.0639 at the loss minimum to 0.0831 by step 840, a 30% relative gain entirely on the far side of the point where held-out loss stopped improving. Because the selection rule was settled beforehand, that number is reported and not gated — which is exactly the post-hoc freedom the scalar comparison was run to remove.

Where this leaves Gate G

Two of the three candidate factors are now eliminated on matched, predeclared comparisons. Data volume is not the constraint: v2. Model capacity is not the constraint: this run. The remaining candidate is the one the render probe pointed at — the corruption schedule.

That probe showed x_t at corruption probability 0.35 is already visually destroyed: a third of the geometry is simply gone, and no amount of model or data recovers information that is not there. The probability has been fixed at 0.35 since the Gate F four-icon fixtures, chosen for a four-icon experiment and never revisited at corpus scale. The held-out loss floor near 7.35 that both v2 and v3 converge to looks like a property of that regime rather than of either model.

The next experiment should vary it — ideally training across a range of corruption levels rather than a single fixed point, which is also what a denoiser facing many levels at sampling time would need.

State transitions

  1. planned2026-09-20T18:30:00Z
  2. completed2026-09-20T18:45:00Zcapacity_is_not_the_binding_constraint_same_loss_floor_reached_sooner_then_overfits

Run record

Verbatim from runs/openmoji-g1-capacity-v3-3c252d5-ee32665b-9b9b1699/run.yaml, the record committed before launch.

schema_version
1
run_id
openmoji-g1-capacity-v3-3c252d5-ee32665b-9b9b1699
state
completed
parent_run
openmoji-g1-train-v2-datascale-a50b2e0-c474c94d-9b9b1699
hypothesis
The bottleneck v2 exposed is model capacity. At 2,040,976 parameters, 3.53x v2's 577,552, the same denoiser on the same 2,681-icon split recovers materially more corrupted geometry and renders measurably closer to `x_0` than v2's selected checkpoint.
expected_information_gain
v2 showed that 10.5x the data narrows the train/held-out gap 39.7% but lifts changed-token recovery only 1.307x, and leaves that gap at 0.0942 - the model is close to fitting what it can express. If capacity is the constraint, this is the cheapest test that shows it. If it is not, the remaining candidate is the corruption schedule, and that is a different experiment.
scope_limits
Fixed-topology and geometry-only, single seed, single corruption probability of 0.35. Not unconditional generation, not topology generation, not an AR comparison. Capacity is varied by width and depth together, which is how a transformer's capacity is conventionally expressed; head dimension stays at 24 and the feedforward ratio at 2x d_model, so no other shape parameter moves.
changed_factors
primary
d_model 96->192, heads 4->8, layers 2->4, feedforward 192->384. 577,552 -> 2,040,976 parameters, 3.53x.
held_fixed
the full 2,681-icon family-disjoint train split, the 128-icon validation draw and its corruption seeds, seed 3101, corruption probability 0.35, batch size 16, learning rate 0.001, eval_every 60, the 6,300-step cap, the held-out-loss selection policy with patience 8, the codec, and the bucket.
selection_scalar
Held-out loss, the project's settled rule. Adopted by convention after the 128-icon comparison bounded the render-quality difference between loss-selected and accuracy-selected checkpoints below 0.0032 RGBA MAE. Declared here so no reading is chosen after the fact.
code
git_commit
3c252d5
execution_mode
native-local
config
path
configs/learning/openmoji-g1-capacity-v3.yaml
sha256
ee32665be9e3fda5…
dataset
hybrid_sha256
9b9b1699677a6f97…
bucket
bucket-p32-t128
train_samples
2681
validation_samples
128
baseline
run
openmoji-g1-train-v2-datascale-a50b2e0-c474c94d-9b9b1699
read_at
its held-out loss minimum, step 840, the same selection rule this run uses
held_out_loss
7.3332648277282715
held_out_aggregate_accuracy
0.28861727654469105
held_out_changed_accuracy
0.057304335384175134
held_out_retained_accuracy
0.41283134844410296
rendered_median_rgba_mae_72px_128_icons
0.141215
worker
alias
gpubox-4080
gpu
NVIDIA GeForce RTX 4080
driver
595.71.05
torch
2.14.0a0+4fdf77b940.nv26.08
cuda
13.4
execution
seed_set
3101
deterministic_algorithms
true
cublas_workspace_config
:4096:8
optimizer_step_cap
6300
early_stopping
objective
held_out_loss
patience_evals
8
min_delta
0.0
resource_cap
max_steps
20000
max_storage_gb
50
network
none
predeclared_criteria
note
Declared before launch. No preflight was run; v2's wall time of 75.6 s at 290.1 MiB peak CUDA memory for a 3.53x smaller model is ample evidence that this fits the bounds.
evaluated_at
the checkpoint selected by held-out loss
primary
id
capacity_lowers_held_out_loss
statement
Held-out loss at the selected step is below v2's 7.3332648277282715.
id
changed_recovery_improves
statement
Held-out changed-token accuracy is at least 1.25x v2's 0.057304335384175134, that is at least 0.071630.
id
renders_measurably_closer
statement
On the complete 128-icon held-out draw, with identical corruption seeds, the mean per-icon RGBA MAE of `x_hat_0` against `x_0` at 72 px is lower than v2's selected checkpoint and the 95% confidence interval on that paired difference excludes zero. The selection-scalar comparison measured this interval's half-width at 0.0032 on the same draw, so an effect above roughly 0.0064 is detectable and this is a powered test rather than a hopeful one.
id
structural_safety
statement
Locked-path exactness holds and the canonical checkpoint round-trips.
id
reproducibility
statement
A second complete invocation returns an identical JSON result.
standing_gate_not_a_prediction
id
retained_preservation
statement
Held-out retained-token accuracy is at least 0.90.
note
Carried from v1 and v2, where it read 0.3256 and 0.4128. Recorded so the bar is not quietly dropped; 3.53x parameters is not predicted to close it.
falsification_meaning
If capacity does not move these numbers, the model was not the binding constraint either, and the corruption schedule becomes the leading candidate: the render probe showed x_t at probability 0.35 is already visually destroyed, so the training distribution may simply sit in an unachievable regime. That outcome redirects the next experiment and must not be tuned away.
reported_not_gated
the complete held-out trace, the selected step, and whether the cap or early stopping ended it, held-out aggregate and retained accuracy at the selected step, wall time, peak CUDA memory, and parameter count
outputs
local_metadata
runs/openmoji-g1-capacity-v3-3c252d5-ee32665b-9b9b1699
report_root
reports/learning/openmoji-g1-capacity-v3
durable_artifacts
/home/dev/.cache/openmoji-g1-capacity-v3-3c252d5-ee32665b-9b9b1699
planned_at
2026-09-20 18:30:00+00:00
completed_at
2026-09-20 18:45:00+00:00
result
device
cuda
wall_seconds
72.1
peak_cuda_mib
342.5
model_parameters
2040976
ended_by
early_stopping
completed_steps
840
selected_step
360
selected_held_out_loss
7.379046440124512
selected_held_out_aggregate_accuracy
0.3035642871425715
selected_held_out_changed_accuracy
0.06388610673916154
selected_held_out_retained_accuracy
0.43227045716480983
final_step_held_out_changed_accuracy
0.08313063385319788
checkpoint_sha256
93344f19969eb06e…
identical_rerun
true
render_128_icons
control_inputs_identical
true
median_rgba_mae_72px
0.136983
baseline_median_rgba_mae_72px
0.141215
mean_paired_difference_72px
0.001853
ci95_72px
-0.002643, 0.006349
v3_better_on_72px
68
of
128
mean_paired_difference_18px
0.001758
ci95_18px
-0.003065, 0.006581
predeclared_outcome
overall
falsified
capacity_lowers_held_out_loss
passed
false
observed
7.379046440124512
threshold
7.3332648277282715
changed_recovery_improves
passed
false
observed
0.06388610673916154
observed_ratio
1.115
threshold
0.07163
threshold_ratio
1.25
renders_measurably_closer
passed
false
observed_mean_paired_difference
0.001853
ci95
-0.002643, 0.006349
detectable_above
0.0064
note
Powered against a half-width of 0.0032 measured on this exact 128-icon draw by the selection-scalar comparison, so an effect above roughly 0.0064 would have been detected. The interval spans zero.
structural_safety
passed
true
reproducibility
passed
true
retained_preservation
passed
false
observed
0.43227045716480983
threshold
0.9
note
standing bar
not a prediction
None
conclusion
3.53x the parameters did not reach a lower held-out loss. It reached the same floor within 0.6%, in 360 steps instead of 840, then overfitted - the signature of a task-limited rather than capacity-limited problem. The powered render check agrees. Data volume was eliminated by v2 and model capacity by this run, both on matched predeclared comparisons, leaving the corruption schedule as the remaining candidate, exactly as this run's predeclared falsification_meaning anticipated.
artifact_durability
durable_root
/home/dev/.cache/openmoji-g1-capacity-v3-3c252d5-ee32665b-9b9b1699
checkpoint_sha256
93344f19969eb06e…
copied_to_repo
summary.json, metrics.jsonl, validation.jsonl, render-metrics-128.jsonl