MojiDiff

← experiments

openmoji-g1-full-denoiser-v11-27e9956-b9343dcc-9b9b1699

completed falsified

Hypothesis

Every trained model in this project scores below the identity policy - emit the corrupted input unchanged - on held-out token accuracy, the best by 0.27. Four verified defects have since been fixed: the loss averaged over groups so four fields carried 31% of it, no attention padding mask so 29.4% of the sequence was attended as content, corruption drew uniformly so a token-marginal detector scored 2.84 with no context, and coordinates were encoded as unordered categories so no difference was computable. With all four fixed, the detector now beats every free reference at 2.781 against a continuity statistic's 1.893. Running the full objective - the value head training alongside the keep head, with the keep decision gating the decode - should for the first time clear the identity baseline.

Why run it

This is the bar the whole Gate G sequence has failed, and the first time it is attempted with a working detector and a sound training setup. Clearing it would mean the project has a denoiser that does something rather than nothing. Failing it with a detector this good would localise the remaining problem to reconstruction - knowing WHICH field is wrong but not WHAT belongs there - which is a different and narrower question than any asked so far.

Scope limits

Fixed-topology and geometry-only, single seed, marginal-respecting corruption at probability 0.35. Beating identity on token accuracy is not the same as producing a recognizable icon; the render probe is run separately and no run in this project has produced one yet.

Measured behaviour

Held-out loss
lossoptimizer step567805001000selected
Table view
optimizer stepheld-out loss
07.6746
606.2027
1206.0346
1805.8606
2405.7455
3005.6291
3605.5662
4205.5151
4805.4604
5405.4370
6005.4246
6605.3908
7205.3878
7805.4039
8405.3671
9005.3714
9605.3884
10205.3701
10805.4055
11405.3922
12005.4039
12605.4143
13205.4454
Held-out token accuracy
aggregatechanged fieldsretained fields
accuracyoptimizer step00.250.50.75105001000selected
Table view
optimizer stepaggregatechanged fieldsretained fields
00.33410.00130.5121
600.61810.00170.9478
1200.64310.00140.9863
1800.64680.00160.9919
2400.64550.00310.9891
3000.64640.00220.9909
3600.63680.00450.9749
4200.62450.00930.9535
4800.60390.01590.9184
5400.59760.01930.9069
6000.59220.02140.8975
6600.59500.02550.8995
7200.60290.02370.9127
7800.58980.03030.8891
8400.58390.02960.8803
9000.58980.02490.8919
9600.57940.03060.8729
10200.59120.02760.8925
10800.59010.02920.8901
11400.57890.03660.8689
12000.59430.02540.8986
12600.58420.03590.8775
13200.58830.03170.8859
Training token accuracy (per batch)
accuracyoptimizer step00.20.40.60.85001000selected
Table view
optimizer steptrain accuracy
10.3187
20.3767
30.4567
40.4684
50.4901
60.5379
70.5032
80.5309
90.5419
100.5132
110.5139
120.5011
130.4852
140.4928
150.4726
160.5141
170.5191
180.5397
190.5243
200.5518
210.5382
220.5601
230.5669
240.5921
250.5665
260.5887
270.5647
280.5916
290.5734
300.5644
310.5789
320.5545
330.5498
340.5504
350.5496
360.5736
370.5603
380.5843
390.5865
400.6127
410.5921
420.6126
430.5693
440.6162
450.6154
460.5983
470.5912
480.5951
490.5903
500.6072
510.6098
520.6124
530.6093
540.6112
550.6036
560.6096
570.6185
580.6128
590.6200
600.6050
610.6119
620.6012
630.6120
640.6287
650.6406
660.6376
670.6051
680.6361
690.6126
700.6307
710.6375
720.6258
730.6404
740.6195
750.6437
760.6276
770.6367
780.6286
790.5942
800.6393
810.6245
820.6313
830.6127
840.6241
850.6297
860.6111
870.6224
880.6134
890.6366
900.6356
910.6357
920.6329
930.5993
940.6454
950.6437
960.6291
970.6278
980.6501
990.6432
1000.6363
1010.6469
1020.6393
1030.6457
1040.6416
1050.6608
1060.6397
1070.6355
1080.6298
1090.6572
1100.6424
1110.6433
1120.6389
1130.6360
1140.6410
1150.6438
1160.6448
1170.6338
1180.6430
1190.6349
1200.6531
1210.6433
1220.6591
1230.6368
1240.6448
1250.6548
1260.6453
1270.6395
1280.6495
1290.6398
1300.6419
1310.6281
1320.6338
1330.6323
1340.6401
1350.6200
1360.6343
1370.6355
1380.6495
1390.6374
1400.6299
1410.6409
1420.6453
1430.6398
1440.6332
1450.6428
1460.6420
1470.6492
1480.6337
1490.6422
1500.6335
1510.6432
1520.6370
1530.6395
1540.6298
1550.6547
1560.6413
1570.6360
1580.6461
1590.6401
1600.6478
1610.6448
1620.6428
1630.6491
1640.6415
1650.6418
1660.6374
1670.6447
1680.6398
1690.6543
1700.6491
1710.6469
1720.6523
1730.6454
1740.6418
1750.6492
1760.6382
1770.6421
1780.6396
1790.6562
1800.6439
1810.6400
1820.6465
1830.6439
1840.6519
1850.6615
1860.6537
1870.6443
1880.6385
1890.6383
1900.6486
1910.6484
1920.6400
1930.6542
1940.6375
1950.6532
1960.6549
1970.6380
1980.6409
1990.6410
2000.6532
2010.6443
2020.6428
2030.6593
2040.6501
2050.6430
2060.6473
2070.6513
2080.6624
2090.6534
2100.6547
2110.6371
2120.6555
2130.6485
2140.6418
2150.6541
2160.6370
2170.6449
2180.6357
2190.6548
2200.6407
2210.6464
2220.6381
2230.6513
2240.6539
2250.6577
2260.6406
2270.6369
2280.6415
2290.6508
2300.6482
2310.6365
2320.6451
2330.6515
2340.6525
2350.6354
2360.6507
2370.6345
2380.6476
2390.6562
2400.6493
2410.6385
2420.6416
2430.6608
2440.6389
2450.6486
2460.6498
2470.6393
2480.6555
2490.6448
2500.6494
2510.6556
2520.6522
2530.6526
2540.6582
2550.6509
2560.6523
2570.6553
2580.6456
2590.6636
2600.6419
2610.6321
2620.6474
2630.6541
2640.6435
2650.6561
2660.6412
2670.6525
2680.6434
2690.6483
2700.6504
2710.6379
2720.6579
2730.6441
2740.6350
2750.6275
2760.6436
2770.6509
2780.6371
2790.6567
2800.6335
2810.6451
2820.6553
2830.6404
2840.6516
2850.6502
2860.6536
2870.6440
2880.6404
2890.6576
2900.6582
2910.6335
2920.6496
2930.6461
2940.6533
2950.6444
2960.6505
2970.6536
2980.6563
2990.6475
3000.6488
3010.6558
3020.6532
3030.6414
3040.6472
3050.6364
3060.6398
3070.6392
3080.6460
3090.6477
3100.6388
3110.6345
3120.6582
3130.6510
3140.6463
3150.6385
3160.6374
3170.6475
3180.6458
3190.6445
3200.6585
3210.6439
3220.6368
3230.6373
3240.6421
3250.6369
3260.6485
3270.6482
3280.6458
3290.6508
3300.6324
3310.6305
3320.6389
3330.6441
3340.6430
3350.6380
3360.6550
3370.6357
3380.6531
3390.6498
3400.6417
3410.6446
3420.6417
3430.6452
3440.6441
3450.6468
3460.6400
3470.6486
3480.6302
3490.6499
3500.6517
3510.6495
3520.6475
3530.6385
3540.6538
3550.6486
3560.6460
3570.6524
3580.6358
3590.6301
3600.6429
3610.6464
3620.6361
3630.6402
3640.6314
3650.6343
3660.6441
3670.6460
3680.6346
3690.6372
3700.6346
3710.6599
3720.6376
3730.6465
3740.6396
3750.6430
3760.6396
3770.6364
3780.6245
3790.6416
3800.6338
3810.6296
3820.6372
3830.6339
3840.6383
3850.6321
3860.6341
3870.6330
3880.6395
3890.6221
3900.6249
3910.6270
3920.6335
3930.6065
3940.6454
3950.6239
3960.6471
3970.6497
3980.6427
3990.6460
4000.6405
4010.6430
4020.6417
4030.6447
4040.6451
4050.6374
4060.6304
4070.6252
4080.6319
4090.6345
4100.6284
4110.6264
4120.6316
4130.6339
4140.6262
4150.6333
4160.6324
4170.6311
4180.6325
4190.6219
4200.6329
4210.6395
4220.6400
4230.6305
4240.6454
4250.6406
4260.6371
4270.6371
4280.6188
4290.6203
4300.6220
4310.6257
4320.6487
4330.6211
4340.6249
4350.6289
4360.6222
4370.6387
4380.6317
4390.6236
4400.6394
4410.6208
4420.6063
4430.6190
4440.6330
4450.6220
4460.6310
4470.6251
4480.6227
4490.6171
4500.6001
4510.6081
4520.6218
4530.6162
4540.6228
4550.6335
4560.6285
4570.6384
4580.6265
4590.6084
4600.6274
4610.6238
4620.6256
4630.6362
4640.6206
4650.6315
4660.6018
4670.6077
4680.6135
4690.6163
4700.6232
4710.6386
4720.6236
4730.6541
4740.6142
4750.6300
4760.6211
4770.6204
4780.6154
4790.5949
4800.6230
4810.6033
4820.6263
4830.6310
4840.6303
4850.6218
4860.6436
4870.6448
4880.6280
4890.6369
4900.6191
4910.6062
4920.5996
4930.6012
4940.6058
4950.6144
4960.6036
4970.6138
4980.6015
4990.6238
5000.6405
5010.6247
5020.6260
5030.6273
5040.6255
5050.6164
5060.6201
5070.6164
5080.6164
5090.6007
5100.6073
5110.5950
5120.6417
5130.6170
5140.6112
5150.6399
5160.6318
5170.6354
5180.6290
5190.6196
5200.6338
5210.6029
5220.6401
5230.6197
5240.6006
5250.6174
5260.6406
5270.6084
5280.6348
5290.6206
5300.6250
5310.6214
5320.6317
5330.6285
5340.6353
5350.6284
5360.6050
5370.6354
5380.6158
5390.6253
5400.5996
5410.6205
5420.6084
5430.5936
5440.6362
5450.6239
5460.5989
5470.6256
5480.5818
5490.6233
5500.5931
5510.6078
5520.6116
5530.5818
5540.6138
5550.5995
5560.5929
5570.6209
5580.6039
5590.6305
5600.6124
5610.6312
5620.6168
5630.6042
5640.6263
5650.6196
5660.6207
5670.5995
5680.6094
5690.6201
5700.5999
5710.6275
5720.6329
5730.6255
5740.6369
5750.6178
5760.6404
5770.6073
5780.6069
5790.6145
5800.5945
5810.6273
5820.6051
5830.6236
5840.6022
5850.6268
5860.6231
5870.6258
5880.6438
5890.6072
5900.6319
5910.6204
5920.6138
5930.6376
5940.6309
5950.6167
5960.5919
5970.5844
5980.6102
5990.5996
6000.6201
6010.6154
6020.6052
6030.6205
6040.6292
6050.6360
6060.6132
6070.6327
6080.6274
6090.6107
6100.5968
6110.6331
6120.6248
6130.6070
6140.6505
6150.6185
6160.6210
6170.6070
6180.6219
6190.6018
6200.6014
6210.5937
6220.6206
6230.5785
6240.5971
6250.6083
6260.6021
6270.6115
6280.5995
6290.6175
6300.6333
6310.6386
6320.6270
6330.6576
6340.5988
6350.6256
6360.6157
6370.5966
6380.6066
6390.6282
6400.5979
6410.6274
6420.6130
6430.6232
6440.6135
6450.6269
6460.5956
6470.5968
6480.6108
6490.6166
6500.6221
6510.6178
6520.6099
6530.6492
6540.6267
6550.6197
6560.6161
6570.6213
6580.6248
6590.5983
6600.5896
6610.6109
6620.6116
6630.5883
6640.6070
6650.6089
6660.6040
6670.6301
6680.5892
6690.5819
6700.6264
6710.6079
6720.6143
6730.6034
6740.6573
6750.6055
6760.6137
6770.6079
6780.6121
6790.6248
6800.6140
6810.5931
6820.5995
6830.6078
6840.6212
6850.6288
6860.6061
6870.6038
6880.6207
6890.6479
6900.6356
6910.6114
6920.6180
6930.6219
6940.6097
6950.6139
6960.6015
6970.6068
6980.6099
6990.6061
7000.6020
7010.6190
7020.6294
7030.6216
7040.6343
7050.6029
7060.6301
7070.6110
7080.6071
7090.6201
7100.6093
7110.5960
7120.6163
7130.5920
7140.6103
7150.6228
7160.6125
7170.6104
7180.6011
7190.6284
7200.6216
7210.6277
7220.6177
7230.5886
7240.6017
7250.6132
7260.6159
7270.6060
7280.6126
7290.6250
7300.6197
7310.6200
7320.6140
7330.6368
7340.5943
7350.6213
7360.6309
7370.5983
7380.6070
7390.6306
7400.6060
7410.6293
7420.6465
7430.6066
7440.6309
7450.5854
7460.6079
7470.6163
7480.6147
7490.5852
7500.5873
7510.6205
7520.6114
7530.6139
7540.6098
7550.6128
7560.6293
7570.6240
7580.6039
7590.6307
7600.6294
7610.6562
7620.6402
7630.6160
7640.5926
7650.6065
7660.6140
7670.6226
7680.6031
7690.6026
7700.6035
7710.6027
7720.6075
7730.6149
7740.6326
7750.6088
7760.6113
7770.6407
7780.6330
7790.6359
7800.6171
7810.6188
7820.6161
7830.6131
7840.6290
7850.5835
7860.6151
7870.6010
7880.6403
7890.5928
7900.6149
7910.6186
7920.6060
7930.5990
7940.5999
7950.6056
7960.5980
7970.6073
7980.6101
7990.6271
8000.6435
8010.6058
8020.6167
8030.6214
8040.6189
8050.6036
8060.6213
8070.6047
8080.6173
8090.6094
8100.6118
8110.6224
8120.6139
8130.6189
8140.5874
8150.6201
8160.5968
8170.6131
8180.5979
8190.6188
8200.6171
8210.6266
8220.6332
8230.6346
8240.6111
8250.6049
8260.6288
8270.5805
8280.5962
8290.5921
8300.6011
8310.5822
8320.6078
8330.5937
8340.6125
8350.6246
8360.5966
8370.5873
8380.6214
8390.6071
8400.6204
8410.6240
8420.6178
8430.6018
8440.5983
8450.6194
8460.6071
8470.6253
8480.6303
8490.5975
8500.6619
8510.6020
8520.6266
8530.6401
8540.6151
8550.6218
8560.6253
8570.6390
8580.6049
8590.6239
8600.5962
8610.6154
8620.5900
8630.6131
8640.6109
8650.6006
8660.6137
8670.6038
8680.6183
8690.6095
8700.6362
8710.6139
8720.6516
8730.6234
8740.6335
8750.6141
8760.6064
8770.5995
8780.5864
8790.6238
8800.6163
8810.5906
8820.6277
8830.6128
8840.6206
8850.5805
8860.6235
8870.6317
8880.6096
8890.6141
8900.6039
8910.5992
8920.6202
8930.6369
8940.5731
8950.6277
8960.6134
8970.6182
8980.6218
8990.6448
9000.6177
9010.6378
9020.6358
9030.6198
9040.6309
9050.6096
9060.6269
9070.6161
9080.6264
9090.6270
9100.6310
9110.6123
9120.6232
9130.5893
9140.6101
9150.6017
9160.6075
9170.6026
9180.5927
9190.6196
9200.6102
9210.6291
9220.6243
9230.6266
9240.6612
9250.6080
9260.6053
9270.6125
9280.6349
9290.6536
9300.6123
9310.5884
9320.5901
9330.6074
9340.6045
9350.6167
9360.6162
9370.6017
9380.6351
9390.6164
9400.6304
9410.5878
9420.6354
9430.6227
9440.6190
9450.6221
9460.6583
9470.6326
9480.6075
9490.6517
9500.6282
9510.6043
9520.6046
9530.6119
9540.5988
9550.6095
9560.6010
9570.6231
9580.5940
9590.6184
9600.6012
9610.5930
9620.5992
9630.6243
9640.6073
9650.6227
9660.6399
9670.6338
9680.6526
9690.5891
9700.6236
9710.5908
9720.6026
9730.6039
9740.6389
9750.5929
9760.6152
9770.5925
9780.6238
9790.6229
9800.6257
9810.6233
9820.6117
9830.6079
9840.6352
9850.6158
9860.6334
9870.5947
9880.6095
9890.6348
9900.5979
9910.6162
9920.6038
9930.6240
9940.6088
9950.6144
9960.6161
9970.6073
9980.5899
9990.6126
10000.6085
10010.5948
10020.6254
10030.6015
10040.5766
10050.6097
10060.6109
10070.6279
10080.6031
10090.6332
10100.6303
10110.5908
10120.6154
10130.6199
10140.6051
10150.6216
10160.6023
10170.6030
10180.6386
10190.6051
10200.6517
10210.6326
10220.6158
10230.6227
10240.6389
10250.6423
10260.5938
10270.6080
10280.6319
10290.6052
10300.6161
10310.6290
10320.6250
10330.6292
10340.6329
10350.6127
10360.6255
10370.6113
10380.6255
10390.6075
10400.6299
10410.6377
10420.6341
10430.6235
10440.6169
10450.6111
10460.6109
10470.6428
10480.5987
10490.6315
10500.6224
10510.5995
10520.6102
10530.5887
10540.6104
10550.6153
10560.6065
10570.6151
10580.5840
10590.6041
10600.6129
10610.6259
10620.6229
10630.6117
10640.6240
10650.6166
10660.6064
10670.6204
10680.6279
10690.6058
10700.6360
10710.6266
10720.5929
10730.5977
10740.6280
10750.6245
10760.6280
10770.6418
10780.6134
10790.6640
10800.5955
10810.5943
10820.6222
10830.6116
10840.6061
10850.6161
10860.6254
10870.6136
10880.6603
10890.6247
10900.6352
10910.6214
10920.6213
10930.5871
10940.6367
10950.6240
10960.6174
10970.6630
10980.5999
10990.5890
11000.5961
11010.6265
11020.6033
11030.6295
11040.6086
11050.6223
11060.6132
11070.6017
11080.6272
11090.6402
11100.6044
11110.6170
11120.6137
11130.6432
11140.6480
11150.6375
11160.6256
11170.6384
11180.6174
11190.6201
11200.6105
11210.6212
11220.5869
11230.5997
11240.5942
11250.6113
11260.5992
11270.6146
11280.6116
11290.6031
11300.6316
11310.6035
11320.6197
11330.6111
11340.6379
11350.6394
11360.6328
11370.5948
11380.5830
11390.6230
11400.6067
11410.6297
11420.6006
11430.6167
11440.6299
11450.5916
11460.6405
11470.6282
11480.6244
11490.5927
11500.6251
11510.6248
11520.6217
11530.6116
11540.6190
11550.5933
11560.6211
11570.6245
11580.5986
11590.6165
11600.5894
11610.6177
11620.6127
11630.6031
11640.6149
11650.6118
11660.6035
11670.6153
11680.6016
11690.6256
11700.6335
11710.5874
11720.5756
11730.6286
11740.5771
11750.6006
11760.6203
11770.6374
11780.6020
11790.5952
11800.6171
11810.6164
11820.6308
11830.6008
11840.5911
11850.6457
11860.6144
11870.6628
11880.6496
11890.6228
11900.6209
11910.6021
11920.6604
11930.6411
11940.5945
11950.6146
11960.6478
11970.6105
11980.6391
11990.6229
12000.6184
12010.6470
12020.5969
12030.6251
12040.6242
12050.6390
12060.6000
12070.6390
12080.6244
12090.6378
12100.6011
12110.6417
12120.6118
12130.6232
12140.6293
12150.6345
12160.5960
12170.6394
12180.6055
12190.6140
12200.6007
12210.6339
12220.6425
12230.6143
12240.6392
12250.6181
12260.5899
12270.6408
12280.6518
12290.5636
12300.6120
12310.5961
12320.6243
12330.6056
12340.6435
12350.6008
12360.6329
12370.6165
12380.6493
12390.6352
12400.5951
12410.6540
12420.6177
12430.6451
12440.6279
12450.6389
12460.6272
12470.6181
12480.6010
12490.6219
12500.5973
12510.6303
12520.5854
12530.5983
12540.5856
12550.6132
12560.6233
12570.6243
12580.6191
12590.6417
12600.6123
12610.6104
12620.6361
12630.6337
12640.6534
12650.6530
12660.5911
12670.5947
12680.6082
12690.6088
12700.6167
12710.6088
12720.6025
12730.6035
12740.6494
12750.6198
12760.6208
12770.6198
12780.6387
12790.6054
12800.6156
12810.6342
12820.6413
12830.5910
12840.6602
12850.5965
12860.6050
12870.5965
12880.6341
12890.5966
12900.6137
12910.6196
12920.6429
12930.5788
12940.6063
12950.6490
12960.6149
12970.6086
12980.6084
12990.5967
13000.5952
13010.6339
13020.6044
13030.6673
13040.5817
13050.6296
13060.6092
13070.6134
13080.6190
13090.6424
13100.6043
13110.6277
13120.5988
13130.6373
13140.6432
13150.6448
13160.6317
13170.5962
13180.6160
13190.6094
13200.6396

Measured result

Verbatim from the run's summary.json.

bucket
bucket-p32-t128
bucket_icons
3359
checkpoint_bytes
6463493
checkpoint_round_trip
true
checkpoint_sha256
7d14aacdb7b9ef91…
config_sha256
b9343dcc4a2e0012…
corruption
factorized_marginal_respecting_geometry
cuda_version
13.4
deterministic_algorithms
true
device
cuda
eval_every
60
final_train
loss
3.872609853744507
step
1320
train_token_accuracy
0.6396255850234009
group_vocabulary_size
12
locked_path_exact
true
metrics_sha256
bea69b3bb9abefb2…
model_parameters
525152
schema_version
1
scope
dominant-bucket fixed-topology geometry pilot; not unconditional generation
selected_train_rows
color/svg/1F3F4-E0064-E0065-E0062-E0065-E007F.svg, color/svg/1F468-1F3FC-200D-1F9BC.svg, color/svg/1F561.svg, color/svg/1F469-1F3FE-200D-1F9BC.svg, color/svg/1F469-1F3FD-200D-2764-FE0F-200D-1F48B-200D-1F469-1F3FE.svg, color/svg/1F6B5-1F3FB-200D-2642-FE0F.svg … and 2675 more
selected_validation_rows
color/svg/1F994.svg, color/svg/E30A.svg, color/svg/1F6BE.svg, color/svg/1F3CA-1F3FB-200D-2642-FE0F.svg, color/svg/1F3CB-1F3FC-200D-2642-FE0F.svg, color/svg/1F93D-1F3FC.svg … and 122 more
selection
completed_steps
1320
evals_without_improvement
8
min_delta
0.0
objective
held_out_loss
patience_evals
8
selected_held_out_loss
5.367093086242676
selected_step
840
stopped_early
true
steps
6300
study_version
openmoji-g1-full-denoiser-v11
subgroup_vocabulary_size
118
torch_version
2.14.0a0+4fdf77b940.nv26.08
validation
accuracy
0.5838582283543291
changed_accuracy
0.029553116706118644
changed_total
13941
loss
5.367093086242676
retained_accuracy
0.8803084359535044
retained_total
26067
validation_final_step
accuracy
0.588257348530294
changed_accuracy
0.031705042679865146
changed_total
13941
loss
5.445407867431641
retained_accuracy
0.8859093873479879
retained_total
26067
validation_trace_sha256
feaceee701a37550…
validation_untrained
accuracy
0.3340831833633273
changed_accuracy
0.0012911555842479018
changed_total
13941
loss
7.674570083618164
retained_accuracy
0.5120650631066099
retained_total
26067

Written result

v10 produced the project's first trained model to beat a zero-parameter heuristic on a leak-free task: detector lift 2.781 against a continuity statistic's 1.893. This run turns the value head back on, so the model reconstructs as well as detects, and tests the bar every model here has failed — the identity policy of emitting the input unchanged.

Result

v11identitybest prior (v6)
aggregate token accuracy0.58390.65150.6119
changed-token accuracy0.02960.00000.0132
retained-token accuracy0.88031.00000.9335
criterionthresholdobservedoutcome
beats_identity> 0.65150.5839falsified
recovers_more_than_identity> 0.01320.0296pass
preserves_what_identity_preserves>= 0.900.8803falsified
structural_safetylocked-path exact, round-tripsboth truepass
reproducibilityidentical rerunidenticalpass

Changed-token recovery is 2.24x the best any earlier model managed, which is real. Everything else falls short, and the reason is exact.

The value objective destroys the detector

modelobjectivedetector liftdetector precision
v10detection only2.7810.9691
v11joint1.7170.5983
—free continuity statistic1.893—

Adding the value head costs 1.064 of detector lift and drops it back below the free statistic. The two heads share one encoder, and a 289- or 417-way exact-token objective against a 2-class decision is not a fair fight: the encoder is shaped by the harder task, which it performs at 0.1087, and the easier one it had solved is collateral damage. The same pattern appeared at small scale between v6 and v7 (1.234 against 1.365); with a genuinely good detector to lose, it is now an order of magnitude larger.

Why no threshold rescues it

Flagging a field is only worth it if the expected gain beats keeping it. Keeping a retained field is always right; flagging one is right only if the value head happens to re-predict the same token. With the value head at 0.1087 on genuinely corrupted fields, the break-even detection confidence is

p > 1 / (1 + 0.1087) = 0.9019

The model flags 19.15% of fields, far past where it is that confident. But sweeping the threshold does not save it either — every flag rate loses to identity, monotonically:

flag rateaggregatevs identity
0.00 (identity)0.6515—
0.010.6496−0.0020
0.050.6374−0.0141
0.1915 (the model's own gate)0.5839−0.0668
0.350.5137−0.1378

Working backwards from the 1% point, this detector's precision at its most confident 1% is only about 0.73 — not the 0.9691 v10 achieved, because this is the degraded joint detector. With a detector at v10's precision and a value head at 0.1087, flagging the top percentile would have paid.

So identity is not beaten, and the two ingredients that would beat it have both been demonstrated — just never at the same time.

What this localises

The remaining problem is not detection: v10 solved that. It is that detection and reconstruction cannot currently be learned together, and that reconstruction itself is weak at 0.1087 exact-token accuracy on a 289- or 417-way vocabulary.

Both have concrete, measured next steps, and they are separable:

  1. Stop the two objectives competing. Weight them by field count rather than letting an exact-token softmax dominate, or give the heads separate encoders. v10 and v11 bracket exactly what is at stake: 2.781 against 1.717 on the identical setup.
  2. Make reconstruction learnable. The value head predicts an exact bin on a quarter-unit metric lattice through a categorical softmax. The task-formulation lens argued for a distance-kernel target — mass spread over nearby bins in proportion to their distance — so that being close earns gradient. That lens also measured that the model is a calibrated localiser being graded pass/fail at ±0.125 units.

Scope

Fixed-topology and geometry-only, single seed, marginal-respecting corruption at 0.35. Selected step 840 of 1,320 before early stopping, against v10's 5,340 — the joint objective also collapses training far sooner. No render was produced for this run; the token numbers do not warrant one, and that is itself consistent with the standing rule that no recovery claim is made without one.

State transitions

  1. planned2026-09-21T01:55:00Z
  2. completed2026-09-21T02:05:00Zthe_value_objective_costs_the_detector_1.064_of_lift_and_identity_is_still_not_beaten

Run record

Verbatim from runs/openmoji-g1-full-denoiser-v11-27e9956-b9343dcc-9b9b1699/run.yaml, the record committed before launch.

schema_version
1
run_id
openmoji-g1-full-denoiser-v11-27e9956-b9343dcc-9b9b1699
state
completed
parent_run
openmoji-g1-metric-coordinates-v10-abec130-d7044dc2-9b9b1699
hypothesis
Every trained model in this project scores below the identity policy - emit the corrupted input unchanged - on held-out token accuracy, the best by 0.27. Four verified defects have since been fixed: the loss averaged over groups so four fields carried 31% of it, no attention padding mask so 29.4% of the sequence was attended as content, corruption drew uniformly so a token-marginal detector scored 2.84 with no context, and coordinates were encoded as unordered categories so no difference was computable. With all four fixed, the detector now beats every free reference at 2.781 against a continuity statistic's 1.893. Running the full objective - the value head training alongside the keep head, with the keep decision gating the decode - should for the first time clear the identity baseline.
expected_information_gain
This is the bar the whole Gate G sequence has failed, and the first time it is attempted with a working detector and a sound training setup. Clearing it would mean the project has a denoiser that does something rather than nothing. Failing it with a detector this good would localise the remaining problem to reconstruction - knowing WHICH field is wrong but not WHAT belongs there - which is a different and narrower question than any asked so far.
scope_limits
Fixed-topology and geometry-only, single seed, marginal-respecting corruption at probability 0.35. Beating identity on token accuracy is not the same as producing a recognizable icon; the render probe is run separately and no run in this project has produced one yet.
changed_factors
primary
training.detection_only true -> false, so the value head trains.
parameters
unchanged at 525,152
held_fixed
v10 exactly otherwise: metric coordinates, marginal corruption, pooled loss, padding mask, edit-mask heads, no slot binding, the full 2,681-icon split, the 128-icon validation draw and its corruption seeds, seed 3101, probability 0.35, batch 16, lr 0.001, eval_every 60, the 6,300-step cap, held-out-loss selection, patience 8.
code
git_commit
27e9956
execution_mode
native-local
config
path
configs/learning/openmoji-g1-full-denoiser-v11.yaml
sha256
b9343dcc4a2e0012…
dataset
hybrid_sha256
9b9b1699677a6f97…
bucket
bucket-p32-t128
train_samples
2681
validation_samples
128
baselines
identity
definition
Emit x_t unchanged. Legal under argmax decoding, and exactly representable here as predict-keep-everywhere. Its aggregate accuracy is retained_total divided by (changed_total + retained_total) for this run's own corruption draw, so it is computed from the run's reported totals rather than assumed. Under uniform corruption at this probability it was 0.6506; under marginal corruption the corrupted rate is 0.3482, so it is expected near 0.6518.
changed_accuracy
0.0
retained_accuracy
1.0
best_prior_model
run
v6
aggregate
0.6119
changed
0.0132
retained
0.9335
v10_detector_same_setup
lift
2.781
note
detection only
no value head
None
predeclared_criteria
evaluated_at
the checkpoint selected by held-out loss
primary
id
beats_identity
statement
Held-out aggregate token accuracy exceeds the identity policy's, computed as retained_total / (changed_total + retained_total) from this run's own totals. No model in this project has done this.
id
recovers_more_than_identity
statement
Held-out changed-token accuracy exceeds v6's 0.0132, the best any edit-mask model achieved. Identity scores 0.0 here, so passing beats_identity while scoring near zero on changed fields would be collapse, not success.
id
preserves_what_identity_preserves
statement
Held-out retained-token accuracy is at least 0.90.
id
structural_safety
statement
Locked-path exactness holds and the canonical checkpoint round-trips.
id
reproducibility
statement
A second complete invocation returns an identical JSON result.
standing_gate_not_a_prediction
id
retained_preservation_full
statement
Held-out retained-token accuracy is at least 0.95.
note
A stricter form of the original Gate G bar, unmet since v1.
degenerate_pass_watch
Predicting keep everywhere scores exactly identity. Report the fraction of held-out fields predicted keep; a pass on beats_identity with that fraction approaching 1.0 is collapse to the trivial policy and must be reported as such.
falsification_meaning
If a model that detects corrupted fields at 2.781 lift still cannot beat emitting its input unchanged, the remaining problem is reconstruction rather than detection: it knows which field is wrong and not what belongs there. That is narrower than anything asked so far and points at the value head's 289- or 417-way exact-token softmax over a metric lattice, which the task-formulation lens argued should be a distance-kernel target instead.
reported_not_gated
the keep fraction, detector lift, and the full held-out breakdown beside identity, the selected step and whether early stopping or the cap ended the run, renders, which are run separately and reported whatever the token numbers say
outputs
local_metadata
runs/openmoji-g1-full-denoiser-v11-27e9956-b9343dcc-9b9b1699
report_root
reports/learning/openmoji-g1-full-denoiser-v11
durable_artifacts
/home/dev/.cache/openmoji-g1-full-denoiser-v11-27e9956-b9343dcc-9b9b1699
planned_at
2026-09-21 01:55:00+00:00
completed_at
2026-09-21 02:05:00+00:00
result
device
cuda
model_parameters
525152
selected_step
840
completed_steps
1320
ended_by
early_stopping
identical_rerun
true
held_out
aggregate
0.5839
changed
0.0296
retained
0.8803
identity_baseline
aggregate
0.6515
changed
0.0
retained
1.0
detector_lift
1.717
detector_lift_detection_only_same_setup
2.781
value_head_exact_accuracy_on_corrupted_fields
0.1087
break_even_detection_confidence
0.9019
threshold_sweep_best
flag_rate
0.0
aggregate
0.6515
note
no flag rate beats identity
predeclared_outcome
overall
falsified
beats_identity
passed
false
observed
0.5839
threshold
0.6515
recovers_more_than_identity
passed
true
observed
0.0296
threshold
0.0132
preserves_what_identity_preserves
passed
false
observed
0.8803
threshold
0.9
structural_safety
passed
true
reproducibility
passed
true
degenerate_pass_watch
fired
false
flag_fraction
0.1915
note
the model over-flags rather than collapsing to keep
conclusion
Changed-token recovery is 2.24x the best any earlier model managed, but identity is not beaten and the reason is exact: adding the value head costs 1.064 of detector lift, 2.781 down to 1.717, dropping it back below the free continuity statistic. The two heads share one encoder and a 289/417-way exact-token objective against a 2-class decision is not a fair fight. No flag rate rescues it - every threshold loses to identity monotonically - because this degraded detector reaches only about 0.73 precision at its most confident 1%, where v10 reached 0.9691.
localises
Not detection, which v10 solved. Detection and reconstruction cannot currently be learned together, and reconstruction itself is weak at 0.1087 exact-token accuracy. Two separable next steps: weight the two objectives by field count or give them separate encoders, and replace the exact-token value target with a distance kernel over the quarter-unit lattice so being close earns gradient.