masked-overfit-l6-metric-head-5475283-4icons-9b9b1699
completed —
Measured behaviour
modelposition-marginal floor
Table view
| optimizer step | model | position-marginal floor |
|---|---|---|
| 0 | 6.7085 | 4.9159 |
| 100 | 5.0807 | 4.9159 |
| 200 | 4.6296 | 4.9159 |
| 300 | 3.7614 | 4.9159 |
| 400 | 2.8905 | 4.9159 |
| 500 | 2.4920 | 4.9159 |
| 600 | 2.2785 | 4.9159 |
| 700 | 2.2434 | 4.9159 |
| 800 | 2.2406 | 4.9159 |
| 900 | 2.2185 | 4.9159 |
| 1000 | 2.2086 | 4.9159 |
| 1100 | 2.2067 | 4.9159 |
| 1200 | 2.1957 | 4.9159 |
| 1300 | 2.1853 | 4.9159 |
| 1400 | 2.1769 | 4.9159 |
| 1500 | 2.1620 | 4.9159 |
| 1600 | 2.1498 | 4.9159 |
| 1700 | 2.1395 | 4.9159 |
| 1800 | 2.1230 | 4.9159 |
| 1900 | 2.1289 | 4.9159 |
| 2000 | 2.2029 | 4.9159 |
| 2100 | 2.1445 | 4.9159 |
| 2200 | 2.1224 | 4.9159 |
| 2300 | 2.1171 | 4.9159 |
| 2400 | 2.1057 | 4.9159 |
| 2500 | 2.1069 | 4.9159 |
| 2600 | 2.0963 | 4.9159 |
| 2700 | 2.0939 | 4.9159 |
| 2800 | 2.0861 | 4.9159 |
| 2900 | 2.0873 | 4.9159 |
| 3000 | 2.0826 | 4.9159 |
modelposition-marginal floor
Table view
| optimizer step | model | position-marginal floor |
|---|---|---|
| 0 | 0.0196 | 0.5470 |
| 100 | 0.0519 | 0.5470 |
| 200 | 0.0912 | 0.5470 |
| 300 | 0.1585 | 0.5470 |
| 400 | 0.2244 | 0.5470 |
| 500 | 0.3408 | 0.5470 |
| 600 | 0.4727 | 0.5470 |
| 700 | 0.5091 | 0.5470 |
| 800 | 0.5554 | 0.5470 |
| 900 | 0.5933 | 0.5470 |
| 1000 | 0.6353 | 0.5470 |
| 1100 | 0.6339 | 0.5470 |
| 1200 | 0.6396 | 0.5470 |
| 1300 | 0.6760 | 0.5470 |
| 1400 | 0.7293 | 0.5470 |
| 1500 | 0.7447 | 0.5470 |
| 1600 | 0.7616 | 0.5470 |
| 1700 | 0.6971 | 0.5470 |
| 1800 | 0.8359 | 0.5470 |
| 1900 | 0.7854 | 0.5470 |
| 2000 | 0.5820 | 0.5470 |
| 2100 | 0.8359 | 0.5470 |
| 2200 | 0.8569 | 0.5470 |
| 2300 | 0.8654 | 0.5470 |
| 2400 | 0.9074 | 0.5470 |
| 2500 | 0.8892 | 0.5470 |
| 2600 | 0.9187 | 0.5470 |
| 2700 | 0.9481 | 0.5470 |
| 2800 | 0.9607 | 0.5470 |
| 2900 | 0.9313 | 0.5470 |
| 3000 | 0.9579 | 0.5470 |
Measured result
Verbatim from the run's summary.json.
checkpoint_round_trip
true
checkpoint_sha256
4b122fbc5f025b85…checks
all_valid
true
exact_path_reproduction
true
masked_token_accuracy
false
config_sha256
668a5d357336fb76…continuity
copy_previous
mean_bins_off
28.125
median_bins_off
9.25
n
4
marginal
mean_bins_off
116.125
median_bins_off
113.5
n
4
model
mean_bins_off
0.0
median_bins_off
0.0
n
4
coordinate_tau
1.0
criteria
min_exact_path_reproduction_rate
1.0
min_masked_token_accuracy
0.99
require_all_valid
true
deterministic_algorithms
true
device
cuda
evaluation_icons
4
evaluation_split
train
first_train_loss
5.12129020690918
gpu
NVIDIA GeForce RTX 4080
held_out_nll_per_masked_token
2.082627733827139
icon_presentations
12000
initial_held_out_nll
6.7085084206258765
inpainting
all_valid
true
attempted
4
decoding
greedy_for_criteria
true
iterations
8
samples
1
drop
median_rgba_mae
0.0034825291999515855
exact_reproduction_rate
1.0
icons
4
marginal
helped
0
mean_recovery
-1.8921178315003433
median_recovery
-2.0224896560213343
median_rgba_mae
0.010612782921810698
model
helped
4
mean_recovery
1.0
median_recovery
1.0
median_rgba_mae
0.0
paired
drop_minus_model
interval
-0.01649456796035647, 0.03891612145829885
interval_excludes_zero_above
false
mean
0.011210776748971191
n
4
positive
4
marginal_minus_model
interval
-0.0119139906416794, 0.04538188204206671
interval_excludes_zero_above
false
mean
0.016733945700193657
n
4
positive
4
rows_sha256
6f9b9087277b9470…sheet_sha256
693ac2cfbd6464a2…last_train_loss
1.941217303276062
loss_reduction_factor
3.2211750144602136
marginal_masked_token_accuracy
0.5469845722300141
marginal_nll_per_masked_token
4.9158920138280155
mask_mixture
drawn
geometry
1203
path
3603
random
3605
span
2422
style
1167
max_paths
2
random_rate
0.05, 0.95
span_max
None
weights
geometry
0.1
path
0.3
random
0.3
span
0.2
style
0.1
masked_token_accuracy
0.9579242636746143
metric_head
true
metrics_sha256
eec445629b3929de…model_parameters
528244
nll_ratio_to_marginal
0.42365205093376174
path_binding
false
predeclared_outcome
falsified
schema_version
1
selected_step
3000
sequence_length
1376
steps_run
3000
study_version
masked-overfit-l6-metric-head
timing
peak_vram_gib
0.49811792373657227
prepare_seconds
0.2582809339801315
train_seconds
107.14212694301386
torch_version
2.14.0a0+4fdf77b940.nv26.08
train_icons
4
vocabulary
418
Visual output
State transitions
- planned2026-09-21T09:03:05Z
- completed2026-09-21T09:06:02Zthe strong criterion passes - 4 of 4 masked whole paths reproduced token for token with pixel-identical renders, every completion valid, continuity error 0.0 bins on the training icons - and masked-token accuracy is 0.958 against 0.99, still rising at the last evaluation (0.907, 0.919, 0.961, 0.958 over the last four)
Run record
This run has no run.yaml. What follows is the identity and configuration carried by its rows in state/runs.jsonl, the append-only registry.
run_id
masked-overfit-l6-metric-head-5475283-4icons-9b9b1699
config
configs/learning/masked-overfit-l6-metric-head.yaml
outputs
checkpoint_root
data/processed/masked-overfit-l6-metric-head
report_root
reports/learning/masked-overfit-l6-metric-head
