masked-overfit-l1-f51c119-4icons-9b9b1699
completed —
Measured behaviour
modelposition-marginal floor
Table view
| optimizer step | model | position-marginal floor |
|---|---|---|
| 0 | 5.7808 | 4.9159 |
| 100 | 4.8317 | 4.9159 |
| 200 | 3.6715 | 4.9159 |
| 300 | 2.4638 | 4.9159 |
| 400 | 2.2152 | 4.9159 |
| 500 | 2.1695 | 4.9159 |
| 600 | 2.1246 | 4.9159 |
| 700 | 2.0997 | 4.9159 |
| 800 | 2.0844 | 4.9159 |
| 900 | 2.0658 | 4.9159 |
| 1000 | 2.0500 | 4.9159 |
| 1100 | 2.0401 | 4.9159 |
| 1200 | 2.0320 | 4.9159 |
| 1300 | 2.0244 | 4.9159 |
| 1400 | 2.0200 | 4.9159 |
| 1500 | 2.0076 | 4.9159 |
| 1600 | 2.0014 | 4.9159 |
| 1700 | 1.9965 | 4.9159 |
| 1800 | 1.9928 | 4.9159 |
| 1900 | 1.9969 | 4.9159 |
| 2000 | 1.9858 | 4.9159 |
| 2100 | 1.9882 | 4.9159 |
| 2200 | 1.9782 | 4.9159 |
| 2300 | 1.9798 | 4.9159 |
| 2400 | 1.9776 | 4.9159 |
| 2500 | 1.9788 | 4.9159 |
| 2600 | 1.9788 | 4.9159 |
| 2700 | 1.9703 | 4.9159 |
| 2800 | 1.9729 | 4.9159 |
| 2900 | 1.9748 | 4.9159 |
| 3000 | 1.9698 | 4.9159 |
modelposition-marginal floor
Table view
| optimizer step | model | position-marginal floor |
|---|---|---|
| 0 | 0.0238 | 0.5470 |
| 100 | 0.0715 | 0.5470 |
| 200 | 0.2062 | 0.5470 |
| 300 | 0.4418 | 0.5470 |
| 400 | 0.6241 | 0.5470 |
| 500 | 0.7363 | 0.5470 |
| 600 | 0.8205 | 0.5470 |
| 700 | 0.8794 | 0.5470 |
| 800 | 0.9046 | 0.5470 |
| 900 | 0.9467 | 0.5470 |
| 1000 | 0.9621 | 0.5470 |
| 1100 | 0.9776 | 0.5470 |
| 1200 | 0.9832 | 0.5470 |
| 1300 | 0.9916 | 0.5470 |
| 1400 | 0.9916 | 0.5470 |
| 1500 | 0.9972 | 0.5470 |
| 1600 | 0.9972 | 0.5470 |
| 1700 | 0.9972 | 0.5470 |
| 1800 | 0.9986 | 0.5470 |
| 1900 | 0.9888 | 0.5470 |
| 2000 | 0.9986 | 0.5470 |
| 2100 | 1.0000 | 0.5470 |
| 2200 | 1.0000 | 0.5470 |
| 2300 | 0.9972 | 0.5470 |
| 2400 | 1.0000 | 0.5470 |
| 2500 | 0.9986 | 0.5470 |
| 2600 | 1.0000 | 0.5470 |
| 2700 | 1.0000 | 0.5470 |
| 2800 | 1.0000 | 0.5470 |
| 2900 | 0.9986 | 0.5470 |
| 3000 | 1.0000 | 0.5470 |
Measured result
Verbatim from the run's summary.json.
checkpoint_round_trip
true
checkpoint_sha256
d67d6ae3d447d38e…checks
all_valid
true
exact_path_reproduction
true
loss_reduction
false
masked_token_accuracy
true
config_sha256
64bf0b75e86afc24…coordinate_tau
1.0
criteria
min_exact_path_reproduction_rate
1.0
min_loss_reduction_factor
100.0
min_masked_token_accuracy
0.99
require_all_valid
true
deterministic_algorithms
true
device
cuda
evaluation_icons
4
evaluation_split
train
first_train_loss
4.840734004974365
gpu
NVIDIA GeForce RTX 4080
held_out_nll_per_masked_token
1.9697834030943198
icon_presentations
12000
initial_held_out_nll
5.7808438979225105
inpainting
all_valid
true
attempted
4
decoding
greedy_for_criteria
true
iterations
8
samples
1
drop
median_rgba_mae
0.0034825291999515855
exact_reproduction_rate
1.0
icons
4
marginal
helped
0
mean_recovery
-1.8921178315003433
median_recovery
-2.0224896560213343
median_rgba_mae
0.010612782921810698
model
helped
4
mean_recovery
1.0
median_recovery
1.0
median_rgba_mae
0.0
paired
drop_minus_model
interval
-0.01649456796035647, 0.03891612145829885
interval_excludes_zero_above
false
mean
0.011210776748971191
n
4
positive
4
marginal_minus_model
interval
-0.0119139906416794, 0.04538188204206671
interval_excludes_zero_above
false
mean
0.016733945700193657
n
4
positive
4
rows_sha256
b737fc98c371338e…sheet_sha256
0f12c5ea2b77f8ee…last_train_loss
1.938839077949524
loss_reduction_factor
2.9347611970135503
marginal_masked_token_accuracy
0.5469845722300141
marginal_nll_per_masked_token
4.9158920138280155
mask_mixture
drawn
geometry
1203
path
3603
random
3605
span
2422
style
1167
max_paths
2
random_rate
0.05, 0.95
weights
geometry
0.1
path
0.3
random
0.3
span
0.2
style
0.1
masked_token_accuracy
1.0
metrics_sha256
b91e20d2f8b79b6d…model_parameters
526498
nll_ratio_to_marginal
0.4006970449215473
predeclared_outcome
falsified
schema_version
1
selected_step
3000
sequence_length
1376
steps_run
3000
study_version
masked-overfit-l1
timing
peak_vram_gib
0.4974055290222168
prepare_seconds
0.25998882099520415
train_seconds
106.27825578599004
torch_version
2.14.0a0+4fdf77b940.nv26.08
train_icons
4
vocabulary
418
Visual output
State transitions
- planned2026-09-21T08:12:54Z
- completed2026-09-21T08:16:16Zthe plumbing is sound: masked-token accuracy 1.000 over the free positions under the fixed masks, 4 of 4 masked whole paths reproduced token for token by greedy grammar-ordered completion with pixel-identical renders, every greedy, sampled and marginal completion valid; the loss-reduction check is falsified as written because under a soft coordinate target the exact-token likelihood cannot fall below the kernel's own entropy, about 2 nats per coordinate, so 2.9x is a ceiling set by the target rather than a fact about learning
Run record
This run has no run.yaml. What follows is the identity and configuration carried by its rows in state/runs.jsonl, the append-only registry.
run_id
masked-overfit-l1-f51c119-4icons-9b9b1699
config
configs/learning/masked-overfit-l1.yaml
outputs
checkpoint_root
data/processed/masked-overfit-l1
report_root
reports/learning/masked-overfit-l1
