masked-overfit-l1-01917d5-4icons-9b9b1699
completed —
Measured behaviour
modelposition-marginal floor
Table view
| optimizer step | model | position-marginal floor |
|---|---|---|
| 0 | 5.7808 | 4.9159 |
| 100 | 4.9008 | 4.9159 |
| 200 | 3.7819 | 4.9159 |
| 300 | 2.6064 | 4.9159 |
| 400 | 2.3416 | 4.9159 |
| 500 | 2.3067 | 4.9159 |
| 600 | 2.2775 | 4.9159 |
| 700 | 2.2607 | 4.9159 |
| 800 | 2.2491 | 4.9159 |
| 900 | 2.2376 | 4.9159 |
| 1000 | 2.2246 | 4.9159 |
| 1100 | 2.2210 | 4.9159 |
| 1200 | 2.2146 | 4.9159 |
| 1300 | 2.2071 | 4.9159 |
| 1400 | 2.2036 | 4.9159 |
| 1500 | 2.1974 | 4.9159 |
| 1600 | 2.1943 | 4.9159 |
| 1700 | 2.1936 | 4.9159 |
| 1800 | 2.1901 | 4.9159 |
| 1900 | 2.2020 | 4.9159 |
| 2000 | 2.1854 | 4.9159 |
| 2100 | 2.1887 | 4.9159 |
| 2200 | 2.1830 | 4.9159 |
| 2300 | 2.1815 | 4.9159 |
| 2400 | 2.1833 | 4.9159 |
| 2500 | 2.1831 | 4.9159 |
| 2600 | 2.1817 | 4.9159 |
| 2700 | 2.1743 | 4.9159 |
| 2800 | 2.1807 | 4.9159 |
| 2900 | 2.1850 | 4.9159 |
| 3000 | 2.1778 | 4.9159 |
modelposition-marginal floor
Table view
| optimizer step | model | position-marginal floor |
|---|---|---|
| 0 | 0.0238 | 0.5470 |
| 100 | 0.0687 | 0.5470 |
| 200 | 0.1374 | 0.5470 |
| 300 | 0.2118 | 0.5470 |
| 400 | 0.2482 | 0.5470 |
| 500 | 0.1907 | 0.5470 |
| 600 | 0.1388 | 0.5470 |
| 700 | 0.1150 | 0.5470 |
| 800 | 0.1080 | 0.5470 |
| 900 | 0.0912 | 0.5470 |
| 1000 | 0.0743 | 0.5470 |
| 1100 | 0.0743 | 0.5470 |
| 1200 | 0.0729 | 0.5470 |
| 1300 | 0.0715 | 0.5470 |
| 1400 | 0.0729 | 0.5470 |
| 1500 | 0.0757 | 0.5470 |
| 1600 | 0.0687 | 0.5470 |
| 1700 | 0.0687 | 0.5470 |
| 1800 | 0.0687 | 0.5470 |
| 1900 | 0.0687 | 0.5470 |
| 2000 | 0.0687 | 0.5470 |
| 2100 | 0.0687 | 0.5470 |
| 2200 | 0.0687 | 0.5470 |
| 2300 | 0.0687 | 0.5470 |
| 2400 | 0.0687 | 0.5470 |
| 2500 | 0.0687 | 0.5470 |
| 2600 | 0.0687 | 0.5470 |
| 2700 | 0.0687 | 0.5470 |
| 2800 | 0.0687 | 0.5470 |
| 2900 | 0.0687 | 0.5470 |
| 3000 | 0.0687 | 0.5470 |
Measured result
Verbatim from the run's summary.json.
checkpoint_round_trip
true
checkpoint_sha256
3da5b0f8f824588d…checks
all_valid
true
exact_path_reproduction
false
loss_reduction
false
masked_token_accuracy
false
config_sha256
98000972b1c6898b…coordinate_tau
1.0
criteria
min_exact_path_reproduction_rate
1.0
min_loss_reduction_factor
100.0
min_masked_token_accuracy
0.99
require_all_valid
true
deterministic_algorithms
true
device
cuda
evaluation_icons
4
evaluation_split
train
first_train_loss
4.846034049987793
gpu
NVIDIA GeForce RTX 4080
held_out_nll_per_masked_token
2.174339690348878
icon_presentations
12000
initial_held_out_nll
5.7808438979225105
inpainting
all_valid
true
attempted
4
decoding
greedy_for_criteria
true
iterations
8
samples
1
drop
median_rgba_mae
0.0034825291999515855
exact_reproduction_rate
0.0
icons
4
marginal
helped
0
mean_recovery
-1.8921178315003433
median_recovery
-2.0224896560213343
median_rgba_mae
0.010612782921810698
model
helped
4
mean_recovery
0.8152479893889422
median_recovery
0.7915168120171914
median_rgba_mae
0.0007350104393609296
paired
drop_minus_model
interval
-0.016603383721697625, 0.03754687878947782
interval_excludes_zero_above
false
mean
0.010471747533890099
n
4
positive
4
marginal_minus_model
interval
-0.011947317879671328, 0.04393715084989645
interval_excludes_zero_above
false
mean
0.015994916485112563
n
4
positive
4
rows_sha256
31a4c874aa5445cb…sheet_sha256
8600d06f82c56323…last_train_loss
1.9388209581375122
loss_reduction_factor
2.6586664096606545
marginal_masked_token_accuracy
0.5469845722300141
marginal_nll_per_masked_token
4.9158920138280155
mask_mixture
drawn
geometry
1203
path
3603
random
3605
span
2422
style
1167
max_paths
2
random_rate
0.05, 0.95
weights
geometry
0.1
path
0.3
random
0.3
span
0.2
style
0.1
masked_token_accuracy
0.06872370266479663
metrics_sha256
7ba8cb44028dbaf3…model_parameters
526498
nll_ratio_to_marginal
0.4423082696350189
predeclared_outcome
falsified
schema_version
1
selected_step
2700
sequence_length
1376
steps_run
3000
study_version
masked-overfit-l1
timing
peak_vram_gib
0.4974055290222168
prepare_seconds
0.2613500489969738
train_seconds
106.223827386013
torch_version
2.14.0a0+4fdf77b940.nv26.08
train_icons
4
vocabulary
418
Visual output
State transitions
- planned2026-09-21T08:07:04Z
- completed2026-09-21T08:12:33Zfalsified by a harness defect the test exists to catch: every non-coordinate token type is memorised exactly (lengths, layers, styles, kinds at accuracy 1.000, nll ~0) and every one of 407 masked coordinates is wrong by exactly one bin, because the soft coordinate target was centred on token truth-1 - the denoiser's bin-index convention copied into a loss whose logits index tokens; eval and train forwards agree to 1e-5, so it is not the encoder
Run record
This run has no run.yaml. What follows is the identity and configuration carried by its rows in state/runs.jsonl, the append-only registry.
run_id
masked-overfit-l1-01917d5-4icons-9b9b1699
config
configs/learning/masked-overfit-l1.yaml
artifacts
checkpoint_root
data/processed/masked-overfit-l1-attempt1
report_root
reports/learning/masked-overfit-l1-attempt1
