masked-span-l20-specialist-4-span4-9202129-2681icons-9b9b1699
completed —
Measured behaviour
modelposition-marginal floor
Table view
| optimizer step | model | position-marginal floor |
|---|---|---|
| 0 | 6.2190 | 4.7651 |
| 300 | 4.4374 | 4.7651 |
| 600 | 4.3458 | 4.7651 |
| 900 | 4.2691 | 4.7651 |
| 1200 | 4.2303 | 4.7651 |
| 1500 | 4.2036 | 4.7651 |
| 1800 | 4.1980 | 4.7651 |
| 2100 | 4.1683 | 4.7651 |
| 2400 | 4.1551 | 4.7651 |
| 2700 | 4.1356 | 4.7651 |
| 3000 | 4.1411 | 4.7651 |
| 3300 | 4.1260 | 4.7651 |
| 3600 | 4.1054 | 4.7651 |
| 3900 | 4.0552 | 4.7651 |
| 4200 | 4.0108 | 4.7651 |
| 4500 | 3.9464 | 4.7651 |
| 4800 | 3.9314 | 4.7651 |
| 5100 | 3.8987 | 4.7651 |
| 5400 | 3.9052 | 4.7651 |
| 5700 | 3.8996 | 4.7651 |
| 6000 | 3.8908 | 4.7651 |
| 6300 | 3.8552 | 4.7651 |
modelposition-marginal floor
Table view
| optimizer step | model | position-marginal floor |
|---|---|---|
| 0 | 0.0520 | 0.1280 |
| 300 | 0.1386 | 0.1280 |
| 600 | 0.1406 | 0.1280 |
| 900 | 0.1467 | 0.1280 |
| 1200 | 0.1532 | 0.1280 |
| 1500 | 0.1608 | 0.1280 |
| 1800 | 0.1511 | 0.1280 |
| 2100 | 0.1602 | 0.1280 |
| 2400 | 0.1587 | 0.1280 |
| 2700 | 0.1646 | 0.1280 |
| 3000 | 0.1590 | 0.1280 |
| 3300 | 0.1614 | 0.1280 |
| 3600 | 0.1631 | 0.1280 |
| 3900 | 0.1611 | 0.1280 |
| 4200 | 0.1631 | 0.1280 |
| 4500 | 0.1789 | 0.1280 |
| 4800 | 0.1628 | 0.1280 |
| 5100 | 0.1663 | 0.1280 |
| 5400 | 0.1786 | 0.1280 |
| 5700 | 0.1730 | 0.1280 |
| 6000 | 0.1713 | 0.1280 |
| 6300 | 0.1774 | 0.1280 |
Measured result
Verbatim from the run's summary.json.
checkpoint_round_trip
true
checkpoint_sha256
8e2413d580f59d78…checkpoint_source
None
checkpoint_source_sha256
None
checks
all_valid
true
beats_drop_baseline
false
beats_marginal_baseline
true
median_recovery
false
config_sha256
91f01db8a48aeeeb…continuity
copy_previous
mean_bins_off
25.64749262536873
median_bins_off
17.0
n
339
marginal
mean_bins_off
140.05162241887905
median_bins_off
138.5
n
339
model
mean_bins_off
22.073746312684367
median_bins_off
15.0
n
339
coordinate_tau
1.0
criteria
beats_drop_baseline
true
beats_marginal_baseline
true
min_median_recovery
0.3
require_all_valid
true
deterministic_algorithms
true
device
cuda
evaluation_icons
339
evaluation_split
validation
first_train_loss
6.352187633514404
gpu
NVIDIA GeForce RTX 4080
held_out_nll_per_masked_token
3.855228947045963
icon_presentations
100800
initial_held_out_nll
6.218960549879478
inpainting
all_valid
true
attempted
56
close_reproduction_rate
0.0
decoding
chain_order
true
greedy_for_criteria
true
iterations
8
samples
1
drop
median_rgba_mae
0.005432628298232873
exact_reproduction_rate
0.0
icons
56
identity_policy
join_span
inpaint_seed
3501
marginal
helped
4
mean_recovery
-9243201.564606853
median_recovery
-1.5514862525194055
median_rgba_mae
0.02051230407286371
model
helped
24
mean_recovery
-1.1948874697817782
median_recovery
-0.013644876652820558
median_rgba_mae
0.004619322651900267
paired
drop_minus_model
interval
-0.0016402854628366566, 0.0005194522916052009
interval_excludes_zero_above
false
mean
-0.0005604165856157279
n
56
positive
24
marginal_minus_model
interval
0.004541531249884808, 0.015153157968395097
interval_excludes_zero_above
true
mean
0.009847344609139953
n
56
positive
49
rows_sha256
3ff53db42f479f0f…sheet_sha256
07790ad9e994c3f2…span_length
4
task
span
last_train_loss
3.964060068130493
loss_reduction_factor
1.6131235356708724
marginal_masked_token_accuracy
0.12803273896521486
marginal_nll_per_masked_token
4.7650980387679045
mask_mixture
drawn
span
100800
max_paths
2
random_rate
0.05, 0.95
span_max
4
weights
span
1.0
masked_token_accuracy
0.17743349897690733
metric_head
true
metrics_sha256
ef088d9ebaeecb83…model_parameters
531892
nll_ratio_to_marginal
0.8090555358317028
path_binding
false
predeclared_outcome
falsified
schema_version
1
selected_step
6300
sequence_length
1376
start_features
true
steps_run
6300
study_version
masked-span-l20-specialist-4-span4
timing
peak_vram_gib
1.9460439682006836
prepare_seconds
102.55899478000356
train_seconds
418.81431883000187
torch_version
2.14.0a0+4fdf77b940.nv26.08
train_icons
2681
trained_here
true
vocabulary
418
Visual output
State transitions
- planned2026-09-21T10:30:42Z
- completed2026-09-21T10:44:16Ztrained on spans of one to four and read at four, the specialist ties the join fill - mean -0.0006, interval [-0.0016, +0.0005] spanning zero, 24 of 56 helped, median recovery -0.01 - beats the marginal policy on 49 of 56, and is indistinguishable span for span from the model trained on shorter spans (30 better, 24 worse); held-out likelihood still falling at step 6,300; every completion valid
Run record
This run has no run.yaml. What follows is the identity and configuration carried by its rows in state/runs.jsonl, the append-only registry.
run_id
masked-span-l20-specialist-4-span4-9202129-2681icons-9b9b1699
config
configs/learning/masked-span-l20-specialist-4-span4.yaml
outputs
checkpoint_root
data/processed/masked-span-l20-specialist-4-span4
report_root
reports/learning/masked-span-l20-specialist-4-span4
