masked-span-l17-specialist-3-aea91c2-2681icons-9b9b1699
completed —
Measured behaviour
modelposition-marginal floor
Table view
| optimizer step | model | position-marginal floor |
|---|---|---|
| 0 | 6.2439 | 4.7574 |
| 300 | 4.4178 | 4.7574 |
| 600 | 4.3294 | 4.7574 |
| 900 | 4.2511 | 4.7574 |
| 1200 | 4.2074 | 4.7574 |
| 1500 | 4.1586 | 4.7574 |
| 1800 | 4.1498 | 4.7574 |
| 2100 | 4.1306 | 4.7574 |
| 2400 | 4.1164 | 4.7574 |
| 2700 | 4.1079 | 4.7574 |
| 3000 | 4.1026 | 4.7574 |
| 3300 | 4.0784 | 4.7574 |
| 3600 | 4.0459 | 4.7574 |
| 3900 | 3.9901 | 4.7574 |
| 4200 | 3.9548 | 4.7574 |
| 4500 | 3.9159 | 4.7574 |
| 4800 | 3.8494 | 4.7574 |
| 5100 | 3.8084 | 4.7574 |
| 5400 | 3.7595 | 4.7574 |
| 5700 | 3.7524 | 4.7574 |
| 6000 | 3.7504 | 4.7574 |
| 6300 | 3.7201 | 4.7574 |
modelposition-marginal floor
Table view
| optimizer step | model | position-marginal floor |
|---|---|---|
| 0 | 0.0525 | 0.1303 |
| 300 | 0.1388 | 0.1303 |
| 600 | 0.1451 | 0.1303 |
| 900 | 0.1483 | 0.1303 |
| 1200 | 0.1493 | 0.1303 |
| 1500 | 0.1654 | 0.1303 |
| 1800 | 0.1575 | 0.1303 |
| 2100 | 0.1661 | 0.1303 |
| 2400 | 0.1569 | 0.1303 |
| 2700 | 0.1621 | 0.1303 |
| 3000 | 0.1602 | 0.1303 |
| 3300 | 0.1569 | 0.1303 |
| 3600 | 0.1641 | 0.1303 |
| 3900 | 0.1667 | 0.1303 |
| 4200 | 0.1654 | 0.1303 |
| 4500 | 0.1707 | 0.1303 |
| 4800 | 0.1674 | 0.1303 |
| 5100 | 0.1634 | 0.1303 |
| 5400 | 0.1739 | 0.1303 |
| 5700 | 0.1772 | 0.1303 |
| 6000 | 0.1749 | 0.1303 |
| 6300 | 0.1772 | 0.1303 |
Measured result
Verbatim from the run's summary.json.
checkpoint_round_trip
true
checkpoint_sha256
baca10153f43debb…checkpoint_source
None
checkpoint_source_sha256
None
checks
all_valid
true
beats_drop_baseline
true
beats_marginal_baseline
true
median_recovery
false
config_sha256
8e182eaa07a806aa…continuity
copy_previous
mean_bins_off
25.64749262536873
median_bins_off
17.0
n
339
marginal
mean_bins_off
140.05162241887905
median_bins_off
138.5
n
339
model
mean_bins_off
19.874631268436577
median_bins_off
12.5
n
339
coordinate_tau
1.0
criteria
beats_drop_baseline
true
beats_marginal_baseline
true
min_median_recovery
0.3
require_all_valid
true
deterministic_algorithms
true
device
cuda
evaluation_icons
339
evaluation_split
validation
first_train_loss
6.4128499031066895
gpu
NVIDIA GeForce RTX 4080
held_out_nll_per_masked_token
3.720063364462732
icon_presentations
100800
initial_held_out_nll
6.243883809614541
inpainting
all_valid
true
attempted
64
close_reproduction_rate
0.0625
decoding
chain_order
true
greedy_for_criteria
true
iterations
8
samples
1
drop
median_rgba_mae
0.002642463235294118
exact_reproduction_rate
0.015625
icons
64
identity_policy
join_span
marginal
helped
8
mean_recovery
-99992183.0543034
median_recovery
-2.3124597462758185
median_rgba_mae
0.012874739015976761
model
helped
26
mean_recovery
-2963.2846921082414
median_recovery
-0.02560687842182008
median_rgba_mae
0.001739609809973372
paired
drop_minus_model
interval
0.0004174575220972001, 0.017932260596245815
interval_excludes_zero_above
true
mean
0.009174859059171507
n
64
positive
26
marginal_minus_model
interval
0.0083093994938134, 0.025966874552649323
interval_excludes_zero_above
true
mean
0.017138137023231362
n
64
positive
58
rows_sha256
64f8493453bd945d…sheet_sha256
8d7adcfea5369673…span_length
2
task
span
last_train_loss
3.775001049041748
loss_reduction_factor
1.6784348001331182
marginal_masked_token_accuracy
0.13029209058089924
marginal_nll_per_masked_token
4.757360254931849
mask_mixture
drawn
span
100800
max_paths
2
random_rate
0.05, 0.95
span_max
3
weights
span
1.0
masked_token_accuracy
0.1772234985231375
metric_head
true
metrics_sha256
59a92bb4559d1d7f…model_parameters
531892
nll_ratio_to_marginal
0.7819595668850652
path_binding
false
predeclared_outcome
falsified
schema_version
1
selected_step
6300
sequence_length
1376
start_features
true
steps_run
6300
study_version
masked-span-l17-specialist-3
timing
peak_vram_gib
1.9460439682006836
prepare_seconds
103.20381241000723
train_seconds
418.0104524580238
torch_version
2.14.0a0+4fdf77b940.nv26.08
train_icons
2681
trained_here
true
vocabulary
418
Visual output
State transitions
- planned2026-09-21T10:14:57Z
- completed2026-09-21T10:26:16Ztrained on spans of one to three segments, the specialist beats the join fill on paired render error - mean +0.0092, interval [+0.0004, +0.0179] excluding zero - and the marginal policy on 58 of 64; 26 of 64 spans helped and the median recovery is -0.03 against the 0.30 bar, so the run is falsified as predeclared while both paired tests pass; every completion valid; continuity 19.9 bins against the copy policy's 25.6
Run record
This run has no run.yaml. What follows is the identity and configuration carried by its rows in state/runs.jsonl, the append-only registry.
run_id
masked-span-l17-specialist-3-aea91c2-2681icons-9b9b1699
config
configs/learning/masked-span-l17-specialist-3.yaml
outputs
checkpoint_root
data/processed/masked-span-l17-specialist-3
report_root
reports/learning/masked-span-l17-specialist-3
