masked-span-l22-specialist-4-large-a83419d-2681icons-9b9b1699
completed —
Measured behaviour
modelposition-marginal floor
Table view
| optimizer step | model | position-marginal floor |
|---|---|---|
| 0 | 6.2408 | 4.7651 |
| 300 | 4.4632 | 4.7651 |
| 600 | 4.3202 | 4.7651 |
| 900 | 4.2758 | 4.7651 |
| 1200 | 4.2417 | 4.7651 |
| 1500 | 4.1717 | 4.7651 |
| 1800 | 4.2091 | 4.7651 |
| 2100 | 4.1518 | 4.7651 |
| 2400 | 4.1412 | 4.7651 |
| 2700 | 4.0684 | 4.7651 |
| 3000 | 4.0098 | 4.7651 |
| 3300 | 3.9611 | 4.7651 |
| 3600 | 3.9325 | 4.7651 |
| 3900 | 3.9313 | 4.7651 |
| 4200 | 3.9028 | 4.7651 |
| 4500 | 3.8827 | 4.7651 |
| 4800 | 3.8610 | 4.7651 |
| 5100 | 3.8598 | 4.7651 |
| 5400 | 3.8310 | 4.7651 |
| 5700 | 3.7946 | 4.7651 |
| 6000 | 3.7967 | 4.7651 |
| 6300 | 3.7753 | 4.7651 |
modelposition-marginal floor
Table view
| optimizer step | model | position-marginal floor |
|---|---|---|
| 0 | 0.0734 | 0.1280 |
| 300 | 0.1377 | 0.1280 |
| 600 | 0.1397 | 0.1280 |
| 900 | 0.1485 | 0.1280 |
| 1200 | 0.1523 | 0.1280 |
| 1500 | 0.1570 | 0.1280 |
| 1800 | 0.1502 | 0.1280 |
| 2100 | 0.1608 | 0.1280 |
| 2400 | 0.1584 | 0.1280 |
| 2700 | 0.1713 | 0.1280 |
| 3000 | 0.1625 | 0.1280 |
| 3300 | 0.1570 | 0.1280 |
| 3600 | 0.1643 | 0.1280 |
| 3900 | 0.1695 | 0.1280 |
| 4200 | 0.1666 | 0.1280 |
| 4500 | 0.1757 | 0.1280 |
| 4800 | 0.1646 | 0.1280 |
| 5100 | 0.1801 | 0.1280 |
| 5400 | 0.1766 | 0.1280 |
| 5700 | 0.1748 | 0.1280 |
| 6000 | 0.1663 | 0.1280 |
| 6300 | 0.1757 | 0.1280 |
Measured result
Verbatim from the run's summary.json.
checkpoint_round_trip
true
checkpoint_sha256
d402223ff9f3c943…checkpoint_source
None
checkpoint_source_sha256
None
checks
all_valid
true
beats_drop_baseline
false
beats_marginal_baseline
true
median_recovery
false
config_sha256
f1b3aa67cf3ff306…continuity
copy_previous
mean_bins_off
25.64749262536873
median_bins_off
17.0
n
339
marginal
mean_bins_off
140.05162241887905
median_bins_off
138.5
n
339
model
mean_bins_off
19.775811209439528
median_bins_off
13.0
n
339
coordinate_tau
1.0
criteria
beats_drop_baseline
true
beats_marginal_baseline
true
min_median_recovery
0.3
require_all_valid
true
deterministic_algorithms
true
device
cuda
evaluation_icons
339
evaluation_split
validation
first_train_loss
5.863951206207275
gpu
NVIDIA GeForce RTX 4080
held_out_nll_per_masked_token
3.77532209871945
icon_presentations
100800
initial_held_out_nll
6.240840703344805
inpainting
all_valid
true
attempted
56
close_reproduction_rate
0.0
decoding
chain_order
true
greedy_for_criteria
true
iterations
8
samples
1
drop
median_rgba_mae
0.005432628298232873
exact_reproduction_rate
0.0
icons
56
identity_policy
join_span
inpaint_seed
3501
marginal
helped
4
mean_recovery
-9243201.564606853
median_recovery
-1.5514862525194055
median_rgba_mae
0.02051230407286371
model
helped
28
mean_recovery
-2269426.702740625
median_recovery
-0.0012699399968391285
median_rgba_mae
0.0045676932038247395
paired
drop_minus_model
interval
-0.003170476520215973, 0.0013023401838913196
interval_excludes_zero_above
false
mean
-0.0009340681681623268
n
56
positive
28
marginal_minus_model
interval
0.0043722298206685046, 0.014575156232518205
interval_excludes_zero_above
true
mean
0.009473693026593354
n
56
positive
49
rows_sha256
d8903420d9914cc2…sheet_sha256
bd0c1891df4862a4…span_length
4
task
span
last_train_loss
3.9231743812561035
loss_reduction_factor
1.6530617892077695
marginal_masked_token_accuracy
0.12803273896521486
marginal_nll_per_masked_token
4.7650980387679045
mask_mixture
drawn
span
100800
max_paths
2
random_rate
0.05, 0.95
span_max
4
weights
span
1.0
masked_token_accuracy
0.17567962584039754
metric_head
true
metrics_sha256
dea0e16f36be21a3…model_parameters
5358516
nll_ratio_to_marginal
0.7922863429050502
path_binding
false
predeclared_outcome
falsified
schema_version
1
selected_step
6300
sequence_length
1376
start_features
true
steps_run
6300
study_version
masked-span-l22-specialist-4-large
timing
peak_vram_gib
3.903975486755371
prepare_seconds
101.63444957701722
train_seconds
1236.324762987002
torch_version
2.14.0a0+4fdf77b940.nv26.08
train_icons
2681
trained_here
true
vocabulary
418
Visual output
State transitions
- planned2026-09-21T10:44:17Z
- completed2026-09-21T11:31:23Zat 5,358,516 parameters with dropout 0.1 - ten times arm 20 - the specialist ties the join fill on four-segment spans: mean -0.0009, interval [-0.0032, +0.0013] spanning zero, 28 of 56 helped, median recovery -0.001; the marginal policy beaten on 49 of 56; span for span against arm 20, 28 better and 27 worse; continuity 19.8 bins; every completion valid; 1,236 s to train, 3.9 GiB peak
Run record
This run has no run.yaml. What follows is the identity and configuration carried by its rows in state/runs.jsonl, the append-only registry.
run_id
masked-span-l22-specialist-4-large-a83419d-2681icons-9b9b1699
config
configs/learning/masked-span-l22-specialist-4-large.yaml
outputs
checkpoint_root
data/processed/masked-span-l22-specialist-4-large
report_root
reports/learning/masked-span-l22-specialist-4-large
