masked-span-l16-specialist-bde5d24-2681icons-9b9b1699
completed —
Measured behaviour
modelposition-marginal floor
Table view
| optimizer step | model | position-marginal floor |
|---|---|---|
| 0 | 5.2606 | 4.1110 |
| 6300 | 4.6595 | 4.1110 |
modelposition-marginal floor
Table view
| optimizer step | model | position-marginal floor |
|---|---|---|
| 0 | 0.0660 | 0.2413 |
| 6300 | 0.1107 | 0.2413 |
Measured result
Verbatim from the run's summary.json.
checkpoint_round_trip
true
checkpoint_sha256
bc0d3ace435a1e55…checkpoint_source
data/processed/masked-continuity-l13-start-head-full/checkpoint.zip
checkpoint_source_sha256
3b9e15b39acbe4fe…checks
all_valid
true
beats_drop_baseline
false
beats_marginal_baseline
true
median_recovery
false
config_sha256
599bd5f16ea56970…continuity
copy_previous
mean_bins_off
25.64749262536873
median_bins_off
17.0
n
339
marginal
mean_bins_off
140.05162241887905
median_bins_off
138.5
n
339
model
mean_bins_off
17.86283185840708
median_bins_off
10.5
n
339
coordinate_tau
1.0
criteria
beats_drop_baseline
true
beats_marginal_baseline
true
min_median_recovery
0.3
require_all_valid
true
deterministic_algorithms
true
device
cuda
evaluation_icons
339
evaluation_split
validation
first_train_loss
None
gpu
NVIDIA GeForce RTX 4080
held_out_nll_per_masked_token
4.6594607245982225
icon_presentations
None
initial_held_out_nll
5.260572421212839
inpainting
all_valid
true
attempted
64
close_reproduction_rate
0.015625
decoding
chain_order
true
greedy_for_criteria
true
iterations
8
samples
1
drop
median_rgba_mae
0.0014610377329944322
exact_reproduction_rate
0.015625
icons
64
identity_policy
join_span
marginal
helped
4
mean_recovery
-253049919.98605102
median_recovery
-2.743950679036371
median_rgba_mae
0.01059368191721133
model
helped
26
mean_recovery
-7183562.797062205
median_recovery
-0.03515278425328994
median_rgba_mae
0.0011201509924957638
paired
drop_minus_model
interval
-0.0003637274372127817, 0.0026232486260895672
interval_excludes_zero_above
false
mean
0.0011297605944383927
n
64
positive
26
marginal_minus_model
interval
0.005864020789901072, 0.010699446080257486
interval_excludes_zero_above
true
mean
0.00828173343507928
n
64
positive
60
rows_sha256
fb0509cf7958ca28…sheet_sha256
e8f5ba681b223e9f…span_length
2
task
span
last_train_loss
None
loss_reduction_factor
None
marginal_masked_token_accuracy
0.2413130873428418
marginal_nll_per_masked_token
4.1110007453812765
mask_mixture
drawn
max_paths
2
random_rate
0.05, 0.95
span_max
None
weights
geometry
0.1
path
0.3
random
0.3
span
0.2
style
0.1
masked_token_accuracy
0.11067011497196118
metric_head
true
metrics_sha256
15e15f3d6f770a70…model_parameters
531892
nll_ratio_to_marginal
1.1334127656953463
path_binding
false
predeclared_outcome
falsified
schema_version
1
selected_step
6300
sequence_length
1376
start_features
true
steps_run
6300
study_version
masked-span-l16-specialist
timing
peak_vram_gib
1.8976726531982422
prepare_seconds
101.6064741309965
train_seconds
0.8970100769947749
torch_version
2.14.0a0+4fdf77b940.nv26.08
train_icons
2681
trained_here
false
vocabulary
418
Visual output
State transitions
- planned2026-09-21T10:08:04Z
- completed2026-09-21T10:14:56Zarm 13's span specialist, trained on single-segment masks and read on two-segment spans it never saw, is at parity with joining the visible ends: mean +0.0011 with the interval [-0.0004, +0.0026] spanning zero, 26 of 64 helped, median recovery -0.035, its median error 0.00112 below the join's 0.00146; it beats the marginal policy on 60 of 64 with the interval excluding zero; every completion valid; the fills are indistinguishable from the clean icons in nearly every row of the sheet
Run record
This run has no run.yaml. What follows is the identity and configuration carried by its rows in state/runs.jsonl, the append-only registry.
run_id
masked-span-l16-specialist-bde5d24-2681icons-9b9b1699
config
configs/learning/masked-span-l16-specialist.yaml
outputs
report_root
reports/learning/masked-span-l16-specialist
