masked-span-l18-specialist-3-l16-spans-f7fd711-2681icons-9b9b1699
completed —
Measured behaviour
modelposition-marginal floor
Table view
| optimizer step | model | position-marginal floor |
|---|---|---|
| 0 | 5.2606 | 4.1110 |
| 6300 | 4.6358 | 4.1110 |
modelposition-marginal floor
Table view
| optimizer step | model | position-marginal floor |
|---|---|---|
| 0 | 0.0660 | 0.2413 |
| 6300 | 0.1038 | 0.2413 |
Measured result
Verbatim from the run's summary.json.
checkpoint_round_trip
true
checkpoint_sha256
ffacae3e96bc85f3…checkpoint_source
data/processed/masked-span-l17-specialist-3/checkpoint.zip
checkpoint_source_sha256
baca10153f43debb…checks
all_valid
true
beats_drop_baseline
false
beats_marginal_baseline
true
median_recovery
false
config_sha256
272a40bf1aa73954…continuity
copy_previous
mean_bins_off
25.64749262536873
median_bins_off
17.0
n
339
marginal
mean_bins_off
140.05162241887905
median_bins_off
138.5
n
339
model
mean_bins_off
19.874631268436577
median_bins_off
12.5
n
339
coordinate_tau
1.0
criteria
beats_drop_baseline
true
beats_marginal_baseline
true
min_median_recovery
0.3
require_all_valid
true
deterministic_algorithms
true
device
cuda
evaluation_icons
339
evaluation_split
validation
first_train_loss
None
gpu
NVIDIA GeForce RTX 4080
held_out_nll_per_masked_token
4.635817549578057
icon_presentations
None
initial_held_out_nll
5.260572421212839
inpainting
all_valid
true
attempted
64
close_reproduction_rate
0.0625
decoding
chain_order
true
greedy_for_criteria
true
iterations
8
samples
1
drop
median_rgba_mae
0.0014610377329944322
exact_reproduction_rate
0.03125
icons
64
identity_policy
join_span
marginal
helped
4
mean_recovery
-253049919.98605102
median_recovery
-2.743950679036371
median_rgba_mae
0.01059368191721133
model
helped
30
mean_recovery
-13616562.376142139
median_recovery
0.0
median_rgba_mae
0.0017904827826192204
paired
drop_minus_model
interval
-0.0016157232228196628, 0.0030064501106040974
interval_excludes_zero_above
false
mean
0.0006953634438922173
n
64
positive
30
marginal_minus_model
interval
0.00487410411374075, 0.010820568455325457
interval_excludes_zero_above
true
mean
0.007847336284533104
n
64
positive
58
rows_sha256
2bfa1d25034518ec…sheet_sha256
5a21f819a9aefce7…span_length
2
task
span
last_train_loss
None
loss_reduction_factor
None
marginal_masked_token_accuracy
0.2413130873428418
marginal_nll_per_masked_token
4.1110007453812765
mask_mixture
drawn
max_paths
2
random_rate
0.05, 0.95
span_max
None
weights
geometry
0.1
path
0.3
random
0.3
span
0.2
style
0.1
masked_token_accuracy
0.10375643356464292
metric_head
true
metrics_sha256
a7ac9b387572a53d…model_parameters
531892
nll_ratio_to_marginal
1.1276615687278613
path_binding
false
predeclared_outcome
falsified
schema_version
1
selected_step
6300
sequence_length
1376
start_features
true
steps_run
6300
study_version
masked-span-l18-specialist-3-l16-spans
timing
peak_vram_gib
1.8976726531982422
prepare_seconds
103.75202324200654
train_seconds
0.8970639029867016
torch_version
2.14.0a0+4fdf77b940.nv26.08
train_icons
2681
trained_here
false
vocabulary
418
Visual output
State transitions
- planned2026-09-21T10:26:18Z
- completed2026-09-21T10:30:41Zread on arm 16's exact spans, arm 17's checkpoint ties the join fill - mean +0.0007, interval [-0.0016, +0.0030] spanning zero, 30 of 64 helped, median recovery 0.0 - and ties the single-segment specialist span for span (31 better, 30 worse, 3 equal); it beats the marginal policy on 58 of 64 with the interval excluding zero; every completion valid
Run record
This run has no run.yaml. What follows is the identity and configuration carried by its rows in state/runs.jsonl, the append-only registry.
run_id
masked-span-l18-specialist-3-l16-spans-f7fd711-2681icons-9b9b1699
config
configs/learning/masked-span-l18-specialist-3-l16-spans.yaml
outputs
report_root
reports/learning/masked-span-l18-specialist-3-l16-spans
