masked-span-l19-specialist-3-span4-9202129-2681icons-9b9b1699
completed —
Measured behaviour
modelposition-marginal floor
Table view
| optimizer step | model | position-marginal floor |
|---|---|---|
| 0 | 5.2606 | 4.1110 |
| 6300 | 4.6358 | 4.1110 |
modelposition-marginal floor
Table view
| optimizer step | model | position-marginal floor |
|---|---|---|
| 0 | 0.0660 | 0.2413 |
| 6300 | 0.1038 | 0.2413 |
Measured result
Verbatim from the run's summary.json.
checkpoint_round_trip
true
checkpoint_sha256
ffacae3e96bc85f3…checkpoint_source
data/processed/masked-span-l17-specialist-3/checkpoint.zip
checkpoint_source_sha256
baca10153f43debb…checks
all_valid
true
beats_drop_baseline
false
beats_marginal_baseline
true
median_recovery
false
config_sha256
f5d236f54ffda0e2…continuity
copy_previous
mean_bins_off
25.64749262536873
median_bins_off
17.0
n
339
marginal
mean_bins_off
140.05162241887905
median_bins_off
138.5
n
339
model
mean_bins_off
19.874631268436577
median_bins_off
12.5
n
339
coordinate_tau
1.0
criteria
beats_drop_baseline
true
beats_marginal_baseline
true
min_median_recovery
0.3
require_all_valid
true
deterministic_algorithms
true
device
cuda
evaluation_icons
339
evaluation_split
validation
first_train_loss
None
gpu
NVIDIA GeForce RTX 4080
held_out_nll_per_masked_token
4.635817549578057
icon_presentations
None
initial_held_out_nll
5.260572421212839
inpainting
all_valid
true
attempted
56
close_reproduction_rate
0.0
decoding
chain_order
true
greedy_for_criteria
true
iterations
8
samples
1
drop
median_rgba_mae
0.005432628298232873
exact_reproduction_rate
0.0
icons
56
identity_policy
join_span
inpaint_seed
3501
marginal
helped
4
mean_recovery
-9243201.564606853
median_recovery
-1.5514862525194055
median_rgba_mae
0.02051230407286371
model
helped
26
mean_recovery
-10131.962719963649
median_recovery
-0.013909173659141304
median_rgba_mae
0.0045958719135802475
paired
drop_minus_model
interval
-0.0010316583610373577, 0.0014010142438225708
interval_excludes_zero_above
false
mean
0.0001846779413926065
n
56
positive
26
marginal_minus_model
interval
0.005256827214532227, 0.015928051057764344
interval_excludes_zero_above
true
mean
0.010592439136148286
n
56
positive
49
rows_sha256
29ac520a207414c6…sheet_sha256
da21fcb35f728099…span_length
4
task
span
last_train_loss
None
loss_reduction_factor
None
marginal_masked_token_accuracy
0.2413130873428418
marginal_nll_per_masked_token
4.1110007453812765
mask_mixture
drawn
max_paths
2
random_rate
0.05, 0.95
span_max
None
weights
geometry
0.1
path
0.3
random
0.3
span
0.2
style
0.1
masked_token_accuracy
0.10375643356464292
metric_head
true
metrics_sha256
a7ac9b387572a53d…model_parameters
531892
nll_ratio_to_marginal
1.1276615687278613
path_binding
false
predeclared_outcome
falsified
schema_version
1
selected_step
6300
sequence_length
1376
start_features
true
steps_run
6300
study_version
masked-span-l19-specialist-3-span4
timing
peak_vram_gib
1.8976726531982422
prepare_seconds
101.67533485300373
train_seconds
0.8960046049905941
torch_version
2.14.0a0+4fdf77b940.nv26.08
train_icons
2681
trained_here
false
vocabulary
418
Visual output
State transitions
- planned2026-09-21T10:30:42Z
- completed2026-09-21T10:34:27Zread on four-segment spans of 56 held-out icons with a path long enough, the specialist trained on spans of one to three ties the join fill - mean +0.0002, interval [-0.0010, +0.0014] spanning zero, 26 of 56 helped, its median error 0.0046 below the join's 0.0054 - and beats the marginal policy on 49 of 56 with the interval excluding zero; every completion valid
Run record
This run has no run.yaml. What follows is the identity and configuration carried by its rows in state/runs.jsonl, the append-only registry.
run_id
masked-span-l19-specialist-3-span4-9202129-2681icons-9b9b1699
config
configs/learning/masked-span-l19-specialist-3-span4.yaml
outputs
report_root
reports/learning/masked-span-l19-specialist-3-span4
