masked-span-l15-mixture-bde5d24-2681icons-9b9b1699
completed —
Measured behaviour
modelposition-marginal floor
Table view
| optimizer step | model | position-marginal floor |
|---|---|---|
| 0 | 5.2606 | 4.1110 |
| 4200 | 3.8755 | 4.1110 |
modelposition-marginal floor
Table view
| optimizer step | model | position-marginal floor |
|---|---|---|
| 0 | 0.0660 | 0.2413 |
| 4200 | 0.2617 | 0.2413 |
Measured result
Verbatim from the run's summary.json.
checkpoint_round_trip
true
checkpoint_sha256
baa5647f55cb6fc2…checkpoint_source
data/processed/masked-inpaint-l14-start-head/checkpoint.zip
checkpoint_source_sha256
ec454fbde6509b8b…checks
all_valid
true
beats_drop_baseline
false
beats_marginal_baseline
true
median_recovery
false
config_sha256
37aa3d2312507616…continuity
copy_previous
mean_bins_off
25.64749262536873
median_bins_off
17.0
n
339
marginal
mean_bins_off
140.05162241887905
median_bins_off
138.5
n
339
model
mean_bins_off
41.79203539823009
median_bins_off
36.0
n
339
coordinate_tau
1.0
criteria
beats_drop_baseline
true
beats_marginal_baseline
true
min_median_recovery
0.3
require_all_valid
true
deterministic_algorithms
true
device
cuda
evaluation_icons
339
evaluation_split
validation
first_train_loss
None
gpu
NVIDIA GeForce RTX 4080
held_out_nll_per_masked_token
3.875509736467909
icon_presentations
None
initial_held_out_nll
5.260572421212839
inpainting
all_valid
true
attempted
64
close_reproduction_rate
0.0
decoding
chain_order
true
greedy_for_criteria
true
iterations
8
samples
1
drop
median_rgba_mae
0.0014610377329944322
exact_reproduction_rate
0.0
icons
64
identity_policy
join_span
marginal
helped
4
mean_recovery
-253049919.98605102
median_recovery
-2.743950679036371
median_rgba_mae
0.01059368191721133
model
helped
9
mean_recovery
-45752014.78243475
median_recovery
-1.3453651432695155
median_rgba_mae
0.00561181841563786
paired
drop_minus_model
interval
-0.007376429164583229, -0.0022944000081352406
interval_excludes_zero_above
false
mean
-0.004835414586359235
n
64
positive
9
marginal_minus_model
interval
0.00016999600779433023, 0.004463120500768972
interval_excludes_zero_above
true
mean
0.002316558254281651
n
64
positive
40
rows_sha256
6455a03b8e1d2078…sheet_sha256
0fa36c0ed127b975…span_length
2
task
span
last_train_loss
None
loss_reduction_factor
None
marginal_masked_token_accuracy
0.2413130873428418
marginal_nll_per_masked_token
4.1110007453812765
mask_mixture
drawn
max_paths
2
random_rate
0.05, 0.95
span_max
None
weights
geometry
0.1
path
0.3
random
0.3
span
0.2
style
0.1
masked_token_accuracy
0.2616700381532789
metric_head
true
metrics_sha256
f302878ced46c748…model_parameters
531892
nll_ratio_to_marginal
0.9427168654304083
path_binding
false
predeclared_outcome
falsified
schema_version
1
selected_step
4200
sequence_length
1376
start_features
true
steps_run
4200
study_version
masked-span-l15-mixture
timing
peak_vram_gib
1.8976726531982422
prepare_seconds
102.80136108299484
train_seconds
0.897731217002729
torch_version
2.14.0a0+4fdf77b940.nv26.08
train_icons
2681
trained_here
false
vocabulary
418
Visual output
State transitions
- planned2026-09-21T10:08:04Z
- completed2026-09-21T10:14:56Zarm 14's mixture model, read on two-segment spans, is worse than joining the visible ends: mean -0.0048 RGBA MAE with the interval excluding zero, 9 of 64 helped, median recovery -1.35; it beats the marginal policy narrowly (40 of 64); every completion valid
Run record
This run has no run.yaml. What follows is the identity and configuration carried by its rows in state/runs.jsonl, the append-only registry.
run_id
masked-span-l15-mixture-bde5d24-2681icons-9b9b1699
config
configs/learning/masked-span-l15-mixture.yaml
outputs
report_root
reports/learning/masked-span-l15-mixture
