masked-continuity-l10-start-4c9d1e3-2681icons-9b9b1699
completed —
Measured behaviour
modelposition-marginal floor
Table view
| optimizer step | model | position-marginal floor |
|---|---|---|
| 0 | 5.1856 | 4.7865 |
| 300 | 4.3929 | 4.7865 |
| 600 | 4.2906 | 4.7865 |
| 900 | 4.2283 | 4.7865 |
| 1200 | 4.1834 | 4.7865 |
| 1500 | 4.1580 | 4.7865 |
| 1800 | 4.1412 | 4.7865 |
| 2100 | 4.1194 | 4.7865 |
modelposition-marginal floor
Table view
| optimizer step | model | position-marginal floor |
|---|---|---|
| 0 | 0.0479 | 0.1307 |
| 300 | 0.1318 | 0.1307 |
| 600 | 0.1469 | 0.1307 |
| 900 | 0.1463 | 0.1307 |
| 1200 | 0.1506 | 0.1307 |
| 1500 | 0.1538 | 0.1307 |
| 1800 | 0.1668 | 0.1307 |
| 2100 | 0.1662 | 0.1307 |
Measured result
Verbatim from the run's summary.json.
checkpoint_round_trip
true
checkpoint_sha256
0ba23ad51a8e334d…checks
all_valid
true
beats_copy_previous
false
config_sha256
907a0f4aa201358f…continuity
copy_previous
mean_bins_off
25.64749262536873
median_bins_off
17.0
n
339
marginal
mean_bins_off
140.05162241887905
median_bins_off
138.5
n
339
model
mean_bins_off
47.836283185840706
median_bins_off
38.5
n
339
coordinate_tau
1.0
criteria
beats_copy_previous
true
require_all_valid
true
deterministic_algorithms
true
device
cuda
evaluation_icons
339
evaluation_split
validation
first_train_loss
5.011204242706299
gpu
NVIDIA GeForce RTX 4080
held_out_nll_per_masked_token
4.119351624544748
icon_presentations
33600
initial_held_out_nll
5.185607692642582
inpainting
all_valid
true
attempted
16
decoding
greedy_for_criteria
true
iterations
8
samples
1
drop
median_rgba_mae
0.009685249485596707
exact_reproduction_rate
0.0
icons
16
marginal
helped
2
mean_recovery
-29608927.159026496
median_recovery
-1.3892468653280403
median_rgba_mae
0.026510208635923502
model
helped
1
mean_recovery
-124676619.44439356
median_recovery
-2.1593629049290963
median_rgba_mae
0.018349257141128054
paired
drop_minus_model
interval
-0.04475168936197965, -0.0006796954797656969
interval_excludes_zero_above
false
mean
-0.022715692420872673
n
16
positive
1
marginal_minus_model
interval
-0.036510210316283806, 0.004721127013209498
interval_excludes_zero_above
false
mean
-0.015894541651537156
n
16
positive
8
rows_sha256
a71d06f61fd5bdc4…sheet_sha256
f3cb8537bb3bb8f5…last_train_loss
3.9136998653411865
loss_reduction_factor
1.2588407509921349
marginal_masked_token_accuracy
0.1307154384077461
marginal_nll_per_masked_token
4.786513732286344
mask_mixture
drawn
span
33600
max_paths
2
random_rate
0.05, 0.95
span_max
1
weights
span
1.0
masked_token_accuracy
0.16621839698762775
metric_head
false
metrics_sha256
309158c80b33f99b…model_parameters
530146
nll_ratio_to_marginal
0.8606162762593984
path_binding
false
predeclared_outcome
falsified
schema_version
1
selected_step
2100
sequence_length
1376
start_features
true
steps_run
2100
study_version
masked-continuity-l10-start
timing
peak_vram_gib
1.945967674255371
prepare_seconds
105.03419460298028
train_seconds
137.91324080701452
torch_version
2.14.0a0+4fdf77b940.nv26.08
train_icons
2681
vocabulary
418
Visual output
State transitions
- planned2026-09-21T09:23:39Z
- completed2026-09-21T09:32:43Zthe first arm to move anything: held-out likelihood 4.119 against 4.532 for arm 5 and both head arms, continuity error 47.8 bins (median 38.5) against 58-60, and the held-out curve still falling at the last evaluation where every earlier arm was flat from step 300; still short of the copy policy's 25.6
Run record
This run has no run.yaml. What follows is the identity and configuration carried by its rows in state/runs.jsonl, the append-only registry.
run_id
masked-continuity-l10-start-4c9d1e3-2681icons-9b9b1699
config
configs/learning/masked-continuity-l10-start.yaml
outputs
checkpoint_root
data/processed/masked-continuity-l10-start
report_root
reports/learning/masked-continuity-l10-start
