masked-continuity-l7-metric-head-5475283-2681icons-9b9b1699
completed —
Measured behaviour
modelposition-marginal floor
Table view
| optimizer step | model | position-marginal floor |
|---|---|---|
| 0 | 5.1984 | 4.7865 |
| 300 | 4.6120 | 4.7865 |
| 600 | 4.5955 | 4.7865 |
| 900 | 4.5885 | 4.7865 |
| 1200 | 4.5715 | 4.7865 |
| 1500 | 4.5521 | 4.7865 |
| 1800 | 4.5364 | 4.7865 |
| 2100 | 4.5318 | 4.7865 |
modelposition-marginal floor
Table view
| optimizer step | model | position-marginal floor |
|---|---|---|
| 0 | 0.0559 | 0.1307 |
| 300 | 0.1189 | 0.1307 |
| 600 | 0.1458 | 0.1307 |
| 900 | 0.1415 | 0.1307 |
| 1200 | 0.1447 | 0.1307 |
| 1500 | 0.1420 | 0.1307 |
| 1800 | 0.1490 | 0.1307 |
| 2100 | 0.1447 | 0.1307 |
Measured result
Verbatim from the run's summary.json.
checkpoint_round_trip
true
checkpoint_sha256
863e99743684c78b…checks
all_valid
true
beats_copy_previous
false
config_sha256
11296fcaa4384b3b…continuity
copy_previous
mean_bins_off
25.64749262536873
median_bins_off
17.0
n
339
marginal
mean_bins_off
140.05162241887905
median_bins_off
138.5
n
339
model
mean_bins_off
64.60324483775811
median_bins_off
60.5
n
339
coordinate_tau
1.0
criteria
beats_copy_previous
true
require_all_valid
true
deterministic_algorithms
true
device
cuda
evaluation_icons
339
evaluation_split
validation
first_train_loss
4.990401268005371
gpu
NVIDIA GeForce RTX 4080
held_out_nll_per_masked_token
4.531823072592503
icon_presentations
33600
initial_held_out_nll
5.198353921809204
inpainting
all_valid
true
attempted
16
decoding
greedy_for_criteria
true
iterations
8
samples
1
drop
median_rgba_mae
0.009685249485596707
exact_reproduction_rate
0.0
icons
16
marginal
helped
2
mean_recovery
-29608927.159026496
median_recovery
-1.3892468653280403
median_rgba_mae
0.026510208635923502
model
helped
0
mean_recovery
-1.5068213827303556
median_recovery
-0.022224810486944607
median_rgba_mae
0.009822266097797143
paired
drop_minus_model
interval
-0.0050247677741438995, 0.0002545519138800402
interval_excludes_zero_above
false
mean
-0.00238510793013193
n
16
positive
0
marginal_minus_model
interval
-0.01024353672312386, 0.01911562240153103
interval_excludes_zero_above
false
mean
0.0044360428392035845
n
16
positive
14
rows_sha256
a471327051078a5e…sheet_sha256
6b670c834af2bfd0…last_train_loss
4.410138130187988
loss_reduction_factor
1.1470778621627435
marginal_masked_token_accuracy
0.1307154384077461
marginal_nll_per_masked_token
4.786513732286344
mask_mixture
drawn
span
33600
max_paths
2
random_rate
0.05, 0.95
span_max
1
weights
span
1.0
masked_token_accuracy
0.1447014523937601
metric_head
true
metrics_sha256
fde290171ab5a42d…model_parameters
528244
nll_ratio_to_marginal
0.9467899448452675
path_binding
false
predeclared_outcome
falsified
schema_version
1
selected_step
2100
sequence_length
1376
steps_run
2100
study_version
masked-continuity-l7-metric-head
timing
peak_vram_gib
1.9458246231079102
prepare_seconds
102.21235602800152
train_seconds
131.47408982599154
torch_version
2.14.0a0+4fdf77b940.nv26.08
train_icons
2681
vocabulary
418
Visual output
State transitions
- planned2026-09-21T09:03:05Z
- completed2026-09-21T09:11:48Zthe held-out trace is identical to arm 5's to four decimals at every step, so the change under test never took effect: the metric head decides a position's role from the kind token in the input, and under span and path masks the kind is hidden with its coordinates, so every target position used the categorical head; the probe, where the kind is visible, then ran an untrained projection and scored 64.6 bins
Run record
This run has no run.yaml. What follows is the identity and configuration carried by its rows in state/runs.jsonl, the append-only registry.
run_id
masked-continuity-l7-metric-head-5475283-2681icons-9b9b1699
config
configs/learning/masked-continuity-l7-metric-head.yaml
outputs
checkpoint_root
data/processed/masked-continuity-l7-metric-head
report_root
reports/learning/masked-continuity-l7-metric-head
