MojiDiff

← experiments

masked-span-l17-specialist-3-aea91c2-2681icons-9b9b1699

completed —

Measured behaviour

Held-out negative log likelihood per free token
modelposition-marginal floor
nats per tokenoptimizer step024680200040006000selected
Table view
optimizer stepmodelposition-marginal floor
06.24394.7574
3004.41784.7574
6004.32944.7574
9004.25114.7574
12004.20744.7574
15004.15864.7574
18004.14984.7574
21004.13064.7574
24004.11644.7574
27004.10794.7574
30004.10264.7574
33004.07844.7574
36004.04594.7574
39003.99014.7574
42003.95484.7574
45003.91594.7574
48003.84944.7574
51003.80844.7574
54003.75954.7574
57003.75244.7574
60003.75044.7574
63003.72014.7574
Masked-token accuracy over unforced positions
modelposition-marginal floor
accuracyoptimizer step00.050.10.150.20200040006000selected
Table view
optimizer stepmodelposition-marginal floor
00.05250.1303
3000.13880.1303
6000.14510.1303
9000.14830.1303
12000.14930.1303
15000.16540.1303
18000.15750.1303
21000.16610.1303
24000.15690.1303
27000.16210.1303
30000.16020.1303
33000.15690.1303
36000.16410.1303
39000.16670.1303
42000.16540.1303
45000.17070.1303
48000.16740.1303
51000.16340.1303
54000.17390.1303
57000.17720.1303
60000.17490.1303
63000.17720.1303

Measured result

Verbatim from the run's summary.json.

checkpoint_round_trip
true
checkpoint_sha256
baca10153f43debb…
checkpoint_source
None
checkpoint_source_sha256
None
checks
all_valid
true
beats_drop_baseline
true
beats_marginal_baseline
true
median_recovery
false
config_sha256
8e182eaa07a806aa…
continuity
copy_previous
mean_bins_off
25.64749262536873
median_bins_off
17.0
n
339
marginal
mean_bins_off
140.05162241887905
median_bins_off
138.5
n
339
model
mean_bins_off
19.874631268436577
median_bins_off
12.5
n
339
coordinate_tau
1.0
criteria
beats_drop_baseline
true
beats_marginal_baseline
true
min_median_recovery
0.3
require_all_valid
true
deterministic_algorithms
true
device
cuda
evaluation_icons
339
evaluation_split
validation
first_train_loss
6.4128499031066895
gpu
NVIDIA GeForce RTX 4080
held_out_nll_per_masked_token
3.720063364462732
icon_presentations
100800
initial_held_out_nll
6.243883809614541
inpainting
all_valid
true
attempted
64
close_reproduction_rate
0.0625
decoding
chain_order
true
greedy_for_criteria
true
iterations
8
samples
1
drop
median_rgba_mae
0.002642463235294118
exact_reproduction_rate
0.015625
icons
64
identity_policy
join_span
marginal
helped
8
mean_recovery
-99992183.0543034
median_recovery
-2.3124597462758185
median_rgba_mae
0.012874739015976761
model
helped
26
mean_recovery
-2963.2846921082414
median_recovery
-0.02560687842182008
median_rgba_mae
0.001739609809973372
paired
drop_minus_model
interval
0.0004174575220972001, 0.017932260596245815
interval_excludes_zero_above
true
mean
0.009174859059171507
n
64
positive
26
marginal_minus_model
interval
0.0083093994938134, 0.025966874552649323
interval_excludes_zero_above
true
mean
0.017138137023231362
n
64
positive
58
rows_sha256
64f8493453bd945d…
sheet_sha256
8d7adcfea5369673…
span_length
2
task
span
last_train_loss
3.775001049041748
loss_reduction_factor
1.6784348001331182
marginal_masked_token_accuracy
0.13029209058089924
marginal_nll_per_masked_token
4.757360254931849
mask_mixture
drawn
span
100800
max_paths
2
random_rate
0.05, 0.95
span_max
3
weights
span
1.0
masked_token_accuracy
0.1772234985231375
metric_head
true
metrics_sha256
59a92bb4559d1d7f…
model_parameters
531892
nll_ratio_to_marginal
0.7819595668850652
path_binding
false
predeclared_outcome
falsified
schema_version
1
selected_step
6300
sequence_length
1376
start_features
true
steps_run
6300
study_version
masked-span-l17-specialist-3
timing
peak_vram_gib
1.9460439682006836
prepare_seconds
103.20381241000723
train_seconds
418.0104524580238
torch_version
2.14.0a0+4fdf77b940.nv26.08
train_icons
2681
trained_here
true
vocabulary
418

Visual output

inpainting
inpainting.png

State transitions

  1. planned2026-09-21T10:14:57Z
  2. completed2026-09-21T10:26:16Ztrained on spans of one to three segments, the specialist beats the join fill on paired render error - mean +0.0092, interval [+0.0004, +0.0179] excluding zero - and the marginal policy on 58 of 64; 26 of 64 spans helped and the median recovery is -0.03 against the 0.30 bar, so the run is falsified as predeclared while both paired tests pass; every completion valid; continuity 19.9 bins against the copy policy's 25.6

Run record

This run has no run.yaml. What follows is the identity and configuration carried by its rows in state/runs.jsonl, the append-only registry.

run_id
masked-span-l17-specialist-3-aea91c2-2681icons-9b9b1699
config
configs/learning/masked-span-l17-specialist-3.yaml
outputs
checkpoint_root
data/processed/masked-span-l17-specialist-3
report_root
reports/learning/masked-span-l17-specialist-3