MojiDiff

← experiments

masked-span-l21-specialist-4-long-5a50b51-2681icons-9b9b1699

completed —

Measured behaviour

Held-out negative log likelihood per free token
modelposition-marginal floor
nats per tokenoptimizer step02468050001000015000selected
Table view
optimizer stepmodelposition-marginal floor
06.21904.7651
3004.43744.7651
6004.34584.7651
9004.26914.7651
12004.23034.7651
15004.20364.7651
18004.19804.7651
21004.16834.7651
24004.15514.7651
27004.13564.7651
30004.14114.7651
33004.12604.7651
36004.10544.7651
39004.05524.7651
42004.01084.7651
45003.94644.7651
48003.93144.7651
51003.89874.7651
54003.90524.7651
57003.89964.7651
60003.89084.7651
63003.85524.7651
66003.83974.7651
69003.84064.7651
72003.80234.7651
75003.79944.7651
78003.76834.7651
81003.76374.7651
84003.78384.7651
87003.76074.7651
90003.72784.7651
93003.73374.7651
96003.70294.7651
99003.73254.7651
102003.71274.7651
105003.69974.7651
108003.67164.7651
111003.70834.7651
114003.67524.7651
117003.69904.7651
120003.67014.7651
123003.64534.7651
126003.70054.7651
129003.66624.7651
132003.65004.7651
135003.68064.7651
138003.64434.7651
141003.62494.7651
144003.62284.7651
147003.62884.7651
150003.60664.7651
153003.61064.7651
156003.59314.7651
159003.59504.7651
162003.62704.7651
165003.60994.7651
168003.61894.7651
171003.61684.7651
174003.60984.7651
177003.61954.7651
180003.59874.7651
Masked-token accuracy over unforced positions
modelposition-marginal floor
accuracyoptimizer step00.10.20.3050001000015000selected
Table view
optimizer stepmodelposition-marginal floor
00.05200.1280
3000.13860.1280
6000.14060.1280
9000.14670.1280
12000.15320.1280
15000.16080.1280
18000.15110.1280
21000.16020.1280
24000.15870.1280
27000.16460.1280
30000.15900.1280
33000.16140.1280
36000.16310.1280
39000.16110.1280
42000.16310.1280
45000.17890.1280
48000.16280.1280
51000.16630.1280
54000.17860.1280
57000.17300.1280
60000.17130.1280
63000.17740.1280
66000.17190.1280
69000.16900.1280
72000.18210.1280
75000.17830.1280
78000.17100.1280
81000.17890.1280
84000.19060.1280
87000.18240.1280
90000.18820.1280
93000.18850.1280
96000.18500.1280
99000.17600.1280
102000.18500.1280
105000.19380.1280
108000.19440.1280
111000.18530.1280
114000.18940.1280
117000.19350.1280
120000.19320.1280
123000.19880.1280
126000.19150.1280
129000.18470.1280
132000.19030.1280
135000.18800.1280
138000.19700.1280
141000.19940.1280
144000.19260.1280
147000.19670.1280
150000.19990.1280
153000.19880.1280
156000.20110.1280
159000.20260.1280
162000.20230.1280
165000.20050.1280
168000.20320.1280
171000.19640.1280
174000.20110.1280
177000.20080.1280
180000.19940.1280

Measured result

Verbatim from the run's summary.json.

checkpoint_round_trip
true
checkpoint_sha256
6a847ce6179567b0…
checkpoint_source
None
checkpoint_source_sha256
None
checks
all_valid
true
beats_drop_baseline
false
beats_marginal_baseline
true
median_recovery
false
config_sha256
792fcbd457542727…
continuity
copy_previous
mean_bins_off
25.64749262536873
median_bins_off
17.0
n
339
marginal
mean_bins_off
140.05162241887905
median_bins_off
138.5
n
339
model
mean_bins_off
17.5811209439528
median_bins_off
11.0
n
339
coordinate_tau
1.0
criteria
beats_drop_baseline
true
beats_marginal_baseline
true
min_median_recovery
0.3
require_all_valid
true
deterministic_algorithms
true
device
cuda
evaluation_icons
339
evaluation_split
validation
first_train_loss
6.352187633514404
gpu
NVIDIA GeForce RTX 4080
held_out_nll_per_masked_token
3.593099948154908
icon_presentations
288000
initial_held_out_nll
6.218960549879478
inpainting
all_valid
true
attempted
56
close_reproduction_rate
0.0
decoding
chain_order
true
greedy_for_criteria
true
iterations
8
samples
1
drop
median_rgba_mae
0.005432628298232873
exact_reproduction_rate
0.0
icons
56
identity_policy
join_span
inpaint_seed
3501
marginal
helped
4
mean_recovery
-9243201.564606853
median_recovery
-1.5514862525194055
median_rgba_mae
0.02051230407286371
model
helped
22
mean_recovery
-2958359.976244374
median_recovery
-0.044644347957101746
median_rgba_mae
0.004989428255870249
paired
drop_minus_model
interval
-0.0022500233027615414, 0.0018352586314609215
interval_excludes_zero_above
false
mean
-0.0002073823356503099
n
56
positive
22
marginal_minus_model
interval
0.004811872256668611, 0.015588885461542124
interval_excludes_zero_above
true
mean
0.010200378859105368
n
56
positive
49
rows_sha256
6c4fed0fbf6ea80f…
sheet_sha256
a3f65aeadbaf6096…
span_length
4
task
span
last_train_loss
3.5865390300750732
loss_reduction_factor
1.7308064455799441
marginal_masked_token_accuracy
0.12803273896521486
marginal_nll_per_masked_token
4.7650980387679045
mask_mixture
drawn
span
288000
max_paths
2
random_rate
0.05, 0.95
span_max
4
weights
span
1.0
masked_token_accuracy
0.20111078631978954
metric_head
true
metrics_sha256
eeca047ef8232e16…
model_parameters
531892
nll_ratio_to_marginal
0.7540453352527378
path_binding
false
predeclared_outcome
falsified
schema_version
1
selected_step
15600
sequence_length
1376
start_features
true
steps_run
18000
study_version
masked-span-l21-specialist-4-long
timing
peak_vram_gib
1.9460439682006836
prepare_seconds
102.04778214800172
train_seconds
1199.1558452319878
torch_version
2.14.0a0+4fdf77b940.nv26.08
train_icons
2681
trained_here
true
vocabulary
418

Visual output

inpainting
inpainting.png

State transitions

  1. planned2026-09-21T10:34:28Z
  2. completed2026-09-21T11:07:07Ztrained three times longer, the specialist's held-out likelihood keeps improving - 3.593, 0.754 of the floor, selected at step 15,600 - and its continuity reaches 17.6 bins against the copy policy's 25.6, and the edit still ties the join: mean -0.0002, interval [-0.0023, +0.0018] spanning zero, 22 of 56 helped, indistinguishable span for span from arm 20 (24 better, 32 worse); the marginal policy beaten on 49 of 56; every completion valid

Run record

This run has no run.yaml. What follows is the identity and configuration carried by its rows in state/runs.jsonl, the append-only registry.

run_id
masked-span-l21-specialist-4-long-5a50b51-2681icons-9b9b1699
config
configs/learning/masked-span-l21-specialist-4-long.yaml
outputs
checkpoint_root
data/processed/masked-span-l21-specialist-4-long
report_root
reports/learning/masked-span-l21-specialist-4-long