MojiDiff

← experiments

masked-span-l20-specialist-4-span4-9202129-2681icons-9b9b1699

completed —

Measured behaviour

Held-out negative log likelihood per free token
modelposition-marginal floor
nats per tokenoptimizer step024680200040006000selected
Table view
optimizer stepmodelposition-marginal floor
06.21904.7651
3004.43744.7651
6004.34584.7651
9004.26914.7651
12004.23034.7651
15004.20364.7651
18004.19804.7651
21004.16834.7651
24004.15514.7651
27004.13564.7651
30004.14114.7651
33004.12604.7651
36004.10544.7651
39004.05524.7651
42004.01084.7651
45003.94644.7651
48003.93144.7651
51003.89874.7651
54003.90524.7651
57003.89964.7651
60003.89084.7651
63003.85524.7651
Masked-token accuracy over unforced positions
modelposition-marginal floor
accuracyoptimizer step00.050.10.150.20200040006000selected
Table view
optimizer stepmodelposition-marginal floor
00.05200.1280
3000.13860.1280
6000.14060.1280
9000.14670.1280
12000.15320.1280
15000.16080.1280
18000.15110.1280
21000.16020.1280
24000.15870.1280
27000.16460.1280
30000.15900.1280
33000.16140.1280
36000.16310.1280
39000.16110.1280
42000.16310.1280
45000.17890.1280
48000.16280.1280
51000.16630.1280
54000.17860.1280
57000.17300.1280
60000.17130.1280
63000.17740.1280

Measured result

Verbatim from the run's summary.json.

checkpoint_round_trip
true
checkpoint_sha256
8e2413d580f59d78…
checkpoint_source
None
checkpoint_source_sha256
None
checks
all_valid
true
beats_drop_baseline
false
beats_marginal_baseline
true
median_recovery
false
config_sha256
91f01db8a48aeeeb…
continuity
copy_previous
mean_bins_off
25.64749262536873
median_bins_off
17.0
n
339
marginal
mean_bins_off
140.05162241887905
median_bins_off
138.5
n
339
model
mean_bins_off
22.073746312684367
median_bins_off
15.0
n
339
coordinate_tau
1.0
criteria
beats_drop_baseline
true
beats_marginal_baseline
true
min_median_recovery
0.3
require_all_valid
true
deterministic_algorithms
true
device
cuda
evaluation_icons
339
evaluation_split
validation
first_train_loss
6.352187633514404
gpu
NVIDIA GeForce RTX 4080
held_out_nll_per_masked_token
3.855228947045963
icon_presentations
100800
initial_held_out_nll
6.218960549879478
inpainting
all_valid
true
attempted
56
close_reproduction_rate
0.0
decoding
chain_order
true
greedy_for_criteria
true
iterations
8
samples
1
drop
median_rgba_mae
0.005432628298232873
exact_reproduction_rate
0.0
icons
56
identity_policy
join_span
inpaint_seed
3501
marginal
helped
4
mean_recovery
-9243201.564606853
median_recovery
-1.5514862525194055
median_rgba_mae
0.02051230407286371
model
helped
24
mean_recovery
-1.1948874697817782
median_recovery
-0.013644876652820558
median_rgba_mae
0.004619322651900267
paired
drop_minus_model
interval
-0.0016402854628366566, 0.0005194522916052009
interval_excludes_zero_above
false
mean
-0.0005604165856157279
n
56
positive
24
marginal_minus_model
interval
0.004541531249884808, 0.015153157968395097
interval_excludes_zero_above
true
mean
0.009847344609139953
n
56
positive
49
rows_sha256
3ff53db42f479f0f…
sheet_sha256
07790ad9e994c3f2…
span_length
4
task
span
last_train_loss
3.964060068130493
loss_reduction_factor
1.6131235356708724
marginal_masked_token_accuracy
0.12803273896521486
marginal_nll_per_masked_token
4.7650980387679045
mask_mixture
drawn
span
100800
max_paths
2
random_rate
0.05, 0.95
span_max
4
weights
span
1.0
masked_token_accuracy
0.17743349897690733
metric_head
true
metrics_sha256
ef088d9ebaeecb83…
model_parameters
531892
nll_ratio_to_marginal
0.8090555358317028
path_binding
false
predeclared_outcome
falsified
schema_version
1
selected_step
6300
sequence_length
1376
start_features
true
steps_run
6300
study_version
masked-span-l20-specialist-4-span4
timing
peak_vram_gib
1.9460439682006836
prepare_seconds
102.55899478000356
train_seconds
418.81431883000187
torch_version
2.14.0a0+4fdf77b940.nv26.08
train_icons
2681
trained_here
true
vocabulary
418

Visual output

inpainting
inpainting.png

State transitions

  1. planned2026-09-21T10:30:42Z
  2. completed2026-09-21T10:44:16Ztrained on spans of one to four and read at four, the specialist ties the join fill - mean -0.0006, interval [-0.0016, +0.0005] spanning zero, 24 of 56 helped, median recovery -0.01 - beats the marginal policy on 49 of 56, and is indistinguishable span for span from the model trained on shorter spans (30 better, 24 worse); held-out likelihood still falling at step 6,300; every completion valid

Run record

This run has no run.yaml. What follows is the identity and configuration carried by its rows in state/runs.jsonl, the append-only registry.

run_id
masked-span-l20-specialist-4-span4-9202129-2681icons-9b9b1699
config
configs/learning/masked-span-l20-specialist-4-span4.yaml
outputs
checkpoint_root
data/processed/masked-span-l20-specialist-4-span4
report_root
reports/learning/masked-span-l20-specialist-4-span4