MojiDiff

← experiments

masked-continuity-l5-78d4545-2681icons-9b9b1699

completed —

Measured behaviour

Held-out negative log likelihood per free token
modelposition-marginal floor
nats per tokenoptimizer step02460500100015002000selected
Table view
optimizer stepmodelposition-marginal floor
05.19844.7865
3004.61204.7865
6004.59554.7865
9004.58854.7865
12004.57154.7865
15004.55214.7865
18004.53644.7865
21004.53184.7865
Masked-token accuracy over unforced positions
modelposition-marginal floor
accuracyoptimizer step00.050.10.150500100015002000selected
Table view
optimizer stepmodelposition-marginal floor
00.05590.1307
3000.11890.1307
6000.14580.1307
9000.14150.1307
12000.14470.1307
15000.14200.1307
18000.14900.1307
21000.14470.1307

Measured result

Verbatim from the run's summary.json.

checkpoint_round_trip
true
checkpoint_sha256
fb614557426bf5f0…
checks
all_valid
true
beats_copy_previous
false
config_sha256
25feed4f0cae8e4d…
continuity
copy_previous
mean_bins_off
25.64749262536873
median_bins_off
17.0
n
339
marginal
mean_bins_off
140.05162241887905
median_bins_off
138.5
n
339
model
mean_bins_off
59.690265486725664
median_bins_off
54.5
n
339
coordinate_tau
1.0
criteria
beats_copy_previous
true
require_all_valid
true
deterministic_algorithms
true
device
cuda
evaluation_icons
339
evaluation_split
validation
first_train_loss
4.990401268005371
gpu
NVIDIA GeForce RTX 4080
held_out_nll_per_masked_token
4.531823072592503
icon_presentations
33600
initial_held_out_nll
5.198353921809204
inpainting
all_valid
true
attempted
16
decoding
greedy_for_criteria
true
iterations
8
samples
1
drop
median_rgba_mae
0.009685249485596707
exact_reproduction_rate
0.0
icons
16
marginal
helped
2
mean_recovery
-29608927.159026496
median_recovery
-1.3892468653280403
median_rgba_mae
0.026510208635923502
model
helped
2
mean_recovery
-251693562.62832725
median_recovery
-0.3468446812109242
median_rgba_mae
0.014632977033405955
paired
drop_minus_model
interval
-0.009900194960851053, 0.008203940648941099
interval_excludes_zero_above
false
mean
-0.000848127155954977
n
16
positive
2
marginal_minus_model
interval
-0.0021257274622927484, 0.014071774689053825
interval_excludes_zero_above
false
mean
0.005973023613380538
n
16
positive
11
rows_sha256
027ba80aef0c7b20…
sheet_sha256
3d37f1333ff2b6dd…
last_train_loss
4.410138130187988
loss_reduction_factor
1.1470778621627435
marginal_masked_token_accuracy
0.1307154384077461
marginal_nll_per_masked_token
4.786513732286344
mask_mixture
drawn
span
33600
max_paths
2
random_rate
0.05, 0.95
span_max
1
weights
span
1.0
masked_token_accuracy
0.1447014523937601
metrics_sha256
fde290171ab5a42d…
model_parameters
526498
nll_ratio_to_marginal
0.9467899448452675
path_binding
false
predeclared_outcome
falsified
schema_version
1
selected_step
2100
sequence_length
1376
steps_run
2100
study_version
masked-continuity-l5
timing
peak_vram_gib
1.9457483291625977
prepare_seconds
101.73797533099423
train_seconds
128.32428999000695
torch_version
2.14.0a0+4fdf77b940.nv26.08
train_icons
2681
vocabulary
418

Visual output

inpainting
inpainting.png

State transitions

  1. planned2026-09-21T08:57:15Z
  2. completed2026-09-21T09:03:04Ztrained on single-segment masks alone, the model's predicted endpoint sits 59.7 bins from the truth (median 54.5) against 25.6 (median 17.0) for the copy-the-previous-endpoint policy and 140 for the marginal argmax; held-out likelihood 0.947 of the floor, training loss flat from step 300

Run record

This run has no run.yaml. What follows is the identity and configuration carried by its rows in state/runs.jsonl, the append-only registry.

run_id
masked-continuity-l5-78d4545-2681icons-9b9b1699
config
configs/learning/masked-continuity-l5.yaml
outputs
checkpoint_root
data/processed/masked-continuity-l5
report_root
reports/learning/masked-continuity-l5