MojiDiff

← experiments

masked-inpaint-l14-start-head-ba54da5-2681icons-9b9b1699

completed —

Measured behaviour

Held-out negative log likelihood per free token
modelposition-marginal floor
nats per tokenoptimizer step02460200040006000selected
Table view
optimizer stepmodelposition-marginal floor
05.26064.1110
3004.03764.1110
6003.92344.1110
9003.90534.1110
12003.90504.1110
15003.90014.1110
18003.90474.1110
21003.90404.1110
24003.89094.1110
27003.89804.1110
30003.89134.1110
33003.88654.1110
36003.87794.1110
39003.89624.1110
42003.87554.1110
45003.90244.1110
48003.88004.1110
51003.88004.1110
54003.92604.1110
57003.89524.1110
60003.88734.1110
63003.90474.1110
Masked-token accuracy over unforced positions
modelposition-marginal floor
accuracyoptimizer step00.10.20.30200040006000selected
Table view
optimizer stepmodelposition-marginal floor
00.06600.2413
3000.21290.2413
6000.24240.2413
9000.24900.2413
12000.24950.2413
15000.25440.2413
18000.25380.2413
21000.25780.2413
24000.25680.2413
27000.26140.2413
30000.26050.2413
33000.26070.2413
36000.26300.2413
39000.26250.2413
42000.26170.2413
45000.26320.2413
48000.26670.2413
51000.26430.2413
54000.26370.2413
57000.26090.2413
60000.26320.2413
63000.26750.2413

Measured result

Verbatim from the run's summary.json.

checkpoint_round_trip
true
checkpoint_sha256
ec454fbde6509b8b…
checks
all_valid
true
beats_drop_baseline
false
beats_marginal_baseline
true
median_recovery
false
config_sha256
0257980a69c181e9…
continuity
copy_previous
mean_bins_off
25.64749262536873
median_bins_off
17.0
n
339
marginal
mean_bins_off
140.05162241887905
median_bins_off
138.5
n
339
model
mean_bins_off
41.79203539823009
median_bins_off
36.0
n
339
coordinate_tau
1.0
criteria
beats_drop_baseline
true
beats_marginal_baseline
true
min_median_recovery
0.3
require_all_valid
true
deterministic_algorithms
true
device
cuda
evaluation_icons
339
evaluation_split
validation
first_train_loss
5.358754634857178
gpu
NVIDIA GeForce RTX 4080
held_out_nll_per_masked_token
3.875509736467909
icon_presentations
100800
initial_held_out_nll
5.260572421212839
inpainting
all_valid
true
attempted
64
close_reproduction_rate
0.0
decoding
chain_order
true
greedy_for_criteria
true
iterations
8
samples
1
drop
median_rgba_mae
0.006453019093439845
exact_reproduction_rate
0.0
icons
64
marginal
helped
0
mean_recovery
-112236156.97001305
median_recovery
-1.7608338149263418
median_rgba_mae
0.019917714385136773
model
helped
6
mean_recovery
-5055975.0345634995
median_recovery
-0.1087351943803935
median_rgba_mae
0.007291666666666667
paired
drop_minus_model
interval
-0.0016653814081545757, -0.0006072185317056271
interval_excludes_zero_above
false
mean
-0.0011362999699301014
n
64
positive
6
marginal_minus_model
interval
0.01057050893873887, 0.017620381253178944
interval_excludes_zero_above
true
mean
0.014095445095958907
n
64
positive
64
rows_sha256
a02b0457847a55fa…
sheet_sha256
3b0f9e3f18bba2a7…
last_train_loss
3.8391597270965576
loss_reduction_factor
1.3573885189119042
marginal_masked_token_accuracy
0.2413130873428418
marginal_nll_per_masked_token
4.1110007453812765
mask_mixture
drawn
geometry
10032
path
30461
random
30235
span
20052
style
10020
max_paths
2
random_rate
0.05, 0.95
span_max
None
weights
geometry
0.1
path
0.3
random
0.3
span
0.2
style
0.1
masked_token_accuracy
0.2616700381532789
metric_head
true
metrics_sha256
fb1c4decc67f8d00…
model_parameters
531892
nll_ratio_to_marginal
0.9427168654304083
path_binding
false
predeclared_outcome
falsified
schema_version
1
selected_step
4200
sequence_length
1376
start_features
true
steps_run
6300
study_version
masked-inpaint-l14-start-head
timing
peak_vram_gib
1.9460439682006836
prepare_seconds
103.3546569510072
train_seconds
421.51113984399126
torch_version
2.14.0a0+4fdf77b940.nv26.08
train_icons
2681
vocabulary
418

Visual output

inpainting
inpainting.png

State transitions

  1. planned2026-09-21T09:48:14Z
  2. completed2026-09-21T10:08:03Zwith start features, the metric head and chain-order decoding, whole-path completion still loses to the hole - mean paired difference -0.0011, interval [-0.0017, -0.0006], 6 of 64 helped, median recovery -0.11 - but the gap to the hole is a quarter of l2's -0.0040, the model beats the marginal policy on all 64 icons, and icon by icon it beats l2 on 44 of 64 by 0.0029; the fills are quiet rather than scribbles; every completion valid

Run record

This run has no run.yaml. What follows is the identity and configuration carried by its rows in state/runs.jsonl, the append-only registry.

run_id
masked-inpaint-l14-start-head-ba54da5-2681icons-9b9b1699
config
configs/learning/masked-inpaint-l14-start-head.yaml
outputs
checkpoint_root
data/processed/masked-inpaint-l14-start-head
report_root
reports/learning/masked-inpaint-l14-start-head