MojiDiff

← experiments

masked-inpaint-l2-54061b2-2681icons-9b9b1699

completed —

Measured behaviour

Held-out negative log likelihood per free token
modelposition-marginal floor
nats per tokenoptimizer step0246020004000selected
Table view
optimizer stepmodelposition-marginal floor
04.80504.1110
3003.98214.1110
6003.92634.1110
9003.90504.1110
12003.90544.1110
15003.89834.1110
18003.89854.1110
21003.89324.1110
24003.89824.1110
27003.88694.1110
30003.89894.1110
33003.86884.1110
36003.89784.1110
39003.90964.1110
42003.91604.1110
45003.90504.1110
48003.89144.1110
51003.88744.1110
54003.98474.1110
57003.88144.1110
Masked-token accuracy over unforced positions
modelposition-marginal floor
accuracyoptimizer step00.10.20.3020004000selected
Table view
optimizer stepmodelposition-marginal floor
00.04900.2413
3000.23870.2413
6000.24630.2413
9000.25250.2413
12000.25050.2413
15000.25450.2413
18000.25620.2413
21000.25960.2413
24000.25860.2413
27000.26290.2413
30000.26170.2413
33000.26610.2413
36000.26370.2413
39000.26430.2413
42000.26570.2413
45000.26700.2413
48000.27010.2413
51000.26890.2413
54000.27070.2413
57000.27010.2413

Measured result

Verbatim from the run's summary.json.

checkpoint_round_trip
true
checkpoint_sha256
45d32c26f2dba0af…
checks
all_valid
true
beats_drop_baseline
false
beats_marginal_baseline
true
median_recovery
false
config_sha256
2a8a71d97359a9d9…
coordinate_tau
1.0
criteria
beats_drop_baseline
true
beats_marginal_baseline
true
min_median_recovery
0.3
require_all_valid
true
deterministic_algorithms
true
device
cuda
evaluation_icons
339
evaluation_split
validation
first_train_loss
4.784581661224365
gpu
NVIDIA GeForce RTX 4080
held_out_nll_per_masked_token
3.8688370011586817
icon_presentations
91200
initial_held_out_nll
4.805001695162264
inpainting
all_valid
true
attempted
64
decoding
greedy_for_criteria
true
iterations
8
samples
1
drop
median_rgba_mae
0.006453019093439845
exact_reproduction_rate
0.0
icons
64
marginal
helped
0
mean_recovery
-112236156.97001305
median_recovery
-1.7608338149263418
median_rgba_mae
0.019917714385136773
model
helped
7
mean_recovery
-165497.3872923598
median_recovery
-0.6925183889869049
median_rgba_mae
0.01281734144274994
paired
drop_minus_model
interval
-0.006853281016532452, -0.0012139148046972728
interval_excludes_zero_above
false
mean
-0.004033597910614862
n
64
positive
7
marginal_minus_model
interval
0.0065920744050887315, 0.015804219905459563
interval_excludes_zero_above
true
mean
0.011198147155274146
n
64
positive
54
rows_sha256
7924e3ba6d562112…
sheet_sha256
521d238cb1d1d221…
last_train_loss
3.75107479095459
loss_reduction_factor
1.2419757394077882
marginal_masked_token_accuracy
0.2413130873428418
marginal_nll_per_masked_token
4.1110007453812765
mask_mixture
drawn
geometry
9084
path
27531
random
27359
span
18210
style
9016
max_paths
2
random_rate
0.05, 0.95
weights
geometry
0.1
path
0.3
random
0.3
span
0.2
style
0.1
masked_token_accuracy
0.266074309272015
metrics_sha256
cbbad9bb94a87b3e…
model_parameters
526498
nll_ratio_to_marginal
0.941093724078093
predeclared_outcome
falsified
schema_version
1
selected_step
3300
sequence_length
1376
steps_run
5700
study_version
masked-inpaint-l2
timing
peak_vram_gib
1.9457483291625977
prepare_seconds
101.68371957101044
train_seconds
345.6692114850157
torch_version
2.14.0a0+4fdf77b940.nv26.08
train_icons
2681
vocabulary
418

Visual output

inpainting
inpainting.png

State transitions

  1. planned2026-09-21T08:16:40Z
  2. completed2026-09-21T08:30:48Zthe model beats the position-marginal policy on 54 of 64 held-out icons with the interval excluding zero, and is worse than leaving the hole on 57 of 64: mean paired difference -0.0040 RGBA MAE, interval [-0.0069, -0.0012]; median recovery -0.69 against a bar of 0.30; every completion valid

Run record

This run has no run.yaml. What follows is the identity and configuration carried by its rows in state/runs.jsonl, the append-only registry.

run_id
masked-inpaint-l2-54061b2-2681icons-9b9b1699
config
configs/learning/masked-inpaint-l2.yaml
outputs
checkpoint_root
data/processed/masked-inpaint-l2
report_root
reports/learning/masked-inpaint-l2