MojiDiff

← experiments

masked-overfit-l1-01917d5-4icons-9b9b1699

completed —

Measured behaviour

Held-out negative log likelihood per free token
modelposition-marginal floor
nats per tokenoptimizer step02460100020003000selected
Table view
optimizer stepmodelposition-marginal floor
05.78084.9159
1004.90084.9159
2003.78194.9159
3002.60644.9159
4002.34164.9159
5002.30674.9159
6002.27754.9159
7002.26074.9159
8002.24914.9159
9002.23764.9159
10002.22464.9159
11002.22104.9159
12002.21464.9159
13002.20714.9159
14002.20364.9159
15002.19744.9159
16002.19434.9159
17002.19364.9159
18002.19014.9159
19002.20204.9159
20002.18544.9159
21002.18874.9159
22002.18304.9159
23002.18154.9159
24002.18334.9159
25002.18314.9159
26002.18174.9159
27002.17434.9159
28002.18074.9159
29002.18504.9159
30002.17784.9159
Masked-token accuracy over unforced positions
modelposition-marginal floor
accuracyoptimizer step00.20.40.60100020003000selected
Table view
optimizer stepmodelposition-marginal floor
00.02380.5470
1000.06870.5470
2000.13740.5470
3000.21180.5470
4000.24820.5470
5000.19070.5470
6000.13880.5470
7000.11500.5470
8000.10800.5470
9000.09120.5470
10000.07430.5470
11000.07430.5470
12000.07290.5470
13000.07150.5470
14000.07290.5470
15000.07570.5470
16000.06870.5470
17000.06870.5470
18000.06870.5470
19000.06870.5470
20000.06870.5470
21000.06870.5470
22000.06870.5470
23000.06870.5470
24000.06870.5470
25000.06870.5470
26000.06870.5470
27000.06870.5470
28000.06870.5470
29000.06870.5470
30000.06870.5470

Measured result

Verbatim from the run's summary.json.

checkpoint_round_trip
true
checkpoint_sha256
3da5b0f8f824588d…
checks
all_valid
true
exact_path_reproduction
false
loss_reduction
false
masked_token_accuracy
false
config_sha256
98000972b1c6898b…
coordinate_tau
1.0
criteria
min_exact_path_reproduction_rate
1.0
min_loss_reduction_factor
100.0
min_masked_token_accuracy
0.99
require_all_valid
true
deterministic_algorithms
true
device
cuda
evaluation_icons
4
evaluation_split
train
first_train_loss
4.846034049987793
gpu
NVIDIA GeForce RTX 4080
held_out_nll_per_masked_token
2.174339690348878
icon_presentations
12000
initial_held_out_nll
5.7808438979225105
inpainting
all_valid
true
attempted
4
decoding
greedy_for_criteria
true
iterations
8
samples
1
drop
median_rgba_mae
0.0034825291999515855
exact_reproduction_rate
0.0
icons
4
marginal
helped
0
mean_recovery
-1.8921178315003433
median_recovery
-2.0224896560213343
median_rgba_mae
0.010612782921810698
model
helped
4
mean_recovery
0.8152479893889422
median_recovery
0.7915168120171914
median_rgba_mae
0.0007350104393609296
paired
drop_minus_model
interval
-0.016603383721697625, 0.03754687878947782
interval_excludes_zero_above
false
mean
0.010471747533890099
n
4
positive
4
marginal_minus_model
interval
-0.011947317879671328, 0.04393715084989645
interval_excludes_zero_above
false
mean
0.015994916485112563
n
4
positive
4
rows_sha256
31a4c874aa5445cb…
sheet_sha256
8600d06f82c56323…
last_train_loss
1.9388209581375122
loss_reduction_factor
2.6586664096606545
marginal_masked_token_accuracy
0.5469845722300141
marginal_nll_per_masked_token
4.9158920138280155
mask_mixture
drawn
geometry
1203
path
3603
random
3605
span
2422
style
1167
max_paths
2
random_rate
0.05, 0.95
weights
geometry
0.1
path
0.3
random
0.3
span
0.2
style
0.1
masked_token_accuracy
0.06872370266479663
metrics_sha256
7ba8cb44028dbaf3…
model_parameters
526498
nll_ratio_to_marginal
0.4423082696350189
predeclared_outcome
falsified
schema_version
1
selected_step
2700
sequence_length
1376
steps_run
3000
study_version
masked-overfit-l1
timing
peak_vram_gib
0.4974055290222168
prepare_seconds
0.2613500489969738
train_seconds
106.223827386013
torch_version
2.14.0a0+4fdf77b940.nv26.08
train_icons
4
vocabulary
418

Visual output

inpainting
inpainting.png

State transitions

  1. planned2026-09-21T08:07:04Z
  2. completed2026-09-21T08:12:33Zfalsified by a harness defect the test exists to catch: every non-coordinate token type is memorised exactly (lengths, layers, styles, kinds at accuracy 1.000, nll ~0) and every one of 407 masked coordinates is wrong by exactly one bin, because the soft coordinate target was centred on token truth-1 - the denoiser's bin-index convention copied into a loss whose logits index tokens; eval and train forwards agree to 1e-5, so it is not the encoder

Run record

This run has no run.yaml. What follows is the identity and configuration carried by its rows in state/runs.jsonl, the append-only registry.

run_id
masked-overfit-l1-01917d5-4icons-9b9b1699
config
configs/learning/masked-overfit-l1.yaml
artifacts
checkpoint_root
data/processed/masked-overfit-l1-attempt1
report_root
reports/learning/masked-overfit-l1-attempt1