MojiDiff

← experiments

masked-span-l15-mixture-bde5d24-2681icons-9b9b1699

completed —

Measured behaviour

Held-out negative log likelihood per free token
modelposition-marginal floor
nats per tokenoptimizer step024601000200030004000selected
Table view
optimizer stepmodelposition-marginal floor
05.26064.1110
42003.87554.1110
Masked-token accuracy over unforced positions
modelposition-marginal floor
accuracyoptimizer step00.10.20.301000200030004000selected
Table view
optimizer stepmodelposition-marginal floor
00.06600.2413
42000.26170.2413

Measured result

Verbatim from the run's summary.json.

checkpoint_round_trip
true
checkpoint_sha256
baa5647f55cb6fc2…
checkpoint_source
data/processed/masked-inpaint-l14-start-head/checkpoint.zip
checkpoint_source_sha256
ec454fbde6509b8b…
checks
all_valid
true
beats_drop_baseline
false
beats_marginal_baseline
true
median_recovery
false
config_sha256
37aa3d2312507616…
continuity
copy_previous
mean_bins_off
25.64749262536873
median_bins_off
17.0
n
339
marginal
mean_bins_off
140.05162241887905
median_bins_off
138.5
n
339
model
mean_bins_off
41.79203539823009
median_bins_off
36.0
n
339
coordinate_tau
1.0
criteria
beats_drop_baseline
true
beats_marginal_baseline
true
min_median_recovery
0.3
require_all_valid
true
deterministic_algorithms
true
device
cuda
evaluation_icons
339
evaluation_split
validation
first_train_loss
None
gpu
NVIDIA GeForce RTX 4080
held_out_nll_per_masked_token
3.875509736467909
icon_presentations
None
initial_held_out_nll
5.260572421212839
inpainting
all_valid
true
attempted
64
close_reproduction_rate
0.0
decoding
chain_order
true
greedy_for_criteria
true
iterations
8
samples
1
drop
median_rgba_mae
0.0014610377329944322
exact_reproduction_rate
0.0
icons
64
identity_policy
join_span
marginal
helped
4
mean_recovery
-253049919.98605102
median_recovery
-2.743950679036371
median_rgba_mae
0.01059368191721133
model
helped
9
mean_recovery
-45752014.78243475
median_recovery
-1.3453651432695155
median_rgba_mae
0.00561181841563786
paired
drop_minus_model
interval
-0.007376429164583229, -0.0022944000081352406
interval_excludes_zero_above
false
mean
-0.004835414586359235
n
64
positive
9
marginal_minus_model
interval
0.00016999600779433023, 0.004463120500768972
interval_excludes_zero_above
true
mean
0.002316558254281651
n
64
positive
40
rows_sha256
6455a03b8e1d2078…
sheet_sha256
0fa36c0ed127b975…
span_length
2
task
span
last_train_loss
None
loss_reduction_factor
None
marginal_masked_token_accuracy
0.2413130873428418
marginal_nll_per_masked_token
4.1110007453812765
mask_mixture
drawn
max_paths
2
random_rate
0.05, 0.95
span_max
None
weights
geometry
0.1
path
0.3
random
0.3
span
0.2
style
0.1
masked_token_accuracy
0.2616700381532789
metric_head
true
metrics_sha256
f302878ced46c748…
model_parameters
531892
nll_ratio_to_marginal
0.9427168654304083
path_binding
false
predeclared_outcome
falsified
schema_version
1
selected_step
4200
sequence_length
1376
start_features
true
steps_run
4200
study_version
masked-span-l15-mixture
timing
peak_vram_gib
1.8976726531982422
prepare_seconds
102.80136108299484
train_seconds
0.897731217002729
torch_version
2.14.0a0+4fdf77b940.nv26.08
train_icons
2681
trained_here
false
vocabulary
418

Visual output

inpainting
inpainting.png

State transitions

  1. planned2026-09-21T10:08:04Z
  2. completed2026-09-21T10:14:56Zarm 14's mixture model, read on two-segment spans, is worse than joining the visible ends: mean -0.0048 RGBA MAE with the interval excluding zero, 9 of 64 helped, median recovery -1.35; it beats the marginal policy narrowly (40 of 64); every completion valid

Run record

This run has no run.yaml. What follows is the identity and configuration carried by its rows in state/runs.jsonl, the append-only registry.

run_id
masked-span-l15-mixture-bde5d24-2681icons-9b9b1699
config
configs/learning/masked-span-l15-mixture.yaml
outputs
report_root
reports/learning/masked-span-l15-mixture