MojiDiff

← experiments

masked-continuity-l11-start-head-7d18a83-2681icons-9b9b1699

completed —

Measured behaviour

Held-out negative log likelihood per free token
modelposition-marginal floor
nats per tokenoptimizer step024680500100015002000selected
Table view
optimizer stepmodelposition-marginal floor
06.32174.7865
3004.32654.7865
6004.16904.7865
9004.12844.7865
12004.07424.7865
15004.04034.7865
18004.02234.7865
21004.00484.7865
Masked-token accuracy over unforced positions
modelposition-marginal floor
accuracyoptimizer step00.050.10.150.20500100015002000selected
Table view
optimizer stepmodelposition-marginal floor
00.04570.1307
3000.11300.1307
6000.13720.1307
9000.13820.1307
12000.15170.1307
15000.15600.1307
18000.16680.1307
21000.16140.1307

Measured result

Verbatim from the run's summary.json.

checkpoint_round_trip
true
checkpoint_sha256
0d630880b7a803fc…
checks
all_valid
true
beats_copy_previous
false
config_sha256
7c02e7c62f73fad9…
continuity
copy_previous
mean_bins_off
25.64749262536873
median_bins_off
17.0
n
339
marginal
mean_bins_off
140.05162241887905
median_bins_off
138.5
n
339
model
mean_bins_off
41.733038348082594
median_bins_off
37.0
n
339
coordinate_tau
1.0
criteria
beats_copy_previous
true
require_all_valid
true
deterministic_algorithms
true
device
cuda
evaluation_icons
339
evaluation_split
validation
first_train_loss
6.082370281219482
gpu
NVIDIA GeForce RTX 4080
held_out_nll_per_masked_token
4.004803353320405
icon_presentations
33600
initial_held_out_nll
6.321668299633434
inpainting
all_valid
true
attempted
16
decoding
chain_order
false
greedy_for_criteria
true
iterations
8
samples
1
drop
median_rgba_mae
0.009685249485596707
exact_reproduction_rate
0.0
icons
16
marginal
helped
2
mean_recovery
-29608927.159026496
median_recovery
-1.3892468653280403
median_rgba_mae
0.026510208635923502
model
helped
2
mean_recovery
-3.94284278682398
median_recovery
-0.365004512018558
median_rgba_mae
0.0136214748244977
paired
drop_minus_model
interval
-0.01016122114532422, -0.000247904406902851
interval_excludes_zero_above
false
mean
-0.005204562776113535
n
16
positive
2
marginal_minus_model
interval
-0.012795680612311434, 0.016028856598755394
interval_excludes_zero_above
false
mean
0.0016165879932219804
n
16
positive
14
rows_sha256
257bb55357d423bd…
sheet_sha256
2a320b45a42262fd…
last_train_loss
3.88089656829834
loss_reduction_factor
1.5785215257553418
marginal_masked_token_accuracy
0.1307154384077461
marginal_nll_per_masked_token
4.786513732286344
mask_mixture
drawn
span
33600
max_paths
2
random_rate
0.05, 0.95
span_max
1
weights
span
1.0
masked_token_accuracy
0.16137708445400753
metric_head
true
metrics_sha256
3b1cfbc164054087…
model_parameters
531892
nll_ratio_to_marginal
0.8366848143162969
path_binding
false
predeclared_outcome
falsified
schema_version
1
selected_step
2100
sequence_length
1376
start_features
true
steps_run
2100
study_version
masked-continuity-l11-start-head
timing
peak_vram_gib
1.9460439682006836
prepare_seconds
101.56017234697356
train_seconds
140.54242861099192
torch_version
2.14.0a0+4fdf77b940.nv26.08
train_icons
2681
vocabulary
418

Visual output

inpainting
inpainting.png

State transitions

  1. planned2026-09-21T09:32:44Z
  2. completed2026-09-21T09:38:59Zwith the start point in the input the metric head now shows: continuity error 41.7 bins (median 37.0) against arm 10's 47.8, held-out likelihood 4.005 against 4.119 - 0.837 of the floor - and still falling at the last evaluation; short of the copy policy's 25.6 at a third of the budget

Run record

This run has no run.yaml. What follows is the identity and configuration carried by its rows in state/runs.jsonl, the append-only registry.

run_id
masked-continuity-l11-start-head-7d18a83-2681icons-9b9b1699
config
configs/learning/masked-continuity-l11-start-head.yaml
outputs
checkpoint_root
data/processed/masked-continuity-l11-start-head
report_root
reports/learning/masked-continuity-l11-start-head