MojiDiff

← experiments

masked-span-l22-specialist-4-large-a83419d-2681icons-9b9b1699

completed —

Measured behaviour

Held-out negative log likelihood per free token
modelposition-marginal floor
nats per tokenoptimizer step024680200040006000selected
Table view
optimizer stepmodelposition-marginal floor
06.24084.7651
3004.46324.7651
6004.32024.7651
9004.27584.7651
12004.24174.7651
15004.17174.7651
18004.20914.7651
21004.15184.7651
24004.14124.7651
27004.06844.7651
30004.00984.7651
33003.96114.7651
36003.93254.7651
39003.93134.7651
42003.90284.7651
45003.88274.7651
48003.86104.7651
51003.85984.7651
54003.83104.7651
57003.79464.7651
60003.79674.7651
63003.77534.7651
Masked-token accuracy over unforced positions
modelposition-marginal floor
accuracyoptimizer step00.050.10.150.20200040006000selected
Table view
optimizer stepmodelposition-marginal floor
00.07340.1280
3000.13770.1280
6000.13970.1280
9000.14850.1280
12000.15230.1280
15000.15700.1280
18000.15020.1280
21000.16080.1280
24000.15840.1280
27000.17130.1280
30000.16250.1280
33000.15700.1280
36000.16430.1280
39000.16950.1280
42000.16660.1280
45000.17570.1280
48000.16460.1280
51000.18010.1280
54000.17660.1280
57000.17480.1280
60000.16630.1280
63000.17570.1280

Measured result

Verbatim from the run's summary.json.

checkpoint_round_trip
true
checkpoint_sha256
d402223ff9f3c943…
checkpoint_source
None
checkpoint_source_sha256
None
checks
all_valid
true
beats_drop_baseline
false
beats_marginal_baseline
true
median_recovery
false
config_sha256
f1b3aa67cf3ff306…
continuity
copy_previous
mean_bins_off
25.64749262536873
median_bins_off
17.0
n
339
marginal
mean_bins_off
140.05162241887905
median_bins_off
138.5
n
339
model
mean_bins_off
19.775811209439528
median_bins_off
13.0
n
339
coordinate_tau
1.0
criteria
beats_drop_baseline
true
beats_marginal_baseline
true
min_median_recovery
0.3
require_all_valid
true
deterministic_algorithms
true
device
cuda
evaluation_icons
339
evaluation_split
validation
first_train_loss
5.863951206207275
gpu
NVIDIA GeForce RTX 4080
held_out_nll_per_masked_token
3.77532209871945
icon_presentations
100800
initial_held_out_nll
6.240840703344805
inpainting
all_valid
true
attempted
56
close_reproduction_rate
0.0
decoding
chain_order
true
greedy_for_criteria
true
iterations
8
samples
1
drop
median_rgba_mae
0.005432628298232873
exact_reproduction_rate
0.0
icons
56
identity_policy
join_span
inpaint_seed
3501
marginal
helped
4
mean_recovery
-9243201.564606853
median_recovery
-1.5514862525194055
median_rgba_mae
0.02051230407286371
model
helped
28
mean_recovery
-2269426.702740625
median_recovery
-0.0012699399968391285
median_rgba_mae
0.0045676932038247395
paired
drop_minus_model
interval
-0.003170476520215973, 0.0013023401838913196
interval_excludes_zero_above
false
mean
-0.0009340681681623268
n
56
positive
28
marginal_minus_model
interval
0.0043722298206685046, 0.014575156232518205
interval_excludes_zero_above
true
mean
0.009473693026593354
n
56
positive
49
rows_sha256
d8903420d9914cc2…
sheet_sha256
bd0c1891df4862a4…
span_length
4
task
span
last_train_loss
3.9231743812561035
loss_reduction_factor
1.6530617892077695
marginal_masked_token_accuracy
0.12803273896521486
marginal_nll_per_masked_token
4.7650980387679045
mask_mixture
drawn
span
100800
max_paths
2
random_rate
0.05, 0.95
span_max
4
weights
span
1.0
masked_token_accuracy
0.17567962584039754
metric_head
true
metrics_sha256
dea0e16f36be21a3…
model_parameters
5358516
nll_ratio_to_marginal
0.7922863429050502
path_binding
false
predeclared_outcome
falsified
schema_version
1
selected_step
6300
sequence_length
1376
start_features
true
steps_run
6300
study_version
masked-span-l22-specialist-4-large
timing
peak_vram_gib
3.903975486755371
prepare_seconds
101.63444957701722
train_seconds
1236.324762987002
torch_version
2.14.0a0+4fdf77b940.nv26.08
train_icons
2681
trained_here
true
vocabulary
418

Visual output

inpainting
inpainting.png

State transitions

  1. planned2026-09-21T10:44:17Z
  2. completed2026-09-21T11:31:23Zat 5,358,516 parameters with dropout 0.1 - ten times arm 20 - the specialist ties the join fill on four-segment spans: mean -0.0009, interval [-0.0032, +0.0013] spanning zero, 28 of 56 helped, median recovery -0.001; the marginal policy beaten on 49 of 56; span for span against arm 20, 28 better and 27 worse; continuity 19.8 bins; every completion valid; 1,236 s to train, 3.9 GiB peak

Run record

This run has no run.yaml. What follows is the identity and configuration carried by its rows in state/runs.jsonl, the append-only registry.

run_id
masked-span-l22-specialist-4-large-a83419d-2681icons-9b9b1699
config
configs/learning/masked-span-l22-specialist-4-large.yaml
outputs
checkpoint_root
data/processed/masked-span-l22-specialist-4-large
report_root
reports/learning/masked-span-l22-specialist-4-large