MojiDiff

← experiments

ar-regularisation-i6-86094df-3levels-9b9b1699

completed —

Measured behaviour

Held-out negative log likelihood per free token
modelposition-marginal floor
nats per tokenoptimizer step0246100020003000
Table view
optimizer stepmodelposition-marginal floor
3003.66123.9290
6003.64803.9290
9003.83843.9290
12003.98493.9290
15004.21913.9290
18004.60093.9290
21004.74263.9290
24005.00713.9290
27005.25503.9290
30005.49423.9290
3003.67323.9290
6003.59773.9290
9003.64033.9290
12003.71513.9290
15003.80833.9290
18003.86733.9290
21003.89803.9290
24003.94463.9290
27004.05043.9290
30004.04803.9290
3003.70903.9290
6003.61643.9290
9003.59213.9290
12003.59943.9290
15003.62533.9290
18003.68013.9290
21003.67353.9290
24003.69243.9290
27003.73893.9290
30003.74913.9290
33003.77363.9290

Measured result

Verbatim from the run's summary.json.

arms
dropout
0.0
epochs_to_best
3.5807534502051475
fraction
1.0
held_out_nll
3.64796488139269
icons
2681
label
none
marginal_nll
3.9289882210903344
parameters
524674
ratio_to_own_floor
0.9284743746012917
selected_step
600
steps_run
3000
trace
fraction
1.0
held_out_nll
3.6611640545050923
label
none
marginal_nll
3.9289882210903344
step
300
train_loss
3.2960057258605957
fraction
1.0
held_out_nll
3.64796488139269
label
none
marginal_nll
3.9289882210903344
step
600
train_loss
3.3200597763061523
fraction
1.0
held_out_nll
3.8384189372997466
label
none
marginal_nll
3.9289882210903344
step
900
train_loss
2.903935194015503
fraction
1.0
held_out_nll
3.984894900990683
label
none
marginal_nll
3.9289882210903344
step
1200
train_loss
1.603338360786438
fraction
1.0
held_out_nll
4.219101442814329
label
none
marginal_nll
3.9289882210903344
step
1500
train_loss
2.4615190029144287
fraction
1.0
held_out_nll
4.600927306373807
label
none
marginal_nll
3.9289882210903344
step
1800
train_loss
1.678338646888733
… and 4 more
train_seconds
305.60176272000535
weight_decay
0.01
dropout
0.1
epochs_to_best
3.5807534502051475
fraction
1.0
held_out_nll
3.5977242662007285
icons
2681
label
dropout-0.1
marginal_nll
3.9289882210903344
parameters
524674
ratio_to_own_floor
0.9156872109945708
selected_step
600
steps_run
3000
trace
fraction
1.0
held_out_nll
3.6731803649732435
label
dropout-0.1
marginal_nll
3.9289882210903344
step
300
train_loss
3.421252727508545
fraction
1.0
held_out_nll
3.5977242662007285
label
dropout-0.1
marginal_nll
3.9289882210903344
step
600
train_loss
3.4300286769866943
fraction
1.0
held_out_nll
3.640337257682916
label
dropout-0.1
marginal_nll
3.9289882210903344
step
900
train_loss
3.2496330738067627
fraction
1.0
held_out_nll
3.7151111867983517
label
dropout-0.1
marginal_nll
3.9289882210903344
step
1200
train_loss
2.272975444793701
fraction
1.0
held_out_nll
3.808261396217343
label
dropout-0.1
marginal_nll
3.9289882210903344
step
1500
train_loss
3.0613815784454346
fraction
1.0
held_out_nll
3.867266619247664
label
dropout-0.1
marginal_nll
3.9289882210903344
step
1800
train_loss
2.3356008529663086
… and 4 more
train_seconds
306.15317193399824
weight_decay
0.01
dropout
0.3
epochs_to_best
5.371130175307721
fraction
1.0
held_out_nll
3.592118464805465
icons
2681
label
dropout-0.3
marginal_nll
3.9289882210903344
parameters
524674
ratio_to_own_floor
0.9142604311011692
selected_step
900
steps_run
3300
trace
fraction
1.0
held_out_nll
3.709010581521979
label
dropout-0.3
marginal_nll
3.9289882210903344
step
300
train_loss
3.567981481552124
fraction
1.0
held_out_nll
3.616420899002616
label
dropout-0.3
marginal_nll
3.9289882210903344
step
600
train_loss
3.518282413482666
fraction
1.0
held_out_nll
3.592118464805465
label
dropout-0.3
marginal_nll
3.9289882210903344
step
900
train_loss
3.475635290145874
fraction
1.0
held_out_nll
3.599358769742663
label
dropout-0.3
marginal_nll
3.9289882210903344
step
1200
train_loss
2.826960802078247
fraction
1.0
held_out_nll
3.6253005883014704
label
dropout-0.3
marginal_nll
3.9289882210903344
step
1500
train_loss
3.457867383956909
fraction
1.0
held_out_nll
3.680139994520633
label
dropout-0.3
marginal_nll
3.9289882210903344
step
1800
train_loss
2.757734775543213
… and 5 more
train_seconds
336.6650172630034
weight_decay
0.01
axis
regularisation
best_ratio_to_own_floor
0.9142604311011692
checks
best_arm_clears_magnitude_bar
false
config_sha256
2680463269a80492…
criteria
expect_monotone
false
max_best_ratio
0.85
min_second_doubling_share
0.25
deterministic_algorithms
true
device
cuda
gpu
NVIDIA GeForce RTX 4080
improvement_first_doubling
0.050240615191961435
improvement_second_doubling
0.0056058013952635655
predeclared_outcome
not_limited_by_regularisation
prepare_seconds
102.37748816399835
schema_version
1
second_doubling_share
0.11157907549190403
study_version
ar-regularisation-i6
torch_version
2.14.0a0+4fdf77b940.nv26.08
validation_icons
339

State transitions

  1. planned2026-09-21T13:55:00Z
  2. completed2026-09-21T14:30:00Zdropout is the largest single effect measured in this gate and is still an order of magnitude too small; 0.9143 against a standing bar of 0.5

Run record

This run has no run.yaml. What follows is the identity and configuration carried by its rows in state/runs.jsonl, the append-only registry.

run_id
ar-regularisation-i6-86094df-3levels-9b9b1699
config
configs/learning/ar-regularisation-i6.yaml