MojiDiff

← experiments

merge-o1-finetune-cf747da-48targets

completed —

Measured behaviour

Held-out negative log likelihood per free token
nats per tokenoptimizer step00.10.20.30.4050100150200
Table view
optimizer stepmodel
00.2936
250.2487
500.2359
750.2339
1000.2411
1250.2727
1500.2584
1750.2797
2000.3014
2250.2998

Measured result

Verbatim from the run's summary.json.

adapter_sha256
a0f9a9e79da0cd7e…
arms
first_component
chance_top1
0.020833333333333332
clip_to_target_mean
0.9397472254931927
codec_valid_rate
1.0
rendered_rate
1.0
rgba_mae_to_target_mean
0.10492695078331356
rgba_mae_to_target_median
0.11425123736262321
target_top1_rate
0.5416666666666666
targets
48
model
chance_top1
0.020833333333333332
clip_to_target_mean
0.9341366610875944
codec_valid_rate
0.7916666666666666
ended_rate
0.8541666666666666
median_tokens
1464.0
memorised_exactly
0
rendered_rate
0.8541666666666666
rgba_mae_to_target_mean
0.11430112521232265
rgba_mae_to_target_median
0.1164664551615715
target_top1_rate
0.3958333333333333
targets
48
overlay
chance_top1
0.020833333333333332
clip_to_target_mean
0.8477968350052834
codec_valid_rate
0.5
rendered_rate
1.0
rgba_mae_to_target_mean
0.15553298220038414
rgba_mae_to_target_median
0.1617993712425232
target_top1_rate
0.16666666666666666
targets
48
checks
closer_than_first_component
false
closer_than_overlay
true
codec_valid_rate
false
memorisation
true
target_top1_rate
false
config_sha256
058d5cb238ed7c84…
criteria
closer_than_first_component_ci_excludes_zero
true
closer_than_overlay_ci_excludes_zero
true
max_memorised_exactly
0
min_codec_valid_rate
0.8
min_target_top1_rate
0.5
generation
max_new_tokens
3072
seconds_per_merge
35.254253748748546
mode
finetune
model
fine_tuned
true
parameters
1898644288
revision
b1485b2fa6dfa1287294f269f5fb618e03d52d7c
source
Qwen/Qwen3.5-2B-Base
trainable_parameters
16819200
model_vs_baselines
first_component
error_reduction_ci95
-0.002761572472205976, 0.001424268394617772
error_reduction_mean
-0.0004178822585722295
paired_targets
41
targets_closer
10
overlay
error_reduction_ci95
0.035743292953063756, 0.05551476822427769
error_reduction_mean
0.045742181254687105
paired_targets
41
targets_closer
40
pairs
primary/test
91
primary/train
817
primary/validation
107
predeclared_outcome
falsified
rows_sha256
a4a6519fee9a35e1…
schema_version
1
sheet_sha256
82a9dd11f20a93b7…
study_version
merge-o1-finetune
targets
48
timing
load_seconds
4.438463274971582
training
excluded_selection_over_max_tokens
7
excluded_train_over_max_tokens
90
held_out_icons
25
held_out_nll
0.23391149293048769
initial_held_out_nll
0.29355869781147814
max_tokens
8192
peak_vram_gib
6.69413948059082
selected_step
75
selection_targets
25
sequences_per_step
8
steps_run
225
train_icons
727
train_seconds
3448.24761878798
train_targets
727

Visual output

samples
samples.png

State transitions

  1. planned2026-09-24T19:56:29Z
  2. completed2026-09-24T21:27:58Z727 training pairs under 8,192 tokens, held-out likelihood 0.294 to 0.234 at step 75 and rising after, stopped on patience at 225 (57 min); on 48 held-out targets the model's merges are taken by the codec 79% of the time (85% end, median 1,464 tokens), reach pixel error 0.114 to the true merged icon against 0.156 for the overlay (closer on 40 of 41, interval +0.036 to +0.056) and 0.105 for the first component alone (closer on 10 of 41, interval -0.003 to +0.001), and retrieve the true target first for 0.40 (the first component alone: 0.54); no training target reproduced

Run record

This run has no run.yaml. What follows is the identity and configuration carried by its rows in state/runs.jsonl, the append-only registry.

run_id
merge-o1-finetune-cf747da-48targets
config
configs/learning/merge-o1-finetune.yaml
outputs
adapter_sha256
a0f9a9e79da0cd7e…
report_root
reports/learning/merge-o1-finetune