merge-o1-finetune-cf747da-48targets
completed —
Measured behaviour
Table view
| optimizer step | model |
|---|---|
| 0 | 0.2936 |
| 25 | 0.2487 |
| 50 | 0.2359 |
| 75 | 0.2339 |
| 100 | 0.2411 |
| 125 | 0.2727 |
| 150 | 0.2584 |
| 175 | 0.2797 |
| 200 | 0.3014 |
| 225 | 0.2998 |
Measured result
Verbatim from the run's summary.json.
adapter_sha256
a0f9a9e79da0cd7e…arms
first_component
chance_top1
0.020833333333333332
clip_to_target_mean
0.9397472254931927
codec_valid_rate
1.0
rendered_rate
1.0
rgba_mae_to_target_mean
0.10492695078331356
rgba_mae_to_target_median
0.11425123736262321
target_top1_rate
0.5416666666666666
targets
48
model
chance_top1
0.020833333333333332
clip_to_target_mean
0.9341366610875944
codec_valid_rate
0.7916666666666666
ended_rate
0.8541666666666666
median_tokens
1464.0
memorised_exactly
0
rendered_rate
0.8541666666666666
rgba_mae_to_target_mean
0.11430112521232265
rgba_mae_to_target_median
0.1164664551615715
target_top1_rate
0.3958333333333333
targets
48
overlay
chance_top1
0.020833333333333332
clip_to_target_mean
0.8477968350052834
codec_valid_rate
0.5
rendered_rate
1.0
rgba_mae_to_target_mean
0.15553298220038414
rgba_mae_to_target_median
0.1617993712425232
target_top1_rate
0.16666666666666666
targets
48
checks
closer_than_first_component
false
closer_than_overlay
true
codec_valid_rate
false
memorisation
true
target_top1_rate
false
config_sha256
058d5cb238ed7c84…criteria
closer_than_first_component_ci_excludes_zero
true
closer_than_overlay_ci_excludes_zero
true
max_memorised_exactly
0
min_codec_valid_rate
0.8
min_target_top1_rate
0.5
generation
max_new_tokens
3072
seconds_per_merge
35.254253748748546
mode
finetune
model
fine_tuned
true
parameters
1898644288
revision
b1485b2fa6dfa1287294f269f5fb618e03d52d7c
source
Qwen/Qwen3.5-2B-Base
trainable_parameters
16819200
model_vs_baselines
first_component
error_reduction_ci95
-0.002761572472205976, 0.001424268394617772
error_reduction_mean
-0.0004178822585722295
paired_targets
41
targets_closer
10
overlay
error_reduction_ci95
0.035743292953063756, 0.05551476822427769
error_reduction_mean
0.045742181254687105
paired_targets
41
targets_closer
40
pairs
primary/test
91
primary/train
817
primary/validation
107
predeclared_outcome
falsified
rows_sha256
a4a6519fee9a35e1…schema_version
1
sheet_sha256
82a9dd11f20a93b7…study_version
merge-o1-finetune
targets
48
timing
load_seconds
4.438463274971582
training
excluded_selection_over_max_tokens
7
excluded_train_over_max_tokens
90
held_out_icons
25
held_out_nll
0.23391149293048769
initial_held_out_nll
0.29355869781147814
max_tokens
8192
peak_vram_gib
6.69413948059082
selected_step
75
selection_targets
25
sequences_per_step
8
steps_run
225
train_icons
727
train_seconds
3448.24761878798
train_targets
727
Visual output
State transitions
- planned2026-09-24T19:56:29Z
- completed2026-09-24T21:27:58Z727 training pairs under 8,192 tokens, held-out likelihood 0.294 to 0.234 at step 75 and rising after, stopped on patience at 225 (57 min); on 48 held-out targets the model's merges are taken by the codec 79% of the time (85% end, median 1,464 tokens), reach pixel error 0.114 to the true merged icon against 0.156 for the overlay (closer on 40 of 41, interval +0.036 to +0.056) and 0.105 for the first component alone (closer on 10 of 41, interval -0.003 to +0.001), and retrieve the true target first for 0.40 (the first component alone: 0.54); no training target reproduced
Run record
This run has no run.yaml. What follows is the identity and configuration carried by its rows in state/runs.jsonl, the append-only registry.
run_id
merge-o1-finetune-cf747da-48targets
config
configs/learning/merge-o1-finetune.yaml
outputs
adapter_sha256
a0f9a9e79da0cd7e…report_root
reports/learning/merge-o1-finetune
