MojiDiff

← experiments

prior-m2-finetune-4k-evaluate-6a8b77f-32icons-9b9b1699

completed —

Measured result

Verbatim from the run's summary.json.

adapter
data/processed/prior/prior-m2-finetune-4k/adapter.zip
adapter_sha256
f3c0f43274d7fc8c…
checks
clip_gain_over_control
false
codec_valid_rate
false
memorisation
true
reference_top1_rate
false
clip_to_caption
mean
0.22406816379777317
median
0.2249937802553177
n
29
clip_to_reference
mean
0.8267076817052118
median
0.8228124380111694
n
29
closed_svg_rate
0.453125
codec_valid_rate
0.421875
config_sha256
63a2da0a371a8d4f…
control
clip_to_reference_mean
0.8119077861309052
control_clip_to_reference_mean
0.7987703651189804
control_codec_valid_rate
0.046875
control_reference_top1_rate
None
control_rows_sha256
ca33f9adbb2b2fc0…
gain_ci95
-0.008706659078598022, 0.03297443956136704
icons_improved
7
mean_gain
0.013137421011924744
paired_icons
10
report_root
reports/learning/prior-m2-control
criteria
clip_gain_over_control_ci_excludes_zero
true
max_memorised_exactly
0
min_codec_valid_rate
0.8
min_reference_top1_rate
0.25
drawings
64
ended_rate
0.453125
failures
encode:endpoint coordinate is outside 0..72 and clamping is disabled
1
no_closed_svg
35
pack:program needs 138 packed segments but capacity is 128
1
generation
max_new_tokens
4096
seconds_per_drawing
54.20806298617299
icons
32
median_tokens
4096.0
memorised_exactly
0
mode
evaluate
model
fine_tuned
true
parameters
1881825088
revision
b1485b2fa6dfa1287294f269f5fb618e03d52d7c
source
Qwen/Qwen3.5-2B-Base
trainable_parameters
0
predeclared_outcome
falsified
raw_render_rate
0.4375
reference_rank
chance_top1
0.03125
mean
12.379310344827585
scored
29
top1_rate
0.03125
top5_rate
0.203125
rows_sha256
c8ec9df0f09d710d…
samples_per_icon
2
schema_version
1
sheet_sha256
2783fabe804b3e80…
study_version
prior-m2-finetune-4k-evaluate
timing
load_seconds
4.172816991980653

Visual output

samples
samples.png

State transitions

  1. planned2026-09-21T19:58:41Z
  2. completed2026-09-21T21:00:56Zthe adapter trained on all 2,681 icons (held-out likelihood 0.391 to 0.271, selected at step 325) draws longer and closes less: 29 of 64 drawings close (the 2,048-token arm closed 47), 27 enter the codec (44), CLIP-to-reference 0.827 over the 29 (0.839 over 47), paired gain over the 10 pairable icons +0.013 with an interval of -0.009 to +0.033, the right icon first for 2 of 64 - chance; no training icon reproduced; no parse or render exceeded the 60 s guard

Run record

This run has no run.yaml. What follows is the identity and configuration carried by its rows in state/runs.jsonl, the append-only registry.

run_id
prior-m2-finetune-4k-evaluate-6a8b77f-32icons-9b9b1699
config
configs/learning/prior-m2-finetune-4k-evaluate.yaml
outputs
adapter_sha256
f3c0f43274d7fc8c…
report_root
reports/learning/prior-m2-finetune-4k-evaluate