semif-m3-qwen3.5-0.8b-279cc3d-52decisions
completed —
Measured result
Verbatim from the run's summary.json.
accuracy
0.40384615384615385
checks
accuracy
false
clarify_recall
false
order_invariance
false
clarify_false_alarms
23
clarify_recall
0.5714285714285714
confusion
clarify
clarify
8
generate
0
recolor
5
restyle
1
simplify
0
generate
clarify
9
generate
1
recolor
0
restyle
0
simplify
0
recolor
clarify
0
generate
0
recolor
10
restyle
0
simplify
0
restyle
clarify
8
generate
0
recolor
0
restyle
2
simplify
0
simplify
clarify
6
generate
1
recolor
0
restyle
1
simplify
0
criteria
min_accuracy
0.8
min_clarify_recall
0.7
min_order_invariance
0.9
decisions
52
decisions_sha256
e3de27630e72d14e…model
dtype
bfloat16
parameters
752393024
revision
2fc06364715b967f1860aea9cf38778875588b17
source
Qwen/Qwen3.5-0.8B
torch_version
2.14.0a0+4fdf77b940.nv26.08
transformers_version
5.17.0
operations
generate, recolor, restyle, simplify, clarify
order_invariance
0.8076923076923077
per_kind
base
0.3157894736842105
distractor
0.2
missing_target
0.5714285714285714
negation
0.5
paraphrase
0.4
per_operation
clarify
0.5714285714285714
generate
0.1
recolor
1.0
restyle
0.2
simplify
0.0
predeclared_outcome
falsified
readout
native final-position logits restricted to the answer letters; softmax over letters
schema_version
1
study
semif-decision-readout
timing
load_seconds
2.560997088003205
mean_decision_seconds
0.02920984846115551
State transitions
- planned2026-09-21T12:57:48Z
- completed2026-09-21T12:57:48Zaccuracy 0.404 over 52 decisions: recolour 1.0, ask 0.57, restyle 0.2, generate 0.1, simplify 0.0; 23 of the 38 non-ask requests were routed to ask; the choice survived reversing the option order on 81% of decisions
Run record
This run has no run.yaml. What follows is the identity and configuration carried by its rows in state/runs.jsonl, the append-only registry.
run_id
semif-m3-qwen3.5-0.8b-279cc3d-52decisions
outputs
report_root
reports/learning/semif-m3-qwen3.5-0.8b