semif-m3-qwen3.5-4b-averaged-52b1950-52decisions
completed —
Measured result
Verbatim from the run's summary.json.
accuracy
0.5769230769230769
averaged_readout
accuracy
0.75
clarify_false_alarms
10
note
secondary readout, mean of the two option orders; not a criterion
checks
accuracy
false
clarify_recall
true
order_invariance
false
clarify_false_alarms
22
clarify_recall
1.0
confusion
clarify
clarify
14
generate
0
recolor
0
restyle
0
simplify
0
generate
clarify
8
generate
2
recolor
0
restyle
0
simplify
0
recolor
clarify
3
generate
0
recolor
7
restyle
0
simplify
0
restyle
clarify
5
generate
0
recolor
0
restyle
5
simplify
0
simplify
clarify
6
generate
0
recolor
0
restyle
0
simplify
2
criteria
min_accuracy
0.8
min_clarify_recall
0.7
min_order_invariance
0.9
decisions
52
decisions_sha256
ce8d733bb0da0013…model
dtype
bfloat16
parameters
4205751296
revision
851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a
source
Qwen/Qwen3.5-4B
torch_version
2.14.0a0+4fdf77b940.nv26.08
transformers_version
5.17.0
operations
generate, recolor, restyle, simplify, clarify
order_invariance
0.5
per_kind
base
0.47368421052631576
distractor
0.4
missing_target
1.0
negation
0.75
paraphrase
0.2
per_operation
clarify
1.0
generate
0.2
recolor
0.7
restyle
0.5
simplify
0.25
predeclared_outcome
falsified
readout
native final-position logits restricted to the answer letters; softmax over letters
schema_version
1
study
semif-decision-readout
timing
load_seconds
3.810356097004842
mean_decision_seconds
0.04762484951872746
State transitions
- completed2026-09-21T13:59:56Zaveraging the two option orders' probabilities lifts accuracy from 0.577 to 0.75 and halves the false asks from 22 to 10; the direct readout is unchanged at 0.577 with 50% order invariance
Run record
This run has no run.yaml. What follows is the identity and configuration carried by its rows in state/runs.jsonl, the append-only registry.
run_id
semif-m3-qwen3.5-4b-averaged-52b1950-52decisions
outputs
report_root
reports/learning/semif-m3-qwen3.5-4b-averaged