MojiDiff

← experiments

tiny-geometry-v4-ede1009-7ccbc64e-32a80ab5

completed declared

Hypothesis

Deterministic per-step corruption resampling for diverse-four will raise held-out changed-token accuracy to at least 0.75 and retained-token accuracy to at least 0.90 without increasing model size, batch size, corruption probability, or optimizer steps.

Why run it

Tests whether v3 is limited primarily by seeing only four fixed corruptions per icon, before spending evidence budget on more capacity or longer training.

Measured behaviour

Training token accuracy (per batch)
accuracyoptimizer step00.250.50.751100200300
Table view
optimizer steptrain accuracy
10.0012
100.2892
200.6366
300.8676
400.9694
500.9975
601.0000
701.0000
801.0000
901.0000
1001.0000
1101.0000
1201.0000
1301.0000
1401.0000
1501.0000
1601.0000
10.0020
100.0804
200.3095
300.5623
400.7269
500.7879
600.8441
700.8789
800.9031
900.9262
1000.9294
1100.9299
1200.9503
1300.9456
1400.9608
1500.9639
1600.9682
1700.9671
1800.9658
1900.9678
2000.9793
2100.9699
2200.9680
2300.9795
2400.9760
2500.9811
2600.9791
2700.9782
2800.9782
2900.9854
3000.9837
3100.9863
3200.9852

Measured result

Verbatim from the run's summary.json.

cases
diverse-four
checkpoint_sha256
e975f17c15497260…
config
corruptions_per_icon
4
heldout_corruptions_per_icon
4
icon_count
4
max_loss_ratio
0.2
min_heldout_accuracy
0.85
min_train_accuracy
0.98
name
diverse-four
resample_each_step
true
resume_step
160
seed
1702
steps
320
continuous_loss_sha256
1c909c13042da5ba…
final
accuracy
0.9856271777003485
changed_accuracy
0.9664634146341463
changed_correct
1585
changed_total
1640
correct
4526
loss
0.04804584034718573
retained_accuracy
0.9962737127371274
retained_correct
2941
retained_total
2952
total
4592
final_model_sha256
d7972e2a714662fd…
heldout_examples
16
initial
accuracy
0.002395470383275261
changed_accuracy
0.0024390243902439024
changed_correct
4
changed_total
1640
correct
11
loss
14.943475902080536
retained_accuracy
0.0023712737127371273
retained_correct
7
retained_total
2952
total
4592
loss_ratio
0.0029922704661883926
passes_predeclared_criteria
true
resume_exact
true
resumed_loss_sha256
1c909c13042da5ba…
train_examples
16
train_final
accuracy
0.985191637630662
changed_accuracy
0.9636251541307028
changed_correct
1563
changed_total
1622
correct
4524
loss
0.05133143730927259
retained_accuracy
0.996969696969697
retained_correct
2961
retained_total
2970
total
4592
one-icon
checkpoint_sha256
None
config
corruptions_per_icon
8
heldout_corruptions_per_icon
8
icon_count
1
max_loss_ratio
0.1
min_heldout_accuracy
0.95
min_train_accuracy
0.99
name
one-icon
resample_each_step
false
resume_step
None
seed
1701
steps
160
continuous_loss_sha256
None
final
accuracy
0.9356617647058824
changed_accuracy
0.8858695652173914
changed_correct
489
changed_total
552
correct
1527
loss
0.2106002140790224
retained_accuracy
0.9611111111111111
retained_correct
1038
retained_total
1080
total
1632
final_model_sha256
8e0a9878092058a0…
heldout_examples
8
initial
accuracy
0.001838235294117647
changed_accuracy
0.005434782608695652
changed_correct
3
changed_total
552
correct
3
loss
15.358022212982178
retained_accuracy
0.0
retained_correct
0
retained_total
1080
total
1632
loss_ratio
9.539973069621893e-05
passes_predeclared_criteria
false
resume_exact
None
resumed_loss_sha256
None
train_examples
8
train_final
accuracy
1.0
changed_accuracy
1.0
changed_correct
599
changed_total
599
correct
1632
loss
0.001458597937016748
retained_accuracy
1.0
retained_correct
1033
retained_total
1033
total
1632
code_identity
geometry_sha256
8f30cf4c834873bd…
git_commit
ede100987cb4b9dc8a179f36122abf319da9340c
packed_sha256
9b12879ba3a18236…
program_sha256
0d02b532b1dbe573…
tiny_study_sha256
8e09bccf8d3c1b90…
config_sha256
7ccbc64e51a0f009…
deterministic_algorithms
true
device
cpu
fixture
data/manifests/openmoji-17.0.0-tiny-learning-fixture-v1.json
fixture_sha256
32a80ab576a6d5a4…
metrics_sha256
e11c8598e51d6913…
model
d_model
48
feedforward
96
heads
4
layers
2
parameters
241072
scope
fixed-topology geometry-only diagnostic
passes_predeclared_gate_candidate
false
render_metrics_sha256
8eafdf8ce49f84c2…
schema_version
1
study_version
tiny-geometry-v4-resampled-corruption
torch_version
2.8.0+cpu

Visual output

trajectory
trajectory.png

Written result

Completed reproducibly; the diverse corruption-coverage hypothesis passed.

At unchanged batch size, model, probability, and step count, deterministic per-step resampling raised diverse-four held-out accuracy from 73.82% to 98.56%. Changed-token accuracy rose from 58.17% to 96.65%; retained-token accuracy rose from 82.52% to 99.63%. Checkpoint continuation and the full artifact rerun are exact. The paired renders are recognizable at both target sizes.

The unchanged static one-icon control remains below its added held-out threshold, so the generic combined gate flag remains false. Apply the same treatment there before closing Gate E.

State transitions

  1. planned2026-08-30T06:05:24.460858Z
  2. completed2026-08-30T06:11:11.007462Zper_step_corruption_coverage_recovers_diverse_heldout_geometry_without_more_model_batch_or_steps

Run record

Verbatim from runs/tiny-geometry-v4-ede1009-7ccbc64e-32a80ab5/run.yaml, the record committed before launch.

schema_version
1
run_id
tiny-geometry-v4-ede1009-7ccbc64e-32a80ab5
state
completed
hypothesis
Deterministic per-step corruption resampling for diverse-four will raise held-out changed-token accuracy to at least 0.75 and retained-token accuracy to at least 0.90 without increasing model size, batch size, corruption probability, or optimizer steps.
expected_information_gain
Tests whether v3 is limited primarily by seeing only four fixed corruptions per icon, before spending evidence budget on more capacity or longer training.
parent_run
tiny-geometry-v3-93ed354-ff42f4dd-32a80ab5
git_commit
ede100987cb4b9dc8a179f36122abf319da9340c
dirty_patch_sha256
e3b0c44298fc1c14…
config
configs/learning/tiny-geometry-v4-resampled-corruption.yaml
config_sha256
7ccbc64e51a0f009…
dataset_manifest
data/manifests/openmoji-17.0.0-tiny-learning-fixture-v1.json
dataset_manifest_sha256
32a80ab576a6d5a4…
worker
gtc-local-cpu
environment
hostname
gtc
architecture
x86_64
python
3.12.3
torch
2.8.0+cpu
numpy
2.5.2
device
cpu
training_threads
2
gpu
None
precision
float32
seeds
1701, 1702
controlled_factor
diverse-four corruption draw coverage
held_constant
one-icon control, fixture and held-out draws, model and initialization, batch size and corruption probability, optimizer and 320-step diverse budget, checkpoint boundary
predeclared_criteria
diverse_four_changed_accuracy_minimum
0.75
diverse_four_retained_accuracy_minimum
0.9
checkpoint_resume_exact
true
artifact_identity_rerun
true
resource_cap
external_spend_usd
0
optimizer_steps
800
batch_size_diverse
16
distinct_corruptions_per_icon_continuous
1280
heldout_examples
24
isolated_renders
24
max_local_storage_gb
0.1
command
.venv/bin/python -m mojidiff.learning.tiny_study --config configs/learning/tiny-geometry-v4-resampled-corruption.yaml
outputs
report
reports/learning/tiny-geometry-v4-resampled-corruption
derived
data/processed/tiny-geometry-v4-resampled-corruption
run_record
runs/tiny-geometry-v4-ede1009-7ccbc64e-32a80ab5
artifact_durability
compact-report-pending-local-git;checkpoint-local-only
result
treatment_hypothesis_passed
true
artifact_identity_rerun
true
checkpoint_resume_exact
true
diverse_four
aggregate_accuracy
0.9856271777003485
changed_accuracy
0.9664634146341463
retained_accuracy
0.9962737127371274
one_icon_control_aggregate_accuracy
0.9356617647058824
generic_combined_gate_flag
false