MojiDiff

← experiments

tiny-geometry-v3-93ed354-ff42f4dd-32a80ab5

completed —

Hypothesis

The diverse-four aggregate held-out accuracy materially overstates genuine denoising: accuracy on actually changed coordinate tokens will trail aggregate accuracy by at least 0.15, while all v2 learning outputs remain otherwise unchanged.

Why run it

Determines whether the next Gate E experiment should target corrupted-token recovery rather than model capacity, additional optimization steps, or topology.

Measured behaviour

Training token accuracy (per batch)
accuracyoptimizer step00.250.50.751100200300
Table view
optimizer steptrain accuracy
10.0012
100.2892
200.6366
300.8676
400.9694
500.9975
601.0000
701.0000
801.0000
901.0000
1001.0000
1101.0000
1201.0000
1301.0000
1401.0000
1501.0000
1601.0000
10.0028
100.1141
200.4179
300.7195
400.8920
500.9649
600.9948
700.9993
801.0000
901.0000
1001.0000
1101.0000
1201.0000
1301.0000
1401.0000
1501.0000
1601.0000
1701.0000
1801.0000
1901.0000
2001.0000
2101.0000
2201.0000
2301.0000
2401.0000
2501.0000
2601.0000
2701.0000
2801.0000
2901.0000
3001.0000
3101.0000
3201.0000

Measured result

Verbatim from the run's summary.json.

cases
diverse-four
checkpoint_sha256
52aee590c65f81f5…
config
corruptions_per_icon
4
heldout_corruptions_per_icon
4
icon_count
4
max_loss_ratio
0.2
min_heldout_accuracy
0.85
min_train_accuracy
0.98
name
diverse-four
resume_step
160
seed
1702
steps
320
continuous_loss_sha256
240bd2ca34fd1b1a…
final
accuracy
0.7382404181184669
changed_accuracy
0.5817073170731707
changed_correct
954
changed_total
1640
correct
3390
loss
1.3494370430707932
retained_accuracy
0.8252032520325203
retained_correct
2436
retained_total
2952
total
4592
final_model_sha256
bd0b0814a19c7402…
heldout_examples
16
initial
accuracy
0.002395470383275261
changed_accuracy
0.0024390243902439024
changed_correct
4
changed_total
1640
correct
11
loss
14.943475902080536
retained_accuracy
0.0023712737127371273
retained_correct
7
retained_total
2952
total
4592
loss_ratio
5.3330747233888907e-05
passes_predeclared_criteria
false
resume_exact
true
resumed_loss_sha256
240bd2ca34fd1b1a…
train_examples
16
train_final
accuracy
1.0
changed_accuracy
1.0
changed_correct
1622
changed_total
1622
correct
4592
loss
0.0007702835428062826
retained_accuracy
1.0
retained_correct
2970
retained_total
2970
total
4592
one-icon
checkpoint_sha256
None
config
corruptions_per_icon
8
heldout_corruptions_per_icon
8
icon_count
1
max_loss_ratio
0.1
min_heldout_accuracy
0.95
min_train_accuracy
0.99
name
one-icon
resume_step
None
seed
1701
steps
160
continuous_loss_sha256
None
final
accuracy
0.9356617647058824
changed_accuracy
0.8858695652173914
changed_correct
489
changed_total
552
correct
1527
loss
0.2106002140790224
retained_accuracy
0.9611111111111111
retained_correct
1038
retained_total
1080
total
1632
final_model_sha256
8e0a9878092058a0…
heldout_examples
8
initial
accuracy
0.001838235294117647
changed_accuracy
0.005434782608695652
changed_correct
3
changed_total
552
correct
3
loss
15.358022212982178
retained_accuracy
0.0
retained_correct
0
retained_total
1080
total
1632
loss_ratio
9.539973069621893e-05
passes_predeclared_criteria
false
resume_exact
None
resumed_loss_sha256
None
train_examples
8
train_final
accuracy
1.0
changed_accuracy
1.0
changed_correct
599
changed_total
599
correct
1632
loss
0.001458597937016748
retained_accuracy
1.0
retained_correct
1033
retained_total
1033
total
1632
code_identity
geometry_sha256
8f30cf4c834873bd…
git_commit
93ed3542a69f1697c4d3ca13ba1982cc2cda5828
packed_sha256
9b12879ba3a18236…
program_sha256
0d02b532b1dbe573…
tiny_study_sha256
578ec1d402fdb88f…
config_sha256
ff42f4dd0b8915f0…
deterministic_algorithms
true
device
cpu
fixture
data/manifests/openmoji-17.0.0-tiny-learning-fixture-v1.json
fixture_sha256
32a80ab576a6d5a4…
metrics_sha256
c4e0d9a8c7873466…
model
d_model
48
feedforward
96
heads
4
layers
2
parameters
241072
scope
fixed-topology geometry-only diagnostic
passes_predeclared_gate_candidate
false
render_metrics_sha256
d4c36c75381689fb…
schema_version
1
study_version
tiny-geometry-v3-corruption-split
torch_version
2.8.0+cpu

Visual output

trajectory
trajectory.png

Written result

Completed reproducibly; the predeclared diagnostic passed.

For diverse-four, actually-changed-token accuracy is 58.17%, 15.65 points below the 73.82% aggregate. Retained-token accuracy is also limited at 82.52%, so the model often overwrites coordinates already correct in x_t. The one-icon case is much stronger: 88.59% changed and 96.11% retained accuracy. A complete rerun reproduced all artifacts.

This supports testing broader deterministic corruption coverage at the same batch size, model, probability, and optimizer-step count before increasing capacity or steps.

State transitions

  1. planned2026-08-30T05:55:54.629673Z
  2. completed2026-08-30T06:01:12.713902Zaggregate_accuracy_materially_overstates_changed_token_recovery_and_diverse_model_also_overwrites_retained_coordinates

Run record

Verbatim from runs/tiny-geometry-v3-93ed354-ff42f4dd-32a80ab5/run.yaml, the record committed before launch.

schema_version
1
run_id
tiny-geometry-v3-93ed354-ff42f4dd-32a80ab5
state
completed
hypothesis
The diverse-four aggregate held-out accuracy materially overstates genuine denoising: accuracy on actually changed coordinate tokens will trail aggregate accuracy by at least 0.15, while all v2 learning outputs remain otherwise unchanged.
expected_information_gain
Determines whether the next Gate E experiment should target corrupted-token recovery rather than model capacity, additional optimization steps, or topology.
parent_run
tiny-geometry-v2-47811d0-75d558c9-32a80ab5
git_commit
93ed3542a69f1697c4d3ca13ba1982cc2cda5828
dirty_patch_sha256
e3b0c44298fc1c14…
config
configs/learning/tiny-geometry-v3-corruption-split.yaml
config_sha256
ff42f4dd0b8915f0…
dataset_manifest
data/manifests/openmoji-17.0.0-tiny-learning-fixture-v1.json
dataset_manifest_sha256
32a80ab576a6d5a4…
worker
gtc-local-cpu
environment
hostname
gtc
architecture
x86_64
python
3.12.3
torch
2.8.0+cpu
numpy
2.5.2
device
cpu
training_threads
2
gpu
None
precision
float32
seeds
1701, 1702
controlled_factor
evaluation metric partition only
held_constant
fixture, model and initialization, corruption draws, optimizer and batches, training steps, aggregate acceptance thresholds
predeclared_diagnostic
diverse_four_changed_accuracy_gap_minimum
0.15
resource_cap
external_spend_usd
0
optimizer_steps
800
train_examples
24
heldout_examples
24
isolated_renders
24
max_local_storage_gb
0.1
command
.venv/bin/python -m mojidiff.learning.tiny_study --config configs/learning/tiny-geometry-v3-corruption-split.yaml
outputs
report
reports/learning/tiny-geometry-v3-corruption-split
derived
data/processed/tiny-geometry-v3-corruption-split
run_record
runs/tiny-geometry-v3-93ed354-ff42f4dd-32a80ab5
artifact_durability
compact-report-pending-local-git;checkpoint-local-only
result
diagnostic_hypothesis_passed
true
artifact_identity_rerun
true
diverse_four
aggregate_accuracy
0.7382404181184669
changed_accuracy
0.5817073170731707
retained_accuracy
0.8252032520325203
aggregate_minus_changed
0.1565331010452962
one_icon
aggregate_accuracy
0.9356617647058824
changed_accuracy
0.8858695652173914
retained_accuracy
0.9611111111111111