MojiDiff

MojiDiff · categorical denoising over typed SVG programs

Research weblog

Gate L is complete and the premise has changed. Twenty-two arms showed what a model trained from nothing on 2,681 icons learns - colour, style, continuity - and what it cannot: what an unseen icon's parts look like. Gate M starts from a model that already knows what shapes look like: Qwen3-4B-Base, Apache 2.0 at a pinned revision, fine-tuned with LoRA to write the project's canonical SVG text from a caption, every output parsed by the typed codec and rendered by the safe renderer, read against its own zero-shot output and against memorisation on family-disjoint held-out icons.

1planned
133completed
15failed
1cancelled
155rendered artifacts

Latest decisions

Newest first, verbatim openings from reports/findings.md.

2026-10-04 — Public release; in-browser inference is feasible: v9 greedy at about 1.2 s per icon on WASM

Release. The GitHub repo cpietsch/emojidiff is public under CC BY-SA 4.0 (OpenMoji and Twemoji attributed). The weblog deploys to https://cpietsch.github.io/emojidiff/ from GitHub Actions on every push to main (it needs only PyYAML). All checkpoints are mirrored at https://huggingface.co/chrispie/mojidiff-checkpoints (sha256 index reports/checkpoints-hf.json, verified against the remote). The demos were packa

2026-10-03 — vt-v2-c16 fails A1 again: the canvas line stops; the exploratory flow prior fails B1, B2, B4, B5

Both jobs finished on 2026-09-29 before a power failure; recorded on 2026-10-03 from the run directory and the harness reports, which survived intact (run-record audit consistent).

2026-09-29 — Blinding fixes the collapse; the canvas latent reconstructs near v9

The harness (latent_metrics, copy rule without-twins; reports in reports/latent/) scored the three program-latent models on 339 prior samples, 339 reconstructions and 32 validation pairs x 9 frames, against reference rows. Plain N(0, I) prior:

Full decision trail →

Evidence gates

The project advances from cheap tests to expensive ones. A gate opens only when the preceding evidence justifies it.

Gate Acomplete

Orchestration and recovery

Would a worker loss erase the experiment ledger, or hide an active job?

Run registry, durable artifact sink, and detached launch/status/sync adapters are in place and exercised. The owned-worker adapter smoke round-tripped a hashed artifact twice, and three preceding adapter failures are preserved with reason codes.

Gate Bcomplete

Corpus and rights audit

Is every sample, split, and exclusion traceable, hashed, and justified?

The reviewed OpenMoji 17.0.0 primary manifest has 4,006 rows: 3,902 ordinary includes and 104 visually reviewed overrides. It excludes 270 metadata flags, keeps 218 noncanonical aliases in a reversible split, and retains one renderer-failing SVG as a defect row. Independent verification passes.

Gate Ccomplete

Representation study

Which typed SVG representation, with what explicit tradeoffs and fallback?

Semantic strokes with a 60-case reason-coded outlined fallback, role-typed q289 endpoints and q417 control handles, K48+1 stroke widths, and packed P80/T1216 capacity. Selected against measured alternatives; the falsified viewBox-clamping and unaugmented-K48 tails are kept as negative results.

Gate Dcomplete

Invariant and renderer stress test

Do exposed states serialize safely and render reliably, with failures classified?

200 stable random packed round trips, 2,000/2,000 invalid mutations rejected across ten families, three malformed typed-XML cases classified, and 25 isolated renders including an 80-path/1,216-segment boundary program. The first mutator-boundary failure is preserved.

Gate Ecomplete

Tiny learning proof

Is the representation learnable end to end, with exact checkpoint resume?

Closed from a controlled v1-v5 sequence: 100% one-icon overfit, one-icon and diverse held-out recovery, recognizable renders at 72 and 18 px, and exact resume with byte-stable artifacts. Fixed-topology geometry only.

Gate Fcomplete

Corruption comparison

Which corruption process is the primary, on measured grounds?

Factorized role-uniform corruption is the primary. Across two seeds it beats path-correlated corruption on changed-token recovery by 12.07 and 12.60 points on a matched v3 fixture. Whole-path replacement reaches 100% but changes far fewer held-out fields, so it is retained as a lighter feasibility ablation.

Gate Gcomplete

OpenMoji generation and editing

Does the selected formulation actually generate and edit OpenMoji well, under structured conditioning and locked-path inpainting?

Answered with the honest assessment the gate asks for. After a five-lens diagnostic found four verified defects present throughout - a loss averaged over per-(kind, slot) groups, no attention padding mask, a corruption process that leaked a density shortcut, and coordinates as unordered categories - and after three further corrections (the identity baseline, the distance-kernel target and a decode gate chosen by geometric gain), v16 became the first model here to beat emitting its input unchanged: paired render recovery +0.2511 at p=0.35, +0.3902 at p=0.10, helping 26 of 32 lightly corrupted icons significantly. Data volume matters; capacity, noise conditioning and the training schedule do not, all retested on the repaired measurement. What v16 is: a fixed-topology geometry denoiser that removes some stray strokes from a mostly intact icon. It cannot add, remove or restyle a path and it never predicted topology or style. The assessment: generation is not compelling at this scale (Gate I measured why), blind denoising helps only for light corruption, and editing is the compelling candidate. Locked-path and style inpainting move to Gate L as its whole subject. Every earlier plateau in this gate was measured under the broken loss and is not admissible as read.

Gate Hnot-started

Broad pretraining question

Does broader licensed pretraining help, or dilute the style?

Deferred behind Gate L. It is the only route to generation - Gate I's data-scaling arm says the OpenMoji corpus cannot get there at any size it reaches - but it is at most an order of magnitude of data with a licensing audit in front of it, and the next result is an editor. Reopen only for generation, only if Gate L's editor works.

Gate Icomplete

Cached AR comparison

How does the denoiser compare with a matched, KV-cached autoregressive baseline - and does the representation support generation at all?

Answered, negatively and thoroughly. The causal model is built and learns: it memorises four icons exactly through the KV cache and the dynamic legal-token masks. At corpus scale, matched on every term the plan names, it reaches 3.6249 nats per free token against a zero-parameter position-marginal floor of 3.9290 - a ratio of 0.923 where the criterion was 0.500 - and from step 1200 it is worse than the floor while its training loss keeps falling. Five arms then measured every factor available: the coordinate encoding that was decisive for the denoiser on this same codec does nothing, 4x the data buys 0.0055 of ratio, 9x the capacity is monotonically worse, and dropout is the largest effect at 0.0142 and still an order of magnitude short. The renders settle it - eight subgroups, three samples each, beside real icons: valid programs in corpus palette colours, median 8.5 active paths, ink coverage 0.388 against 0.253, and no recognisable shape anywhere. The numbers and the pictures agree. On gpubox-4080: 301 s to train, 3.63 GiB peak, and the KV cache is worth 16% rather than an order of magnitude because at this size the decode is bound by kernel launches rather than arithmetic - a measurement that contradicts the asymptotic argument. The quality half of the plan's comparison is not well-posed between these two models and that is recorded as the result rather than papered over: no metric applies to both without putting one out of distribution. The branch this recommends is an any-order masked model, which would share this codec and harness and could be scored against the denoiser on one task with one metric.

Gate Lcomplete

Structured editing with a masked any-order model

Can one bidirectional model over the same codec complete held-out icons - a missing path, a span, a style block - visibly and measurably better than doing nothing and than a zero-parameter sampler that knows the corpus statistics?

Answered in twenty-two registered arms on 2026-09-21. The first corpus model beat the position-marginal policy on 54 of 64 held-out icons and lost to leaving the hole on 57 of 64 because it never read its neighbours - at the floor for coordinates on its own training icons, a 40-bin shift of a visible endpoint moving its prediction by a median of zero. Ownership binding, the training mixture and the output head alone changed nothing. Putting each segment's start point in its own input did, and the metric head on top of it: continuity error 58-60 bins to 17.6-19.9 against the copy-the-previous-endpoint policy's 25.6 over 339 held-out icons, in every run that carried the mechanism - the first learned geometry in this project. Carried into editing: whole-path completion stops damaging (-0.0011 against the hole from -0.0040) and still loses to it; span completion at two and four segments ties a zero-parameter fill that joins the visible ends - five readings on pinned spans with three checkpoints, a three-times-longer budget and a ten-times-larger model, the interval spanning zero every time - and beats the marginal policy on 49 of 56 or better every time. Likelihood and continuity kept improving through all of it; the render-level edit did not follow, because what remains is shape, which the corpus cannot teach for an unseen icon. Two harness defects, two mis-specified criteria, one inherited seed and one wrong reading were caught and preserved. The frozen specialist of arm 21 is the checkpoint for Gate K's viewer.

Gate Jnot-started

Deployment optimization

Do the optimizations preserve numerical and visual parity, and pay off?

Gated on a frozen checkpoint worth optimizing.

Gate Mcomplete

A pretrained SVG prior, fine-tuned on OpenMoji

Does a pretrained prior, fine-tuned with LoRA on the training split, draw recognisable held-out icons from a caption - read against its own zero-shot output, against chance, and against memorisation?

Answered no on 2026-09-21, three experiments, every run registered. OmniSVG 1.1 4B fine-tuned in its own token language learns OpenMoji's palette and stock parts, not icons: falsified under three decoders. Qwen3.5-2B-Base fine-tuned to write the codec's canonical SVG is the strongest arm - 44 of 64 drawings enter the codec against 3 for its control, a paired CLIP gain of +0.040 with an interval excluding zero - and is falsified on codec validity (0.69 against 0.8) and retrieval (0.094 against 0.25); training on the full corpus at 4,096 tokens is worse. SemIf's decision readout is falsified on both models; averaged over option orders it reaches 0.75, a secondary signal. Retrieval of the right icon among the 32 held-out renders never rose above 0.094 in any arm (chance 0.031). The open direction is conditioning on more than a caption, with the text prior as the vehicle.

Gate Ncomplete

A render as the condition

Given an icon's render, does a pretrained prior return the icon closely enough that the right icon is retrieved first among the held-out renders and the codec takes the result, and does fine-tuning the image branch on OpenMoji's own render-to-program pairs close the remaining gap?

Answered 2026-09-23. Yes to the first: OmniSVG's released image branch retrieves the right icon first among 32 held-out renders for 61% of drawings zero-shot and 77% with best-of-six decoding against the input render and 83% with twelve (similarity 0.944 and 0.954, paired gains with intervals excluding zero, no training); edits made in pixels and re-vectorised the same way identify the edited icon 71% of the time with half the pixel error of plain decoding. No to the second: three fine-tuning arms in two dialects, with the merger trainable and selection on free-running drawings, all fitted the corpus and drew worse from their first 50 steps - the failure is the first point and the close, sequence-level decisions teacher forcing cannot reach; a pipeline test on the model's own drawings proves the code sound. Codec validity is capped at 0.25 in this language by outlining. Next: more candidates and a loop stop in decoding; a sequence-level objective if compact programs are needed.

Gate Ocomplete

Merging icons, vectors in and vectors out

Fine-tuned on OpenMoji's ZWJ pairs written as text, does the Qwen3.5-2B text prior merge held-out components into the true merged icon better than stacking them and better than keeping the first component?

Answered no on 2026-09-24. The fine-tune beats the overlay (40 of 41 targets closer) and ties the first component alone (interval -0.003 to +0.001), retrieving the true target first for 0.40 where the first component does 0.54: it copies the base figure. The path provenance of all 1,015 merged icons explains it: only 17% of their paths are exact copies of component paths and 40% have no structural match, a median 0.895 of an icon's paths unexplained by its parts - OpenMoji redraws a merge rather than composing it. The kitbash tool covers the quarter that is composition, exactly and instantly.

Gate Kopen

Research interface and narrative

Does every displayed claim link to a reproducible run and environment?

This weblog is the first piece: a static view over the committed run registry, run records, findings, and renders, served on the machine's Tailscale address, with every Gate L sheet beside its numbers. The editing viewer over Gate L's frozen specialist - clean, join, model and marginal for each held-out span, with the per-span numbers - is the next piece; the paired trajectory viewer and the benchmark explorer are not built.

Recent experiments

Newest first. Failed and superseded runs are kept and shown; a run whose predeclared criterion was falsified is a result, not a mistake.

runstatepredeclared outcomerecorded
vt-v2-c16-0f4d5de-1497d4b5-47646604completed—2026-09-29 16:44
lfp-v1x-exploratory-81eb2e2-5ae4d777-47646604completed—2026-09-29 15:48
vt-v1-74dfbd2-46380052-47646604completed—2026-09-29 14:44
latent-v3-blind-e254c16-9c7db490-47646604completed—2026-09-29 13:20
r2s-full-v9-colour-621bc8e-b26ad95e-47646604completed—2026-09-28 06:31
latent-v2-kl-efbd2e2-451268e6-47646604completed—2026-09-28 05:03
r2s-full-v8-twemoji-f529d3f-5f240ad1-47646604completed—2026-09-28 03:31
latent-v1-8b0d3f9-6e874509-47646604completed—2026-09-28 01:59
r2s-full-v7-systems-42762c2-e9842024-47646604completed—2026-09-28 00:40
r2s-full-v6-compose-7505e5b-6472dcf0-47646604completed—2026-09-27 23:03
r2s-full-v5-path-7a72de9-12766043-47646604completed—2026-09-27 22:28
r2s-full-v4-online-c0fe6d3-a854be2e-47646604completed—2026-09-27 21:57

All 150 runs →

Visual evidence

Renders the claims are checkable against: corpus contact sheets, codec worst cases, and paired corruption/prediction trajectories.

grid-comparison-semantic-18
grid-comparison-semantic-18.png
worst-style-delta-18
worst-style-delta-18.png
worst-style-delta-18
worst-style-delta-18.png
worst-18-q256
worst-18-q256.png
hist-alpha_fraction_18
hist-alpha_fraction_18.png
hist-alpha_fraction_18
hist-alpha_fraction_18.png
hist-alpha_fraction_18
hist-alpha_fraction_18.png
hist-alpha_fraction_18
hist-alpha_fraction_18.png

Full gallery →