The project advances from cheap tests to expensive ones. A gate opens only when the preceding evidence justifies it.
Orchestration and recovery
Would a worker loss erase the experiment ledger, or hide an active job?
Run registry, durable artifact sink, and detached launch/status/sync adapters are in place and exercised. The owned-worker adapter smoke round-tripped a hashed artifact twice, and three preceding adapter failures are preserved with reason codes.
Corpus and rights audit
Is every sample, split, and exclusion traceable, hashed, and justified?
The reviewed OpenMoji 17.0.0 primary manifest has 4,006 rows: 3,902 ordinary includes and 104 visually reviewed overrides. It excludes 270 metadata flags, keeps 218 noncanonical aliases in a reversible split, and retains one renderer-failing SVG as a defect row. Independent verification passes.
Representation study
Which typed SVG representation, with what explicit tradeoffs and fallback?
Semantic strokes with a 60-case reason-coded outlined fallback, role-typed q289 endpoints and q417 control handles, K48+1 stroke widths, and packed P80/T1216 capacity. Selected against measured alternatives; the falsified viewBox-clamping and unaugmented-K48 tails are kept as negative results.
Invariant and renderer stress test
Do exposed states serialize safely and render reliably, with failures classified?
200 stable random packed round trips, 2,000/2,000 invalid mutations rejected across ten families, three malformed typed-XML cases classified, and 25 isolated renders including an 80-path/1,216-segment boundary program. The first mutator-boundary failure is preserved.
Tiny learning proof
Is the representation learnable end to end, with exact checkpoint resume?
Closed from a controlled v1-v5 sequence: 100% one-icon overfit, one-icon and diverse held-out recovery, recognizable renders at 72 and 18 px, and exact resume with byte-stable artifacts. Fixed-topology geometry only.
Corruption comparison
Which corruption process is the primary, on measured grounds?
Factorized role-uniform corruption is the primary. Across two seeds it beats path-correlated corruption on changed-token recovery by 12.07 and 12.60 points on a matched v3 fixture. Whole-path replacement reaches 100% but changes far fewer held-out fields, so it is retained as a lighter feasibility ablation.
OpenMoji generation and editing
Does the selected formulation actually generate and edit OpenMoji well, under structured conditioning and locked-path inpainting?
Answered with the honest assessment the gate asks for. After a five-lens diagnostic found four verified defects present throughout - a loss averaged over per-(kind, slot) groups, no attention padding mask, a corruption process that leaked a density shortcut, and coordinates as unordered categories - and after three further corrections (the identity baseline, the distance-kernel target and a decode gate chosen by geometric gain), v16 became the first model here to beat emitting its input unchanged: paired render recovery +0.2511 at p=0.35, +0.3902 at p=0.10, helping 26 of 32 lightly corrupted icons significantly. Data volume matters; capacity, noise conditioning and the training schedule do not, all retested on the repaired measurement. What v16 is: a fixed-topology geometry denoiser that removes some stray strokes from a mostly intact icon. It cannot add, remove or restyle a path and it never predicted topology or style. The assessment: generation is not compelling at this scale (Gate I measured why), blind denoising helps only for light corruption, and editing is the compelling candidate. Locked-path and style inpainting move to Gate L as its whole subject. Every earlier plateau in this gate was measured under the broken loss and is not admissible as read.
Broad pretraining question
Does broader licensed pretraining help, or dilute the style?
Deferred behind Gate L. It is the only route to generation - Gate I's data-scaling arm says the OpenMoji corpus cannot get there at any size it reaches - but it is at most an order of magnitude of data with a licensing audit in front of it, and the next result is an editor. Reopen only for generation, only if Gate L's editor works.
Cached AR comparison
How does the denoiser compare with a matched, KV-cached autoregressive baseline - and does the representation support generation at all?
Answered, negatively and thoroughly. The causal model is built and learns: it memorises four icons exactly through the KV cache and the dynamic legal-token masks. At corpus scale, matched on every term the plan names, it reaches 3.6249 nats per free token against a zero-parameter position-marginal floor of 3.9290 - a ratio of 0.923 where the criterion was 0.500 - and from step 1200 it is worse than the floor while its training loss keeps falling. Five arms then measured every factor available: the coordinate encoding that was decisive for the denoiser on this same codec does nothing, 4x the data buys 0.0055 of ratio, 9x the capacity is monotonically worse, and dropout is the largest effect at 0.0142 and still an order of magnitude short. The renders settle it - eight subgroups, three samples each, beside real icons: valid programs in corpus palette colours, median 8.5 active paths, ink coverage 0.388 against 0.253, and no recognisable shape anywhere. The numbers and the pictures agree. On gpubox-4080: 301 s to train, 3.63 GiB peak, and the KV cache is worth 16% rather than an order of magnitude because at this size the decode is bound by kernel launches rather than arithmetic - a measurement that contradicts the asymptotic argument. The quality half of the plan's comparison is not well-posed between these two models and that is recorded as the result rather than papered over: no metric applies to both without putting one out of distribution. The branch this recommends is an any-order masked model, which would share this codec and harness and could be scored against the denoiser on one task with one metric.
Structured editing with a masked any-order model
Can one bidirectional model over the same codec complete held-out icons - a missing path, a span, a style block - visibly and measurably better than doing nothing and than a zero-parameter sampler that knows the corpus statistics?
Answered in twenty-two registered arms on 2026-09-21. The first corpus model beat the position-marginal policy on 54 of 64 held-out icons and lost to leaving the hole on 57 of 64 because it never read its neighbours - at the floor for coordinates on its own training icons, a 40-bin shift of a visible endpoint moving its prediction by a median of zero. Ownership binding, the training mixture and the output head alone changed nothing. Putting each segment's start point in its own input did, and the metric head on top of it: continuity error 58-60 bins to 17.6-19.9 against the copy-the-previous-endpoint policy's 25.6 over 339 held-out icons, in every run that carried the mechanism - the first learned geometry in this project. Carried into editing: whole-path completion stops damaging (-0.0011 against the hole from -0.0040) and still loses to it; span completion at two and four segments ties a zero-parameter fill that joins the visible ends - five readings on pinned spans with three checkpoints, a three-times-longer budget and a ten-times-larger model, the interval spanning zero every time - and beats the marginal policy on 49 of 56 or better every time. Likelihood and continuity kept improving through all of it; the render-level edit did not follow, because what remains is shape, which the corpus cannot teach for an unseen icon. Two harness defects, two mis-specified criteria, one inherited seed and one wrong reading were caught and preserved. The frozen specialist of arm 21 is the checkpoint for Gate K's viewer.
Deployment optimization
Do the optimizations preserve numerical and visual parity, and pay off?
Gated on a frozen checkpoint worth optimizing.
A pretrained SVG prior, fine-tuned on OpenMoji
Does a pretrained prior, fine-tuned with LoRA on the training split, draw recognisable held-out icons from a caption - read against its own zero-shot output, against chance, and against memorisation?
Answered no on 2026-09-21, three experiments, every run registered. OmniSVG 1.1 4B fine-tuned in its own token language learns OpenMoji's palette and stock parts, not icons: falsified under three decoders. Qwen3.5-2B-Base fine-tuned to write the codec's canonical SVG is the strongest arm - 44 of 64 drawings enter the codec against 3 for its control, a paired CLIP gain of +0.040 with an interval excluding zero - and is falsified on codec validity (0.69 against 0.8) and retrieval (0.094 against 0.25); training on the full corpus at 4,096 tokens is worse. SemIf's decision readout is falsified on both models; averaged over option orders it reaches 0.75, a secondary signal. Retrieval of the right icon among the 32 held-out renders never rose above 0.094 in any arm (chance 0.031). The open direction is conditioning on more than a caption, with the text prior as the vehicle.
A render as the condition
Given an icon's render, does a pretrained prior return the icon closely enough that the right icon is retrieved first among the held-out renders and the codec takes the result, and does fine-tuning the image branch on OpenMoji's own render-to-program pairs close the remaining gap?
Answered 2026-09-23. Yes to the first: OmniSVG's released image branch retrieves the right icon first among 32 held-out renders for 61% of drawings zero-shot and 77% with best-of-six decoding against the input render and 83% with twelve (similarity 0.944 and 0.954, paired gains with intervals excluding zero, no training); edits made in pixels and re-vectorised the same way identify the edited icon 71% of the time with half the pixel error of plain decoding. No to the second: three fine-tuning arms in two dialects, with the merger trainable and selection on free-running drawings, all fitted the corpus and drew worse from their first 50 steps - the failure is the first point and the close, sequence-level decisions teacher forcing cannot reach; a pipeline test on the model's own drawings proves the code sound. Codec validity is capped at 0.25 in this language by outlining. Next: more candidates and a loop stop in decoding; a sequence-level objective if compact programs are needed.
Merging icons, vectors in and vectors out
Fine-tuned on OpenMoji's ZWJ pairs written as text, does the Qwen3.5-2B text prior merge held-out components into the true merged icon better than stacking them and better than keeping the first component?
Answered no on 2026-09-24. The fine-tune beats the overlay (40 of 41 targets closer) and ties the first component alone (interval -0.003 to +0.001), retrieving the true target first for 0.40 where the first component does 0.54: it copies the base figure. The path provenance of all 1,015 merged icons explains it: only 17% of their paths are exact copies of component paths and 40% have no structural match, a median 0.895 of an icon's paths unexplained by its parts - OpenMoji redraws a merge rather than composing it. The kitbash tool covers the quarter that is composition, exactly and instantly.
Research interface and narrative
Does every displayed claim link to a reproducible run and environment?
This weblog is the first piece: a static view over the committed run registry, run records, findings, and renders, served on the machine's Tailscale address, with every Gate L sheet beside its numbers. The editing viewer over Gate L's frozen specialist - clean, join, model and marginal for each held-out span, with the per-span numbers - is the next piece; the paired trajectory viewer and the benchmark explorer are not built.