r2s-full-v3-metric-5dba9d7-ffcc2ffb-47646604
completed —
Hypothesis
Trained on the 2,681 family-disjoint training icons, the 8.9M-parameter render-to-program model transcribes held-out validation renders better than retrieving the closest training icon. Pass requires all three: (1) mean 72 px pixel error below the nearest-training-icon baseline with the paired 95% interval on the reduction excluding zero; (2) CLIP top-1 on Gate N's 32 validation icons at or above OmniSVG 4B zero-shot, 0.609; (3) median single-icon latency on the RTX 4080 under 500 ms.
Visual output
Written result
Hypothesis. Trained on the 2,681 family-disjoint training icons, the 8.9M-parameter render-to-program model transcribes held-out validation renders better than retrieving the closest training icon. Pass requires all three: (1) mean 72 px pixel error below the nearest-training-icon baseline with the paired 95% interval on the reduction excluding zero; (2) CLIP top-1 on Gate N's 32 validation icons at or above OmniSVG 4B zero-shot, 0.609; (3) median single-icon latency on the RTX 4080 under 500 ms.
Result
| measure | value |
|---|---|
| icons evaluated | 339 (primary/validation) |
| model pixel error, mean [95% CI] | 0.1064 [0.1005, 0.1126] |
| model pixel error, median | 0.1078 |
| exact program rate | 0.006 |
| rendered rate | 1.000 |
| blank canvas pixel error | 0.1720 |
| nearest training icon pixel error | 0.0897 [0.0848, 0.0947] |
| error reduction vs nearest icon | -0.0167 [-0.0218, -0.0116] |
| icons where model beats nearest icon | 138 |
| CLIP top-1, model greedy (32 Gate N icons) | 0.219 |
| CLIP top-1, nearest training icon | 0.344 |
| CLIP top-1, OmniSVG 4B zero-shot (Gate N) | 0.609 |
Resources
| measure | value |
|---|---|
| device | NVIDIA GeForce RTX 4080 |
| parameters | 9,011,746 |
| train seconds | 1841 |
| peak VRAM GiB | 4.340941905975342 |
| latency ms/icon, batch 1, median | 835.0 |
| latency ms/icon, batch 1, p95 | 837.3 |
| throughput ms/icon, batch 64 | 60.2 |
| model calls per icon (decoding) | 445.0 |
| torch / CUDA | 2.14.0a0+4fdf77b940.nv26.08 / 13.4 |
Latency covers encoding and decoding to a validated token program; it excludes rasterising the SVG. Samples: samples.png, rows are reference, model, nearest training icon.
State transitions
- running2026-09-27T20:55:16Z
- completed2026-09-27T21:27:02Z
Run record
Verbatim from runs/r2s-full-v3-metric-5dba9d7-ffcc2ffb-47646604/run.yaml, the record committed before launch.
9d6a9f3415f8e8bf…ffcc2ffb9e6f21c9…476466042da98d72…35c7ef3ccc5bb6e3…
