MojiDiff

← experiments

r2s-full-v8-twemoji-f529d3f-5f240ad1-47646604

completed —

Hypothesis

Adding the Twemoji icons that fit the codec to v7's training data lowers held-out OpenMoji pixel error below v7's at the same budget, greedy decoding, with the paired 95% interval on the per-icon difference excluding zero. It also transcribes held-out validation renders better than retrieving the closest training icon. Pass requires all three: (1) mean 72 px pixel error below the nearest-training-icon baseline with the paired 95% interval on the reduction excluding zero; (2) CLIP top-1 on Gate N's 32 validation icons at or above OmniSVG 4B zero-shot, 0.609; (3) median single-icon latency on the RTX 4080 under 500 ms.

Visual output

eval-validation-rerank8
eval-validation-rerank8.png
eval-validation
eval-validation.png
samples
samples.png

Written result

Hypothesis. Adding the Twemoji icons that fit the codec to v7's training data lowers held-out OpenMoji pixel error below v7's at the same budget, greedy decoding, with the paired 95% interval on the per-icon difference excluding zero. It also transcribes held-out validation renders better than retrieving the closest training icon. Pass requires all three: (1) mean 72 px pixel error below the nearest-training-icon baseline with the paired 95% interval on the reduction excluding zero; (2) CLIP top-1 on Gate N's 32 validation icons at or above OmniSVG 4B zero-shot, 0.609; (3) median single-icon latency on the RTX 4080 under 500 ms.

Result

measurevalue
icons evaluated339 (primary/validation)
model pixel error, mean [95% CI]0.0911 [0.0851, 0.0972]
model pixel error, median0.0890
exact program rate0.015
rendered rate1.000
blank canvas pixel error0.1720
nearest training icon pixel error0.0897 [0.0848, 0.0947]
error reduction vs nearest icon-0.0013 [-0.0062, 0.0035]
icons where model beats nearest icon190
CLIP top-1, model greedy (32 Gate N icons)0.531
CLIP top-1, nearest training icon0.344
CLIP top-1, OmniSVG 4B zero-shot (Gate N)0.609

Resources

measurevalue
deviceNVIDIA GeForce RTX 4080
parameters9,025,826
train seconds5455
peak VRAM GiB4.341529369354248
latency ms/icon, batch 1, median1078.4
latency ms/icon, batch 1, p951082.5
throughput ms/icon, batch 6479.6
model calls per icon (decoding)540.0
torch / CUDA2.14.0a0+4fdf77b940.nv26.08 / 13.4

Latency covers encoding and decoding to a validated token program; it excludes rasterising the SVG. Samples: samples.png, rows are reference, model, nearest training icon.

State transitions

  1. running2026-09-28T01:59:21Z
  2. completed2026-09-28T03:31:33Z

Run record

Verbatim from runs/r2s-full-v8-twemoji-f529d3f-5f240ad1-47646604/run.yaml, the record committed before launch.

schema_version
2
run_id
r2s-full-v8-twemoji-f529d3f-5f240ad1-47646604
state
completed
planned_at
2026-09-28T01:59:21Z
hypothesis
Adding the Twemoji icons that fit the codec to v7's training data lowers held-out OpenMoji pixel error below v7's at the same budget, greedy decoding, with the paired 95% interval on the per-icon difference excluding zero. It also transcribes held-out validation renders better than retrieving the closest training icon. Pass requires all three: (1) mean 72 px pixel error below the nearest-training-icon baseline with the paired 95% interval on the reduction excluding zero; (2) CLIP top-1 on Gate N's 32 validation icons at or above OmniSVG 4B zero-shot, 0.609; (3) median single-icon latency on the RTX 4080 under 500 ms.
parent_run
r2s-full-v7-systems
git_commit
f529d3f89030f90018e0473ef060b33bf01403d0
dirty_patch_sha256
1c3a590cdad4c1ae…
config
configs/render2svg/full-v8-twemoji.yaml
config_sha256
5f240ad15add7591…
config_resolved
model
image_size
144
d_model
256
heads
8
encoder_layers
2
decoder_layers
6
feedforward
1024
dropout
0.1
metric
true
fourier
10
order
path
training
steps
60000
batch_size
32
learning_rate
0.0005
weight_decay
0.01
warmup_steps
1000
seed
7001
eval_every
2000
eval_icons
64
train_icons
None
evaluate_on_train
false
augment_variants
0
augment_seed
9001
augment_mirror
0.5
augment_max_shift
48
augment_colour
0.5
augment_online
true
compose_probability
0.5
compose_parts
4
augment_original
0.1
loader_workers
16
extra_training
twemoji
bf16
true
final_eval_icons
339
clip_icons
32
dataset
pilot_config
configs/learning/openmoji-g1-geometric-gate-v16.yaml
cache_sha256
476466042da98d72…
train_icons
4892
augmented_variants
0
extra_training
format
1
commit
b6b55fef1e8636b540a6d016a4729ca8cdf2e60b
size
144
palette
#000000, #ffffff, #92d3f5, #61b2e4, #ea5a47, #d22f27 … and 22 more
slots
128
excluded_bases
354
candidates
3375
source_files
4009
excluded_as_held_out
634
fit_and_kept
2211
cache
/home/dev/.cache/mojidiff/render2svg-extra-twemoji-23878da09a9c.npz
cache_sha256
b263d60cbb1ed1ad…
augmented_sha256
None
evaluated_on
primary/validation
seed
7001
determinism
seeded; cuDNN and SDPA kernels not forced deterministic
model_parameters
9025826
command
.venv/bin/python -m mojidiff.learning.render2svg --config configs/render2svg/full-v8-twemoji.yaml
outputs
run_dir
runs/r2s-full-v8-twemoji-f529d3f-5f240ad1-47646604
checkpoints
/home/dev/.cache/mojidiff/runs/r2s-full-v8-twemoji-f529d3f-5f240ad1-47646604
notes
completed_at
2026-09-28T03:31:33Z
resource
device
NVIDIA GeForce RTX 4080
torch_version
2.14.0a0+4fdf77b940.nv26.08
cuda_version
13.4
python_version
3.12.3
peak_vram_gib
4.341529369354248
train_seconds
5454.982071467093
inference_ms_per_icon
1078.350827039685
inference_p95_ms_per_icon
1082.5091790175065
icons_per_second
0.927341988270389
batched_ms_per_icon
79.62540853077371
batched_icons_per_call
64
excludes
rendering the output SVG to pixels
baselines
blank_canvas_pixel_error
0.1720480552069557
nearest_training_icon_pixel_error
0.08973124639576185
gate_n_omnisvg_zero_shot_top1
0.609375
nearest_training_icon_clip_top1
0.34375
result
icons
339
model_pixel_error
0.0910644995288577, 0.08508873070793221, 0.0972212718477717
model_pixel_error_median
0.08895394206047058
exact_program_rate
0.014749262536873156
rendered_rate
1.0
blank_pixel_error
0.1720480552069557
selected_step
52000
nearest_training_icon_with_extra_pixel_error
0.0875778230639301, 0.08281926003772477, 0.09234717615560617
nearest_training_icon_pixel_error
0.08973124639576185, 0.0848205749792233, 0.09466393787711656
model_minus_baseline_error_reduction
-0.0013332531330958496, -0.006230565753120892, 0.003473313955883553
icons_model_beats_baseline
190
clip_retrieval
model_greedy
top1_rate
0.53125
top5_rate
0.78125
mean_rank
4.40625
clip_to_reference_mean
0.8974534403532743
chance_top1
0.03125
icons
32
nearest_training_icon
top1_rate
0.34375
top5_rate
0.59375
mean_rank
7.40625
clip_to_reference_mean
0.8630629815161228
chance_top1
0.03125
icons
32
gate_n_omnisvg_zero_shot_top1
0.609375
gate_n_omnisvg_best_of_12_top1
0.828125
model_calls_per_icon
540