MojiDiff

← experiments

lfp-v1x-exploratory-81eb2e2-5ae4d777-47646604

completed —

Hypothesis

A rectified-flow prior fitted to the aggregate posterior of the frozen canvas latent (vt-v1, c = 8) replaces its N(0, I) prior, whose draws decode to fragments, and so generates whole, novel OpenMoji-like drawings and interpolates through plausible drawings. The prior is a 9.7M-parameter DiT over standardised (8, 18, 18) grids (x_t = (1 - t) x_0 + t eps, velocity target, logit-normal t, EMA 0.999), trained on posterior draws of 32 exact variants per training icon (no compositions) and selected by EMA flow loss on the validation icons' latents. Samples: 50 Euler steps from noise scale 1 to z_hat_0, greedy IEEE float32 decoding by the VT. Scored afterwards by latent_metrics through the gallery backend canvas-flow on its fixed sets (339 samples, seed 23; 32 validation pairs x 9 frames, rng 17; CLIP ViT-B/32 at 72 px against the 339 validation renders; bootstrap intervals, paired where paired; copy rule without-twins): B1-B3 on the report with default settings (its samples block); B4 and B5 on the report run with --interpolation backend, whose interpolation block is canvas-flow's own path, the slerp below (the report records interpolation.path slerp-through-prior-noise). The default report's interpolation block is the frozen VT's lerp of posterior means and scores neither B4 nor B5. Pass requires B1-B3; B4 and B5 are the interpolation claims. (B1) fragment rate at most 10%. (B2) CLIP k-NN precision at least 2x the best latent-v2 prior row (N(0, I) or refit), and recall above it, both with intervals excluding zero. (B3) copy rate at most 10% and all 339 samples distinct. (B4) slerp through the prior's noise space (reverse-ODE inversion of both posterior means, slerp, forward integration): interior fragment rate at most 15%, jump share below latent-v2's (paired over pairs, interval excluding zero), detour rate at most 20%. (B5) slerp interior precision above v9 on pixel crossfades of the same pairs, paired, interval excluding zero; a failure is recorded as "the latent adds nothing over crossfade plus transcriber". Decision: B1-B3 pass - the gallery's main latent model. B3 fails - add compositions to the prior's data, then an earlier-stopped prior. B1 or B2 fails while B3 passes - a larger prior. Only B5 fails - keep the model; its interpolation is a dissolve. Reported without a criterion: validation flow loss against the Gaussian velocity, vt-v1's N(0, I) samples on the same seeds, lerp strips, watch-it-draw strips, ms per sample.

Visual output

lerp-interpolations
lerp-interpolations.png
prior-samples
prior-samples.png
slerp-interpolations
slerp-interpolations.png
vt-normal-samples
vt-normal-samples.png
watch-it-draw
watch-it-draw.png

Written result

Hypothesis. A rectified-flow prior fitted to the aggregate posterior of the frozen canvas latent (vt-v1, c = 8) replaces its N(0, I) prior, whose draws decode to fragments, and so generates whole, novel OpenMoji-like drawings and interpolates through plausible drawings. The prior is a 9.7M-parameter DiT over standardised (8, 18, 18) grids (x_t = (1 - t) x_0 + t eps, velocity target, logit-normal t, EMA 0.999), trained on posterior draws of 32 exact variants per training icon (no compositions) and selected by EMA flow loss on the validation icons' latents. Samples: 50 Euler steps from noise scale 1 to z_hat_0, greedy IEEE float32 decoding by the VT. Scored afterwards by latent_metrics through the gallery backend canvas-flow on its fixed sets (339 samples, seed 23; 32 validation pairs x 9 frames, rng 17; CLIP ViT-B/32 at 72 px against the 339 validation renders; bootstrap intervals, paired where paired; copy rule without-twins): B1-B3 on the report with default settings (its samples block); B4 and B5 on the report run with --interpolation backend, whose interpolation block is canvas-flow's own path, the slerp below (the report records interpolation.path slerp-through-prior-noise). The default report's interpolation block is the frozen VT's lerp of posterior means and scores neither B4 nor B5. Pass requires B1-B3; B4 and B5 are the interpolation claims. (B1) fragment rate at most 10%. (B2) CLIP k-NN precision at least 2x the best latent-v2 prior row (N(0, I) or refit), and recall above it, both with intervals excluding zero. (B3) copy rate at most 10% and all 339 samples distinct. (B4) slerp through the prior's noise space (reverse-ODE inversion of both posterior means, slerp, forward integration): interior fragment rate at most 15%, jump share below latent-v2's (paired over pairs, interval excluding zero), detour rate at most 20%. (B5) slerp interior precision above v9 on pixel crossfades of the same pairs, paired, interval excluding zero; a failure is recorded as "the latent adds nothing over crossfade plus transcriber". Decision: B1-B3 pass - the gallery's main latent model. B3 fails - add compositions to the prior's data, then an earlier-stopped prior. B1 or B2 fails while B3 passes - a larger prior. Only B5 fails - keep the model; its interpolation is a dissolve. Reported without a criterion: validation flow loss against the Gaussian velocity, vt-v1's N(0, I) samples on the same seeds, lerp strips, watch-it-draw strips, ms per sample.

Criteria

B1-B5 are scored after this run by latent_metrics through the gallery backend canvas-flow (339 samples, 32 pairs, CLIP, copy rule without-twins): B1-B3 on its default report, B4 and B5 on its --interpolation backend report, whose interpolation block is the slerp through the prior's noise (the default report's is the frozen VT's lerp). The numbers below are this run's own checks, not the criteria.

criterionclaim
B1fragment rate at most 10%
B2CLIP precision at least 2x the best latent-v2 prior row, and recall above it, both with intervals excluding zero
B3copy rate at most 10% under copy rule without-twins, and all samples distinct
B4slerp interior fragments at most 15%, jump share below latent-v2's (paired), detours at most 20%
B5slerp interior CLIP precision above v9 on crossfades, paired

Result (IEEE float32 greedy decoding (TF32 off: matmul and cuDNN); pixel metrics at 72 px)

measurevalue
validation flow loss, EMA (served)1.2587
validation flow loss, Gaussian velocity (zero parameters)1.7629
selected step38000
flow samples distinct / rendered64 / 64 of 64
flow samples with ink < 0.100.141
parent VT, same N(0, I) seeds, ink < 0.100.328
validation icons with ink < 0.100.021
slerp interior frames with ink < 0.100.089
lerp interior frames with ink < 0.100.071
slerp endpoint pixel error (inversion round trip)0.0923 [0.0678, 0.1188]
lerp endpoint pixel error (VT from mu)0.0875 [0.0627, 0.1159]

Resources

measurevalue
deviceNVIDIA GeForce RTX 4080
flow parameters (trained)9,747,744
served parameters (flow + frozen VT)18,779,986
train seconds1179
peak VRAM GiB3.4414305686950684
flow + graph decode, ms per sample, batch 11048.4
flow alone, ms per sample, batch 160.3
batched, ms per sample143.0
torch / CUDA2.14.0a0+4fdf77b940.nv26.08 / 13.4

Sheets: prior-samples.png, vt-normal-samples.png (the parent's own prior on the same seeds), slerp-interpolations.png and lerp-interpolations.png (A, 9 frames, B), watch-it-draw.png (z_hat_0 at 8 flow times, then the sample).

State transitions

  1. running2026-09-29T15:27:09Z
  2. completed2026-09-29T15:48:01Z

Run record

Verbatim from runs/lfp-v1x-exploratory-81eb2e2-5ae4d777-47646604/run.yaml, the record committed before launch.

schema_version
2
run_id
lfp-v1x-exploratory-81eb2e2-5ae4d777-47646604
state
completed
planned_at
2026-09-29T15:27:09Z
hypothesis
A rectified-flow prior fitted to the aggregate posterior of the frozen canvas latent (vt-v1, c = 8) replaces its N(0, I) prior, whose draws decode to fragments, and so generates whole, novel OpenMoji-like drawings and interpolates through plausible drawings. The prior is a 9.7M-parameter DiT over standardised (8, 18, 18) grids (x_t = (1 - t) x_0 + t eps, velocity target, logit-normal t, EMA 0.999), trained on posterior draws of 32 exact variants per training icon (no compositions) and selected by EMA flow loss on the validation icons' latents. Samples: 50 Euler steps from noise scale 1 to z_hat_0, greedy IEEE float32 decoding by the VT. Scored afterwards by latent_metrics through the gallery backend canvas-flow on its fixed sets (339 samples, seed 23; 32 validation pairs x 9 frames, rng 17; CLIP ViT-B/32 at 72 px against the 339 validation renders; bootstrap intervals, paired where paired; copy rule without-twins): B1-B3 on the report with default settings (its samples block); B4 and B5 on the report run with --interpolation backend, whose interpolation block is canvas-flow's own path, the slerp below (the report records interpolation.path slerp-through-prior-noise). The default report's interpolation block is the frozen VT's lerp of posterior means and scores neither B4 nor B5. Pass requires B1-B3; B4 and B5 are the interpolation claims. (B1) fragment rate at most 10%. (B2) CLIP k-NN precision at least 2x the best latent-v2 prior row (N(0, I) or refit), and recall above it, both with intervals excluding zero. (B3) copy rate at most 10% and all 339 samples distinct. (B4) slerp through the prior's noise space (reverse-ODE inversion of both posterior means, slerp, forward integration): interior fragment rate at most 15%, jump share below latent-v2's (paired over pairs, interval excluding zero), detour rate at most 20%. (B5) slerp interior precision above v9 on pixel crossfades of the same pairs, paired, interval excluding zero; a failure is recorded as "the latent adds nothing over crossfade plus transcriber". Decision: B1-B3 pass - the gallery's main latent model. B3 fails - add compositions to the prior's data, then an earlier-stopped prior. B1 or B2 fails while B3 passes - a larger prior. Only B5 fails - keep the model; its interpolation is a dissolve. Reported without a criterion: validation flow loss against the Gaussian velocity, vt-v1's N(0, I) samples on the same seeds, lerp strips, watch-it-draw strips, ms per sample.
parent_run
vt-v1-74dfbd2-46380052-47646604
git_commit
81eb2e24f5d7557a63284b3e011d13260b643cc4
dirty_patch_sha256
config_sha256
5ae4d777ae3f07b9…
uncommitted_sources
config
configs/latent/lfp-v1x-exploratory.yaml
config_resolved
parent_gate
A3
latents
variants
32
originals
true
mirror
0.5
max_shift
48
colour
0.1
seed
9101
chunk_icons
8
train_icons
None
validation_icons
339
encode_batch
64
workers
16
network
channels
8
grid
18
patch
2
d_model
256
layers
8
heads
8
mlp_ratio
4.0
frequency_dim
256
training
steps
40000
batch_size
256
learning_rate
0.0003
weight_decay
0.0
warmup_steps
1000
schedule
constant
grad_clip
1.0
ema_decay
0.999
time_mean
0.0
time_std
1.0
posterior_sampling
true
seed
7101
eval_every
1000
validation_draws
16
trace_every
100
bf16
true
sampling
steps
50
trajectory_frames
8
eval_icons
339
prior_samples
64
interpolation_pairs
8
interpolation_frames
9
trajectories
4
parent
run
vt-v1-74dfbd2-46380052-47646604
checkpoint
/home/dev/.cache/mojidiff/runs/vt-v1-74dfbd2-46380052-47646604/best.pt
checkpoint_sha256
b43c20240ac47c38…
latent_shape
8, 18, 18
gate
record
/home/dev/workspace/mojidiff/runs/vt-v1-74dfbd2-46380052-47646604/run.yaml
state
completed
gate
A3
true
frozen
true
latents
cache
/home/dev/.cache/mojidiff/latent-flow/vt-v1-74dfbd2-46380052-47646604-7b18709c9403.pt
cache_sha256
2e9dfcbeb9a7882f…
created_by_this_run
true
settings
format
1
vt_checkpoint_sha256
b43c20240ac47c38…
dataset_sha256
476466042da98d72…
latent_shape
8, 18, 18
image_size
144
precision
IEEE float32 encoding (TF32 off), stored float16
variants
32
originals
true
mirror
0.5
max_shift
48
colour
0.1
seed
9101
chunk_icons
8
train_icons
None
validation_icons
339
encode_batch
64
train_latents
85792
train_icons
2681
dropped_variants
0
validation_latents
339
precompute_seconds
115.95677558891475
channel_mean
-0.035881295800209045, 0.01895667426288128, -0.019481094554066658, -0.23405441641807556, 0.06478327512741089, 0.009515517391264439 … and 2 more
channel_std
1.0073826313018799, 0.9975386261940002, 0.9573033452033997, 0.9334527254104614, 0.9612548351287842, 0.9861312508583069 … and 2 more
note
test icons are never encoded; validation latents select the checkpoint
dataset
pilot_config
configs/learning/openmoji-g1-geometric-gate-v16.yaml
cache_sha256
476466042da98d72…
train_icons
2681
evaluated_on
primary/validation
seed
7101
determinism
seeded (torch generator on the device for batches, t and eps; fixed CPU draws for validation); cuDNN and SDPA kernels not forced deterministic
model_parameters
18779986
parameters
flow
9747744
parent_vt_frozen
9032242
total
18779986
note
model_parameters is the served sampler: the flow and the frozen VT that decodes its samples, as the gallery's canvas-flow backend counts it; only the flow trains
command
.venv/bin/python -m mojidiff.learning.latent_flow --config configs/latent/lfp-v1x-exploratory.yaml
outputs
run_dir
runs/lfp-v1x-exploratory-81eb2e2-5ae4d777-47646604
checkpoints
/home/dev/.cache/mojidiff/runs/lfp-v1x-exploratory-81eb2e2-5ae4d777-47646604
notes
completed_at
2026-09-29T15:48:01Z
resource
device
NVIDIA GeForce RTX 4080
torch_version
2.14.0a0+4fdf77b940.nv26.08
cuda_version
13.4
python_version
3.12.3
peak_vram_gib
3.4414305686950684
train_seconds
1178.589293822064
inference_ms_per_icon
1048.441516526509
inference_p95_ms_per_icon
1049.014563090168
icons_per_second
0.9537966441018131
inference
50 Euler steps of the EMA flow at batch 1 plus the VT's float32 CUDA-graph greedy decode, IEEE float32 (TF32 off)
excludes
rasterising the output SVG
note
measured while other GPU work may be running
baselines
gaussian_velocity_validation_flow_loss
1.7629378424999667
vt_standard_normal_prior_fragment_rate
0.328125
validation_icons_fragment_rate
0.02064896755162242
note
zero-parameter controls on the same draws: the exact velocity if the standardised latents were N(0, I); the parent VT decoding the same N(0, I) seeds as latents (its own prior); the validation icons
result
precision
IEEE float32 greedy decoding (TF32 off: matmul and cuDNN); pixel metrics at 72 px
validation_flow_loss
ema_served
1.2587471028056298
gaussian_velocity_baseline
1.7629378424999667
icons
339
draws_per_icon
16
gaussian_velocity_baseline_at_selection_time
1.7629378424999667
prior_samples
count
64
distinct
64
rendered
64
ink_coverage_median
0.2279128086419753
fragment_rate_ink_below_0.10
0.140625
near_blank_rate_ink_below_0.02
0.0
noise_scale
1.0
seed
23
euler_steps
50
vt_standard_normal
count
64
distinct
64
rendered
64
ink_coverage_median
0.11882716049382716
fragment_rate_ink_below_0.10
0.328125
near_blank_rate_ink_below_0.02
0.0
validation_fragment_rate_ink_below_0.10
0.02064896755162242
interpolation
pairs
286, 250
54, 155
124, 257
154, 72
309, 145
206, 30
… and 2 more
frames
9
inversion_steps
50
slerp
interior_fragment_rate_ink_below_0.10
0.08928571428571429
endpoint_pixel_error
0.09233731415588409, 0.06777590333949775, 0.1187708532641409
distinct_programs_per_strip
9.0
lerp
interior_fragment_rate_ink_below_0.10
0.07142857142857142
endpoint_pixel_error
0.08751612238120288, 0.06271789727034047, 0.11594126901181882
distinct_programs_per_strip
9.0
watch_it_draw
trajectories
4
seed
31
times
1.0, 0.86, 0.72, 0.58, 0.46, 0.32 … and 2 more
step_indices
0, 7, 14, 21, 27, 34 … and 2 more
columns
z_hat_0 at each time, then the final sample
latency
end_to_end_ms_batch_1
1048.441516526509
end_to_end_p95_ms_batch_1
1049.014563090168
flow_only_ms_batch_1
60.31673902180046
batched_ms_per_sample
143.0008201568853
batched_note
64 samples: flow then greedy_decode in batches of 64
criteria
B1
claim
fragment rate at most 10%
status
scored after the run by latent_metrics through the gallery backend canvas-flow, copy rule without-twins, default settings: its samples block
B2
claim
CLIP precision at least 2x the best latent-v2 prior row, and recall above it, both with intervals excluding zero
status
scored after the run by latent_metrics through the gallery backend canvas-flow, copy rule without-twins, default settings: its samples block
B3
claim
copy rate at most 10% under copy rule without-twins, and all samples distinct
status
scored after the run by latent_metrics through the gallery backend canvas-flow, copy rule without-twins, default settings: its samples block
B4
claim
slerp interior fragments at most 15%, jump share below latent-v2's (paired), detours at most 20%
status
scored after the run by latent_metrics through the gallery backend canvas-flow, copy rule without-twins, with --interpolation backend: its interpolation block, the backend's slerp through the prior's noise (the report records interpolation.path slerp-through-prior-noise). The default report's interpolation block is the frozen VT's lerp of posterior means and scores neither B4 nor B5
B5
claim
slerp interior CLIP precision above v9 on crossfades, paired
status
scored after the run by latent_metrics through the gallery backend canvas-flow, copy rule without-twins, with --interpolation backend: its interpolation block, the backend's slerp through the prior's noise (the report records interpolation.path slerp-through-prior-noise). The default report's interpolation block is the frozen VT's lerp of posterior means and scores neither B4 nor B5
float32_flags
matmul_allow_tf32
false
cudnn_allow_tf32
false
float32_matmul_precision
highest
tf32_cublas_override_variable
1
selection
rule
lowest EMA validation flow loss (fixed t and eps draws, 16 per validation latent) over evaluations every 1000 steps
step
38000
validation_flow_loss_ema
1.2587482705409831
selected_step
38000