MojiDiff

← experiments

openmoji-g1-gpu-e7dc920-0bafd5c-9b9b1699

failed —

Hypothesis

The exact locally verified Gate G pipeline completes one CUDA train step on the owned RTX 4080, restores a canonical checkpoint in durable storage, and preserves locked paths exactly.

Why run it

Validates the real data, selected normalizers, conditioned model, optimizer, checkpoint, and locked-edit path in the target GPU environment before larger work. Also tests the single corrected factor: the staged wrapper now resolves relative input paths against the immutable source root rather than the image working directory.

Measured behaviour

Training token accuracy (per batch)
accuracyoptimizer step00.0020.0040.00611.21.41.61.82
Table view
optimizer steptrain accuracy
10.0051

Measured result

Verbatim from the run's summary.json.

bucket
bucket-p32-t128
bucket_icons
3359
checkpoint_bytes
7075309
checkpoint_round_trip
true
checkpoint_sha256
d11efa006657dafc…
config_sha256
0bafd5c3b3de29da…
corruption
factorized_role_uniform_geometry
cuda_version
None
deterministic_algorithms
true
device
cpu
final_train
loss
15.216715812683105
step
1
train_token_accuracy
0.005067567567567568
group_vocabulary_size
12
locked_path_exact
true
metrics_sha256
71ee91a4c99b9f53…
model_parameters
577552
schema_version
1
scope
dominant-bucket fixed-topology geometry pilot; not unconditional generation
selected_train_rows
color/svg/1F3F4-E0064-E0065-E0062-E0065-E007F.svg, color/svg/1F468-1F3FC-200D-1F9BC.svg, color/svg/1F561.svg, color/svg/1F469-1F3FE-200D-1F9BC.svg
selected_validation_rows
color/svg/1F994.svg, color/svg/E30A.svg
steps
1
study_version
openmoji-g1-dominant-bucket-smoke
subgroup_vocabulary_size
118
torch_version
2.8.0+cpu
validation
accuracy
0.009404388714733543
changed_accuracy
0.008849557522123894
changed_total
226
loss
15.344958305358887
retained_accuracy
0.009708737864077669
retained_total
412

Written result

Status: failed during the first CUDA step; no checkpoint written.

This run is bounded to the exact one-step local pilot and stages only its six selected raw SVGs plus the pinned palette with the committed source snapshot.

It retries openmoji-g1-gpu-cd3250e-0bafd5c-9b9b1699, which verified its stage but failed before model construction because the staged pilot resolved relative input paths against the image working directory. The only changed factor is the wrapper's working directory; the config, fixture, seed, step budget, image, and cap are unchanged.

The working-directory correction worked: the staged config, pinned palette, six raw SVGs, selected normalizers, conditioned model, and optimizer all loaded from the immutable source root. The run then failed inside the first CUDA step because the pipeline declares torch.use_deterministic_algorithms(True) and CUDA >= 10.2 cuBLAS needs CUBLAS_WORKSPACE_CONFIG to honor that declaration. The container left an empty artifact directory, no checkpoint, and no running container.

This is a launcher environment gap, not a weakening of the determinism contract. The correction sets CUBLAS_WORKSPACE_CONFIG=:4096:8 explicitly in the owned Docker smoke launcher, which makes the declared determinism achievable rather than relaxing it.

State transitions

  1. planned2026-09-20T11:01:52Z
  2. staged2026-09-20T11:03:17Z
  3. failed2026-09-20T11:04:35Zcuda_deterministic_algorithms_require_cublas_workspace_config

Run record

Verbatim from runs/openmoji-g1-gpu-e7dc920-0bafd5c-9b9b1699/run.yaml, the record committed before launch.

schema_version
1
run_id
openmoji-g1-gpu-e7dc920-0bafd5c-9b9b1699
state
failed
parent_run
openmoji-g1-smoke-da44363-0bafd5c-9b9b1699
retry_of
openmoji-g1-gpu-cd3250e-0bafd5c-9b9b1699
hypothesis
The exact locally verified Gate G pipeline completes one CUDA train step on the owned RTX 4080, restores a canonical checkpoint in durable storage, and preserves locked paths exactly.
expected_information_gain
Validates the real data, selected normalizers, conditioned model, optimizer, checkpoint, and locked-edit path in the target GPU environment before larger work. Also tests the single corrected factor: the staged wrapper now resolves relative input paths against the immutable source root rather than the image working directory.
code
git_commit
e7dc9209d53e1e53cdfe1c5273f8f5857416b125
snapshot_mode
git-archive-plus-hash-pinned-six-svg-smoke-fixture
archive_sha256
7c8c71136016300d…
tree_sha256
f74a8a69ede13b32…
change_from_retry_parent
scripts/remote/openmoji_pilot_smoke.py chdir to staged source root
config
path
configs/learning/openmoji-g1-dominant-bucket-smoke.yaml
sha256
0bafd5c3b3de29da…
dataset
hybrid_sha256
9b9b1699677a6f97…
assignments_sha256
e0cdb2a3cc8f00df…
staged_raw_svg_count
6
bucket
bucket-p32-t128
worker
alias
owned-gpu
gpu
NVIDIA GeForce RTX 4080
driver
595.71.05
image
mojidiff/owned-gpu-smoke:cd3250e
image_digest
sha256:f42d0acc2a768b6929a982dbb1afcfb342b382751ea0e974ba00052e5dff38ca
torch
2.8.0a0+5228986c39.nv25.06
cuda
12.9
execution
smoke_id
openmoji-g1-pipeline-v1
seed_set
3101
deterministic_algorithms
true
optimizer_steps
1
resource_cap
max_steps
20000
max_storage_gb
50
network
none
checkpoint_policy
retain all bounded smoke outputs in persistent artifact volume
outputs
local_metadata
runs/openmoji-g1-gpu-e7dc920-0bafd5c-9b9b1699
remote_workspace
/home/dev/workspace/openmoji-g1-gpu-e7dc920-0bafd5c-9b9b1699
durable_artifacts
/home/dev/.cache/openmoji-g1-gpu-e7dc920-0bafd5c-9b9b1699
planned_at
2026-09-20 11:01:52+00:00
staged_at
2026-09-20 11:03:17+00:00
failed_at
2026-09-20 11:04:35+00:00
failure
reason
cuda_deterministic_algorithms_require_cublas_workspace_config
gpu_step_started
true
model_constructed
true
artifact_bytes
0