MojiDiff

← experiments

openmoji-g1-gpu-215bcb8-0bafd5c-9b9b1699

failed —

Hypothesis

The exact locally verified Gate G pipeline completes one CUDA train step on the owned RTX 4080, restores a canonical checkpoint in durable storage, and preserves locked paths exactly.

Why run it

Validates the real data, selected normalizers, conditioned model, optimizer, checkpoint, and locked-edit path in the target GPU environment before larger work. The single corrected factor is the launcher environment: the owned Docker smoke now sets CUBLAS_WORKSPACE_CONFIG so the pipeline's declared deterministic algorithms are achievable on CUDA rather than rejected.

Measured behaviour

Training token accuracy (per batch)
accuracyoptimizer step00.0020.0040.00611.21.41.61.82
Table view
optimizer steptrain accuracy
10.0051

Measured result

Verbatim from the run's summary.json.

bucket
bucket-p32-t128
bucket_icons
3359
checkpoint_bytes
7075309
checkpoint_round_trip
true
checkpoint_sha256
d11efa006657dafc…
config_sha256
0bafd5c3b3de29da…
corruption
factorized_role_uniform_geometry
cuda_version
None
deterministic_algorithms
true
device
cpu
final_train
loss
15.216715812683105
step
1
train_token_accuracy
0.005067567567567568
group_vocabulary_size
12
locked_path_exact
true
metrics_sha256
71ee91a4c99b9f53…
model_parameters
577552
schema_version
1
scope
dominant-bucket fixed-topology geometry pilot; not unconditional generation
selected_train_rows
color/svg/1F3F4-E0064-E0065-E0062-E0065-E007F.svg, color/svg/1F468-1F3FC-200D-1F9BC.svg, color/svg/1F561.svg, color/svg/1F469-1F3FE-200D-1F9BC.svg
selected_validation_rows
color/svg/1F994.svg, color/svg/E30A.svg
steps
1
study_version
openmoji-g1-dominant-bucket-smoke
subgroup_vocabulary_size
118
torch_version
2.8.0+cpu
validation
accuracy
0.009404388714733543
changed_accuracy
0.008849557522123894
changed_total
226
loss
15.344958305358887
retained_accuracy
0.009708737864077669
retained_total
412

Written result

Status: failed on the adapter result contract after the GPU step completed.

This run is bounded to the exact one-step local pilot and stages only its six selected raw SVGs plus the pinned palette with the committed source snapshot.

It retries openmoji-g1-gpu-e7dc920-0bafd5c-9b9b1699, which loaded the staged config, data, normalizers, model, and optimizer correctly but was rejected inside the first CUDA step because torch.use_deterministic_algorithms(True) requires CUBLAS_WORKSPACE_CONFIG on CUDA >= 10.2. The only changed factor is that launcher environment variable. The config, fixture, seed, step budget, image, and cap are unchanged.

The deterministic-cuBLAS correction worked. The pipeline ran to completion on the RTX 4080: device is cuda, deterministic_algorithms is true, one optimizer step produced train loss 15.216644 and token accuracy 0.005068, validation loss is 15.344566 with 0.009404 aggregate / 0.008850 changed / 0.009709 retained accuracy over 226 changed and 412 retained fields, checkpoint_round_trip is true, and locked_path_exact is true. Near-random first-step accuracy is expected and is not learning evidence.

The run is still recorded as failed because the adapter's result contract failed: the staged wrapper computed its summary digest from the output root rather than the pilot's report root, so it raised before printing the JSON result. The GPU work and its artifacts are intact and durable on the worker and are listed in artifacts.json; nothing was deleted. The correction is confined to the wrapper's summary path and requires a new code and run identity.

Recovered evidence on the worker:

The local CPU pilot checkpoint is the same 7,075,309 bytes but hashes differently (d11efa00...), which is the expected CPU/GPU floating-point difference. This run was never intended to establish cross-device artifact identity.

State transitions

  1. planned2026-09-20T11:05:30Z
  2. staged2026-09-20T11:06:50Z
  3. failed2026-09-20T11:08:34Zwrapper_read_summary_from_output_root_instead_of_report_root

Run record

Verbatim from runs/openmoji-g1-gpu-215bcb8-0bafd5c-9b9b1699/run.yaml, the record committed before launch.

schema_version
1
run_id
openmoji-g1-gpu-215bcb8-0bafd5c-9b9b1699
state
failed
parent_run
openmoji-g1-smoke-da44363-0bafd5c-9b9b1699
retry_of
openmoji-g1-gpu-e7dc920-0bafd5c-9b9b1699
hypothesis
The exact locally verified Gate G pipeline completes one CUDA train step on the owned RTX 4080, restores a canonical checkpoint in durable storage, and preserves locked paths exactly.
expected_information_gain
Validates the real data, selected normalizers, conditioned model, optimizer, checkpoint, and locked-edit path in the target GPU environment before larger work. The single corrected factor is the launcher environment: the owned Docker smoke now sets CUBLAS_WORKSPACE_CONFIG so the pipeline's declared deterministic algorithms are achievable on CUDA rather than rejected.
code
git_commit
215bcb8c8d67eefe4651c4807f45fab91c2971c2
snapshot_mode
git-archive-plus-hash-pinned-six-svg-smoke-fixture
archive_sha256
fa27c3b501e2abe7…
tree_sha256
706315fd65867443…
change_from_retry_parent
owned Docker smoke launcher sets CUBLAS_WORKSPACE_CONFIG=:4096:8
config
path
configs/learning/openmoji-g1-dominant-bucket-smoke.yaml
sha256
0bafd5c3b3de29da…
dataset
hybrid_sha256
9b9b1699677a6f97…
assignments_sha256
e0cdb2a3cc8f00df…
staged_raw_svg_count
6
bucket
bucket-p32-t128
worker
alias
owned-gpu
gpu
NVIDIA GeForce RTX 4080
driver
595.71.05
image
mojidiff/owned-gpu-smoke:cd3250e
image_digest
sha256:f42d0acc2a768b6929a982dbb1afcfb342b382751ea0e974ba00052e5dff38ca
torch
2.8.0a0+5228986c39.nv25.06
cuda
12.9
execution
smoke_id
openmoji-g1-pipeline-v1
seed_set
3101
deterministic_algorithms
true
cublas_workspace_config
:4096:8
optimizer_steps
1
resource_cap
max_steps
20000
max_storage_gb
50
network
none
checkpoint_policy
retain all bounded smoke outputs in persistent artifact volume
outputs
local_metadata
runs/openmoji-g1-gpu-215bcb8-0bafd5c-9b9b1699
remote_workspace
/home/dev/workspace/openmoji-g1-gpu-215bcb8-0bafd5c-9b9b1699
durable_artifacts
/home/dev/.cache/openmoji-g1-gpu-215bcb8-0bafd5c-9b9b1699
planned_at
2026-09-20 11:05:30+00:00
staged_at
2026-09-20 11:06:50+00:00
failed_at
2026-09-20 11:08:34+00:00
failure
reason
wrapper_read_summary_from_output_root_instead_of_report_root
gpu_step_started
true
gpu_step_completed
true
artifacts_written
true
artifacts_durable
true