Skip to content
Octopus Research Institute
Status: ReleasedData ownership: Institute-ownedLicence: CC BY 4.0

Cross-chip byte-identical decode traces

Per-token traces quantifying how far greedy decode stays byte-identical across two different chips — the basis of mid-stream failover.

Purpose

Establish how reliably decode can hand off between chips without diverging.

Source

Captured live during a 2-chip run of the same quantized model under greedy decode.

Composition

Downloadable CSV — measured cross-backend consistency (backend, comparison, cosine, argmax) plus the CUDA↔ROCm head-to-head (Qwen2.5-0.5B, 31/32 token-identical). The full per-token trace set is larger and available on request.

Known biases

  • Greedy decode only (not sampled).

Limitations

  • A single 2-chip session.
  • The disagreement residual concentrates at sub-f16-ULP ties and is irreducible at f16.

Access conditions

The measured cross-backend consistency tables are available for download below, under CC-BY-4.0. (The full per-token trace set is larger and available on request.)

Data preview
backendcomparisoncosineargmax_matchargmax_totalsovereignty
CUDA (AWS L4) MoE enginefp32 vs int8-KV teacher-forced0.9997611616ldd = libcudart only (no cuBLAS/cuDNN/cuTLASS)
Metal (Mac M-series) q4 residentprod fused-MoE vs CPU-glue ref0.999999999999441616runtime cBLAS count = 0
Metal single-encoder decode (short ctx)short 72-tok ctx0.99999999999989byte-identicalgreedy
Metal single-encoder decode (long ctx)long 72-tok ctx0.99999999999992byte-identicalgreedy
CUDA-Silicon directchip-to-chip 30B q40.99999999both sovereign
First 5 of 5 rows.
Download · CC BY 4.0