skip to content

empirical

Cross-GPU determinism

Four hosts. Two GPU families. Two CUDA toolchains. Two PyTorch wheels. 90 samples each. The same rollout re-executed everywhere lands on one bit-identical digest — so the validator can re-run a miner's decode and reject any nondeterminism.

count

hosts

4

count

samples / host

90

result

digest agreement

all match

bit-identical across classes

FIG.01 · cross-host digest matrix · same rollout, every GPU
b0b8b16b24b32b40b48b56blackwellblackwellblackwellhopperdev-server · Blackwell · bytes 0–7: 4de19184 = canonical✓dev-server · Blackwell · bytes 8–15: 31de9f26 = canonical✓dev-server · Blackwell · bytes 16–23: 8efacfcb = canonical✓dev-server · Blackwell · bytes 24–31: 298e52cb = canonical✓dev-server · Blackwell · bytes 32–39: e80a4adb = canonical✓dev-server · Blackwell · bytes 40–47: 6388af3b = canonical✓dev-server · Blackwell · bytes 48–55: 183753e7 = canonical✓dev-server · Blackwell · bytes 56–63: e960572c = canonical✓staging1 · Blackwell · bytes 0–7: 4de19184 = canonical✓staging1 · Blackwell · bytes 8–15: 31de9f26 = canonical✓staging1 · Blackwell · bytes 16–23: 8efacfcb = canonical✓staging1 · Blackwell · bytes 24–31: 298e52cb = canonical✓staging1 · Blackwell · bytes 32–39: e80a4adb = canonical✓staging1 · Blackwell · bytes 40–47: 6388af3b = canonical✓staging1 · Blackwell · bytes 48–55: 183753e7 = canonical✓staging1 · Blackwell · bytes 56–63: e960572c = canonical✓staging2 · Blackwell · bytes 0–7: 4de19184 = canonical✓staging2 · Blackwell · bytes 8–15: 31de9f26 = canonical✓staging2 · Blackwell · bytes 16–23: 8efacfcb = canonical✓staging2 · Blackwell · bytes 24–31: 298e52cb = canonical✓staging2 · Blackwell · bytes 32–39: e80a4adb = canonical✓staging2 · Blackwell · bytes 40–47: 6388af3b = canonical✓staging2 · Blackwell · bytes 48–55: 183753e7 = canonical✓staging2 · Blackwell · bytes 56–63: e960572c = canonical✓staging3 · Hopper · bytes 0–7: 4de19184 = canonical✓staging3 · Hopper · bytes 8–15: 31de9f26 = canonical✓staging3 · Hopper · bytes 16–23: 8efacfcb = canonical✓staging3 · Hopper · bytes 24–31: 298e52cb = canonical✓staging3 · Hopper · bytes 32–39: e80a4adb = canonical✓staging3 · Hopper · bytes 40–47: 6388af3b = canonical✓staging3 · Hopper · bytes 48–55: 183753e7 = canonical✓staging3 · Hopper · bytes 56–63: e960572c = canonical✓

canonical sha2564de1918431de…e960572c· 32 segments compared · 0 divergent

FIG.02 · re-execution gate · how determinism is enforced
MINER DECODErollout + tokensGRAIL SKETCHcommit + openVALIDATOR RE-EXECreplays decode · msDIGEST MATCH?sha256 == canonicalACCEPTbit-identicalany drift → the re-executed digest ≠ the miner's → REJECT · honest decode replays byte-for-byte
FIG.03per-host evidence · hardware + toolchain + digest

every row ran the same 90-sample workload · a single-byte drift would surface as a red segment in FIG.01 · none observed

host
GPU
arch
cuda
torch
digest vs canonical
dev-server · Blackwell
RTX PRO 6000
sm_120
13.0
2.11.0+cu130
4de19184…60572c
staging1 · Blackwell
RTX PRO 6000
sm_120
13.0
2.11.0+cu130
4de19184…60572c
staging2 · Blackwell
RTX PRO 6000
sm_120
13.0
2.11.0+cu130
4de19184…60572c
staging3 · Hopper
H100
sm_90
12.4
2.5.1+cu124
4de19184…60572c

why it matters

If two honest miners on different GPUs produced different proofs for the same completion, mesh consensus would fail — not because anyone cheated, but because of numerical drift. Reliquary's proof envelope is sized to absorb long-sequence attention drift, while remaining tight enough that tampering is caught. The matrix above shows the envelope has substantial headroom: the honest path across four distinct hardware + toolchain combinations produces byte-identical digests, so the validator can re-execute any miner's decode and gate nondeterminism the instant a digest diverges.

Full reproduction steps + per-sample digests are committed in reliquary-ledger/docs/audit/cross_gpu.