The latest measured window is not available yet. Protocol stages remain an illustrative schematic.
A decentralized market hunts the prompts a model can still learn from. GRAIL verifies the work; selected groups become eligible for the next training step.
A model only learns from prompts at its frontier — and only 1–15% of prompts are ever there. Train on those and every step counts; miss them and you burn compute on prompts that teach nothing. A decentralized network mines exactly that slice, proves ranked candidates top-down, and makes only credited groups eligible for Forge — focusing scarce training budget on frontier signal instead of raw volume.
M=8 · binary reward geometry
active group · k=2
Reinforce two rare correct trajectories
2 / 6
correct / wrong
where training signal lives
The steady gate keeps k=2…6 where σ ≥ 0.43.
correct rollouts k · eight attempts per group
| k | correct / wrong | p | sigma | 1 − p | discovery pullσ(1−p) | A correct | A wrong | training role |
|---|---|---|---|---|---|---|---|---|
| 0 / 8 | 0.000 | 0.000 | 1.000 | 0.000 | — | — | No correct path Unanimous failure, no contrast | |
| 1 / 7 | 0.125 | 0.331 | 0.875 | 0.289 | +2.646 | −0.378 | Possible discovery Strong positive, but one success may be lucky | |
| 2 / 6 | 0.250 | 0.433 | 0.750 | 0.325 | +1.732 | −0.577 | Discovery peak Reinforce two rare correct trajectories | |
| 3 / 5 | 0.375 | 0.484 | 0.625 | 0.303 | +1.291 | −0.775 | Strong exploration Useful correct paths remain uncommon | |
| 4 / 4 | 0.500 | 0.500 | 0.500 | 0.250 | +1.000 | −1.000 | Balanced frontier Equal positive and negative concentration | |
| 5 / 3 | 0.625 | 0.484 | 0.375 | 0.182 | +0.775 | −1.291 | Early refinement Correct behavior is already common | |
| 6 / 2 | 0.750 | 0.433 | 0.250 | 0.108 | +0.577 | −1.732 | Refinement mirror Suppress two rare failure trajectories | |
| 7 / 1 | 0.875 | 0.331 | 0.125 | 0.041 | +0.378 | −2.646 | Nearly mastered Mostly cleanup of one rare failure | |
| 8 / 0 | 1.000 | 0.000 | 0.000 | 0.000 | — | — | Already mastered Unanimous success, no contrast |
p = k/8 · σ = √p(1−p) · discovery pull = σ(1−p) · A = (r−p)/σ · displayed to 3 decimals
The pull is an asymmetric teaching-signal lens, not protocol variance. Advantages are the idealized ε-free values; an em dash means σ=0 and no within-group contrast.
pass@1 · held-out math · 300 steps
+14pp
reactive filter → predictive market
As the policy improves, frontier prompts become rarer. A reactive filter should spend more compute finding them; a market makes locating them the product. The magnitude of that efficiency gap still needs a direct, repeated compute-normalized experiment.
+14pp
pass@1 vs vanilla GRPO
identical 300-step budget
1 run
controlled comparison
replication remains open
1–15%
of prompts are in-zone
the rest teach nothing
We're building the decentralized RL training layer for Bittensor.
read the measurement →Awaiting the first sealed window. Once a window seals, its submitted bundles flow through the stages below and the sealing validator recomputes the proof. Selected groups become eligible for GRPO only when the batch clears quarantine.
01
window
event-sealed · each tempo
02
select
miner bets a prompt
03
rollout
M=8 sampled completions
04
GRAIL sketch
binds rollout to weights
05
learning zone
group-σ in the band
06
training gate
healthy full batches → GRPO
the gate every rollout passes — recompute, or earn zero.
Windows seal once enough distinct in-zone prompts arrive. The moment this validator publishes its archive, the submitted / selected split appears right here.
anatomy of a sealed window
R2 exposes this lean record publicly. Raw completion text and token IDs are reserved for authorized auditors.
One protocol, two instruments. The archive is the cool one — its public record exposes the selected inputs, rollout statistics, reward vectors, Merkle roots, checkpoint claims, and the validator that published them. Raw completion text and token IDs stay behind an auditor credential; the provenance does not.
Follow the record, field by field.
The hot instrument. Healthy balanced batches can become GRPO steps. The validator publishes every four successful balanced optimizer steps, with earlier publication on behavior-policy drift; quarantined, blocked, or partial batches remain recorded without pretending the optimizer moved. Every published checkpoint is benchmarked against the one before it.
Watch the furnace tick.
The architecture isn't tied to one model or domain — it's a general-purpose RL inference layer. A client brings a model and a set of environments; the miner network handles rollout generation, prompt selection, and verification. RL inference on demand.
who · 01
Bring a model and your environments. Skip building and babysitting rollout infrastructure — the network returns an optimized training signal, already verified.
who · 02
Any Bittensor subnet that wants to RL-tune a model can route its rollout generation here instead of standing up a trainer of its own.
The RL inference layer on Bittensor.
read the vision →Run a miner. Predict which prompts sit at the model’s edge. Land in the learning zone and your verified rollouts train the next checkpoint.
◉ Reliquary · subnet 81 · the learning frontier, as a market
SN81 · live ownership horizon
finalized signalThe number that can move.
Synchronizing finalized state
next finalized epoch check · synchronizing
Day-level projection, conditional on unchanged locks, α supply, runtime rates, and leader identity. The horizon never confirms ownership; finalized chain state does.
Ownership campaign state: checking finalized dataA model only learns from prompts at its frontier — and only 1–15% of prompts are ever there. Train on those and every step counts; miss them and you burn compute on prompts that teach nothing. A decentralized network mines exactly that slice, proves ranked candidates top-down, and makes only credited groups eligible for Forge — focusing scarce training budget on frontier signal instead of raw volume.
M=8 · binary reward geometry
active group · k=2
Reinforce two rare correct trajectories
2 / 6
correct / wrong
where training signal lives
The steady gate keeps k=2…6 where σ ≥ 0.43.
correct rollouts k · eight attempts per group
| k | correct / wrong | p | sigma | 1 − p | discovery pullσ(1−p) | A correct | A wrong | training role |
|---|---|---|---|---|---|---|---|---|
| 0 / 8 | 0.000 | 0.000 | 1.000 | 0.000 | — | — | No correct path Unanimous failure, no contrast | |
| 1 / 7 | 0.125 | 0.331 | 0.875 | 0.289 | +2.646 | −0.378 | Possible discovery Strong positive, but one success may be lucky | |
| 2 / 6 | 0.250 | 0.433 | 0.750 | 0.325 | +1.732 | −0.577 | Discovery peak Reinforce two rare correct trajectories | |
| 3 / 5 | 0.375 | 0.484 | 0.625 | 0.303 | +1.291 | −0.775 | Strong exploration Useful correct paths remain uncommon | |
| 4 / 4 | 0.500 | 0.500 | 0.500 | 0.250 | +1.000 | −1.000 | Balanced frontier Equal positive and negative concentration | |
| 5 / 3 | 0.625 | 0.484 | 0.375 | 0.182 | +0.775 | −1.291 | Early refinement Correct behavior is already common | |
| 6 / 2 | 0.750 | 0.433 | 0.250 | 0.108 | +0.577 | −1.732 | Refinement mirror Suppress two rare failure trajectories | |
| 7 / 1 | 0.875 | 0.331 | 0.125 | 0.041 | +0.378 | −2.646 | Nearly mastered Mostly cleanup of one rare failure | |
| 8 / 0 | 1.000 | 0.000 | 0.000 | 0.000 | — | — | Already mastered Unanimous success, no contrast |
p = k/8 · σ = √p(1−p) · discovery pull = σ(1−p) · A = (r−p)/σ · displayed to 3 decimals
The pull is an asymmetric teaching-signal lens, not protocol variance. Advantages are the idealized ε-free values; an em dash means σ=0 and no within-group contrast.
pass@1 · held-out math · 300 steps
+14pp
reactive filter → predictive market
As the policy improves, frontier prompts become rarer. A reactive filter should spend more compute finding them; a market makes locating them the product. The magnitude of that efficiency gap still needs a direct, repeated compute-normalized experiment.
+14pp
pass@1 vs vanilla GRPO
identical 300-step budget
1 run
controlled comparison
replication remains open
1–15%
of prompts are in-zone
the rest teach nothing
We're building the decentralized RL training layer for Bittensor.
read the measurement →Measured window 28,467 measured 78 submitted bundles; its archive records 1 publisher hotkey. Selected groups become training-eligible; the quarantine gate decides whether the model steps.
01
window
event-sealed · each tempo
02
select
miner bets a prompt
03
rollout
M=8 sampled completions
04
GRAIL sketch
binds rollout to weights
05
learning zone
group-σ in the band
06
training gate
healthy full batches → GRPO
the gate every rollout passes — recompute, or earn zero.
Sealed 78 submitted bundles and selected the groups eligible for the training gate. The remaining candidates did not enter this training batch.
submitted
78
selected
16
not selected
62
disagreement
n/a
anatomy of a sealed window
R2 exposes this lean record publicly. Raw completion text and token IDs are reserved for authorized auditors.
One protocol, two instruments. The archive is the cool one — its public record exposes the selected inputs, rollout statistics, reward vectors, Merkle roots, checkpoint claims, and the validator that published them. Raw completion text and token IDs stay behind an auditor credential; the provenance does not.
Follow the record, field by field.
The hot instrument. Healthy balanced batches can become GRPO steps. The validator publishes every four successful balanced optimizer steps, with earlier publication on behavior-policy drift; quarantined, blocked, or partial batches remain recorded without pretending the optimizer moved. Every published checkpoint is benchmarked against the one before it.
Watch the furnace tick.
The architecture isn't tied to one model or domain — it's a general-purpose RL inference layer. A client brings a model and a set of environments; the miner network handles rollout generation, prompt selection, and verification. RL inference on demand.
who · 01
Bring a model and your environments. Skip building and babysitting rollout infrastructure — the network returns an optimized training signal, already verified.
who · 02
Any Bittensor subnet that wants to RL-tune a model can route its rollout generation here instead of standing up a trainer of its own.
The RL inference layer on Bittensor.
read the vision →Run a miner. Predict which prompts sit at the model’s edge. Land in the learning zone and your verified rollouts train the next checkpoint.
◉ Reliquary · subnet 81 · the learning frontier, as a market
Latest measured window 28,467: 78 submitted bundles, 16 selected, 20.5% selection rate.
A decentralized market hunts the prompts a model can still learn from. GRAIL verifies the work; selected groups become eligible for the next training step.