skip to content
syncing conviction
conviction watch · syncing finalized state
syncing
→
Conviction state: syncing conviction
Reliquary
dashboard
conviction
research
roadmap
docs
source
x
menu
Forge · live training · team only · Reliquary
forge · live training
5H47sFL6-recovery-checkpoint-182
·
step 28,467
finished
▸ open in wandb
status
finished
runtime
4d 03h
last seen
—
loss
7.03e-5
kl
2.98e-3
grad_norm
0.6211
reward μ
0.5156
steps / h
286
target met
lr
5.00e-7
gpu util
100.0%
gpu mem
100.0%
ai advisor
reading the last 160 points…
model quality
computing quality signals…
validator rejections
tailing validator logs over ssh…
PPO loss
primary objective
KL divergence
budget kl_beta = 0.01
grad_norm
clip @ 1
learning rate
cosine schedule
rewards
mean ± std
degenerate-group ratio
zero-variance reward groups
valid rollout ratio
GRAIL accepted / submitted
model improvement · checkpoint evals
held-out pass@1 · math + code
no data
No checkpoint evals yet
Held-out benchmarks appear here as soon as the eval pipeline publishes its first checkpoint for the current base. Expected shortly after a base reset.
accepted rollouts · rolling window
verified into training · last 3 windows
accepted rollouts
verified miner rollouts stream once the validator seals a window
gpu util
100.0%
gpu mem
100.0%
sm occupancy
55.9%
gpu temp
79°C
power
349 W
run config
22 keys · hide
b_batch
8
grad_clip_norm
1
grad_norm_skip_threshold
50
kl_base_model
Qwen/Qwen3.5-4B@851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a
kl_beta
0.01
kl_reference_mode
fixed
kl_reference_repo_id
Qwen/Qwen3.5-4B
kl_reference_revision
851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a
kl_reference_storage_bytes
9078531392
learning_rate
0.000001
lr_cosine_max_windows
10000
lr_warmup_windows
10
m_rollouts_per_prompt
8
pi_old_source
verify_model
ppo_clip_epsilon
0.2
ppo_ratio_outside_clip_skip_threshold
0.1
reliquary_version
0.1.0
shape_len_frac
0.5
shape_penalty
0
train_until_checkpoint_n
0
wandb_training_version
recovery-checkpoint-182
window_length
5