Trained Verifier Models
Surrogate code verifiers across three model sizes trained using multiple different algorithms as described in the Aletheia paper
Text Generation • 2B • Updated • 263Note Our flagship verifiers, trained using a GRPO-style algorithm with a 16k generation limit
Aletheia-Bench/GRPO-Think-7B-16k
Text Generation • 8B • Updated • 283Note Our flagship verifiers, trained using a GRPO-style algorithm with a 16k generation limit
Aletheia-Bench/GRPO-Think-14B-16k
Text Generation • 15B • Updated • 268Note Our flagship verifiers, trained using a GRPO-style algorithm with a 16k generation limit
Aletheia-Bench/GRPO-Think-1.5B-4k
Text Generation • 2B • Updated • 253Note A variant of the flagship verifiers trained with a 4k generation limit
Aletheia-Bench/GRPO-Think-7B-4k
Text Generation • 8B • Updated • 283Note A variant of the flagship verifiers trained with a 4k generation limit
Aletheia-Bench/GRPO-Think-14B-4k
Text Generation • 15B • Updated • 271Note A variant of the flagship verifiers trained with a 4k generation limit
Aletheia-Bench/GRPO-Think-1.5B-8k
Text Generation • 2B • Updated • 255Note A variant of our flagship verifiers, trained with a 8k generation limit
Aletheia-Bench/GRPO-Think-7B-8k
Text Generation • 8B • Updated • 274Note A variant of our flagship verifiers, trained with a 8k generation limit
Aletheia-Bench/GRPO-Think-14B-8k
Text Generation • 15B • Updated • 281 • 1Note A variant of our flagship verifiers, trained with a 8k generation limit
Aletheia-Bench/GRPO-Instruct-1.5B
Text Generation • 2B • Updated • 271Note A GRPO training ablation that does not generate thinking traces
Aletheia-Bench/GRPO-Instruct-7B
Text Generation • 8B • Updated • 272Note A GRPO training ablation that does not generate thinking traces
Aletheia-Bench/GRPO-Instruct-14B
Text Generation • 15B • Updated • 281Note A GRPO training ablation that does not generate thinking traces
Aletheia-Bench/DPO-Think-1.5B
Text Generation • 2B • Updated • 256Note A verifier trained completely offline using DPO
Aletheia-Bench/DPO-Think-7B
Text Generation • 8B • Updated • 258Note A verifier trained completely offline using DPO
Aletheia-Bench/DPO-Think-14B
Text Generation • 15B • Updated • 332 • 2Note A verifier trained completely offline using DPO
Aletheia-Bench/BatchOnline-GRPO-1.5B
Text Generation • 2B • Updated • 240Note A variant of our flagship verifiers, where the generation policy is synced every 4 gradient updates
Aletheia-Bench/BatchOnline-GRPO-7B
Text Generation • 8B • Updated • 249 • 1Note A variant of our flagship verifiers, where the generation policy is synced every 4 gradient updates
Aletheia-Bench/BatchOnline-GRPO-14B
Text Generation • 15B • Updated • 250 • 1Note A variant of our flagship verifiers, where the generation policy is synced every 4 gradient updates
Aletheia-Bench/RAFT-1.5B
Text Generation • 2B • Updated • 285Note A verifier trained using on-policy rejection sampling on only positive samples
Aletheia-Bench/RAFT-7B
Text Generation • 8B • Updated • 283Note A verifier trained using on-policy rejection sampling on only positive samples
Aletheia-Bench/RAFT-14B
Text Generation • 15B • Updated • 282Note A verifier trained using on-policy rejection sampling on only positive samples
-
Aletheia: What Makes RLVR For Code Verifiers Tick?
Paper • 2601.12186 • Published • 1