Something I really like when I study a subject is understanding its history, how it reached the point where it is today
I did that exercise for RL in post-training: from RLHF and PPO, to verifiable rewards, to the GRPO family of variants, to agents acting in environments. Everything is backed by what the labs themselves say in their public reports (DeepSeek, Qwen, Kimi, GLM-5, Nemotron, Mistral and more), in their own words
This is the companion piece to Class 3 of our Training Agents series with @burtenshaw. The class explains how GRPO works, with three hands-on experiments. The article shows where the same ideas appear at frontier scale
we just released a new blog "Training a coding agent using the OpenCode harness in remote HF sandboxes with TRL and OpenEnv"
you can take a real coding agent (OpenCode), let it run its own tool loop against real coding problems, and train it with RL on the exact tokens it produced
and every rollout runs in its own remote HF sandbox, so rollouts scale out beyond one machine
the loop: - OpenCode owns its tool loop inside an OpenEnv sandbox - an in-sandbox proxy records the real token ids + logprobs, per turn - a hidden-test verifier scores the result, and that is the reward - TRL trains with AsyncGRPO, weights sync back to vLLM over NCCL
and the @liquidai blog comes with some nice details about the training procedure, so let's analyze it.
basically, a full agent training pipeline but compressed into 2.6B
base model β SFT β specialized teachers per domain (SFT + RLVR) β on-policy distillation back into one student β agentic RL
the two most interesting stages
β MOPD: the student generates, each prompt routes to its domain teacher for token-level feedback. teachers branch from the same SFT checkpoint, so their signal stays close to the student's distribution
β agentic RL: multi-turn GRPO inside real harnesses (OpenClaw, Hermes Agent), one sandbox per rollout, a proxy captures token-level trajectories while the harness stays a black box
this makes a 2.6B that beats much larger models on instruction following and tool use
SFT, distillation, RL, RL envs: exactly what we're covering in our Training Agents livestream series (next one coming soon!)
tomorrow (Tuesday, July 28), we're back with Class 3 of the Training Agents live series
π§ what: reinforcement learning for training agents (GRPO): how it works, how to implement it in TRL, and end-to-end examples ποΈ when: Tuesday, July 28 - π 5:00 PM CEST / 8:30 PM IST π where: Live on @huggingface's X, YouTube, and LinkedIn
you can now train your own coding agents with trl + openenv, starting with opencode
we just added end-to-end support for training agent harnesses:
> TRL: a loop-owning training path (AsyncGRPOTrainer + HarnessRolloutWorker) that launches the agent in an OpenEnv session, reads back its trace, reconstructs the training samples, and trains with AsyncGRPO > OpenEnv: the OpenCode harness environment plus a transparent proxy that forwards the agent's model calls and records each turn's token ids and logprobs
you train the actual opencode agent as is, it runs its own loop and tools and the policy learns from the exact tokens it produced
we're shipping a self-contained example: local subprocess sandbox, DeepCoder problems, validated on Qwen3-8B.