SEED SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Paper • 2607.14777 • Published 14 days ago • 103 Jinyang23/Seed-AlfWorld-3B Text Generation • 3B • Updated 13 days ago • 572 • 2
SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Paper • 2607.14777 • Published 14 days ago • 103
Spark Spark Jinyang23/Spark-1.5B-WebShop Reinforcement Learning • 2B • Updated Jan 30 • 10 • 1 Jinyang23/Spark-1.5B-ALFWorld Reinforcement Learning • 2B • Updated Jan 30 • 8 Jinyang23/Spark-1.5B-ScienceWorld Reinforcement Learning • 2B • Updated Jan 30 • 10 Spark: Strategic Policy-Aware Exploration via Dynamic Branching for Long-Horizon Agentic Learning Paper • 2601.20209 • Published Jan 28 • 24
Spark: Strategic Policy-Aware Exploration via Dynamic Branching for Long-Horizon Agentic Learning Paper • 2601.20209 • Published Jan 28 • 24
OPID OPID Jinyang23/OPID-ALFWorld-1.7B Reinforcement Learning • 2B • Updated Jun 26 • 41 • 3 OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning Paper • 2606.26790 • Published Jun 25 • 57
OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning Paper • 2606.26790 • Published Jun 25 • 57
SEED SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Paper • 2607.14777 • Published 14 days ago • 103 Jinyang23/Seed-AlfWorld-3B Text Generation • 3B • Updated 13 days ago • 572 • 2
SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Paper • 2607.14777 • Published 14 days ago • 103
OPID OPID Jinyang23/OPID-ALFWorld-1.7B Reinforcement Learning • 2B • Updated Jun 26 • 41 • 3 OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning Paper • 2606.26790 • Published Jun 25 • 57
OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning Paper • 2606.26790 • Published Jun 25 • 57
Spark Spark Jinyang23/Spark-1.5B-WebShop Reinforcement Learning • 2B • Updated Jan 30 • 10 • 1 Jinyang23/Spark-1.5B-ALFWorld Reinforcement Learning • 2B • Updated Jan 30 • 8 Jinyang23/Spark-1.5B-ScienceWorld Reinforcement Learning • 2B • Updated Jan 30 • 10 Spark: Strategic Policy-Aware Exploration via Dynamic Branching for Long-Horizon Agentic Learning Paper • 2601.20209 • Published Jan 28 • 24
Spark: Strategic Policy-Aware Exploration via Dynamic Branching for Long-Horizon Agentic Learning Paper • 2601.20209 • Published Jan 28 • 24