DiagEvo: Diagnosis-Guided Self-Evolution via Hierarchical Error Memory Paper • 2609.00768 • Published 12 days ago • 22
Procedural Graphs: Self-Evolving Execution Structures for LLM Agents Paper • 2609.09153 • Published 5 days ago • 37
Steering Geometry: Validating Human Value Geometry in LLM Steering Space Paper • 2609.06289 • Published 8 days ago • 29
Evaluating the Hidden Costs of Personalization in Large Language Models Paper • 2608.28833 • Published 16 days ago • 30
BeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model Inference Paper • 2609.04971 • Published 9 days ago • 36
Hi-Q: Hierarchical Evidence-guided Query Refinement for Multi-Hop Question Answering Paper • 2608.30468 • Published 13 days ago • 37
S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement? Paper • 2608.31100 • Published 13 days ago • 40
EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction Paper • 2609.02783 • Published 11 days ago • 120
Unlocking Lossless Speedups in LLMs via Discrete Diffusion Paper • 2609.04010 • Published 10 days ago • 144
Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning Paper • 2609.03430 • Published 10 days ago • 180
Knowing When Not to Reuse: Conditional Experience Transfer in Autonomous LLM Post-Training Paper • 2608.26730 • Published 17 days ago • 155
HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness? Paper • 2609.01437 • Published 12 days ago • 267
Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments Paper • 2609.04148 • Published 10 days ago • 292
Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving Paper • 2609.00111 • Published 13 days ago • 383
NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness Paper • 2609.08183 • Published 5 days ago • 407
Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills Paper • 2609.02749 • Published 11 days ago • 545