HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness? Paper • 2609.01437 • Published 8 days ago • 259
Knowing When Not to Reuse: Conditional Experience Transfer in Autonomous LLM Post-Training Paper • 2608.26730 • Published 13 days ago • 150
Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments Paper • 2609.04148 • Published 6 days ago • 288
Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM Paper • 2609.04098 • Published 6 days ago • 78
Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning Paper • 2609.03430 • Published 6 days ago • 170
microsoft/VibeVoice-ASR-Streaming-1.5B Automatic Speech Recognition • 3B • Updated 5 days ago • 1.96k • 44
Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving Paper • 2609.00111 • Published 9 days ago • 376
Post-Training Language Models for Gold-Medal Performance in Coding Competitions Paper • 2609.02849 • Published 7 days ago • 10
CAR-bench: Evaluating the Consistency and Limit-Awareness of LLM Agents under Real-World Uncertainty Paper • 2601.22027 • Published Jan 29 • 86
Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090 Paper • 2608.27370 • Published 13 days ago • 38
VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction Paper • 2608.26005 • Published 14 days ago • 178
Running on CPU Upgrade Agents Featured 1.45k Open ASR Leaderboard 🏆 1.45k Explore ASR model performance across datasets
Apodex 1.1: Scaling Agentic Intelligence for Complex Work Paper • 2608.23283 • Published 16 days ago • 206
OmniAssistBench: Assistant-style Interaction Benchmark for Omni-LLMs Paper • 2608.21360 • Published 19 days ago • 31
Towards Quantifying Benchmark Optimization in ASR Models Paper • 2608.19936 • Published 20 days ago • 12
SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science? Paper • 2608.19799 • Published 20 days ago • 66