The Embedder's Dilemma: LLMs Are Better, but at What Cost? Paper • 2608.12875 • Published 23 days ago • 15
SWE-Marathon: Can Agents Autonomously Complete Ultra-Long-Horizon Software Work? Paper • 2606.07682 • Published Jun 5 • 2
HUME: Measuring the Human-Model Performance Gap in Text Embedding Task Paper • 2510.10062 • Published Oct 11, 2025 • 10