Demystifying On-Policy Distillation: Roles, Pathologies, and Regulations Paper • 2607.13399 • Published 6 days ago • 20
Length Penalties Make Chain-of-Thought Less Monitorable Paper • 2607.09786 • Published 13 days ago • 11
What LLM Forecasters Know but Don't Say: Probing Internal Representations for Calibration and Faithfulness Paper • 2607.08046 • Published 12 days ago • 14
ShortOPD: Recovering Pruned LLMs with Short-to-Long On-Policy Distillation Paper • 2607.13124 • Published 7 days ago • 19
Towards Autonomous and Auditable Medical Imaging Model Development Paper • 2607.10522 • Published 9 days ago • 20
MuScriptor: An Open Model for Multi-Instrument Music Transcription Paper • 2607.08168 • Published 12 days ago • 21
Know Before Fix: QA-Driven Repository Knowledge Acquisition for Software Issue Resolution Paper • 2607.11111 • Published 8 days ago • 21
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models Paper • 2607.08317 • Published 12 days ago • 34
SynthDocBench: Controlled Benchmark for Long-Context Visual Document Understanding Paper • 2607.10400 • Published 10 days ago • 70
Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation Paper • 2607.11886 • Published 8 days ago • 83
Phone Segmentation and Recognition through Phonological Activation Mapping Paper • 2607.09020 • Published 11 days ago • 9
Towards Mechanistically Understanding Why Memorized Knowledge Fails to Generalize in Large Language Model Finetuning Paper • 2607.08393 • Published 11 days ago • 18
Self-Guided Test-Time Training for Long-Context LLMs Paper • 2607.09415 • Published 11 days ago • 19
KronQ: LLM Quantization via Kronecker-Factored Hessian Paper • 2607.07964 • Published 13 days ago • 32
Can Dialects Be Steered Like Languages? Sparse Neurons and Distributed Directions in Arabic LLMs Paper • 2607.03936 • Published 17 days ago • 4
Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE Paper • 2607.07740 • Published 13 days ago • 23