PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails Paper • 2607.05910 • Published 15 days ago • 38
$π$-Bench: Evaluating Proactive Personal Assistant Agents in Long-Horizon Workflows Paper • 2605.14678 • Published May 19 • 108
AdaReasoner: Dynamic Tool Orchestration for Iterative Visual Reasoning Paper • 2601.18631 • Published Jan 26 • 48
RxSafeBench: Identifying Medication Safety Issues of Large Language Models in Simulated Consultation Paper • 2511.04328 • Published Nov 6, 2025 • 1
X-Boundary: Establishing Exact Safety Boundary to Shield LLMs from Multi-Turn Jailbreaks without Compromising Usability Paper • 2502.09990 • Published Feb 14, 2025 • 2
Bridging Interleaved Multi-Modal Reasoning as a Unified Decision Process Paper • 2607.03748 • Published 18 days ago • 39
Look Before You Leap: Distilling Tree Search into Action Evaluation for Frozen VLA Models Paper • 2607.03751 • Published 18 days ago • 19
Look Before You Leap: Distilling Tree Search into Action Evaluation for Frozen VLA Models Paper • 2607.03751 • Published 18 days ago • 19
Bridging Interleaved Multi-Modal Reasoning as a Unified Decision Process Paper • 2607.03748 • Published 18 days ago • 39
AdaReasoner: Dynamic Tool Orchestration for Iterative Visual Reasoning Paper • 2601.18631 • Published Jan 26 • 48
Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Paper • 2501.05444 • Published Jan 9, 2025 • 3
ThinkMorph: Emergent Properties in Multimodal Interleaved Chain-of-Thought Reasoning Paper • 2510.27492 • Published Oct 30, 2025 • 88
ThinkMorph: Emergent Properties in Multimodal Interleaved Chain-of-Thought Reasoning Paper • 2510.27492 • Published Oct 30, 2025 • 88
Diversity-Incentivized Exploration for Versatile Reasoning Paper • 2509.26209 • Published Sep 30, 2025 • 18
Divide and Conquer: Grounding LLMs as Efficient Decision-Making Agents via Offline Hierarchical Reinforcement Learning Paper • 2505.19761 • Published May 26, 2025
Text-to-Decision Agent: Offline Meta-Reinforcement Learning from Natural Language Supervision Paper • 2504.15046 • Published Apr 21, 2025
Attention-Guided Contrastive Role Representations for Multi-Agent Reinforcement Learning Paper • 2312.04819 • Published Dec 8, 2023
Mixture-of-Experts Meets In-Context Reinforcement Learning Paper • 2506.05426 • Published Jun 5, 2025 • 5
Unfolding Spatial Cognition: Evaluating Multimodal Models on Visual Simulations Paper • 2506.04633 • Published Jun 5, 2025 • 21
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Paper • 2505.14810 • Published May 20, 2025 • 63