GraphVid: Interactive Graph-Controllable Video Generation Paper • 2607.21580 • Published 6 days ago • 7
SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation Paper • 2607.21553 • Published 6 days ago • 37
Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers Paper • 2607.21594 • Published 6 days ago • 14
AREX: Towards a Recursively Self-Improving Agent for Deep Research Paper • 2607.21461 • Published 6 days ago • 146
ATSplat: Compact Feed-forward 3D Gaussian Splatting with Adaptive Token Expansion Paper • 2607.20417 • Published 7 days ago • 9
SeerGuard: A Safety Framework for Mobile GUI Agents via World Model Prediction Paper • 2607.15550 • Published 12 days ago • 26
Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning Paper • 2607.14183 • Published 14 days ago • 68
Hierarchical Denoising For Multi-Step Visual Reasoning Paper • 2607.15278 • Published 13 days ago • 6
ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU Paper • 2607.19191 • Published 7 days ago • 302
AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report Paper • 2607.18367 • Published 8 days ago • 57
Apple-π: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence Paper • 2607.16401 • Published 12 days ago • 43
FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications Paper • 2607.18171 • Published 9 days ago • 6
HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enchancement Paper • 2607.18217 • Published 9 days ago • 59
EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary World Paper • 2607.17250 • Published 10 days ago • 92
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Paper • 2607.17423 • Published 10 days ago • 165
DSWorld: A Data Science World Model for Efficient Autonomous Agents Paper • 2607.15901 • Published 12 days ago • 12
RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Paper • 2607.14187 • Published 14 days ago • 31