ReactVAU: A Slow-Fast Decoupled Framework for Streaming Video Anomaly Understanding Paper • 2609.07941 • Published 5 days ago • 15
Scal3R: Learning Efficient Multi-Relative Pose Query for Scalable Online 3D Reconstruction Paper • 2609.04201 • Published 9 days ago • 46
ReactVAU: A Slow-Fast Decoupled Framework for Streaming Video Anomaly Understanding Paper • 2609.07941 • Published 5 days ago • 15
PhysCaP: Grounding Code-as-Policy Agent with Physics-Informed Exploration Paper • 2608.21031 • Published 22 days ago • 4
PhysCaP: Grounding Code-as-Policy Agent with Physics-Informed Exploration Paper • 2608.21031 • Published 22 days ago • 4
SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning Paper • 2606.13673 • Published Jun 11 • 113
See What I See, Know What I Think: Dense Latent Communication Across Heterogeneous Agents Paper • 2606.13594 • Published Jun 11 • 6
SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning Paper • 2606.13673 • Published Jun 11 • 113
SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning Paper • 2606.13673 • Published Jun 11 • 113
VoLo: A Physical Orchestrator for Open-Vocabulary Long-Horizon Manipulation Paper • 2606.07723 • Published Jun 5 • 6
FineBench: Benchmarking and Enhancing Vision-Language Models for Fine-grained Human Activity Understanding Paper • 2605.19846 • Published May 20 • 3
Physics in 2-Steps: Locking Motion Priors Before Visual Refinement Erases Them Paper • 2606.06361 • Published Jun 4 • 16
Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models Paper • 2501.14818 • Published Jan 20, 2025 • 10
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Paper • 2502.05178 • Published Feb 7, 2025 • 10
GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Paper • 2503.14734 • Published Mar 18, 2025 • 9
Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models Paper • 2504.03624 • Published Apr 4, 2025 • 20
OpenVision 3: A Family of Unified Visual Encoder for Both Understanding and Generation Paper • 2601.15369 • Published Jan 21 • 22
Fully Attentional Networks with Self-emerging Token Labeling Paper • 2401.03844 • Published Jan 8, 2024