Block3D: Efficient Text-to-3D Generation via Block-Wise Diffusion Paper • 2608.19567 • Published 6 days ago • 27
Unlocking the Potential of Image Editing via Concept Scaling and Dense Supervision Paper • 2608.16812 • Published 9 days ago • 47
EviRank: Structured Relevance Evidence for Multimodal Image Re-ranking Paper • 2608.20886 • Published 5 days ago • 11
EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement Learning Paper • 2608.06197 • Published 20 days ago • 46
Memory Decoder at Scale: A Pretrained, Parametric Long-Term Memory Paper • 2607.27919 • Published 27 days ago • 59
Beyond Relevance-Centric Retrieval: Rubric-Oriented Document Set Selection and Ranking Paper • 2607.19747 • Published Jul 22 • 30
WeaveBench: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces Paper • 2606.09426 • Published Jun 8 • 107
LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding Paper • 2605.27365 • Published May 26 • 147
TriSplat: Simulation-Ready Feed-Forward 3D Scene Reconstruction Paper • 2605.26115 • Published May 25 • 53
Flash-GRPO: Efficient Alignment for Video Diffusion via One-Step Policy Optimization Paper • 2605.15980 • Published May 15 • 36
Warp-as-History: Generalizable Camera-Controlled Video Generation from One Training Video Paper • 2605.15182 • Published May 14 • 39
Rewarding the Scientific Process: Process-Level Reward Modeling for Agentic Data Analysis Paper • 2604.24198 • Published Apr 27 • 23
World-R1: Reinforcing 3D Constraints for Text-to-Video Generation Paper • 2604.24764 • Published Apr 27 • 120
Feed-Forward 3D Scene Modeling: A Problem-Driven Perspective Paper • 2604.14025 • Published Apr 15 • 15