SynthDocBench: Controlled Benchmark for Long-Context Visual Document Understanding Paper • 2607.10400 • Published 21 days ago • 71
japanese-asr/whisper_transcriptions.reazon_speech_all Viewer • Updated Sep 14, 2024 • 17.3M • 30.1k • 15
A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook Paper • 2605.20266 • Published May 18 • 56
Evaluating Cognitive Age Alignment in Interactive AI Agents Paper • 2605.17894 • Published May 18 • 5
Learning to Explore: Scaling Agentic Reasoning via Exploration-Aware Policy Optimization Paper • 2605.08978 • Published May 12 • 2
SEIF: Self-Evolving Reinforcement Learning for Instruction Following Paper • 2605.07465 • Published May 8 • 30
Mean Mode Screaming: Mean--Variance Split Residuals for 1000-Layer Diffusion Transformers Paper • 2605.06169 • Published May 7 • 238