Let Confidence Change, Not the Prediction: Prediction-Preserving Repair for Post-hoc Calibration Paper • 2609.01072 • Published 4 days ago • 13
LatentPress: Context Compression Beyond Text and Vision Paper • 2609.01507 • Published 5 days ago • 108
SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation Paper • 2608.17426 • Published 19 days ago • 158
Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination Paper • 2608.14391 • Published 23 days ago • 282
OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution Paper • 2608.00677 • Published Aug 1 • 263
FactorJEPA: Factorizing Monolithic Futures into Layout-Agent-Interaction Channels for Crowded and Chaotic Global South Urban Worlds Paper • 2608.01049 • Published Aug 2 • 13
GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks Paper • 2608.03764 • Published Aug 4 • 27
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Paper • 2608.02023 • Published Aug 3 • 158
ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction Paper • 2607.29677 • Published Jul 31 • 25