WarpSAC: Towards the Pinnacle of Scalable Off-policy RL by Rethinking Exploration and Exploitation Paper • 2608.24479 • Published 13 days ago • 142
Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill Paper • 2608.11924 • Published 26 days ago • 291
Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA Paper • 2608.09819 • Published 28 days ago • 342
trl-internal-testing/tiny-Qwen2ForCausalLM-2.5 Text Generation • 2.43M • Updated Dec 19, 2025 • 17.6M • 21
Quickest Detection of Hallucination Onset: Delay Bounds and Learned CUSUM Statistics Paper • 2606.12476 • Published Jun 10 • 2