Parameter Exploration for RLVR via Variational Learning Paper • 2608.09805 • Published 3 days ago • 1
3PO Models Collection 3PO family methods trained on DapoMath-17k using Olmo3-IVON-SFT-7B and Qwen2.5Math-IVON-SFT-7B • 10 items • Updated about 2 hours ago
Parameter Exploration for RLVR via Variational Learning Paper • 2608.09805 • Published 3 days ago • 1