QwenGyre: An Elastic Reinforcement Learning Framework for Training xLong-Horizon Agents Paper • 2609.33848 • Published 11 days ago • 43
Outcome Accuracy is Not Enough: Aligning the Reasoning Process of Reward Models Paper • 2602.04649 • Published Feb 4 • 14
QwenGyre: An Elastic Reinforcement Learning Framework for Training xLong-Horizon Agents Paper • 2609.33848 • Published 11 days ago • 43
Deeper is Not Always Better: Mitigating the Alignment Tax via Confident Layer Decoding Paper • 2606.21906 • Published Jun 20 • 27
Deeper is Not Always Better: Mitigating the Alignment Tax via Confident Layer Decoding Paper • 2606.21906 • Published Jun 20 • 27
HopChain: Multi-Hop Data Synthesis for Generalizable Vision-Language Reasoning Paper • 2603.17024 • Published Mar 17 • 111