Towards Looped Models Done Right, Part II: Rethinking at Fixed Points Paper • 2610.06833 • Published 4 days ago • 30
Base Models Can Reason By Taking a Cue From Training Data Paper • 2610.06851 • Published 4 days ago • 22
LoGRA: Scaling LLM Reinforcement Learning with Low-Rank Gradient Sketches Paper • 2610.06647 • Published 4 days ago • 121
Rethinking Long-Video Efficiency: A Joint Allocation Perspective on Frames, Pixels, and Front-End Latency Paper • 2610.04318 • Published 6 days ago • 18
LexReward: A Taxonomy-Driven Reward Framework for Legal Language Models Paper • 2609.39071 • Published 9 days ago • 63
MetaRubric: Learning to Reward for Rubric-Based Reinforcement Learning Paper • 2610.02824 • Published 7 days ago • 32
Diptych: Scoped, AI-Interpreted Comparison for Reference Listening in Music Production Paper • 2609.39963 • Published 9 days ago • 21
OmniSeek: Native Tool Integration for Multi-turn Audio-Visual Reasoning Paper • 2610.02181 • Published 8 days ago • 21
On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics Paper • 2609.35259 • Published 11 days ago • 198
Make Sparse Rewards Count: Density-Aware Reward Aggregation for Multi-Reward RL Paper • 2610.00574 • Published 9 days ago • 64
Adaptive Reward Routing: Dynamic Multi-Reward Optimization for Joint Audio-Video Diffusion via Forward-Process RL Paper • 2609.37200 • Published 10 days ago • 138
Persona Dosing: Calibrated Activation Steering for Graded Trait Control Paper • 2609.36388 • Published 11 days ago • 54
Thinking Outside the Box: Can Language Models Rely on External Guidance Selectively? Paper • 2609.39578 • Published 9 days ago • 68
TERRA: Terrain-Aware Reconstruction, Retargeting and Control for Musculoskeletal Locomotion Paper • 2609.38653 • Published 10 days ago • 37
The Teacher Is a Direction, Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation Paper • 2609.36484 • Published 10 days ago • 581
AdviSD: Learning to Advise Frontier LLMs via Targeted Multi-Turn Self-Distillation Paper • 2609.38142 • Published 10 days ago • 12
UniEvo-VL: An On-policy Self-Distillation Training Recipe for Multimodal Model Self-improvement Paper • 2609.38721 • Published 9 days ago • 293