ReSPO: Reshaped Sequence Policy Optimization for Gradient Starvation in Off-Policy Learning Paper • 2609.35433 • Published 13 days ago • 12
Post-Training Frontier Text-to-Image Models by Composing Preference and Rubric Rewards Paper • 2610.02967 • Published 9 days ago • 29
Post-Training Frontier Text-to-Image Models by Composing Preference and Rubric Rewards Paper • 2610.02967 • Published 9 days ago • 29
ReCAST: Reward Credit Assignment across Timesteps for Online Diffusion Reinforcement Paper • 2609.13425 • Published about 1 month ago • 10
Stream4D: 4D-Consistency for Streaming Autoregressive Diffusion Video Models Paper • 2608.19556 • Published Aug 20 • 5
AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment Paper • 2605.17602 • Published May 20 • 17