andersonbcdefg/sharegpt_reward_modeling_pairwise_no_as_an_ai Viewer • Updated Jun 6, 2023 • 11.8k • 141 • 5
ROSS: Relearning from Self-Generated Rollouts through Selective Supervision Paper • 2609.35954 • Published 7 days ago • 49
PackLab: A Comprehensive Framework for Developing, Training, and Evaluating MLLMs in Robotic Bin Packing Paper • 2609.23784 • Published 15 days ago • 16
Marathoner: Ultra-Long-Horizon Autonomous Intelligence Paper • 2609.34378 • Published 7 days ago • 42
Raven: The Harness of Harnesses for Composable Agentic Intelligence Paper • 2609.33439 • Published 8 days ago • 564
Think Before You Score: Thinking Reward Model for Visual Generation Paper • 2609.37372 • Published 6 days ago • 101
CompoWorld: Compositional Environment Scaling for General Agents Paper • 2609.33665 • Published 8 days ago • 42
Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL Paper • 2609.32577 • Published 9 days ago • 130
FocusVTC: Efficient and High-Performance Visual Text Compression with Adaptive Resolution Paper • 2609.36651 • Published 6 days ago • 33