SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Paper • 2608.03092 • Published Aug 4 • 9
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation Paper • 2608.03092 • Published Aug 4 • 9 • 1
DanhVuiVe/ChartQA_Benetech_PlotQa_DVQA_combined_matcha_complete Viewer • Updated Oct 29, 2024 • 535k • 60 • 2
alimama-creative/FLUX.1-dev-Controlnet-Inpainting-Beta Image-to-Image • 2B • Updated Oct 12, 2024 • 8.61k • 430
openai/whisper-large-v3-turbo Automatic Speech Recognition • 0.8B • Updated Oct 4, 2024 • 6.87M • • 3.3k