Inference Providers
Active filters: grpo
ByteDance/UniVR-34B-Planning
Image-Text-to-Text
• Updated • 15
• 11
Danau5tin/ai-trains-ai-trainer
Text Generation
• Updated • 12
• 4
kings-crown/ProofSeeker_v1
7B • Updated • 1
trentmkelly/Llama-3.1-8b-Instruct-Pangram
Text Generation
• Updated • 5
DATEXIS/DeepICD-R1-Llama-8B
Text Generation
• 8B • Updated • 6
• 2
Aion2/llama3.2-3b-grpo-v1
Text Generation
• Updated • 1
wheattoast11/OmniCoder-9B-Zero-Phase2
Text Generation
• Updated • 44
• 1
sajan-sarker/Qwen3.5-4B-LoRA-GRPO-CyberSec-Reasoner
Text Generation
• 4B • Updated • 3
• 1
Text Generation
• 4B • Updated • 856
• • 29
alireza7/GrepSeek-Qwen3.5-9B-GRPO
Text Generation
• 9B • Updated • 353
• 6
Text Generation
• 8B • Updated • 25
• 1
mradermacher/AAPA-8B-GGUF
8B • Updated • 524
• 1
mradermacher/AAPA-8B-i1-GGUF
8B • Updated • 1.07k
• 2
lxazjk/qwen2.5-1.5b-24game-grpo
Text Generation
• 2B • Updated • 308
• 1
artichoke42/Qwen3.6-27B-KR-MTP-GGUF
Text Generation
• 3.05M • Updated • 555
• 1
Text Generation
• 9B • Updated • 638
• 1
Image-Text-to-Text
• 8B • Updated • 23
• 1
sandeeprdy1729/TIMPS-Coder-7B
Text Generation
• 8B • Updated • 758
• 1
WilliamSheltonWu/Confession-Qwen3_5_9B
Reinforcement Learning
• 9B • Updated • 102
• 1
SESPOIR/ReGround-Qwen2.5-VL-7B
Image-Text-to-Text
• 849k • Updated • 21
• 1
Text Generation
• 8B • Updated • 16
• 1
Text Generation
• 0.1B • Updated • 23
8B • Updated • 3
sergiopaniego/Qwen2-0.5B-GRPO-test
Updated
Novaciano/ESP-NSFW-GRPO-1B-Sin_Censura-GGUF
1B • Updated • 178
• 6
nbd22/Llama-3.1-8B-Instruct-GRPO-gsm8k-ft-lora
Updated
sergiopaniego/Qwen2-0.5B-GRPO
Updated
philschmid/qwen-2.5-3b-r1-countdown
Text Generation
• 3B • Updated • 14
• 8
spinech/qwen-2.5-3b-r1-countdown
Text Generation
• 3B • Updated • 5
Dongwei/Qwen2.5-1.5B-Open-R1-GRPO
Text Generation
• 2B • Updated • 5
• 1