Inference Providers
Active filters: RL
SII-Enigma/Qwen2.5-7B-Ins-AMPO
Text Generation
• 8B • Updated • 46
SII-Enigma/Qwen2.5-7B-Ins-SFT-GRPO
Text Generation
• 8B • Updated • 29
SII-Enigma/Llama3.2-8B-Ins-GRPO
Text Generation
• 2B • Updated • 15
• 1
mradermacher/Llama3.2-8B-Ins-GRPO-GGUF
8B • Updated • 415
• 1
SII-Enigma/Qwen2.5-7B-Ins-GRPO
Text Generation
• 2B • Updated • 23
SII-Enigma/Qwen2.5-1.5B-Ins-AMPO
Text Generation
• 2B • Updated • 30
SII-Enigma/Llama3.2-8B-Ins-AMPO
Text Generation
• 8B • Updated • 60
SII-Enigma/Qwen2.5-1.5B-Ins-GRPO
Text Generation
• 2B • Updated • 19
Text Generation
• 2B • Updated • 17
mradermacher/GCPO-R1-1.5B-GGUF
2B • Updated • 256
mradermacher/GCPO-R1-1.5B-i1-GGUF
2B • Updated • 386
mradermacher/DeepHermes-Egregore-8B-131K-GGUF
Reinforcement Learning
• 8B • Updated • 533
• 1
mradermacher/DeepHermes-Egregore-8B-131K-i1-GGUF
Reinforcement Learning
• 8B • Updated • 853
• 1
stephenchungmh/thinker_r1_5b
2B • Updated • 10
• 1
stephenchungmh/thinker_q1_5b
2B • Updated • 5
• 1
stephenchungmh/thinker_r7b
8B • Updated • 12
• 1
8B • Updated • 19
• 1
mradermacher/RENT-Qwen-7B-GGUF
8B • Updated • 216
• 1
mradermacher/RENT-Qwen-7B-i1-GGUF
8B • Updated • 1.02k
• 1
beyoru/MinCoder-4B-Expert
Text Generation
• 4B • Updated • 43
• 1
mradermacher/MinCoder-4B-Expert-GGUF
4B • Updated • 221
• 2
mradermacher/MinCoder-4B-Expert-i1-GGUF
4B • Updated • 745
• 1
Text Generation
• 4B • Updated • 1
aryan-kolapkar/MathReasoner-Mini-1.5b
Text Generation
• 2B • Updated • 18
• 1
mradermacher/MathReasoner-Mini-1.5b-GGUF
2B • Updated • 269
ryota39/Qwen3-8B-math-RL-ja
8B • Updated • 16
nvidia/Nemotron-Cascade-8B-Thinking
Text Generation
• 8B • Updated • 1.08k
• • 41
8B • Updated • 372
Reinforcement Learning
• Updated • 1
nvidia/Nemotron-Cascade-8B
Text Generation
• 8B • Updated • 607
• • 67