GRPO
Found in 1 package: unsloth
unsloth(4 releases)
v0.1.511-betaThis release introduces support for Kimi K3 and DeepSeek v4 Flash with Unsloth Dynamic GGUFs, enables parallel chat generation, and adds a Deep Research mode. It also brings significant improvements to AMD and Intel GPU support, DoRA training, and various installer, MLX, export, and inference fixes.
v0.1.512-betaThis release introduces local support for Kimi K3 and DeepSeek v4 Flash models via Unsloth Dynamic GGUFs, alongside parallel chat capabilities and a new Deep Research mode. It also brings significant improvements to AMD and Intel GPU support, DoRA training, and various installer, MLX, export, and inference fixes.
v0.1.48-betaThis release significantly enhances Unsloth Studio with new export formats (NVFP4, FP8, imatrix GGUF) and introduces an OpenAI-compatible API serving system with automatic model swapping. Core improvements include substantial speedups for GRPO and MoE training, alongside numerous reliability fixes across installation, RAG, and training workflows.
July-2025This release focuses heavily on stability, VRAM reduction (10-25% less), and broad model compatibility, including full fixes for Gemma 3N Vision and support for new models like Devstral 1.1 and MedGemma.
Track Symbol Changes
Use the Change8 MCP server or GitHub Action to get notified when GRPO changes.
Learn More