ROCm
Found in 4 packages: unsloth, vllm, invokeai, ollama
unsloth(5 releases)
v0.1.801-betaThis release introduces experimental Auto Compaction for longer chats and a preview of Remote & LAN Access. It also brings significant speed improvements to chat, support for custom llama.cpp builds, and enhanced features for the Unsloth Dynamic GGUFs.
v0.1.800-betaBreakingUnsloth now supports running Qwen3.8 models locally with reduced RAM requirements and offers significant performance improvements, including faster inference for GGUFs and MiniMax-H3. New features enhance chat capabilities, tool integration, and hardware compatibility.
v0.1.511-betaThis release introduces support for Kimi K3 and DeepSeek v4 Flash with Unsloth Dynamic GGUFs, enables parallel chat generation, and adds a Deep Research mode. It also brings significant improvements to AMD and Intel GPU support, DoRA training, and various installer, MLX, export, and inference fixes.
v0.1.51-betaThis release introduces Kimi K3 local execution, parallel chat capabilities, and a Deep Research mode. It also brings significant improvements to AMD and Intel GPU support, DoRA training, and various installer, MLX, export, and inference fixes.
v0.1.512-betaThis release introduces local support for Kimi K3 and DeepSeek v4 Flash models via Unsloth Dynamic GGUFs, alongside parallel chat capabilities and a new Deep Research mode. It also brings significant improvements to AMD and Intel GPU support, DoRA training, and various installer, MLX, export, and inference fixes.
vllm(3 releases)
v0.27.0BreakingvLLM v0.27.0 introduces Kimi K3 support, new model additions like Qwen3.5 and VaultGemma, and a significant upgrade to PyTorch 2.13.0. The release also enhances performance and features across various areas including FlashAttention 4, Model Runner V2, KV offloading, and hardware enablement.
v0.26.0vLLM v0.26.0 introduces the Inkling model family, significant performance boosts for DeepSeek-V4, flexible attention backends, and enhanced KV offloading. The release also includes a Rust frontend with multimodal capabilities and deeper integration with Transformers 5.13.0.
v0.25.0BreakingvLLM v0.25.0 introduces Model Runner V2 as the default for dense models, significantly improving performance and adding support for new features like EVS and realtime embeddings. The release also deprecates PagedAttention and enhances the Transformers backend to match native vLLM speed, alongside numerous model additions and performance optimizations across various hardware platforms.
invokeai(2 releases)
v6.13.5This maintenance release focuses on stability and bug fixes across various components, including UI responsiveness, model loading stability (especially with FP8 and Flux.2), and CI/CD pipeline improvements. Future major features like video generation are noted for version 6.14.0.
v6.13.5.rc1This maintenance release focuses on stability and bug fixes across the UI, model loading, and CI/CD pipelines. Key updates include support for PEFT named-adapter LoRAs and dependency upgrades for React, ROCm, and Transformers.
ollama(3 releases)
v0.12.7This release introduces support for Qwen3-VL and MiniMax-M2 models, adds file attachments and thinking level adjustments to the app, and provides updated API documentation alongside several embedding and backend bug fixes.
v0.12.5BreakingThis release introduces structured output support for thinking models and improves app startup behavior, while removing support for older macOS versions and specific AMD GPU architectures.
v0.12.4BreakingThis release enables Flash Attention by default for Qwen 3 models and improves VRAM detection, while dropping support for older macOS versions and specific AMD GPU architectures.
Track Symbol Changes
Use the Change8 MCP server or GitHub Action to get notified when ROCm changes.
Learn More