Change8

vLLM

AI & LLMs

A high-throughput and memory-efficient inference and serving engine for LLMs

Latest: v0.28.025 releases7 breaking changes23 common errorsView on GitHub

Release History

View all versions →
v0.28.0Breaking5 fixes29 features
Aug 26, 2026

This release introduces significant performance optimizations for Kimi-K3 and DeepSeek V4, alongside advancements in speculative decoding and Model Runner V2 maturation. It also features a new Rust frontend and gRPC capabilities, tiered KV cache offloading, and expanded model support.

v0.27.2rc01 feature
Aug 12, 2026

This release introduces confidence-scheduled verification for DSpark spec decoding, enhancing its verification capabilities.

v0.27.11 feature
Aug 11, 2026

This is a patch release that adds support for quantized DSpark Markov heads.

v0.27.0Breaking7 fixes24 features
Aug 10, 2026

vLLM v0.27.0 introduces Kimi K3 support, new model additions like Qwen3.5 and VaultGemma, and a significant upgrade to PyTorch 2.13.0. The release also enhances performance and features across various areas including FlashAttention 4, Model Runner V2, KV offloading, and hardware enablement.

v0.27.0rc21 feature
Aug 9, 2026

This release allows TPU to import kimi_k3.common, addressing issue #51529.

v0.27.0rc11 fix
Aug 7, 2026

This release includes a bug fix that preserves ModelOpt FP8 weight dimensions. This change ensures the correct handling of FP8 weights in model optimization.

v0.26.1rc01 fix
Jul 27, 2026

This release includes a fix for the `test_ocp_mx_wikitext_correctness` reference value within the CI/ROCm environment.

v0.26.03 fixes82 features
Jul 25, 2026

vLLM v0.26.0 introduces the Inkling model family, significant performance boosts for DeepSeek-V4, flexible attention backends, and enhanced KV offloading. The release also includes a Rust frontend with multimodal capabilities and deeper integration with Transformers 5.13.0.

v0.26.0rc11 fix
Jul 23, 2026

This release includes a bug fix that prevents engine crashes by properly handling grammar compilation failures.

v0.25.12 fixes
Jul 14, 2026

v0.25.1 is a patch release that fixes two critical bugs. The first prevents model launching from being blocked when FFmpeg is not available for TorchCodec, and the second guards mixed-dtype allreduce RMSNorm quant fusions to prevent corrupted outputs.

v0.25.0Breaking15 fixes33 features
Jul 11, 2026

vLLM v0.25.0 introduces Model Runner V2 as the default for dense models, significantly improving performance and adding support for new features like EVS and realtime embeddings. The release also deprecates PagedAttention and enhances the Transformers backend to match native vLLM speed, alongside numerous model additions and performance optimizations across various hardware platforms.

v0.25.0rc31 fix
Jul 9, 2026

This release includes a bug fix for PD async KV load lookahead handling in MTP spec decode.

v0.25.0rc21 fix
Jul 9, 2026

This release fixes issues with embed scaling and CUDA graphs in the Transformers modeling backend.

v0.25.0rc11 fix
Jul 8, 2026

This release includes a bug fix for a flaky test on ARM architectures related to ShortConv prefill. The fix addresses issues with uninitialized weights.

v0.24.0Breaking19 fixes36 features
Jun 29, 2026

v0.24.0 introduces extensive support and performance optimizations for new models like MiniMax-M3 and DeepSeek-V4, matures the Model Runner V2 with default quantization support, and overhauls device selection by removing internal use of CUDA_VISIBLE_DEVICES.

v0.24.0rc21 fix
Jun 25, 2026

This release includes a bug fix for issues related to P/D with DP Supervisor.

v0.24.0rc11 fix
Jun 24, 2026

This release addresses a bug in the CI/Build process, specifically fixing the topk histogram build on SM75 hardware.

v0.23.0Breaking30 fixes25 features
Jun 12, 2026

v0.23.0 brings significant hardening and optimization for DeepSeek-V4, expands Model Runner V2 to Llama/Mistral models, and advances the experimental Rust frontend. This release also mandates compatibility with Transformers v5.

v0.22.16 fixes2 features
Jun 5, 2026

v0.22.1 is a patch release introducing support for Mellum v2 and enabling quantized inference acceleration on AMD Zen CPUs. It also includes several critical fixes for model initialization, Ray serving stability, and build issues.

v0.22.023 fixes95 features
May 29, 2026

This release focuses heavily on DeepSeek V4 maturity with new kernel support and packaging, significant advancements in Model Runner V2, and the introduction of an experimental Rust frontend. Performance saw notable gains from batch-invariant inference with Cutlass FP8 and the rollout of multi-tier KV cache offloading.

v0.21.0Breaking15 fixes22 features
May 14, 2026

This release introduces significant performance and stability improvements, notably integrating KV offloading with the Hybrid Memory Allocator and enabling speculative decoding with thinking budgets. It also formally deprecates support for older versions of the Transformers library.

v0.20.24 fixes
May 10, 2026

vLLM v0.20.2 is a small patch release focused on bug fixes for DeepSeek V4, gpt-oss, and Qwen3-VL models.

v0.20.111 fixes6 features
May 3, 2026

vLLM v0.20.1 is a patch release focused on stabilizing and improving performance for DeepSeek V4, including various kernel optimizations and critical bug fixes across CUDA and ROCm platforms.

v0.20.0Breaking6 fixes10 features
Apr 23, 2026

v0.20.0 introduces major infrastructure upgrades, including a default switch to CUDA 13.0 and PyTorch 2.11, alongside significant performance enhancements like TurboQuant 2-bit KV cache and the re-enabling of FlashAttention 4 as default prefill backend.

v0.19.17 fixes2 features
Apr 18, 2026

This patch release upgrades to Transformers v5.5.4 and delivers numerous bug fixes specifically targeting Gemma4 streaming, tool calls, and model loading, alongside adding support for Gemma4 Eagle3 and quantized MoE.

Common Errors

Related AI & LLMs Packages

Subscribe to Updates

Get notified when new versions are released

RSS Feed