AsyncLLM
Found in 1 package: vllm
vllm(3 releases)
v0.27.0BreakingvLLM v0.27.0 introduces Kimi K3 support, new model additions like Qwen3.5 and VaultGemma, and a significant upgrade to PyTorch 2.13.0. The release also enhances performance and features across various areas including FlashAttention 4, Model Runner V2, KV offloading, and hardware enablement.
v0.10.1Breakingv0.10.1 introduces support for Blackwell and RTX 5090 GPUs, expands vision-language model compatibility, and adds a plugin system for model loaders. It also deprecates V0 FA3 support and removes AQLM quantization.
v0.8.0rc1BreakingThis release introduces Expert Parallelism for DeepSeek models, a new /score endpoint for embeddings, and significant V1 engine enhancements including parallel sampling. It also removes global seed setting, requiring users to manually define seeds for reproducibility.
Track Symbol Changes
Use the Change8 MCP server or GitHub Action to get notified when AsyncLLM changes.
Learn More