BF16
Found in 3 packages: unsloth, vllm, llama-cpp
unsloth(1 releases)
vllm(2 releases)
v0.26.0vLLM v0.26.0 introduces the Inkling model family, significant performance boosts for DeepSeek-V4, flexible attention backends, and enhanced KV offloading. The release also includes a Rust frontend with multimodal capabilities and deeper integration with Transformers 5.13.0.
v0.25.0BreakingvLLM v0.25.0 introduces Model Runner V2 as the default for dense models, significantly improving performance and adding support for new features like EVS and realtime embeddings. The release also deprecates PagedAttention and enhances the Transformers backend to match native vLLM speed, alongside numerous model additions and performance optimizations across various hardware platforms.
llama-cpp(2 releases)
b9890This release focuses on bug fixes across CUDA, CDNA, and BF16 logic, alongside a refactoring of the cuBLAS implementation. It also provides numerous pre-compiled binaries for different operating systems and hardware configurations.
b7306BreakingThis release fixes RDNA3 matrix multiplication for HIP and announces a transition in Linux release packaging from .zip to .tar.gz.
Track Symbol Changes
Use the Change8 MCP server or GitHub Action to get notified when BF16 changes.
Learn More