CUDA
Found in 4 packages: llama-cpp, invokeai, unsloth, ollama
llama-cpp(8 releases)
v0.2.0This release bumps the llama.cpp version to 0.2.0 and ggml to 0.21.0, incorporating numerous backend improvements, bug fixes, and new model support across various hardware accelerators like SYCL, OpenCL, Metal, and Vulkan.
b10353BreakingThis release addresses an issue with the ggml ROLL operation on CUDA and Metal, ensuring correct results by requiring contiguous source tensors. It also includes various pre-compiled binaries for different platforms and hardware accelerators.
b10106This release includes a fix for external compilation of q1_0 MMQ for CUDA. It also provides pre-compiled binaries for various platforms including macOS, Linux, Android, and Windows with different hardware acceleration options.
b7619This release introduces a CUDA optimization to reduce memory overhead by conditionally allocating the Flash Attention temporary buffer. It includes a wide range of pre-built binaries for multiple operating systems and hardware architectures.
b7591This release primarily addresses a bug in the CUDA backend related to KQ max calculation and provides updated binaries for multiple platforms including Windows, macOS, and Linux.
b7567This release introduces NVIDIA Blackwell architecture support for non-native CUDA builds and provides updated binaries across multiple platforms including Windows, macOS, and Linux.
b7345BreakingThis release fixes a stride padding issue in the CUDA MMA Flash Attention kernel and announces a transition for Linux release binaries from .zip to .tar.gz format.
b7339BreakingThis release introduces DIAG support for CUDA and announces a transition in Linux package formats from .zip to .tar.gz.
invokeai(1 releases)
unsloth(3 releases)
v0.1.61-betaThis release introduces Meta's Muse Glimmer 30B open model and adds video generation and preliminary image diffusion support. It also includes numerous bug fixes and UI enhancements for the Unsloth Studio and Desktop applications.
v0.1.60-betaThis release introduces Meta's Muse Glimmer 30B model, optimized for local agentic and coding workflows, and enhances Unsloth's Studio and Desktop applications with numerous bug fixes and UI improvements.
v0.1.527-betaThis release includes numerous Studio and Desktop improvements, bug fixes, and feature enhancements. Key updates involve better handling of cached pipelines, improved UI elements in Studio, and more robust desktop application behavior.
ollama(2 releases)
v0.13.1This release introduces support for Ministral-3 and Mistral-Large-3 models, adds tool calling for cogito-v2.1, and includes several fixes for CUDA detection and error reporting.
v0.13.0This release introduces support for DeepSeek-OCR, Cogito-V2.1, and DeepSeek-V3.1 architecture, alongside a new performance benchmarking tool and significant engine optimizations for KV caching and GPU detection.
Track Symbol Changes
Use the Change8 MCP server or GitHub Action to get notified when CUDA changes.
Learn More