Change8

Ollama

AI & LLMs

Get up and running with OpenAI gpt-oss, DeepSeek-R1, Gemma 3 and other models.

Latest: v0.34.0-rc125 releases2 breaking changes1 common errorsView on GitHub

Release History

View all versions →
v0.34.0-rc11 fix3 features
Sep 5, 2026

This release introduces the ability to use Ollama models directly in ChatGPT Desktop and improves structured output performance on Apple Silicon. It also adds support for OpenAI-compatible client tool search and response compaction, with bug fixes for image handling in compacted responses.

v0.33.3-rc01 feature
Sep 2, 2026

This release updates MLX, MLX-C, and llama.cpp dependencies and ensures GGUF models honor their default parameters.

v0.33.3-rc22 features
Sep 2, 2026

This release introduces reporting of cached prompt tokens and ensures GGUF models respect their default parameters. It also includes updates to MLX, MLX-C, and llama.cpp.

v0.33.3-rc12 features
Sep 2, 2026

This release introduces features for reporting cached prompt tokens and honoring GGUF model default parameters. It also includes updates to MLX, MLX-C, and llama.cpp.

v0.33.33 features
Sep 2, 2026

This release introduces support for images and audio with Gemma 4 on the MLX engine, reports cached prompt tokens, and honors GGUF model default parameters. It also includes updates to MLX, MLX-C, and llama.cpp.

v0.33.23 fixes
Aug 27, 2026

This release restores dark mode support for the Ollama app, fixes a macOS specific bug with instance handling, and improves the Claude Desktop proxy's stability during model catalog updates.

v0.33.2-rc13 features
Aug 27, 2026

This release restores system dark mode for the app, allows proxy requests to continue during model catalog changes, and synchronizes macOS app handoff.

v0.33.15 features
Aug 26, 2026

This release brings updates to MLX and llama.cpp, adds structured output support to mlxrunner, and improves GPU timeout handling.

v0.33.1-rc11 fix3 features
Aug 26, 2026

This release introduces Qwen3.8 Flash Next support for MLX and adds structured output capabilities to mlxrunner. It also includes improvements to cmake compatibility patches and addresses Metal GPU timeouts.

v0.33.0-rc03 fixes4 features
Aug 21, 2026

This release introduces the Claude desktop app and enhances the user experience with polished onboarding and a new 'Connect your apps' feature. It also includes several bug fixes for MLX and the DeepSeek Harness.

v0.33.0-rc35 fixes6 features
Aug 21, 2026

This release introduces new features for managing Ollama models within Claude Desktop, improves prefill caching for more reliable retries, and includes several bug fixes for packaging and UI elements.

v0.33.0-rc22 fixes4 features
Aug 21, 2026

This release introduces several new features for the desktop application, including Claude support and a "Connect your apps" experience. It also includes bug fixes for the MLX runner and general linting improvements.

v0.33.05 fixes6 features
Aug 21, 2026

This release introduces new features for managing Ollama models within Claude Desktop, improves prefill caching for more reliable retries, and includes several bug fixes for cross-platform compatibility and UI elements.

v0.32.151 feature
Aug 19, 2026

Introduced a model metadata cache to improve Ollama's performance by reducing per-request overhead. This release also acknowledges a new contributor.

v0.32.15-rc11 feature
Aug 19, 2026

Introduced a model metadata cache to improve Ollama's performance by reducing per-request overhead. This release also acknowledges a new contributor.

v0.32.14-rc01 fix1 feature
Aug 15, 2026

This release introduces WebP image transcoding for llama-server and improves handling of non-leading system messages in Qwen renderers.

v0.32.141 fix1 feature
Aug 15, 2026

This release introduces WebP image transcoding for llama-server and improves handling of non-leading system messages in Qwen renderers.

v0.32.131 feature
Aug 14, 2026

This release adds support for developer instructions for the qwen3.8 model. It provides a link to the full changelog for more details.

v0.32.122 features
Aug 14, 2026

This release introduces support for the Qwen 3.8 27B model, with specific optimizations for Apple Silicon devices to enhance performance for agents and repeated tasks.

v0.32.112 features
Aug 14, 2026

This release introduces new integrations for Muse Code and DeepSeek Harness, enhancing the launch capabilities of Ollama.

v0.32.10Breaking1 fix1 feature
Aug 12, 2026

This release adjusts the default repeat penalty for models to 1.0, improving speculative decoding speed, and introduces faster prefill performance for certain MLX models. A bug in blob verification has also been fixed.

v0.32.10-rc1Breaking1 fix1 feature
Aug 12, 2026

This release adjusts the default repeat penalty for models and introduces performance improvements for NVFP4 MLX models. A bug related to blob verification has also been fixed.

v0.32.10-rc0
Aug 12, 2026

Optimized prefill performance for double-scale nvfp4 models by compiling multiply and cast operations into a single kernel, reducing kernel launches and intermediate materialization.

v0.32.91 fix1 feature
Aug 11, 2026

Introduced the Nemotron 3 architecture and fixed a boundary condition in the Muse Glimmer function calling parser.

v0.32.83 features
Aug 10, 2026

Ollama now supports the Muse Glimmer model across various platforms, including NVIDIA and AMD, enhancing capabilities for coding agents and personal assistants. MLX engine on Apple Silicon also sees improvements.

Common Errors

Related AI & LLMs Packages

Subscribe to Updates

Get notified when new versions are released

RSS Feed