Ollama
AI & LLMsGet up and running with OpenAI gpt-oss, DeepSeek-R1, Gemma 3 and other models.
Release History
View all versions →v0.34.0-rc11 fix3 featuresThis release introduces the ability to use Ollama models directly in ChatGPT Desktop and improves structured output performance on Apple Silicon. It also adds support for OpenAI-compatible client tool search and response compaction, with bug fixes for image handling in compacted responses.
v0.33.3-rc01 featureThis release updates MLX, MLX-C, and llama.cpp dependencies and ensures GGUF models honor their default parameters.
v0.33.3-rc22 featuresThis release introduces reporting of cached prompt tokens and ensures GGUF models respect their default parameters. It also includes updates to MLX, MLX-C, and llama.cpp.
v0.33.3-rc12 featuresThis release introduces features for reporting cached prompt tokens and honoring GGUF model default parameters. It also includes updates to MLX, MLX-C, and llama.cpp.
v0.33.33 featuresThis release introduces support for images and audio with Gemma 4 on the MLX engine, reports cached prompt tokens, and honors GGUF model default parameters. It also includes updates to MLX, MLX-C, and llama.cpp.
v0.33.23 fixesThis release restores dark mode support for the Ollama app, fixes a macOS specific bug with instance handling, and improves the Claude Desktop proxy's stability during model catalog updates.
v0.33.2-rc13 featuresThis release restores system dark mode for the app, allows proxy requests to continue during model catalog changes, and synchronizes macOS app handoff.
v0.33.15 featuresThis release brings updates to MLX and llama.cpp, adds structured output support to mlxrunner, and improves GPU timeout handling.
v0.33.1-rc11 fix3 featuresThis release introduces Qwen3.8 Flash Next support for MLX and adds structured output capabilities to mlxrunner. It also includes improvements to cmake compatibility patches and addresses Metal GPU timeouts.
v0.33.0-rc03 fixes4 featuresThis release introduces the Claude desktop app and enhances the user experience with polished onboarding and a new 'Connect your apps' feature. It also includes several bug fixes for MLX and the DeepSeek Harness.
v0.33.0-rc35 fixes6 featuresThis release introduces new features for managing Ollama models within Claude Desktop, improves prefill caching for more reliable retries, and includes several bug fixes for packaging and UI elements.
v0.33.0-rc22 fixes4 featuresThis release introduces several new features for the desktop application, including Claude support and a "Connect your apps" experience. It also includes bug fixes for the MLX runner and general linting improvements.
v0.33.05 fixes6 featuresThis release introduces new features for managing Ollama models within Claude Desktop, improves prefill caching for more reliable retries, and includes several bug fixes for cross-platform compatibility and UI elements.
v0.32.151 featureIntroduced a model metadata cache to improve Ollama's performance by reducing per-request overhead. This release also acknowledges a new contributor.
v0.32.15-rc11 featureIntroduced a model metadata cache to improve Ollama's performance by reducing per-request overhead. This release also acknowledges a new contributor.
v0.32.14-rc01 fix1 featureThis release introduces WebP image transcoding for llama-server and improves handling of non-leading system messages in Qwen renderers.
v0.32.141 fix1 featureThis release introduces WebP image transcoding for llama-server and improves handling of non-leading system messages in Qwen renderers.
v0.32.131 featureThis release adds support for developer instructions for the qwen3.8 model. It provides a link to the full changelog for more details.
v0.32.122 featuresThis release introduces support for the Qwen 3.8 27B model, with specific optimizations for Apple Silicon devices to enhance performance for agents and repeated tasks.
v0.32.112 featuresThis release introduces new integrations for Muse Code and DeepSeek Harness, enhancing the launch capabilities of Ollama.
v0.32.10Breaking1 fix1 featureThis release adjusts the default repeat penalty for models to 1.0, improving speculative decoding speed, and introduces faster prefill performance for certain MLX models. A bug in blob verification has also been fixed.
v0.32.10-rc1Breaking1 fix1 featureThis release adjusts the default repeat penalty for models and introduces performance improvements for NVFP4 MLX models. A bug related to blob verification has also been fixed.
v0.32.10-rc0Optimized prefill performance for double-scale nvfp4 models by compiling multiply and cast operations into a single kernel, reducing kernel launches and intermediate materialization.
v0.32.91 fix1 featureIntroduced the Nemotron 3 architecture and fixed a boundary condition in the Muse Glimmer function calling parser.
v0.32.83 featuresOllama now supports the Muse Glimmer model across various platforms, including NVIDIA and AMD, enhancing capabilities for coding agents and personal assistants. MLX engine on Apple Silicon also sees improvements.
Common Errors
Related AI & LLMs Packages
AutoGPT is the vision of accessible AI for everyone, to use and to build on. Our mission is to provide the tools, so that you can focus on what matters.
🦜🔗 The platform for reliable agents.
The most powerful and modular diffusion model GUI, api and backend with a graph/nodes interface.
LLM inference in C/C++
GPT4All: Run Local LLMs on Any Device. Open-source and available for commercial use.
A high-throughput and memory-efficient inference and serving engine for LLMs
Subscribe to Updates
Get notified when new versions are released