MLX engine
Found in 1 package: ollama
ollama(6 releases)
v0.33.3This release introduces support for images and audio with Gemma 4 on the MLX engine, reports cached prompt tokens, and honors GGUF model default parameters. It also includes updates to MLX, MLX-C, and llama.cpp.
v0.32.6-rc0BreakingThis release enhances performance for Qwen3.5 on Apple GPUs, aligns the `/v1/chat/completions` streaming API with OpenAI's format, and includes several TUI improvements. Experimental image generation has been temporarily removed.
v0.32.6BreakingThis release enhances performance for Qwen3.5 on Apple GPUs and standardizes the `/v1/chat/completions` streaming API to match OpenAI's format. It also includes several TUI improvements and bug fixes.
v0.32.3This release addresses several bugs, including stalled model downloads and dropped GLM tool calls. It also introduces expanded GPU support and new features for Laguna 2.1 models.
v0.31.1This release introduces significant performance improvements for Gemma 4 on Apple Silicon by leveraging multi-token prediction (MTP). It also includes updates to the underlying MLX and llama.cpp engines.
v0.17.5This release focuses on stability and performance improvements for Qwen 3.5 models, particularly when running across multiple devices or using the MLX engine, and introduces peak memory reporting.
Track Symbol Changes
Use the Change8 MCP server or GitHub Action to get notified when MLX engine changes.
Learn More