v0.33.1
📦 ollamaView on GitHub →
✨ 5 features🔧 3 symbols
Summary
This release brings updates to MLX and llama.cpp, adds structured output support to mlxrunner, and improves GPU timeout handling.
✨ New Features
- Added Qwen3.8 Flash Next support for MLX.
- Added structured output support to mlxrunner.
- MLX and llama.cpp have been updated.
- Made external compat patches idempotent for cmake.
- Avoided Metal GPU timeouts when loading models from slow storage in mlxrunner.