Change8

v0.33.1

📦 ollamaView on GitHub →
5 features🔧 3 symbols

Summary

This release brings updates to MLX and llama.cpp, adds structured output support to mlxrunner, and improves GPU timeout handling.

✨ New Features

  • Added Qwen3.8 Flash Next support for MLX.
  • Added structured output support to mlxrunner.
  • MLX and llama.cpp have been updated.
  • Made external compat patches idempotent for cmake.
  • Avoided Metal GPU timeouts when loading models from slow storage in mlxrunner.

Affected Symbols