llama-server
Found in 3 packages: llama-cpp, unsloth, ollama
llama-cpp(13 releases)
b10649Introduced benchmark-only synthetic speculative acceptance options for llama-server and llama-cli. This release also provides various pre-compiled binaries for different operating systems and hardware configurations.
b10094This release introduces automatic speculative type inference from draft repository sidecars, simplifying the `llama-server` command. It also provides numerous pre-compiled binaries for various platforms and hardware accelerators.
b8748This release fixes an issue in llama-server where the --alias flag conflicted with model presets, and provides updated pre-compiled binaries for broad platform compatibility.
b8696This release addresses a bug in llama-server where model parameters were not being propagated correctly. It also provides numerous pre-compiled binaries for different operating systems and hardware setups.
b8668This release updates the llama-server startup logging to remove redundancy and include build/commit information. It also provides numerous pre-compiled binaries for various operating systems and hardware configurations.
b8560This release enhances the server component by adding control over SO_REUSEPORT socket options and includes a fix for Windows compatibility.
b7488This release fixes a logging bug in llama-server's speculative decoding where the number of drafted tokens was incorrectly reported as zero.
b7487This release introduces model autoloading on startup and preset-only options for the llama.cpp server.
b7481BreakingThis release improves llama-server error handling for context overflows and announces a transition in Linux package distribution formats from .zip to .tar.gz.
b7352BreakingThis release introduces recursive model loading and an INI-based preset system for llama-server, while transitioning Linux release artifacts to .tar.gz format.
b7248BreakingThis release fixes a duplicate HTTP header bug in llama-server's multi-model mode and announces a transition for Linux release artifacts from .zip to .tar.gz.
b7243BreakingThis release introduces a new --media-path flag for the server to handle local media files and announces a change in the Linux distribution format from .zip to .tar.gz.
b7229BreakingThis release introduces explicit execution path handling for the server and announces a transition in Linux package formats from ZIP to TAR.GZ.
unsloth(7 releases)
v0.1.803-betaThis bug fix release introduces significant improvements including experimental auto compaction for longer chats, preview of remote & LAN access, and enhanced chat/hardware/API functionalities. It also includes over 170 PRs addressing bugs, reliability, and performance.
v0.1.800-betaBreakingUnsloth now supports running Qwen3.8 models locally with reduced RAM requirements and offers significant performance improvements, including faster inference for GGUFs and MiniMax-H3. New features enhance chat capabilities, tool integration, and hardware compatibility.
v0.1.47-betaThis release introduces significant enhancements to Unsloth Studio, including 3x longer context lengths, support for GLM 5.2 GGUFs, and new parallel operational capabilities. Key new features focus on chat experience (forking, queueing) and secure cloud access via Cloudflare.
v0.1.461-betaThis release addresses several issues related to local GGUF vision handling in llama-server, including maintaining vision, adjusting log levels, and improving companion file discovery.
v0.1.39-betaThis release introduces powerful local LLM integration via a self-hosted API endpoint supporting advanced tooling, code execution, and web search. Numerous bug fixes were applied across the Studio UI, training stability (DPO hangs), and installation scripts.
v0.1.38-betaThis release introduces powerful local LLM API serving via `llama-server` with features like self-healing tool calling, code execution, and web search. Numerous stability and usability improvements were made across the Unsloth Studio interface and training pipeline.
v0.1.2-betaThis release focuses heavily on Unsloth Studio improvements, including major performance gains via pre-compiled binaries, enhanced tool calling, and robust fixes for Windows and Colab environments. Installation sizes have been significantly reduced, and UX features like persistent settings and multi-file uploads have been added.
ollama(1 releases)
Track Symbol Changes
Use the Change8 MCP server or GitHub Action to get notified when llama-server changes.
Learn More