hex-ops
Found in 1 package: llama-cpp
llama-cpp(4 releases)
b10643BreakingThis release introduces significant enhancements to the Hexagon backend, including multi-NPU device support, a fully asynchronous backend, and improved performance for various operations like ALLREDUCE and matrix multiplications. It also includes numerous bug fixes and optimizations across different Hexagon modules.
b9483This release includes fixes for the hexagon profiler output and updates to the profiling script to handle the tot.usec column.
b9470This release focuses heavily on cleanup and performance optimizations across Hexagon (HEX) and HMX backends for matrix multiplication (MUL_MAT, MUL_MAT_ID), Flash Attention, and GDN, introducing initial F32 matmul support and fixing several fusion and stride bugs.
b8754This release significantly improves hexagon performance through op request batching, buffer management rewrite, and explicit L2 cache control. It also removes the deprecated GGML_HEXAGON_EXPERIMENTAL environment variable.
Track Symbol Changes
Use the Change8 MCP server or GitHub Action to get notified when hex-ops changes.
Learn More