hex-fa
Found in 1 package: llama-cpp
llama-cpp(5 releases)
b10643BreakingThis release introduces significant enhancements to the Hexagon backend, including multi-NPU device support, a fully asynchronous backend, and improved performance for various operations like ALLREDUCE and matrix multiplications. It also includes numerous bug fixes and optimizations across different Hexagon modules.
b9928This release focuses heavily on internal optimizations and robustness improvements for the hexagon backend, particularly for matrix multiplication and attention kernels, alongside asynchronous queue handling and workpool enhancements.
b9857This release focuses heavily on reworking the hexagon flash attention implementation, bringing significant optimizations and accuracy improvements across various internal components (hex-mm, hex-fa, hmx-fa). Numerous bug fixes and performance enhancements related to tracing, memory alignment, and kernel usage were also implemented.
b8578This release focuses on DMA optimizations within the hexagon backend, primarily fixing performance regressions by introducing a mask cache and disabling unnecessary in-order descriptor processing.
b8204This release focuses heavily on Flash Attention optimizations for hexagon, including DMA, mpyacc, and multi-row improvements, alongside various MatMul updates and bug fixes in hexagon vector operations.
Track Symbol Changes
Use the Change8 MCP server or GitHub Action to get notified when hex-fa changes.
Learn More