Change8

b10605

📦 llama-cppView on GitHub →
🐛 1 fixes🔧 1 symbols

Summary

Optimized mamba2 by flattening projections for GEMM dispatch and removed redundant reshape operations. This release also provides pre-compiled binaries for various platforms and hardware accelerators.

🐛 Bug Fixes

  • mamba2: Flattened in/out projections to dispatch GEMM instead of GEMV, improving performance. Removed redundant output reshape.

Affected Symbols