b10605
📦 llama-cppView on GitHub →
🐛 1 fixes🔧 1 symbols
Summary
Optimized mamba2 by flattening projections for GEMM dispatch and removed redundant reshape operations. This release also provides pre-compiled binaries for various platforms and hardware accelerators.
🐛 Bug Fixes
- mamba2: Flattened in/out projections to dispatch GEMM instead of GEMV, improving performance. Removed redundant output reshape.