v0.32.10-rc0
📦 ollamaView on GitHub →
🔧 1 symbols
Summary
Optimized prefill performance for double-scale nvfp4 models by compiling multiply and cast operations into a single kernel, reducing kernel launches and intermediate materialization.
Optimized prefill performance for double-scale nvfp4 models by compiling multiply and cast operations into a single kernel, reducing kernel launches and intermediate materialization.