Ollama's Gemma 4 now uses multi-token prediction with auto-tuning on MLX/llama.cpp backends, delivering ~90% token speedup on macOS without configuration or output changes.
Summary
Local LLM inference on Apple Silicon just became practical for coding agents and real-time tasks. Zero-config speedup means existing Ollama workflows get faster automatically—no model retraining or prompt engineering required.
Why it matters
Local LLM inference on Apple Silicon just became practical for coding agents and real-time tasks. Zero-config speedup means existing Ollama workflows get faster automatically—no model retraining or prompt engineering required.
Implementation verdict
Drop-in improvement for Ollama users running Gemma 4 on M-series chips. Requires updating to Ollama v0.31.1+; no code changes needed. Worth upgrading immediately if you use this model locally.
Sources
Dev Signal
Get briefs like this in your inbox — free, every weekday.
100+ sources compressed into one 4-minute read. Ranked, cited, implementation-ready.