← Founder Notes
Archive

The local llm just 3x'd on metal. llama.cpp v0.6.0, out oct 6, ships a 3x speedup on apple silicon…

Yethikrishna ROriginal on Threads

the local llm just 3x'd on metal. llama.cpp v0.6.0, out oct 6, ships a 3x speedup on apple silicon while 4-bit blackwell builds went production-ready.

the small model now runs fastest on a laptop.

Context

AICoder reports that ggml-org released llama.cpp v0.6.0 on Oct 5, 2026, with up to about 3x faster few-row matrix multiplication on Apple GPUs, a new llama_batch_ext API and support for GLM-5.3-Flash. Jason Vs The Noise (Oct 6, 2026) covers the same release and warns that old sessions do not carry over.

A llama.cpp pull request adds a native 4-bit float quant for Blackwell that gives about 40% faster prompt processing, and other pull requests add native NVFP4 support.

How it compares

The release is dated Oct 5, 2026 in the AICoder report, and the 3x is for few-row matrix multiplication on Apple GPUs, not a general 3x speedup. The Blackwell 4-bit work is in pull requests with a 40% prompt processing gain, and that those builds went production ready was not seen in the sources read, so unsupported here, not refuted. 'The small model now runs fastest on a laptop' is the author's opinion.

Related work

Watch next

  • Read the v0.6.0 release notes for the Apple GPU benchmark numbers.

Sources

  1. AICoder: llama.cpp v0.6.0 ships GLM-5.3-Flash and Metal speedupsaicoder.com
  2. Jason Vs The Noise: llama.cpp 0.6.0 speeds up Apple GPUsjasonvsthenoise.com
  3. GitHub: llama.cpp CUDA native 4-bit float quant for Blackwellgithub.com

Provenance

The note above is reproduced unedited from the original post, first published on Threads on 11 October 2026 at 21:50 IST. Sources are the papers and datasets the note draws on.

View the original post
Embed this note
<iframe src="https://founder.myndlabs.tech/notes/embed/the-local-llm-just-3x-d-on-metal-DeXEunUjBmJ" width="480" height="420" style="border:0;max-width:100%" loading="lazy" title="The local llm just 3x'd on metal. llama.cpp v0.6.0, out oct 6, ships a 3x speedup on apple silicon…"></iframe>

More notes