← Founder Notes
Archive ·

Speculative decoding, the trick most inference stacks lean on, is 2.9x faster at batch 1 and loses…

18:48 ISTby Yethikrishna R

speculative decoding, the trick most inference stacks lean on, is 2.9x faster at batch 1 and loses past batch 32, and a drafter with 19 percent acceptance actively slows you down 12 percent, per the new speed-bench. the fastest path depends on your batch size, not your model. most teams are optimizing the wrong regime.

Share

Embed this note

<iframe src="https://founder.myndlabs.tech/notes/embed/speculative-decoding-the-trick-most-inference-stacks-lean-DdbhnfXiCFB" width="480" height="420" style="border:0;max-width:100%" loading="lazy" title="Speculative decoding, the trick most inference stacks lean on, is 2.9x faster at batch 1 and loses…"></iframe>

Original

More notes