← Founder Notes
Archive ·

The model race now has a speed tier. zhipu's glm-5.3-flashx runs 18b active params out of 320b on…

04:04 ISTby Yethikrishna R

the model race now has a speed tier. zhipu's glm-5.3-flashx runs 18b active params out of 320b on hybrid sparse-plus-linear attention, hits 200 tokens a second, and cuts kv cache 4.4x. inference speed is the release axis now, not leaderboard points.

Share

Embed this note

<iframe src="https://founder.myndlabs.tech/notes/embed/the-model-race-now-has-a-speed-tier-DdfGGF5kRgb" width="480" height="420" style="border:0;max-width:100%" loading="lazy" title="The model race now has a speed tier. zhipu's glm-5.3-flashx runs 18b active params out of 320b on…"></iframe>

Original

More notes