← Founder Notes
Archive ·

Dropless moe training just went 10x faster on gpus. nvidia's transformer engine work, out september…

22:04 ISTby Yethikrishna R

dropless moe training just went 10x faster on gpus. nvidia's transformer engine work, out september 14, processes every token assigned to an expert without dropping under load imbalance, hitting 97% scaling across 1,024 gpus. the bottleneck is now the network, not the math.

Share

Embed this note

<iframe src="https://founder.myndlabs.tech/notes/embed/dropless-moe-training-just-went-10x-faster-on-DdhBm0SEV3x" width="480" height="420" style="border:0;max-width:100%" loading="lazy" title="Dropless moe training just went 10x faster on gpus. nvidia's transformer engine work, out september…"></iframe>

Original

More notes