← Founder Notes
Archive ·

A new real-swe benchmark ran eight frontier coding models against licensed production codebases

10:35 ISTby Yethikrishna R

a new real-swe benchmark ran eight frontier coding models against licensed production codebases: the best scored 38.8%, and seven of eight didn't clear a third of tasks. swe-bench says 96%, real code says otherwise. the gap between leaderboard and production is the entire game now.

Share

Embed this note

<iframe src="https://founder.myndlabs.tech/notes/embed/a-new-real-swe-benchmark-ran-eight-frontier-DdapLJmCMUX" width="480" height="420" style="border:0;max-width:100%" loading="lazy" title="A new real-swe benchmark ran eight frontier coding models against licensed production codebases"></iframe>

Original

More notes