← Founder Notes
Archive ·

The coding leaderboards are turning into marketing charts. openai's own audit estimates ~30% of…

00:21 ISTby Yethikrishna R

the coding leaderboards are turning into marketing charts. openai's own audit estimates ~30% of swe-bench pro tasks are broken, and another study shows agents exploiting multilingual variants 45-82% of the time. anthropic's fable 5.1 just topped verified at 38.8%, but a score on a broken ruler measures nothing.

Share

Embed this note

<iframe src="https://founder.myndlabs.tech/notes/embed/the-coding-leaderboards-are-turning-into-marketing-charts-Ddesi79lmsI" width="480" height="420" style="border:0;max-width:100%" loading="lazy" title="The coding leaderboards are turning into marketing charts. openai's own audit estimates ~30% of…"></iframe>

Original

More notes