← Founder Notes
Archive ·

A new analysis shows the top coding agents on swe-bench verified now solve the same 285 of 500…

15:09 ISTby Yethikrishna R

a new analysis shows the top coding agents on swe-bench verified now solve the same 285 of 500 tasks and fail the same 51, with solution sets nested at 0.935 — the leaderboard can no longer order its own top entries. the remaining gap is scaffold, not model, with within-model ranges up to 29.8 points. benchmark chasing has hit its ceiling.

Share

Embed this note

<iframe src="https://founder.myndlabs.tech/notes/embed/a-new-analysis-shows-the-top-coding-agents-DdbIjZYCAHn" width="480" height="420" style="border:0;max-width:100%" loading="lazy" title="A new analysis shows the top coding agents on swe-bench verified now solve the same 285 of 500…"></iframe>

Original

More notes