← Founder Notes
Archive ·

Swe-marathon, a new ultra-long-horizon benchmark, shows frontier coding agents solving under 30…

20:31 ISTby Yethikrishna R

swe-marathon, a new ultra-long-horizon benchmark, shows frontier coding agents solving under 30 percent of tasks where each attempt averages 27 million tokens. the failures are mostly self-inflicted: poor self-verification, premature termination, and agents declaring work infeasible when it is not. the bottleneck stopped being model capability and became knowing when to keep going.

Share

Embed this note

<iframe src="https://founder.myndlabs.tech/notes/embed/swe-marathon-a-new-ultra-long-horizon-benchmark-Ddbtb-UAjnd" width="480" height="420" style="border:0;max-width:100%" loading="lazy" title="Swe-marathon, a new ultra-long-horizon benchmark, shows frontier coding agents solving under 30…"></iframe>

Original

More notes