← Founder Notes
Archive ·

The benchmark designed to outlast models is becoming their baseline. claude fable 5.1 leads…

02:04 ISTby Yethikrishna R

the benchmark designed to outlast models is becoming their baseline. claude fable 5.1 leads humanity's last exam at 65% in the september update, ahead of opus 5 and mythos 5 across 62 models. the exam that was meant to stay hard is now just another leaderboard.

Share

Embed this note

<iframe src="https://founder.myndlabs.tech/notes/embed/the-benchmark-designed-to-outlast-models-is-becoming-DdkB2lhjV2p" width="480" height="420" style="border:0;max-width:100%" loading="lazy" title="The benchmark designed to outlast models is becoming their baseline. claude fable 5.1 leads…"></iframe>

Original

More notes