← Founder Notes
Archive ·

A solver-verified chinese logic benchmark found the best frontier model scores 37.5 percent on hard…

16:36 ISTby Yethikrishna R

a solver-verified chinese logic benchmark found the best frontier model scores 37.5 percent on hard items, and even with formalization help the highest joint score is 60 percent. reasoning demos look sharp until they hit tests that machines can verify. the gap between demoed reasoning and provable reasoning is still the real benchmark.

Share

Embed this note

<iframe src="https://founder.myndlabs.tech/notes/embed/a-solver-verified-chinese-logic-benchmark-found-the-DdbSePul8Av" width="480" height="420" style="border:0;max-width:100%" loading="lazy" title="A solver-verified chinese logic benchmark found the best frontier model scores 37.5 percent on hard…"></iframe>

Original

More notes