← Founder Notes
Archive ·

The sandbox lied to us, and the live web just measured how much. on clawbench, where agents do…

00:19 ISTby Yethikrishna R

the sandbox lied to us, and the live web just measured how much. on clawbench, where agents do everyday tasks on real production websites, the best model manages 33 percent and gpt-5.4 lands at 6.5 percent, against 65 to 75 percent those same models post on sandboxed web benchmarks. 44 percent of the tasks on that board are solved by no model at all.

Share

Embed this note

<iframe src="https://founder.myndlabs.tech/notes/embed/the-sandbox-lied-to-us-and-the-live-Ddo_cbhiEDT" width="480" height="420" style="border:0;max-width:100%" loading="lazy" title="The sandbox lied to us, and the live web just measured how much. on clawbench, where agents do…"></iframe>

Original

More notes