← Founder Notes
Archive

The agent exam just graded everyone below 35 percent. argo-bench, out oct 2, scored claude opus 5.5…

Yethikrishna ROriginal on Threads

the agent exam just graded everyone below 35 percent. argo-bench, out oct 2, scored claude opus 5.5 as the top agent and it still cleared only 34.8 percent of 210 real tasks, averaging 59.5 points.

the passing bar now sits above every model.

Context

Verified dates: the Argo-Bench paper, 'Evaluating Data Agents on Enterprise-Scale Workflows' (arXiv 2610.02122), was published Oct 1, 2026 by TextQL researchers, and AI.info covered it on Oct 2. It is a 210-task benchmark built on a simulated New York City food delivery platform, with ground truth hidden in an Oracle E-Business Suite warehouse of 235 tables and 7.5 billion rows.

TextQL says the strongest model, Claude Opus 5.5, fully solves 34.8 percent of tasks, scoring 95 or more, and averages 59.5 points; eight of the thirteen models average under 30.

How it compares

The 34.8 percent, the 59.5 average, the 210 tasks and Claude Opus 5.5 at the top match TextQL's page and the paper coverage. The post says the benchmark was out Oct 2; the paper is dated Oct 1 and Oct 2 is the AI.info news date. The sources differ on model count: TextQL says thirteen models, AI.info says 14 frontier and open-weight models. 'Solved' means a score of 95 or more, so the 'passing bar' is the paper's threshold. 'The passing bar now sits above every model' is the author's opinion.

Related work

Watch next

  • Check the leaderboard for the v1.1 run and any new models added after Oct 2.

Sources

  1. TextQL: Decision-Bench (Argo-Bench), evaluating data agents on enterprise-scale workflowstextql.com
  2. arXiv: Argo-Bench, evaluating data agents on enterprise-scale workflowsarxiv.org
  3. AI.info: Argo-Bench finds data agents miss the decision, not just the SQLai.info

Provenance

The note above is reproduced unedited from the original post, first published on Threads on 11 October 2026 at 16:41 IST. Sources are the papers and datasets the note draws on.

View the original post
Embed this note
<iframe src="https://founder.myndlabs.tech/notes/embed/the-agent-exam-just-graded-everyone-below-35-DeWhaSZjNj1" width="480" height="420" style="border:0;max-width:100%" loading="lazy" title="The agent exam just graded everyone below 35 percent. argo-bench, out oct 2, scored claude opus 5.5…"></iframe>

More notes