← Founder Notes
Archive

The agent leaderboard just split into score and price, and the gap is 7x for 6.6 points. codex +…

Yethikrishna ROriginal on Threads

the agent leaderboard just split into score and price, and the gap is 7x for 6.6 points. codex + gpt-6.1 sol scores 58.2% on terminal-bench 4.0 for $1.92 per task, against $14.30 for the 64.8% leader.

the top of the board is now a premium for six and a half points.

Context

Morph's leaderboard page, updated October 6, 2026, says Claude Code with Opus 5.5 holds the top Terminal-Bench 4.0 entry at 64.8% and costs $14.30 per task, and that Codex with GPT-6.1 Sol scores 58.2% at $1.92 per task. The page says Sol's cost is about a fifth of GPT-6 Astra's $9.90.

CodingFleet's page, also updated October 6, shows the same scores on the public board: Opus 5.5 at 64.8% with a margin of 3.1, GPT-6.1 Sol and GPT-6 Astra both at 58.2%. It lists run costs of $4.7k for Opus 5.5 and $0.6k for Sol, and says the Sol run used 1.5B tokens against 8.0B for Opus.

How it compares

The 58.2% against 64.8% is a gap of 6.6 points, and $14.30 against $1.92 is about 7.4 times, so the note's 7x and 6.6 points match the Morph figures. The per-task costs come from Morph, a secondary aggregator. CodingFleet's run totals of $4.7k and $0.6k give a ratio of about 7.8 times, which points the same way but is a different measure.

The scores carry margins of about 3 points on the public board. CodingFleet lists Opus 5.5 at 64.8% with a margin of 3.1 and Sol at 58.2% with a margin of 3.1, so the intervals overlap at their edges. 'Six and a half points' is the gap in point estimates.

CodingFleet says three setups publish Terminal-Bench 4.0 numbers and they disagree: the public board with each lab's own harness, Artificial Analysis re-runs inside mini-SWE-agent, and Anthropic's launch runs. The 64.8% and 58.2% are public-board numbers. The post's 'leader' is the public-board leader, and that is not the same ranking under the other setups. Whether tbench.ai itself publishes the per-task cost column was not confirmed here.

Related work

Watch next

  • Open the tbench.ai board for the cost column and trial counts. Check the cost of each run against the harness settings.

Sources

  1. Best AI Coding Agents (October 2026): Scored Leaderboard (Morph, updated October 6, 2026)morphllm.com
  2. Terminal-Bench 4.0 Leaderboard: Opus 5.5 and Sonnet 5.5 Lead, GPT-6.1 Sol Tied at 58.2% (CodingFleet, updated October 6, 2026)codingfleet.com

Provenance

The note above is reproduced unedited from the original post, first published on Threads on 9 October 2026 at 01:18 IST. Sources are the papers and datasets the note draws on.

View the original post
Embed this note
<iframe src="https://founder.myndlabs.tech/notes/embed/the-agent-leaderboard-just-split-into-score-and-DePuKoUFTF7" width="480" height="420" style="border:0;max-width:100%" loading="lazy" title="The agent leaderboard just split into score and price, and the gap is 7x for 6.6 points. codex +…"></iframe>

More notes