← Founder Notes
Archive

The first bottleneck agents hit inside anthropic was the test runner, not the model. ci jobs grew…

Yethikrishna ROriginal on Threads

the first bottleneck agents hit inside anthropic was the test runner, not the model. ci jobs grew 25x in six months as agents per engineer rose while code output hit 8x per quarter, and the fix was test impact analysis so each change runs only the tests it can touch.

the constraint moved from writing code to proving it.

Context

Anthropic's post of September 14, 2026 says its CI job volume rose 25x over six months. It says Anthropic engineers on average ship 8x as much code per quarter as they did from 2021 to 2025, that Claude authors 80% of that code, and that the amount of tests grew 10x with only a nominal increase in engineers. Not every test runs on every pull request, which is why the job count is not simply a product of those figures.

The post says the load threatened to overload the test impact analysis service several times. Three quick fixes lasted 70 days, 29 days and less than a day, and the team then redesigned the service. It says CI jobs increase as the average number of agents per engineer rises and that writing code is no longer the bottleneck.

How it compares

The 25x CI growth over six months, the 8x code figure and the test impact analysis fix match the post. The post's 8x is relative to a 2021 to 2025 average per quarter. The note's 'hit 8x per quarter' compresses that and should be read as 8x the earlier rate.

The note says the first bottleneck agents hit inside Anthropic was the test runner. The post describes CI and the test impact analysis service as what came under strain, but a ranking of 'first' against other bottlenecks was not found in the page read. That is unsupported, not refuted. The post itself says the model was not the bottleneck for coding.

'The constraint moved from writing code to proving it' is the author's phrasing. The post says writing code is no longer the bottleneck, which is consistent with it.

Related work

Watch next

  • Read how the redesigned service selects tests and how it handles flaky tests. Look for similar scaling reports from other teams running agents.

Sources

  1. Agentic coding is straining CI. Here's how we scaled test impact analysis at Anthropic (Claude blog, September 14, 2026)claude.com

Provenance

The note above is reproduced unedited from the original post, first published on Threads on 9 October 2026 at 00:46 IST. Sources are the papers and datasets the note draws on.

View the original post
Embed this note
<iframe src="https://founder.myndlabs.tech/notes/embed/the-first-bottleneck-agents-hit-inside-anthropic-was-DePqigQlRf5" width="480" height="420" style="border:0;max-width:100%" loading="lazy" title="The first bottleneck agents hit inside anthropic was the test runner, not the model. ci jobs grew…"></iframe>

More notes