The ai coding race moved to review while everyone was still arguing about generation. alibaba's…
the ai coding race moved to review while everyone was still arguing about generation. alibaba's open-code-review hit 44k stars in a month, github launched reviewbench on oct 7, and the tool that reviews alibaba's own code is a go cli anyone can run.
the moat in coding ai stopped being the model.
Context
The alibaba/open-code-review repository describes an AI code review CLI that reads Git diffs and has a model agent produce line-level comments, with a built-in multi-language ruleset. When read, the page showed 43,914 stars, 3,168 forks, an Apache 2.0 license and a creation date of May 18, 2026. Its README says the tool originated as Alibaba's internal code review assistant and has served tens of thousands of developers.
The README's own benchmark uses 200 pull requests from 50 repositories in 10 languages with 1,505 annotated issues. It claims higher precision and F1 than a general agent (Claude Code) with the same underlying model at about one ninth of the tokens, and says its recall is lower. These are the project's own numbers.
GitHub's blog post of October 5, 2026 introduces ReviewBench: 219 public pull requests from 187 repositories in 19 languages, shaped by an analysis of 103.9 million GitHub pull requests. It reports precision, recall and F1 in a grounded form and an augmented form, publishes the dataset and a leaderboard, and says independent senior engineers who re-labelled the findings agreed with the benchmark 96.6% of the time.
The note says open-code-review 'hit 44k stars in a month'. The page shows 43,914 stars and a creation date of May 18, 2026, which is more than four months before the note. How fast the stars arrived is not in the sources read, so 'in a month' is unsupported, not refuted.
The note dates the ReviewBench launch to October 7. GitHub's own post is dated October 5, and heise's coverage appeared October 7, so the date likely comes from press coverage.
The README lists nine languages for the repository, including Go, alongside TypeScript, Kotlin and others, and calls it a CLI. 'A Go cli' is partly supported and the sources read do not say the whole tool is written in Go. That it reviews Alibaba's own code rests on the README's statement that the tool began as an internal assistant. The benchmark numbers in that README are vendor numbers.
That the moat in coding AI 'stopped being the model' is the author's opinion. The README's comparison holds the model fixed and varies the pipeline, which is consistent with that view but is the project's own result.
Related work
- alibaba/open-code-review (GitHub) ↗Source for the stars, license, language list and the project's own benchmark claims.
- ReviewBench: An open benchmark for AI code review (The GitHub Blog) ↗Primary source for the ReviewBench design and its October 5 date.
- GitHub ReviewBench: Benchmark for AI-assisted Code Reviews (heise online) ↗Press coverage of the ReviewBench release.
Watch next
- Check whether open-code-review appears on the ReviewBench leaderboard. Find the star history to see how fast the repository grew.
Sources
Provenance
The note above is reproduced unedited from the original post, first published on Threads on 8 October 2026 at 22:04 IST. Sources are the papers and datasets the note draws on.
View the original post ↗Embed this note
More notes
The air is now being asked to keep its own ledger
the air is now being asked to keep its own ledger: ecmwf’s aifs compo becomes the first ai model to forecast atmospheric composition globally every three hours, cleanair simulates 365 days of pm2.5 over china in ten seconds, and a unified framework maps six pollutants at one kilometer across the whole country. the air now files its own composition report.
read the note →The current is now being asked to draw its own map
the current is now being asked to draw its own map: china’s langya 2.0 predicts six ocean phenomena including internal waves and mesoscale eddies, a deep net called wenhai resolves eddies globally with air sea flux formulas built in, and scripps infers surface currents from the way temperature patterns deform in satellite images. the ocean now files its own circulation report.
read the note →The soil is now being asked to report its own carbon
the soil is now being asked to report its own carbon: a nix color sensor paired with generative data augmentation predicts soil organic carbon without a lab, random forest drives 74 percent of soil health mapping studies, and sentinel 2 tracks five year carbon change across france and italy from 922 samples. the dirt now files its own carbon account.
read the note →