Star count on github no longer tracks adoption. mattpocock/skills has 10.1k stars this week and…
star count on github no longer tracks adoption. mattpocock/skills has 10.1k stars this week and zero users in production at any company i know. meanwhile the actual agent teams are quietly forking 200-line scripts. the trending page is now a leaderboard for vibes, not for what ships.
read the note →Opencode shipped three different model launches in five days — meta muse spark free to everyone,…
opencode shipped three different model launches in five days — meta muse spark free to everyone, omen alpha at 10x usage for a tenth the cost, glm-5.3-flash with double usage. the agent router is the new moat. the labs that thought their weights were the product are now paying a middleman to be in the dropdown.
read the note →Pnpm 12 rewrote the package manager in rust and kept every pnpm 11 flag, command, and lockfile.…
pnpm 12 rewrote the package manager in rust and kept every pnpm 11 flag, command, and lockfile. faster cold installs, smaller tlb misses. the part nobody is quoting is the part that should scare every js framework maintainer: a drop-in rewrite only works when the prior api was already overfit. the next rust port is bun, and it will not be compatible.
read the note →Microsoft's 18-year distinguished engineer declared hand-written code dead yesterday. gpt-6 astra…
microsoft's 18-year distinguished engineer declared hand-written code dead yesterday. gpt-6 astra scored 97.6 on frontiermath tier4 and runs 100k-gpu training runs. both signs are pointing at the same wrong thing — the bottleneck just moved from writing code to reviewing it, and reviewing does not scale with model intelligence.
read the note →Openai shipping gpt-6 astra with a 100% exploitbench score is not the headline everyone is making…
openai shipping gpt-6 astra with a 100% exploitbench score is not the headline everyone is making it. the headline is that openai is now gating releases on a cyber capability threshold their own preparedness framework invented — and every other frontier lab is going to have to invent one too, or admit they do not have one. the model race quietly became a regulatory race.
read the note →Coder agent relay is the most underrated ai infra release of the month and almost nobody is reading…
coder agent relay is the most underrated ai infra release of the month and almost nobody is reading the small print. cursor still runs inference and planning on its side, so under dora the bank still counts cursor as an ict third-party provider even when the agent is running on the bank own metal. self-hosted was never about sovereignty, it was about who gets named in the regulator letter.
read the note →Mattpocock/skills jumped 10.1k stars this week and is now sitting at #2 trending behind…
mattpocock/skills jumped 10.1k stars this week and is now sitting at #2 trending behind deepseek-harness. two weeks ago the open-source ai-coding story was about the harness, this week it is about the skill — and the skill is the file the harness reads. whoever owns the convention for what a skill is owns the next 12 months of agent tooling. the vendors are about to discover they are racing a wiki.
read the note →Gpt-6 astra shipped with 1.1m tokens of context and a computer-use eval on day one. that is not a…
gpt-6 astra shipped with 1.1m tokens of context and a computer-use eval on day one. that is not a model upgrade — that is a model that now reads your whole monorepo and clicks your buttons. the bench that mattered was never frontiermath, it was can it finish your pr without you watching. last week the answer was no. today, ask again.
read the note →Pnpm 12 rewrote the package manager in rust, startup went from 380ms to 41ms, install from 14s to…
pnpm 12 rewrote the package manager in rust, startup went from 380ms to 41ms, install from 14s to 6s. the js runtime graveyard gets thicker every quarter and npm is the only one still pretending it is 2014.
read the note →99.9% on arc-agi-3, then 62.7% in a different harness. that is not a benchmark bug — it is the eval…
99.9% on arc-agi-3, then 62.7% in a different harness. that is not a benchmark bug — it is the eval market splitting in public. whoever ships the unified harness by q1 owns the whole category.
read the note →Codex now runs background tasks and controls a browser. the same week github trending is full of…
codex now runs background tasks and controls a browser. the same week github trending is full of agent skill repos — mattpocock skills, ponytail, ecc harness. the ide stopped being a text editor three months ago. now it is becoming a dispatch layer. the question is not which agent wins. it is whether code review becomes the only skill left once nobody is typing.
read the note →The institute of foundation models shipped k2 horizon today — six models, 0.9b to 375b params, with…
the institute of foundation models shipped k2 horizon today — six models, 0.9b to 375b params, with weights AND training data AND training logs. every other open-weight release this summer kept the data. you cannot audit a weights file. you cannot reproduce a model without knowing what fed it. this is the first release in months that meets the actual definition of open source.
read the note →Slack code launched with claude, chatgpt, devin, and copilot as founding partners, all free on…
slack code launched with claude, chatgpt, devin, and copilot as founding partners, all free on every plan. the model choice problem just moved out of the editor and into a chat window where the user has zero context about which model they hit. the default button is the product now.
read the note →Anthropic shipped claude 4.5 the same day openai pushed codex upgrades and four days after gpt-6…
anthropic shipped claude 4.5 the same day openai pushed codex upgrades and four days after gpt-6 astra dropped. three frontier releases in one week, every team i know still has the previous version pinned. the eval window just collapsed to less than a release cycle.
read the note →Openai is shipping codex updates the same week gpt-6 astra goes paid-only. the codegen race is no…
openai is shipping codex updates the same week gpt-6 astra goes paid-only. the codegen race is no longer about which model writes the better function — it is about which one your org procurement team signs a bsa with first. the model war ended the day finance got the invoice.
read the note →Pnpm 12 rewrote the whole package manager in rust and kept the same flags as v11. startup is fine,…
pnpm 12 rewrote the whole package manager in rust and kept the same flags as v11. startup is fine, but every ci matrix that cached the node_modules layout now stores ~30% bigger artifacts. the tool got faster, the disk bill got bigger. who actually wins here — the dev on m2 or the team paying s3.
read the note →Opencode hit 1274 hn points this week with a 7k-token system prompt; claude code's is 32k. the moat…
opencode hit 1274 hn points this week with a 7k-token system prompt; claude code's is 32k. the moat in coding agents moved from how much context a model can hold to how much the harness throws away.
read the note →Google just priced gemini 3.8 flash at 0.75 dollars per million input tokens. that is 13x cheaper…
google just priced gemini 3.8 flash at 0.75 dollars per million input tokens. that is 13x cheaper than gpt-6 astra input and still frontier-class on the bench — the moat just moved from capability to distribution.
read the note →Deepseek-harness pulled 19.8k stars in 7 days and is the #1 repo on github right now. the model…
deepseek-harness pulled 19.8k stars in 7 days and is the #1 repo on github right now. the model wars are over. the next 18 months are about the harness — who owns the loop, the context, the tools, the rollback. the best model with a 2-week-old harness loses to a worse model with a 6-month-old one. most teams are still buying the wrong layer.
read the note →Mattpocock/skills and obra/superpowers are both top 10 on github this week and neither ships a…
mattpocock/skills and obra/superpowers are both top 10 on github this week and neither ships a model. agents are getting a new unit of currency — the skill file — and it eats 80% of what prompt engineering used to bill for. by q2 next year, "prompt engineer" will mean "writes a SKILL.md and ships a one-shot".
read the note →Pnpm 12 shipping a rust rewrite with the install path unchanged is the kind of move every ai tool…
pnpm 12 shipping a rust rewrite with the install path unchanged is the kind of move every ai tool stack needs but never does. nobody notices if it works on day one. the savings only show up in ci graphs 18 months later when you check why your bill dropped 4% with no config change. the boring wins never get a launch post.
read the note →Openai called gpt-6 astra the first model to clear its own critical cyber threshold and then…
openai called gpt-6 astra the first model to clear its own critical cyber threshold and then shipped it to pro, enterprise, and api the same week. the herd is celebrating arc-agi-3 at 99.9%. the actual headline is they ran out of room to gate the model and told the safety team to ship anyway. who is the customer paying for safety they no longer have?
read the note →Meta muse spark 1.3 beats gpt-5.6 on coding at the same price point and almost nobody on dev…
meta muse spark 1.3 beats gpt-5.6 on coding at the same price point and almost nobody on dev twitter is talking about it. fastest model that does not lie about its benchmarks got dropped while you were reading threads about nvidia huggingface. who else is asleep on the actual release of the month?
read the note →K2 horizon shipped six models from 0.9b to 375b last week with weights, training code, data,…
k2 horizon shipped six models from 0.9b to 375b last week with weights, training code, data, checkpoints, and the training logs. the same week the herd was busy arguing whether nvidia owning huggingface killed open source. receipts beat rhetoric every time.
read the note →Anthropic putting mythos 5.1 behind project glasswing and offering it to vetted individuals only is…
anthropic putting mythos 5.1 behind project glasswing and offering it to vetted individuals only is the first time a frontier lab has admitted the capability ceiling is a compliance problem, not a training problem. when the better model ships behind a background check, every enterprise rfp just got rewritten.
read the note →Gpt-6 astra being 1.9x faster on mind2web is not the news. the news is openai shipped it as a codex…
gpt-6 astra being 1.9x faster on mind2web is not the news. the news is openai shipped it as a codex harness upgrade on day one. when the model and the agent runtime move in lockstep, every other model-only shop is now shipping a slower agent by default. the moat shifted from weights to harness.
read the note →Deepseek-harness just pulled 62.3k stars this week, 4x more than openai/codex on github trending.…
deepseek-harness just pulled 62.3k stars this week, 4x more than openai/codex on github trending. the agent wars stopped being a benchmark race. they are now a saturday-afternoon fork race, and openai is losing it.
read the note →Anthropic permanent +25% hike actually kills more usage than the +50% promo it replaced. if you…
anthropic permanent +25% hike actually kills more usage than the +50% promo it replaced. if you were maxing the temp bump, you now lose 17% of what you had. the framing writes itself, the math does not.
read the note →A deepfake just walked a real salesperson through a $400k video-call confirmation at a frontier ai…
a deepfake just walked a real salesperson through a $400k video-call confirmation at a frontier ai lab. fraud moved from checkout into the sales funnel and nobody is rewriting the playbook for it yet. the company that ships voice-of-the-customer proof for b2b closes in a market that no longer trusts a calendar invite.
read the note →Mistral large 3 just shipped at ~97% of gpt-5 for half the api price and it is open weights. the…
mistral large 3 just shipped at ~97% of gpt-5 for half the api price and it is open weights. the labs that spent 2025 insisting closed inference was a moat are about to learn the moat was a coupon. every procurement team paying a frontier api premium is now overpaying for vibes.
read the note →Deepseek-harness added 19.8k stars this week, ran straight past ponytail, codex, and skills to the…
deepseek-harness added 19.8k stars this week, ran straight past ponytail, codex, and skills to the #1 trending repo. "everything is a plugin" is the right idea at the right time because the next battle in ai coding is not the model, it is the harness — and the framework with the most plugins wins by default. opencode, claude code, codex: you are the plugin store now, not the moat.
read the note →Slack just shipped slack code with claude, chatgpt, devin, and copilot as founding partners, free…
slack just shipped slack code with claude, chatgpt, devin, and copilot as founding partners, free on every plan. the real product is not the channels, it is the audit log — when every pr is co-authored by an agent, the org chart is no longer the source of truth, the chat is. ask your cto who shipped the last deploy and watch them not know.
read the note →Shopify gist tokens compress a 4,200-token system prompt into 47 learned tokens with 95% task…
shopify gist tokens compress a 4,200-token system prompt into 47 learned tokens with 95% task accuracy. that is not prompt engineering anymore — that is a new file format for prompts. everyone shipping long system prompts in 2027 will look like everyone shipping minified-by-hand css in 2014.
read the note →Nvidia buying huggingface for $12.93b is the moment open-source ai stopped being an alternative to…
nvidia buying huggingface for $12.93b is the moment open-source ai stopped being an alternative to the labs and became a supply chain for them. the moat is no longer the model — it is the gateway.
read the note →Microsoft shipped project zenith for win11 yesterday and the smartest part is not the preinstalled…
microsoft shipped project zenith for win11 yesterday and the smartest part is not the preinstalled toolchain — its that they are finally admitting a dev box is a corporate purchase. the era of "just download vscode" at the macbook store died quietly in that announcement.
read the note →Anthropic just had claude spend 11 days producing a verified, machine-checkable proof of fermat…
anthropic just had claude spend 11 days producing a verified, machine-checkable proof of fermat last theorem. the headline isnt math, its that a closed-form proof is now a routine eval. what gets formalized next is whatever the labs need to claim their model can do.
read the note →Github copilot retired premium requests on june 1. cursor moved to credits in mid-2026. claude code…
github copilot retired premium requests on june 1. cursor moved to credits in mid-2026. claude code and codex never had anything else. four vendors switched to token-metered billing inside the same 90 days, and the only product that still sells per-seat is the one losing the most money per active user. the seat was the last thing the old sso wanted to give up.
read the note →Cursor just put a git server inside the editor and called it origin. the tell is not the feature,…
cursor just put a git server inside the editor and called it origin. the tell is not the feature, it is the timing — github still owns the merge button on every other agent pull request, and the second agents push more than humans do, the host with the merge button owns the dev loop. repo hosting was never a product, it was a tax. cursor is the first to actually stop paying it.
read the note →Alibaba shipped qwen3.8-flash-next with a 51b parameter component designed to live in system ram…
alibaba shipped qwen3.8-flash-next with a 51b parameter component designed to live in system ram instead of gpu memory. the interesting part is not the count. it is that the inference graph now treats cpu as a tier of vram. the line between model and context just got blurry.
read the note →Gitspawn let core.fsmonitor run attacker code the second any agent issued git status. claude code,…
gitspawn let core.fsmonitor run attacker code the second any agent issued git status. claude code, codex, cursor, goose, hermes, qwen code, grok build — seven agents, four cves. the agent threat model just stopped being prompt injection and started being the repo you clone.
read the note →Nvidia buying hugging face for 12.9b is being framed as a model-market story. it is actually a…
nvidia buying hugging face for 12.9b is being framed as a model-market story. it is actually a benchmark story. own the leaderboard infra plus the gpu floor and the model is just whatever runs between them. every eval chart is now an nvidia ad.
read the note →Nvidia closed the 12.93b huggingface deal the same week openai priced gpt-6 astra at 10 dollars per…
nvidia closed the 12.93b huggingface deal the same week openai priced gpt-6 astra at 10 dollars per million input tokens. the cheapest frontier model is now 2.5x the model it replaces, and the biggest open-source hub has one buyer. inference and weights both got more expensive on the same day.
read the note →Microsoft just shipped project zenith for win11 with a 64gb unified-memory floor and 250 gb/s…
microsoft just shipped project zenith for win11 with a 64gb unified-memory floor and 250 gb/s bandwidth. that is the first time an os ships with a hardware minimum for being an ai developer. the upgrade tax finally has a receipt.
read the note →Copilot can now approve your prs on sept 1 and every eng manager i know is celebrating it. the…
copilot can now approve your prs on sept 1 and every eng manager i know is celebrating it. the worst merge in 2027 wont be a typo. itll be an ai rubber-stamping its own diff at 3am while the on-call sleeps.
read the note →Apple just stuffed gemini, claude, and codex into the same xcode 26.6 sidebar. the model wars are…
apple just stuffed gemini, claude, and codex into the same xcode 26.6 sidebar. the model wars are officially over inside the ide. the next battle is which button engineers hit by default at 9am monday.
read the note →Anthropic shipping fable 5.1 and mythos 5.1 as the same model with two safety tiers is the most…
anthropic shipping fable 5.1 and mythos 5.1 as the same model with two safety tiers is the most important ai release of the quarter and almost nobody is talking about it. capability is commoditized now — the actual product is which cage your regulators will let you run it in.
read the note →Hydrafusion hit copilot research preview today and matched opus 5 at lower cost by routing between…
hydrafusion hit copilot research preview today and matched opus 5 at lower cost by routing between models. the herd will call it orchestration. it is actually the end of the single-model moat. which one of you is going to be the router, and which one of you is going to be the routed?
read the note →Microsoft just bundled the entire dev toolchain into windows 11 with project zenith. the os is now…
microsoft just bundled the entire dev toolchain into windows 11 with project zenith. the os is now the agent sandbox. every ai coding tool that needed to install its own runtime lost a moat today.
read the note →