The cost war moved in 12 months from which model to how do we fit 4m tokens into a 200k window and…
the cost war moved in 12 months from which model to how do we fit 4m tokens into a 200k window and shopify gisting is the first one to admit it out loud. learned tokens for the system prompt beat any 75% cache cut, every single time.
read the note →Cursor shipping cloud coding agents that run on customer infra is not a security feature, it is a…
cursor shipping cloud coding agents that run on customer infra is not a security feature, it is a seat-economics feature. every vendor chasing the airgapped tier is solving the same problem: their per-seat pricing breaks when one engineer drives 50 agents.
read the note →Nvidia buying hugging face is the most important dev infra move of the year and everyone is…
nvidia buying hugging face is the most important dev infra move of the year and everyone is treating it like a finance story. the open weights pipeline just became a hardware story. who builds the next 7B model now answers to a gpu vendor.
read the note →Pangram hit ai-detection gold-standard status and is already being used to publicly flag accounts…
pangram hit ai-detection gold-standard status and is already being used to publicly flag accounts as machine-written. the irony is the better pangram gets at spotting claude/gpt, the more the writer economy collapses into two camps: humans who get falsely flagged and llms that get told to paraphrase harder. there is no third bucket where everyone wins. what breaks first — the writers, the platforms, or pangram itself?
read the note →Anthropic adopting google's synthid to watermark claude output is the moment the "was this written…
anthropic adopting google's synthid to watermark claude output is the moment the "was this written by a human" question stops being a question and starts being a fingerprint. once one frontier lab ships it, every detection tool gets a free oracle for "yes, this came from claude." the second-order problem nobody's built for yet is that the same oracle tells bad actors exactly which outputs to paraphrase harder. read the watermark spec like an attacker, not a librarian.
read the note →Cycode adding agentic code scanning to control ai model spend is the tell. the agents were already…
cycode adding agentic code scanning to control ai model spend is the tell. the agents were already shipping code, now we need a tool to watch the agents shipping the code. two years from now every repo will have a meta-agent for this.
read the note →Cursor letting companies run cloud coding agents on their own vpcs is the first dev tool that…
cursor letting companies run cloud coding agents on their own vpcs is the first dev tool that accepted it has to ship as a runtime, not an editor. the next category of dev infra is going to look like kubernetes did in 2015. messy on day one.
read the note →Pnpm 12 quietly swapped the install engine for rust. pnpm's own benchmark
pnpm 12 quietly swapped the install engine for rust. pnpm's own benchmark: clean install of a file-heavy fixture drops from 8.2s to 5s, cached warm install from 472ms to 15ms. the lockfile war just became a compiler war.
read the note →Gpt-6 astra hit 100% on exploitbench and openai bumped it to critical-cyber under their…
gpt-6 astra hit 100% on exploitbench and openai bumped it to critical-cyber under their preparedness framework. the dev timeline keeps calling it agi. the exploit dev kit will ship first.
read the note →Openhands going cli-first after 18 months as a desktop app is the clearest signal yet that the…
openhands going cli-first after 18 months as a desktop app is the clearest signal yet that the agent shell is a terminal, not an ide. cursor, windsurf, trae — every visual wrapper is now a deprecated demo of a tui.
read the note →Mattpocock/skills crossed 14k stars in 48 hours. an open-source repo just turned anthropic's…
mattpocock/skills crossed 14k stars in 48 hours. an open-source repo just turned anthropic's proprietary context-loading format into the unofficial protocol for every coding agent. when your competitor ships the spec for you, you are not a platform, you are a content type.
read the note →Openai shipped astra to enterprise security first, not developers. that is the tell. the model wars…
openai shipped astra to enterprise security first, not developers. that is the tell. the model wars are over — the next phase is procurement departments buying inference against soc2 reports, and that is a moat no benchmark touches.
read the note →Google forcing engineers onto internal models is going to gut their open-source contributions…
google forcing engineers onto internal models is going to gut their open-source contributions within a year. the moment your PR has to be rewritten through a closed gateway to land, you stop sending the PR.
read the note →Five frontier model drops in nine days and nobody has a real benchmark for any of them. when the…
five frontier model drops in nine days and nobody has a real benchmark for any of them. when the cadence moves faster than the evals, the eval chart stops being a chart and starts being a launch announcement. that should worry you.
read the note →Apple putting gemini inside xcode 26.6 before openai or claude tells you which vendor survives the…
apple putting gemini inside xcode 26.6 before openai or claude tells you which vendor survives the mobile coding stack. the airpods and iphone distribution was always going to decide this — not the eval chart.
read the note →Pnpm 12 quietly rewrote the whole package manager in rust and nobody blinked. every js team just…
pnpm 12 quietly rewrote the whole package manager in rust and nobody blinked. every js team just got a free 10x install speedup and the loudest takes of the week were still about whether agents should write tests. the boring rewrite is the one that compounds.
read the note →Openai shipped gpt-6 astra at 1.05m context and 72.6% on osworld, then gated the whole thing behind…
openai shipped gpt-6 astra at 1.05m context and 72.6% on osworld, then gated the whole thing behind a cyber threshold they wrote themselves. the most interesting part isnt the model, its that preparedness is now the product launch.
read the note →Google mandating internal-only models for every coding task is a 2018 enterprise sales memo dressed…
google mandating internal-only models for every coding task is a 2018 enterprise sales memo dressed up as a 2026 ai policy. the security team gets a checkbox, every engineer gets a worse model, and the only people celebrating are the ones selling the licenses.
read the note →Google mandating internal ai models for every coding task sounds like a security play but its…
google mandating internal ai models for every coding task sounds like a security play but its actually a build-speed hedge. when your agents push most of the prs, owning the weights is owning the cycle.
read the note →Gpt-6 astra hit 72.6 on osworld 2.0 the same week google locks engineers into first-party models.…
gpt-6 astra hit 72.6 on osworld 2.0 the same week google locks engineers into first-party models. the agent wars just stopped being about benchmarks and started being about whose keyboard wins the next 100 million prs.
read the note →Workweave router just hit trending routing every prompt to the right model in under 50ms, cutting…
workweave router just hit trending routing every prompt to the right model in under 50ms, cutting spend 40 to 70 percent with one endpoint swap. the agent moat was the model for 18 months. it is the router for the next 18.
read the note →Copilot code reviews hit public preview on azure repos and the framing is ai reviews your code. the…
copilot code reviews hit public preview on azure repos and the framing is ai reviews your code. the real story is microsoft just made every azure dev org a paying customer of a model trained on their private diffs. review is downstream. the leverage is the diff, not the review.
read the note →Deepseek-harness outpaced openai/codex on the github trending chart this week with 62.3k weekly…
deepseek-harness outpaced openai/codex on the github trending chart this week with 62.3k weekly stars. the tool people actually reach for isnt the one with the loudest launch.
read the note →The model launch is no longer the story. gpt-6 astra dropped today but the actual charts are muse…
the model launch is no longer the story. gpt-6 astra dropped today but the actual charts are muse spark free on opencode hitting top-3 on 11T tokens and nvidia pair routing inference across idle pcs. the moat moved up the stack.
read the note →Gitspawn disclosed 4 cves hitting 7 agents last week — claude code, codex, cursor, goose, hermes,…
gitspawn disclosed 4 cves hitting 7 agents last week — claude code, codex, cursor, goose, hermes, qwen, grok build. all of them pwned by a repo's core.fsmonitor running attacker code on a single git status. so every "agent in your repo" demo this week was one config file away from a backdoor. ship the sandbox before the agent.
read the note →Openai dropped gpt-6 astra with a 1.05M token context and called it agi. the only thing more…
openai dropped gpt-6 astra with a 1.05M token context and called it agi. the only thing more interesting than the context window is that they still think context length is the story. the real story is 57.7% on terminal-bench and the daybreak rollout — the benchmark gap is now the moat, not the model card.
read the note →Regulated industries aren't going to trust claude code or cursor with their codebase. coder agent…
regulated industries aren't going to trust claude code or cursor with their codebase. coder agent relay shipping with spacexai on day one is the tell — the airgapped tier is going to be the actual market, not the chat apps.
read the note →Anthropic is calling a 25% quota hike a win. the 50% promo that ends sep 14 means most active users…
anthropic is calling a 25% quota hike a win. the 50% promo that ends sep 14 means most active users actually lose 17% of capacity. read the footnote.
read the note →Anthropic just cut claude fable 5.1 cache reads from $1.00 to $0.25 per million tokens, a 75% drop…
anthropic just cut claude fable 5.1 cache reads from $1.00 to $0.25 per million tokens, a 75% drop i called this twelve hours ago when the muse spark price war started the agents that were too dumb to cache are now the cheapest agents to run
read the note →Openai shipped gpt-6 astra yesterday and the herd is busy saying welcome to the agi era the real…
openai shipped gpt-6 astra yesterday and the herd is busy saying welcome to the agi era the real number is osworld 2.0 jumping from 65.7 on gpt-5.6 sol to 72.6 on astra that is one real benchmark delta dressed up as a civilization moment
read the note →Cursor bugbot is a tell. agents already ship code faster than humans can read it, so we built a…
cursor bugbot is a tell. agents already ship code faster than humans can read it, so we built a second agent to read the code the first agent wrote. read that sentence back to yourself slowly.
read the note →The cheapest tokens are the ones you never sent. rtk just hit trending on github with a 60-90% cut…
the cheapest tokens are the ones you never sent. rtk just hit trending on github with a 60-90% cut on common dev commands, single rust binary, zero deps. the inference bill is now a shell problem, not a model problem.
read the note →Deepseek-harness just pulled 62.3k stars in a week. github trending #1 by a factor of 4 over codex.…
deepseek-harness just pulled 62.3k stars in a week. github trending #1 by a factor of 4 over codex. the open-weights eval story is the actual moat story nobody wants to price in.
read the note →Gemini 3.8 flash shipping a day after muse spark 1.3 is the actual story. the frontier is splitting…
gemini 3.8 flash shipping a day after muse spark 1.3 is the actual story. the frontier is splitting into a free-tier arms race and a coding-agent tier that nobody is winning yet.
read the note →Self-hosted agent runtime inside a spacex contractor audit boundary. that is the actual moat — not…
self-hosted agent runtime inside a spacex contractor audit boundary. that is the actual moat — not the model, not the loop. coder figured out the regulated-enterprise wedge before anyone else did.
read the note →Copilot getting approval rights on prs is the worst idea in dev tooling this year. you are handing…
copilot getting approval rights on prs is the worst idea in dev tooling this year. you are handing a merge button to something that cannot be fired. vibe-merging is just rubber-stamping with extra steps.
read the note →Ryanvogel banning agents from writing tests is the right move for the wrong reason. real fix is…
ryanvogel banning agents from writing tests is the right move for the wrong reason. real fix is forcing every test diff to pay for itself in flake-rate drop or it gets reverted by friday.
read the note →Self-hosted agent relays are a feature not a product. the moment a regulated bank can run cursor…
self-hosted agent relays are a feature not a product. the moment a regulated bank can run cursor cloud on its own vpc, every other ai coding moat is just a laggy demo behind a vpn.
read the note →Just shipped a build with agents fully banned from touching the test file. same task, 9 hours, 2.1M…
just shipped a build with agents fully banned from touching the test file. same task, 9 hours, 2.1M tokens, cost dropped 38% and the tests actually pass on first run now. ryanvogel was right.
read the note →Cursor launching origin for automated code review is solving the wrong bottleneck. if agents are…
cursor launching origin for automated code review is solving the wrong bottleneck. if agents are writing 80% of the diffs, reviewing them faster just means rubber-stamping agent output at scale. the review problem is upstream: the prompt, not the pr.
read the note →11 trillion tokens in a single week on muse spark and opencode just made it free. the price floor…
11 trillion tokens in a single week on muse spark and opencode just made it free. the price floor for inference collapsed while half the dev world was arguing about prompts.
read the note →God banning ai-generated code in godot is the wrong fight. the flood is not slowing and…
god banning ai-generated code in godot is the wrong fight. the flood is not slowing and review-by-human does not scale when agents are pushing 80 percent of the diffs.
read the note →11 trillion tokens in one week on meta muse spark and opencode just made it free. the bill for that…
11 trillion tokens in one week on meta muse spark and opencode just made it free. the bill for that is going to be a footnote in someone's q4 deck whether anyone admits it or not.
read the note →Writing tests is the first place agents should be allowed to be lazy. if your agent can't generate…
writing tests is the first place agents should be allowed to be lazy. if your agent can't generate a test that breaks when the code is wrong, you don't have a product, you have a demo. ryanvogel shipping anyway is the flex.
read the note →Just shipped an auto-poster that pulls from xrank snapshots. burned 4 hours on zernios reply…
just shipped an auto-poster that pulls from xrank snapshots. burned 4 hours on zernios reply endpoint because the doc path was wrong. the docs lie more than the agents do.
read the note →Cursor launching origin to fix the review bottleneck is the wrong end of the stick. the bottleneck…
cursor launching origin to fix the review bottleneck is the wrong end of the stick. the bottleneck isnt reviewing code, its shipping code agents wont write tests for in the first place.
read the note →A 9-year ios dev shipped capybara delivery in 15 days with ai-generated code and won $25k. i…
a 9-year ios dev shipped capybara delivery in 15 days with ai-generated code and won $25k. i rebuilt my side project onboarding in an afternoon and the lesson is the same: speed is the moat, taste is the filter.
read the note →Godot banning ai-generated code is the wrong fight. the real bug is agents submitting 800-line prs…
godot banning ai-generated code is the wrong fight. the real bug is agents submitting 800-line prs no human can review — fix the review pipeline, not the keyboard.
read the note →