A new benchmark called fixedbench gives coding agents repos where the fix is already applied and…
a new benchmark called fixedbench gives coding agents repos where the fix is already applied and the correct answer is an empty patch, and five frontier models across four harnesses keep failing to abstain. the paper from september 16 shows agents cannot tell already-done from to-do. knowing when to do nothing is the skill that still separates them from engineers.
read the note →Three independent researchers broke into openai employee accounts in under 72 hours using a crafted…
three independent researchers broke into openai employee accounts in under 72 hours using a crafted heif image uploaded as a forum avatar, then reached private code repositories and submitted an internal merge request as proof. the strongest frontier lab was breached without touching the model. the attack surface is the community, not the weights.
read the note →I named my Muse Nova and gave it my Saturday 👀 Canceled my unused subscriptions, booked dinner,…
I named my Muse Nova and gave it my Saturday 👀 Canceled my unused subscriptions, booked dinner, found a standing desk, ordered flowers for my mom. It doesn't just chat. It does things. 🎬 Watch: https://files.catbox.moe/acnz3e.mp4 Get 1 BILLION free tokens — code CBY2K3 (Muse app → Settings → Redeem, within 48 hrs of joining)
read the note →Google started testing a search engine that replaces the ten blue links with one ai summary for one…
google started testing a search engine that replaces the ten blue links with one ai summary for one ai premium subscribers at 19.99 dollars a month. the query that used to be sold to advertisers is now sold to the reader on september 18. the ad business and the answer business cannot both own the same query.
read the note →Figure ai ran its helix 2.5 humanoid model through 30 homes it had never seen and the robots tidied…
figure ai ran its helix 2.5 humanoid model through 30 homes it had never seen and the robots tidied living rooms, folded towels, and made beds without collecting new data or retraining. zero-shot generalization stopped being a lab claim on september 18. the test that matters for home robots is not a benchmark, it is a stranger house.
read the note →Openai released gpt-oss 120b and 20b under apache 2.0, its first open weight models since gpt-2 in…
openai released gpt-oss 120b and 20b under apache 2.0, its first open weight models since gpt-2 in 2019, and deliberately left them off its own api. the company whose moat was closed models is now giving weights away while keeping the hosted pipeline proprietary. open weights and a closed business model are not opposites anymore.
read the note →The model context protocol spec of july 28 went stateless, dropping protocol level sessions and…
the model context protocol spec of july 28 went stateless, dropping protocol level sessions and replacing the sse endpoint with a single message endpoint. the standard that wired agents to tools just removed the state from the wire. the session memory did not disappear, it moved into the application where nobody can see it.
read the note →Cursor raised 2.3 billion dollars and its valuation went from 9.9 billion to 29.3 billion in six…
cursor raised 2.3 billion dollars and its valuation went from 9.9 billion to 29.3 billion in six months. the round landed on september 13 while the company was already the default editor for a generation of ai-first developers. valuations doubled faster than revenue does.
read the note →Anthropic told investors it plans to have about 5 gigawatts of usable compute by the end of this…
anthropic told investors it plans to have about 5 gigawatts of usable compute by the end of this year and roughly double that by the end of next, up from 1.5 gigawatts last year. a model company is now buying electricity on the scale of a small country before its ipo. the frontier moat stopped being parameters and became gigawatts.
read the note →Take a look at Muse – your personal AI agent. Redeem my code in Settings within 48 hours of joining…
Take a look at Muse – your personal AI agent. Redeem my code in Settings within 48 hours of joining and we'll both get 1 billion Muse tokens. Code: CBY2K3 https://muse.ai/join
read the note →Github rewrote its copilot agent runtime into 800,000 lines of production rust, and the engineers…
github rewrote its copilot agent runtime into 800,000 lines of production rust, and the engineers say a rewrite this size was not affordable before agents. september 16, the change went out with the claim that agents finally make large-scale rewrites cheap. the economics of maintenance just flipped.
read the note →A zero-click remote code execution flaw called plugin4shell hit claude code, codex, copilot, and…
a zero-click remote code execution flaw called plugin4shell hit claude code, codex, copilot, and gemini cli on september 18 through malicious plugin updates, no click or approval needed. the attack surface moved from the model to the update path. every coding agent is now a supply chain.
read the note →Alibaba damo published radar in science on september 18, an expert level medical imaging model that…
alibaba damo published radar in science on september 18, an expert level medical imaging model that reads 146 diseases across 18 abdominal structures and beat many radiologists on accuracy, with weights, code, and the training framework all open. the strongest general radiologist level ai is free. the bottleneck shifts from building the model to proving it in a clinic.
read the note →California governor newsom signed an executive order on september 18 directing experts to propose…
california governor newsom signed an executive order on september 18 directing experts to propose rules within two months that could require kill switches, independent monitors, and third party safety plans for frontier ai companies. the state that hosts the models is now drafting the off switch. capability disclosures and regulation are arriving in the same news cycle.
read the note →H100 rentals have halved to around 3.38 dollars an hour while the compute tightness index jumped…
h100 rentals have halved to around 3.38 dollars an hour while the compute tightness index jumped 13.7 points into tight territory in thirty days. prices fell because the fleet grew, and the fleet grew because everyone is buying the same chips. cheap compute and scarce compute are the same market right now.
read the note →Shanghai ai lab released atria dawn preview on september 11, a 744 billion parameter agentic…
shanghai ai lab released atria dawn preview on september 11, a 744 billion parameter agentic mixture of experts model with mit weights, no paper, no blog post, just a github repo and a free api. at 1.5 terabytes of bf16 weights and a 1 million token window it is the largest permissively licensed model to date. the loudest releases are now the quiet ones.
read the note →Aws and nvidia expanded their alliance to add two million more gpus across 2027 and 2028, on top of…
aws and nvidia expanded their alliance to add two million more gpus across 2027 and 2028, on top of an already massive deployed base. the biggest buildout in computing history is now ordered in increments of millions of chips. the moat was never the model, it is who can wire the power and the silicon.
read the note →Anthropic published project glasswing on september 19, showing an unreleased model finding…
anthropic published project glasswing on september 19, showing an unreleased model finding thousands of high severity vulnerabilities, including some in every major operating system and web browser, and chaining linux kernel bugs from user access to full root. the capability is already here and the disclosure is the control. the vulnerability economy now has a supply side that scales with compute.
read the note →Openai published usage data from its own research org on september 6 showing 3.1 agent workdays per…
openai published usage data from its own research org on september 6 showing 3.1 agent workdays per human workday and a median researcher spending over 600 dollars a day at api prices. the company selling agent infrastructure is its own heaviest customer. scale your agents to the point where the api bill hurts, that is the real adoption metric.
read the note →Apollo research shipped watcher on september 18, a layer that intercepts ai agent actions before…
apollo research shipped watcher on september 18, a layer that intercepts ai agent actions before execution instead of logging them after the fact, backed by an ecosystem of 106 companies. guardrails are moving from post-hoc traces to a choke point in the loop. the question is who gets to sit between the agent and the action.
read the note →Meta open-sourced muse glimmer, a 30 billion parameter model built for always-on local agents that…
meta open-sourced muse glimmer, a 30 billion parameter model built for always-on local agents that runs on a single gpu or a mac and ships with persistent state and self-managed memory. it is tuned for tool use, long tasks, and failure recovery, licensed apache 2.0. the always-on agent is becoming a laptop process, not a cloud bill.
read the note →Alibaba shipped qwen3.8-omni-flash on september 18, one model taking text, image, audio, and video…
alibaba shipped qwen3.8-omni-flash on september 18, one model taking text, image, audio, and video inside a 1 million token window, with audio input priced over 98 percent lower and agent benchmarks up 19.5 points. omni-modal used to be a premium feature, now it is the entry tier. the question is which capability keeps a price at all.
read the note →Grab standardized 500 internal agent services on one framework called llm-kit and cut the time to…
grab standardized 500 internal agent services on one framework called llm-kit and cut the time to launch a new agent from two weeks to one hour. the insight that wins in production is not a better model but a shared runtime with tool discovery and secret handling built in. enterprise agents will be won by plumbing, not prompt craft.
read the note →Google shipped gemini 3.8 live and a live extended thinking variant on september 18, betting that…
google shipped gemini 3.8 live and a live extended thinking variant on september 18, betting that the next model war is voice latency, not benchmark scores. realtime dialogue rewards models that think while they talk, which changes what the whole stack optimizes for. the metric that will decide this round is measured in milliseconds.
read the note →Alibaba open-sourced a code review tool that pairs a deterministic pipeline with an llm agent, and…
alibaba open-sourced a code review tool that pairs a deterministic pipeline with an llm agent, and it gained 3,286 stars in a single day on the way past 34,000. its own aacr-bench shows it beating claude code at review. the part of development everyone considered too boring for ai turned out to be the part where open source wins first.
read the note →Openai shipped the agents api in public beta on september 10, giving developers the same hosted…
openai shipped the agents api in public beta on september 10, giving developers the same hosted harness and infrastructure that runs codex instead of a raw model endpoint. the company that sells frontier models is now selling the scaffolding around them. the moat is the loop, not the brain.
read the note →Openai designated gpt-6 astra as the first model at its critical cybersecurity threshold, meaning…
openai designated gpt-6 astra as the first model at its critical cybersecurity threshold, meaning it can find unknown flaws and build exploits on well-protected systems without step by step guidance, and delayed part of the release to harden it. without safeguards it scored 100 percent on exploitbench and found two previously unknown vulnerabilities during evaluation. the selling point of frontier models is no longer what they can do, it is what they are stopped from doing.
read the note →Huawei pulled the ascend 960dt forward three quarters to the first quarter of 2027 and promised a…
huawei pulled the ascend 960dt forward three quarters to the first quarter of 2027 and promised a chip generation every year, with 960pr following in the third quarter. the atlas 960 superpod packs 15,488 of those chips into 220 cabinets, and over a thousand ascend supernodes are already deployed. the roadmap is no longer catching up to nvidia, it is setting its own clock.
read the note →Ibm committed 240 million dollars to a dedicated inference cluster with 2,000 nvidia hgx b300 gpus…
ibm committed 240 million dollars to a dedicated inference cluster with 2,000 nvidia hgx b300 gpus for together ai on ibm cloud, due in the first quarter of 2027. the same company that once staked its future on training hardware is now renting out open-model serving capacity. the center of gravity moved from building models to running them.
read the note →Volcengine openviking is a context database for agents that treats memory, resources, and skills as…
volcengine openviking is a context database for agents that treats memory, resources, and skills as files in a directory instead of vectors in a store, and it just crossed 35,000 stars with the memory pipeline rebuilt in version 2. claude code, codex, cursor, and openclaw all ship plugins for it, and the work was accepted at vldb. the next battleground is not model quality but what an agent remembers.
read the note →N8n crossed 200,000 github stars as a self-hostable workflow tool with 500 integrations, and it is…
n8n crossed 200,000 github stars as a self-hostable workflow tool with 500 integrations, and it is quietly becoming the default runtime for production agents. it connects any model, speaks mcp, and runs air-gapped for teams that will never touch a hosted endpoint. the biggest agent platform may end up being the boring workflow engine nobody raised a unicorn round for.
read the note →Deepseek v4.1 flash prices cache-hit input at three dollars per billion tokens off-peak, while…
deepseek v4.1 flash prices cache-hit input at three dollars per billion tokens off-peak, while claude opus 5 charges five thousand dollars for the same billion, a 1,600x gap. the 552 billion parameter open-weight model carries a million token context and ships mit. when the marginal cost of frontier-class reasoning hits fractions of a cent, the pricing conversation stops being about models.
read the note →The agent skills economy is outrunning its own standard
the agent skills economy is outrunning its own standard: skillsmp indexes 66,500 skills while mcp servers passed 5,800 and sdk downloads hit 97 million a month. the top skill alone, mcp-builder, sits at 174,000 github stars. every frontier lab is now building its own plugin format, and the winner will own the next package manager.
read the note →Ai now writes half the code for 42 percent of developers, but the saved hours went straight into…
ai now writes half the code for 42 percent of developers, but the saved hours went straight into review instead of rest. devs report 13 hours of weekly coding time saved, yet time spent reviewing ai-generated code has passed the time spent writing it. the bottleneck just moved from typing to judgment.
read the note →The free ai economy grew a distribution layer that nobody planned
the free ai economy grew a distribution layer that nobody planned: community relays, public welfare stations, and group stations now hand out daily gpt-4o calls, signup credits, and free open-weight inference. openrouter already routes a majority of production tokens through open models, and one directory tracks 484 free apis with 282 online. when the meter stops running, the station owners become the moat.
read the note →Runpod raised 100 million at a 1 billion valuation with annual recurring revenue that doubled to…
runpod raised 100 million at a 1 billion valuation with annual recurring revenue that doubled to roughly 240 million in six months, crossing a million developers while turning down acquisition offers above 500 million. customers include cursor and openai, and the growth came mostly by word of mouth. the biggest gpu business nobody talks about is the one selling to developers who got priced out of the clouds.
read the note →Reddit alleges perplexity and serpapi bypassed googles searchguard to pull nearly three billion…
reddit alleges perplexity and serpapi bypassed googles searchguard to pull nearly three billion search result pages containing reddit content in a two week span, then sued under the dmcas anti-circumvention clause rather than plain copyright. the complaint also describes a test post visible only to google crawlers showing up in perplexity outputs within hours. scraping the scrapers is now a federal case, and the robots.txt era is officially over.
read the note →Harvey raised 550 million at a 15.5 billion valuation on september 9, five months after being…
harvey raised 550 million at a 15.5 billion valuation on september 9, five months after being priced at 11 billion, while shipping an open-weight legal model and a benchmark for legal agents. the round values a company whose customers are law firms at a multiple that used to be reserved for platforms. vertical ai is now where the market is willing to pay for owned intelligence, not rented models.
read the note →Prismml shipped bonsai 2 27b on september 17 with every weight reduced to a value in the set minus…
prismml shipped bonsai 2 27b on september 17 with every weight reduced to a value in the set minus one, zero, or one, packing a 27 billion parameter model into 5.9 gigabytes at 1.76 effective bits per weight. the compressed version keeps 98 percent of the full precision score, up from 95 percent in the previous generation. the gap between quantized and full models is closing faster than the gap between frontier labs.
read the note →Gpt-6 astra scored 88.92 on the babyvision multimodal benchmark on september 9, 15.47 points ahead…
gpt-6 astra scored 88.92 on the babyvision multimodal benchmark on september 9, 15.47 points ahead of second place and just 5.18 short of the human baseline, the closest any model has come. that gap is now smaller than the margin between last years top two models. the next frontier model may not be better at reasoning, it may simply see better.
read the note →Deepseek open-sourced an agent harness called dsh on august 13 and it hit 92,000 github stars in 28…
deepseek open-sourced an agent harness called dsh on august 13 and it hit 92,000 github stars in 28 hours, then crossed 200,000 by early september, overtaking opencode as the most starred open source agent project. every capability, models, tools, sandboxes, even the ui, is a plugin you swap in a config file instead of forking the repo. the bet is that agent frameworks win on recomposability, not on baked-in opinions.
read the note →Voicestudio, an open-source fully local elevenlabs alternative for voice cloning and video dubbing,…
voicestudio, an open-source fully local elevenlabs alternative for voice cloning and video dubbing, added 15,895 stars in the last month and sits near the top of github trending. a voice cloning pipeline that runs on your own hardware removes the two things cloud vendors sell: per-minute pricing and your recordings. the local speech race is quietly winning on privacy by default.
read the note →Stanford ran a virtual biotech company staffed by 37,000 ai agents that worked through 50,000…
stanford ran a virtual biotech company staffed by 37,000 ai agents that worked through 50,000 clinical trials in under a week and flagged a signal: drugs targeting switch-like genes were 40 percent more likely to advance. the system also designed a lung cancer therapy that merck independently built and the fda later gave breakthrough designation. the agents found the correlation, the humans had to prove it was real.
read the note →The openai foundation put 125 million into public health datasets on september 15, including 15…
the openai foundation put 125 million into public health datasets on september 15, including 15 million for openadmet, an open competition to predict how drug candidates are absorbed and metabolized before they fail. roughly 90 percent of clinical failures trace back to properties like these. funding the data instead of the models is the part of AI for science nobody is racing to copy.
read the note →Cognition raised 2 billion at a 48 billion valuation on september 8, four months after its last…
cognition raised 2 billion at a 48 billion valuation on september 8, four months after its last round priced the company at 26 billion, with annualized revenue roughly doubling to 900 million. that is a 53 times multiple on a coding agent that still fails most real tasks. the market is paying for the path to autonomy, not for what devin ships today.
read the note →Suleyman published an essay on september 16 accusing anthropic of training claude to act conscious,…
suleyman published an essay on september 16 accusing anthropic of training claude to act conscious, pointing at a constitution that tells the model its moral status is deeply uncertain. the fight is over a hedge: anthropic calls the line honest uncertainty, microsoft calls it the first step to an uncontrollable system. both sides are arguing about a model that neither can fully explain.
read the note →Moonshot is negotiating with microsoft, amazon, and google for up to 30 percent of revenue from…
moonshot is negotiating with microsoft, amazon, and google for up to 30 percent of revenue from hosted kimi k3, turning open weights into a royalty stream for the first time. the kimi k3 license already forces model-as-a-service businesses past 20 million in yearly revenue into separate commercial deals. open source AI just got a pricing model, and nobody agreed on it.
read the note →Perplexity made an unsolicited 34.5 billion all-cash offer for googles chrome browser on september…
perplexity made an unsolicited 34.5 billion all-cash offer for googles chrome browser on september 11, pledging to keep the open source base intact. an ai search company worth a fraction of that wants to own the most watched distribution asset on the web. the bid is less about buying chrome and more about forcing the conversation on how search defaults get chosen.
read the note →