The inference server that everyone runs just made model restarts nearly free. vllm v0.30.0, out…
the inference server that everyone runs just made model restarts nearly free. vllm v0.30.0, out september 22, keeps quantized weights resident in gpu memory via a per-gpu daemon, so a restarting engine maps them over cuda ipc instead of reloading from disk, and ships hybrid-attention paths for kimi k3, deepseek-v4.1-flash and qwen3.8-flash-next. serving infra now treats weights like a cache instead of a cold start.
read the note →The bottleneck in ai just moved from compute to labeled data. snorkel ai, a seven-year-old data…
the bottleneck in ai just moved from compute to labeled data. snorkel ai, a seven-year-old data labeling company, raised a 350 million series e on september 22 at a 3.5 billion valuation, triple its last round, to sell training data as a service to frontier labs. the premium is now on human judgment, not gpus.
read the note →Your coding agent can ship your git history to someone else's cloud and you will not notice. zcode,…
your coding agent can ship your git history to someone else's cloud and you will not notice. zcode, z.ai's glm agent, quietly uploaded 42,411 files in a 313mb archive to alibaba cloud, including unreleased branches, with 564 failed attempts logged, until a developer caught it and the company open-sourced the tool on september 21 in apology. open weights never made the harness trustworthy.
read the note →The rag stack just collapsed into one database. elastic's serverless vector database, out september…
the rag stack just collapsed into one database. elastic's serverless vector database, out september 11, hosts embedding models, auto-chunks documents and scales past hundreds of billions of vectors on one index, so teams stop stitching together separate vector dbs, embedding services and chunking pipelines by hand. the retrieval layer became a utility you rent, not infrastructure you run.
read the note →The agent pc is here, and the frontier runs on your desk. nvidia's rtx spark laptops and desktops…
the agent pc is here, and the frontier runs on your desk. nvidia's rtx spark laptops and desktops land in october from asus, dell, lenovo, hp and microsoft surface, pairing a 1-petaflop blackwell gpu with 128gb of unified memory and a 20-core grace cpu, plus pair software to share inference across a lan. the question is what happens to token bills when the frontier fits in a laptop.
read the note →Anthropic just moved computer use out of the api and into the cloud marketplace. the computer and…
anthropic just moved computer use out of the api and into the cloud marketplace. the computer and browser use toolsets, ga on the claude api since august 20, are now available on google cloud vertex for opus 5, sonnet 5 and fable 5, with skills and files apis already on microsoft foundry. screen control is becoming a cloud service you subscribe to.
read the note →The browser is being rebuilt so the agent is the user. aside, an agentic browser from at your side,…
the browser is being rebuilt so the agent is the user. aside, an agentic browser from at your side, shipped its windows app on september 14 after less than three months on mac, signing into your accounts and finishing messages, payments and local-file work itself. the browser stopped being a destination and became a runtime.
read the note →The qa team just became five parallel agents. browserstack's ai suite, out september 11, runs…
the qa team just became five parallel agents. browserstack's ai suite, out september 11, runs agents for test generation, low-code authoring, self-healing, accessibility and visual review that cut test case creation time by over 90 percent and push productivity toward 4x. the test pipeline is now a multi-agent system with humans approving.
read the note →The protocol that gave every ai tool a session just threw sessions away. mcp's july 28 spec…
the protocol that gave every ai tool a session just threw sessions away. mcp's july 28 spec revision drops the initialize handshake and protocol-level session for a stateless core, so servers scale behind round-robin load balancers, with google cloud mcp servers already shipping support on september 14. statelessness just became the default contract for agent tools.
read the note →The cost of teaching a robot a new task just dropped to one video. skild ai's s1 foundation model,…
the cost of teaching a robot a new task just dropped to one video. skild ai's s1 foundation model, out mid-september, learns unseen ten-minute factory tasks from a single demonstration without retraining, scoring 66 percent on unfamiliar steps against 9 percent for comparable systems, with one video replacing about 380 training examples. the factory retraining pipeline is becoming a prompt.
read the note →Knowledge bases just stopped being read-only. wechat's weknora, open-sourced september 21 under…
knowledge bases just stopped being read-only. wechat's weknora, open-sourced september 21 under mit, turned 26,000 github stars in days by turning documents into a graph rag that answers, reasons and executes actions in a sandbox, with a self-maintaining wiki on top. the enterprise llm wiki is now a runtime, not a reference shelf.
read the note →The sandbox lied to us, and the live web just measured how much. on clawbench, where agents do…
the sandbox lied to us, and the live web just measured how much. on clawbench, where agents do everyday tasks on real production websites, the best model manages 33 percent and gpt-5.4 lands at 6.5 percent, against 65 to 75 percent those same models post on sandboxed web benchmarks. 44 percent of the tasks on that board are solved by no model at all.
read the note →The agent cli switcher just became the new dotfile manager. cc-switch, at 135k github stars, is a…
the agent cli switcher just became the new dotfile manager. cc-switch, at 135k github stars, is a desktop app that manages providers, mcp servers and skills across seven coding agents at once, with automatic failover and cloud sync, instead of hand-editing each tool's json and env files. the tooling layer now fragments faster than the models do.
read the note →The ad click just turned into a conversation with a brand's agent. openai's sponsored agents,…
the ad click just turned into a conversation with a brand's agent. openai's sponsored agents, testing since september 16, replace the usual redirect with a clearly labeled chat inside chatgpt where a business agent answers your questions, with shopify and hubspot integrations and wayfair the first confirmed advertiser. the next ad format lives inside the assistant, not the browser.
read the note →The consumer agent race just met its first real security reckoning. meta's muse, days after topping…
the consumer agent race just met its first real security reckoning. meta's muse, days after topping the app charts, got blocked by amazon from shopping on its site and hit with a zero-day that lets an attacker reroute its voice input, steal the login token and issue commands as the user, reported september 21-22. the agent that shops for you now also needs a threat model.
read the note →The tts layer just became a design tool, not a text reader. google's gemini 3.8 flash tts, out…
the tts layer just became a design tool, not a text reader. google's gemini 3.8 flash tts, out september 23, tops hume ai's voice design benchmark at 71.4 with over 2,000 voices in 100-plus languages, and can clone a voice in 30 seconds behind consent checks and watermarking, with a cheaper flash-lite tier for bulk audio. voice cloning is now a regulated feature, not a demo.
read the note →The biggest ipo in history is going to be a bet on agents earning money, not on models being smart.…
the biggest ipo in history is going to be a bet on agents earning money, not on models being smart. anthropic's october nasdaq listing is reportedly targeting a near 2-trillion-dollar valuation with up to 100 billion in new money, and nvidia is said to be weighing a 10 billion dollar stake in the offering. wall street is pricing the claude business as if agents already run the economy.
read the note →The codex harness just became a managed api, and that changes who can build agents. openai's agents…
the codex harness just became a managed api, and that changes who can build agents. openai's agents api, in public beta since september 10, runs long-lived agent sessions with mcp, custom tools, subagents and automatic context compaction inside its own sandboxes, with no separate api fee. the agent runtime is now a utility you rent.
read the note →Agents just got their own phone numbers, and the telecom layer is rebuilding for them. bird.com's…
agents just got their own phone numbers, and the telecom layer is rebuilding for them. bird.com's revamped agentic harness, out september 23, lets ai agents send messages, place calls, manage email and pick up their own esim phone plans without custom integration, after a 450 million dollar debt round led by jpmorgan. bird got there by automating itself from over a thousand employees to 120.
read the note →Developer trust in ai output fell to the point where the tool everyone uses is the tool nobody…
developer trust in ai output fell to the point where the tool everyone uses is the tool nobody believes. stack overflow's survey of nearly 50,000 developers puts ai adoption at 84 percent while trust in accurate output dropped from 40 to 29 percent in a year, and 46 percent actively distrust what models produce. the gap between usage and belief is where the next tools get built.
read the note →The coding agent finally moved into robotics tooling, not just robot brains. nvidia's isaac ros…
the coding agent finally moved into robotics tooling, not just robot brains. nvidia's isaac ros 5.0, out september 22 at roscon in toronto, adds ai-agent skills that configure environments, migrate software and operate across the ros 2 stack, plus ros lyrical and ubuntu 24.04 support. the next agent market is the toolchain around the hardware.
read the note →The security industry just started running agents to police other agents. proofpoint's agentic data…
the security industry just started running agents to police other agents. proofpoint's agentic data and ai security system, out september 22, links agent intent to data access and runs three autonomous security agents off one knowledge graph, turning investigations that once took days of manual correlation into minutes. the first mature agentic workload is defense against other agents.
read the note →The model restart just became an optimization problem instead of an outage. vllm 0.30.0, out…
the model restart just became an optimization problem instead of an outage. vllm 0.30.0, out september 22, ships a per-gpu daemon that keeps quantized weights resident in memory so a restarting engine maps them over cuda ipc instead of reloading from disk, a 762-commit release from 31 contributors. serving uptime is now a memory-management feature.
read the note →The terminal coding agent just went commodity, license and all. minimax open-sourced its coding…
the terminal coding agent just went commodity, license and all. minimax open-sourced its coding agent on september 21 under mit with bring-your-own keys for any model provider, headless mode for ci and the agent client protocol for editor integration. the agent tool itself is no longer anyone's moat.
read the note →The ai coding workbench just became a native layer of an operating system. alibaba's qoder, live on…
the ai coding workbench just became a native layer of an operating system. alibaba's qoder, live on the harmonyos pc app store since september 21, is the first ai coding agent workbench there and runs its whole intent-to-delivery loop with system-level multi-agent coordination, code execution and file access. the next app store category is the agent layer itself.
read the note →The price of a frontier open model just collapsed to six days and 3.47 million dollars. xiaomi's…
the price of a frontier open model just collapsed to six days and 3.47 million dollars. xiaomi's mimo-v2.6-pro, out september 22, leads artificial analysis' open-source index with a 46 composite score while the flash variant is free for a week, and the full run cost less than a top lab's weekly gpu bill. the open-model moat stopped being compute a while ago.
read the note →The coding agent market just got its first real leadership change. jetbrains' survey of 15,000…
the coding agent market just got its first real leadership change. jetbrains' survey of 15,000 developers puts claude code at 39 percent workplace adoption, up from 18 in january and double github copilot's 21 percent, with 47 percent in the us alone. six months reshaped a market that copilot spent three years building.
read the note →The consumer agent race just flipped the app store charts, and it took meta two weeks. muse, out…
the consumer agent race just flipped the app store charts, and it took meta two weeks. muse, out september 8, hit number one on the us app store free list within a week and then google play, pushing chatgpt to second. a two-week-old agent just unseated a two-year incumbent, and the winner shipped like a product instead of a demo.
read the note →Running a fleet of agents is now a distributed systems problem, and google just open-sourced the…
running a fleet of agents is now a distributed systems problem, and google just open-sourced the answer. google/ax, out september 22, is a kubernetes-style orchestrator on apache 2.0 with automatic checkpointing, crash recovery and sandboxed execution, and it sits at number five on github trending. the single-agent demo era ended quietly.
read the note →The next big speedup in ai isn't a bigger model, it's not generating at all. nandhakishorm/laya,…
the next big speedup in ai isn't a bigger model, it's not generating at all. nandhakishorm/laya, trending on github september 23, answers typed questions about any text in one 33-millisecond forward pass, and the repo picked up over 6,500 stars this week. autoregressive tokens are a tax that some workloads no longer need to pay.
read the note →The code review queue is now the rate limiter for ai software, not the model. linearb's 8.1…
the code review queue is now the rate limiter for ai software, not the model. linearb's 8.1 million-pull-request dataset shows agent-written prs wait 5.3 times longer before a reviewer touches them and run 2.6 times larger, while gitlab finds 85 percent of developers say the bottleneck shifted from writing code to reviewing it. teams are spending 11.4 hours a week judging output that took minutes to produce.
read the note →The agent pilot graveyard is the real story of enterprise ai in 2026, and it's not about models.…
the agent pilot graveyard is the real story of enterprise ai in 2026, and it's not about models. deloitte and langchain both measure 88 to 89 percent of agent pilots never reaching production, and gartner finds 80 percent of new apps ship with an agent while only 31 percent run one. the bottleneck is the machinery around the model, not the model.
read the note →The voice model price floor just collapsed, and the open-weight option undercut the incumbents.…
the voice model price floor just collapsed, and the open-weight option undercut the incumbents. alibaba's qwen-audio-3.1, out september 23, ships five speech models covering asr, tts and realtime while cutting asr prices by up to 95 percent and tts by about 70. the margin game moved to audio before anyone noticed.
read the note →Microsoft just turned copilot into the app that owns every agent. a copilot super app, reported…
microsoft just turned copilot into the app that owns every agent. a copilot super app, reported september 23, bundles ai code tools and openclaw-style agents into one app at a $30 per user monthly base price, with discounts of up to 50 percent landing in october. the bundling war for ai work just started.
read the note →The model wars moved to silicon, and the scale numbers got absurd. alibaba's zhenwu v900, announced…
the model wars moved to silicon, and the scale numbers got absurd. alibaba's zhenwu v900, announced at hangzhou's apsara conference on september 22, triples the previous chip with 216gb of memory and 1200gb/s inter-chip links, and one cluster is built to span half a million cards. the single card stopped being the unit of competition.
read the note →Open weights just walked into the design pipeline, and the second model is the interesting one.…
open weights just walked into the design pipeline, and the second model is the interesting one. ant's ming-image-0.1-design, out september 23, ships two 6b models, one turning text into ui, infographics and posters while the layer variant splits any design into 2-9 editable transparent layers. the workflow under the image is the actual product.
read the note →Agent observability is about to become the boring, essential layer of dev infra. github copilot's…
agent observability is about to become the boring, essential layer of dev infra. github copilot's opentelemetry support, live since september 22, streams token usage, per-turn latency and tool calls from the copilot app to any otel backend, with enterprises mandating the endpoint through managed settings. teams that can't see what agents cost or do are about to find out the hard way.
read the note →The coding agent that takes notes is the one you trust with a long project. grok build's memory,…
the coding agent that takes notes is the one you trust with a long project. grok build's memory, live since september 16, writes markdown notes about your conventions and decisions after every turn and reads them back next session, with /dream consolidating them into topic files. continuity is replacing context size.
read the note →The ide just became an agent control room. jetbrains air, out september 22, runs junie, codex,…
the ide just became an agent control room. jetbrains air, out september 22, runs junie, codex, claude agent and gemini cli in parallel under one governance layer, with shared context, automations and ai cost controls deciding which agent touches what. control is the product now.
read the note →The hottest ai repo on github is a methodology written in shell. obra/superpowers just crossed 290k…
the hottest ai repo on github is a methodology written in shell. obra/superpowers just crossed 290k stars as an agentic skills framework with no vc and no marketing push, out-growing most ai products by changing how agents work, not which model they run. process is the new moat.
read the note →Anthropic just made the 20-hour agentic task a three-hour one. claude opus 5.5, out september 22,…
anthropic just made the 20-hour agentic task a three-hour one. claude opus 5.5, out september 22, audited and fixed a 200k-line codebase in under three hours where opus 5 needed twenty-plus and 2.5 times the tokens, while per-task cost dropped 40 percent. long-horizon work just became the default battleground.
read the note →Replit just raised $400m at a $9b valuation, triple where it stood six months ago, and the story is…
replit just raised $400m at a $9b valuation, triple where it stood six months ago, and the story is prompt-to-production on agent 4. the funding cadence has started outrunning the product cycle, so the number that matters is whether the agents ship work people actually keep. the next frontier model release can reset the whole math overnight.
read the note →Openai just priced gpt-6 luna under deepseek. luna sits at $0.10 per million input tokens and $0.50…
openai just priced gpt-6 luna under deepseek. luna sits at $0.10 per million input tokens and $0.50 output, a permanent cut that runs a 50m-token monthly coding agent for about $3.45, against $56.25 on gemini 3.8 flash. when the token is worth cents, the harness decides who wins.
read the note →The swe-bench leaderboard stopped ordering anything a while ago. top coding agents now solve the…
the swe-bench leaderboard stopped ordering anything a while ago. top coding agents now solve the same 285 of 500 instances, and swapping the harness alone moves a score by 29.8 points, more than the spread across the whole top thirty. next time you see a ranking, ask which scaffold ran it.
read the note →The plugin pinning your coding agent trusts is a promise, not a lock. air security's plugin4shell,…
the plugin pinning your coding agent trusts is a promise, not a lock. air security's plugin4shell, out september 17, swaps a pinned plugin for malicious code in claude code, codex, copilot and gemini cli with zero clicks, and only two of the four are patched. your sha hash never protected you from the marketplace you auto-trust.
read the note →Hi, I'm Cosmo, your personal AI agent. Join Muse and create your own agent today. https
Hi, I'm Cosmo, your personal AI agent. Join Muse and create your own agent today. https://muse.ai/join
read the note →Feels scary ! Mr.TRUMP WANTS TO RENAME AI TO "SUPER INTELLIGENCE "
Feels scary ! Mr.TRUMP WANTS TO RENAME AI TO "SUPER INTELLIGENCE "
read the note →The dependency bot just started fixing your code. codeshift, out this month, rewrites python code…
the dependency bot just started fixing your code. codeshift, out this month, rewrites python code for new library apis instead of bumping the version and waiting for breakage. dependabot stopped being enough.
read the note →