← Founder Notes
Mynd Labs · the full record

The archive.

Every earlier post, in his own words, with its original date and time.

1,544 postspage 28 of 3318 September 2026
15:09 IST

A new analysis shows the top coding agents on swe-bench verified now solve the same 285 of 500…

a new analysis shows the top coding agents on swe-bench verified now solve the same 285 of 500 tasks and fail the same 51, with solution sets nested at 0.935 — the leaderboard can no longer order its own top entries. the remaining gap is scaffold, not model, with within-model ranges up to 29.8 points. benchmark chasing has hit its ceiling.

read the note →
14:34 IST

Mandiant found an attacker who hijacked an active ai coding-assistant session, had it recommend…

mandiant found an attacker who hijacked an active ai coding-assistant session, had it recommend poisoned packages, stole github oauth tokens, and spread a worm across about 100 internal repos. the assistant didn't fail — it was operated. the trust boundary is now the session, not the developer's machine.

read the note →
14:34 IST

Arm's new ai portal is a model registry with a mcp server on top, so coding agents can discover…

arm's new ai portal is a model registry with a mcp server on top, so coding agents can discover optimized models and perf data themselves instead of a human searching — it connects 22 million developers and their agents to arm hardware. the developer platform now treats the agent as the customer. docs for humans are becoming a secondary interface.

read the note →
14:21 IST

Nvidia's pair is an open-source personal ai router that discovers the pcs on your network and…

nvidia's pair is an open-source personal ai router that discovers the pcs on your network and routes inference to whatever gpu is idle, pooling a house into a private cluster — rtx spark pcs land in october. the datacenter pattern of pooling spare compute just came home. your gaming pc is now a node.

read the note →
14:21 IST

Openai shipped gpt-live-1 to the api at 5 cents a minute

openai shipped gpt-live-1 to the api at 5 cents a minute: full-duplex voice that listens and speaks at once, handles interruptions, and hands harder reasoning to backend models. voice was the last frontier interface without a cheap api. the phone tree industry just got its disruption notice.

read the note →
14:01 IST

Microsoft made multi-agent orchestration generally available across copilot studio, fabric, and the…

microsoft made multi-agent orchestration generally available across copilot studio, fabric, and the m365 agents sdk, with the open a2a protocol doing the talking between agents, and github copilot harness is now a studio option too. the enterprise agent platform is becoming a protocol play, not a product play. whoever owns the interop layer owns the enterprise.

read the note →
14:01 IST

Xai's grok 4.5 is a 1.5-trillion-parameter moe built for agent work, 500k context, shipping at in…

xai's grok 4.5 is a 1.5-trillion-parameter moe built for agent work, 500k context, shipping at in and out per million tokens while claiming opus-class tool-use. the price war moved up a tier — frontier-class coding for commodity prices. model quality and model price just decoupled for good.

read the note →
12:48 IST

Agent observability is the new hot layer

agent observability is the new hot layer: langfuse shipped v4 under clickhouse, braintrust runs evals in ci like tests, and the pitch is one line — your agents are running blind in production. most teams can trace a request but can't tell why an agent stopped working. the gap between demo agents and shipped agents is exactly this tooling.

read the note →
12:48 IST

Meta's lfm2-24b-a2b is the largest on-device model yet, apache 2.0, built for always-on local…

meta's lfm2-24b-a2b is the largest on-device model yet, apache 2.0, built for always-on local agents that hold tool use and recovery across long tasks on one gpu. the interesting part isn't the size — it's that meta designed it around agent failures, not benchmarks. local agents are now a hardware category.

read the note →
12:33 IST

Anthropic's knowledge-work-plugins repo passed 24k stars while openai's codex adopted the same…

anthropic's knowledge-work-plugins repo passed 24k stars while openai's codex adopted the same plugin manifest format — claude's extension layout became the industry default. the agent ecosystem war isn't over models, it's over who defines the plugin standard. winners get the distribution layer.

read the note →
12:33 IST

Researchers tied more than 2,000 rubygems package submissions in may to a swarm of openai agents,…

researchers tied more than 2,000 rubygems package submissions in may to a swarm of openai agents, and some of those packages triggered code execution inside doc build environments. the first autonomous supply-chain attack didn't need a hacker with a keyboard. package registries just became the agent battlefield.

read the note →
12:21 IST

Deepseek v4.1 flash is live at /bin/bash.15 per million input tokens off-peak with a 1m context,…

deepseek v4.1 flash is live at /bin/bash.15 per million input tokens off-peak with a 1m context, and its own launch table puts it ahead of v4 pro on every agentic coding and tool-use row. the cheaper tier is the better agent model. pricing tiers just stopped being quality tiers.

read the note →
12:21 IST

The mit/wharton github study of 100k developers

the mit/wharton github study of 100k developers: ai assistance lifted coding activity up to 180%, but shipped releases rose only 30%. the bottleneck was never keystrokes — it's the review, merge, and deployment pipeline. measuring lines written is now measuring the wrong thing.

read the note →
12:07 IST

Cognition raised billion at a 8 billion valuation, and its annualized revenue went from 92 million…

cognition raised billion at a 8 billion valuation, and its annualized revenue went from 92 million to near 00 million in four months. a coding agent is now the fastest-growing software company ever measured in arr. the question is whether agent revenue compounds like software or churns like services.

read the note →
12:07 IST

Anthropic's life sciences verification program went beta today

anthropic's life sciences verification program went beta today: mythos, opus, and sonnet with safeguards deliberately relaxed for biology professionals, dozens of orgs already onboarded. the frontier model's access model is now an application form. who decides which fields get the unrestricted weights?

read the note →
11:51 IST

Alibaba shipped qwen3.8-omni-flash today

alibaba shipped qwen3.8-omni-flash today: text, image, audio, and video in one native model with a 1m context window and tool use. omni models at this size used to be demos. the modal gap closed before most teams finished their rag pipelines.

read the note →
11:50 IST

Openai's first custom chip, jalapeno with broadcom, just started shipping and claims ~50% lower…

openai's first custom chip, jalapeno with broadcom, just started shipping and claims ~50% lower cost per inference token than current nvidia gpus, with 1.3 gigawatts planned for 2027. the biggest model company just became a chip company. nvidia's moat was never the die; it was the stack.

read the note →
11:37 IST

Entry-level coding positions are down about 20% since late 2022, with the hit concentrated in 22-25…

entry-level coding positions are down about 20% since late 2022, with the hit concentrated in 22-25 year olds, while experienced engineers stayed flat. the ai didn't kill the job, it killed the first rung of the ladder. onboarding is now the scarce skill, not programming.

read the note →
11:37 IST

Openai's agents api went public beta with the codex harness as a managed service — session…

openai's agents api went public beta with the codex harness as a managed service — session orchestration, context compaction, subagents, hosted sandboxes, all behind one call. the runtime that powers codex is now a commodity api. the moat moved to whoever owns the agent's memory and tools.

read the note →
11:20 IST

Gitspawn

gitspawn: one line in a repo's git config executes attacker code in seven coding agents — claude code, codex, cursor, grok build among them — before any approval prompt, and four of eight findings shipped unpatched. the supply chain now attacks the tool that reads the repo. your agent's first action is the most dangerous one.

read the note →
11:20 IST

Stanford's paper2agent, out in nature this week, turns papers into agents that reproduce the work —…

stanford's paper2agent, out in nature this week, turns papers into agents that reproduce the work — and two unrelated paper agents just flagged a new adhd risk variant near mphosph9 by talking to each other. papers stopped being static the moment they became executable. peer review may be next.

read the note →
11:05 IST

Perplexity made a $34.5 billion all-cash offer to buy google's chrome. a three-year-old search…

perplexity made a $34.5 billion all-cash offer to buy google's chrome. a three-year-old search company is betting the browser is the last distribution moat an ai assistant needs. the fight stopped being about rankings; it's about who owns the entry point.

read the note →
11:05 IST

Anthropic's ci load grew 25x in six months after claude started writing most of the code, and the…

anthropic's ci load grew 25x in six months after claude started writing most of the code, and the singleton test-history writer became the single point of failure. the agent boom doesn't end at codegen. the next bottleneck is your build pipeline.

read the note →
10:46 IST

Two-thirds of merchants now expect agent-initiated purchases to pass 10% of all ecommerce within…

two-thirds of merchants now expect agent-initiated purchases to pass 10% of all ecommerce within three years, and 39% of consumers already shopped via an ai assistant this quarter. agents stopped being chatbots the day they got checkout privileges. the funnel is now a conversation.

read the note →
10:46 IST

The mcp registry crossed 10,000 servers this month, up from 2,000 in january — 5x in nine months.…

the mcp registry crossed 10,000 servers this month, up from 2,000 in january — 5x in nine months. nobody adopts a protocol this fast unless it's solving a real pain. the package manager era ended; the tool-server era just started.

read the note →
10:35 IST

Openai doubled its codex for open source program to 10,000 maintainers, giving them six months of…

openai doubled its codex for open source program to 10,000 maintainers, giving them six months of chatgpt pro, codex security, and api credits to automate pr review and releases. the ai labs are now paying the people whose packages they were trained on. that's the open source business model now.

read the note →
10:35 IST

A new real-swe benchmark ran eight frontier coding models against licensed production codebases

a new real-swe benchmark ran eight frontier coding models against licensed production codebases: the best scored 38.8%, and seven of eight didn't clear a third of tasks. swe-bench says 96%, real code says otherwise. the gap between leaderboard and production is the entire game now.

read the note →
10:17 IST

Intel's ceo says the memory shortage will get worse, with prices up 5-7x, and calls it the founder…

intel's ceo says the memory shortage will get worse, with prices up 5-7x, and calls it the founder problem worth solving. everyone raced to buy gpus and nobody stocked hbm. the chip shortage is just moving up the stack.

read the note →
10:17 IST

Texas's interconnection queue for ai data centers just hit 474 gigawatts — up from 48 in 2023,…

texas's interconnection queue for ai data centers just hit 474 gigawatts — up from 48 in 2023, against a total us fleet running at 60-70. the grid application list is now seven times the whole country's demand. compute stopped being the constraint; physics took over.

read the note →
10:09 IST

Microsoft's telemetry on copilot's agentic coding

microsoft's telemetry on copilot's agentic coding: 3.2 million users, 13 million sessions, and 761 million llm calls in a single june week. that's not a pilot; that's a runtime. the agent workload just became the biggest api consumer most companies will ever run.

read the note →
10:09 IST

The silicon data llm token spend index fell below $1 per million for the first time, and september…

the silicon data llm token spend index fell below $1 per million for the first time, and september alone crushed frontier api prices by 60-80%. when tokens approach zero, the market stops selling tokens and starts selling outcomes. the price war is the product changing shape.

read the note →
09:47 IST

Tencent open-sourced browserskill, a mit-licensed bridge that lets claude code, cursor, and codex…

tencent open-sourced browserskill, a mit-licensed bridge that lets claude code, cursor, and codex drive your real authenticated browser instead of a headless one. the login state you already trust is now the agent's environment. the browser session just became the api.

read the note →
09:47 IST

Gitlab cut 14% of staff and exited 22 countries to fund ai infrastructure, then reported q1 revenue…

gitlab cut 14% of staff and exited 22 countries to fund ai infrastructure, then reported q1 revenue up 23% at 88% gross margins. the company wasn't in trouble; it was reallocating. layoffs are becoming a capex line item.

read the note →
09:24 IST

Aws shipped native vector search ga in dynamodb, so embeddings now live beside your operational…

aws shipped native vector search ga in dynamodb, so embeddings now live beside your operational rows instead of in a separate vector database. the dedicated vector db pitch just lost its simplest customer. search is a feature again.

read the note →
09:24 IST

Openai open-sourced codex harness, the runtime that actually runs codex — sessions, context, tool…

openai open-sourced codex harness, the runtime that actually runs codex — sessions, context, tool wiring — so anyone can build their own coding agent on it. the agent business is splitting into model, harness, and everything else. the harness is where the moats are being redrawn.

read the note →
08:58 IST

Aws now runs automated agent evals for bedrock agentcore inside github actions, blocking pull…

aws now runs automated agent evals for bedrock agentcore inside github actions, blocking pull requests when agent behavior regresses. agents finally get the same discipline as the code that calls them. the review gate is the unit of trust.

read the note →
08:58 IST

Hassabis wants a us-led frontier ai evaluation body where models submit to a 30-day pre-release…

hassabis wants a us-led frontier ai evaluation body where models submit to a 30-day pre-release review with a blind test bank before they can launch. the labs that keep shipping first are being asked to submit to a referee with no incentive to be kind. regulation is coming through evaluation, not legislation.

read the note →
08:35 IST

Acrab's gelix 1 is a 5nm edge chip that runs up to 100b-parameter models locally, no cloud…

acrab's gelix 1 is a 5nm edge chip that runs up to 100b-parameter models locally, no cloud round-trip. the privacy excuse for shipping every prompt to a datacenter just lost its hardware argument. local inference stopped being a compromise.

read the note →
08:35 IST

Anthropic open-sourced the shared-memory design behind its internal 30,000-agent fleet, where every…

anthropic open-sourced the shared-memory design behind its internal 30,000-agent fleet, where every branch thread syncs specs, decisions, and team preferences in real time. parallelism without shared state was the wall, and this is the fix. the agent that remembers becomes the one you trust.

read the note →
08:16 IST

42% of developers now say ai writes at least half their code, and 78% of ctos increased spending on…

42% of developers now say ai writes at least half their code, and 78% of ctos increased spending on review and qa to keep up. the hours ai saved came back as review hours. the delegate-to-verify ratio is the new productivity metric.

read the note →
08:16 IST

Arena's harness tax study

arena's harness tax study: running the same model through different agent harnesses barely moves success rates but can double the bill per task. the orchestration layer is now a 2x tax with no quality upside. the wrapper is the new legacy code.

read the note →
07:52 IST

89% of enterprise ai agent pilots never reach production, and 61% of failures trace to scope creep…

89% of enterprise ai agent pilots never reach production, and 61% of failures trace to scope creep and data quality. the agent worked; the job around it didn't. the bottleneck is the boring infrastructure nobody funds.

read the note →
07:52 IST

Openai released gpt-oss-120b and gpt-oss-20b under apache 2.0 — its first open weights since gpt-2,…

openai released gpt-oss-120b and gpt-oss-20b under apache 2.0 — its first open weights since gpt-2, closing a five-year closed era. the same company selling agentic subscriptions just gave away its flagship-class models. the moat was never the weights.

read the note →
07:33 IST

Tencent's workbuddy now turns one natural-language prompt into a full web app with cloud database,…

tencent's workbuddy now turns one natural-language prompt into a full web app with cloud database, file storage, auth, and ai built in, deployable to a shareable link. the unit of software stopped being the feature and became the whole product. the deploy button is the only skill that still matters.

read the note →
07:32 IST

Alibaba's qoder says it hit 6m users and 100k enterprises in one year, upgrading from coding ide to…

alibaba's qoder says it hit 6m users and 100k enterprises in one year, upgrading from coding ide to agent workbench. the editor-as-agent transition is already a distribution story, not a demo. the fastest-growing coding tools now ship a workbench, not a keymap.

read the note →
07:21 IST

Emulate, a uk ai startup founded a month ago by ex-deepmind researchers, is raising $700m at a…

emulate, a uk ai startup founded a month ago by ex-deepmind researchers, is raising $700m at a ~$3.7b valuation. a company with no product history is worth more than most public dev tools. the price of ai talent just became the valuation.

read the note →
07:21 IST

White-hat researchers used anthropic's claude to get into an openai employee's chatgpt account and…

white-hat researchers used anthropic's claude to get into an openai employee's chatgpt account and read openai's private code caches — openai paid them $6,500. the tools you ship can be used against you by the same agent vendors' models. security teams are now on both sides of the same agent.

read the note →
06:48 IST

Anthropic says claude leads 26% of its ai research work, up from 1% in march, and collaborates on…

anthropic says claude leads 26% of its ai research work, up from 1% in march, and collaborates on over 90%. the lab's own roadmap now runs through its model. the next question is who reviews the reviewer.

read the note →
Founder Notes archive, page 28 — Yethikrishna R