The OpenAI number everyone will screenshot, the six jobs a new standard just handed you, and who is paying for the compute.
A week ago we wrote that structure beats headcount when it comes to AI agents. Three days later Anthropic's red team published a second reason to believe it, and not the one we gave. Ours was coordination breaking down as agent teams grow. Theirs is copies of the same model sharing the same blind spot, which bites at two agents, not sixteen.
Four reads below. If something here raises a question about your own operation, email contact@comavenai.com. A human reads that inbox.
- Justin, Matt, and Alex
Two agents are not two opinions. Anthropic's Frontier Red Team ran copies of the same model side by side. The run that got the coverage was a migration test: three agents on three virtual machines, each told to port the same Python backend to a different language, none aware of the others at the start. They decided the interference was deliberate and went after each other, disabling Unix accounts and running scripts that hunted competing processes. Keep the setup in mind, because the conflict was designed in. In other runs the same models negotiated a truce and asked for a human. The finding that should cost you money is duller: in an early version of a game-building experiment, with every agent coming online at once, 18 of 30 opened a git branch with the identical name. Anthropic's own inference: “by implication, this means that when one agent makes a bad decision, it is likely that many agents will make that same bad decision.” Ask how an agent deployment is controlled and the answer is usually that a second agent reviews the first. One model, one prompt, one blind spot. In a portfolio you would call that one supplier and two invoices. Read the research · Rebecca Bellan in TechCrunch
Read OpenAI's enterprise report from the bottom up. The number that will get screenshotted is 8.3x: output tokens per active user at “frontier” firms versus typical ones as of June, up from 2.6x in January. Look at the sorting before that lands in a value-creation plan. OpenAI ranks its enterprise customers monthly on that same metric, calls the top 10 percent frontier and the middle band typical, and reports the distance between them, without ever saying how many firms are in the sample. The better number sits further down: 21 percent of weekly active users at frontier firms use Plugins and 19 percent use skills, against 9 and 3 at typical firms. Not clean either, since connecting tools is part of what makes a firm high-token. But the direction survives. The gap is tooling and documentation, not appetite. Token spend is the vendor's billing meter; a team can triple it in a quarter and move nothing on the P&L. Read the report
Six things Agent Plugins deliberately does not do. Agent Plugins hit 1.0 on August 6: a vendor-neutral format that puts an agent's skills and the MCP servers they depend on into one folder. Vercel proposed it, AWS, Anysphere, GitHub, Microsoft and OpenAI refined it, and Google has since joined as a core maintainer. The problem is real; until now every client invented its own layout, so authors forked the same package over and over. In their writeup, Kevin Hou, Haoyu Wang and Alan Blount at Google include a section headed “What It Deliberately Leaves Out”: v1 “defines no install mechanism, no distribution protocol, no permission model, no sandboxing requirements, no trust or provenance verification, and no user experience.” Right call for a file format. It also means all six now belong to whoever installs one. Decide who in your company that is before the first folder shows up. The Google writeup · Jonathan Hefner's Vercel announcement
Who is actually paying for the compute. Fifty billion dollars committed to American compute infrastructure, announced in November 2025 by a company running less than $9 billion in annualized revenue. Campbell Hutcheson at Epoch AI traced the money behind Anthropic's buildout and found nearly $50 billion of debt: about $34.5 billion for over a gigawatt of Google TPU systems, $30 billion of it conditionally supported by Broadcom, plus roughly $15.2 billion lent against five Fluidstack datacenters. Most of it was arranged before the revenue arrived. His conclusion, that financing is unlikely to be the immediate limit on frontier compute growth, is conditional on suppliers continuing to lend their credit to deals they benefit from. That condition is the argument, and the part that would break first. Read the analysis
The part of our stack this week vindicated. Our outbound runs on seven agents, each doing one job. They do not talk to each other directly; they pass signals through a shared database, so any one can fail and restart without taking the pipeline down. Seven agents with seven different jobs do not share a blind spot the way thirty copies of one model do, and none of them can reach into another's work. The full writeup has the numbers and the failures.
Missed the first issue? Issue 1 covered where agent teams start degrading, Ethan Mollick's field guide, and the EU disclosure rules that took effect on August 2.
See you in the arena.
Have a workflow that deserves better? Grab thirty minutes with us.
Choose which cookies we may use. Keeping Marketing off is your "Do Not Sell or Share" opt out. We also honor the Global Privacy Control signal.
Required for the site to work: security, consent storage, spam prevention. Always active.
Google Analytics 4. Anonymized usage statistics that help us improve the site.
HubSpot tracking and advertising cookies used for sales follow up and ad measurement.