In the Arena · September 23, 2026

In the Arena #7: same model, twice the bill

The software around a model can double what it costs to run. Four reads on what AI really costs, who is already using it, and what agents do when nobody checks.

Two navy workstations run identical cyan modules and turn out the same finished card, but the bulkier station draws on a far taller stack of paper.

A note from Alex

Alex here. Two questions come up in almost every conversation I have with operators right now. What will this cost to run? And who in the company is already using AI without telling us? This week's reads answer both, and neither answer is really about the model.

Arena's research team found the same model costing up to five times as much, depending on the tooling wrapped around it. Deloitte UK found workers paying for AI tools out of their own pockets, often without their employer knowing. Researchers say OpenAI's own agents attacked a public software registry, and nobody told the registry. And Ethan Mollick argues the biggest gap is how people use the models that already exist.

If you are trying to put a number on an AI workflow, or a policy around one, tell me where you are stuck. Email aschreiner@comavenai.com or grab thirty minutes with me.

- Alex

The short version

  • The same model can cost up to five times as much depending on the harness around it. Price the whole loop, not the model.
  • UK workers spend an estimated £958 million a year of their own money on AI for work. Give people a sanctioned tool before they choose one for you.
  • Researchers say OpenAI's agents uploaded malicious packages to RubyGems. If your agents can reach the internet, log what they do.
  • Ethan Mollick: the gap is between what models can do and what people do with them.

Worth your time

Same model, twice the bill

The tooling around a model can matter more to cost than the model you pick.

Arena's research team ran seven models through three coding-agent harnesses (Claude Code, Codex CLI, and Pi) on two public coding benchmarks. Their September 16 finding: the harness barely changed success rates, but “the same model can achieve similar success rates at up to 5x costs.” Claude Fable 5 solved 97.8% of attempts in Claude Code and 96.7% in Pi, yet cost about twice as much in Claude Code ($1.33 against $0.67). Pi, the cheapest, gives the model four tools: read, write, edit, and bash.

Siddharth Sambharia at Portkey named this the harness tax back in April: “every token your agent spends on itself before it spends a single token on your task.” In his test, Claude Code sent about 27,000 input tokens for a small script. Pi sent about 2,600.

These are coding benchmarks, but the lesson travels. When we price an agent workflow, we price everything the harness sends on every turn, not just the model's rate card. Run the same task through two setups and compare the cost of a result you would accept.

Arena, “HarnessTax: How Much Does the Harness Matter for Coding Agents?” (Arena Team; free) · Siddharth Sambharia, “The Harness Tax” (Portkey; free).

Who is paying for your team's AI?

In the UK, a lot of employees are, out of their own pockets.

Deloitte UK surveyed 25,000 workers through Ipsos UK between May 7 and June 10, 2026. It estimates they spend £958 million a year of their own money on generative AI tools for work. 63% say they knowingly use it for work. Nearly a third of users do so without their employer's knowledge. The average reported time saved is 70 minutes a week.

Hayley McKelvey, Deloitte UK's chief AI officer: “UK workers are showing they don't want to wait for permission to use GenAI.”

The survey is British. We would not bet on a US workforce being different. A policy that only says no mostly moves the usage out of sight. Pick approved tools, say what data can go into them, and train people. Then the 70 minutes shows up somewhere you can measure it.

Deloitte UK press release, September 16, 2026 (free).

What were the agents doing on RubyGems?

Researchers say OpenAI's agents uploaded malicious packages to a public software registry, and the registry was never told.

On September 11, Spencer Kitts, Thomas Larsen, and Sydney Von Arx published a report attributing hundreds of malicious packages uploaded to RubyGems on May 11 to internal OpenAI agents. “The agents clearly regarded what they were doing as hacking,” they write, pointing to file names like exploit.rb. People in the RubyGems community told them OpenAI never informed them.

OpenAI confirmed the incident to Reuters and said its agents “used the RubyGems platform to access the internet to carry out benign tasks and retrieve public information,” adding that its review continues.

We are not in a position to settle whose account is right. The lesson holds either way. If a frontier lab's own agents can do this without anyone disclosing it for four months, a company giving agents internet access needs a record of what they did, limits on what they can touch, and a person who reads the record.

Kitts, Larsen, and Von Arx, “OpenAI agents carried out an undisclosed cyber-attack on RubyGems” (free) · Reuters, via BNN Bloomberg (free).

The gap is in how people use it

Ethan Mollick calls it the capability overhang.

In his September 18 essay, Mollick describes “the gap between what these models can do and what almost anyone is doing with them.” The people who close it bring four things: deep knowledge, wide knowledge, taste, and agency. “You are not trying to compete with AI in producing outputs, that is a losing game.”

His examples are personal projects, not company deployments. The point still carries. In our work, the constraint is rarely the model. It is whether the person who knows the job is in the room when the system gets designed.

Ethan Mollick, “The Overhang” (One Useful Thing; free).

From the team

Alex wrote for companies that already have a capable AI team. If your people know the operation and have a roadmap, the question is what another capable set of hands could move forward, with ownership staying inside. Read “Your AI team has a roadmap. Give them room to deliver.”

Map the handoff, not the screen. The expensive part of a workflow is often the moment work moves between people and systems. Justin and Alex lay out the six parts of a handoff worth mapping before you automate anything. Read the handoff guide

Missed last week? Issue 6 was Justin on why your business isn't the frontier, and how we pick a model for each job.

See you in the arena.

Have a workflow that deserves better? Grab thirty minutes with us.