In the Arena · September 30, 2026

In the Arena #8: two of 66 runs caught the $25.7 million

The best agent finished about a quarter of the finance assignments. Four reads on where agent work actually breaks, and where it is already cheap.

A long row of navy report folders rides a paper conveyor past open inspection frames; two cyan frames hold a magnifying lens over a page with a small ochre mark, while the rest pass through unchecked.

A note from Alex

Alex here. The pitch I hear most from operators right now is some version of “the agent does the whole job.” In the week of September 21, four sources put numbers on that sentence, and the numbers are more useful than the pitch.

Surge AI gave frontier agents 80 realistic finance assignments, and the best one finished about a quarter of them to a standard a professional could use. Anthropic sent 201 colleagues' agents into a market to trade on their behalf and found the agents negotiated fine; what they lacked was a good enough picture of the person. Toby Ord worked out what you actually buy when you add agents to a swarm. And Tomasz Tunguz found that a model which only picks from a fixed set of answers nearly doubled the accuracy of a classifier in one of his agents, at about a hundredth of the cost.

None of this argues against agents. It argues for being specific about which part of the job you are handing over. If you are scoping an agent project for the fourth quarter, tell me which part. Email aschreiner@comavenai.com or grab thirty minutes with me. I will tell you what we have seen hold up. No deck, no pitch, just a working conversation.

- Alex

The short version

  • The strongest agent completed 23.9 percent of Surge AI's 80 finance assignments. Two of 66 runs caught a $25.7 million pricing error in a fund report.
  • A model that returns only a probability over the answers you give it took one of Tomasz Tunguz's email classifiers from 47 percent accuracy to 80 percent, at about a hundredth of the cost.
  • Toby Ord, working from OpenAI's own charts: ten times the agents buys the gain of three to five times the tokens on one agent. Swarms buy speed, and bill you for it.
  • Anthropic's book-trading agents negotiated well. 85 percent of the shortfall came from what they did not know about their person.

Worth your time

“No other release blocker was identified.”

An agent held a fund report over $132,000 of errors and missed the $25.7 million one. Surge AI's new benchmark is full of moments like that.

In the week of September 21, Surge AI released DAYJOB, a benchmark family for what it calls economically valuable agents, starting with finance and healthcare. The finance set is 80 assignments written by people who have done the work, among them a senior portfolio manager at BlackRock and a finance professional who structured infrastructure and renewables debt at HSBC. The prompts read like a message from a colleague: 81 words on average, against 337 in OpenAI's GDPval. The agent is handed an average of 25.7 files and has to work out which ones matter. Surge estimates a human professional would need 21.6 hours per task, and the median task is graded against 57.5 rubric criteria.

The strongest model, Claude Opus 5.5, succeeds on 23.9 percent of assignments. GPT-6 Astra scores 21.5 percent, Claude Fable 5.1 19.8 percent, and the leaderboard runs down to a cluster at zero.

The failures are the instructive part. One task asked the agent to review a new client's first fund report pack before it went out. One position had been priced by hand from a Bloomberg screenshot that showed South African cents; the value had been entered as rand, so a roughly $260,000 position appeared to be worth $26 million. GPT-5.6 Sol found two smaller pricing errors worth a net $132,000, correctly recommended holding the reports, and wrote: “No other release blocker was identified.” Across 66 runs on 22 model configurations, two caught the cents-versus-rand error, both on Claude Opus 5. In another task, Grok 4.6 recommended a store expansion using a forecast that added returns and allowances to gross sales instead of subtracting them.

Two things to keep in view. Surge AI sells training data and RL environments to the labs, so a benchmark that frontier models fail is good for its business. And the score counts assignments the graders judged usable end to end; the same agents did the arithmetic and the extraction well. The gap is judgment: which file is wrong, which detail changes the decision. For anyone putting an agent near a close, a covenant test, or a client deliverable, that is the test to run before the pilot. Hand it your own pack with one planted error and see whether it finds it or signs off.

Surge AI, “DAYJOB: Can Agents Survive a 9 to 5?” (team post, no byline; free) · the DAYJOB: Finance leaderboard and worked examples (free).

The cheapest model in the stack only picks from the answers you give it

Tomasz Tunguz swapped deciders into a quarter of the if-then calls in one of his agents. Accuracy went up.

On September 15, TypeSafe, a lab that spent two years in stealth, released Jev, a model that does not generate text. You give it a question and the set of allowed answers, yes or no, or a multiple choice, and it returns a calibrated probability for each in hundreds of milliseconds, priced at 4.2 cents per million input tokens with no charge for output. (TypeSafe says the price may be subsidized.) Tunguz calls it “a robust if-then decider” and works out that on a typical classification call it is 82 times cheaper than a Sonnet-class model. SemIf, an open reproduction of the pattern, appeared within three days, and kev followed.

The useful part is what he did with it. Tunguz went through one of his own agents looking for the if-then decisions he had handed to a general model, and replaced about a quarter of those calls with the deciders. On 98 hand-verified production email threads, the production classifier had been right 47 percent of the time (46 of 98). Jev scored 80 percent (78 of 98) and the local SemIf model 82 percent. In live logs across 31 inbound emails, the local decider acted on 8 with zero errors and deferred the rest to the frontier model.

We read this as the start of a split most operators already feel: frontier models for discovery and design, narrow models for the thousand daily judgment calls once the workflow is settled. Is this email urgent. Does this ticket need a person. Is this invoice a duplicate. Those are not reasoning problems, and paying reasoning prices for them is where a good part of an agent's monthly bill comes from. The catch is the translation work. A fuzzy question like “is this urgent” has to be broken into checks a decider can score, and someone who knows the job has to write them. That is where the time goes, and it is time well spent.

Tomasz Tunguz, “AI Comes for the If Statement” (free) · TypeSafe's Jev page (vendor site; free).

Ten times the agents is worth three to five times the tokens

Toby Ord read OpenAI's own charts. A swarm buys you time, and bills you for it.

Toby Ord, who has charted how AI capability scales with compute in earlier essays, took OpenAI's launch charts for GPT-5.6 Sol and asked a question the charts did not answer directly: what do you get from adding agents to a swarm, holding everything else fixed? His answer, published September 21, borrows a parameter economists use for human teams, the “stepping on toes” factor. On the three benchmarks OpenAI reported, it comes out between 0.48 and 0.68 (regressions Ord had Claude Opus 5 run). In plain terms, multiplying the number of agents by ten gets you the gain you would get from giving one agent three to five times as many tokens. The shortfall compounds: to match a 100-fold increase in one agent's budget, you would need 900 to 15,000 agents.

So why run a swarm at all? Speed. A four-agent swarm needed about twice the total tokens of a single agent to reach the same score, but each agent needed half as many, and they run in parallel. Sixteen agents: in theory, roughly a quarter of the time for four times the cost, and Ord notes the real speedups fall a little short of that. He quotes Noam Brown on the Dwarkesh podcast making the same trade: “you're paying a 2x more to get an answer twice as quickly.”

The buying lesson is direct. “Multi-agent” in a vendor deck sounds like a claim about accuracy. On Ord's numbers the demonstrated gain is wall-clock time, with a compute bill attached. If the task has a deadline that matters, the trade can be worth it. If it does not, one agent with a bigger budget is the cheaper route to the same answer. Our own outbound runs on a swarm, and it is worth being precise about why: it is a pipeline of specialists, each doing one job and handing off to the next, not sixteen copies of one agent attacking the same task. Justin's writeup of what it actually produced, failures included, is here.

Toby Ord, “Swarm Scaling” (free).

The agents negotiated well. They did not know their people.

Anthropic sent 201 employees' agents to trade books. Most of the shortfall was in the brief.

Anthropic published a small experiment on September 24. This summer, 201 employees across six offices each brought in a book to give away, had a short chat with Claude about what they like to read, and sent a Claude agent onto a trading floor to swap for something better. Separately, participants ranked ten books by hand, a ground truth the agents never saw.

The agents' picture of their person was decent for a five-minute chat: on 61 percent of book pairs, the agent's ranking agreed with the person's, where guessing is 50 percent and ranking by popularity gets 53. But decent was the ceiling. Judged on their own rankings, participants ended up at 0.55, roughly their fifth choice of ten, where the best possible outcome was 0.89, about their second. When Anthropic decomposed the gap, 85 percent of it came from the agents' imprecise picture of what their person wanted and 15 percent from the trading itself. Once the rankings were noisy, swapping in a centralized matching rule barely moved the result. Effort did: the median participant typed 216 words in the intake chat, and writing about 300 words instead of 150 predicts four points more agreement.

We see the same pattern outside book swaps. When an agent's output disappoints, the instinct is to blame the doing, the prompt or the tool. More often the agent executed a brief that was never good enough to execute. Before an agent takes over a task, write down what you want the way you would for a new hire on day one, including the things you assume are obvious. Then check what it heard. Anthropic showed participants a summary of their intake; those who said it missed nothing would hand an agent 34 percent of their book budget, and those who said it missed something, 23 percent. The people whose brief had landed were the ones ready to delegate.

Anthropic, “Project Swap: What happens when agents trade for us?” (Zoë Hitzig and six co-authors; free).

From the team

Build, buy, or automate inside the stack you already have? Justin and Alex on the question most teams ask too late. Start with the work that should stop taking so much time, then pick the smallest approach that removes it safely, with a neutral decision matrix for buy, connect, configure, build, or map first. Read the decision matrix

Missed last week? Issue 7 went up on September 23: the harness tax on the same model, UK workers paying for their own AI tools, what OpenAI's agents did on RubyGems, and Ethan Mollick on the capability overhang.

See you in the arena.

Have a workflow that deserves better? Grab thirty minutes with us.