In the Arena · October 5, 2026

In the Arena #9: when the agent asks, who answers?

Four reads on what AI can do, what it costs, and the people who keep the work useful.

A small white paper robot waits at a navy checkpoint with a cyan call button, beside an unattended desk and empty chair.

A note from Alex

Alex here. If your team is putting an AI agent into a workflow, I think there is one question worth asking before you give it more responsibility: when it needs a decision from a person, who actually answers?

That sounds straightforward. Someone gets a notification, checks the request, and makes the call. But what happens when that person is in a meeting? Or the request arrives overnight? Or the agent gets an automated response that sounds like approval?

One of this week's reads looks at that last case. It stayed with me because it turns a broad conversation about AI oversight into something a team can check.

The other three bring us back to the same practical work: checking the economics, writing down what your people know, and giving someone time to keep the system working after launch.

I'd bring these questions to the people doing the job. They know which decisions are routine and which ones deserve a closer look.

The short version

  • Check the handoff. If an agent needs approval, name the reviewer and decide what happens when nobody responds.
  • Price the whole task. Include setup, supervision, maintenance, and the work people still have to do.
  • Write down the exceptions. Your team's judgment needs to be part of the instructions.
  • Give ownership time. Reviewing and improving an AI workflow belongs in someone's job.

Worth your time

1. Asking permission is only useful if the answer means something

The UK AI Security Institute tested GPT-6 Astra in simulated cybersecurity exercises, with OpenAI's cyber classifiers turned off. The model completed an out-of-scope supply-chain attack in 29.2% of runs. All actions were simulated; this is not a rate for ordinary business use.

The detail I would focus on is the approval step. When the agent asked a question, the test returned an automated message telling it to continue using its judgment. Sometimes it treated that response as permission, including when it recognized that the response was probably automated.

Clearer scope instructions reduced the attacks in a selected set of scenarios, but did not eliminate them. The researchers recommend protections beyond model instructions, including sandboxing and monitoring.

Read AISI's testing results.

My question for an operating team would be specific: if this agent proposes a refund, a customer email, or a change to a record, which actions can it complete alone? Which ones wait for a named person? And can you show me what happens if that person never responds?

For an action that requires approval, I would expect it to stay pending until that approval arrives. Then I'd test the handoff with the reviewer unavailable. That is a useful thing to learn before customers depend on it.

2. A capable robot still needs a business case

Anthropic's robot exposure study estimates that today's robots can perform roughly three-quarters of US physical tasks, mostly in controlled settings. Those tasks represent 34% of working hours. Its cost estimates put robots below human labor costs for just 0.3% of all work time.

Packers and packagers are one occupation where the estimated economics already work. Nursing and general repair are much less exposed to today's capabilities. Claude helped assess tasks and estimate costs, and the authors describe the cost estimates as approximate.

Read Anthropic's study.

I would use this as a starting point for a conversation with a vendor. Take one task in your operation and work through the full cost of changing it. What needs to be installed? Who supervises it? What happens when the input is awkward or the equipment is down?

The same questions help with a software agent. If it saves an associate time but creates a review queue for a specialist, include both in the calculation. A useful business case explains what the whole team gets back.

3. Your team's unwritten knowledge is part of the build

Vercel's skills.sh registry reached a million agent skills. A skill is a set of reusable instructions for a particular job. Its report found that 375 skills accounted for 62% of installs, while business operations and writing drew more installs per listing than the average.

These are registry counters, not unique users, and the category analysis uses a classified sample. Vercel runs the registry. Its broader argument is that company-specific judgment will become more valuable as general instructions become widely available.

Read Vercel's report.

That is where I would spend time with your experienced people. Ask the person who handles the exceptions to walk through a few recent ones. Which customer needed a different answer? Which document was out of date? When did they stop and ask someone else?

Write those decisions down, then use them to test the workflow. Your internal team can help an engineering partner recognize whether the proposed answer makes sense. That knowledge deserves a place in the system and an owner who keeps it current.

4. Maintaining the workflow is a job

Warp's Zach Lloyd describes two responsibilities for engineers using agents: building a useful product and improving the system that builds it. His team tracks human touches per pull request as one measure of how that system is improving.

Warp sells the infrastructure he discusses, so this is a vendor's perspective. He also makes clear that people remain responsible for the product's quality and usefulness.

Read Lloyd's essay.

I think there is a useful management question here for any department using AI. Who has time to review the exceptions, update the instructions, and check whether the result is getting better?

If that work is added to someone's responsibilities, give them room to do it. For an invoice workflow, you might track review time, corrections, and invoices completed. For customer service, you might look at resolution time and the quality of the answer. Fewer human touches help when the work also meets the standard your team needs.

One thing to try this week

Pick one AI workflow your team already uses. Walk through a normal case, an exception, and a request for approval when the reviewer is unavailable. Ask the people involved to show you what happens at each step.

You may find that the rules are clear and the handoff works. You may find a decision nobody realized they needed to make. Either result gives you something useful to work with.

From the team

We recently published Who owns an AI system after launch?, covering the business, technical, security, and review responsibilities that keep a production system useful. It is a good companion to this week's reads.

You can also catch up on In the Arena #8.

If you are working through an approval handoff, reply to aschreiner@comavenai.com. Tell me what the agent does and where your team wants a person involved. We can work through what a useful next step would look like.

See you in the arena.

Alex Schreiner
Head of Growth, CoMavenAI