In the Arena #5: when should the agent ask a human?
Chatbots for everyone changed very little. Four reads on where people belong once agents do the work.

A note from Alex
Alex here. Most of the operators I talk with have already done the obvious thing. They gave the whole company a chatbot. Almost none of them would call that a transformation, and this week's reads explain why.
Benedict Evans starts with recognizing the task worth changing. Ethan Mollick asks when an agent should involve a person. Palo Alto Networks' incident team describes an intrusion where agents carried out attack steps while a human retained consequential decisions. The final read asks what an AI vendor means by an outcome.
If you are deciding where people belong in a workflow you are about to automate, tell me how you are drawing the line — email aschreiner@comavenai.com or grab thirty minutes with me. No deck, no pitch, just a working conversation.
- Alex Schreiner, Head of Growth
The short version
- Find the task worth changing, then work through how people will adopt the change.
- Design human involvement around approval, expertise, varied ideas, and meaningful decisions.
- Agent-assisted attacks can move quickly. Keep an inventory of your AI endpoints and access keys.
- Vendors are moving to pay-per-outcome. Read the definition of “outcome” before you sign.
Worth your time
The hard part was never the tool
The difficult part is finding the task worth changing.
Benedict Evans argues that useful AI work starts before code. People focused on cases, clients, or operations may not recognize an automation opportunity. Even when they do, improving it can require changing how several teams and systems work together. A working prototype still needs people to use it.
His September 3 essay asks three questions: how a company will buy, build, and deploy AI; how operations change; and what happens to economics or competition. The forward-deployed engineer helps connect technical possibility to the work people actually do.
That is the job we hire for, and the walking-around part is most of it. Here is how we describe it.
The tool is the easy part.
Benedict Evans, “AI, tools and transformation” (free, no account required).
When should the agent ask a human?
For approval, expertise, varied ideas, and decisions people find meaningful.
Ethan Mollick's August 31 essay uses the Hugging Face Incident to ask when agents should involve humans. The test agents coordinated their work but had no path to ask a person for help.
He proposes a facilitator that recognizes four reasons to involve people: consequential approvals, gaps in expertise, diversity of ideas, and the value of interesting work. Full automation can remove the very decisions through which people develop judgment.
We would put that list in the design review. Before an agent goes near a system of record, write down which trigger requires a hand-off, who receives it, and how that person answers. Those decisions belong in the workflow design.
Ethan Mollick, “Agency and Agents” · Dwarkesh Patel's account of the incident. Both are freely readable.
Two weeks of intrusion in less than ten hours
Routine weaknesses become more dangerous when attack steps run quickly.
Unit 42's September 2 investigation describes an agent-assisted intrusion involving more than fifty techniques in under ten hours, compared with roughly two weeks for human teams. It progressed from a public web service through exposed credentials and cloud infrastructure, without needing a novel zero-day. Branch protection stopped attempted Terraform backdoors.
The team's four recommendations: coordinate containment across systems; inventory and govern AI infrastructure; watch for unusual automated activity; and require multiple reviewers and protected branches for infrastructure code.
If your company has adopted AI tools faster than it has inventoried them, start with the second one. Knowing what exists and who can access it gives the rest of the response somewhere to start.
Unit 42 investigation, by Renzon Cruz, Nicolas Bareil, Eric Semaan, and Omar Jbari (free).
What should an AI invoice count?
Seats are being disqualified, consumption is preferred, and outcomes are where the market is heading. The contract question is what an outcome is.
Futurum Group surveyed enterprise software buyers in the first half of 2026: 43 percent prefer consumption-based pricing, 27 percent prefer outcome-based pricing, and fewer than one in five still prefer paying per user. Analyst Keith Kirkpatrick's summary: “Outcome-based pricing is becoming a market standard.” Customer support got there first. Intercom charges $0.99 per Fin outcome, and Zendesk charges for successful AI resolutions instead of AI seats.
Read the definition. Intercom counts an outcome when the customer confirms the issue is resolved, when Fin completes a workflow (including a handoff to a person), or, in Intercom's wording, when “They don’t ask for more help after Fin responds.” A customer who goes quiet counts the same as a customer who was helped, and so does a handoff. That is not a complaint about the price. It is the clause a CFO should read before signing.
Outcome pricing can move the cost of failed attempts from you to the vendor, which is the right direction. It also makes the vendor's definition of success your invoice. Ask who verifies it, how, and what you can audit.
Futurum Group, May 12, 2026 (Keith Kirkpatrick; free) · Intercom's pricing page (the definition is in the FAQ).
From the team
Justin and Alex wrote down the five AI projects we tell companies not to build. The project with no named outcome, the one built on a process nobody can explain, the one with unsafe write access, the integration problem wearing an AI costume, and the hardest problem first. A good no creates a better yes. Read the five
A free AI ROI calculator, no account required. Describe one recurring task in plain language and get a conditional estimate with assumptions you can question, including the work that stays with a person. Built for a solo owner, a nonprofit, or a team inside a larger company. Try it on one workflow or read why we start with one.
Missed last week? Issue 4 covered the twentieth run, which models businesses are actually buying, cost per accepted change, and how to get a team to think bigger.
See you in the arena.
Have a workflow that deserves better? Grab thirty minutes with us.
