Skip to main content
Newsletter
ROAI — Return on Applied AI, Edition 12

Nobody puts that in the slide deck.

McKinsey published something this week that's worth your time if anyone is currently pitching you AI agents.

They costed out what it actually takes to run one. Not the license. The whole thing.

For a customer service agent at a bank, the tokens come to about a quarter of the cost of running it. Everything else is people. Checking the work, catching the exceptions, deciding whether what the agent said was right.

Human oversight was 70 to 75 percent.

That should change how you hear the next pitch. When a vendor quotes you a price per conversation, they are quoting you the cheap part. Someone on your side still has to read what the agent said and decide whether it was right.

Nobody puts that in the slide deck.

There's a second number in there worth sitting with if you run more than a few locations. A conversational agent onboarding 2,500 new customers a year cost between $10,000 and $15,000. Doubling the customers only took it to $20,000.

Most of the cost is fixed. So the same agent that barely makes sense at one location can be obvious at thirty. Scale is not a nice-to-have with this stuff. Scale is the thing that makes the math work at all.

The number they say to watch isn't cost per conversation. It's what it costs to finish the whole job, people included, against what finishing it is worth to you. In their bank example, opening one account takes five to seven agents, a stack of older systems, and two to four teams of humans watching it.

That's still cheaper than what it replaced. It's just nothing like the pitch.

So ask a different question this week. Not what does it cost per conversation. Ask what it costs to finish the job, and who on your team has to check it.

Brian Holmes
Brian Holmes
CEO & Founder
Applied AI Series

Learn how to actually use AI in your business.

Two practical sessions for people building useful AI into the work.

Make on-brand creative with Claude Design
AI for Marketers
Wednesday, September 16 · 12:00 PM CT
Make on-brand creative with Claude Design
If you need more creative, aren't using AI to make assets, or don't have a Design System set up in Claude for your brand yet, don't miss this.

Your marketing person will thank you for showing this to them.

Plus, we'll send them the post-session guide if they register.
Save Your Seat
AI for Builders
In case you missed it
Start using open source models
If you weren't able to make it to this month's session, don't worry — we made a guide and recorded everything. This is your shortcut to using an open source model same-day.
View Guide & Recording
Worth Sharing

The stories worth carrying into next week.

Practical signals about what it takes to make AI work after the pitch.

Study
The people running agents in production keep them on a short leash
Researchers at Berkeley ran 20 in-depth interviews and surveyed 86 agent systems actually running in production, across 26 industries.

68% stop the agent after ten steps or fewer and hand back to a person. 74% rely primarily on human evaluation to judge whether the work was any good. Reliability, not cost or talent, was the top challenge they named.

Our Take: This is about where we are at too in our daily work. It is kinda scary that 26% don't primarily rely on human judgement. That's where workslop is coming from.
Read: Berkeley
Story
How your vendor defines "resolved" is where the money moves
A benchmark of 18 AI agent vendors found that outcome-based pricing only works when the definition of the outcome is honest, and it usually isn't. Some vendors bill an assumed resolution when a customer stops replying, which counts the people who gave up. One vendor's own billing docs confirm an AI-resolved conversation gets charged twice, as a ticket and as a resolution, unless the customer reaches a human within 72 hours.

Our Take: Get the definition in writing before you sign anything. Ask what counts as resolved, who decides, and what happens when the customer just stops answering. If the vendor writes the definition and also sends the invoice, you are not buying outcomes, you are buying their opinion of one.
Read: Aissist
Survey
People are buying less from texts, and tolerating fewer of them
EZ Texting's 2026 consumer report found the share of people who say they've purchased in response to a business text dropped from 49 percent to 36 percent. Fewer consumers say they're texting more this year than said so last year. Speed expectations went the other direction: over 80 percent read a business text within 15 minutes, and about the same share reply that fast.

Our Take: People still read your messages. They just do less about them. Speed is the one thing still working in your favor, and most businesses waste it. Check how long it takes yours to answer a new lead. If it is measured in hours, that is the number to fix before you touch anything else.
Read: EZ Texting
Numbers
It takes 17% more messages to get a booking than it did last year
Across the businesses running our agents, it now takes 17% more messages to land one booking than a year ago. That tracks with the survey above — people are reading just as fast and acting on it less.

Whether that costs you anything depends on what your vendor charges for. Published rates run about $3 to $25 per qualified lead and $8 to $40 per booked appointment. Priced per message instead, it's a penny and a half for a raw text and up to 75 cents once an AI platform is writing them.

Billed per lead or per booking, the extra messages are free to you. Billed per message, you're paying 17% more for the same number of bookings.

Our Take: Work out what you pay per booking, not per message. Take last month's bill, divide it by bookings, and compare it to what a booking is worth to you. That one number tells you more than any vendor's rate card.

Source: AI Front Desk Internal Reporting 2025-2026
Try This
Stop reading all of it
If oversight is most of what an agent costs, the obvious lever is reading less of what it does. Banks running agents in production don't review everything. Expert review sits at 10 to 20 percent of runs, chosen deliberately rather than at random.

Pick 20 of last week's conversations. Not the first 20, and not the ones that went wrong. Take every conversation that ended without a booking, plus five that ended with one. The failures tell you what to fix, and the successes tell you what not to break.

Reading everything feels responsible and mostly it's expensive. If someone on your team is reviewing every transcript, they are spending real hours confirming that most things went fine. Sample instead, write down what you find, and check the same 20 next week.

Then work out what that review time is actually costing you. We built a calculator for it: vendor cost, checking cost, and what share of your spend is the person doing the checking.

We've spent thousands of hours doing this on our own agents. Reading conversations, finding where they go wrong, and fixing them. That's the part we take on so our customers don't have to, and it's why the checking cost shows up in our column instead of theirs.
Work out your cost per booking
The AI Front Desk team

That's it for today! See you soon.

Brian, Murphy, John, Beth, and Jake — some of the humans behind AI Front Desk.

Not subscribed yet?

Get ROAI in your inbox every week.

Subscribe →
ROAI Newsletter · Practical AI, every week
Get practical AI tips that actually move the needle.
No spam. Unsubscribe anytime. Privacy Policy.