Skip to main content
Newsletter
ROAI — Return on Applied AI, Edition 09

Moment of brilliance, then trash.

I had AI put together two market reports for me last week. Both came back sharp — well-organized, accurate, genuinely useful. I forwarded them to the team feeling pretty good about where we'd landed with this stuff.

Then I ran the third one. Same prompt, same model, same everything.

Garbage. Confidently stated numbers that didn't hold up, a structure that looked right but led nowhere, and conclusions I wouldn't have signed off on. I spent more time picking it apart than I would have spent doing it from scratch.

The frustrating part wasn't the bad result. It was that two good results had already made me lower my guard.

There's a Harvard study in this edition that explains exactly what happened — not to me specifically, but to a whole group of consultants who got worse results when they used AI, while their colleagues got better ones. The difference wasn't how smart they were. It was what they expected the tool to do for them.

Two good results don't tell you the third will be good. The output is variable. That's just true, and it's not going to change anytime soon — there's a research piece in this edition explaining why, at a technical level, the same prompt genuinely does give you different answers.

My recommendation: treat AI output as a range, not a verdict. The third report wasn't a fluke — it was the range showing up. Build your review step accordingly, every time, not just when something feels off.

Brian Holmes
Brian Holmes
CEO & Founder, AI Front Desk
Stories Worth Sharing

Here are the best insights our team found from across the web this week.

There is more AI news than any operator has time for. So we read it so you do not have to.

Study
The 19% who did worse
A Harvard/SSRN study gave a group of consultants access to AI for their work. Most performed better. But 19% — one in five — performed worse than colleagues who never used the tool at all. The researchers found that this group applied AI to tasks where it was poorly suited, over-trusted the output, and stopped applying their own judgment at the exact moments it was most needed.

Our take: The frontier is invisible from the outside. Two good results in a row don't tell you where the edge is. If you can't explain why AI did well on a task, you can't predict when it will fail — and that's when the 19% happen.
Read: The Harvard/SSRN Study
Try This
Run your own job through it, one task at a time
JobsGPT, built by the Marketing AI Institute, breaks any role into its component tasks and rates each one by AI exposure — how much of that specific task AI can realistically handle today. You can run it inside ChatGPT. It takes about ten minutes and gives you a map of where AI genuinely helps versus where you're still the irreplaceable part.

Our take: This is the practical answer to the study above. Instead of wondering whether AI is right for your job in general, find out which specific tasks it's actually suited for — and build your workflow around that honest assessment.
Try: JobsGPT — Marketing AI Institute
Research
The same prompt really does give you different answers
Thinking Machines Lab published research on non-determinism in large language models — the technical reason why identical prompts produce different outputs. The short version: how requests get batched and processed on the backend introduces variation that isn't visible to the user and can't be controlled from the prompt side.

Our take: Inconsistency isn't your imagination, and it isn't something you're doing wrong. The model genuinely produces a range of answers to the same question. Plan for that range — review critically every time, not just when something feels off.
Read: LLM Non-Determinism — Thinking Machines Lab
Guide
Two prompting handbooks, from companies that compete with each other
Google published a Gemini prompting guide for Workspace users. Anthropic published their best practices for prompt engineering with Claude. They were written independently, by teams that are competing for the same customers. When you read them side by side, they land on the same four moves: give context, be specific about format, show an example, and iterate rather than give up after one try.

Our take: Where two competitors agree is what to trust. These four moves aren't marketing — they're what actually works, confirmed independently by both sides of the market.
Read: Google Prompting Guide (PDF)
Survey
Half of American workers never use AI at work
Gallup surveyed roughly 22,000 American workers and found that 49% never use AI at work. 12% use it daily. The daily users aren't necessarily in tech — they're spread across industries and roles. What they share is that they kept going after the first bad result and figured out what the tool is actually good for.

Our take: The 12% isn't a different kind of person. They just didn't stop after a result like that third report. Joining them costs nothing except a willingness to keep going when the output is garbage — and to review it carefully when it isn't.
Read: Gallup AI at Work Survey
The AI Front Desk team

That's it for today. See you soon.

Brian, Murphy, John, Beth, and Jake — some of the humans behind AI Front Desk.

Not subscribed yet?

Get ROAI in your inbox every week.

Subscribe →
ROAI Newsletter · Practical AI, every week
Get practical AI tips that actually move the needle.
No spam. Unsubscribe anytime. Privacy Policy.