We use AI for a lot of the work now. Research, first drafts, pulling figures out of ad accounts and analytics, building reports. It is quicker than doing it all by hand, and a lot of the time the result is better. But there is one rule we do not bend: a piece of AI work is not finished until someone has checked it. This week we caught a handful of small mistakes in work AI had produced. None of them were dramatic. Every one of them would have looked careless if it had gone out.
We are also noticing the same kinds of mistake going live elsewhere, on websites, in emails and in adverts, and we think trust is a big part of why.
The mistakes are small, and that is the problem
Nobody needs telling to check an answer that is obviously wrong. The mistakes that get through are the ones that look fine at a glance. These are the kinds we see most often:
- Text that appears twice. A sentence, a paragraph or a whole section pasted in twice. It reads normally until you reach the repeat, and by then a customer has noticed it too.
- Symbols that change on the way. A pound sign or a bullet point that turns into a jumble of characters when text moves from one app or system to another. It looks right where it was written and wrong where it lands.
- “Checked against the data” when it was not. The AI says a figure has been checked, or recommends an action because the data supports it. Sometimes it checked an old saved copy instead of the live account. Sometimes it did not check at all. The sentence reads the same either way.
- Times and dates slightly out. The wrong time zone, a date range a week off, “last month” meaning a different month from the one you meant.
- “Done” when it is only half done. A task reported as complete when only part of it happened, or when something slightly different from what was asked was done instead.
“Checked” is a claim, not a fact
AI tools write “verified”, “confirmed” and “based on the data” as fluently as they write anything else. Those words are part of the output. They are not proof that anything was checked.
So when a tool tells you a figure is right, or that the numbers say you should raise a budget, pause a campaign or change a price, treat that as something to confirm. Open the source yourself: the live ad account, the analytics report, the order, the page as a customer sees it. And check it by a different route from the one the AI took. If it pulled the number from a report, look at the account. If it read a saved file, look at the live page. A check that repeats the same steps will usually repeat the same mistake.
Why trusted work gets waved through
The research points at trust. The more someone trusts the tool, the less they check it.
- More trust, less checking. Researchers at Microsoft Research and Carnegie Mellon University surveyed 319 knowledge workers who use generative AI at work. Between them they shared 936 real examples. Their finding, in their own words: “higher confidence in GenAI is associated with less critical thinking, while higher self-confidence is associated with more critical thinking.”
- Brilliant on some tasks, a handicap on others. In a pre-registered experiment with 758 Boston Consulting Group consultants, people using GPT-4 on tasks it handled well completed 12.2% more tasks, 25.1% more quickly, at more than 40% higher quality. On a task chosen to be outside what the AI could do well, consultants using it were 19 percentage points less likely to reach the correct answer than consultants working without it.
- A mistake that went through a whole chain. In May 2025 a summer reading list ran in a syndicated supplement carried by the Chicago Sun-Times and The Philadelphia Inquirer. Ten of the fifteen books on it did not exist. The freelance writer said he had used AI and had not checked what it produced. It passed through a writer, a syndication company and at least two newspapers, and was printed before anyone noticed.
When work looks finished and the tool has been right before, checking feels like a waste of time. That is when it matters most. In Why AI needs a manager we looked at the same effect in medicine, where experienced radiologists became much less accurate when they were shown deliberately wrong AI suggestions.
How we check AI work before it goes out
None of this means doing the work twice. Checking usually takes a fraction of the time the work took. It just has to actually happen. This is the routine we use:
- Look at it where the customer will see it. On the live page, in the email preview, on a phone. Not in the editor or the AI chat window. Repeated text and broken symbols are much easier to spot in the finished place.
- Search for the obvious. Repeated sentences, odd characters, notes that were meant to be removed, a placeholder still sitting in square brackets.
- Re-check every figure at its source. Every price, date, percentage and claim, against the live account or the original document, not the copy the AI worked from.
- Use a second check that is looking for mistakes. For anything going to a client or going live, a separate review whose only job is to find what is wrong, not to agree that it is right. We use a second, independent AI pass for this, and then a person makes the final call.
- Treat “done” as unconfirmed until you have seen the result. If the tool says a page was updated or an email was sent, open the page or the sent folder and look.
- Write down what went wrong. A short note each time a mistake is caught. Patterns show up quickly, and most of them can be fixed for good with a clearer instruction or an extra check.
The short version
Let AI do the work. Then check it, properly, before it goes anywhere a customer can see it. The time AI saves is real, it is just not all of the time. The part it gives back is best spent on the checking and on the decisions a tool should not be making on its own.
FAQ
Can AI check its own work?
Partly. In our experience a second, separate AI pass that is told to look for mistakes catches things the first pass missed. But one AI checking another can repeat the same error, especially if both read the same source. The final check should be a person looking at the original source.
What are the most common mistakes in AI-produced marketing work?
The ones we see most are repeated text, characters that break when copied between systems, figures checked against an out-of-date copy, dates and time zones slightly out, and tasks reported as finished when they are not. Claims that sound right but have no source behind them are the other big one.
How long should checking take?
It depends on what is going out. A social post might need a couple of minutes. A report with figures in it needs every number traced back to its source, which takes longer. The more people will see it and the more money it affects, the more checking it gets.
Is it still worth using AI for marketing work?
Yes. The same research that shows the risks also shows large gains on the right tasks, and the time saved is real. The aim is not to use less AI. It is to build the checking into the work, so the speed does not come with mistakes attached.
Who is responsible when AI gets something wrong?
The business that publishes it. For AI agents specifically, the Competition and Markets Authority’s guidance says businesses are responsible for what an AI agent does in the same way they are responsible for what an employee does, and that this applies even when someone else designed or supplies the tool.
Sources: Lee et al., The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers, CHI ’25 (Microsoft Research and Carnegie Mellon University, 2025); Dell’Acqua et al., Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of AI on Knowledge Worker Productivity and Quality, Harvard Business School Working Paper 24-013 (2023); NPR, How an AI-generated summer reading list got published in major newspapers (May 2025); CMA, Complying with consumer law when using AI agents (March 2026).