2026-07-14
How we stress-test an AI-generated ad recommendation before anyone acts on it
An AI system can produce a plausible-sounding recommendation in seconds. The discipline that matters is what happens between that output and a change to a live account.
By Melvin Salas, Director & Co-founder, Riibon
Plausible is not the same thing as correct
Ask an AI system to look at an ad account and it will find something. Point it at a Meta or Google account with enough history and it can always produce a recommendation that sounds reasonable: a metric moved, a cause is proposed, an action is suggested. The output reads as confident and specific, and that's exactly the problem. Confidence and specificity are properties of the writing, not the analysis.
The failure mode that gives AI-assisted marketing analysis its bad reputation isn't that the AI is wrong sometimes, every analyst is wrong sometimes. It's treating the AI's first pass as a finished verdict instead of what it actually is, which is a hypothesis generated fast by a system that hasn't yet been asked to defend it. A recommendation that hasn't been checked against the mechanism it claims, the data it cites, and the ways it could be an artifact isn't an insight. It's a guess with good sentence structure.
The actual discipline
We run every AI-generated recommendation through the same scrutiny we'd apply to a proposal from a human analyst, before it's presented as a real finding. That means checking the specific mechanism being claimed: not "CPA went up," but why, mechanically, would this specific change to this specific campaign produce this specific effect. A recommendation that can't name a mechanism doesn't get a pass just because the numbers moved in a suggestive direction.
It also means checking whether the data actually supports the claim rather than merely correlating with it, and running the recommendation against a checklist of the ways this kind of analysis commonly goes wrong: metric definitions that don't actually match across the comparison being made, data windows that haven't matured yet and will look different in three days, two changes that happened at the same time so the effect can't be cleanly attributed to either one, and small-sample noise dressed up as a trend. None of these are exotic failure modes. They're the ordinary ways a directionally-correct-looking analysis turns out to be wrong, and an AI system is exactly as capable of tripping over them as a person working too fast is. After that, a human reviews it before anything is acted on for a live account. No exceptions for how confident the output sounds.
Verification isn't a disclaimer, it's the operating model
It would be easy to frame this as a safety net bolted onto an otherwise-automated process, the thing you add so you can say you added it. That's not how we think about it. Our position is not "AI helps us go faster and we hope it's right." It's that AI expands the number of accounts, campaigns, and time windows we can actually look at closely, and the verification discipline is the thing that makes that expanded surface area trustworthy instead of just louder.
That distinction matters for how the human review step functions day to day. It isn't a bottleneck fighting against the AI, slowing down what would otherwise be a clean automated pipeline. It's the component that makes running AI analysis at this scale a reasonable thing to do in the first place. Take it away and you don't get the same output faster, you get a different, worse product: a lot of confident-sounding text with no mechanism to catch the fraction of it that's wrong.
This exists because we've been wrong before
We're not claiming this process makes AI-generated analysis infallible. It doesn't, and treating it as though it does would just move the overconfidence problem one layer up, from the AI's first draft to our own review process. AI-generated analysis, including our own, can be wrong in specific, describable ways: it can build a coherent story around a coincidence, tie two unrelated changes together in a satisfying causal narrative, or take a real change out of context in exactly the pattern of an AI advisor blaming a competitor for a shift that came from somewhere else entirely (we wrote about a version of that on this site, see the piece on why your platform's AI advisor blames a competitor).
That's the point of writing this down as a discipline rather than a claim. The value isn't that our AI doesn't make mistakes, it's that we've built a step whose entire job is to catch the mistakes before they reach your account. Verification is what separates AI-assisted analysis that's actually useful from AI output that just sounds useful, and we'd rather be explicit about needing that step than let the fluency of the output imply we don't.