The grace conversions rule: how to cut a keyword before the data is significant
Small accounts never reach statistical significance. Here is the rule we use instead: give a keyword the conversions it has not earned, for free, and see whether it is still worse than its peers.
By Melvin Salas, Director & Co-founder, Riibon · Last verified: 2026-09-18
The advice to wait is written for accounts that will get there
Every article about pausing keywords tells you to wait for significance. Get to 30 conversions, get to 100, run the test properly, do not act on noise.
That advice is written for accounts that will get there.
Take a small local services account: 110 a day of budget, a cost per click around 9, one enquiry a day as the target. Round numbers, but the shape is the common one. That is twelve clicks a day across the whole account. Split across two ad groups and fifteen keywords, a single keyword sees one or two clicks a day. At a 5% conversion rate, one keyword produces a conversion every three weeks.
Wait for thirty conversions on that keyword and you will be waiting for two years. Meanwhile it is spending every day.
So the real question is not when the data becomes significant. It is what the least reckless thing is to do with data that will never be significant. That question has an answer, and we have been using it long enough that it is now written down, reviewed, broken twice, fixed, and running as code.
The rule in one line
Give the keyword the conversions it has not earned, for free, and ask whether it would still be worse than its peers.
That is the whole idea. Everything else is detail.
If a keyword has spent 1,400 and produced one conversion, the temptation is to say it only needs one more and it is fine. So grant it that. One more conversion, at zero extra spend, the most generous assumption available. Two conversions on 1,400 is 700 a conversion. Still terrible. The argument for keeping it has just been made in its favour and it still lost.
Now run the same test on a keyword that spent 20 with one conversion. Give it a free one: 10 a conversion. If the rest of the ad group averages 10, that keyword is now sitting exactly at the average. You cannot stop it. There is no version of the future you can construct in which it is clearly bad.
And a keyword at 30 with one conversion: one free conversion puts it at 15 against a peer average of 10. That is a clear one to stop when you need to optimise.
Three numbers, three verdicts, no p-values.
Compare to peers, not to your target
The benchmark is not your cost per acquisition goal. It is the average of the set the entity actually sits in: ads inside their ad group, keywords inside theirs, ad groups inside the campaign.
Your goal is a business number. The peer average is the real alternative use of the money. If you pause a keyword, its budget does not go to your goal, it goes to its siblings. So the only honest question is whether the siblings would do more with it.
But the peer average on its own is a trap, because half of any set is below average by definition, and if you cut everything below average every week you eventually cut everything.
Two thresholds, and the band between them
So there are two numbers, not one. The peer average, which is what the siblings actually deliver, and the profit ceiling, which is the highest cost per conversion that still makes money. At or under the peer average, keep. Above the profit ceiling, be harsh, because there is no case for an unprofitable keyword. Between the two, keep.
That middle band is where most people over-optimise. If your peer average is 10 and your ceiling is 15, a keyword at 11 is not a problem. It is diversification. Three keywords at 9, 10 and 11 is a better position than one keyword at 10, because you have less audience saturation, three sources of volume instead of one, and a hedge for the week the 10 stops working.
Cutting the 11 to protect the 10 is how accounts end up depending entirely on one term.
The margin has to shrink as the sample grows
That said, at some point an 11 against a 10 is a real difference and you should act on it. The question is when.
The answer is that the gap you can act on shrinks as conversions accumulate. Concretely, we treat a gap as real when it is larger than 1.3 divided by the square root of the conversion count. At one conversion that means the entity has to be 130% worse than its peers before the gap means anything. At ten conversions, 41%. At twenty, 29%. At a hundred, 13%. At a thousand, 4%.
Against a peer average of 10, that is the difference between acting above 23 when you have one conversion and acting above 10.4 when you have a thousand.
At twenty conversions each, you stop the 15 and you leave the 11 alone. At a thousand conversions each, the 11 is now a genuine difference and it goes. Same two numbers, different verdicts, because the evidence behind them is different. That is the part most rules of thumb miss: they fix a threshold and let the sample size float, when it should be the other way round.
How many free conversions is fair
One free conversion is the fast version, for when money is visibly burning and you need to carve. It is deliberately harsh.
Normally the grace should scale with how much the entity has already spent. A keyword that has spent 40 against a 10 benchmark has had four expected conversions worth of chances. A keyword that has spent 400 has had forty. Giving both of them the same two free conversions means the small one is being judged on almost nothing.
So the grace is the larger of 2 and 1.3 times the square root of spend divided by the peer average. The best case is then spend divided by conversions plus that grace, and the verdict falls out of four tests in order. Not judgeable, if it has no conversions and has not yet spent enough to fail. Keep, if its cost per conversion is at or under the peer average. In the band between average and ceiling, keep, unless it has twenty or more conversions and the gap beats the margin above. Above the ceiling, pause if the best case is still above the ceiling.
Plus one rule that overrides all of them: pause, never remove. A removed keyword takes its history with it. A paused one is still there to be read next quarter.
We tried to break it, and it broke
This is the part that would normally stay internal, and it is the most useful part of the article.
Once the rule was written as code, we had it reviewed adversarially: a reviewer whose only job was to destroy it rather than improve it, one pass on the statistics and one on the platform behaviour. It found four things.
One more wasted unit of spend made the keyword look better. Rounding the grace up to a whole conversion meant that at 94 spent the verdict was pause, and at 95 spent it was watch. A rule that rewards waste cannot ship. Fixed by making the grace continuous.
Above the ceiling was softer than below it. The specification said be harsher above the profit ceiling. The code did the opposite: at twenty conversions a keyword at 13 was paused while one at 20 survived, because the free conversions rescued the worse keyword. Fixed by applying both tests above the ceiling.
The pooled average let the worst entity hide. Averaging the whole set means one catastrophic keyword drags the benchmark up until it is no longer worse than the benchmark. Give the test five ad groups, four of them fine and one at 200 times the click price for two conversions, and the average rises enough that an ad group at five times the ceiling comes out as a keep. Fixed by trimming the best and worst before pooling, and by requiring the profit ceiling to be supplied as a business number rather than derived from the peers.
The error rate we were quoting was honest at exactly one point. We had said the fast mode pauses a perfectly average keyword about 20% of the time. Simulated properly across spend levels and true rates, it is 20% to 46%.
What it costs to act early
That last one is worth sitting with. When you carve fast, roughly one in three of the things you pause did not deserve it.
A rule that acts on thin data is wrong a meaningful share of the time by construction. The argument for it is not that it is accurate. It is that the alternative, waiting, is also a decision, and on a small account it is usually the worse one.
Pausing a fine keyword costs you a keyword you can unpause. Waiting six weeks costs you six weeks of budget at a cost per conversion you already suspect.
Three checks before any pause goes live
The arithmetic tells you what to pause. It does not tell you whether pausing will help. Three things decide that.
Is the campaign budget limited or rank limited? Check lost impression share to budget. If you are losing 15% or more to budget, pausing moves that spend to better keywords and you gain conversions. If you are rank limited, pausing saves money and adds nothing, because the remaining keywords cannot absorb the budget anyway. Both are valid, but say which one you are doing.
Where do the searches go? Pausing a keyword does not stop the platform matching those searches. Close variants route them to its siblings, and there is no opt out. If you pause without adding the non converting search terms as negatives, the spend reappears next week under a different keyword and you conclude the pause did nothing.
Can the survivors take the money? If the pause moves more than about 10% of campaign spend, check that the kept keywords have impression share headroom. Otherwise you have not reallocated the budget, you have just underspent.
The whole thing, in formulas
For anyone who wants to implement it. Let S be spend, c be conversions, B the peer average and P the profit ceiling.
The margin that counts as real is 1.3 divided by the square root of c. The grace, k, is 1 in carve mode and otherwise the larger of 2 and 1.3 times the square root of S divided by B. The cost per conversion is S divided by c. The best case is S divided by c plus k. A gap counts as real when c is at least 20 and the cost per conversion exceeds B by more than the margin. The spend floor is three times B in carve mode and otherwise k times P.
Then: not judgeable if c is zero and S is below the floor. Keep if the cost per conversion is at or below B. Keep if it is at or below P and the gap is not real, which is the diversification band. Pause if the gap is real. Pause if the best case is still above P. Watch otherwise.
B is the trimmed pooled cost per conversion of the peer set, frozen before you judge anything, and including the entity you are judging. Never recompute it from the survivors: a benchmark defined as the average of what passed shrinks every week until one keyword is left standing.
All of it runs from one window, and the window starts the day after your last structural change. Landing page, tracker, bidding strategy, budget, audience exclusions: any of those and the history before it belongs to a different account.
Why this exists
Most of what an ads agency does is judgement applied repeatedly to small numbers, and judgement applied repeatedly is exactly what software is for. A rule like this one is the unit of that: a thing a good practitioner does in their head on a Tuesday, written down precisely enough that it can be argued with, tested against its own author's examples, broken by a reviewer, fixed, and then run on every account every week without anyone getting tired.
The rule above now runs as a function. It produced its verdicts on a live account in about a second, and the keywords it flagged were the ones I would have flagged, for reasons I could not have written down a month earlier.
That is the whole business, really. Not automating the ads. Writing down the judgement.