Google's "Ad Strength" label doesn't predict which ads actually perform

Ad Strength grades how complete and varied your ad's assets are, not whether the ad actually converts, and chasing the label can quietly make your ads worse.

By Melvin Salas, Director & Co-founder, Riibon · Last verified: 2026-07-24

What Ad Strength actually measures

If you've built a Search ad in Google Ads, you've seen the little meter next to your ad: Poor, Average, Good, or Excellent. It's easy to read that as a performance grade, something like a report card telling you how well the ad will do. It isn't. Ad Strength is Google's assessment of how complete and varied your ad's raw material is, specifically for Responsive Search Ads (RSAs), the ad format where you supply a pool of headlines and descriptions and Google's system mixes and matches them per auction.

The label is built from three checklist-style inputs: how many distinct headline and description variations you've provided (Google recommends up to 15 headlines and 4 descriptions), how different those variations are from each other (five near-identical headlines score worse than five genuinely distinct ones), and whether the ad follows Google's formatting guidance, like working relevant keywords into headlines and avoiding excessive punctuation or ALL CAPS. That's the whole recipe. Nowhere in that calculation does Google ask whether your specific audience finds the message compelling, whether the offer is any good, or how the ad has actually performed for real users in the auction.

Why "Poor" doesn't mean it won't work, and "Excellent" doesn't mean it will

Here's the part that trips people up: Ad Strength is calculated the moment you save the ad, before a single real person has seen it or clicked it. And critically, it keeps being calculated the same way even after the ad has run for months and built up a genuine track record of clicks, conversions, and cost data. The label doesn't upgrade itself based on results. It stays anchored to asset completeness and formatting, whether the ad is brand new or has driven thousands of conversions.

That disconnect matters in practice. Picture a hypothetical example: an ad with three carefully written headlines, each speaking directly to a specific customer problem, sitting at "Good" because it doesn't hit Google's variation count. Next to it, an ad with twelve headlines, several of them generic filler added purely to check boxes, sitting at "Excellent." There's no rule that says the tightly-written ad performs worse. In fact it's entirely plausible for the opposite to happen, because relevance and message-match to the searcher tend to matter more than raw variation count. Industry practitioners have long noted, anecdotally and without a fixed statistic attached, that lower-labeled ads sometimes outperform higher-labeled ones for exactly this reason. Nobody, including Google, has published a hard number for how often that happens, and you shouldn't trust anyone who claims to have one. The honest takeaway is simpler: the label and the outcome are measuring different things, so they don't have to move together.

The trap: optimizing for the label instead of the ad

Because the Ad Strength meter sits right there in the interface, glowing orange or green, it's tempting to treat moving it as a task in itself. Google Ads will often nudge you toward exactly that, suggesting you add more headline variations to climb from "Good" to "Excellent." The problem is that once you've already written the two or three headlines that actually say something specific about your product, the fastest way to add more variation is to pad the pool with generic filler: vague benefit statements, restated versions of the same headline, phrases that could belong to almost any business in your category.

That's the trap. You're not adding new persuasive material, you're diluting the deliberate message you already had with noise, purely to satisfy a checklist. The label goes up. The ad, if anything, gets worse, because Google's system is now more likely to serve a combination that includes your weaker filler headlines alongside your strong ones. Chasing the score can actively work against the goal the score is supposedly a proxy for.

How to actually judge whether an ad is working

The only real signal is real traffic: within-group conversion rate and cost per conversion, compared across your own ad variations, over a sample size and time window large enough to not be noise. That means giving an RSA (or a set of ads you're testing against each other) enough real impressions and conversions before drawing conclusions, then looking at what your own account's data says about clicks, cost, and conversions, not at the label Google assigned before any of that data existed.

That doesn't mean ignore Ad Strength entirely. It's genuinely useful as a completeness reminder while you're first building an ad. If it says "Poor" because you've only written a single headline variation, that's worth fixing since a bigger, more diverse pool gives Google's system more combinations to test against real searchers. The distinction is timing and purpose: treat the label as a pre-launch checklist, not a post-launch verdict. Once an ad has real performance data behind it, that data is the only thing that should decide whether to keep it, change it, or kill it.

The same caution applies to every auto-generated score

Ad Strength isn't the only platform-assigned grade that gets over-trusted. Google's account-level "Optimization Score" and its recommendations panel, and equivalent features on Meta and other ad platforms, are all built the same way: from the platform's own general heuristics about what a well-configured account looks like, not from validated evidence about what will work for your specific business, audience, or offer. They're worth glancing at as a source of ideas, in the same way Ad Strength is worth glancing at as a completeness check. They're not worth treating as instructions, and they're not a substitute for looking at what your own conversion data is actually telling you.

← All articles