Artificial Intelligence

How to Measure AI Search Visibility: A Practical Methodology

Your agency says your AI visibility improved 40% last quarter.
Ask them one question: against which query set?

If there is not a fixed, documented list, the number means nothing. A visibility figure measured against a query set that changed between periods is not a measurement, it is a chart.

This article is the methodology we use. It is not proprietary and there is no reason to keep it so, because the difficulty here is in running it consistently rather than in knowing what to run.

Disclosure: GrowthRocks sells AI search visibility work. We are publishing the method partly because it is the right way to compete in a market full of unfalsifiable claims, and partly because most teams who read this will conclude they would rather someone else run it.

Why this is hard

Four properties of AI systems make measurement genuinely difficult, and any methodology that does not acknowledge them is overselling.

Non-determinism. The same query returns different answers to the same person minutes apart. There is no single correct result to record.

Personalisation. Since I/O 2026, personal context is a global feature across nearly 200 countries. What a logged-in user sees reflects their Gmail, location and history. Your measurement necessarily describes a baseline that few real users experience.

No official reporting. Search Console has begun surfacing some generative data, but there is no equivalent of the rank tracking infrastructure that classic SEO relies on. Everything is sampled from outside.

Platform divergence. ChatGPT, Google AI Mode, Perplexity and the rest use different retrieval and different source preferences. A single number across all of them hides more than it reveals.

None of this makes measurement worthless. It makes precision claims worthless. Directional measurement, run consistently, is genuinely useful.

The methodology

Step 1. Build a fixed query set

The foundation. Everything downstream depends on this list not changing.

Aim for 30 to 60 queries across four types:

  • Category queries. “best [category] software”, “[category] agency“. Where you want to be considered.
  • Problem queries. How your buyer describes the problem before they know your category exists.
  • Comparison queries. “[you] vs [competitor]”, “alternatives to [competitor]”.
  • Branded queries. Your name, and your name plus a modifier. Tests whether AI systems describe you correctly, which is a separate and underrated problem.

Two rules. Write them as a person would speak to an assistant, because AI Mode’s redesigned search box is built for long conversational input and two-word keywords no longer reflect real behaviour. And freeze the list. Add to it at most quarterly, and when you do, report the old and new sets separately until you have enough history on the new one.

Step 2. Choose your platforms and be explicit

At minimum: Google AI Overviews, Google AI Mode, ChatGPT. Ideally add Gemini, Perplexity, Copilot and Grok.

Weight by where your buyers are, not by platform size. A B2B SaaS buyer’s research behaviour differs from a consumer’s. If you do not know, measure all of them for a quarter and let the data tell you.

Step 3. Define what counts as a citation

This is where most reporting quietly falls apart, because two vendors counting differently produce incomparable numbers.

Record three things separately:

  • Brand mention. Your name appears in the generated answer.
  • Linked citation. Your domain appears as a source.
  • Position. Where in the answer, since first-mentioned behaves differently from listed sixth.

A brand mention without a link still shapes the buyer’s shortlist. A link without a mention still sends traffic. They are different outcomes and collapsing them into one number destroys information.

Step 4. Sample enough to survive non-determinism

Minimum three runs per query per platform per measurement period. Five is better.

Standardise conditions and document them: logged-out or fresh accounts, stated geography, stated time window. Then keep them identical between periods. A change in sampling conditions produces a change in results that looks exactly like a change in performance.

Step 5. Track these five metrics

Citation frequency. How many of your query set returned a citation. Your headline number.

Share of voice. Your citations as a percentage of all brand citations on your query set. More honest than a raw count, because it moves when competitors move.

Sentiment and accuracy. Are you described correctly? Being cited inaccurately is a problem that a citation count will never surface, and it is more common than you would expect.

Source path. Which pages get cited. Frequently a review site or a forum rather than your own domain, which changes what you should work on.

Position within the answer. First-named versus buried in a list.

Step 6. Always measure against named competitors

Your citation count in isolation is close to meaningless. Against three named competitors on the identical query set, it becomes a decision-making input.

Pick competitors you actually lose deals to, not aspirational ones.

Step 7. Set the cadence and hold it

Monthly for most companies. Weekly only if you are running an active programme and can act on the data. Anything less frequent than quarterly and you cannot separate signal from platform drift.

Same day of month, same conditions, same query set. Consistency beats sophistication.

The tools

Ahrefs Brand Radar. Covers the Google surfaces and the major assistants, integrated with keyword data you probably already have. The most practical starting point for most teams.

Dedicated GEO platforms. A crowded and fast-moving category in 2026. Evaluate on whether they publish their sampling methodology. Any tool that will not tell you how many samples it takes per query is asking for trust it has not earned.

Manual sampling. Underrated. Thirty queries run manually once a month across three platforms takes a few hours and produces data you fully understand. Start here if you are unsure, because it teaches you what the automated tools are actually doing.

Search Console. Now surfacing some generative data. Worth watching as reporting improves.

What to avoid: any tool producing a single proprietary “AI visibility score” with no published method. That is a number you cannot audit, cannot explain to leadership, and cannot compare to anything.

What good and bad reporting looks like

Bad: “AI visibility up 40%.” Against what set, on which platforms, with what sampling, versus which competitors.

Good: “Across our fixed 45-query set, sampled three times per platform on 1 August: cited on 12 of 45 queries in AI Overviews, up from 8 in July. Share of voice 14%, versus Competitor A at 22%. Seven of our 12 citations traced to our Clutch profile rather than our own site.”

The second version is auditable, comparable, and points at what to do next. That last sentence in particular is worth more than the percentage.

What this does not tell you

It does not tell you about revenue. Citation share is a leading indicator at best. Connect it to branded search volume and to pipeline before treating it as a business metric.

It does not reflect personalised results. You are measuring an unpersonalised baseline. Real users see something shaped by their own context.

It does not survive platform changes cleanly. When a platform changes its model, your time series has a discontinuity. Annotate them, the way you would annotate a core update.

It does not prove causation. Citations rising after you published something is correlation. Treat it accordingly.

Frequently asked questions

How do you measure AI search visibility? Build a fixed query set of 30 to 60 queries, run it against your chosen AI platforms on a fixed schedule with at least three samples per query, and record brand mentions, linked citations and position separately. Compare against three named competitors on the identical set.

What is a good AI citation rate? There is no established benchmark, and anyone quoting one is guessing. Measure your own baseline and your competitors’ on the same query set. Relative position is the only meaningful reading right now.

How often should I measure AI visibility? Monthly for most companies. Weekly only if an active programme is running and you can act on the results. Less than quarterly and platform drift is indistinguishable from performance change.

Can I track AI citations in Google Search Console? Partially, and improving. Search Console has begun surfacing generative data, but there is still no equivalent of full rank tracking. Most measurement is sampled externally.

Why do AI visibility tools disagree with each other? Different query sets, different sampling frequencies, different definitions of a citation, and genuine non-determinism in the underlying systems. This is why the methodology matters more than the tool.

What is the difference between AI visibility and GEO? AI visibility is the measurement. GEO, generative engine optimisation, is the work you do to improve it. Buying the second without the first is how agencies sell unfalsifiable outcomes.

Should I measure ChatGPT separately from Google? Yes. Source preferences differ substantially between platforms, and a combined number hides which surface is actually working for you.

Was this article useful?

Share
Published by
Theodore Moulos

Recent Posts

I Don’t Read My Emails Anymore

Built an AI routine that checks who is emailing me, understands the business context, filters…

3 days ago

Building Community 2.0: Where Humans and AI Grow Together

Traditional online communities were built for a different internet. Community 2.0 brings humans and AI…

4 days ago

MCP-First Apps: When the UI Is No Longer the Product

MCP-first apps change how software is built and used. Discover why the UI is becoming…

5 days ago

Growth Hacking Techniques That Still Work in 2026 (and the Ones That Stopped)

In 2026, growth hacking is no longer about finding single-turn shortcuts; it has evolved into…

2 weeks ago

Visualizing Results: From Markdown Tables to MCP Apps

Accessing AI data is only half the challenge. Learn how to transition from static Markdown…

2 weeks ago

The 360° Campaign Is Still Mostly a Myth

Learn what is stopping large firms from building truly integrated campaigns! it's not budget!

4 weeks ago