AI search visibility metrics: the KPIs that can actually be measured
Every AI-visibility tool reports “your score.” Very few of them say what the number is made of. That gap is where most of the bad decisions in this category get made: a brand sees a score move four points and rewrites a homepage, when the four points were three lucky runs and one failed API call counted as a zero.
This guide defines the six things that can honestly be measured about a brand’s presence in AI answers, the ones that cannot be measured at all, and the specific ways a dashboard can be technically accurate and still leave you with a false picture. Read it before you buy any tool in this category, including ours.
The six metrics
There are exactly six states worth reporting per prompt, per engine, per run. Everything else on a visibility dashboard is one of these six aggregated, averaged, or renamed.
1 mentioned
The answer names your brand. The model produced the name from what it learned in training or from a source it read elsewhere. Your website may never have been fetched, may not be indexed by that engine at all, and may not exist any more. Mentioned is a reputation signal. It tells you the model associates your name with the category. It tells you nothing about your site.
2 cited
The engine retrieved a page and attributed the answer to it, with a link. This is a different fact, produced by a different mechanism, and it is the one that sends a visitor. Cited means a real fetch happened against a real URL, which means your content, your crawler access and your page structure were all in the path.
3 missing
The answer covered the question, named someone, and did not name you. This is the metric that should drive the work, because it comes with the useful half attached: who got the slot. “Missing” with three competitor names in the same answer is a brief. “Missing” on its own is a mood.
4 not measured
The check did not run. A missing API key, a failed call, a rate limit, a quota exhausted, an engine with no programmatic route. This is a real state and it must appear in the report as itself. The moment a tool renders an unrun check as a zero, every aggregate on the page becomes fiction, and it fails in the direction that flatters the tool: a broken integration looks like your competitor winning.
5 share of answer
Of the runs where anyone in your category was named, the proportion that named you. This is the only one of the six that is a ratio, and a ratio needs its denominator stated. Share of answer computed over five runs is not a percentage, it is a fraction with a small number on the bottom pretending to be a percentage.
6 position
Where you appear when the answer lists several options: first recommendation, third, or the hedge at the end. Position is the most volatile of the six. Report it as a distribution across runs, or do not report it.
Mentioned is not cited, and merging them is the easiest lie to get caught in
If you take one thing from this page, take this one.
A language model can name your brand with no access to your website whatsoever. It learned the name. Ask it for the best project management tools and it will produce a list from memory, confidently, whether or not any of those companies still exist in the shape it describes. That is a mention.
A citation is a different event in a different part of the system. The engine decided the question needed current information, ran a retrieval, fetched pages, and linked one. Everything you control lives on that path: whether your robots.txt lets the crawler in, whether your content renders without JavaScript, whether the page answers the question in a liftable form, whether the site is fast enough to be fetched inside the engine’s patience.
So the two numbers have opposite implications for what you do on Monday:
- Mentioned but never cited. The model knows the name and reaches for it out of memory. Your site is not in the retrieval path. This is a technical and structural problem first: crawler access, rendering, page structure, schema. Rewriting more blog posts will not fix a blocked user agent.
- Cited but rarely mentioned. Engines will fetch and quote you when the question is specific enough to trigger retrieval, but you are not part of the model’s default picture of the category. This is a presence problem: the pages, roundups, comparisons and third-party sources the models have already read do not include you.
- Both. The good case, and also the one that decays quietly if nothing is watching.
- Neither, while competitors are both. The whole reason this category exists.
A dashboard that reports one number called “visibility” or “mentions” over both events cannot distinguish these four situations. Neither can you, if that is the only number you have. Ask any vendor, in these words: does this number count answers that named me but never fetched my site? If they cannot answer immediately, the number merges them.
Why single runs lie (and what averaging actually fixes)
Ask an engine the same question five times and you will not get the same answer five times. Sampling is not a bug in these systems, it is how they generate text. Retrieval adds a second layer of variance on top: what gets fetched depends on the phrasing, the moment, and sometimes the region.
A realistic five-run result for one prompt looks like this:
| Run | 1 | 2 | 3 | 4 | 5 |
|---|---|---|---|---|---|
| Result | cited | missing | cited | mentioned | cited |
Honest reporting of that row is “cited in 3 of 5 runs.” Not “you rank #1.” Not “you are cited.” And certainly not a screenshot of run 1.
What averaging fixes is precisely this and nothing more: it converts a coin flip into a rate. What it does not fix is a bad prompt set, a bad denominator, or an engine that was down. Averaging noise still gives you an average of noise, just with more decimal places and more apparent authority.
Three practical consequences:
- A single-check score is not a measurement, it is an anecdote. Anyone selling one has chosen a cheaper API bill over a defensible number.
- Movement between two runs is not a trend. Only a line across weeks is a trend, and only if the prompt set and the run count held constant between the points. Changing the prompt set and then showing the same line is the oldest trick in reporting.
- Volatility is itself a finding. A brand cited in 5 of 5 runs and a brand cited in 3 of 5 runs are in genuinely different positions, and both average above zero. A tool that reports only the mean throws away the more interesting half.
What is not measurable, stated plainly
Some surfaces cannot be measured programmatically. Not “hard to measure.” Not “coming soon.” Cannot.
Engines with no API. If there is no programmatic route to a surface, there is no measurement of it, only spot-checks. Amazon’s Rufus is the clearest example: no API, so no honest tracking. A tool that claims to track it is either scraping a consumer interface or a human is looking at it occasionally and the report is not saying so.
Google’s AI answers, without a paid intermediary. Google publishes no API for AI Overviews or AI Mode. Reaching them at all means a paid SERP vendor sitting in the middle, on quotas, and any tool that includes them is carrying that dependency whether it mentions it or not. That is workable. Pretending it is a first-party measurement is not.
The dedicated shopping interfaces. The chat APIs do not expose the shopping carousels and product tiles. What is answerable is the buyer question itself: ask an engine which running shoes it recommends and the answer is measurable through the normal API. “Your product’s tile in a shopping carousel” is a different object and it is not exposed.
Anything behind a logged-in session. A personalised answer, shaped by a user’s account, chat history, memory and settings, is not reproducible by anyone else. It cannot be measured because there is no stable thing to measure. If you see your brand in your own ChatGPT and not in a report, that is not necessarily a contradiction, and neither observation invalidates the other.
Prompt volume. No engine publishes how many people ask a given question. Any volume figure a tool shows you next to a prompt is modeled unless it comes from a measured consumer panel, and modeled numbers get quoted in pitch decks long after the footnote falls off. Prompt sets are worth generating; volume columns beside them are worth ignoring.
Not measurable
The correct output for all of these is “not measurable,” and a tool willing to write that sentence is telling you something about how it treats the numbers it does show.
How to read a visibility report without fooling yourself
Six questions, in order, before you act on anything a dashboard tells you.
What is the denominator? Every percentage on the page is a fraction. Runs of what prompt set, across how many engines, over how many repetitions. A share-of-voice figure with no visible denominator is decoration.
Which of the six states produced this number? If the answer is “mentions,” find out whether citations are folded in.
What happened to the checks that failed? Look for an explicit unrun or not-measured state. If the report has no way of expressing “this did not run,” it is expressing failures as zeros.
Did the prompt set change between the two dates I am comparing? If yes, the comparison is void. Add prompts, by all means, but the trend line has to restart or split.
Is this the mean, or the distribution? Ask for run-level results on at least one prompt. If the tool cannot show them, it may not be keeping them.
Does the recommended action follow from the metric? “You are missing from 12 buyer questions” is a diagnosis. It is not a task. The next artifact after it should be the pages that answer those questions, or the technical reason the engine cannot read the pages you already have. A report that ends at the number ends one step short of the work.
The metrics vendors dress up, and the one question that unmasks each
None of these are lies exactly. All of them are numbers arranged to look larger, steadier or more scientific than the underlying data supports.
| The dressed-up metric | What it hides | The question that unmasks it |
|---|---|---|
| A single “AI visibility score” | An unpublished weighting of the six states, tuned so the number moves | “Which of the six states go into this, and with what weight?” |
| “Mentions” as one figure | Whether the site was ever fetched | “Does this count answers that named me but never cited me?” |
| A zero | A failed call, a rate limit, or a missing key | “How does this report show a check that did not run?” |
| “Ranked #1 in ChatGPT” | One run, screenshotted | “Over how many runs, and what was the spread?” |
| Share of voice | The denominator, and often a competitor set the tool chose | “Share of what, over which prompts, and who picked the competitors?” |
| A sentiment score | A classifier’s judgment presented as a measurement | “What model classified this, and can I see the sentences it scored?” |
| Prompt volume | That nobody publishes it | “Measured from what, or modeled?” |
| Engine count (“we track 12 engines”) | That several of them have no API | “Which of those have a programmatic route, and which are spot-checked?” |
The unmasking question
One question covers most of it: measured how, and over how many runs? A tool that answers cleanly is worth trusting on the numbers you cannot check. A tool that reaches for adjectives is telling you the number is decoration.
What to do with all of this
Measure the six states separately, average across repeated runs, keep unrun checks visible as unrun, and treat missing plus the competitor who holds the slot as the actual work queue. Then fix the thing the split points at: crawler access and page structure when you are mentioned but not cited, presence in the sources engines already read when you are cited but not mentioned, and pages that answer the buyer questions when you are neither.
The number is not the deliverable. The page you did not have is.
FAQs
What is the difference between being mentioned and being cited in AI answers?
A mention is the model naming your brand from what it already knows, with no guarantee your website was ever fetched. A citation means the engine retrieved a specific page and linked it as the source of the answer. They come from different parts of the system, they have different fixes, and a tool that reports them as one number cannot tell you which situation you are in.
How many AI visibility metrics are there really?
Six states are honestly measurable per prompt, per engine, per run: mentioned, cited, missing, not measured, share of answer, and position. Everything else on a visibility dashboard is one of those six aggregated, averaged, weighted or renamed. If a metric cannot be traced back to them, ask what it is made of.
Why do AI visibility scores change when nothing on my site changed?
Because these engines sample. The same prompt asked five times can return five different answers, and retrieval adds more variance on top. That is why a single check is closer to a coin flip than a measurement, and why any number worth acting on is an average across repeated runs with a trend line behind it.
Can AI visibility in ChatGPT, Perplexity and Gemini be measured accurately?
Those three expose APIs, so answers to a fixed prompt set can be queried repeatedly and reported as rates rather than snapshots. Accuracy comes from the run count and the honesty of the labels, not from the engine list. Surfaces with no API cannot be measured at all, only spot-checked, and should be labelled that way.
What does “not measured” mean in a visibility report?
That the check did not run: a failed call, a rate limit, an exhausted quota, or an engine with no programmatic route. It is not a zero and it is not evidence of absence. A report that has no way of displaying “this did not run” is displaying its own failures as your losses, and always in the direction that makes the tool look busier.
Why can’t anyone track Google AI Overviews or Amazon Rufus directly?
Neither publishes an API. AI Overviews is reachable only through a paid third-party SERP vendor, on a quota, which is a real dependency rather than a first-party measurement. Rufus has no compliant programmatic route at all, so anything reported about it is a human spot-check or a scrape. “Not measurable” is the honest output for both.
How often should AI visibility be checked?
Often enough that a change is visible above the run-to-run noise, which in practice means a fixed prompt set re-run on a regular schedule and compared over weeks, not days.
These exact metrics, on your brand.
The six metrics, defined precisely, averaged across repeated runs — free.