Measurement · AI search
AI search KPIs: what to measure, and how many times to measure it
Picking the metrics is the easy part — the industry broadly agrees on them. The part almost nobody publishes is how many samples each one needs before the number means anything.
By Manas Dasgupta, Founder & Lead Engineer · Published 7 Oct 2026 · Last reviewed 7 Oct 2026 · 10 min read
What KPIs should you track for AI search visibility?
Eight metrics across three tiers: mention rate, citation rate, share of voice, accuracy rate and win rate at the channel level; AI referral sessions and assisted conversions at the performance level; and retrieval mode, which tells you whether any of the others can be moved by on-site work.
That list is not contentious. Ahrefs, iPullRank and most AI visibility tools converge on roughly the same set, and the three-tier structure — inputs, channel, performance — comes from iPullRank's framework.
What is contentious, or at least unexamined, is everything that happens after you pick them. A mention rate computed from one run of one prompt is a coin toss reported as a fact. Most AI visibility dashboards show exactly that, with a sparkline attached, and the movement in the sparkline is mostly sampling noise.
So this post covers the metrics briefly and the sampling at length, because the sampling is what decides whether your reporting survives a sceptical question from a CFO.
Why don't traditional SEO metrics work for AI search?
Because three of the four quantities conventional SEO reporting is built on do not exist in a generative answer.
There is no rank. A generative engine composes one answer from several retrieved passages. There is no ordered list, so there is no position one. Order of mention inside an answer is sometimes recorded as a position-weighted score, but it is not a ranking and it moves between runs of the same prompt.
There are no impressions. No engine publishes how often an answer containing your brand was shown. Search Console reports Google's AI surfaces inconsistently and nothing equivalent exists for ChatGPT, Perplexity or Claude.
There is no click-through rate on an answer the user never clicks. The whole design intent of an AI answer is that the user does not need to visit the source.
iPullRank calls the result the Measurement Chasm: “A user's query might retrieve passages from your site, merge them with content from other sources, and synthesize an answer, but unless that answer includes a citation and the user clicks it, you have no direct evidence of your involvement.”
The fourth quantity, sessions, does survive — but it is small enough to be statistically thin by construction, which is the first reason you need channel metrics rather than traffic metrics as the primary KPI.
The set
What does each KPI actually tell you?
Eight metrics. The fourth column is the one usually missing from vendor documentation: what breaks the number.
| KPI | What it answers | How it is computed | What breaks it |
|---|---|---|---|
| Mention rate | Are we in the conversation at all? | Answers naming the brand ÷ total sampled answers | Reported without a confidence interval, or from too few runs |
| Citation rate | Is our own page being used as a source? | Answers attributing a URL on your domain ÷ total sampled answers | Counting any domain mention as a citation |
| Share of voice | Are we winning or losing against named rivals? | Your mentions ÷ all vendor mentions in the same answers | Changing the competitor set between periods |
| Win rate | When compared directly, do we get picked? | Head-to-head prompts where you are recommended first | Too few comparison prompts to be stable |
| Accuracy rate | Is what AI says about us true? | Checkable claims matching approved facts ÷ all checkable claims | No maintained fact list to check against |
| AI referral sessions | Is anyone actually arriving? | Sessions matched by referring domain in analytics | Referrer stripping; volume too low for monthly significance |
| Assisted conversions | Does it reach revenue? | Conversions with an AI referral anywhere in the path | Last-click attribution hiding the assist entirely |
| Retrieval mode | Did the engine search, or answer from memory? | Presence of live citations vs. unsourced recall, per answer | Not recorded at all by most tools |
The two rows in Signal are the ones that change decisions. Accuracy rate is the only metric here where a rising number can mean a worsening problem — more mentions means more chances to be described wrongly. Retrieval mode decides whether on-site work can move anything: if an engine is answering from training memory rather than searching, your content changes will not reach it this quarter regardless of what you publish.
How many times must you ask before a number means anything?
A mention rate is a binomial proportion, so the arithmetic is settled and borrowed directly from survey statistics. At a true rate near 30%, ten runs give you a 95% interval of roughly ±28 percentage points. The number you are reporting is almost entirely noise.
Every one of the channel metrics above is a proportion: some number of sampled answers out of a total. That makes the precision of the estimate a function of one thing you control — how many answers you sampled.
The half-width of a 95% confidence interval for a proportion is approximately 1.96 × √(p(1−p) / n). At p = 0.3, which is a realistic mention rate for a mid-market brand in its own category, that produces this:
Confidence interval half-width, by number of observations
95% interval at a true mention rate of 30%. Shorter is better. Bars are scaled against the widest interval.
Normal approximation to the binomial, α = 0.05. Percentage points. Precision improves with the square root of n, so quadrupling the runs halves the interval.
What this means for your prompt set
Observations are prompts × runs × engines. A practical floor for portfolio-level reporting is 50 prompts, run 5 times each, per engine, per period — 250 observations, giving roughly ±6 percentage points.
The consequence most teams have not absorbed: individual prompts can almost never be reported on their own. One prompt run five times gives you five observations and an interval of roughly ±40 points. A dashboard that shows per-prompt movement week on week is showing you noise in a presentation layer.
What counts as a real change?
Detecting a difference between two periods is harder than estimating one period, because both estimates carry error. As a working rule, the minimum detectable change is roughly 1.4 times the half-width of a single period's interval.
At 250 observations per period, that puts the minimum detectable effect at about 8 percentage points. A move from 30% to 35% is, on that sample, no detected change — and should be reported in exactly those words rather than as a 5-point gain.
The reporting rule that follows. Report every channel metric as a rate with its interval, and state the minimum detectable effect alongside it. Any movement smaller than that is recorded as “no detected change”. This is less satisfying than a rising line, and it is the only version that survives someone checking it.
Three conditions that must be frozen
- The prompt set. Adding or removing prompts changes the population being sampled, which invalidates the comparison. Version it, and when you revise it, report the old and new sets in parallel for one period.
- The competitor set. Share of voice has a denominator. Changing who is in it changes the metric without anything changing in the world.
- The model version, where it is exposed. A model update can move every number overnight. Record what you queried, so a step change can be attributed to the engine rather than to your work.
Evidence
Does AI referral traffic convert better than organic search?
Slightly, on average, and not for everyone. The popular “four to five times better” figure does not appear in the best published data.
-
1.18×
AI referral conversion rate against organic search — 7.18% versus 6.08%. Pooled median 1.21×. 95 websites, 12 months to August 2026.
-
2.6%
AI referrals as a share of organic session volume. Organic delivered 39 times more traffic across the same properties.
-
23 / 63
Websites in the conversion analysis where AI traffic converted worse than organic search, not better.
The claim that AI traffic converts four to five times better than organic circulates widely and is usually presented without a sample size. The most reliable counter-evidence available is a twelve-month benchmark of 95 websites running September 2025 to August 2026, with conversion analysis restricted to the 63 properties clearing a threshold of 100 AI sessions and 1,000 organic sessions.
It found AI traffic converting at 7.18% against 6.08% for organic — a 1.18× ratio, not 4×.
More useful than the average is the spread. Business-to-consumer sites showed 1.51×. Business-to-business product companies showed 0.62× — AI traffic converting materially worse. On 23 of the 63 sites, AI underperformed organic.
If you sell B2B software, the published average is pointing the wrong way for your business. Measure your own ratio before accepting anyone's multiplier, including this one.
| Segment | AI vs organic |
|---|---|
| B2C | 1.51× |
| All sites (mean) | 1.18× |
| Pooled median | 1.21× |
| B2B product | 0.62× |
| Commonly repeated claim | 4–5× |
One benchmark is one benchmark. It is a larger and longer sample than most figures circulating in this field, and its method is stated, which is why it is used here — but a single study is directional, not settled. Published research on AI search behaviour disagrees with itself substantially, and the honest position is to cite the sample size and date every time.
Anti-patterns
What should you stop reporting?
Five things that appear on AI visibility dashboards and cannot support a decision.
- A single-vendor “AI appearance metric” A composite with no published check list and no denominator is not a measurement. Ask what it is out of and which inputs it combines; if the answer is not public, the number cannot be defended to a board or compared against anything.
- Any metric from one observation Screenshots of a single ChatGPT answer are evidence that it happened once. They are not a rate, and they should never appear in a trend line.
- Rank-style position numbers “We are position 2 in ChatGPT” imports a concept that does not exist. Order of mention varies between runs of the same prompt and is not a rank.
- Mention counts without a denominator “412 mentions this month” is unreadable without knowing how many answers were sampled. Rates, not counts.
- Week-on-week movement at the prompt level Five observations cannot show a trend. Reporting it trains stakeholders to react to noise, which is worse than not reporting it at all.
Build
How do you set this up from scratch?
Six steps. The first two determine whether everything after them is interpretable.
-
Build the prompt set from real buyer questions
Fifty prompts is a workable floor. Stratify them: category-level questions where you want to be named, comparison questions against named competitors, and factual questions about your own company for the accuracy metric. Take the wording from sales calls and support tickets, not from a keyword tool — people phrase questions to an assistant differently than they type them into a search box.
-
Freeze the prompt set and the competitor list, and version both
Write them down with a version number and a date. Every reported number carries that version. When you revise, run old and new in parallel for one period so the series stays connected.
-
Decide the sampling plan before the first run
Runs per prompt, engines covered, collection frequency. Compute the resulting confidence interval and minimum detectable effect, and agree them with whoever will read the report. Doing this afterwards means fitting the statistics to a number you have already shown someone.
-
Record retrieval mode on every answer
Note whether the engine cited live sources or answered without them. This is the single most diagnostic field in the dataset and almost no tool captures it. It tells you whether a flat month means your content failed or was never fetched.
-
Maintain an approved-facts list
Pricing, positioning, founding year, locations, named services, leadership. Accuracy rate is uncomputable without it, and assembling it usually surfaces internal disagreement about what the company currently claims, which is useful on its own.
-
Report monthly, collect weekly, escalate accuracy immediately
Weekly collection accumulates observations and catches sudden drops. Monthly reporting keeps the interval meaningful. Factual errors about your company are the exception: those go out when found, because they compound while they sit in a queue.
Common questions
AI search KPIs — answered
What KPIs should you track for AI search visibility?
Eight, grouped in three tiers. Channel metrics are mention rate, citation rate, share of voice, accuracy rate and win rate. Performance metrics are AI referral sessions and assisted conversions. The eighth is retrieval mode, which records whether the engine searched the live web or answered from training memory, because that determines whether on-site work can move the other seven at all. Tracking fewer than five of these produces a report that cannot explain why a number moved.
How many times should you run each prompt?
Enough that the confidence interval is narrower than the change you want to detect. A mention rate is a binomial proportion, so at a true rate near 30 percent, ten observations give a 95 percent interval of roughly plus or minus 28 percentage points, thirty give plus or minus 16, one hundred give plus or minus 9, and four hundred give plus or minus 4.5. A practical floor for portfolio-level reporting is fifty prompts run five times each, which is 250 observations per engine per period. Individual prompts almost never carry enough observations to be reported on their own.
Does AI referral traffic convert better than organic search?
Slightly, on average, and not for everyone. A twelve-month benchmark of 95 websites running from September 2025 to August 2026 measured AI referral traffic converting at 7.18 percent against 6.08 percent for organic search, a ratio of 1.18 times with a pooled median of 1.21 times. The widely repeated claim that AI traffic converts four to five times better does not appear in that data. The average also hides the spread: business-to-consumer sites showed 1.51 times, business-to-business product companies showed 0.62 times, and on 23 of the 63 sites analysed AI traffic converted worse than organic search.
Why can't you track rankings in AI search?
Because there is no ranked list to hold a position in. A generative engine composes one answer from several retrieved passages, so there is no position one, no impression count and no click-through rate for an answer the user never clicks. Order of mention within an answer is sometimes recorded as a position-weighted metric, but it is not a ranking in the search engine sense and it varies between runs of the same prompt.
What is a good AI share of voice?
There is no absolute benchmark, because share of voice is defined against a competitor set you choose. A 20 percent share across four named competitors is parity; the same 20 percent across twelve competitors is leadership. The number is only interpretable alongside the competitor list, the prompt set and the period, and it is only comparable over time if all three stay frozen. Report it as a change against your own prior period, never as a standalone figure.
How often should you report AI visibility KPIs?
Monthly for reporting, weekly for collection. Collect weekly so you accumulate enough observations to narrow the interval and so you can see a sudden drop quickly. Report monthly, because a weekly report on 50 observations will show movement that is almost entirely sampling noise and will train stakeholders to react to it. Accuracy incidents are the exception and should be escalated whenever they are found, not held for the reporting cycle.
Sources
Every figure on this page, with its origin
| Claim | Source | Date |
|---|---|---|
| AI conversion 7.18% vs organic 6.08%; 1.18× mean, 1.21× pooled median | SEO Works AI referral traffic benchmark — 95 sites, 63 in conversion analysis | Sep 2025 – Aug 2026 |
| AI referrals = 2.6% of organic volume; organic 39× larger | SEO Works — 2,825,594 organic vs 72,095 AI sessions | Sep 2025 – Aug 2026 |
| B2C 1.51×, B2B product 0.62×, 23 of 63 sites worse | SEO Works | Aug 2026 |
| Three-tier model; “Measurement Chasm”; repeated sampling requirement | iPullRank, The Measurement Chasm | 2026 |
| Metric definitions: mention rate, win rate, citation stability, position-weighted share of voice | Ahrefs AI visibility workflow | 2026 |
| Confidence interval and minimum detectable effect figures | Computed — normal approximation to the binomial | — |
Published research on AI search behaviour disagrees substantially between studies. Figures above are reported as found, with sample size and date, and should be treated as directional rather than precise.
Author
Manas Dasgupta
Founder & Lead Engineer · Code4X
Founder and Lead Engineer at Code4X. Building tools for Answer Engine Optimization and agent readiness. — two or three sentences of genuine background: what they build, what data they see, why they are qualified to write this. LinkedIn
- Reviewed7 Oct 2026
- Sources cited3
- Next reviewJan 2027
Continue reading
More on AI search and agent readiness
- AEO, GEO, AIO and LLM SEOThree of these four describe the same work. Here is what differs.
- We audited our own website against 37 AEO checks. We scored 23.A published self-audit, including the two failures that were embarrassing.