Twenty questions put to ChatGPT or Google AI Mode about your category give your brand a visibility number. Across 513,813 answers from the query sets our sister company Qvery tracks, that number landed within five points of the brand’s real monthly figure 55% of the time,1 and nothing on the screen says which kind of reading you hold.
The habit is borrowed from search. A Google ranking is one published list, so one look tells you where you stand. An AI answer is drawn fresh every time, which makes the visibility number a poll with a margin of error, and nobody selling a dashboard prints it. Here it is, measured, after the SEO metrics it gets mistaken for.
Empact Partners, a go-to-market consultancy working inside software companies’ marketing teams, sizes that question set in the first weeks of every Generative Engine Optimization (GEO) partnership, our work on getting a brand named in AI answers about its category. If you would rather hand the measurement to us, book a call with me.
SEO Metrics Read a Published List, AI Visibility Metrics Sample Answers
SEO metrics grew up around a ranked list Google publishes for each query, which everyone who searches it sees in roughly the same order. Position, impressions and clicks hang off that list, and search volume says what each query is worth. Check your rank on Monday and you could still trust it on Tuesday.
An AI engine publishes no list. It writes an answer, names a handful of brands, cites a few pages, then writes a different answer for the next person. So AI visibility metrics count answers, and the two that matter most, brand visibility and share of voice, divide the same answers two ways.
| SEO metric | What it reads | AI visibility counterpart | What it reads |
|---|---|---|---|
| Rank position | Your place in one list Google publishes for a query | Visibility | The share of sampled answers that name you |
| Search volume | How often a query is searched, published by Google | Nothing published | The question set somebody chose stands in for demand |
| Clicks and clickthrough | The visits a ranking earned | Mentions and citations | Whether the answer named you or pointed to your page, click or no click |
| Impression share | How often you appeared among the results | Share of voice | Your share of all the brand mentions in the answers |
| Backlinks and authority | The links pointing at your site | Third-party mentions | How many pages the engine reads say your name |
The click, which used to settle every SEO argument, has mostly stopped happening. In Pew Research Center’s March 2025 browsing panel of 900 US adults, Google users clicked a result in 8% of visits with an AI summary, against 15% without one, and clicked a link inside the summary in 1% of visits.2
Both numbers that priced a ranking weakened at once. Ahrefs found a 58% lower clickthrough for the top result on keywords carrying an AI Overview,3 and while ChatGPT took 18 billion messages a week from 700 million users by July 2025,4 none of the engines publishes how often any one question gets asked.
Software buyers got there first. Of the 1,076 B2B software buyers G2 asked in March 2026, 51% of buyers now begin with an AI chatbot more often than with Google, against 29% in April 2025.5 That makes the AI number the one a SaaS board asks about, and unlike a rank, it comes with a margin of error nobody prints.
A Google Ranking Does Not Decide What an AI Engine Says
The AI number is a different measurement of a different contest, because the pages an engine reads are mostly not the pages Google ranks. When Ahrefs ran 15,000 long-tail prompts through the AI assistants and through Google, about 12% of the URLs the assistants cited also sat in Google’s top 10 for the same prompt.6
Qvery is ours, so weigh it accordingly and check the pages underneath before taking our word. It asks buyers’ questions on ChatGPT and Google AI Mode daily, and when it compared AI citations with Google’s organic top 10 across 922 queries, only 13.9% of cited domains ranked there: 10.9% on ChatGPT, 16.7% on Google AI Mode.7
Position still counts for something. In the same study a domain in first position had a 48.8% chance of being cited by an AI engine for that query, falling to 17.3% in tenth.7 A top ranking roughly doubles the odds and still misses more often than it lands.

The authority numbers SEO teams watch explain even less. Qvery set ChatGPT’s naming of SaaS brands against their domain authority in September 2026. On a rank correlation, where 1 means two measures always move together, naming tracked the share of cited roundups listing the brand at 0.64, backlinks rank at 0.35 and organic traffic at 0.27.8

That gap is the mechanism behind GEO = UGC + Mentions, the formula our GEO work runs on. What real users say about you in public, and which independent pages name you, decide the answer well ahead of the links pointing at your site.
And the answer is drawn, not looked up. SparkToro had 600 volunteers run 12 brand-recommendation prompts 2,961 times, and fewer than 1 in 100 pairs of responses gave the same brand list, nearer 1 in 1,000 in the same order.9
Mike Sonders ran 12 B2B software prompts 100 times each through ChatGPT. About 44 brands appeared across a prompt’s 100 answers, about 10 in any single answer, and about 5 in 80% of them or more.10
.png)
Setting the randomness dial to zero does not freeze an answer either. Thinking Machines Lab sampled one prompt 1,000 times at temperature zero on an open model and got 80 different completions.11
Highlights
Read together, they say the question list is the instrument. Too small and the number wanders. Asked daily, it wanders less and keeps whatever bias the list was born with. Drawn from one corner of the category, it measures that corner with great precision.
One Answer Is One Draw, Even From the Same Question
Every visibility number starts with one question asked once, so start there. Qvery asks each tracked question once a day on each engine, and some questions run twice in a day, which gives two clean comparisons across January to August 2026.

The first brand named, the slot most screenshots get cropped to, is no steadier:
The same-day figure matching the next-day figure is the finding. A day passing adds nothing to the disagreement, so most of what people read as AI search volatility is a fresh draw from the same pool.
Every outside test that repeats a question lands in the same place:
That last test scores right answers rather than brand names, and it still leaves more than a quarter of its questions failing at least one of ten identical asks.
Twenty Questions Buy a 20-Point Range
Twenty questions is about what a founder runs on a Thursday afternoon between two calls, so that is where the sampling starts to bite. We drew random subsets from Qvery’s tracked sets, one random day’s answer per question, 400 times for every brand at 5% visibility or more, and set each reading against the brand’s whole month.
Across 3,338 brand-months, with the typical brand near 10%:
For a brand whose real month figure is about 10%, 90% of 20-question readings land anywhere from 0% to 20%. A founder’s reading near the top of that band and a dashboard’s near the bottom can both be looking at the same brand in the same month.

Comparisons are worse, because both numbers wobble. In polling, Pew Research Center’s rule is that a 3-point margin on each candidate becomes about 6 points on the gap between them.16
On Qvery’s sets, two brands 1 to 5 points apart swapped order in 38% of 20-question checks. The month’s leader came out on top in 80% of those checks, against 65% at 10 questions and 94% at 50.1
None of this is news to anyone who has run a survey. Pew’s 2025 methodology table puts the margin on a full sample of 5,022 at ±1.9 points and on a subgroup of 211 at ±8.9.17 Same survey, same questions, and only the number of people answering changed.
OtterlyAI ran the same subset test on its own 320-prompt dataset, one brand over 30 days, and found the same shape:
A Few Hundred Questions Is Where a Single Reading Holds
Qvery’s sets stop at about 100 questions, so the curve past that comes from our own instrument, the way we measure AI search visibility for GEO partners. Each partner gets one weekly study of unbranded buyer questions, asked on ChatGPT and Google AI Mode, every answer kept.
In the B2B software category we measure for one partner, each product line carries 1,000 questions, enough to draw subsets far larger than any tracked set allows:
Two instruments in two categories agree at 20 questions, with about half the readings landing close. The curve flattens past a hundred and is nearly flat by two hundred, which is where a single reading becomes a number you can put in front of a board.

Asking Every Day Fixes the Draw, Not the Question List
The obvious fix for twenty noisy questions is to ask them every day, and it fixes one of the two things wrong with them. A one-off reading wobbles because of which questions were in the set and which answer each one happened to draw. On Qvery’s sets, 41% of the noise came from the questions and 59% from the draw.1
That split sets the price of a usable number. For the median brand, a 95% margin of five points takes about 141 questions asked once, or about 62 asked every day for a month, and even a loose ten-point margin takes 35 questions in a single pass.1

A month of daily runs is also why a dashboard’s daily number deserves no reaction. On sets of about 100 questions, the median brand’s daily figure moved 2.1 points from one day to the next and spanned 11.5 points across the month.1
The noise is not the whole story. In Qvery’s study of ecommerce share of voice, two runs of an identical question overlapped at 0.3329 on cited sources, about a third, against 0.0198, about one in fifty, for two different questions.19
A rerun disagrees with itself and still agrees with itself far more than with any other question, so a large, repeated set measures something real.
A preprint on Swiss-German campaigns reached the same place from the other side. One brand’s detection rate needed 7 same-day runs before its 95% interval narrowed to about ±16 points, and 24 days of daily data to bring its standard error under 0.05.20
Which Questions You Ask Moves the Number More Than How Many
Everything so far assumes the questions were drawn at random from the category, and nobody’s are. Inside the sets Qvery tracks, topics group the questions on one part of the category, and the same brand in the same month read a median 26 points apart between its best and worst topic, 28 in the software sets.1
In the median case the best topic read 31% and the worst under 1%.1 Our panel shows it at scale: across the nine product lines of one software category, a brand averaging 5% or more read a median 40 points apart between its best and worst line.12
Ask only about the corner of the category you are strongest in, and you will measure that corner accurately and call it your visibility.
| What changed | How far a brand’s reading moved | Measured on |
|---|---|---|
| The next day’s answers, same full set | 2.1 points, median daily move | Qvery’s tracked sets |
| Twenty questions instead of the full month | 4.5 points, median miss | Qvery’s tracked sets |
| A different topic in the same set | 26 points, median best minus worst | Qvery’s tracked sets |
| A different product line in the same category | 40 points, median best minus worst | Our panel |
Wording moves it too, in two tests run on paired phrasings:
The rates held while the rosters changed. A number can look steady while it quietly measures different questions.
AAPOR’s disclosure standards ask probability surveys to report their sampling error, and allow a non-probability sample a precision figure only beside a description of the model behind it.23 A question set is a non-probability sample by construction, so an honest AI visibility number says how its questions were chosen.
That build is where a GEO partnership with Empact Partners starts. It opens with an audit of what the engines say in the partner’s category, and the question set is built inside that audit, before any number reaches a board:
How the questions get written stays in-house. What the partner gets every week is the band around every number, read by the senior consultant who owns the partnership over the quarters GEO takes to move.
Report the Sample Size With Every AI Visibility Number
A visibility reading from fewer than about 140 questions asked once is an anecdote, and you can check that against your own data. A move of under five points on fewer than about 60 questions asked daily for a month is noise, and a gap between two brands needs roughly twice the margin of either number.
Put four labels beside any number you circulate: the engine, the window, how many questions and how many runs. A visibility figure without them is a thermometer reading with no units, precise to the decimal and useful for nothing.
At Empact, this is what we specialize in: AI engine optimization. If you believe you need help optimizing for search engines, reach out. We can help you appear more and get recommended more in search.
Sources
- Qvery, our sister company: 513,813 answers to the query sets it tracks daily on ChatGPT and Google AI Mode, 22 sets (18 of them software companies), January to August 2026, random subsets computed for this article, read 5 October 2026.
- Pew Research Center, “Google users are less likely to click on links when an AI summary appears in the results”, 2025.
- Ahrefs, “Update: AI Overviews Reduce Clicks by 58%”, 2026.
- Chatterji, Cunningham, Deming, Hitzig, Ong, Shan and Wadman, “How People Use ChatGPT”, NBER Working Paper 34255, 2025.
- G2, “The Answer Economy: How AI Search Is Rewiring B2B Software Buying”, PR Newswire, 2026.
- Ahrefs, “Only 12% of AI Cited URLs Rank in Google’s Top 10 for the Original Prompt”, 2025.
- Qvery, “AI Engine Citations vs Google Organic SERPs: Only 13.9% Overlap”, 2026.
- Qvery, “Domain Authority vs ChatGPT Naming for SaaS Brands”, 2026.
- SparkToro, “NEW Research: AIs are highly inconsistent when recommending brands or products; marketers should take care when tracking AI visibility”, 2026.
- Mike Sonders, “What repeated ChatGPT runs reveal about brand visibility”, Search Engine Land, 2026.
- Thinking Machines Lab, “Defeating Nondeterminism in LLM Inference”, 2025.
- The Empact Panel, our own: a B2B software category we measure weekly for a partner, nine product lines of 1,000 unbranded buyer questions each, asked on ChatGPT and Google AI Mode on 23 and 30 September 2026, read 5 October 2026.
- Detailed.com, “September 2026: I Compared Prompt Tracking Across 8 AI Visibility Tools”, 2026.
- SE Ranking, “AI Mode research: Volatility, source patterns, and differences from AIO and organic results”, 2025.
- Cicek, Ulu, Uslay and Karniouchina, “Unstable Intelligence: GenAI Struggles with Accuracy and Consistency”, Rutgers Business Review, 2025.
- Pew Research Center, “5 key things to know about the margin of error in election polls”, 2016.
- Pew Research Center, “Social Media Use in 2025: Methodology”, 2025.
- OtterlyAI, “AI Search Visibility: How Stable Are Brand Mentions and Citations?”, undated, read 5 October 2026.
- Qvery, “How To Measure Ecommerce Share Of Voice”, 2026.
- Schulte, Bleeker and Kaufmann, “Don’t Measure Once: Measuring Visibility in AI Search (GEO)”, arXiv, 2026.
- Jack, Lehman, Maloney and Xu, “Paraphrase Brittleness in Production Retrieval-Augmented Commercial Recommendation”, arXiv, 2026.
- Qvery, “How to Generate LLM Tracking Queries Automatically”, 2026.
- American Association for Public Opinion Research, “Disclosure Standards”, AAPOR Code, revised 2021.

.png)
