SEO

44 AI Visibility Sample Size Statistics for SaaS

44 figures on sampling AI answers: a 20-question check lands within five points of a brand’s monthly figure just 55% of the time.

Author:
Vlad Shvets
Contributors
Vlad Shvets
Date:
October 5, 2026

Twenty questions put to ChatGPT or Google AI Mode about your category give your brand a visibility number. Across 513,813 answers from the query sets our sister company Qvery tracks, that number landed within five points of the brand’s real monthly figure 55% of the time,1 and nothing on the screen says which kind of reading you hold.

The habit is borrowed from search. A Google ranking is one published list, so one look tells you where you stand. An AI answer is drawn fresh every time, which makes the visibility number a poll with a margin of error, and nobody selling a dashboard prints it. Here it is, measured, after the SEO metrics it gets mistaken for.

Empact Partners, a go-to-market consultancy working inside software companies’ marketing teams, sizes that question set in the first weeks of every Generative Engine Optimization (GEO) partnership, our work on getting a brand named in AI answers about its category. If you would rather hand the measurement to us, book a call with me.

SEO Metrics Read a Published List, AI Visibility Metrics Sample Answers

SEO metrics grew up around a ranked list Google publishes for each query, which everyone who searches it sees in roughly the same order. Position, impressions and clicks hang off that list, and search volume says what each query is worth. Check your rank on Monday and you could still trust it on Tuesday.

An AI engine publishes no list. It writes an answer, names a handful of brands, cites a few pages, then writes a different answer for the next person. So AI visibility metrics count answers, and the two that matter most, brand visibility and share of voice, divide the same answers two ways.

SEO metricWhat it readsAI visibility counterpartWhat it reads
Rank positionYour place in one list Google publishes for a queryVisibilityThe share of sampled answers that name you
Search volumeHow often a query is searched, published by GoogleNothing publishedThe question set somebody chose stands in for demand
Clicks and clickthroughThe visits a ranking earnedMentions and citationsWhether the answer named you or pointed to your page, click or no click
Impression shareHow often you appeared among the resultsShare of voiceYour share of all the brand mentions in the answers
Backlinks and authorityThe links pointing at your siteThird-party mentionsHow many pages the engine reads say your name

The click, which used to settle every SEO argument, has mostly stopped happening. In Pew Research Center’s March 2025 browsing panel of 900 US adults, Google users clicked a result in 8% of visits with an AI summary, against 15% without one, and clicked a link inside the summary in 1% of visits.2

Both numbers that priced a ranking weakened at once. Ahrefs found a 58% lower clickthrough for the top result on keywords carrying an AI Overview,3 and while ChatGPT took 18 billion messages a week from 700 million users by July 2025,4 none of the engines publishes how often any one question gets asked.

Software buyers got there first. Of the 1,076 B2B software buyers G2 asked in March 2026, 51% of buyers now begin with an AI chatbot more often than with Google, against 29% in April 2025.5 That makes the AI number the one a SaaS board asks about, and unlike a rank, it comes with a margin of error nobody prints.

A Google Ranking Does Not Decide What an AI Engine Says

The AI number is a different measurement of a different contest, because the pages an engine reads are mostly not the pages Google ranks. When Ahrefs ran 15,000 long-tail prompts through the AI assistants and through Google, about 12% of the URLs the assistants cited also sat in Google’s top 10 for the same prompt.6

Qvery is ours, so weigh it accordingly and check the pages underneath before taking our word. It asks buyers’ questions on ChatGPT and Google AI Mode daily, and when it compared AI citations with Google’s organic top 10 across 922 queries, only 13.9% of cited domains ranked there: 10.9% on ChatGPT, 16.7% on Google AI Mode.7

Position still counts for something. In the same study a domain in first position had a 48.8% chance of being cited by an AI engine for that query, falling to 17.3% in tenth.7 A top ranking roughly doubles the odds and still misses more often than it lands.

Bar chart of the chance a domain ranking in Google’s top 10 is also cited by ChatGPT or Google AI Mode for the same query: 39.4% at positions 1 to 3, 24.3% at positions 4 to 7 and 19.1% at positions 8 to 10, across 922 queries.
Qvery, 922 queries, AI engine citations against Google’s organic top 10. Even the top three positions are cited in fewer than four cases in ten.

The authority numbers SEO teams watch explain even less. Qvery set ChatGPT’s naming of SaaS brands against their domain authority in September 2026. On a rank correlation, where 1 means two measures always move together, naming tracked the share of cited roundups listing the brand at 0.64, backlinks rank at 0.35 and organic traffic at 0.27.8

Horizontal bar chart of how closely four measures tracked ChatGPT’s naming of SaaS brands, as rank correlations: roundup occupancy 0.64, referring domains 0.38, backlinks rank 0.35 and US organic traffic 0.27.
Qvery, ChatGPT, September 2026. Roundup occupancy is the share of cited roundup pages that list the brand; its lead over referring domains was inconclusive.

That gap is the mechanism behind GEO = UGC + Mentions, the formula our GEO work runs on. What real users say about you in public, and which independent pages name you, decide the answer well ahead of the links pointing at your site.

And the answer is drawn, not looked up. SparkToro had 600 volunteers run 12 brand-recommendation prompts 2,961 times, and fewer than 1 in 100 pairs of responses gave the same brand list, nearer 1 in 1,000 in the same order.9

Mike Sonders ran 12 B2B software prompts 100 times each through ChatGPT. About 44 brands appeared across a prompt’s 100 answers, about 10 in any single answer, and about 5 in 80% of them or more.10

An AI visibility number is a poll of answers, and every poll has a margin of error, whether anyone prints it or not.

Setting the randomness dial to zero does not freeze an answer either. Thinking Machines Lab sampled one prompt 1,000 times at temperature zero on an open model and got 80 different completions.11

Highlights

55% of 20-question checks land within five points of the brand’s figure for the month, across the query sets Qvery tracks.1
A brand whose month figure is about 10% reads anywhere from 0% to 20% on 20 questions, in 90% of checks.1
A second run of the same query names 49% of the brands the first run named, the same day or the next.1
41% of the noise in a one-off reading comes from which questions were asked, and daily reruns cannot remove it.1
A reading good to five points, 95% of the time, takes about 141 questions asked once, or 62 asked every day for a month.1
In our own panel, 98% of 200-question readings land within five points of the full 1,000 questions, against 48% at 20.12
A 20-question check names the category leader in 80% of checks on Qvery’s sets1 and 82% in our panel.12
The same brand in the same month reads 26 points apart between its best and worst topic.1

Read together, they say the question list is the instrument. Too small and the number wanders. Asked daily, it wanders less and keeps whatever bias the list was born with. Drawn from one corner of the category, it measures that corner with great precision.

One Answer Is One Draw, Even From the Same Question

Every visibility number starts with one question asked once, so start there. Qvery asks each tracked question once a day on each engine, and some questions run twice in a day, which gives two clean comparisons across January to August 2026.

Statcard of three figures from runs of the same query: 49% of brands named again by a second run the same day, 49% named again by the next day’s run, and 5% of same-day reruns repeating the whole brand list.
Qvery, 9,623 same-day and 463,886 next-day pairs of runs of one query.

The first brand named, the slot most screenshots get cropped to, is no steadier:

The first-named brand matched in 47% and 49% of pairs, same day and next day.1
The whole brand list repeated in 5% and 6% of pairs.1
Of the questions that named a brand at least once in a month, 4.9% named it on every run, and 71% on half the runs or fewer.1

The same-day figure matching the next-day figure is the finding. A day passing adds nothing to the disagreement, so most of what people read as AI search volatility is a fresh draw from the same pool.

Every outside test that repeats a question lands in the same place:

Detailed.com tracked 1,300 prompts daily for four weeks, and any two days returned an identical ChatGPT brand set in 0.3% of pairs, while 41% of prompts never repeated a set at all.13
SE Ranking ran 10,000 keywords through Google AI Mode three times in one day, and the cited URLs overlapped 9.2% on average.14
Rutgers Business Review’s authors asked ChatGPT the same 719 true-or-false research questions ten times each, and 72.9% came back right on all ten runs in 2025.15

That last test scores right answers rather than brand names, and it still leaves more than a quarter of its questions failing at least one of ten identical asks.

Vlad Shvets
Founder @ Empact Partners
A screenshot of your brand in a ChatGPT answer proves the brand is in the engine’s pool and says nothing about how often the engine draws it. Report the rate across many questions and many runs, never the screenshot, and never the first brand named, which changes on about half of reruns.

Twenty Questions Buy a 20-Point Range

Twenty questions is about what a founder runs on a Thursday afternoon between two calls, so that is where the sampling starts to bite. We drew random subsets from Qvery’s tracked sets, one random day’s answer per question, 400 times for every brand at 5% visibility or more, and set each reading against the brand’s whole month.

Across 3,338 brand-months, with the typical brand near 10%:

35%, 55% and 60% of readings landed within five points, at 10, 20 and 25 questions.1
The typical miss was 7.4, 4.5 and 3.9 points.1
At 10 questions the typical miss was 94% of the brand’s own figure, and at 20 still 39%.1
The 18 software companies’ sets read the same, at 56% within five points on 20 questions.1

For a brand whose real month figure is about 10%, 90% of 20-question readings land anywhere from 0% to 20%. A founder’s reading near the top of that band and a dashboard’s near the bottom can both be looking at the same brand in the same month.

Spread chart of where 90% of one-off visibility readings land for a brand whose monthly figure is about 10%: 0% to 30% on 10 questions, 0% to 20% on 20 and on 25 questions, 4% to 16% on 50, and 6% to 13% for the whole set of about 100 on one day.
Qvery, 738 brand-months, one random day’s answer per question, 400 draws each. The 50-question and whole-set rows come from sets of about 100 questions.

Comparisons are worse, because both numbers wobble. In polling, Pew Research Center’s rule is that a 3-point margin on each candidate becomes about 6 points on the gap between them.16

On Qvery’s sets, two brands 1 to 5 points apart swapped order in 38% of 20-question checks. The month’s leader came out on top in 80% of those checks, against 65% at 10 questions and 94% at 50.1

Vlad Shvets
Founder @ Empact Partners
A twenty-question check in ChatGPT tells a SaaS team whether its brand exists in the answers. It cannot say whether the brand is beating a close competitor, because a lead of a few points sits inside the noise at that size. Spend those twenty questions finding where you are absent.

None of this is news to anyone who has run a survey. Pew’s 2025 methodology table puts the margin on a full sample of 5,022 at ±1.9 points and on a subgroup of 211 at ±8.9.17 Same survey, same questions, and only the number of people answering changed.

OtterlyAI ran the same subset test on its own 320-prompt dataset, one brand over 30 days, and found the same shape:

39% to 71% held 90% of the brand’s readings at 10 prompts.18
49% to 62% held them at 50 prompts.18
51% to 60% held them at 100 prompts.18

A Few Hundred Questions Is Where a Single Reading Holds

Qvery’s sets stop at about 100 questions, so the curve past that comes from our own instrument, the way we measure AI search visibility for GEO partners. Each partner gets one weekly study of unbranded buyer questions, asked on ChatGPT and Google AI Mode, every answer kept.

In the B2B software category we measure for one partner, each product line carries 1,000 questions, enough to draw subsets far larger than any tracked set allows:

At 20 questions, 48% of readings landed within five points of the full line, 88% at 100 and 98% at 200.12
The typical miss fell from 5.3 points at 20 questions to 2.2 at 100 and 1.5 at 200.12
The line’s leader came out on top in 82% of 20-question draws, 95% at 50 and 99% at 100.12
On the full 1,000 questions, a brand’s figure moved a median 1.4 points from one week to the next.12

Two instruments in two categories agree at 20 questions, with about half the readings landing close. The curve flattens past a hundred and is nearly flat by two hundred, which is where a single reading becomes a number you can put in front of a board.

Line chart of the share of visibility readings within five points of the full 1,000-question line, by number of questions asked: 33% at 10, 48% at 20, 53% at 25, 71% at 50, 88% at 100 and 98% at 200.
A B2B software category we measure for a partner, 533 brand-line readings on ChatGPT and Google AI Mode, 23 and 30 September 2026.

Asking Every Day Fixes the Draw, Not the Question List

The obvious fix for twenty noisy questions is to ask them every day, and it fixes one of the two things wrong with them. A one-off reading wobbles because of which questions were in the set and which answer each one happened to draw. On Qvery’s sets, 41% of the noise came from the questions and 59% from the draw.1

That split sets the price of a usable number. For the median brand, a 95% margin of five points takes about 141 questions asked once, or about 62 asked every day for a month, and even a loose ten-point margin takes 35 questions in a single pass.1

Dumbbell chart of the questions needed for a five-point margin at 95%, asked once against asked daily for a month: 141 and 62 for all brands, 139 and 65 in software sets, 139 and 60 on ChatGPT, 142 and 65 on Google AI Mode, and 175 and 62 for the set’s own tracked brand.
Qvery, medians across brand-months, projected from each brand’s measured variance. A brand nearer 50% visibility needs more questions, one nearer zero fewer.

A month of daily runs is also why a dashboard’s daily number deserves no reaction. On sets of about 100 questions, the median brand’s daily figure moved 2.1 points from one day to the next and spanned 11.5 points across the month.1

Vlad Shvets
Founder @ Empact Partners
Running the same twenty prompts every day for a month fixes the engine’s draw and leaves the twenty prompts exactly as unrepresentative as they were on day one. Build the question list before you track anything, because no amount of repetition repairs a list that asks about the wrong part of the category.

The noise is not the whole story. In Qvery’s study of ecommerce share of voice, two runs of an identical question overlapped at 0.3329 on cited sources, about a third, against 0.0198, about one in fifty, for two different questions.19

A rerun disagrees with itself and still agrees with itself far more than with any other question, so a large, repeated set measures something real.

A preprint on Swiss-German campaigns reached the same place from the other side. One brand’s detection rate needed 7 same-day runs before its 95% interval narrowed to about ±16 points, and 24 days of daily data to bring its standard error under 0.05.20

Which Questions You Ask Moves the Number More Than How Many

Everything so far assumes the questions were drawn at random from the category, and nobody’s are. Inside the sets Qvery tracks, topics group the questions on one part of the category, and the same brand in the same month read a median 26 points apart between its best and worst topic, 28 in the software sets.1

In the median case the best topic read 31% and the worst under 1%.1 Our panel shows it at scale: across the nine product lines of one software category, a brand averaging 5% or more read a median 40 points apart between its best and worst line.12

Ask only about the corner of the category you are strongest in, and you will measure that corner accurately and call it your visibility.

What changedHow far a brand’s reading movedMeasured on
The next day’s answers, same full set2.1 points, median daily moveQvery’s tracked sets
Twenty questions instead of the full month4.5 points, median missQvery’s tracked sets
A different topic in the same set26 points, median best minus worstQvery’s tracked sets
A different product line in the same category40 points, median best minus worstOur panel

Wording moves it too, in two tests run on paired phrasings:

A preprint testing paraphrases on OpenAI and Anthropic models found two rewordings of one buying question shared recommended brands at an overlap of 0.288, or 0.135 when the rewording added a constraint, against 0.50 to 0.61 for the same prompt asked again, where 1 means identical lists.21
On Qvery’s paired tracking questions, one added constraint left the source rates within a median 2.8 points, while the two phrasings’ top ten sources shared only 3.22

The rates held while the rosters changed. A number can look steady while it quietly measures different questions.

AAPOR’s disclosure standards ask probability surveys to report their sampling error, and allow a non-probability sample a precision figure only beside a description of the model behind it.23 A question set is a non-probability sample by construction, so an honest AI visibility number says how its questions were chosen.

That build is where a GEO partnership with Empact Partners starts. It opens with an audit of what the engines say in the partner’s category, and the question set is built inside that audit, before any number reaches a board:

Coverage: every part of the category a buyer asks about, with enough questions in each that no single topic speaks for the brand.
Unbranded questions: the buyer’s question as the buyer types it, never the brand’s name planted in it.
Both engines, every market: ChatGPT and Google AI Mode, in each country the partner sells into.
Frozen wording: the strings stay fixed, so month eight compares with month one.
Size and repeats: enough questions and runs that a five-point move clears the noise.

How the questions get written stays in-house. What the partner gets every week is the band around every number, read by the senior consultant who owns the partnership over the quarters GEO takes to move.

Report the Sample Size With Every AI Visibility Number

A visibility reading from fewer than about 140 questions asked once is an anecdote, and you can check that against your own data. A move of under five points on fewer than about 60 questions asked daily for a month is noise, and a gap between two brands needs roughly twice the margin of either number.

Put four labels beside any number you circulate: the engine, the window, how many questions and how many runs. A visibility figure without them is a thermometer reading with no units, precise to the decimal and useful for nothing.

At Empact, this is what we specialize in: AI engine optimization. If you believe you need help optimizing for search engines, reach out. We can help you appear more and get recommended more in search.

Sources

  1. Qvery, our sister company: 513,813 answers to the query sets it tracks daily on ChatGPT and Google AI Mode, 22 sets (18 of them software companies), January to August 2026, random subsets computed for this article, read 5 October 2026.
  2. Pew Research Center, “Google users are less likely to click on links when an AI summary appears in the results”, 2025.
  3. Ahrefs, “Update: AI Overviews Reduce Clicks by 58%”, 2026.
  4. Chatterji, Cunningham, Deming, Hitzig, Ong, Shan and Wadman, “How People Use ChatGPT”, NBER Working Paper 34255, 2025.
  5. G2, “The Answer Economy: How AI Search Is Rewiring B2B Software Buying”, PR Newswire, 2026.
  6. Ahrefs, “Only 12% of AI Cited URLs Rank in Google’s Top 10 for the Original Prompt”, 2025.
  7. Qvery, “AI Engine Citations vs Google Organic SERPs: Only 13.9% Overlap”, 2026.
  8. Qvery, “Domain Authority vs ChatGPT Naming for SaaS Brands”, 2026.
  9. SparkToro, “NEW Research: AIs are highly inconsistent when recommending brands or products; marketers should take care when tracking AI visibility”, 2026.
  10. Mike Sonders, “What repeated ChatGPT runs reveal about brand visibility”, Search Engine Land, 2026.
  11. Thinking Machines Lab, “Defeating Nondeterminism in LLM Inference”, 2025.
  12. The Empact Panel, our own: a B2B software category we measure weekly for a partner, nine product lines of 1,000 unbranded buyer questions each, asked on ChatGPT and Google AI Mode on 23 and 30 September 2026, read 5 October 2026.
  13. Detailed.com, “September 2026: I Compared Prompt Tracking Across 8 AI Visibility Tools”, 2026.
  14. SE Ranking, “AI Mode research: Volatility, source patterns, and differences from AIO and organic results”, 2025.
  15. Cicek, Ulu, Uslay and Karniouchina, “Unstable Intelligence: GenAI Struggles with Accuracy and Consistency”, Rutgers Business Review, 2025.
  16. Pew Research Center, “5 key things to know about the margin of error in election polls”, 2016.
  17. Pew Research Center, “Social Media Use in 2025: Methodology”, 2025.
  18. OtterlyAI, “AI Search Visibility: How Stable Are Brand Mentions and Citations?”, undated, read 5 October 2026.
  19. Qvery, “How To Measure Ecommerce Share Of Voice”, 2026.
  20. Schulte, Bleeker and Kaufmann, “Don’t Measure Once: Measuring Visibility in AI Search (GEO)”, arXiv, 2026.
  21. Jack, Lehman, Maloney and Xu, “Paraphrase Brittleness in Production Retrieval-Augmented Commercial Recommendation”, arXiv, 2026.
  22. Qvery, “How to Generate LLM Tracking Queries Automatically”, 2026.
  23. American Association for Public Opinion Research, “Disclosure Standards”, AAPOR Code, revised 2021.

Ready
To Connect?

Let's Partner