Why We Ask the Same Question 8 Times a Week
Ask an AI engine the same question twice and it may recommend different products. So a tool that asks once a week and reports that one answer as a verdict can easily be reporting chance.
Bevizio asks each question 8 times a week and reports how many of the answers named you. This post explains the research that motivates it, exactly how the asking works, and what it still can’t tell you.
What the research found
In November and December 2025, SparkToro and Gumshoe.ai had 600 volunteers run 12 prompts through each of ChatGPT, Claude and Google’s AI Overviews (or AI Mode, when no Overview appeared), 2,961 runs in total. Each prompt was run 60 to 100 times. They published it in January 2026, with the raw responses available to browse.
The headline numbers:
- For ChatGPT and Google, there is less than a 1 in 100 chance that any two responses to the same question give the same list of brands. Claude was only slightly more likely to repeat itself.
- Getting two lists in the same order was closer to 1 in 1,000.
- List length varied too: sometimes 2 or 3 recommendations, and just as often 10 or more.
But the study also shows why a rate is worth measuring. In a separate headphones test of 994 responses, Bose, Sony, Sennheiser and Apple showed up 55% to 77% of the time. The lists varied, but the leading brands kept coming back. Rand Fishkin had started out sceptical of visibility metrics. He concluded that a visibility percentage, measured across dozens to hundreds of prompts and run multiple times, is a reasonable metric.
He also wrote that any tool that gives you a “ranking position in AI” is “full of baloney”.
How Bevizio asks
Each question is asked 8 times a week on every plan: ChatGPT three times, Claude twice, and Gemini, Perplexity and Google AI Overviews once each. The asks are spread across the week, not bunched into one run. ChatGPT and Claude get more because their answers moved most from one week to the next in our own earlier weekly data. The other three were steadier. The research supports counting a rate. The schedule itself is our own choice. The two measures differ: SparkToro looked at whether two responses gave the same list of brands, while our own data looked at whether a brand’s named or not-named result flipped from one week to the next.
Results are reported as a count, for example “named in 9 of 60 answers”, over 7, 28 or 84 days, with 28 as the default. There are rules about how much a count can say:
- For one question on one engine, Bevizio shows the raw record and no verdict.
- Words like “most” or “some” need at least 12 answers behind them.
- Percentages in the summary counts appear only on pooled numbers with at least 20 answers. The trend charts plot a weekly point once it has at least 10 answers.
- A data-quality badge reads “Solid” only from 100 answers, and “Thin data” below that.
A change is only reported when it clears a stricter test. Bevizio compares the last 14 days with the 28 days before. The gap has to be at least 40 percentage points and pass a significance test at p < 0.01, meaning a gap that size would be unlikely to arise by chance, with at least 8 answers on each side. It is checked for each question across all engines, and for each engine across all questions, never for one question on one engine. That is why the first change reports can appear only after about three weeks.
What 8 asks a week can’t do
SparkToro ran each prompt 60 to 100 times. Bevizio asks ChatGPT 3 times a week per question, so one question takes about 20 weeks to reach 60 ChatGPT answers. The longest window Bevizio shows is 84 days. That holds at most 36 ChatGPT answers for a question, so one question never reaches 60 on its own. That is why Bevizio shows the raw record there, and why the pooled view across all your questions is where percentages first appear. The SparkToro authors also left open how many runs are enough for statistically sound answers, so we don’t have a settled figure for that either. Not everyone agrees how many runs are enough: one vendor’s own study says once a day is close to ten times a day, and it is discussed in AI visibility tools compared.
There is a second limit. Bevizio collects answers through the engines’ APIs and search results. Those can differ from what a person sees in the app, because of account, location, model version and time of day. The SparkToro authors say they did not research exhaustively whether API calls mimic the variety real users get. How closely API answers track what people see is not fully settled. Bevizio isn’t affiliated with OpenAI, Anthropic, Google or Perplexity.
A new product has one more thing to consider: whether the engines have found its pages yet. What the companies say about that is in how long until AI engines find your new website.
Is this the same as SEO?
For Google’s own AI features, Google says “there are no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary”, and that “the best practices for SEO remain relevant”. It also says these features may use “query fan-out”, issuing several related searches across subtopics to build one answer.
That is Google’s statement about its own features. It doesn’t cover ChatGPT, Claude or Perplexity. Whatever you do to be named more often, you need a count to know whether it worked, and that count needs many answers, not one. To see one answer from each of three engines (ChatGPT, Perplexity and Google AI Overviews) for your own product first, run the free check. To see what the counts look like over time, the live demo is a real account. It is asked daily until 18 October 2026, so its counts build faster than the schedule above. The step-by-step version of checking by hand is in Does ChatGPT recommend your product? For why an engine might name a competitor instead of you, see why AI engines name your competitor instead of you.
Drafted with AI assistance. Every factual claim was checked against the sources above.
See whether ChatGPT, Perplexity and Google AI Overviews name your product. No account needed.
Run a free check