LLM output consistency: why ChatGPT names different brands
LLM output consistency is lower than most AI-visibility tools assume: a 2,961-run study found ChatGPT repeats the same brand list under 1% of the time.
LLM output consistency is lower than most AI-visibility tools assume: a 2,961-run study found ChatGPT repeats the same brand list under 1% of the time.

Ask ChatGPT to recommend a chef's knife twice in the same afternoon and you can get two different lists. That is not a bug someone will patch. A November-December 2025 study from SparkToro and the AI-visibility tracker Gumshoe.ai ran the same 12 prompts 2,961 times across ChatGPT, Claude, and Google AI Overviews and found the exact same brand list came back in under 1% of runs. LLM output consistency, in plain terms, is how often a model gives the same answer to the same question, and on brand recommendations it is low enough that a single check tells you almost nothing.
The headline finding is blunt: ask an AI model the same question twice and you are very unlikely to get the same list back, let alone the same order. That single fact is the reason a lot of AI-visibility reporting is quietly unreliable, and it is worth understanding the study's methodology before deciding what to do about it.
SparkToro and Gumshoe.ai recruited roughly 600 volunteers and ran the study over November and December 2025, producing 2,961 total prompt runs across ChatGPT, Claude, and Google AI Overviews, according to SparkToro's writeup of the research. The prompt set covered 12 recommendation-style questions spanning categories like chef's knives, headphones, cancer care hospitals, digital marketing consultants, science fiction novels, and cloud computing providers, with each prompt run 60 to 100 times per platform, per Search Engine Journal's coverage of the study.
The study went a step further than most AI-visibility research by testing whether "the same question" even means the same thing when 142 different people write it. Those 142 volunteers each wrote their own version of a headphone-recommendation prompt, and when researchers measured how similar the resulting prompts were to each other, they scored just 0.081 semantic similarity, per SparkToro's research post. In other words, real users rarely phrase a query the same way twice, which means the instability compounds before the model has even started generating a response.
Across all 2,961 runs, the exact same list of recommended brands came back in under 1% of cases, and the exact same order in under 0.1%, according to the same study. That is the number that should recalibrate anyone treating a single AI-search screenshot as proof of standing. It does not mean brands are invisible or that the responses are random noise, though. Visibility rate and exact-list repetition are two different, both valid, measurements, and they tell opposite stories. Top headphone brands, Bose, Sony, Sennheiser, and Apple, still appeared in 55 to 77% of 994 total responses despite the list order shuffling almost every time, per Search Engine Journal's report. A brand can be reliably present while the list around it, and its position within that list, changes on nearly every run.
Some individual results were even more stable. City of Hope, a cancer-treatment hospital, showed up in 97% of ChatGPT's responses to its category prompt (69 of 71 runs), according to SparkToro's data. That is a useful reminder that "LLM output consistency is low" is a statement about the aggregate, not a guarantee that every brand-model pairing is equally chaotic. Some queries land on a dominant answer almost every time. Most don't, and you cannot tell which kind of query you have from one run.
None of this is a flaw ChatGPT, Claude, or Google are going to fix, because the mechanisms producing the variation are also the mechanisms that make these models useful in the first place.
A large language model generates each token by sampling from a probability distribution over likely next tokens, not by always picking the single highest-probability option. That sampling step, governed by a temperature parameter, is what keeps responses from reading like the same canned paragraph every time. It is a deliberate design choice, not an oversight, and it means two runs of an identical prompt can diverge from the very first sentence, well before the model has committed to which brands to mention.
ChatGPT, Claude with browsing, and Google AI Overviews don't answer purely from a frozen training snapshot. Many responses pull from a live retrieval step that queries the web or a search index in real time. If the underlying set of pages a model can retrieve from shifts between two runs of the same prompt, even by minutes, the candidate brands available to mention can shift with it. A prompt run at 9am and the same prompt run at 2pm are not guaranteed to be querying the same retrieved context.
ChatGPT's memory and personalization features mean two people running the identical prompt in two separate accounts are not running a controlled experiment, even if the wording matches exactly. Prior conversation history, saved preferences, and account-level signals can all nudge which sources and brands surface. That is one more reason a single tester's screenshot of an AI answer is a weak proxy for what any other user, or that same user tomorrow, would see.
Underlying models and their retrieval indexes update on their own schedule, independent of anything a marketer does. A prompt that reliably surfaced one set of brands last month can return a different set today simply because the model version or its retrieval index changed underneath it. This is a slower-moving source of variation than sampling or live retrieval, but it means even a well-designed tracking method has a shelf life on its baseline.
If the SparkToro/Gumshoe.ai data covers brand-recommendation lists, a second study from Washington State University points at a related but distinct problem: LLM output consistency is low even on straightforward factual judgments, not just open-ended recommendations.
WSU researchers tested 719 business-research hypotheses, each phrased as a true-or-false claim, running 10 identical repeated prompts per hypothesis and comparing ChatGPT-3.5 (2024) against ChatGPT-5 mini (2025), according to WSU Insider's coverage of the study. Associate Professor Mesut Cicek, who led the research, described the failure mode plainly: "We used 10 prompts with the same exact question. Everything was identical. It would answer true. Next, it says it's false," per the same WSU report. ChatGPT gave a consistent answer to only 73% of those identical, 10-times-repeated prompts, meaning more than a quarter of the time, asking the exact same true-or-false question again produced a flipped answer.
The study also tracked accuracy separately from consistency, and the two moved independently. Overall accuracy on the hypothesis set rose from 76.5% with ChatGPT-3.5 in 2024 to 80% with ChatGPT-5 mini in 2025, per WSU's writeup. A newer, more accurate model generation did not close the consistency gap. This matters beyond brand tracking: it means the instability documented in the SparkToro data is not a quirk specific to open-ended recommendation prompts. It shows up on binary factual judgments too, and it has persisted across at least one full model generation.
Put the two studies together and a pattern emerges that should worry anyone buying, or selling, a single "AI ranking position." Rand Fishkin, who co-authored the SparkToro research, did not soften the conclusion: "Any tool that gives a 'ranking position in AI' is full of baloney," he said, per SparkToro's own post on the findings. That line carries extra weight because Fishkin's co-author on the study runs Gumshoe.ai, an AI-visibility tracking tool. Even a party with a commercial interest in the visibility-tracking category is on record saying a raw ranking number is unreliable.
The reason follows directly from the mechanisms above. A ranking position implies a stable underlying order that a single measurement can reveal. But sampling temperature, live retrieval drift, personalization, and model updates all inject variation independently, so what a "position" tool actually reports is one draw from a moving target, dressed up as a fixed rank. It is the AI-search equivalent of calling a single coin flip a batting average.
None of this means AI-visibility measurement is worthless, only that it has to account for the noise instead of ignoring it. The SparkToro/Gumshoe.ai authors' own conclusion, again notable because they partly sell a tracking tool, is that visibility percentage across many repeated runs is the defensible metric, not a single rank.
You do not need enterprise tooling to get a directionally useful read. A minimum viable version looks like this:
This is close to the manual prompt-log approach the SparkToro researchers themselves used to arrive at a defensible number, just scaled down to what a small team can sustain by hand or in a spreadsheet.
The distinction that matters is appearance rate, not position. Instead of asking "where do we rank in ChatGPT," ask "in what percentage of runs of our target prompts do we appear at all." That framing survives the exact instability the two studies documented: it does not assume a stable order exists to be measured, and it treats each run as one sample toward an estimate rather than a verdict. City of Hope's 97% appearance rate and the top headphone brands' 55-77% range are both visibility percentages in this sense, and both are meaningful numbers precisely because they came from dozens of runs, not one.
If a single snapshot is statistical noise, optimizing content to win one specific AI answer on one specific day is optimizing against noise. The more durable strategy is publishing enough accurate, citable pages on your actual areas of expertise that you show up across a wide sample of runs and phrasings, the way City of Hope's dominant appearance rate came from being a genuinely strong, unambiguous answer to its category rather than from gaming one prompt. How to rank in ChatGPT covers the extractability side of that work: answering questions directly, citing sources, and structuring pages so a model can quote them cleanly. Fishkin's "ranking position" critique is really an argument for retiring that framing altogether in favor of visibility rate as the metric you report and act on.
It is also worth separating two failure modes that look similar from the outside. Everything in this post is about run-to-run variation in which brands a model surfaces, a sampling and retrieval problem. That is different from ChatGPT stating something outright wrong about your company, which is a sourcing or training-data problem, and different again from the platform-wide citation volume swings covered in our look at the 2026 ChatGPT citation drop, which was driven by ad rollouts and default-model changes rather than per-prompt sampling noise. A brand can have a consistency problem, a hallucination problem, and a platform-volatility problem at the same time, and each needs a different fix.
None of this is a problem Lyra solves directly. She is not an AI-visibility or citation tracker, and she does not measure how often or consistently an engine cites the pages she writes; that still needs a separate layer, whether that is a GA4 custom channel, a manual prompt log like the one above, or a paid tracker. What she does affect is the other side of the equation this post argues for: breadth of accurate, citable surface area. Lyra discovers a topic, writes it in your blog's existing voice, fact-checks every claim against a live source, and opens a pull request so nothing publishes without your review. Her plans are priced by posts shipped per month, from a free tier at 3 posts up to 150 posts a month on the Scale plan (current as of this post's publish date; check the live pricing page for today's numbers), which maps directly to the conclusion above: more fact-checked, citable pages sampled across more runs beats chasing one stable score that the underlying math will not produce.
A single AI-visibility snapshot can't tell you where you rank, because the data shows there often isn't a stable rank to find. Publishing more fact-checked, citable ground truth is the strategy that survives the noise.
FAQ
Three overlapping causes. The model samples from a probability distribution rather than always picking the top token, so wording and order vary by design. Live retrieval means the underlying set of pages it can cite changes between two runs of the same prompt, sometimes hours apart. And conversation memory or personalization can skew results per user, so two people asking the identical question in two separate sessions are not really running a controlled experiment.
Not very, at the list level. A November-December 2025 study by SparkToro and Gumshoe.ai ran 2,961 prompts across ChatGPT, Claude, and Google AI Overviews and found the exact same list of recommended brands repeated in under 1% of runs, and the exact same order in under 0.1%. Individual brands still showed up reliably: City of Hope appeared in 97% of ChatGPT's cancer-hospital responses. It is the full list and its order that are unstable, not necessarily whether a strong brand shows up at all.
Treat one snapshot as noise, not a ranking. Rand Fishkin, who co-authored the SparkToro study, put it directly: 'Any tool that gives a ranking position in AI is full of baloney.' A single run captures one sample from a distribution that a temperature setting, a live retrieval pass, and a model update can all shift on their own. Repeated sampling and a visibility percentage across many runs is the defensible version of the same measurement.
The two aren't linked the way you'd expect. A Washington State University study comparing ChatGPT-3.5 (2024) to ChatGPT-5 mini (2025) on 719 repeated business-research prompts found accuracy rose from 76.5% to 80%, but consistency didn't track it: the newer model still gave a different answer to the same identical prompt 27% of the time. A more accurate model is not automatically a more repeatable one.
Built by the tool you're reading about
Lyra finds the topics worth ranking for, writes them in your repo's voice, fact-checks every claim, and opens a pull request scored and ready to merge. You review and hit merge. Want to see what she'd write for you? Start free with three posts, no card.
Keep reading

What a blog post actually costs across every model: raw API tokens, flat SaaS tiers, freelancers, agencies and in-house teams, with the payback math.

AI content ROI reporting your CFO will trust: a worked formula for cost per verified AI citation, tokens plus review time divided by citations, not posts.

How to join Google Search Console and GA4 by landing page URL, assign lead value with no revenue event yet, and stop last-click from hiding organic's real ROI.