Blog

Why does ChatGPT give a different answer every time you ask? The mechanics, and what it means for your score

Ask the same buyer question twice and you can get two different shortlists — same brand, same engine, same day. That's not a bug or a bad prompt. Here's the actual mechanism, and the honest way to measure something that keeps moving.

The PingMyBrand team5 min read
On this page
  1. 01It's not a glitch — it's the mechanism
  2. 02Why this matters more for a recommendation than a fact
  3. 03What one manual check actually tells you (and doesn't)
  4. 04The honest way to measure something that keeps moving
  5. 05What to actually do with this

Ask ChatGPT "what's the best help desk software for a 10-person startup" twice in a row and you can get two different lists. Same words, same model, five minutes apart — and the second answer drops a brand the first one led with. If you're checking your own AI visibility by hand, that's the moment it stops feeling like a measurement and starts feeling like a coin flip. It isn't random, and it isn't a broken prompt — it's a direct consequence of how these models generate text at all, and understanding it changes how you should measure visibility in the first place.

It's not a glitch — it's the mechanism

A large language model doesn't look up an answer and print it. At each step, it computes a probability for every possible next word given everything written so far, then samples one — draws from that probability distribution rather than always taking the single highest-scoring option. That sampling step is deliberate, not an accident of engineering: a model that always picked the top-probability word every time would produce flatter, more repetitive, more robotic text. The controlled randomness is part of what makes the output read naturally.

That has a direct consequence for a brand-recommendation question specifically. If two or three brands sit close together in the model's learned associations for "help desk software for a startup" — none of them a runaway favorite — sampling can pick a different one to lead with on different runs, even though nothing about the world changed between the two asks. The model isn't reconsidering your product. It's drawing from the same underlying distribution and landing on a different sample.

This is worth separating from a second, related cause: models get updated. OpenAI, Anthropic, Google, and xAI ship new model versions and retrieval changes on their own schedule, with no changelog for how your category is now described. That's a slower-moving, structural kind of drift. The run-to-run sampling variance above happens even with zero model changes at all — it's baked into how a single, unchanged model generates any answer, every time.

Why this matters more for a recommendation than a fact

Ask a model "what's the capital of France" and sampling variance is invisible — there's one overwhelmingly likely next token, so it says "Paris" every time. Ask it to recommend a brand in a competitive category, and you're asking a genuinely close-call question: several plausible answers, no single one so dominant that sampling can't occasionally land somewhere else. That's precisely the kind of question a buyer asks you about — and precisely the kind of question where run-to-run variance actually shows up in what you'd see if you checked by hand.

Put plainly: the more competitive your category, the more a single check is just one sample of a moving target — not because the model is confused, but because your category is a genuinely close call, and close calls are exactly where sampling has room to land differently.

What one manual check actually tells you (and doesn't)

We've written about checking ChatGPT by hand: open it, ask a real buyer question, read whether you're named. That's still worth doing today — it's a real, honest data point. What it isn't is a verdict. One run is one draw from the distribution above, not the distribution itself. If it named you, that's genuinely good news, but it doesn't mean every buyer asking the same thing this week saw the same answer. If it didn't name you, the same caveat runs the other way — don't assume you're locked out forever off one unlucky draw.

The honest way to measure something that keeps moving

There are two legitimate ways to turn a moving target into a real signal, and it's worth being precise about which one we actually use and which one we deliberately don't.

What doesn't work, on its own: asking the identical question over and over in one sitting. You could, in principle, ask the same prompt fifty times and report the percentage that named you — that's a real statistical technique, but it measures the sampling noise for that one phrasing of that one question, and buyers don't ask you the same nine words fifty times. It would buy false precision on a question that isn't actually the one you care about.

What we do instead: 25 different real buyer questions, asked once each, across four engines, every scan. Our code runs each of those 25 questions exactly once per engine per scan — no repeat-sampling of a single prompt pretending to be more precise than it is. The noise gets smoothed a different, more honest way: by averaging across many different real questions rather than many repeats of one question, and then by re-scanning weekly and reading the trend line, not any single scan, as the signal. A brand named in 18 of 25 questions this week and 19 of 25 next week is a stable, real pattern. A brand named in one manual check and not the next isn't evidence of anything — it's one sample doing exactly what sampling does.

This is also why we never say a fix "caused" a score change — a score built on non-deterministic models, re-run on a schedule, can move for reasons that have nothing to do with your last edit. The honest response to that isn't to hide the variance behind a falsely precise single number. It's to show you the real trend and let repetition, not a single lucky or unlucky draw, do the convincing.

What to actually do with this

Don't panic over one bad manual check, and don't relax over one good one — both are single draws from a distribution that's genuinely still moving. If you want to know where you actually stand, the only honest approach is the same one that applies to any noisy signal: ask many real questions, not one repeated question, and watch the pattern across more than one reading.

Run the free scan and see it done that way: 25 real buyer questions, put once each to ChatGPT, Claude, Gemini, and Perplexity, with the literal answer shown for every one — not a single number pretending the noise doesn't exist. No signup, about a minute, and if you come back next week and rescan, you'll start seeing the trend line instead of just one more sample.

Share the data

Found this useful? Put it in front of your network

One click copies the link or opens a prefilled post — no rewriting required.

https://pingmybrand.com/blog/why-ai-answers-change-every-time-you-ask?utm_source=share

Does AI recommend you, or a competitor?

Enter your domain. We ask 25 real buyer questions across ChatGPT, Claude, Gemini & Perplexity and show you, per question, whether you're named — the exact sentence, not a green dot. Free, no signup, about a minute.

Free · no signup · 4 engines · ~60 seconds

Keep reading