A reader emailed me, annoyed. Our audit told him his brand barely showed up in AI answers. He opened ChatGPT himself and there he was, first line. "Your tool is wrong." Then he added the part that actually mattered: he ran our audit twice and got two different scores.
He was right on both counts, and the second one sent me down a rabbit hole.
The problem: LLMs are non-deterministic. Ask the same question twice and you can get different brands, different ordering, different recommendations. It's not a bug - they pick words by sampling, so output varies by design, and it holds even at temperature 0. Which means every AI-visibility tool that asks each model once (ours included) was handing people one sample from a noisy distribution and calling it a score. Run it again, it moves.
You can't plan content work against a number that swings 25 points on noise. So I stopped trusting the single shot.
What we built (Deep Audit):
Fan one query into ~10 real buyer-intent variations instead of one phrasing
Run all of them across ChatGPT, Claude, Gemini, Perplexity
Report a mention rate with a 95% confidence band - measure many times, report the distribution
Only count a mention if the brand name is literally in the answer (the models were tagging "partial" mentions for brands they never named - self-grading is generous)
The honest limitation: it measures what the models know from training, not your live site. I tested a brand a friend has promoted hard for a year - our tool showed near-zero, and Google's live AI Mode showed near-zero too for that query. Two different methods, same answer. Live grounding is the next problem; I'd rather ship a number you can reproduce than one that looks live and isn't.
Full writeup with the research: https://geolikeapro.com/blog/why-one-llm-audit-isnt-enough
Free brand check (no signup) on the homepage if you want to see where you stand. Happy to answer anything about the build or the measurement approach.