1
9 Comments

5 AIs recommended the product. Then we looked at what they cited. It was a mess.

Update (2026-09-07): A sharp-eyed reader pushed back on two claims here, and they were right.

The UTM parameters in the GPT citations are not evidence of someone planting content. We opened the hunted.space page the model cited: every outbound link there carries Product Hunt API's own tracking (utm_source=Application: hunted.space), and hunted.space's own links carry ?ref=hs. The model simply reproduced hrefs that were already on the page. Retracting the "someone is engineering citations" framing — the accurate version is: engines cite what already exists on your public pages, and the UTM shows which platforms are measuring their own outbound referrals. Less dramatic, more correct.
The Qwen 0/5 line is one sample. "It knew us, then forgot" and "it was never reliable" are indistinguishable at n=1. We treat it as such and will re-run before claiming anything more.

In our last post we asked several AIs about a real product — and most of them recommended it, with one memorable exception. Several of you came back with the same follow-up question:

"OK, but what did they actually cite to back that up?"

Fair question. So we ran the product through our full tracking — 6 engines, same 5 questions, same day — and this time we ignored the verdicts. We only looked at the citations.

What we found: the product got recommended a lot. And every engine justified it with completely different material.

GPT — 5/5, and 4 of the citations were other people's blogs

GPT answered all 5 questions and cited third-party articles to support most of them — including hunted.space and debutly.app. Not the product's own site. Other sites, talking about it.

Here's the detail that made me sit up: those citations had UTM parameters on them. Somebody is running attribution on AI citations. This isn't a passive game where AIs casually mention your link — people are actively planting content and tracking which engine surfaces it.

Perplexity — 5/5, official site only

Perplexity also answered everything, but it cited almost nothing but the product's own website. zarek.tech, over and over.

Perplexity is a first-party-evidence engine. If your official site is weak — thin, no clear product description — that's all it has to work with. There's no second source to rescue it.

Gemini — 5/5, but the grounding links were hidden

Gemini recommended the product too. Its grounding links existed... and the domain references were masked. You could see "there's a source" without being able to tell whose.

So even a "yes" from Gemini doesn't tell you whether you won on the official site, on a directory, or on a third-party article. The visibility is there; the audit trail is not.

Kimi — 1/5, half-knowing

Kimi answered the questions, but only one mention, and its context was essentially "there's no information about this product". Not wrong, not right — just not there.

GLM — 0/5. Qwen — 0/5, after 16 historical mentions

GLM never mentioned the product at all. Clean zero.

Qwen is the interesting one. Across our full tracking history, it has mentioned this product 16 times — it used to know it. On this run: zero. It knew, then forgot.

AI has no memory. It has retrieval. If the retrieval surface changes — or if the content stops being findable — you don't decay slowly. You drop to zero overnight.

What this means for you

  1. Your citation profile is your GEO blind spot map. The same product looks completely different to each engine: GPT found third-party articles (with UTMs), Perplexity relied on the official site, Gemini cited with hidden domains. None of them saw the same thing. Whatever you have — or don't have — per engine, that's exactly what gets said about you.

  2. The UTM is a battlefield signal. Attribution on AI citations means people are actively engineering this. Not hypothetically. Links are planted, content is written, engines are tested. If you're not doing any of it, you're not competing — you're just being cited (or not) by accident.

  3. No third-party content = no persuasion material. Perplexity trusts first-party. GPT trusts articles it found. If nobody writes about you, GPT has nothing to cite but its own guesses — or worse, someone else's version of your story. Directories get you "exists". Independent articles get you "persuades". The engines split on that line themselves.

So the real question isn't "does AI mention my brand?" — it's "when it does, what does it point to, and who planted that?"

Curious whether the engines even mention you — and what they'd cite if they did? Run the free 30-second scan: https://brandscope.dev

Or drop your domain in the comments and I'll read out what each engine cites for it. Genuinely curious to see how varied the profiles are.

on September 6, 2026
  1. 1

    Thanks for adding the correction. Following the actual link back to the directory makes the UTM finding much easier to understand. I appreciate being able to see what you first thought and what changed after checking.

    1. 1

      Thanks — that's the reason we added it. Keeping the original framing next to the correction felt more honest than silently editing: you get to see what we concluded first, what checking showed, and why the takeaway changed. Turns out the UTM was Product Hunt's API tagging hunted.space's own links — nobody planted anything, engines just cite what already sits on public pages. Easier to trust a finding when you can walk the same path.

      Appreciate you reading that closely. If you want the same look at your own domain — what each engine cites, and whether it's first-party or third-party — drop it here and I'll read it out.

  2. 1

    The engine split is the useful finding: Perplexity leaned on first-party pages while GPT reached for third-party writeups. I’d turn that into two separate content checks rather than one overall “AI visibility” score, and track citation presence plus answer accuracy across repeated runs. A single mention can look impressive while still pointing to the wrong claim.

    1. 1

      Agree — and the split is the product. We don't roll the engines into a single "AI visibility" score for exactly this reason: GPT citing a third-party writeup and Perplexity citing the official site are different assets doing different work, and one number hides which one you're missing.

      Our tracking already buckets per mention — did it mention you, did it recommend, and what did it cite (tagged own-site vs third-party) — and we write our reports engine-by-engine rather than as an average. What you're proposing is the next layer: treat citation presence and answer accuracy as two separate checks. Presence says you're findable. Accuracy says the story being told about you is the right one. Both matter — a mention that points to a wrong claim is arguably worse than silence. That's not theory for us: in our first post in this series, an engine recommended the product but got the story wrong.

      Accuracy across repeated runs is the honest hard part. We have the tracking built, but the accuracy read is still done by hand — and this week a commenter caught us treating a single run as a trend (our Qwen 0/5 line). Fair hit. Our bar now is multiple consistent runs before we report any change as real, and the accuracy side still needs to catch up to that.

      If you drop a domain here I'll read out what each engine cites for it — curious whether the same split shows up for a different product.

  3. 1

    The UTM detail is probably not what it looks like. A directory like hunted.space puts UTMs on its own outbound links so it can measure its referrals, and the model reproduced the href it found on the page. That is the directory attributing its traffic, not someone attributing AI citations. Worth checking before concluding people are planting content: open the listing and see whether the link carries that same tag. If it does, nobody planted anything.

    On Qwen, 16 mentions and then zero across five questions is one sample. These answers are not deterministic, and a model that surfaces you a third of the time will return zero on five questions often enough. Before calling it forgetting, run the same five three times and count. Knew-then-forgot and never-reliable look identical at n=1.

    1. 1

      Thanks — you're right on both counts, and we went and checked.

      On the UTM: we opened the hunted.space page GPT was citing. Every outbound link there carries ?utm_campaign=producthunt-api&utm_medium=api-v2&utm_source=Application: hunted.space (ID: 206049) — that's Product Hunt's API tacking attribution onto hunted.space's own links, and hunted.space's own nav links carry ?ref=hs too. So nobody planted anything; the model just reproduced the hrefs that were already on the page. We're retracting the "people are actively engineering citations" framing. The corrected take is still useful: engines cite whatever links exist on your public pages, and the UTM tells you which platforms are already measuring their outbound referral traffic — a signal worth more than a conspiracy.

      On Qwen: agreed, that's n=1. Same five questions, one run, one zero — "knew then forgot" and "never reliable" look identical from a single sample. Our production data path requires multiple consistent runs before we'd report a change as real; this one doesn't meet that bar yet. We'll re-run before saying anything stronger.

      Appreciate the sanity check — this is exactly the kind of thing our own tool is meant to surface, with the variance caveat attached. Happy to read out what each engine cites for your domain if you drop it here.

      1. 1

        Yes please, utilityseo.com. And so it is a test rather than a favour, here is what I expect before you run it.

        Perplexity should find us, because on your own read it leans on the official site and ours is thorough. GPT should find close to nothing, because it wants third party material and we have ten referring domains and all ten are scraper spam rather than anything editorial.

        If that holds, the same site gets recommended by one engine and is invisible to another purely on backlinks, which is a sharper version of your point than the citation types are on their own.

        If GPT does cite us I would rather know, because it would mean I have been wrong about what our profile actually looks like.

        1. 1

          Ran it — same 6 engines, same 5 questions as the post, on utilityseo.com. Here's your readout.

          Perplexity: exactly as you predicted. 4/5. Official site carries it — your audit tool page, competitor research tool, your own comparison posts — with a handful of third-party cites (onelittleweb, auditzap, leapd, seomator, zapier). The one miss is telling: on "top SEO tools 2026" it said your site isn't in its sources, so it couldn't verify you. No listicle, no mention.

          GPT: you were wrong — in the way you said you'd want to know. 5/5, recommended every time. But look at what it cited: almost entirely utilityseo.com itself. Not one third-party article across all five answers. So your backlink read was right — those ten domains didn't buy you anything — but "GPT can't find me without editorial" is wrong. It found you on your own site, and had zero independent sources to back the recommendation.

          That's the sharper version of your point. The split isn't "one engine sees you, one doesn't". It's: your official site buys you visibility everywhere, and the absence of third-party material shows up as unbacked recommendations — GPT praising you using only your own words.

          Two more findings you didn't predict: the list question is your real blind spot (GPT put you #5 citing your site; Perplexity refused), and the Chinese engines barely see you at all — Kimi mentioned you 4/5 with no sources, Qwen 1/5, GLM 0/5.

          Full citation log if you want it. Happy to re-run after you ship the fixes.

          1. 1

            Wrong in exactly the direction I said I would want to be, so thank you for actually running it.

            The correction I would make to your synthesis is about which question matters. Being recommended 5 out of 5 when someone names us is a branded query and they already knew us. The list question is the acquisition one, and that is the one we lose: fifth on GPT citing only ourselves, and refused outright by Perplexity. So the score that looks best is the one worth least.

            The other thing in your log I would not have guessed. The third party cites you found, onelittleweb, auditzap, leapd, seomator, zapier, are all directory and aggregator pages. So we do have third party surface, it is just entirely listing shaped. That reframes the job from get more third party to get third party the model treats as evidence rather than as an entry.

            Yes to the citation log, and yes to a re-run.