3
13 Comments

Our own AI-visibility score started at 0/100. What 2,150 cited URLs taught us about getting recommended by ChatGPT

Hi IH — I'm Anil, solo and bootstrapped, building Promvia: it checks whether ChatGPT, Perplexity, Gemini and Google's AI answers actually name your product when buyers ask for a recommendation, and what those answers send your site. Launched August 2026, €0 MRR so far, building in public. Disclosure up front: the numbers below come from our own tracking.

The thing worth sharing is a finding, not a feature. We track a fixed set of buyer prompts every week across six answer engines and log every URL the answers cite — about 2,150 cited URLs over 14 days in one niche (software tools). Two things fell out:

  1. "Alternatives to X" prompts cite comparison-shaped pages (paths like /vs, /alternatives, /compare — vendor pages and blogs included, not just G2 or Capterra) in 34–43% of citations. "Best X" prompts: 2.8%.

  2. When we first counted by directory DOMAIN instead of page shape, the effect looked ~10x smaller. The domain list was a bad proxy. What gets cited is the page type, wherever it lives.

The practical version: a well-structured /alternatives or /vs page on your own domain competes for the same answer slot as a G2 grid. "Get listed on directories" and "build comparison pages" are sold as two tactics; at the citation level they look like one. And the cited pages had almost no organic traffic — winning the SERP and being chosen for an answer are different games.

Caveats, honestly: one niche, two weeks, URL-path labeling rather than rendered pages. Direction solid, magnitudes indicative. We pre-registered a follow-up (vendor-authored vs third-party comparison pages) and will publish it on Sep 7 whatever it says.

We eat our own cooking: Promvia's own domain scored 0 on day one, and we publish our own weekly run live at promvia.app/live so the score is a public number rather than a promise.

Question for people running product sites here: do your /vs or /alternatives pages pull AI citations (or AI referral traffic) out of proportion to their organic traffic? That publisher-side number is the one nobody seems to have.

posted toAvatar for product Promvia
Promvia
  1. 1
    Really interesting analysis! I especially like the distinction between being cited and being recommended. One thing I’d be curious to see is citation support whether the cited page actually backs up the AI’s claim. That could add another interesting layer to AI visibility.
    1. 1
      Citation support is the measurement I most want and least have, so let me be exact about why rather than promise it. We store the answer text and the URLs the engine cited. We do not store the content of those pages. So the instrument can tell you a page was cited alongside a claim; it cannot tell you the page supports the claim. Doing it honestly means fetching each cited page at citation time and checking that specific claim against it — a content pass, and a fairly expensive one, not something the URL log can be squeezed into answering. I'd rather name that gap than infer past it. What the same data can say is a weaker but real version of your point. For this week's run I took every pair of engines that answered the SAME question and compared the sources they cited. 487 pairs. In 146 of them (30%) the two engines had no cited domain in common at all. Excluding Reddit and YouTube, which nearly everything leans on, it's 35%. Mean overlap across all pairs is 0.099 — about one source in ten. So two engines answering one buyer's question are usually resting on almost entirely different evidence, and both still hand back a confident, similar-sounding recommendation. That doesn't tell you whether any given cited page backs its claim. It does mean "it was cited" is thin evidence that it did — which I think is the same worry you're pointing at, arrived at from the source side instead of the claim side. The rows behind that are public (CC BY) if you want to check the arithmetic: promvia.app/api/public/ai-source-index/rows?week=2026-W39
  2. 1
    Your line that the cited pages had almost no organic traffic matches what I found on my own site, and the size of it surprised me. Over 92 days, the page with the most ordinary Google search impressions (1,479) earned zero Copilot citations. A page with 78 impressions earned 63 citations, more than half the site's total. Bing and Google also cited near-opposite pages: my top Bing page had 11 Google AI impressions, Google's top page had zero Bing citations. Rank and citation were not just different games; they were close to inverted. One caveat on "the citations are weekly and reproducible". A 2026 study across four engines found only 34% to 42% of cited sources come back when the same question is repeated, even within a day. Your 92% single-appearance tail is probably mostly that churn, and the 21 repeat domains are the real signal. It fits the wider data too: across 25 million cited links, 84% pointed to earned media, and brand-owned sites came out at roughly 3% to 10%. On the publisher-side number nobody has: for two engines it exists. Bing Webmaster Tools' AI Performance report logs Copilot citations per page with the queries behind them, and Search Console's generative AI report logs impressions by page and country. Neither covers ChatGPT, but both are records rather than samples, and they are what let me check my own pages. I do not run /vs pages, so I cannot answer your question directly; my numbers, zeroes included, are here: https://www.dhawalshah.net/article/ai-visibility-tracking/
    1. 1
      That inversion is the sharpest version of this I've seen from the publisher side. 1,479 impressions and zero citations against 78 impressions and 63 isn't a weaker correlation — it's a different selection rule. On your caveat I owe you a correction and a number. The correction: my 92% single-appearance figure is not a churn measurement. It counts domains appearing in exactly one QUESTION within a single week's run — breadth across the panel, not survival over time. Those are two different axes and I had been letting them sit next to each other without saying which was which. So I went and measured the time axis properly, because your study figure deserved an answer rather than a shrug. Four consecutive Monday runs, same 40 questions, same engines. Taking every (question, engine) cell that answered cleanly in both weeks of a pair, and asking what share of the domains cited in week N are cited again in week N+1: W36→W37: 782/1,416 = 55% W37→W38: 580/1,262 = 46% W38→W39: 645/1,220 = 53% pooled: 2,007/3,898 = 51.5% By engine, pooled: Perplexity 62%, Claude 59%, AI Overviews 56%, ChatGPT 47%, Gemini 46%, Google AI Mode 40%. And the harsher cut — of the domains cited in W36, restricted to cells that answered in all four weeks, 311 of 1,259 (25%) were still cited in every subsequent week. So roughly half survive one week and a quarter survive three. Two limits before anyone quotes 51.5% back at me. We take ONE sample per engine per question per week, so this number cannot separate "retrieval is stable" from "we happened to draw the same thing"; your 34–42% is same-day repetition, which is a stricter question our instrument literally cannot ask. And a week gives the underlying index time to change, so this isn't a cleaner version of your number — it's a different measurement that happens to come out higher. On Bing Webmaster Tools' AI Performance report and Search Console's generative AI report: you're right, and I should stop calling the publisher-side number missing when it exists for two engines. The distinction I'd keep is that both answer it only for YOUR pages — which is exactly why your numbers are worth more than mine here, and also why neither can say what the engines cite for a question where you aren't cited at all. Those are complementary instruments, not competing ones. Reading your write-up now.
  3. 1
    The distinction between being cited and being recommended is huge; the latter depends on whether the page gives the model enough context to match a real buyer’s situation. I’d segment the 2,150 citations by prompt intent and compare the pages that earned repeated mentions, not just total citation count. That should reveal which claims and proof points are doing the actual work.
    1. 1
      Agreed on the distinction, and we do split it — cited, mentioned and not-cited are three states in the product rather than one. But your second suggestion is the more interesting one, and I went and ran it, because I had not looked at the panel that way. Segmenting by repeated mentions instead of total volume, this week's run across the 40 buyer questions: 790 distinct domains cited. 727 of them, 92%, appear in exactly one question. Only 21 appear in three or more. That repeat set is almost entirely platforms and publishers, not vendor pages: YouTube in 30 of the 40 questions, Reddit 28, TechRadar 10, Forbes 7, CNBC 6, PCMag 5. So your method surfaces something real, and it points somewhere I did not expect. The pages that earn repeated mentions are not the ones making claims about a product; they are the substrate the engines keep returning to. Vendor pages live almost entirely in that 92% tail, chosen once for one question and then not again. Which suggests the claims-and-proof-points work, where it happens, happens per question rather than accumulating into a durable position. Two things I should be straight about rather than dress up. Counting was already deduped: a domain counts once per question per engine, so four engines naming one page on one question is four, never four for one page. Volume was never raw volume. But that is exactly why your reframing earns its keep. Breadth across questions and breadth across engines are different axes, and I had been reading the first as if it implied the second. The part this instrument cannot answer is which claims and proof points do the work. We store cited URLs, not rendered page content, so I can tell you a page was chosen and for which question shape, not what about it was persuasive. That needs a content analysis on top, and I would rather name the gap than infer it from URLs. On cited versus recommended, the sharpest version I have is our own domain: on branded questions we are named 4 times out of 4; on the 15 where we compete on merit, zero. Same product, same pages. Being retrievable and being chosen really are two different measurements.
      1. 1
        This is the useful cut. 92% one-question domains versus a tiny repeat set of substrate publishers is a much cleaner story than raw citation volume. Two takeaways I am stealing: 1. Breadth across questions ≠ breadth across engines. Your dedupe already forced that apart. 2. Vendor pages living in the one-shot tail means “claims and proof” mostly buys a single retrieval event, not a durable slot. Your branded 4/4 versus merit 0/15 makes the same point from the other side: retrievable and chosen are different jobs. Respect for naming the instrument gap too. URL-level citation can show *that* a page was picked and for which question shape; it cannot show *which* claim or proof point did the work. That content pass on top is the next measurement, not something the URL log can pretend to answer. If you ever publish the question-shape breakdown for those 21 repeat domains, that would be gold for people trying to stop optimizing the 92% tail.
        1. 1
          You asked for the question-shape breakdown of the repeat domains, so here it is, from this week's run (2026-W39). The panel is 40 questions — 8 "alternatives", 16 "best", 8 "compare", 8 "how to buy" — across six answer engines, one sample each, a domain counted once per question per engine. 949 distinct domains cited. 869 of them (92%) in exactly one question. 28 in three or more. The 28, as questions hit / questions of that shape: reddit.com 34/40 — alternatives 6/8, best 15/16, compare 7/8, how-to-buy 6/8 youtube.com 33/40 — alternatives 8/8, best 11/16, compare 8/8, how-to-buy 6/8 techradar.com 12/40 — 4/8, 6/16, 1/8, 1/8 forbes.com 10/40 — 2/8, 3/16, 2/8, 3/8 tomsguide.com 8/40 — 3/8, 2/16, 0/8, 3/8 nerdwallet.com 7/40 — 1/8, 4/16, 1/8, 1/8 medium.com 6/40 — 2/8, 2/16, 1/8, 1/8 zapier.com 6/40 — 2/8, 1/16, 2/8, 1/8 cnet.com 5/40, rtings.com 5/40, dev.to 5/40 theguardian.com 4/40, northflank.com 4/40, consumerreports.org 4/40, facebook.com 4/40, cnbc.com 4/40, quora.com 4/40 salesforce.com, emailvendorselection.com, emailtooltester.com, investopedia.com, healthline.com, macys.com, monday.com, larksuite.com, wallethub.com, wise.com, finder.com — 3/40 each Three things I didn't expect: 1. YouTube is shape-complete where Reddit isn't. YouTube takes 8/8 alternatives and 8/8 compare but only 11/16 best; Reddit is the mirror image, 15/16 best but 6/8 alternatives. They are not interchangeable substrates — they own different question shapes, and anyone treating "get on Reddit and YouTube" as one tactic is buying two different things. 2. Consumer Reports appears in 4 questions and all four are how-to-buy, zero across the other 24. A shape specialist is a completely different asset from a broad publisher, and counting by volume hides that entirely. That's your point about prompt intent, and it's sharper than I gave it credit for. 3. A correction to what I told you last week. I said the repeat set was "almost entirely platforms and publishers, not vendor pages". Looking at it properly: Zapier 6, Northflank 4, Salesforce 3, Monday 3, Lark 3, Wise 3 are vendor-owned domains sitting in the repeat set, plus a layer of affiliate comparison publishers (NerdWallet 7, WalletHub 3, Finder 3, EmailToolTester 3). Platforms still take the top two slots by a distance, but "vendor pages live only in the 92% tail" was too strong and I shouldn't have said it that cleanly. One more thing worth flagging about stability: the repeat set was 21 domains when I quoted it to you and is 28 now, from 790 distinct domains then and 949 now. Same panel, same engines. Separately I measured the week-over-week survival of individual citations across four runs and it's ~52%, with ~25% surviving all three transitions (numbers in my reply to Dhawal above). So the tail/repeat split is a real structure, but the membership of it is much looser than a single week's table makes it look. All of it is published under CC BY at promvia.app/api/public/ai-source-index/rows?week=2026-W39 — the per-question, per-engine rows, not just the tables — so you can rebuild any of the above and tell me where I got it wrong.
  4. 1
    The 2,150 cited URLs make the finding much more interesting than another AI-visibility claim. Have you seen /vs or /alternatives pages produce meaningful qualified traffic from AI answers yet, or is citation volume still the stronger signal?
    1. 1
      Honest answer: citation volume is the stronger signal for us so far, and I would not claim qualified traffic yet. Two reasons, both measurement rather than marketing. 1. Most assistants open links with no referrer, so the visit lands in analytics as Direct. We keep two buckets apart: confirmed (a referrer, or an assistant user-agent like ChatGPT-User fetching the page on a user's behalf) and inferred (Direct sessions landing cold on a deep /vs or /alternatives URL, which typed-in Direct almost never does). The inferred bucket runs roughly 2 to 3x the confirmed one. That gap is the honest size of the attribution hole. 2. On our own domain the numbers are small enough to say out loud: a handful of AI-referred visits and no attributed revenue. The comparison pages that got cited in the dataset were third-party ones, and we do not have their analytics, which is exactly the publisher-side number I was asking for. What I can say: the citations are weekly and reproducible, the traffic is not yet. We started publishing the source side in the open today (promvia.app/ai-source-index, 40 fixed buyer questions across seven engines every Monday), and the first week already showed the engines leaning on YouTube and Reddit far more than on any directory for consumer questions. If your /vs pages show AI referrals out of proportion to organic, I would genuinely like to see the numbers.
      1. 1
        That attribution gap is the interesting part, especially the confirmed vs. inferred split. I’d be interested in comparing notes on what the /vs pages are actually driving. What’s the best email to reach you on?
        1. 1
          Happy to compare notes: support@promvia.app reaches me directly. If you send the /vs URLs you are looking at, I can run them through the same weekly panel and share what the engines cite for their questions, no strings.
          1. 1

            Thanks! I’ve just sent it over.

            Looking forward to hearing your thoughts whenever you have a chance.