I build Orbator. It measures whether AI assistants recommend your product, and traces which sources they read before answering.
A founder emailed last week asking which two or three placements would move the needle for her consumer money app. I assumed the answer was directories and review sites, because that is the advice everyone gives. My own data should have warned me. Across the software categories I track, the most cited domains are reddit and youtube. G2 is the biggest review site, close to four times Capterra, but it sits eighth overall and reddit gets cited almost nine times as often.
Looking into the detail of her actual report changed that. For the question that matches her differentiator, an app that works without connecting your bank, the engines had cited three pages. After checking all three, every one was a competitor's own website. Not a review site, not a listicle. Just products like hers, being used as the source. One of them runs almost exactly her pitch, a 0 to 100 money score with no bank linking, and it gets cited while she gets nothing.
So the assumption was wrong. The gatekeeper was not an editor at a review site. It was her own homepage. Those competitor pages get cited because they say it plainly, right there on the page: "never link a bank", "no bank linking required". Her site says it once, as a feature bullet.
Here is what I had missed. The directory playbook is real, it is just a B2B playbook. When someone is buying a tool for work there is a G2 page to lean on. For consumer apps there is no G2, so the engines fall back on whoever explains the category best, and that is usually a competitor. I had one dataset, I had not even read it closely, and I assumed it applied everywhere.
The test I use now: look at who is already getting named. All household names, skip that question. Apps you have never heard of in there, the slot is open.
Has anyone here checked which sources the AI answers in your category are actually built on?
The useful reframe is that the citation went to whoever stated the category in the same words a person would type. A feature bullet does not do that; a page whose heading is literally "no bank linking required" does. The tactic I would hand her is to list the ten questions she wants to win, then confirm each one has a page answering that exact question in that exact phrasing, because right now a competitor is getting paid for her positioning.
The B2C vs B2B directory split is the one I got wrong first. My assumption was the same — get listed on the right sites, citations follow. That's a B2B playbook and it doesn't port to consumer.
The "plain sentence on the homepage" finding matches what I saw when I diagnosed my own sites. I ran an automated content pipeline for two weeks —18 GEO-optimized articles, technically correct, proper markup, sensible internal links. Zero AI citations. When I finally looked at who was getting cited for my category query, it was a competitor's homepage — one clear sentence stating the specific capability as the lede, not buried in a feature list.
The query class split is worth isolating too. I've been tracking the same URLs across a narrow-intent query ("tool that does X specific thing") vs. a broad-category query ("best tools for Y"). Narrow queries cite whoever states the capability most directly — that's a homepage rewrite problem. Broad ones are messier: Reddit threads, YouTube, sometimes nobody recognizable. Different gatekeepers, different work.
Your test — look at who's already getting named — is the right start. If it's apps you've never heard of, the slot is open. If it's all household names, the citation behavior is already locked to brand signals and homepage rewrites probably won't move it.
We label every prompt by question type before we ask it, so I could run your narrow-versus-broad split this morning. Across 28 days, broad category questions pull 5.2% of their citations from reddit or youtube and narrow problem-solving ones pull 6.9%. Review sites invert it, 1.3% broad against 0.5% narrow. The mixes do differ, just not the way I expected from your description, and it is an average across 295 categories so it cannot see which page wins inside any one of them.
On the pipeline result, the closest thing I have from our side is that 46.4% of the pages the engines cite are owned by a vendor competing in that same race. Consistent with the citation not coming from volume around the category.
Your household-names line is sharper than the test I wrote. Where do you actually draw it, recognition alone, or is there a point where you stop bothering with the rewrite?
The B2B/B2C split might be the wrong cut of that data. The query you ran was her differentiator - works without linking a bank - which is product-shaped, so product pages get cited. Your own aggregate says reddit gets cited nine times more than G2 across the categories you track. That is what a category query pulls: best budgeting app.
Same product, two query classes, two different source mixes. She needs the plain sentence on her homepage for the narrow one and presence in threads for the broad one.
Worth running her category query before she rewrites anything. If those three cited pages flip to reddit and youtube, the variable is query shape.
That is a better cut than mine and I should have separated the two, so I went and measured it. Our prompts carry an intent label already, so I split the last 28 days of runs into broad category questions and narrow problem-solving ones and counted only citations that appeared inline in an answer. Broad questions: 5.2% of citations from reddit or youtube, 1.3% from review sites. Narrow questions: 6.9% reddit or youtube, 0.5% review sites. So no flip toward reddit on the broad ones, if anything slightly more community content on the narrow ones. The review-site half of your intuition does hold, broad questions pull about 2.6 times the review-site share.
This is an average across about 295 categories, so I cannot see which specific page wins inside any one of them. On your feedback, I added a consumer budgeting category with both question shapes in the same prompt set, including the exact narrow one from the post. And I shipped the split itself,share by question type instead of blended. First category we ran, one product holds 95% on the broad question and 40% on the narrow one. First budgeting runs land tomorrow on the next daily batch. I'll keep you posted on the results.
The homepage-as-gatekeeper finding matches what we've seen — engines cite pages that say the claim plainly enough to quote. Her competitor's 'no bank linking required' line is exactly the kind of copy that gets lifted verbatim. We hit the same wall auditing AI citations for our analytics product (https://amami.dev): the pages that get cited are written as answers, not feature lists. Worth rewriting your homepage around the one sentence you want quoted.
Written as answers rather than feature lists matches what I see. What I would add is that it has to be close to the words a buyer would actually type. Her competitor's line reads almost exactly like the query, which is probably why it gets lifted whole. What did your audit turn up on your own domain, did the cited pages skew to your docs or to third parties?
This is a good example of how AI citation behavior doesn't map cleanly onto traditional SEO rankings. Was it purely about content depth on the competitors' homepages, or did you notice structural things too, like how the info was organized or phrased, that seemed to make AI more likely to pull from them?
Phrasing more than depth. The competitor pages that got cited were not longer or more thorough than hers, and a couple were thinner. They stated the constraint in one plain sentence near the top. She had the same fact, as a feature bullet halfway down. I cannot see inside the ranking, so that is a pattern across the pages that got cited rather than a mechanism I can prove.
The shift from the original assumption to what the actual report showed is interesting. Curious how often the cited sources differ from what you would expect going in.
Often enough that I stopped guessing first. I have a restaurant in Miami so I ran, best pancakes in Miami, 3 of the 6 restaurants the engines recommend sit on a single listicle published in 2020, which turned up in 4 separate runs. I would have bet on recent reviews or maps.
That’s a useful example. The fact that one old listicle keeps appearing across separate runs is probably more important than the individual result. If you’re open to continuing the conversation, what’s the best email to reach you at?
This is measurement system lag in action. Founders built their intuition when review sites = distribution. That measurement system told them "get on the review sites" because that's what worked. But the distribution channel shifted (AI is now the amplifier) and their measurement system didn't update. Classic founder blindspot: you can't see the problem your measurement system can't measure. Until you can't compete.
The lag is nastier than just missing data, because the old number still goes up. Listings get added, the directory dashboard looks healthy, and the answer names someone else the whole time. Nothing feels broken, which is why it runs so long before anyone checks.