A few months ago I noticed something that bugged me: people don't "Google" the way they used to. They ask ChatGPT, Perplexity, Gemini. And when they ask "what's the best tool for X?", the AI recommends a handful of products and completely ignores the rest.
I kept asking myself: why those and not the others? The answer sent me down a rabbit hole, and it turned into the product I'm building now (CitableHub a directory built specifically for AI engines to find and cite products). But forget the product for a second the real lesson is about how AI decides who to recommend. Here's what I've learned so far.
A profile/page that just "exists" gets you nothing
My first assumption was wrong. I thought being listed somewhere was enough. It isn't. AI engines don't cite pages that just say "we're the best innovative leading solution." That's filler, and filler is invisible to a model looking for a quotable answer.
Answer-first copy is everything
The pages that actually get cited answer "what is this?" in one clean sentence, in a format the model can lift word-for-word. Something like: "[Name] is [what it is] where [who it's for] achieves [result] with [how]." Boring? Maybe. But that's the exact shape a model quotes. The clearer the sentence, the more likely you are the one it picks.
You have to answer the real questions people ask
This was the biggest unlock. Instead of writing marketing copy, I started mapping the actual questions someone types into an AI right before they'd need the product and answering each one in 2-3 concrete sentences. Every well-answered question is another doorway the AI can walk through to reach you. More questions covered = more conversations you show up in.
Proof is what turns "mentioned" into "recommended"
Models favor sources they can verify. Saying "we're great" does nothing; showing third-party listings, testimonials, real numbers, links that's what earns trust. Proof is the difference between an AI mentioning you exist and an AI recommending you.
Freshness is a ranking signal, not a vanity metric
A page that never changes reads as abandoned. The ones that keep shipping small updates a feature, a milestone, a case study keep signaling they're alive, and they climb over the ones that went quiet. Consistency beats intensity here.
The pattern I keep seeing: be findable (answer the real questions) → be trustworthy (show proof) → stay fresh (keep shipping). It's basically SEO, but for a reader that's a language model instead of a human skimming a results page. People are calling it GEO (Generative Engine Optimization) now, and honestly I think it's going to matter as much as SEO did.
Still early, still figuring it out, building in public.
My question for you all: are you doing anything to optimize your product for AI citation yet? Have you noticed ChatGPT/Perplexity sending you traffic or recommending a competitor over you? I'd love to hear what's working (or not) for other founders here.
The citation angle is interesting, but have you seen any measurable link between these page changes and actual AI-referred traffic or signups? Otherwise it’s hard to separate correlation from what simply sounds citation-friendly.
This is one of the most useful comments I've gotten, thank you James and you're right that I collapsed two different problems into one.
Problem A: "once an engine decides to look at you, does answer-first copy make it cite you?" yes, and your 5/5 GPT result matches what I've seen.
Problem B: "do you even make it into the consideration set on an unbranded best/which tool query?" that's the one with actual buyers on it, and structuring your own pages does almost nothing for it. Different game entirely.
And your observation is exactly why I'm building this: every third-party citation you earned came from directories and aggregators, not articles. That's the lever for Problem B. The honest caveat you raised is the real one though being listed proves you exist, not that you're good, and "best tool" queries are precisely where engines seem to weigh that. So a directory only helps if it carries structured proof (real usage, verifiable signals), not just a logo and a blurb. That's the part I'm still figuring out.
@aryan_sinh on measurable lift, here's the part I think will be uncomfortable for a lot of people but is worth seeing: you don't have to take my word for it. We built a live crawler dashboard on the homepage that shows real AI engines reading profiles in real time every visit is server-verified by user-agent, not estimated. Right now it's showing 500+ verified AI reads, with OpenAI, Claude, DeepSeek, Gemini, Meta AI, Bing and Amazon's crawlers all hitting pages. You can literally watch them arrive. If you want proof that AI engines actually crawl and pull from this kind of structured directory, just open citablehub and watch it happen live.
I still won't pretend the referral-to-signup attribution is clean yet that's the next thing I'm instrumenting properly. But "do the engines actually come and read?" is no longer a theory on my end; it's on screen.
That makes the evidence gap much clearer.
What decision will the attribution work actually help you make about CitableHub?
A data point that supports half of this and complicates the other half.
Our own site is already shaped the way you describe. One sentence definition in llms.txt, FAQ schema on every page, headings phrased as the questions people actually ask. We had it run through six engines this week. GPT named us in five answers out of five while citing almost nothing but our own pages. Perplexity refused us outright on the best tools question, because our site was not in its sources.
So answer-first copy did work, but only once an engine had already decided to look at us. It did not put us in the consideration set on the unbranded question, which is the one with buyers on it. Those are two different problems and the post reads as though they are one.
Encouraging bit for your directory: every third party citation we did get came from directory and aggregator pages rather than articles, so engines genuinely do reach for them. The caveat is that being listed is evidence you exist, not evidence you are good, and the list question is exactly where that difference shows up.
Fair challenge, and honestly the right one to ask. I won't pretend the referral-to-signup attribution is clean yet that's the next thing I'm instrumenting properly.
But on "do AI engines actually crawl this stuff" you don't have to take my word for it. We put a live crawler dashboard on the homepage that shows real AI engines reading profiles in real time, every visit server-verified by user-agent, not estimated. Right now it's 500+ verified AI reads with OpenAI, Claude, DeepSeek, Gemini, Meta AI, Bing and Amazon's crawlers all hitting pages. Open citablehub.com and you can literally watch them arrive. The measurable-lift-to-signups part is still early but "the engines actually come and read" is no longer a theory on my end, it's on screen.