1
0 Comments

We scored 50+ prompts across five AI engines: what the spread tells you about brand visibility

We build tools at Inithouse. One of them scores how five AI engines (ChatGPT, Claude, Perplexity, Gemini, and Google AI Overviews) recommend a brand on a 0-100 scale.

After running reports for dozens of brands, here's what the data actually looks like.

The setup

Each report fires 50+ prompts at all five engines. The prompts cover different buying intents: "best X for Y," "alternatives to Z," "which tool should I use for..." The kinds of questions real people ask AI instead of Googling now.

Each response gets scored: is the brand mentioned? Recommended? Positioned as a top pick? Cited with a link? The engine-level scores roll up into a composite 0 to 100.

The number nobody expects

The average composite score across all brands we've measured sits around 31 out of 100.

That means most brands, including ones with strong Google rankings, established domains, solid reviews, barely register across AI engines. The median is even lower. A few outliers pull the average up.

If your brand has never run an AI visibility check, there's a good chance you're somewhere in that 20-to-35 range. Even with a strong web presence. We've seen brands with 50k+ monthly organic visits from Google score single digits on Claude and Gemini. Google rankings and AI recommendations run on completely different rails.

The spread is the story

Here's what caught us off guard early on: the same brand can score 70 on one engine and 12 on another.

We expected some variance. We didn't expect this much. A SaaS tool with solid documentation and active community mentions might get recommended consistently by Perplexity (which tends to favor well-cited, structured content) while Claude barely mentions it. Or Gemini recommends it in one prompt framing and completely ignores it in a near-identical variation.

The engines don't share a recommendation logic. They have different training data cutoffs, different retrieval approaches, and different citation behaviors. Treating "AI visibility" as a single number hides the most actionable signal: where specifically you're visible and where you're invisible.

The citation gap we didn't anticipate

This one changed how we think about the whole space.

Perplexity cites its sources in roughly 97% of responses. If it mentions your brand, it almost always links to you. ChatGPT sits at about 16%.

That gap means "getting recommended by ChatGPT" and "getting recommended by Perplexity" are fundamentally different outcomes. One drives direct referral traffic. The other builds brand awareness inside a conversation the user might never leave.

When we first built the scoring, we weighted citations equally across engines. That was wrong. A Perplexity citation at position 2 is a real click. A ChatGPT mention buried in paragraph 3 of a conversational response might never convert to a visit.

We rebuilt the weighting to account for this. The composite score now factors in each engine's typical citation behavior. Not just "did you get mentioned" but "did you get mentioned in a way that actually sends traffic."

What the 80+ scorers have in common

The brands scoring above 80 share a pattern. None of them gamed their way there.

Structured, factual content that AI engines can parse and attribute. Not marketing pages. Product documentation, comparison pages with real specs, how-to guides that answer specific questions with verifiable detail.

Consistent mentions across independent sources. Not just their own blog. Third-party reviews, forum discussions, integration partner pages, community posts. The engines cross-reference heavily.

Clear category positioning. Brands that try to own too many categories dilute their AI visibility across all of them. The high scorers are specific: "best X for Y" where X and Y are narrow enough that the engine can confidently recommend.

Active technical presence. Open APIs documented on common platforms, GitHub repos, Stack Overflow answers. The engines index these more than most people realize. One brand we measured jumped from 28 to 67 in three months after publishing detailed API docs and a public changelog. No other marketing changes during that period.

What doesn't move scores (and we tested it)

Three approaches we watched brands try that had no measurable impact:

AI-optimized blog posts that just restate the product pitch in Q&A format. The engines recognize thin content. Scores stayed flat or dropped when the content lacked substance behind the formatting.

Buying backlinks for AI visibility. Traditional SEO link-building had near-zero correlation with AI recommendation scores in our data. The engines weight signals differently than Google's PageRank.

Optimizing for one engine only. A brand that optimized exclusively for ChatGPT saw their Perplexity and Gemini scores drop. The optimization created format lock-in that other engines couldn't parse well. Going deep on one engine can actually hurt your visibility on the others.

What we'd build differently

If we started this tool today, we'd skip the composite score entirely for the first version and ship per-engine breakdowns only. The composite works as a headline number, but every actionable insight lives in the gaps between engines.

We'd also track score changes over time from day one. The brands asking for repeat reports taught us that the delta matters more than the absolute number.

Where this leaves most brands

AI engines are becoming a primary discovery channel, and most brands have almost no visibility there. The average score of 31 means there's a real gap between "we rank on Google" and "AI recommends us."

There's no single fix. It takes structured content, third-party mentions, clear category positioning, and understanding that each engine has its own recommendation logic. The spread between engines is the diagnostic, not the average.

on September 16, 2026