1
0 Comments

What should an AI search agency actually measure? A checklist from 84 test runs

Most GEO reports collapse everything into one AI visibility score. After running 42 buyer-intent queries across two engines for our latest B2B SaaS study, I think that number can be dangerously neat.

If you are hiring an AI search agency or GEO agency, here is the measurement checklist I would use before signing a retainer.

1. A disclosed query set

The agency should show the prompts, explain why they represent buyer intent, and state the denominator. Visibility across 40 commercial questions means something very different from one cherry-picked prompt.

2. Engine and configuration details

ChatGPT with browsing disabled and Perplexity with live retrieval are not equivalent tests. Record the engine, model, browsing mode, geography, language, account type, and testing date.

3. Per-engine results

Do not accept one blended score without the underlying lines. In our study, 22 companies appeared in both ChatGPT and Perplexity, but 25 appeared in only one. An average would hide that gap.

4. A precise citation definition

Decide in advance what counts. We counted a company only when it was named as a recommendation, solution, or notable provider in response to a buyer-intent query. Passing mentions and generic alternative lists did not count.

5. Raw evidence

Every claimed citation should be traceable to the prompt, full output, date, and engine. Screenshots without the prompt or configuration are weak evidence.

6. Third-party source mapping

Owned content is only part of the picture. In our sample, companies with Wikipedia pages, G2 profiles, and analyst coverage were overrepresented among cited brands. That does not prove causation, but it is a useful signal to investigate.

7. Uncertainty stated honestly

AI responses are non-deterministic. Our first report ran each query once per engine, so we describe it as a structured snapshot rather than a statistical estimate. A provider should distinguish observed citations from inferred potential and avoid guarantees.

8. A repeatable cadence

A baseline is useful only if the query bank and testing method can be repeated. Changes in engine configuration should be logged rather than quietly folded into a trend line.

Red flags

  • Guaranteed citations or rankings
  • One score with no engine-level detail
  • Selective screenshots instead of full outputs
  • No fixed query bank
  • Inferred recommendations presented as measured results
  • No discussion of model variance or testing limitations

The practical lesson from our test was that visibility is both scarce and fragmented. Of 100 B2B SaaS companies, 53 received zero citations across both engines, and even the highest total was only 3 citations.

We published the full B2B SaaS benchmark and methodology, including limitations and the raw-data link.

What would you add to this checklist before paying someone to improve AI visibility?

on July 30, 2026