2
30 Comments

I built a tool to measure how visible your brand is in AI responses

I've been working on a simple, stateless platform for understanding how brands show up in AI-generated answers.

The idea is pretty straightforward:

Enter your brand name and your competitors.

The platform generates prompts you can run in AI tools.

Paste the AI responses back into the platform.

It analyzes the responses to see whether your brand and competitors are being mentioned and how visible they are.

The interesting part for me is that traditional search visibility is changing. People are increasingly asking AI tools for recommendations, comparisons, products, and services — but it's not always obvious how a brand is represented in those answers.

I wanted to build something lightweight that lets people actually test this themselves using real AI responses, rather than relying on assumptions.

It's currently an early version, and I'm looking for feedback from people.

Would love to hear:

What would you want to measure in AI responses?

What competitor insights would be useful?

Is this something you'd actually use for your own brand? Try it here: https://fancy-bread-1238.signum760.workers.dev/

Feedback is very welcome.

on September 25, 2026
  1. 2

    A useful extension is to separate being named from being cited with evidence. Record the exact question, answer version, and supporting passages for every run, then rerun the same prompt later; otherwise a visibility score can rise because the model guessed once rather than because your content became reliably retrievable.

    1. 1

      Thanks for the suggestion! I’ve updated the analysis to now include the AI tool used and run date for each prompt, so it’s easier to track where and when a result came from and compare visibility over time.

      Appreciate the feedback — this was a useful addition.

      1. 1

        Nice update. Logging the AI tool and run date per prompt is exactly what makes week-over-week compares trustworthy.

    2. 1

      That’s a really interesting point, especially the distinction between being mentioned and being supported by evidence. I hadn’t thought about it that way. Thanks for sharing this!

      1. 1

        Glad it landed. The useful split is a name mention versus an answer that leans on checkable evidence. If you score both the same way, you hide whether the model is guessing or actually citing.

        1. 1

          you can see the change in the app now !🎉

        2. 1

          Exactly! Thanks for this valuable thought 🎉

          1. 1

            Glad it helped. Appreciate you reading it through.

            1. 1

              Thanks alot for giving your valuable feedback 🎉🎊

  2. 1

    One tweak worth trying: phrase the prompts as buying tasks ("recommend a CRM for a 10-person agency") rather than asking about brands directly. Most buyers describe a problem to the model, not a brand, so whoever gets named in those task-shaped queries is who actually captures demand. I'd also rerun each prompt a few times and average the results — session-to-session variance is big enough that a single snapshot can mislead.

    1. 1

      Thanks for the feedback ! i tried to implement your suggestion by adding 5 more prompts centralised around buying prompts

      1. 1

        you can see the change in the app now !🎉

    2. 1

      Thanks for the suggestion! I’ll definitely try iterating on the prompts to make them more buying-task focused, and rerun them multiple times to account for the variance. Really appreciate you taking the time to comment!

  3. 1

    We run an AEO scan at UtilitySEO, so this is adjacent to what we’re building — disclosure upfront.

    The manual copy-paste approach is actually the right call. When we had someone run an AI visibility tracker across five engines on our own domain, the results diverged sharply: GPT cited us 5/5, Perplexity 4/5, Qwen 1/5, GLM 0/5. API responses often differ from what users see in the chat interface, so pasting real responses avoids that gap.

    On your attribution question: we’re at DA 3 with four visits a month, so even if every AI engine mentioned us, we couldn’t measure the conversion. The honest answer is that attribution from AI mentions probably works like word-of-mouth attribution — you know it’s happening because branded search goes up, not because you can trace a click.

    The signal-first framing is the right one. Prove the measurement is stable before claiming it predicts anything.

    1. 1

      Thanks for the detailed comment! The point about API responses differing from what users actually see is especially interesting. And I agree that for now it makes more sense to focus on making the visibility signal useful and reliable before trying to connect it to business outcomes. Appreciate you sharing your experience!

  4. 1

    The tricky part here is that visibility in AI responses is a different signal than whether users actually convert from that visibility. You measure whether you're mentioned when someone asks ChatGPT about your category—that's the metric. But the real question is whether brands that rank higher in AI responses also get more traffic, more leads, or more customers compared to brands that don't appear.

    That's the engagement vs conversion gap. A brand could show up prominently in AI responses and still get zero traffic if those conversations don't lead anywhere. Or the opposite: a brand barely mentioned but highly relevant to the actual purchase decision could drive more revenue per mention.

    How are you thinking about tracking whether visibility correlates to actual business outcomes for early testers?

    1. 1

      Yeah, I think that’s an important distinction. Right now I’m deliberately treating AI visibility as a signal rather than claiming it directly translates into conversions.

      The current version is focused on making that signal measurable first — how often a brand appears, how prominently it appears, what context it’s mentioned in, and how that compares with competitors across the same prompts.

      The business-outcome question is definitely the more interesting layer. Ideally, you’d be able to connect changes in AI visibility with things like branded traffic, referral traffic, leads, or eventually conversions, but I think there’s a real attribution problem there because an AI response can influence someone without producing an obvious click or referral.

      For the early version, I’m more interested in establishing whether the visibility data itself is useful enough for brands to track consistently. If people start seeing meaningful changes in that signal, then connecting it to downstream outcomes becomes a much more interesting experiment.

      Curious how you’d approach that attribution problem, especially for AI interactions that don't result in a directly trackable visit.

  5. 1

    Really relatable. How much time do you put into this each week?

  6. 1

    Interesting approach. What was the hardest part to get right?

  7. 1

    Nice progress. What is the next thing you are focusing on?

  8. 1

    Helpful post. How did you get your first bit of traction?

  9. 1

    What made you pick this stack over the alternatives?

  10. 1

    Solid lesson. Which channel has worked best for you so far?

  11. 1

    What behavior would validate this beyond people finding the visibility score interesting—changing their positioning, monitoring it repeatedly, or making a decision from the result?

    1. 1

      Yeah, I think that’s the real validation question. A visibility score being interesting for 30 seconds doesn’t necessarily mean there’s a useful product behind it.

      For me, the stronger signals would be someone running the analysis repeatedly over time, comparing themselves against competitors, digging into why they’re appearing or not appearing, and then actually changing something based on what they find.

      Even better would be seeing someone come back after making a change and run the same prompts again to see whether the AI responses changed.

      That’s probably the behavior I’d consider much stronger validation than just people checking their score once.

      1. 1

        That post-change recheck is the strongest signal there. Could be useful to compare notes on that as it develops over email sometime, if you’re open to it.

        1. 1

          Yeah, that makes sense. The post-change recheck feels like a much stronger signal than someone just checking their score once. I’d be happy to compare notes as I learn more from people using it. Thanks for the thoughtful discussion!

          1. 1

            Might be easier to compare notes over email as you learn more — open to that?

            1. 1

              Actually, I was thinking it might be better to keep this conversation here on Indie Hackers for now. Happy to compare notes here — hope that works for you!

            2. 1

              Yep ! 👍😊