5
6 Comments

Building in the AI reliability space — looking for feedback

Hey everyone 👋

I'm exploring a problem around AI customer support: teams can deploy an AI chatbot successfully, but understanding its quality across real customer conversations can become difficult as usage grows.

I've been talking to AI founders and collecting different perspectives on how they currently evaluate their systems.

My next step is to build a very small prototype around one specific problem I've identified from these conversations.

I'm intentionally keeping the first version simple rather than trying to build a huge platform.

I'd love to hear from other Indie Hackers:

If you were building an AI customer-support product, what would you want an evaluation tool to tell you every day?

Would really appreciate your thoughts, especially from people who have built or deployed AI products.

on August 12, 2026
  1. 1

    The useful daily output isn’t a score. It’s a ranked action queue: the conversation, business impact, likely root cause, owner, and whether the fix survived replay. I’d start with three buckets, wrong answer, false confidence, and failed handoff, then show expected cost per 100 conversations. I advise AI operators on turning signals like this into an execution loop. If useful, I can outline the first-week test I’d run with three design partners.

  2. 1

    The challenge of understanding AI quality as usage grows is definitely real. It’s easy to get buried in metrics, but the truly hard part, in my experience, is translating those into actionable improvements. When thinking about an evaluation tool, a key thing for me would be not just 'what' the AI did, but 'why' it did it - and, more importantly, 'what specifically needs to be tweaked' to get a different outcome next time. For instance, pinpointing specific conversation segments where the AI consistently misunderstood intent, or where it provided an unhelpful response, rather than just an overall "accuracy" score. That kind of granular, prescriptive feedback is gold.

  3. 1

    The first thing I'd want is a risk-weighted failure queue, not one aggregate score. Tag each conversation by customer impact, model confidence, evidence support, and whether a human later reversed it. Then rank by expected loss: frequency multiplied by severity multiplied by detectability. A 1% hallucination that creates refunds should outrank 10% awkward phrasing.

    For a small prototype, replay the same held-out conversations before and after every prompt, model, or knowledge-base change and show the regression delta. The kill condition is simple: if founders do not change a prompt, handoff rule, or knowledge article after reviewing the daily queue, the evaluation is reporting rather than a product workflow.

  4. 1

    I’d separate the daily evaluation into three layers: outcome, safety, and drift. Outcome asks whether the customer’s issue was resolved; safety asks whether the answer should have been refused or handed to a human; drift asks whether those rates changed after a prompt, model, or knowledge-base update. The dashboard matters less than turning a bad conversation into a prioritized fix with an owner and a before/after check. A useful first prototype could sample ten conversations per day, let a reviewer label those three dimensions, and show only the regressions. Which signal are founders currently willing to review every morning: unresolved intents, unsafe confidence, or quality changes after deploys?

  5. 1

    honestly for me it'd be less about raw metrics and more about "would this response have made me trust the company less." things like, did it actually solve the person's problem or just sound confident while dodging it, did it know when to hand off to a human instead of guessing, and consistency over time (not just one good convo, but not drifting into bad answers after a few hundred conversations). the trust piece is the hard one to quantify but it's usually what actually gets escalated to a founder's inbox lol. curious what you're hearing most from the AI founders you've talked to so far, is it mostly accuracy issues or more the "doesn't know when to stop" problem?

  6. 1

    What you're describing sounds less like an AI reliability problem and more like an evaluation problem that only becomes obvious once teams have enough real conversations. I'd be curious whether the prototype is centered on finding failures, measuring quality over time, or helping teams decide what to fix first.

Trending on Indie Hackers
What 100B+ Claude tokens actually look like inside a tiny company User Avatar 32 comments 4 months to go. Chrome extension live. Web search integrated. 4 users. $0 revenue. Still here. User Avatar 25 comments Solo → Pre-Seed: The Tool Stack Decision That Will Either Save or Sink Your First 18 Months User Avatar 24 comments Two-way is not the same as symmetric User Avatar 20 comments Show IH:GSL Runtime: Moving Beyond .vrp Input User Avatar 7 comments Show IH: Apollodorus Video - browser-based video editor that runs locally User Avatar 6 comments