5
6 Comments

the hardest part of building an AI assessment isn't the AI

been building aisa.to for a while now — its a conversational AI that assesses peoples AI skills through dialogue instead of quizzes.

the AI part (scoring, calibration, persona classification) was honestly the easier problem to solve. the thing that took way longer than expected? defining what "good" actually means.

like, what does it mean to be good at using AI? is it about prompting? output verification? knowing which tool to use when? integrating AI into your workflow vs just using it occasionally?

ended up with 11 criteria across 5 dimensions and 10 distinct user personas. sounds clean on paper but getting there was messy — lots of conversations where someone scored high on prompting but couldnt tell when the AI was making stuff up. or someone whod never heard of prompt engineering but had this incredibly efficient workflow.

the biggest surprise: self-reported skill barely correlates with observed skill. ppl who say theyre 8/10 routinely demonstrate 4/10 practices. and some who say "i barely use AI" turn out to have really thoughtful verification habits.

anyone else building assessment or evaluation tools? curious whether you hit the same "defining the rubric is harder than building the product" problem.

on May 23, 2026
  1. 2

    This is the real challenge with AI products:

    Building the AI is hard.
    Defining what “good AI usage” means is harder.

    The self-rating mismatch is especially fascinating

    Really thoughtful work — aisa.to sounds super interesting

  2. 1

    This hits so close to home. In indie hacking, we often obsess over the tech stack, but mapping human intuition into a structured rubric is where the real existential dread sets in.

    Your insight about self-reported skill vs. actual verification habits is wild, but it makes perfect sense. The people who know the "buzzwords" assume they are 8/10, while the pragmatic folks who treat AI like a flawed intern quietly build the best workflows. You're essentially trying to codify messy human behavior, which is a brutal product challenge.

    Definitely feeling a similar flavor of this pain while building PandasRouter. Good luck with the iteration, 11 criteria across 5 dimensions sounds like a solid foundation!

  3. 1

    This resonates a lot. I think we’re entering a phase where the hard part isn’t model capability anymore, it’s defining meaningful evaluation criteria for real human workflows.

  4. 1

    This is such an underrated problem layer in AI products. The model itself often stops being the hardest part surprisingly fast. The much harder question becomes:
    “what are we actually optimizing for?”

    I’ve noticed something similar while working on outreach/signal analysis systems. Defining what a “high quality lead” actually means in practice gets messy very quickly once you move beyond surface metrics.

    Because real-world quality is usually contextual, behavioral, and hard to reduce into a simple score.

  5. 1

    The rubric problem is the part most teams underestimate.

    For real work, I would separate "can produce a good answer" from "can verify the answer before it reaches a customer or decision." The second one is usually where the skill gap shows up: checking sources, spotting overconfidence, knowing when to escalate, and fitting the output back into the existing workflow.

    That also makes the assessment more useful commercially, because it tells a team which AI tasks are safe to delegate now and which ones still need training or human approval.

  6. 0

    This is a strong insight because the real product is not “AI skill testing.” It is judgment measurement.

    Most people can claim they know how to use AI, but the valuable part is observing whether they can verify outputs, choose the right tool, understand failure modes, and actually integrate AI into real work. That makes the rubric the product, not just the conversation layer.

    I’d be careful with the name before this gets positioned more widely. aisa.to is short, but it feels a little unclear for something that could become a serious AI capability assessment platform for teams, hiring, training, or internal upskilling.

    For this category, the name needs to feel more credible than a quiz tool or AI experiment. Beryxa .com would fit that direction better because it sounds like a serious assessment and intelligence platform, while still leaving room if the product expands beyond individual skill checks into team benchmarking, role-fit analysis, or workforce AI readiness.

    The product is already touching trust, scoring, calibration, and professional credibility. The brand should help carry that seriousness instead of making the homepage do all the work.