1
2 Comments

Our AI gave the same person a 6 and a 7. Here's how we fixed scoring drift

Quick build-in-public from building Aisa (aisa.to), an AI that assesses how well people actually use AI tools.

Early on our biggest problem wasnt the scoring model itself, it was consistency. The same kind of answer could score a 6 or a 7 depending on small wording differences in the conversation. For something people might put on a CV, that drift kills trust fast.

What fixed it was a second pass. After the conversation ends, a separate model re-reads the whole transcript holistically and re-checks every score against quoted evidence, instead of scoring turn by turn in the moment. Turn-by-turn is reactive and noisy. The calibration pass sees the full picture before committing to a number.

It costs more tokens per assessment, but it's the difference between a score people trust and one they argue with.

Anyone else building eval or assessment products solved scoring consistency a different way? Curious what worked for you.

on June 6, 2026
  1. 1

    Hi, sir.
    Profile: https://topstar-ai.github.io
    I’d really appreciate the opportunity to connect and promise good benefit to you.
    Looking forward to your thoughts.
    Best regards.

  2. 1

    not building an eval product myself, but that second-pass instinct shows up anywhere AI output has to be trusted. the model is confident in the moment, and a fresh read of the whole transcript with the evidence in front of it catches what turn-by-turn scoring couldn't see. it's the same reason a separate review pass beats inline checks on AI-written code. one thing i'd watch is that the calibration pass nails consistency but consistency and correctness aren't the same thing. a score can be rock stable and still calibrated to the wrong bar if the rubric has a gap, and killing the variance won't surface that. does your second model ever disagree with itself across runs, or is it locked in now?