1
0 Comments

the scoring problem nobody warns you about when building AI assessment

been building aisa.to for a while now and the one thing that keeps surprising me is how hard it is to define what 'good' actually looks like when measuring AI skills.

there's no industry standard. no benchmark. no agreed-upon definition of what a competent AI user looks like versus an expert one. we had to build the rubric from scratch — 11 criteria across 5 dimensions — and the real challenge isn't the AI that runs the assessment, it's deciding what the scores should mean.

every time we think we've nailed it, we find an edge case. someone who scores low on prompting but high on verification because they've built an entire workflow around cleaning up AI output instead of getting better prompts. are they a 6 or an 8?

if you're building anything that evaluates human performance, the scoring rubric will take 3x longer than you expect. and you'll still be tuning it.

on May 27, 2026