6
12 Comments

The number that made us rethink our entire positioning

We built AISA (aisa.to) as an AI skills assessment. After ~1,900 assessments, one stat changed everything: the correlation between self-reported AI confidence and measured ability is r = 0.07 to 0.24. Basically zero.

People who rate themselves 8/10 on AI skills often score below 50/100 on task-based measures. Average across all users: 48/100. 67% fall below "Proficient."

We started thinking we were building a learning tool. Turns out the real value is the gap itself. Organisations do not know where their teams stand on AI capability, and self-assessments are nearly worthless.

The EU AI Act made this concrete. Article 4 requires "sufficient AI literacy" for anyone deploying AI systems. When we published our compliance checklist (aisa.to/blog), compliance teams reached out because they had no way to measure what Article 4 demands.

That shifted our whole go-to-market. We stopped saying "upskill your team" and started saying "measure what you cannot see." Same product, different conversion entirely.

Curious if anyone else has had a regulation or external event completely reframe how they position their product.

on September 27, 2026
  1. 1

    Yes, and I'm in the middle of it right now. I run a menu photo tool for independent restaurants, and in May DoorDash launched free AI photo retouching inside its merchant portal. Honestly, my first reaction was panic. What's keeping me going is staying niche and listening to what owners actually ask for: real photos of the food they serve, ready for every delivery app they're on, not just one. If I stay focused on that, I think there's room. Fingers crossed.

  2. 2

    The gap is striking. Have you broken the scores out by job role or by how often people actually use AI at work? A company-wide average could hide a team that needs much more support, and that breakdown might be the most useful part of the report for a buyer.

  3. 1

    Yes — a regulation (or a hard outside number) can reframe the same product overnight.

    What usually changes conversion isn’t new features; it’s which decision the buyer is trying to make. “Upskill the team” asks for training budget. “Measure what you cannot see” asks for evidence a compliance owner can attach to a gap. Same assessments; different buyer question.

    If I were testing the new pitch, I’d keep one proof artifact close to the ask: the confidence-vs-score gap chart plus a one-page “what this number means for Article 4 literacy.” That makes the reframe feel like evidence, not a slogan swap.

    — Cameron M Deans

  4. 1

    At n≈1,900, the standard error on r is only ~0.023, so 0.07 is a real signal — just a weak one. Squared, it means self-ratings explain under 1% of the variance in task scores. That's the cleaner way to state it: the self-rating carries almost no information. And your sample is self-selected — people who volunteer for an AI skills test already skew AI-curious — so in the broader workforce the gap is probably wider, not narrower.

  5. 1

    The low correlation is a clearer story than another general AI training pitch. It also changes who has the problem: the person buying an assessment may need a team-level gap, not an individual quiz score. When compliance teams contacted you after the checklist, what specific decision did they need the assessment to support?

  6. 1

    The r=0.07 correlation is devastating and not surprising. We see the same gap in SEO — site owners consistently rate their optimization as "pretty good" and then an objective scan finds sixty or seventy issues they had no idea about. The confidence-ability gap is not unique to AI skills; it shows up wherever the feedback loop is slow or invisible.

    The pivot from "upskill your team" to "measure what you cannot see" is exactly right. Measurement-first positioning works because it creates evidence before asking for behavior change. We went through a similar shift — stopped leading with "improve your SEO" and started with "find out what is actually wrong in 30 seconds, free." The conversion difference was immediate.

    The EU AI Act angle gives you something most indie products never get: regulatory demand pulling buyers toward you. Compliance teams already have a budget line and a deadline. That is a wedge into enterprise that does not require a sales team.

  7. 1

    This is a great example of why the problem you discover can be more valuable than the problem you originally set out to solve. The shift from “help people improve” to “help organizations measure the gap” completely changes the buyer and the urgency. I’d be curious how you’re now identifying and reaching those compliance teams.

  8. 1

    This is a textbook case of the benchmark becoming the product. The assessment itself is fungible — anyone can write task-based questions — but 1,900 scored assessments is a norms table nobody else has, and that is your actual moat. One caution on the Article 4 angle: compliance buyers purchase on a different clock than learning buyers. The budget sits with legal, the decision is risk-driven, and it is often a one-off audit unless you attach it to a recurring requirement like annual retesting or onboarding screening. If you can make the retest the default — "your Article 4 evidence expires in 12 months" — the regulation stops being just a positioning frame and starts being a retention engine.

  9. 1

    Great example of letting the data pick the positioning. That gap between what people say and what they can do shows up in every self-reported survey, which is why task-based measures (or follow-up questions that ask for a concrete example) are so much more useful. The EU AI Act angle is a strong wedge.

  10. 1

    r = 0.07 to 0.24 is a great headline, and "measure what you cannot see" is a much sharper promise than "upskill your team". It also explains why the old pitch was hard: nobody buys training for a gap they don't believe they have.

    Two thoughts on the Article 4 angle:

    • Compliance buyers need evidence they can file. A per-team report with the date, the method, and the score distribution, in a format that drops straight into an audit folder, is probably worth more to them than any individual score. That's also the thing they'd forward internally, which makes it your sales asset.

    • The retest is your retention loop. If the gap is the value, the second assessment three months later showing the gap closing is what renews the contract. I'd build the "before and after" view before anything else on the roadmap.

    One question: how do you keep a task-based test meaningful the second time someone takes it? If people remember the tasks, the improvement you show at renewal could be partly memory.

  11. 1

    r = 0.07 to 0.24 is a wild number to build a pitch around, and it makes sense that "measure what you cannot see" converts better. we saw a tiny version of this with testers who said the app was fine and then never opened it again, so what people said and what they did barely lined up. did the compliance angle change your pricing too or mostly the messaging?

  12. 1

    Really good writeup, thanks for sharing it. What's the next thing you're planning to try here?