Last week an internal audit ranked our product last out of nine. The top complaints, from ten independent reviewer personas: our landing "demo" was a scripted auto-play animation with zero input elements, you couldn't try anything without a full signup, and the site never said how we differ from Zapier-style tools.
All three were true.
So this week we shipped the uncomfortable fixes:
The fake demo is gone. The landing page now runs our real analysis engine — you type a recurring job you do, it scores how repeatable it is. No account. We cap it at 100 real runs/day and show a clearly-labeled sample when we're over budget, because a fake "live" result would just be the old lie in a new place.
We repositioned from "build assets from your work" to measurement: how AI-ready is your work? The score was always the most honest thing our engine produced — now it's the product.
Every diagnosis can become a shareable report URL. And we publish aggregate stats (a Work Reproducibility Index) with one rule stolen from our sibling product's accuracy page: any bucket under n=10 says "collecting", never a percentage.
Traction reality, since this is build-in-public: 6 users, $0 revenue, and a sunset date (Oct 30) if paid usage stays at zero. The pivot is our answer to why those numbers looked like that.
If you want to poke at the live demo (leverageos.dev) or the index and tell me where it's still lying to you, I'd genuinely like that.
Publishing the audit results against yourself is the part I respect most here. Most people would have quietly fixed the demo and never mentioned ranking last. The "collecting" label instead of a fake percentage under n=10 is a small detail that says a lot, that is usually the first thing that gets cut under pressure to look more finished than you are. With 6 users and a hard sunset date, I would watch whether people run the analysis once out of curiosity or come back with a second real task, that repeat use is probably a stronger signal than the score itself.
The audit clearly gave you evidence that the previous experience wasn't working, but repositioning the score from a component into the product is a bigger conclusion.
What evidence made you confident the score itself had demand, rather than simply being the strongest part of a product users still weren't paying for?