We ran our own product through a competitor's validator on 2026-06-10, expecting ammunition. What we got was more uncomfortable and more useful.
It returned 72 / 100 and PROMISING — excellent potential, among the top ideas we have seen. On the same day, on the same one-liner, our own balanced check returned KILL at 0 / 100 and our adversarial mode returned KILL at 15 / 100.
The caveats, because we publish them next to the number rather than under it: that is one dated run on one idea, not an accuracy benchmark. We hand our own modes a founder profile the other tool never asks for, so the inputs were not identical. All of those verdicts are model outputs and carry noise, and re-running it today would not reproduce the numbers exactly. We ship it as a reproducible script for that reason.
The uncomfortable part is that we cannot tell you which verdict was right. Nobody can, about an unbuilt idea. There is no ground truth, and anyone in this category claiming accuracy is claiming something unmeasurable.
So what is the result actually good for? It is a behaviour demonstration. One of those outputs came with the reasons that produced it and one came as a number you cannot argue with. That difference is checkable, and it is the only thing in this comparison that is.
It is also why we killed our own instinct to lead with the score. A score is the most quotable and least useful artifact a tool in this category produces, including ours.
The five criteria we ended up judging the whole category by, ours included: https://whittleos.com/guides/best-startup-idea-validation-tools-2026