Every validator I tried is a feel-good machine. Paste an idea, get a high score, a paragraph
about why it's promising, and permission to spend three months. At one point I pasted this
very product into one of them and it loved it. That told me what the score was worth.
So I built the opposite. The honest way to show it is my own demo run, published unedited:
https://whittleos.com/sample-discovery
One Discovery run, start to finish:
- 20 sub-markets planned → 60 searches → 93 results → 24 pages read in full
- 69 candidate ideas in, 36 killed by the deal-breaker checks, 5 finalists out
- 56 documented problems from 41 sources. 32 carry a link you can open; the other 24 are
labelled as the tool's own estimate instead of being dressed up as evidence
- Best finalist: 52/100, grade C, verdict "test it first" — ship a narrow landing page
before writing any code
The 52 is the point. That's my showcase run, sitting on my own marketing pages, and I left it
at 52 because the alternative was inventing an 88. A tool that never says no isn't protecting
you from anything.
Mechanically: the 0-100 score is computed in code from a fixed weighted rubric rather than
asked of a model. Validation plans are async by design — a landing page and a real CTA, never
"go book 20 discovery calls". There are 17 analysis modes; Discovery, where it sources the
candidates itself instead of scoring one you bring, is the flagship.
What I want from you: open the run and tell me where it's wrong. And the question I actually
can't answer myself — is a kill-biased tool useful, or does everyone secretly want the 88?
Prepaid credits, no subscription. Signup gives 2 free credits, enough for the single-idea
check; Discovery runs cost credits because each one burns real search and model spend.