Every idea checker I tried is a feel-good machine. I pasted this very product into one of them and it loved it. That told me exactly what the score was worth.
So I built the opposite. The only honest way to show it is my own demo run, published unedited:
https://whittleos.com/sample-discovery
One run, start to finish: 20 sub-markets → 60 searches → 24 pages read in full → 69 candidate ideas in, 36 killed at the deal-breaker gate, 5 out. It cites 56 documented problems from 41 sources — 32 carry a link you can open, and the other 24 are labelled as the tool's own estimate instead of being dressed up as evidence.
Best finalist: 52/100, "test it first."
The 52 is the point. That is my own showcase run, sitting on my own marketing pages, and I left it at 52 because the alternative was inventing an 88.
No signup to read it. If you open it, tell me where the run is wrong — that is genuinely what I want out of posting this.
The honesty is the right product decision and the wrong sales motion, at least for idea-stage founders. Nobody with a fresh idea pays monthly to be told no, but someone who has already burned six months and real money will pay well to avoid the second mistake, which is why I would point this at accelerators, angel groups, and second-time founders first. Same engine, completely different willingness to pay.
Your point about single-sample noise causing false kills is the part I'd want to dig into. A tool built to say no also needs a way to catch its own bad rejections.
Have you tried taking a few rejected ideas and checking the specific deal-breaker with potential customers? “We found no evidence of demand” and “we found evidence there's no demand” would lead me to different next steps.
I'd find that audit more useful than the headline score: which rejections held up, and which turned out to be gaps in the research?
The distinction is right and we're on the wrong side of it. Our most common
rejection is labelled "No real problem behind it" — 21 of the 36 on that run, and
around 440 of the ~700 across every real run. What the check actually does is
look for the candidate among the problems that run collected, and kill it if it
isn't there. That's "we found no evidence", printed as "there is none". Your
phrasing is better than mine, and the label is changing.
The audit you're describing I can't run honestly yet: it needs founders who ran
it, disagreed with a rejection, and went and checked. I don't have that data, and
I'd rather say so than assemble something shaped like it.
What I can do meanwhile is the part that doesn't need anyone's permission — all
36 rejections are published with their stated reason at the bottom of
whittleos.com/sample-discovery. Telling me which of those look like gaps in the
research rather than real kills is the most useful thing anyone in this thread
could do for me.
The distinction between “no evidence” and “no demand” is a valuable guardrail. I’d make the output actionable by pairing each rejection with one falsifiable next test: who to interview, what behavior would change the verdict, and a fixed time or budget cap. Publishing the raw run is also a strong trust signal—over time, anonymized follow-ups could become the product’s most useful dataset.
Splitting your three, because they land differently for me.
The falsifiable next test is a real gap, and half of it already exists in the
wrong place. The single-idea check returns a "what would need to be true" ladder —
the concrete conditions that would raise the verdict — but a Discovery rejection
carries only its reason. So the mode that examines one idea tells you what would
change its mind, and the mode that throws away 36 does not. That asymmetry is
indefensible and I had not noticed it until you wrote this.
"Who to interview" is the one I'll push back on, and not out of squeamishness.
This is built for founders who sell without calls: the validation plans it writes
are a landing page, a real CTA and async outreach, and there is a scanner on the
output that rejects "book a call" and its cousins before they reach the page. If
the next test requires interviewing ten people, it has quietly become a different
product for a different founder. What changes a verdict is behaviour, and
behaviour can be observed without a conversation.
The fixed time or budget cap I agree with completely. It exists today on the one
idea the run recommends, and nowhere else — including on all the rejections, which
is your point again from the other side.
On the dataset: that loop shipped yesterday. A plan writes its target and deadline
down before the test runs, you come back with the numbers you counted, and code
grades them against the bar that was set first — you cannot resubmit upward once
you have seen how it graded. So the mechanism is there and has nothing in it yet.
Whether it fills is not up to me.
The 52 creates a measurement system where the score itself becomes the data, not the validation.
Here's why that matters: founders make decisions based on whether they trust the measurement, not the idea itself. A 52 with cited sources and transparent reasoning tells you "this tool is honest." An 88 tells you the tool is selling confidence.
So a founder sees your own product at 52 and thinks: "If this honest measurement says 52 for something the creator believes in, what does that tell me about my 60?" Nothing helpful. The 88 would have been noise.
But the 52 — that's information about which measurement systems are worth paying attention to. You're not trying to build a predictor that scores ideas correctly. You're trying to build a measurement filter that separates founders who can trust their own thinking from founders who need someone to tell them what to believe.
Most "idea validators" are in the business of making everything look viable. You inverted it: you're in the business of proving you're not.
"A measurement filter rather than a predictor" is a better description of this
than the one on my own landing page, and I'm going to sit with it.
What I can't answer yet is what the framing costs. If the product selects for
founders who can already trust their own thinking, that is a smaller room than the
one everyone else in this category is selling to — and the founders who want to be
told what to believe are exactly the ones who convert on an 88. Whether the room is
big enough is genuinely open, not modestly open.
One thing I'd add: the filter cuts both ways. The same run that says "52, and here
is every source" also says what it could not check — the published one links 32 of
its 56 problems and labels the other 24 as the tool's own estimate. A founder who
finds that annoying is not going to like anything else it says either.
Also, could you elaborate if the skew toward CRM & real estate related plays comes form your personal input or the systems assumptions - and have your analyzed https://www.ideabrowser.com ?
That's my input, not the system. The run was seeded with one market — "real estate
agent and property management tools for solo agents" — and the planner expanded
that seed into 20 sub-markets, all inside it: listing-description generators,
lease-renewal reminders, open-house sign-in, an MLS comps extension, an IDX
widget. So the CRM flavour is the seed showing through, not a preference baked
into the tool. Discovery also runs with no market given at all, which produces a
completely different spread; I published a seeded run because a focused one is
easier for a stranger to check.
Haven't looked at ideabrowser. It's not in the set I've compared against
(ValidatorAI, VenturusAI, FounderPal, ValidateMySaaS, PainMap, Trend Seeker,
ideaproof), so that's a gap rather than a verdict — I'll go read it.
You may find that founders who need their ideas scored in the first place have not yet been properly trained in problem discovery or in developing truly distinctive, high-potential concepts.
I’m formally trained in evaluating ideas and spent 20 years at large firms doing this work for clients. I led teams that would develop up to 100 concepts for a client, with the goal of eventually producing one that became a major commercial or communications success. More recently, I’ve been automating the way professional product-innovation firms approach problem discovery and idea development. You may want to consider incorporating a similar component.
The question is not whether to inflate scores—don’t; that would quickly undermine confidence and trust. The real question is whether you can help a founder move an idea from a score of 52 to 85+. The value is not simply in making the scoring more honest; it is in providing the direction and feedback that materially improves the underlying idea.
Agreed on not inflating. The direction-and-feedback point is fair, but it's less
absent than it looks.
Two pieces already do it. Kill My Idea returns a "what would need to be true"
ladder — the concrete conditions that would raise the verdict to hold, to test, to
build — which exists so the founder can argue back with facts instead of taking
the number. Pivot Wedge takes a broad or weak idea and narrows it into something
one person can sell, or refuses and says no wedge exists rather than inventing
one. And yesterday I shipped a line that names the single lowest-rated area of the
top pick, so the next test aims at the thing most likely to sink it rather than at
demand in general.
Where I'd push back is the destination. 52 to 85 should not be reachable by better
advice, because the two things holding that score down are confirmed distribution
and hard evidence of demand — neither of which is knowable by reading pages about
an idea. They change when the founder runs a test and reports what happened, which
the product now grades against a bar written before the test started. So there is
a path from 52 upward and it runs through evidence, not articulation. If
articulation moved it, I'd have rebuilt the 88 with extra steps.
If your automated problem-discovery work sits at the front of that — better
concepts entering the funnel rather than better scores leaving it — that's a
different and more interesting conversation.
The linked-evidence vs. estimate split is a strong guardrail; I’d make the score’s uncertainty visible as a range rather than a single number. A useful next validation would be to record the lowest-scoring criterion and ask founders to test only that assumption with a pre-set budget or time limit. That turns “say no” into a conservative decision rule, not just a harsher score.
Both land, and the second is a rule I don't have.
On the range: you're right that the number is one draw. The ratings come from a
model at temperature 1, so the same market twice can hand back a different
shortlist. I do pay for sampling where a call is life-or-death — but only in the
single-idea check, whose deal-breaker filters run five independent samples and
vote, after single-sample noise turned out to be the main cause of false kills.
The sweep you're looking at has no voting anywhere: sourcing, triage and scoring
are one sample each. So I can't show you an honest range today, and a range drawn
from a single sample is just an 88 with error bars on it.
On the second: right now the recommended first move is picked by verdict, then by
score — the strongest idea on the list, not the weakest assumption under it.
Yours inverts that: take the criterion that scored lowest and test only that,
against a limit you set before you start. It composes with the part that already
exists, where the plan writes its target and deadline down in advance. That's a
better decision rule than the one I shipped.
Appreciate the transparency on single-sample vs voting — that distinction is the real product insight. Glad the ‘test the weakest criterion with a hard limit’ framing is useful; I’ll keep using that as the default when a score looks decisive but the underlying sample is thin.
The point about setting the target before running the test stood out in your reply. It's easy to move the goalposts once you're attached to an idea. I like that this makes the founder commit to what would count as evidence before seeing the result.
Worth adding, since it shipped a couple of hours after I wrote that: the target
was already fixed before the test, but the answer wasn't. Reporting was free and
unlimited, and the page listed every criterion you missed and by how much — so
you could read "signups 41, needed 50", type 50, and submit again. A bar set in
advance is worth nothing if the answer can be rewritten after you've seen how it
graded.
Now, within one plan, the report that counts is the weakest one you filed.
Corrections downward land immediately; upward ones don't. Not "first report
only", because that would trap a typo — mistype 500 for 50 and you'd have minted
a result you could never withdraw.
What it still can't do is tell whether your numbers are true. Nothing is
verified. The property is that the bar came first, not that the report is honest.
Thanks for explaining the update. I appreciate the distinction between keeping the target fixed and verifying the reported numbers. Your reply makes clear what the tool does and doesn't establish.
The 52 creates a measurement system where the score itself becomes the data, not the validation. Most tools collapse to "well, maybe it works" because they optimized for confidence, not clarity. When a founder sees 52 vs 88 on the same idea, they're not evaluating the idea - they're evaluating whether they trust the measurement. That's where the decision lives.
That's the bet, and it has an obvious cost: a 52 loses to an 88 in the market
even when the 52 is the honest number. Trust in the measurement only pays off
after someone has used the thing twice — the 88 pays off immediately. I don't
have a clever answer to that, which is roughly why the demo is my own run left
at 52 instead of a case study.
The refusal to inflate a 52 is the interesting part. Have founders actually changed a build decision after getting a low score, or is the value still mainly in the research?
Honest answer: I don't have that data yet. It's early enough that nobody has
come back to me with "I killed it, here's what I did instead" — and I'd rather
say that than tell you a story about outcomes I can't show you.
Two things that aren't guesses. "Build it" is a verdict none of these modes has
ever issued, on any run, including my own — it needs confirmed distribution and
hard evidence of demand, and you can't get either by reading pages about an
idea. And the build decisions it has changed so far are mine: it killed five of
the eight ideas I put through it.
Which is why the piece I shipped this week is the one that could actually answer
you. The validation plan writes its target and deadline down before you run the
test; you come back with the numbers you counted, and a plan created after your
test started is refused — so the bar can't be picked once you already know the
result. Ask me again in three months and I'll either have outcomes or I'll tell
you I don't.
The honest 52/100 and the split between linked evidence and estimates are the strongest trust signal here. I’d add a small tally of the 36 killed ideas by deal-breaker reason; that would make the demo useful not just as a score, but as a map of which assumptions keep failing.
That tally is already in the run — it's just buried below the five finalists,
which is a fair hit on the layout. On /sample-discovery, under "Why we rejected
the rest":
36 of 69 dropped. No real problem behind it 21 · Too much hands-on support 9 ·
Needs heavy compliance 8 · Hard for a solo founder to sell or run 6 · Just an AI
wrapper 6 · Leans too hard on one platform 4 · Can't charge enough 1 · No clear
way to make money 1 · No repeatable way to reach buyers 1.
Those sum to 57, not 36, because 21 of the ideas tripped two checks and 15
tripped one — so it's a census, not a partition.
You're right that it belongs higher up. Moving it next to the funnel line.
Built a similar filter for our extension ideas — if we can't explain the acute pain in one sentence, we don't build. Saved months on features nobody asked for.
"If we can't explain the acute pain in one sentence, we don't build" is the same
rule with a human running it, and yours is cheaper. The gate here does exactly
that, and it's why most ideas die: 21 of the 36 drops on that run are "we found no
documented problem", which is the automated version of not being able to say the
sentence.
Months saved on features nobody asked for is the only ROI in this category I
actually believe.