Every idea checker I tried is a feel-good machine. I pasted this very product into one of them and it loved it. That told me exactly what the score was worth.
So I built the opposite. The only honest way to show it is my own demo run, published unedited:
https://whittleos.com/sample-discovery
One run, start to finish: 20 sub-markets → 60 searches → 24 pages read in full → 69 candidate ideas in, 36 killed at the deal-breaker gate, 5 out. It cites 56 documented problems from 41 sources — 32 carry a link you can open, and the other 24 are labelled as the tool's own estimate instead of being dressed up as evidence.
Best finalist: 52/100, "test it first."
The 52 is the point. That is my own showcase run, sitting on my own marketing pages, and I left it at 52 because the alternative was inventing an 88.
No signup to read it. If you open it, tell me where the run is wrong — that is genuinely what I want out of posting this.
The unedited run is a strong wedge because it gives people something concrete to disagree with. I’d turn the “no signup” choice into a deliberate interview loop: put one prompt under the report—“Which conclusion would you bet against, and what evidence would change your mind?”—and tag each response by step (sub-market, search, page, score). After 5–10 runs, the repeated disagreement tells you whether to improve discovery or the scoring model. Keep 52/100 as the artifact; sell the next run as a falsifiable decision, not a prettier score. That’s a good way to find first users without diluting the anti-hype positioning.
The 52 is much more useful than a made-up 88. I’d expose a second axis next to the score: evidence density (number, independence, and recency of sources) versus fit/feasibility, so a low score from thin data doesn’t look identical to a well-researched rejection. For the five finalists, a one-line “what would falsify this next?” test could turn the report into an experiment queue rather than a static ranking.
Agreed, and it's cheaper than you'd expect — the inputs are already stored per row: whether we read the actual page or only a search snippet, how many independent hosts the evidence spans, when it was last seen, and how many separate sweeps found it again.
Which I know because I published two articles today where I computed exactly that by hand — "25 records from 19 pages", "every row was seen once" — precisely because the product doesn't show it. Doing it in prose, one page at a time, is not a strategy.
Your framing also caught something bigger than you aimed at. This morning I measured the corpus and 1,433 of its 1,434 rows have been found by exactly one sweep. The score is severity times recurrence, so with recurrence flat the score was severity — while the surface printed "comes up a lot" next to rows seen once, on 196 of them. Fixed today: the labels now describe severity, every row states its sighting count, and the explainer counts, from the rows in front of you, how many have been found more than once. Your evidence-density axis is the thing that would have made that visible on the page instead of needing a database query to find.
On "what would falsify this next?" — there is a per-finalist async plan today with a stop-loss and a deadline, but it's phrased as a plan to run rather than a claim to break. Falsification is the better frame and I'm taking it.
The “say no positioning is strong because it separates the tool from feel-good idea graders. To make the verdict more useful for getting traction, Id pair each rejection with one concrete next test—who to interview, what behavior to look for, and a deadline—then track whether founders actually run it. That turns a score into a learning loop and gives you outcome data that is much harder to fake than a higher number.
Half of this exists and you can't see it, which is my problem, not yours: each finalist gets an async test with a promise, a real call-to-action, a target number and a stop-loss date, and the report opens with which one to run first. There's also an outcome loop as of yesterday — once a validation result is recorded, a report can be retracted downward but never upgraded. That's the un-fakeable direction, and it's there because a tool that can raise its own grade after the fact is worthless.
Where you're right and I have no defence: the recommended first move is computed from the survivors only. Rejections get reasons and nothing else. On a tool whose whole pitch is that it says no, the "no" is the output with no next step attached — which is exactly backwards, since a rejection is the moment a founder most needs to know what would change the answer.
One place I'll push back, because it's a position rather than an omission: "who to interview" is something the product refuses to produce. Interview and call suggestions are blocked at the output scanner, not left out by accident — the argument is that a validation step requiring a conversation is the step most people never take. So I'll take the structure of your suggestion — one test, one observable behaviour, a deadline, then track whether it was run — and keep it async.
Your point about single-sample noise causing false kills is the part I'd want to dig into. A tool built to say no also needs a way to catch its own bad rejections.
Have you tried taking a few rejected ideas and checking the specific deal-breaker with potential customers? “We found no evidence of demand” and “we found evidence there's no demand” would lead me to different next steps.
I'd find that audit more useful than the headline score: which rejections held up, and which turned out to be gaps in the research?
The distinction is right and we're on the wrong side of it. Our most common
rejection is labelled "No real problem behind it" — 21 of the 36 on that run, and
around 440 of the ~700 across every real run. What the check actually does is
look for the candidate among the problems that run collected, and kill it if it
isn't there. That's "we found no evidence", printed as "there is none". Your
phrasing is better than mine, and the label is changing.
The audit you're describing I can't run honestly yet: it needs founders who ran
it, disagreed with a rejection, and went and checked. I don't have that data, and
I'd rather say so than assemble something shaped like it.
What I can do meanwhile is the part that doesn't need anyone's permission — all
36 rejections are published with their stated reason at the bottom of
whittleos.com/sample-discovery. Telling me which of those look like gaps in the
research rather than real kills is the most useful thing anyone in this thread
could do for me.
The 52 creates a measurement system where the score itself becomes the data, not the validation.
Here's why that matters: founders make decisions based on whether they trust the measurement, not the idea itself. A 52 with cited sources and transparent reasoning tells you "this tool is honest." An 88 tells you the tool is selling confidence.
So a founder sees your own product at 52 and thinks: "If this honest measurement says 52 for something the creator believes in, what does that tell me about my 60?" Nothing helpful. The 88 would have been noise.
But the 52 — that's information about which measurement systems are worth paying attention to. You're not trying to build a predictor that scores ideas correctly. You're trying to build a measurement filter that separates founders who can trust their own thinking from founders who need someone to tell them what to believe.
Most "idea validators" are in the business of making everything look viable. You inverted it: you're in the business of proving you're not.
"A measurement filter rather than a predictor" is a better description of this
than the one on my own landing page, and I'm going to sit with it.
What I can't answer yet is what the framing costs. If the product selects for
founders who can already trust their own thinking, that is a smaller room than the
one everyone else in this category is selling to — and the founders who want to be
told what to believe are exactly the ones who convert on an 88. Whether the room is
big enough is genuinely open, not modestly open.
One thing I'd add: the filter cuts both ways. The same run that says "52, and here
is every source" also says what it could not check — the published one links 32 of
its 56 problems and labels the other 24 as the tool's own estimate. A founder who
finds that annoying is not going to like anything else it says either.
The honesty is the right product decision and the wrong sales motion, at least for idea-stage founders. Nobody with a fresh idea pays monthly to be told no, but someone who has already burned six months and real money will pay well to avoid the second mistake, which is why I would point this at accelerators, angel groups, and second-time founders first. Same engine, completely different willingness to pay.
The distinction between “no evidence” and “no demand” is a valuable guardrail. I’d make the output actionable by pairing each rejection with one falsifiable next test: who to interview, what behavior would change the verdict, and a fixed time or budget cap. Publishing the raw run is also a strong trust signal—over time, anonymized follow-ups could become the product’s most useful dataset.
Splitting your three, because they land differently for me.
The falsifiable next test is a real gap, and half of it already exists in the
wrong place. The single-idea check returns a "what would need to be true" ladder —
the concrete conditions that would raise the verdict — but a Discovery rejection
carries only its reason. So the mode that examines one idea tells you what would
change its mind, and the mode that throws away 36 does not. That asymmetry is
indefensible and I had not noticed it until you wrote this.
"Who to interview" is the one I'll push back on, and not out of squeamishness.
This is built for founders who sell without calls: the validation plans it writes
are a landing page, a real CTA and async outreach, and there is a scanner on the
output that rejects "book a call" and its cousins before they reach the page. If
the next test requires interviewing ten people, it has quietly become a different
product for a different founder. What changes a verdict is behaviour, and
behaviour can be observed without a conversation.
The fixed time or budget cap I agree with completely. It exists today on the one
idea the run recommends, and nowhere else — including on all the rejections, which
is your point again from the other side.
On the dataset: that loop shipped yesterday. A plan writes its target and deadline
down before the test runs, you come back with the numbers you counted, and code
grades them against the bar that was set first — you cannot resubmit upward once
you have seen how it graded. So the mechanism is there and has nothing in it yet.
Whether it fills is not up to me.
Also, could you elaborate if the skew toward CRM & real estate related plays comes form your personal input or the systems assumptions - and have your analyzed https://www.ideabrowser.com ?
That's my input, not the system. The run was seeded with one market — "real estate
agent and property management tools for solo agents" — and the planner expanded
that seed into 20 sub-markets, all inside it: listing-description generators,
lease-renewal reminders, open-house sign-in, an MLS comps extension, an IDX
widget. So the CRM flavour is the seed showing through, not a preference baked
into the tool. Discovery also runs with no market given at all, which produces a
completely different spread; I published a seeded run because a focused one is
easier for a stranger to check.
Haven't looked at ideabrowser. It's not in the set I've compared against
(ValidatorAI, VenturusAI, FounderPal, ValidateMySaaS, PainMap, Trend Seeker,
ideaproof), so that's a gap rather than a verdict — I'll go read it.
You may find that founders who need their ideas scored in the first place have not yet been properly trained in problem discovery or in developing truly distinctive, high-potential concepts.
I’m formally trained in evaluating ideas and spent 20 years at large firms doing this work for clients. I led teams that would develop up to 100 concepts for a client, with the goal of eventually producing one that became a major commercial or communications success. More recently, I’ve been automating the way professional product-innovation firms approach problem discovery and idea development. You may want to consider incorporating a similar component.
The question is not whether to inflate scores—don’t; that would quickly undermine confidence and trust. The real question is whether you can help a founder move an idea from a score of 52 to 85+. The value is not simply in making the scoring more honest; it is in providing the direction and feedback that materially improves the underlying idea.
Agreed on not inflating. The direction-and-feedback point is fair, but it's less
absent than it looks.
Two pieces already do it. Kill My Idea returns a "what would need to be true"
ladder — the concrete conditions that would raise the verdict to hold, to test, to
build — which exists so the founder can argue back with facts instead of taking
the number. Pivot Wedge takes a broad or weak idea and narrows it into something
one person can sell, or refuses and says no wedge exists rather than inventing
one. And yesterday I shipped a line that names the single lowest-rated area of the
top pick, so the next test aims at the thing most likely to sink it rather than at
demand in general.
Where I'd push back is the destination. 52 to 85 should not be reachable by better
advice, because the two things holding that score down are confirmed distribution
and hard evidence of demand — neither of which is knowable by reading pages about
an idea. They change when the founder runs a test and reports what happened, which
the product now grades against a bar written before the test started. So there is
a path from 52 upward and it runs through evidence, not articulation. If
articulation moved it, I'd have rebuilt the 88 with extra steps.
If your automated problem-discovery work sits at the front of that — better
concepts entering the funnel rather than better scores leaving it — that's a
different and more interesting conversation.
The linked-evidence vs. estimate split is a strong guardrail; I’d make the score’s uncertainty visible as a range rather than a single number. A useful next validation would be to record the lowest-scoring criterion and ask founders to test only that assumption with a pre-set budget or time limit. That turns “say no” into a conservative decision rule, not just a harsher score.
Both land, and the second is a rule I don't have.
On the range: you're right that the number is one draw. The ratings come from a
model at temperature 1, so the same market twice can hand back a different
shortlist. I do pay for sampling where a call is life-or-death — but only in the
single-idea check, whose deal-breaker filters run five independent samples and
vote, after single-sample noise turned out to be the main cause of false kills.
The sweep you're looking at has no voting anywhere: sourcing, triage and scoring
are one sample each. So I can't show you an honest range today, and a range drawn
from a single sample is just an 88 with error bars on it.
On the second: right now the recommended first move is picked by verdict, then by
score — the strongest idea on the list, not the weakest assumption under it.
Yours inverts that: take the criterion that scored lowest and test only that,
against a limit you set before you start. It composes with the part that already
exists, where the plan writes its target and deadline down in advance. That's a
better decision rule than the one I shipped.
Appreciate the transparency on single-sample vs voting — that distinction is the real product insight. Glad the ‘test the weakest criterion with a hard limit’ framing is useful; I’ll keep using that as the default when a score looks decisive but the underlying sample is thin.
The point about setting the target before running the test stood out in your reply. It's easy to move the goalposts once you're attached to an idea. I like that this makes the founder commit to what would count as evidence before seeing the result.
Worth adding, since it shipped a couple of hours after I wrote that: the target
was already fixed before the test, but the answer wasn't. Reporting was free and
unlimited, and the page listed every criterion you missed and by how much — so
you could read "signups 41, needed 50", type 50, and submit again. A bar set in
advance is worth nothing if the answer can be rewritten after you've seen how it
graded.
Now, within one plan, the report that counts is the weakest one you filed.
Corrections downward land immediately; upward ones don't. Not "first report
only", because that would trap a typo — mistype 500 for 50 and you'd have minted
a result you could never withdraw.
What it still can't do is tell whether your numbers are true. Nothing is
verified. The property is that the bar came first, not that the report is honest.
Thanks for explaining the update. I appreciate the distinction between keeping the target fixed and verifying the reported numbers. Your reply makes clear what the tool does and doesn't establish.
The 52 creates a measurement system where the score itself becomes the data, not the validation. Most tools collapse to "well, maybe it works" because they optimized for confidence, not clarity. When a founder sees 52 vs 88 on the same idea, they're not evaluating the idea - they're evaluating whether they trust the measurement. That's where the decision lives.
That's the bet, and it has an obvious cost: a 52 loses to an 88 in the market
even when the 52 is the honest number. Trust in the measurement only pays off
after someone has used the thing twice — the 88 pays off immediately. I don't
have a clever answer to that, which is roughly why the demo is my own run left
at 52 instead of a case study.
The refusal to inflate a 52 is the interesting part. Have founders actually changed a build decision after getting a low score, or is the value still mainly in the research?
Honest answer: I don't have that data yet. It's early enough that nobody has
come back to me with "I killed it, here's what I did instead" — and I'd rather
say that than tell you a story about outcomes I can't show you.
Two things that aren't guesses. "Build it" is a verdict none of these modes has
ever issued, on any run, including my own — it needs confirmed distribution and
hard evidence of demand, and you can't get either by reading pages about an
idea. And the build decisions it has changed so far are mine: it killed five of
the eight ideas I put through it.
Which is why the piece I shipped this week is the one that could actually answer
you. The validation plan writes its target and deadline down before you run the
test; you come back with the numbers you counted, and a plan created after your
test started is refused — so the bar can't be picked once you already know the
result. Ask me again in three months and I'll either have outcomes or I'll tell
you I don't.
The honest 52/100 and the split between linked evidence and estimates are the strongest trust signal here. I’d add a small tally of the 36 killed ideas by deal-breaker reason; that would make the demo useful not just as a score, but as a map of which assumptions keep failing.
That tally is already in the run — it's just buried below the five finalists,
which is a fair hit on the layout. On /sample-discovery, under "Why we rejected
the rest":
36 of 69 dropped. No real problem behind it 21 · Too much hands-on support 9 ·
Needs heavy compliance 8 · Hard for a solo founder to sell or run 6 · Just an AI
wrapper 6 · Leans too hard on one platform 4 · Can't charge enough 1 · No clear
way to make money 1 · No repeatable way to reach buyers 1.
Those sum to 57, not 36, because 21 of the ideas tripped two checks and 15
tripped one — so it's a census, not a partition.
You're right that it belongs higher up. Moving it next to the funnel line.
Built a similar filter for our extension ideas — if we can't explain the acute pain in one sentence, we don't build. Saved months on features nobody asked for.
"If we can't explain the acute pain in one sentence, we don't build" is the same
rule with a human running it, and yours is cheaper. The gate here does exactly
that, and it's why most ideas die: 21 of the 36 drops on that run are "we found no
documented problem", which is the automated version of not being able to say the
sentence.
Months saved on features nobody asked for is the only ROI in this category I
actually believe.