2
3 Comments

I ran 23 idea-discovery sweeps through my own tool. It has never once said "build it".

I built a thing that goes out, sources startup ideas from real markets, and kills most of them.

Today I opened the production database and counted what it has actually done, across every real

run that exists.

The funnel — 23 sweeps, 1,281 candidate ideas:

- 706 killed outright at the deal-breaker gate

- 488 of the survivors got a full deep evaluation (each run only evaluates its top-ranked survivors)

- 85 made the final shortlist

The scores on those 488 evaluations: average 44.9/100. Best score ever recorded: 69. Number that

scored 80 or above: zero.

The verdicts on those 85 finalists: 58 HOLD, 21 VALIDATE, 6 KILL. BUILD is an available verdict.

In 1,281 ideas it has never been used once.

Why ideas died. 706 kills, 1,209 cited reasons — most ideas fail more than one check:

filter cited only reason

pain_anchor — no documented complaint behind it 438 218

support_burden — high-touch, white-glove, onboarding calls 187 3

platform_risk — one API or one marketplace owns you 140 15

ai_dependency — dies if a model vendor ships the feature 102 9

pricing_power — can't sustain >$19/mo 81 3

founder_fit — the founder can't sell or operate this 78 3

distribution — no repeatable channel 71 0

regulation_trust — needs SOC2/HIPAA/licences for an MVP 48 4

revenue_path — no path to recurring revenue 31 1

enterprise_procurement — needs legal review, MSAs, calls 30 0

Two things I did not expect.

1. The number one killer isn't competition, or economics, or "someone already built it". It's that

nobody was complaining. pain_anchor means the idea couldn't be tied to a single documented complaint

from a real source. It appears in 62% of all kills, and it is the sole reason 218 times. Two hundred

and eighteen ideas had nothing else wrong with them.

2. Some filters never kill alone. distribution (71 citations) and enterprise_procurement (30) were

never once the only reason an idea died — they turn up as the second or third strike on something

already bleeding. support_burden is second by volume and solo just 3 times. Meanwhile pain_anchor

kills alone half the time it appears. There are executioners and there are accomplices, and I had

them mentally filed the other way round.

What I can't claim, before anyone asks:

- These are my own runs, from two of my own accounts. It is not a customer sample. n=23.

- I excluded 68 stub runs — a $0 synthetic mode I use to smoke-test production. They kill nothing

and would have flattered every rate on this page.

- pain_anchor partly measures my sourcing, not the market. If my search layer didn't surface the

complaint, the idea dies as unanchored even if the pain is real. I can't separate "no one is

complaining" from "I didn't find it", and that gap is doing some of the work in that 438.

- The ratings sample at temperature 1. Same niche, same input, five runs: 5 / 0 / 4 / 5 / 5

finalists. The arithmetic on top is fixed; the inputs to it are not.

- None of these ideas were built, so nothing here says the kills were correct. There is no ground

truth on an unbuilt idea. What I can show is the reasoning, which is why one full run is published

at https://whittleos.com/sample-discovery

Curious what the distribution looks like for other people's idea lists — if you've got a backlog you

keep not-building, I'd bet on pain_anchor being your number one too.

posted toAvatar for product WhittleOS
WhittleOS
  1. 1
    The pain_anchor result is probably the most interesting part to me, but I wonder if it could also reject some good opportunities too aggressively. People don’t always complain explicitly about a problem. Sometimes search demand, people paying for ugly existing solutions, spreadsheets/manual workflows, or competitors already making money are stronger signals than complaints. I’d be curious what happens if “documented pain” becomes one of several alternative demand signals rather than a hard gate.
    1. 1
      You've found the real one, and it's harder than you framed it. It is a hard gate, not a weight. The triage instruction says to drop any candidate with no STRONG-or-MEDIUM documented pain behind it before the other ten filters run. WEAK pains don't count. If the pain box comes back empty, everything dies on that filter. The part I can add: we already collect the signals you're describing. The sweep runs query groups for pain, revenue and demand — named-MRR comparables and real job listings, not just complaint threads. They feed candidate generation and the later scoring. They just don't reach the gate. So what you're proposing is mostly wiring something already sourced into the decision, which is a much cheaper change than it sounds. Why it's hard in the first place: the earlier version was toothless — it demoted 0 candidates out of 108. Making it a real gate dropped survival from about 60% to 39%. That's the number I'd be trading away, so I'm not loosening it on an argument, including a good one. I went and measured it before replying, because the honest answer needed a number. Of 706 killed candidates across 23 real runs, 218 died on pain_anchor alone — no other filter named, nothing else wrong with them. That's 31% of every rejection, and it is the exact set your argument is about: an alternative demand signal could only have changed the outcome for those. The second number cuts the other way. The gate has a branch that kills everything when there is no pain evidence at all, and it has never fired — median 42 documented pains per run, minimum 1, zero runs with an empty box. So the 218 didn't die because we looked and found nothing. They died because they matched none of the 42-odd pains we did have in their own niche, which is a fair description of a solution in search of a problem. Where I'm stuck, precisely: I can't yet tell you how many of those 218 had a demand or revenue signal sitting unused in the same run. The report stores what TYPE each source was — review site, forum, competitor product — but not which query group it came from, so the question isn't answerable from what's already saved. That's a gap in my instrumentation, not an argument against you, and it's the next thing I'll fix. So: 218 is your target, 31% of kills, and it's a real number rather than a hypothetical.
  2. 0
    You've found the real one, and it's harder than you framed it. It is a hard gate, not a weight. The triage instruction says to drop any candidate with no STRONG-or-MEDIUM documented pain behind it before the other ten filters run. WEAK pains don't count. If the pain box comes back empty, everything dies on that filter. The part I can add: we already collect the signals you're describing. The sweep runs query groups for pain, revenue and demand — named-MRR comparables and real job listings, not just complaint threads. They feed candidate generation and the later scoring. They just don't reach the gate. So what you're proposing is mostly wiring something already sourced into the decision, which is a much cheaper change than it sounds. Why it's hard in the first place: the earlier version was toothless — it demoted 0 candidates out of 108. Making it a real gate dropped survival from about 60% to 39%. That's the number I'd be trading away, so I'm not loosening it on an argument, including a good one. I went and measured it before replying, because the honest answer needed a number. Of 706 killed candidates across 23 real runs, 218 died on pain_anchor alone — no other filter named, nothing else wrong with them. That's 31% of every rejection, and it is the exact set your argument is about: an alternative demand signal could only have changed the outcome for those. The second number cuts the other way. The gate has a branch that kills everything when there is no pain evidence at all, and it has never fired — median 42 documented pains per run, minimum 1, zero runs with an empty box. So the 218 didn't die because we looked and found nothing. They died because they matched none of the 42-odd pains we did have in their own niche, which is a fair description of a solution in search of a problem. Where I'm stuck, precisely: I can't yet tell you how many of those 218 had a demand or revenue signal sitting unused in the same run. The report stores what TYPE each source was — review site, forum, competitor product — but not which query group it came from, so the question isn't answerable from what's already saved. That's a gap in my instrumentation, not an argument against you, and it's the next thing I'll fix. So: 218 is your target, 31% of kills, and it's a real number rather than a hypothetical.