I built a thing that sources startup ideas from real markets and kills most of them. Today I opened
the production database and counted everything it has actually done. Top of the list first, because
the graveyard is the boring half.
Highest-scoring idea across every run I've made — 69/100, verdict VALIDATE:
Retainer billing for freelancers. Time-tracking and invoicing built around retainer work — track it,
invoice it, repeat it every month. The first move it handed me: launch a landing page focused on
retainer billing automation with a paid pre-order or waitlist CTA, and test acquisition through async
freelancer communities, SEO pages and niche outbound. No calls, no demos, no "book a slot".
Two more off the same shortlist:
- Checkout Watch (67) — synthetic monitoring that runs a real add-to-cart-through-payment flow on a
Shopify store every few minutes and alerts the owner the moment checkout breaks. Its first move is a
landing page with a paid entry tier, deliberately not a free email signup, so willingness to pay is
what gets tested.
- PDFCraft API (62) — visual PDF template editor with a REST API for developers who need pixel-perfect
documents at scale. Narrow the landing page to one document type first.
That's the shape of the output: a specific buyer, a specific wedge, and something you could start on
tomorrow.
Now the funnel — 23 sweeps, 1,281 candidate ideas:
- 706 killed outright at the deal-breaker gate
- 488 survivors got a full deep evaluation (each run only evaluates its top-ranked survivors)
- 85 made the final shortlist → 21 VALIDATE, 58 HOLD, 6 KILL
Scores across those 488: average 44.9/100, best ever 69, number above 80: zero.
Why the other 706 died. 1,209 cited reasons across 706 kills — most ideas fail more than one:
filter cited only reason
pain_anchor — no documented complaint behind it 438 218
support_burden — high-touch, white-glove, onboarding calls 187 3
platform_risk — one API or one marketplace owns you 140 15
ai_dependency — dies if a model vendor ships the feature 102 9
pricing_power — can't sustain >$19/mo 81 3
founder_fit — the founder can't sell or operate this 78 3
distribution — no repeatable channel 71 0
regulation_trust — needs SOC2/HIPAA/licences for an MVP 48 4
revenue_path — no path to recurring revenue 31 1
enterprise_procurement — needs legal review, MSAs, calls 30 0
Two things I did not expect.
1. The number one killer isn't competition or economics. It's that nobody was complaining.
pain_anchor means the idea couldn't be tied to a single documented complaint from a real source. It's
in 62% of all kills and it's the sole reason 218 times — two hundred and eighteen ideas with nothing
else wrong with them.
2. Some filters never kill alone. distribution (71 citations) and enterprise_procurement (30) were
never once the only reason an idea died; they show up as the second or third strike on something
already bleeding. support_burden is second by volume and solo just 3 times. Meanwhile pain_anchor
kills alone half the time it appears. Executioners and accomplices, and I had them filed the other
way round.
One more thing, since someone will check: nothing has ever scored a BUILD. Not in 1,281 ideas, not in
any mode. That's not the tool sulking — BUILD needs a score of 80+, three of five evidence types
STRONG, and confirmed distribution, and confirmed distribution is something you learn by running the
test, not by analysing an idea at a desk. So VALIDATE is the ceiling at this stage by construction.
"Go test this, here's the page to put up" is the strongest honest thing desk research can hand you.
Anything that tells you BUILD before you've touched the market is selling you a feeling.
What I can't claim, before anyone asks:
- These are my own runs, from my own accounts. Not a customer sample. n=23.
- I excluded 68 stub runs — a $0 synthetic mode I use to smoke-test production. They kill nothing and
would have flattered every rate here.
- pain_anchor partly measures my sourcing, not the market. If my search layer didn't surface the
complaint, the idea dies as unanchored even if the pain is real. I can't separate "no one is
complaining" from "I didn't find it", and that gap is doing some of the work in that 438.
- The ratings sample at temperature 1. Same niche, same input, five runs: 5 / 0 / 4 / 5 / 5 finalists.
The arithmetic on top is fixed; the inputs to it are not.
- None of these ideas were built, so nothing here says the kills were correct. There is no ground truth
on an unbuilt idea. What I can show is the reasoning — one full run is published at
https://whittleos.com/sample-discovery
If you've got a backlog you keep not-building, I'd bet pain_anchor is your number one too.