
WhittleOS
Decision OS for Solo Founders
I built a thing that sources startup ideas from real markets and kills most of them. Today I opened
the production database and counted everything it has actually done. Top of the list first, because
the graveyard is the boring half.
Highest-scoring idea across every run I've made — 69/100, verdict VALIDATE:
Retainer billing for freelancers. Time-tracking and invoicing built around retainer work — track it,
invoice it, repeat it every month. The first move it handed me: launch a landing page focused on
retainer billing automation with a paid pre-order or waitlist CTA, and test acquisition through async
freelancer communities, SEO pages and niche outbound. No calls, no demos, no "book a slot".
Two more off the same shortlist:
- Checkout Watch (67) — synthetic monitoring that runs a real add-to-cart-through-payment flow on a
Shopify store every few minutes and alerts the owner the moment checkout breaks. Its first move is a
landing page with a paid entry tier, deliberately not a free email signup, so willingness to pay is
what gets tested.
- PDFCraft API (62) — visual PDF template editor with a REST API for developers who need pixel-perfect
documents at scale. Narrow the landing page to one document type first.
That's the shape of the output: a specific buyer, a specific wedge, and something you could start on
tomorrow.
Now the funnel — 23 sweeps, 1,281 candidate ideas:
- 706 killed outright at the deal-breaker gate
- 488 survivors got a full deep evaluation (each run only evaluates its top-ranked survivors)
- 85 made the final shortlist → 21 VALIDATE, 58 HOLD, 6 KILL
Scores across those 488: average 44.9/100, best ever 69, number above 80: zero.
Why the other 706 died. 1,209 cited reasons across 706 kills — most ideas fail more than one:
filter cited only reason
pain_anchor — no documented complaint behind it 438 218
support_burden — high-touch, white-glove, onboarding calls 187 3
platform_risk — one API or one marketplace owns you 140 15
ai_dependency — dies if a model vendor ships the feature 102 9
pricing_power — can't sustain >$19/mo 81 3
founder_fit — the founder can't sell or operate this 78 3
distribution — no repeatable channel 71 0
regulation_trust — needs SOC2/HIPAA/licences for an MVP 48 4
revenue_path — no path to recurring revenue 31 1
enterprise_procurement — needs legal review, MSAs, calls 30 0
Two things I did not expect.
1. The number one killer isn't competition or economics. It's that nobody was complaining.
pain_anchor means the idea couldn't be tied to a single documented complaint from a real source. It's
in 62% of all kills and it's the sole reason 218 times — two hundred and eighteen ideas with nothing
else wrong with them.
2. Some filters never kill alone. distribution (71 citations) and enterprise_procurement (30) were
never once the only reason an idea died; they show up as the second or third strike on something
already bleeding. support_burden is second by volume and solo just 3 times. Meanwhile pain_anchor
kills alone half the time it appears. Executioners and accomplices, and I had them filed the other
way round.
One more thing, since someone will check: nothing has ever scored a BUILD. Not in 1,281 ideas, not in
any mode. That's not the tool sulking — BUILD needs a score of 80+, three of five evidence types
STRONG, and confirmed distribution, and confirmed distribution is something you learn by running the
test, not by analysing an idea at a desk. So VALIDATE is the ceiling at this stage by construction.
"Go test this, here's the page to put up" is the strongest honest thing desk research can hand you.
Anything that tells you BUILD before you've touched the market is selling you a feeling.
What I can't claim, before anyone asks:
- These are my own runs, from my own accounts. Not a customer sample. n=23.
- I excluded 68 stub runs — a $0 synthetic mode I use to smoke-test production. They kill nothing and
would have flattered every rate here.
- pain_anchor partly measures my sourcing, not the market. If my search layer didn't surface the
complaint, the idea dies as unanchored even if the pain is real. I can't separate "no one is
complaining" from "I didn't find it", and that gap is doing some of the work in that 438.
- The ratings sample at temperature 1. Same niche, same input, five runs: 5 / 0 / 4 / 5 / 5 finalists.
The arithmetic on top is fixed; the inputs to it are not.
- None of these ideas were built, so nothing here says the kills were correct. There is no ground truth
on an unbuilt idea. What I can show is the reasoning — one full run is published at
https://whittleos.com/sample-discovery
If you've got a backlog you keep not-building, I'd bet pain_anchor is your number one too.
I built a thing that goes out, sources startup ideas from real markets, and kills most of them.
Today I opened the production database and counted what it has actually done, across every real
run that exists.
The funnel — 23 sweeps, 1,281 candidate ideas:
- 706 killed outright at the deal-breaker gate
- 488 of the survivors got a full deep evaluation (each run only evaluates its top-ranked survivors)
- 85 made the final shortlist
The scores on those 488 evaluations: average 44.9/100. Best score ever recorded: 69. Number that
scored 80 or above: zero.
The verdicts on those 85 finalists: 58 HOLD, 21 VALIDATE, 6 KILL. BUILD is an available verdict.
In 1,281 ideas it has never been used once.
Why ideas died. 706 kills, 1,209 cited reasons — most ideas fail more than one check:
filter cited only reason
pain_anchor — no documented complaint behind it 438 218
support_burden — high-touch, white-glove, onboarding calls 187 3
platform_risk — one API or one marketplace owns you 140 15
ai_dependency — dies if a model vendor ships the feature 102 9
pricing_power — can't sustain >$19/mo 81 3
founder_fit — the founder can't sell or operate this 78 3
distribution — no repeatable channel 71 0
regulation_trust — needs SOC2/HIPAA/licences for an MVP 48 4
revenue_path — no path to recurring revenue 31 1
enterprise_procurement — needs legal review, MSAs, calls 30 0
Two things I did not expect.
1. The number one killer isn't competition, or economics, or "someone already built it". It's that
nobody was complaining. pain_anchor means the idea couldn't be tied to a single documented complaint
from a real source. It appears in 62% of all kills, and it is the sole reason 218 times. Two hundred
and eighteen ideas had nothing else wrong with them.
2. Some filters never kill alone. distribution (71 citations) and enterprise_procurement (30) were
never once the only reason an idea died — they turn up as the second or third strike on something
already bleeding. support_burden is second by volume and solo just 3 times. Meanwhile pain_anchor
kills alone half the time it appears. There are executioners and there are accomplices, and I had
them mentally filed the other way round.
What I can't claim, before anyone asks:
- These are my own runs, from two of my own accounts. It is not a customer sample. n=23.
- I excluded 68 stub runs — a $0 synthetic mode I use to smoke-test production. They kill nothing
and would have flattered every rate on this page.
- pain_anchor partly measures my sourcing, not the market. If my search layer didn't surface the
complaint, the idea dies as unanchored even if the pain is real. I can't separate "no one is
complaining" from "I didn't find it", and that gap is doing some of the work in that 438.
- The ratings sample at temperature 1. Same niche, same input, five runs: 5 / 0 / 4 / 5 / 5
finalists. The arithmetic on top is fixed; the inputs to it are not.
- None of these ideas were built, so nothing here says the kills were correct. There is no ground
truth on an unbuilt idea. What I can show is the reasoning, which is why one full run is published
at https://whittleos.com/sample-discovery
Curious what the distribution looks like for other people's idea lists — if you've got a backlog you
keep not-building, I'd bet on pain_anchor being your number one too.
1 Like
Comment
Every validator I tried is a feel-good machine. Paste an idea, get a high score, a paragraph
about why it's promising, and permission to spend three months. At one point I pasted this
very product into one of them and it loved it. That told me what the score was worth.
So I built the opposite. The honest way to show it is my own demo run, published unedited:
https://whittleos.com/sample-discovery
One Discovery run, start to finish:
- 20 sub-markets planned → 60 searches → 93 results → 24 pages read in full
- 69 candidate ideas in, 36 killed by the deal-breaker checks, 5 finalists out
- 56 documented problems from 41 sources. 32 carry a link you can open; the other 24 are
labelled as the tool's own estimate instead of being dressed up as evidence
- Best finalist: 52/100, grade C, verdict "test it first" — ship a narrow landing page
before writing any code
The 52 is the point. That's my showcase run, sitting on my own marketing pages, and I left it
at 52 because the alternative was inventing an 88. A tool that never says no isn't protecting
you from anything.
Mechanically: the 0-100 score is computed in code from a fixed weighted rubric rather than
asked of a model. Validation plans are async by design — a landing page and a real CTA, never
"go book 20 discovery calls". There are 17 analysis modes; Discovery, where it sources the
candidates itself instead of scoring one you bring, is the flagship.
What I want from you: open the run and tell me where it's wrong. And the question I actually
can't answer myself — is a kill-biased tool useful, or does everyone secretly want the 88?
Prepaid credits, no subscription. Signup gives 2 free credits, enough for the single-idea
check; Discovery runs cost credits because each one burns real search and model spend.
1 Like
Comment
About
I kept generating startup ideas and had no honest way to choose between them. Every validator I tried was a feel-good machine: paste an idea, get a high score and a paragraph about why it's promising. At one point I past

Comment