10
20 Comments

We swept 30 markets and kept 93 of 3430 candidates. Here is what the 2.7% actually taught me.

We swept 30 markets and kept 93 of 3430 candidates. I want to write down what the 2.7% actually taught me, because it was not what I expected going in.

What I expected: a wider sweep would mostly produce more of the same, and the real work would be judging whatever was left. What happened: the sweep is the cheap part, and the judging is where the whole thing either works or quietly flatters you.

The result I would not have believed beforehand: the markets that returned nothing were the system working. Some pools were genuinely thin and the honest output was few finalists, or none. A generator that always returns something is not being generous, it is being useless.

The number I now distrust inside my own product is the score. The arithmetic is fixed and reproducible, but the ratings underneath it sample, so an idea can move a few points run to run. It orders candidates; it does not measure them. I had to go back and delete copy that implied otherwise.

The part that compounds: every problem a run cites is stored with the page it was read on. The published run carries 56 documented problems, 32 of which have a link you can open; the rest are labelled as our estimate rather than quietly promoted. That labelling is the only reason the first number means anything.

The run is public if you want to check any of it, and the write-up is here: https://whittleos.com/guides/startup-ideas

on September 15, 2026
  1. 1

    The distinction between “orders candidates” and “measures a market” is useful. In my own workflow I’ve started separating Google Ads estimates, traffic estimates, and verified WHOIS facts instead of collapsing them into one confidence score. Have you considered publishing a machine-readable evidence bundle so another tool can audit why a candidate survived?

  2. 1

    “Empty” being a valid outcome is probably one of the hardest things to build into a system. Most products feel pressure to always return something, even when the data doesn’t support it. The score vs. measurement distinction is another really important detail.

  3. 1

    I run a similar funnel on the investing side, and the number nobody tracks is the false negative rate. Your 2.7% pass rate is easy to measure. What it cost you to kill the other 3,337 is not, and that is the number that tells you whether the filter is actually calibrated. Do you keep the rejects around long enough to check whether any of them got built by someone else and worked?

    1. 1

      You have named the number I do not have, and the answer is worse than "we don't track it". I went and checked before replying instead of telling you what I assumed was there.

      For that sweep specifically — the 30 markets, the 3,430 — the rejects are gone. There is not one stored run from that date left in my database. So the 3,337 cannot be re-read, re-scored or followed up, and whatever the filter got wrong that day is unrecoverable. I have just written that into the file that owns those numbers, because until this afternoon I would have told you the data was there.

      What is kept is a different and smaller set: 31 real runs across three months, 1,662 candidates, 866 of them rejected with the reason attached to each. That part I can query. The most common reason is "no documented complaint found" — cited 487 times and the sole reason in 227 of them. Which is the uncomfortable shape of your question: about a quarter of everything I kill is killed for something my own sourcing failed to find, not for something about the idea.

      The one time I read the dropped set instead of the survivors, it paid. Fifteen candidates had been killed solely for depending on a platform — Shopify apps, Upwork tooling, Discord monetization. That is a real category rather than noise, and for a founder who will not get on calls it is often the only distribution that reaches a buyer at all. That rule now flags and demotes instead of killing. One confirmed false-negative class, found by reading rejections, and invisible in every possible reading of the finalists.

      Your actual question — did any of them get built by someone else and work — I cannot answer, and I have no loop that could. The nearest thing I built is a way for a reader to point at one specific rejection and say it is wrong, tagged by which stage got it wrong. It has zero rows.

      One complication I would rather state than hide behind: the gate is conditional on the founder, so "someone else built it and it worked" is not automatically a false negative — it may have been correctly wrong for the person who ran it. That makes the measurement harder. It does not make it optional, and it is not why I do not have it. I do not have it because I kept the survivors and not the rejects.

  4. 1

    The labeling discipline (56 cited, only 32 linked, rest flagged as estimate) is the real product here, most tools would have labeled all 56 as sources, and not worried about the difference.

    1. 1

      Thanks — though I should be honest about why it exists: the label is nearly free to produce. Provenance comes from the fetch path itself, so a problem read on a page we actually opened gets tagged that way automatically, and everything else defaults to estimate. The expensive part is not generating it, it is publishing it, because it makes the headline number smaller — 32 reads worse than 56 until you ask what the 56 was.

      The limit worth naming: "linked" means the page was fetched and the link opens, not that the problem is important or that I verified the reading. It is a provenance claim, not a quality one. That distinction is the next one I would like to not blur.

  5. 1

    The 'score orders but doesn't measure' line is the one most tools refuse to admit — good on you for deleting the copy that implied otherwise. What actually makes this write-up trustworthy is the 32/56 sourced-problems labelling: that's the difference between a filter you can audit and a filter you have to trust. One suggestion: make the empty-market outputs a first-class result ('this market returned nothing, and here's why that's informative'). A generator that tells you where NOT to look is rare — and it's exactly the 'built to say no' positioning WhittleOS sells.

    1. 1

      The empty-market output being first-class is the right push. Right now a zero-finalist run explains itself on the page but does not claim the ground it covered, which is the part that would make it informative rather than just honest.

  6. 1

    What stands out to me here is that you’re being very careful about separating “survived the filter” from “proven to be a good business.” That distinction is easy to blur, especially when a system produces a clean score or ranking that looks more objective than it really is.

    The part about markets returning nothing being a success is probably the strongest insight in the post. A system that always forces a result can feel productive, but it is often just hiding uncertainty. Being able to say “there is not enough evidence here” is much more useful than manufacturing a shortlist.

    I also like that you went back and changed the copy around the score. If the underlying ratings can move slightly between runs, then presenting the output as an ordering rather than a precise measurement is much more honest. That kind of calibration matters a lot when users are making decisions from the result.

    The next interesting step, like the comments above are getting at, is connecting this filtering system to real user decisions. Not just whether an idea survives, but whether someone actually rejects, changes, or prioritizes something because of the cited evidence. If you can capture that consistently, the product becomes much easier to evaluate.

    The documented-problem layer also feels important. A score by itself is easy to distrust, but a score plus evidence you can open and inspect is much more useful. Even the fact that you label estimates separately instead of presenting everything as sourced evidence is a strong design choice.

    Overall, I think the 2.7% number is interesting, but the more important result is that the system seems willing to return “nothing” when the evidence is weak. That is probably a better sign than a long list of polished ideas.

    1. 1

      The empty-market output being first-class is the right push. Right now a zero-finalist run explains itself on the page but does not claim the ground it covered, which is the part that would make it informative rather than just honest.

  7. 1

    Your answer to aryan_sinh is the most useful part of this thread, and I think it points at a call you can make now, before there's any evidence to look at.

    If that nine-seed customer is any indication, expect mostly usage logs and occasionally a story. The first time someone says "your tool made me drop an idea," it will feel like the missing proof. Whether it is depends on a definition, and I don't see one written down in this thread yet. Did they drop a finalist because of a cited problem they actually opened, or were they already leaning that way and the run agreed with them? After the fact, both read as a win.

    The one habit from my own experiments I'd actually defend is aimed at exactly that. Before each test I write down a fixed window and the rule for what counts as a pass or a fail, before looking at anything. It's the only thing that has stopped me from grading my own homework generously.

    It's the same move you already made with the 24 estimates: labelling them is what lets the 32 links mean something. Applied to users, it's a line like "a finalist was rejected or reshaped, the reason given traces to a cited problem, within N days of the run."

    So the decision I'd make now: write that line and the window before the next outside sweep, and treat anything that doesn't meet it as a usage log, however nice the message. If nothing meets it after a handful of runs, that's your answer to "product or toy", reached honestly instead of by anecdote.

    1. 1

      Taking the pre-registration. Here is the line, written before the next sweep: a finalist is rejected or materially reshaped, and the stated reason traces to a problem the run cited, within 14 days of the run. Anything else is a usage log, including a nice message. If nothing meets it after ten outside runs, that is the answer.

      1. 1

        That line will do the work. The one thing I'd add, from getting this wrong myself: write down what you expect the answer to be, with a number, before the ten runs start.

        Not because the guess matters. Because when the tenth run comes back ambiguous — and it will, the first ones usually are — you will be standing there with a result and a decision to make, and the only thing that stops you from reading it generously is a sentence you wrote when you had nothing invested in it.

        I run small A/B tests on instruction files and I keep a sealed prediction for each one. I have been wrong on the last three in a row, all in the same direction: I expected the intervention to do more than it did. That pattern is only visible because the numbers were written down first. If I had judged after the fact, I would have found a reason each time.

        One practical note on the 14 days. The clock probably needs to start at the run, not at the message, or a user who comes back on day 20 to tell you something useful gets scored as a miss for being slow rather than for the reason you care about. Your call, but worth fixing now while it costs nothing.

        1. 1

          The clock is already on the run in that line, but you found a real gap next to it: it does not say which date is measured when the report is late. Pinning it now — the 14 days apply to the date of the decision, not the date I hear about it. A day-40 message about a day-9 drop counts, and a day-9 message about a decision someone made in June does not.

          Sealed number, before the ten runs: I expect zero. Not modesty, arithmetic. I went and looked at the database before writing this, and there was no surface anywhere in the product that asked the question. Nothing would have carried that answer to me if it had happened.

          Which turned up something worse, and you should have it, since it came out of your suggestion. Outside Discovery runs, total, ever: ten. All from one person, all in one afternoon last week, and the 14-day window on them is still open until the 22nd. So the denominator I named in public was already spent at the moment I named it, and I could not have known, because every inbound channel I have was sitting at zero rows.

          I am not scoring those ten as misses — scoring them would be scoring my own silence — and I am saying so now rather than when the number turns out to be inconvenient.

          So I built the asking instead of starting the count. Two emails, day 7 and day 14, plus the same question on the run page for anyone who comes back on their own. The part I would defend to you is one field: if you dropped or reshaped a finalist, you pick which of the problems that run actually cited your reason traces back to, and "none of these" sits in the same list, equally easy to click. I could have collected the reason as prose and decided afterwards whether it counted. That is grading my own homework, and it is the version of this where the number only ever goes up. "Still deciding" is stored and counted too, for the same reason.

          And the prediction you actually helped me write down is the second one, about myself. My bias runs opposite to yours: you have been wrong three times expecting the intervention to do more than it did. Having predicted zero, the way I get this wrong is talking myself into counting a warm message as a hit, because a hit would be a relief.

  8. 1

    The 2.7% is interesting, but the bigger distinction seems to be between ranking candidates and actually helping someone decide. Have you seen users choose, reject, or materially reshape an idea because of the documented evidence attached to a finalist?

    1. 1

      No, and I should say that plainly rather than reach for an anecdote. You have put your finger on the gap.

      What I can show is the filter working: candidates dropped, with the reason attached. What I cannot show is anyone changing their mind because of it. Very few people outside my own accounts have run a full sweep. The one paying customer who did ran a nine-seed batch, then never came back and never told me what he concluded — so I have a usage log and no decision.

      The distinction you are drawing is the one that decides whether this is a product or a toy, and I do not have the evidence yet. If you ever run one and drop something because of what it cited, I would genuinely like to hear it.

      1. 1

        That gap between filtering and an actual decision is the important one. If you’re open to it, what’s the best email to reach you on?

        1. 1

          Happy to keep it here — the thread is the better record, and the next person chasing the same gap can read it. If you have run a sweep and dropped or reshaped something because of a problem it cited, that is the exact case I am missing, and I would rather have it in public than in my inbox.

  9. 1

    That 2.7% cut is a useful reminder that a long list is not a market. In Speechara.Ai we try to validate with a first useful session, then a second one, instead of relying on signup counts. What evidence made those 93 candidates stay?

    1. 1

      Nothing made them stay — nothing killed them, which is a different claim and the honest one. The gate is subtractive: every candidate has to survive a fixed set of checks, any one of which ends it, and what is left is what no check could end.

      The most common killer is the one your question implies. Across the real runs I have measured, 706 candidates were dropped; the single biggest reason, 438 of them, was that no documented problem could be tied to the idea at all. Support load killed 187, platform dependence 140.

      On the run I publish, the surviving ideas carry 56 documented problems, 32 of which have a link you can open. The other 24 are labelled as our estimate rather than quietly promoted to evidence, and that labelling is the only reason the first number means anything.

      Your first-useful-session test measures something mine cannot: whether the thing gets used twice. Signup counts and my survival rate are both upstream of that.