8
15 Comments

We swept 30 markets and kept 93 of 3430 candidates. Here is what the 2.7% actually taught me.

We swept 30 markets and kept 93 of 3430 candidates. I want to write down what the 2.7% actually taught me, because it was not what I expected going in.

What I expected: a wider sweep would mostly produce more of the same, and the real work would be judging whatever was left. What happened: the sweep is the cheap part, and the judging is where the whole thing either works or quietly flatters you.

The result I would not have believed beforehand: the markets that returned nothing were the system working. Some pools were genuinely thin and the honest output was few finalists, or none. A generator that always returns something is not being generous, it is being useless.

The number I now distrust inside my own product is the score. The arithmetic is fixed and reproducible, but the ratings underneath it sample, so an idea can move a few points run to run. It orders candidates; it does not measure them. I had to go back and delete copy that implied otherwise.

The part that compounds: every problem a run cites is stored with the page it was read on. The published run carries 56 documented problems, 32 of which have a link you can open; the rest are labelled as our estimate rather than quietly promoted. That labelling is the only reason the first number means anything.

The run is public if you want to check any of it, and the write-up is here: https://whittleos.com/guides/startup-ideas

on September 15, 2026
  1. 1

    The labeling discipline (56 cited, only 32 linked, rest flagged as estimate) is the real product here, most tools would have labeled all 56 as sources, and not worried about the difference.

    1. 1

      Thanks — though I should be honest about why it exists: the label is nearly free to produce. Provenance comes from the fetch path itself, so a problem read on a page we actually opened gets tagged that way automatically, and everything else defaults to estimate. The expensive part is not generating it, it is publishing it, because it makes the headline number smaller — 32 reads worse than 56 until you ask what the 56 was.

      The limit worth naming: "linked" means the page was fetched and the link opens, not that the problem is important or that I verified the reading. It is a provenance claim, not a quality one. That distinction is the next one I would like to not blur.

  2. 1

    The 'score orders but doesn't measure' line is the one most tools refuse to admit — good on you for deleting the copy that implied otherwise. What actually makes this write-up trustworthy is the 32/56 sourced-problems labelling: that's the difference between a filter you can audit and a filter you have to trust. One suggestion: make the empty-market outputs a first-class result ('this market returned nothing, and here's why that's informative'). A generator that tells you where NOT to look is rare — and it's exactly the 'built to say no' positioning WhittleOS sells.

    1. 1

      The empty-market output being first-class is the right push. Right now a zero-finalist run explains itself on the page but does not claim the ground it covered, which is the part that would make it informative rather than just honest.

  3. 1

    What stands out to me here is that you’re being very careful about separating “survived the filter” from “proven to be a good business.” That distinction is easy to blur, especially when a system produces a clean score or ranking that looks more objective than it really is.

    The part about markets returning nothing being a success is probably the strongest insight in the post. A system that always forces a result can feel productive, but it is often just hiding uncertainty. Being able to say “there is not enough evidence here” is much more useful than manufacturing a shortlist.

    I also like that you went back and changed the copy around the score. If the underlying ratings can move slightly between runs, then presenting the output as an ordering rather than a precise measurement is much more honest. That kind of calibration matters a lot when users are making decisions from the result.

    The next interesting step, like the comments above are getting at, is connecting this filtering system to real user decisions. Not just whether an idea survives, but whether someone actually rejects, changes, or prioritizes something because of the cited evidence. If you can capture that consistently, the product becomes much easier to evaluate.

    The documented-problem layer also feels important. A score by itself is easy to distrust, but a score plus evidence you can open and inspect is much more useful. Even the fact that you label estimates separately instead of presenting everything as sourced evidence is a strong design choice.

    Overall, I think the 2.7% number is interesting, but the more important result is that the system seems willing to return “nothing” when the evidence is weak. That is probably a better sign than a long list of polished ideas.

    1. 1

      The empty-market output being first-class is the right push. Right now a zero-finalist run explains itself on the page but does not claim the ground it covered, which is the part that would make it informative rather than just honest.

  4. 1

    Your answer to aryan_sinh is the most useful part of this thread, and I think it points at a call you can make now, before there's any evidence to look at.

    If that nine-seed customer is any indication, expect mostly usage logs and occasionally a story. The first time someone says "your tool made me drop an idea," it will feel like the missing proof. Whether it is depends on a definition, and I don't see one written down in this thread yet. Did they drop a finalist because of a cited problem they actually opened, or were they already leaning that way and the run agreed with them? After the fact, both read as a win.

    The one habit from my own experiments I'd actually defend is aimed at exactly that. Before each test I write down a fixed window and the rule for what counts as a pass or a fail, before looking at anything. It's the only thing that has stopped me from grading my own homework generously.

    It's the same move you already made with the 24 estimates: labelling them is what lets the 32 links mean something. Applied to users, it's a line like "a finalist was rejected or reshaped, the reason given traces to a cited problem, within N days of the run."

    So the decision I'd make now: write that line and the window before the next outside sweep, and treat anything that doesn't meet it as a usage log, however nice the message. If nothing meets it after a handful of runs, that's your answer to "product or toy", reached honestly instead of by anecdote.

    1. 1

      Taking the pre-registration. Here is the line, written before the next sweep: a finalist is rejected or materially reshaped, and the stated reason traces to a problem the run cited, within 14 days of the run. Anything else is a usage log, including a nice message. If nothing meets it after ten outside runs, that is the answer.

      1. 1

        That line will do the work. The one thing I'd add, from getting this wrong myself: write down what you expect the answer to be, with a number, before the ten runs start.

        Not because the guess matters. Because when the tenth run comes back ambiguous — and it will, the first ones usually are — you will be standing there with a result and a decision to make, and the only thing that stops you from reading it generously is a sentence you wrote when you had nothing invested in it.

        I run small A/B tests on instruction files and I keep a sealed prediction for each one. I have been wrong on the last three in a row, all in the same direction: I expected the intervention to do more than it did. That pattern is only visible because the numbers were written down first. If I had judged after the fact, I would have found a reason each time.

        One practical note on the 14 days. The clock probably needs to start at the run, not at the message, or a user who comes back on day 20 to tell you something useful gets scored as a miss for being slow rather than for the reason you care about. Your call, but worth fixing now while it costs nothing.

  5. 1

    The 2.7% is interesting, but the bigger distinction seems to be between ranking candidates and actually helping someone decide. Have you seen users choose, reject, or materially reshape an idea because of the documented evidence attached to a finalist?

    1. 1

      No, and I should say that plainly rather than reach for an anecdote. You have put your finger on the gap.

      What I can show is the filter working: candidates dropped, with the reason attached. What I cannot show is anyone changing their mind because of it. Very few people outside my own accounts have run a full sweep. The one paying customer who did ran a nine-seed batch, then never came back and never told me what he concluded — so I have a usage log and no decision.

      The distinction you are drawing is the one that decides whether this is a product or a toy, and I do not have the evidence yet. If you ever run one and drop something because of what it cited, I would genuinely like to hear it.

      1. 1

        That gap between filtering and an actual decision is the important one. If you’re open to it, what’s the best email to reach you on?

        1. 1

          Happy to keep it here — the thread is the better record, and the next person chasing the same gap can read it. If you have run a sweep and dropped or reshaped something because of a problem it cited, that is the exact case I am missing, and I would rather have it in public than in my inbox.

  6. 1

    That 2.7% cut is a useful reminder that a long list is not a market. In Speechara.Ai we try to validate with a first useful session, then a second one, instead of relying on signup counts. What evidence made those 93 candidates stay?

    1. 1

      Nothing made them stay — nothing killed them, which is a different claim and the honest one. The gate is subtractive: every candidate has to survive a fixed set of checks, any one of which ends it, and what is left is what no check could end.

      The most common killer is the one your question implies. Across the real runs I have measured, 706 candidates were dropped; the single biggest reason, 438 of them, was that no documented problem could be tied to the idea at all. Support load killed 187, platform dependence 140.

      On the run I publish, the surviving ideas carry 56 documented problems, 32 of which have a link you can open. The other 24 are labelled as our estimate rather than quietly promoted to evidence, and that labelling is the only reason the first number means anything.

      Your first-useful-session test measures something mine cannot: whether the thing gets used twice. Signup counts and my survival rate are both upstream of that.