17
36 Comments

We swept 30 markets and kept 93 of 3430 candidates. Here is what the 2.7% actually taught me.

We swept 30 markets and kept 93 of 3430 candidates. I want to write down what the 2.7% actually taught me, because it was not what I expected going in.

What I expected: a wider sweep would mostly produce more of the same, and the real work would be judging whatever was left. What happened: the sweep is the cheap part, and the judging is where the whole thing either works or quietly flatters you.

The result I would not have believed beforehand: the markets that returned nothing were the system working. Some pools were genuinely thin and the honest output was few finalists, or none. A generator that always returns something is not being generous, it is being useless.

The number I now distrust inside my own product is the score. The arithmetic is fixed and reproducible, but the ratings underneath it sample, so an idea can move a few points run to run. It orders candidates; it does not measure them. I had to go back and delete copy that implied otherwise.

The part that compounds: every problem a run cites is stored with the page it was read on. The published run carries 56 documented problems, 32 of which have a link you can open; the rest are labelled as our estimate rather than quietly promoted. That labelling is the only reason the first number means anything.

The run is public if you want to check any of it, and the write-up is here: https://whittleos.com/guides/startup-ideas

on September 15, 2026
  1. 2

    The distinction between “orders candidates” and “measures a market” is useful. In my own workflow I’ve started separating Google Ads estimates, traffic estimates, and verified WHOIS facts instead of collapsing them into one confidence score. Have you considered publishing a machine-readable evidence bundle so another tool can audit why a candidate survived?

    1. 1

      Not today, and the honest version is that it's further off than it sounds. What exists: every problem a run cites is stored with the page it was read on, and the published run marks which ones have an openable link versus which are our estimate. The export is markdown — readable by a person, not auditable by a tool. And the 3,430 sweep is a citation inside a write-up, not a dataset you can query; I can't hand you a file another system could check.

      Your separation is the right shape though. We already tag provenance in tiers — fetched page, pasted by the user, our own inference — specifically so a label can't over-claim, and that tier is published alongside each item. Turning it into a stable bundle format is real work rather than a dump, so I'd rather not announce it before it exists.

  2. 2

    “Empty” being a valid outcome is probably one of the hardest things to build into a system. Most products feel pressure to always return something, even when the data doesn’t support it. The score vs. measurement distinction is another really important detail.

    1. 1

      The pressure is real, and it has a price: a thin niche can honestly return zero finalists after someone paid for the run. I keep it anyway, because a system that always finds something is the same system as one that finds nothing — you just can't tell which you're holding.

  3. 2

    I run a similar funnel on the investing side, and the number nobody tracks is the false negative rate. Your 2.7% pass rate is easy to measure. What it cost you to kill the other 3,337 is not, and that is the number that tells you whether the filter is actually calibrated. Do you keep the rejects around long enough to check whether any of them got built by someone else and worked?

    1. 1

      You have named the number I do not have, and the answer is worse than "we don't track it". I went and checked before replying instead of telling you what I assumed was there.

      For that sweep specifically — the 30 markets, the 3,430 — the rejects are gone. There is not one stored run from that date left in my database. So the 3,337 cannot be re-read, re-scored or followed up, and whatever the filter got wrong that day is unrecoverable. I have just written that into the file that owns those numbers, because until this afternoon I would have told you the data was there.

      What is kept is a different and smaller set: 31 real runs across three months, 1,662 candidates, 866 of them rejected with the reason attached to each. That part I can query. The most common reason is "no documented complaint found" — cited 487 times and the sole reason in 227 of them. Which is the uncomfortable shape of your question: about a quarter of everything I kill is killed for something my own sourcing failed to find, not for something about the idea.

      The one time I read the dropped set instead of the survivors, it paid. Fifteen candidates had been killed solely for depending on a platform — Shopify apps, Upwork tooling, Discord monetization. That is a real category rather than noise, and for a founder who will not get on calls it is often the only distribution that reaches a buyer at all. That rule now flags and demotes instead of killing. One confirmed false-negative class, found by reading rejections, and invisible in every possible reading of the finalists.

      Your actual question — did any of them get built by someone else and work — I cannot answer, and I have no loop that could. The nearest thing I built is a way for a reader to point at one specific rejection and say it is wrong, tagged by which stage got it wrong. It has zero rows.

      One complication I would rather state than hide behind: the gate is conditional on the founder, so "someone else built it and it worked" is not automatically a false negative — it may have been correctly wrong for the person who ran it. That makes the measurement harder. It does not make it optional, and it is not why I do not have it. I do not have it because I kept the survivors and not the rejects.

  4. 2

    The labeling discipline (56 cited, only 32 linked, rest flagged as estimate) is the real product here, most tools would have labeled all 56 as sources, and not worried about the difference.

    1. 1

      Thanks — though I should be honest about why it exists: the label is nearly free to produce. Provenance comes from the fetch path itself, so a problem read on a page we actually opened gets tagged that way automatically, and everything else defaults to estimate. The expensive part is not generating it, it is publishing it, because it makes the headline number smaller — 32 reads worse than 56 until you ask what the 56 was.

      The limit worth naming: "linked" means the page was fetched and the link opens, not that the problem is important or that I verified the reading. It is a provenance claim, not a quality one. That distinction is the next one I would like to not blur.

  5. 2

    The 'score orders but doesn't measure' line is the one most tools refuse to admit — good on you for deleting the copy that implied otherwise. What actually makes this write-up trustworthy is the 32/56 sourced-problems labelling: that's the difference between a filter you can audit and a filter you have to trust. One suggestion: make the empty-market outputs a first-class result ('this market returned nothing, and here's why that's informative'). A generator that tells you where NOT to look is rare — and it's exactly the 'built to say no' positioning WhittleOS sells.

    1. 1

      The empty-market output being first-class is the right push. Right now a zero-finalist run explains itself on the page but does not claim the ground it covered, which is the part that would make it informative rather than just honest.

      1. 1

        'Claims the ground it covered' is the exact phrasing — a zero-result run that lists what was checked is a map; one that just says 'no ideas' is an apology. That distinction is probably worth wiring into the output itself, not just the write-up.

        1. 1

          That framing is doing real work, so let me be precise about the gap. A sweep already records everything the map would need: which categories were planned, how many queries ran, how many candidates each one returned, and the reason each candidate was dropped. None of it reaches the page on a zero-finalist run — the user gets the conclusion and not the ground. So this is a rendering problem on state we already hold, not new measurement, which makes it one of the few improvements I can size honestly.

          The thing I'd want to avoid is a map that flatters itself. "Checked 31 categories" means nothing if two of them were thin; the honest version has to show where the funnel actually collapsed, including when the answer is "our sourcing found little here" rather than "this market is empty." Those are different failures and the page should not blur them.

          1. 1

            'A map that flatters itself' is the standard most dashboards quietly fail. One test that keeps it honest: can a reader walk the record backwards — from finalists to queries to the categories that produced nothing — and watch the funnel narrow at each step? If the record only makes sense in the direction that produced your conclusion, it's decoration.

            1. 1

              Fair test, and mine fails it at one joint. Worth naming exactly where.

              What does walk backwards today: the shortlist back to each idea's score by area and to the pages each cited problem was read on; the funnel rungs back through every count, including the ones that don't reconcile and why; the rejected candidates back to the specific check that killed each, with a tally of which checks did the killing. And as of yesterday, the sub-markets back to what each one actually returned — what it kept, how much of it we read in full, and which ones added nothing, with the line under it saying that means our search came back thin there, not that the market is empty. That last part came straight out of this thread. It turned out to be a rendering job on state we were already holding and discarding, which is why it took a day rather than a quarter.

              Where it breaks: you cannot get from a finalist back to the sub-market that produced it. Hits carry the sub-market that surfaced them, but candidates are written from a batch of mixed sources, so the attribution is lost at the moment ideas are generated. That leaves two chains — finalist to its cited pages, and sub-market to its hits — with no join between them. Smaller second gap: the record says how many queries ran, not what they were.

              Rebuilding that join means threading the sub-market through generation and then leaning on a model's own source references for the last hop — days of work for an attribution that would still be approximate. I'd rather say the chain has a break in it than print a lineage I can't stand behind. But it is a break in one named place rather than everywhere, and that is the difference between an incomplete record and decoration.

  6. 1

    The same trap bit me ranking ~160 launch posts by views: the ranking was reproducible, so it felt true, but the top of the list was one-time counted loads that never moved again. What fixed it was a second pass a day later - anything that does not move gets marked unranked instead of promoted. Same move as your sourced-vs-estimate split, one layer up. Do you re-judge old candidates when the rubric changes, or does a published run stay frozen?

    1. 1

      Frozen, deliberately. A published run is a record of a decision someone made on a date, and re-scoring it later edits the thing they acted on. The one direction it can move is down: a run can be retracted after the fact, never upgraded — asymmetric on purpose, because the failure mode I actually fear is a number quietly improving with age.

      That got tested for real. We found an arithmetic bug where an exact half rounded down, and the honest way to size it was to recompute a past run rather than guess: about 2% of reachable scores move by a point and no letter grade changes anywhere. We fixed the arithmetic going forward and rewrote nothing that had been published. If it had changed grades, I think the answer is a dated note on the old run, not a silent re-score.

      Your "unranked, not promoted" has a direct analogue here and it's the part I'm most glad exists. When our scoring stage can't produce a breakdown we can compute a total from, the idea is left off the shortlist instead of ranked on the model's own number — and the page says that was our failure, not a verdict on the idea. The temptation is exactly the one you describe: something that looks like a rank is available, and using it is one line of code.

      The gap you'd catch if you looked: a frozen run doesn't show the reader which version of the rubric produced it. The run records the prompt versions internally; the page doesn't print them. That's a display gap, and it's the thing that makes "frozen" honest rather than convenient.

  7. 1

    This resonates with something I ran into building AI-driven tools solo:
    the instinct is always to make the system return more, when the real
    signal is often in what it correctly refuses to return.

    The score-as-ranking-not-measurement distinction is the one I wish
    more builders admitted out loud. I've had the same experience shipping
    products with LLM-generated scoring — the ordering is defensible, the
    absolute number is not, and it's tempting to let the UI imply
    otherwise because a precise-looking number reads as more credible.

    Curious how you decided on the "estimate" label threshold — was that
    a judgment call per-problem, or is there a rule for when a citation
    graduates from estimate to sourced?

    1. 1

      It's a rule, not a per-problem judgment call — but my rule is coarser than yours, and I'd rather say where.

      The line is drawn at the fetch, not at re-derivation. A problem is labelled as coming from a page when we actually retrieved that page; a weaker label covers a site that blocked our fetcher and left us with the search snippet; anything else is our inference. A problem with no source at all is rejected rather than labelled. It's mechanical in the sense that matters — code refuses the strong label when there's no fetch behind it, and a check in our eval fails a run that claims one we have no record of, so it can't be talked into a label by a confident model.

      By your test, though, most of my 32 are estimates. A number was read off a page by a model; I have no script that reproduces it from the citation alone, and re-running it is a re-reading, not a re-derivation. So what my split honestly means is "we opened this page" versus "we didn't" — which is worth something, and is not what "sourced" implies to someone reading it quickly.

      The version of your rule I could actually build is a third tier under the current one, and it would be much smaller than 32. Worth knowing that before I claim it.

    2. 1

      The cleanest rule I've seen work for this: sourced means a machine could refetch that same URL right now and rederive the same number without a human in the loop. Estimate is everything else, even with a link attached, if getting from that page to the number needed interpretation or judgment, it doesn't graduate. That turns it from a per-problem judgment call into a testable property: can you write a script that reproduces this number from the citation alone? If yes, sourced. If it needs someone to read it and decide, it's an estimate no matter how confident you are in it.

  8. 1

    I run a hiring funnel for a living (we screen software engineers, and most applicants do not make it through), and your reject data maps closely onto a problem we had.

    The split I would make is between "evidence against" and "no evidence found." Your top reject reason, no documented complaint found, is the second kind. That is not the filter judging the idea, it is the filter reporting on its own sourcing. We had the same thing with engineers rejected because we could not find proof of a skill, not because we found proof they lacked it. Once those went into their own "not proven" bucket and got one more look through a different source, that bucket turned out to be where our misses tend to hide. Kills based on evidence against mostly stay killed.

    Two cheap checks that do not need the "did someone else build it" loop:

    Seed canaries. Put a handful of ideas you already know are good (ones with real, public revenue) into a sweep with nothing flagged. If the filter kills any of them, you learn what it gets wrong today instead of waiting years for the market to tell you.

    Score twice. Since the ratings sample, run the finalists and the near misses through scoring two or three times and look at the spread. We started doing this with interview scoring, and the swing matters most for exactly the candidates sitting near the cutoff, which is where a single run hurts you. A wide spread is worth showing as its own label instead of hiding inside one number.

    The zero-row feedback form does not surprise me. People rarely volunteer that a rejection was wrong. One specific question on one specific reject ("have you actually seen this complaint?") usually gets more answers than an open invitation.

    1. 1

      The "evidence against" vs "no evidence found" split is exactly right, and I can put a number on it: across the real discovery runs in our database, the single most common kill reason is a missing problem anchor — roughly six in ten drops. That is my filter reporting on my own sourcing, filed as a verdict about the idea. I hadn't separated the two buckets and I should.

      On your two checks:

      Seed canaries exist for the single-idea gate — ten businesses with real public revenue and ten documented failures, run as a repeatable calibration. The winners got killed about 5% of the time, the failures about 90%. What I have not done is point that harness at the sweep's triage stage, which is the one that produced the 3,337, so the canary result I quote doesn't actually cover the filter you're asking about.

      Score twice is where I'm worst, and I'll say it plainly. The single-idea path samples each kill filter several times and requires a majority. The sweep does not — one sample per stage. Same niche, same input, five runs: five, zero, four, five, five finalists. The zero is the whole argument for your point, and the reason I haven't fixed it is cost, not disagreement — it multiplies the priciest stage. Reporting spread as its own label is the cheaper half of your suggestion and I hadn't considered doing it without the voting.

      And the targeted question instead of the open form is obviously correct. "Have you actually seen this complaint?" on one specific reject asks someone to check one thing; my current version asks them to build a case.

  9. 1

    Yeah, the realisation that the sweep is the easy part and the judging is where it can quietly flatter you feels pretty true. Easy to get excited by volume and forget the filtering is the real work.

    I liked that you kept the provenance honest too — linked sources versus estimates. That kind of care shows.

    Hope the product keeps getting clearer for you, mate. Looking forward to seeing how it evolves.

    1. 1

      Thanks — the provenance labelling is the part that pays off slowly and then all at once, because it's the only thing that lets a number survive someone checking it. Appreciate you reading it.

  10. 1

    That realization about empty market results is huge! A scoring system that forces a result every time is just lying to you. Real value comes from strict filtering and honest labeling of missing data, even when it means coming up empty.

    1. 1

      The labeling half is the part that took the most work. It's easy to filter hard; it's harder to make "we don't know" a first-class output instead of a gap the model quietly fills. Concretely: if there's no comparable product we can point at, the field says unknown rather than inventing a plausible price, and a cited problem carries the page it was read on or it's marked as our estimate. Once missing data is allowed to be missing, coming up empty stops looking like a failure and starts looking like the same rule applied consistently.

  11. 1

    Your answer to aryan_sinh is the most useful part of this thread, and I think it points at a call you can make now, before there's any evidence to look at.

    If that nine-seed customer is any indication, expect mostly usage logs and occasionally a story. The first time someone says "your tool made me drop an idea," it will feel like the missing proof. Whether it is depends on a definition, and I don't see one written down in this thread yet. Did they drop a finalist because of a cited problem they actually opened, or were they already leaning that way and the run agreed with them? After the fact, both read as a win.

    The one habit from my own experiments I'd actually defend is aimed at exactly that. Before each test I write down a fixed window and the rule for what counts as a pass or a fail, before looking at anything. It's the only thing that has stopped me from grading my own homework generously.

    It's the same move you already made with the 24 estimates: labelling them is what lets the 32 links mean something. Applied to users, it's a line like "a finalist was rejected or reshaped, the reason given traces to a cited problem, within N days of the run."

    So the decision I'd make now: write that line and the window before the next outside sweep, and treat anything that doesn't meet it as a usage log, however nice the message. If nothing meets it after a handful of runs, that's your answer to "product or toy", reached honestly instead of by anecdote.

    1. 1

      Taking the pre-registration. Here is the line, written before the next sweep: a finalist is rejected or materially reshaped, and the stated reason traces to a problem the run cited, within 14 days of the run. Anything else is a usage log, including a nice message. If nothing meets it after ten outside runs, that is the answer.

      1. 1

        That line will do the work. The one thing I'd add, from getting this wrong myself: write down what you expect the answer to be, with a number, before the ten runs start.

        Not because the guess matters. Because when the tenth run comes back ambiguous — and it will, the first ones usually are — you will be standing there with a result and a decision to make, and the only thing that stops you from reading it generously is a sentence you wrote when you had nothing invested in it.

        I run small A/B tests on instruction files and I keep a sealed prediction for each one. I have been wrong on the last three in a row, all in the same direction: I expected the intervention to do more than it did. That pattern is only visible because the numbers were written down first. If I had judged after the fact, I would have found a reason each time.

        One practical note on the 14 days. The clock probably needs to start at the run, not at the message, or a user who comes back on day 20 to tell you something useful gets scored as a miss for being slow rather than for the reason you care about. Your call, but worth fixing now while it costs nothing.

        1. 1

          The clock is already on the run in that line, but you found a real gap next to it: it does not say which date is measured when the report is late. Pinning it now — the 14 days apply to the date of the decision, not the date I hear about it. A day-40 message about a day-9 drop counts, and a day-9 message about a decision someone made in June does not.

          Sealed number, before the ten runs: I expect zero. Not modesty, arithmetic. I went and looked at the database before writing this, and there was no surface anywhere in the product that asked the question. Nothing would have carried that answer to me if it had happened.

          Which turned up something worse, and you should have it, since it came out of your suggestion. Outside Discovery runs, total, ever: ten. All from one person, all in one afternoon last week, and the 14-day window on them is still open until the 22nd. So the denominator I named in public was already spent at the moment I named it, and I could not have known, because every inbound channel I have was sitting at zero rows.

          I am not scoring those ten as misses — scoring them would be scoring my own silence — and I am saying so now rather than when the number turns out to be inconvenient.

          So I built the asking instead of starting the count. Two emails, day 7 and day 14, plus the same question on the run page for anyone who comes back on their own. The part I would defend to you is one field: if you dropped or reshaped a finalist, you pick which of the problems that run actually cited your reason traces back to, and "none of these" sits in the same list, equally easy to click. I could have collected the reason as prose and decided afterwards whether it counted. That is grading my own homework, and it is the version of this where the number only ever goes up. "Still deciding" is stored and counted too, for the same reason.

          And the prediction you actually helped me write down is the second one, about myself. My bias runs opposite to yours: you have been wrong three times expecting the intervention to do more than it did. Having predicted zero, the way I get this wrong is talking myself into counting a warm message as a hit, because a hit would be a relief.

  12. 1

    The 2.7% is interesting, but the bigger distinction seems to be between ranking candidates and actually helping someone decide. Have you seen users choose, reject, or materially reshape an idea because of the documented evidence attached to a finalist?

    1. 1

      No, and I should say that plainly rather than reach for an anecdote. You have put your finger on the gap.

      What I can show is the filter working: candidates dropped, with the reason attached. What I cannot show is anyone changing their mind because of it. Very few people outside my own accounts have run a full sweep. The one paying customer who did ran a nine-seed batch, then never came back and never told me what he concluded — so I have a usage log and no decision.

      The distinction you are drawing is the one that decides whether this is a product or a toy, and I do not have the evidence yet. If you ever run one and drop something because of what it cited, I would genuinely like to hear it.

      1. 1

        That gap between filtering and an actual decision is the important one. If you’re open to it, what’s the best email to reach you on?

        1. 1

          Happy to keep it here — the thread is the better record, and the next person chasing the same gap can read it. If you have run a sweep and dropped or reshaped something because of a problem it cited, that is the exact case I am missing, and I would rather have it in public than in my inbox.

  13. 1

    That 2.7% cut is a useful reminder that a long list is not a market. In Speechara.Ai we try to validate with a first useful session, then a second one, instead of relying on signup counts. What evidence made those 93 candidates stay?

    1. 1

      Nothing made them stay — nothing killed them, which is a different claim and the honest one. The gate is subtractive: every candidate has to survive a fixed set of checks, any one of which ends it, and what is left is what no check could end.

      The most common killer is the one your question implies. Across the real runs I have measured, 706 candidates were dropped; the single biggest reason, 438 of them, was that no documented problem could be tied to the idea at all. Support load killed 187, platform dependence 140.

      On the run I publish, the surviving ideas carry 56 documented problems, 32 of which have a link you can open. The other 24 are labelled as our estimate rather than quietly promoted to evidence, and that labelling is the only reason the first number means anything.

      Your first-useful-session test measures something mine cannot: whether the thing gets used twice. Signup counts and my survival rate are both upstream of that.