We swept 30 markets and kept 93 of 3430 candidates. I want to write down what the 2.7% actually taught me, because it was not what I expected going in.
What I expected: a wider sweep would mostly produce more of the same, and the real work would be judging whatever was left. What happened: the sweep is the cheap part, and the judging is where the whole thing either works or quietly flatters you.
The result I would not have believed beforehand: the markets that returned nothing were the system working. Some pools were genuinely thin and the honest output was few finalists, or none. A generator that always returns something is not being generous, it is being useless.
The number I now distrust inside my own product is the score. The arithmetic is fixed and reproducible, but the ratings underneath it sample, so an idea can move a few points run to run. It orders candidates; it does not measure them. I had to go back and delete copy that implied otherwise.
The part that compounds: every problem a run cites is stored with the page it was read on. The published run carries 56 documented problems, 32 of which have a link you can open; the rest are labelled as our estimate rather than quietly promoted. That labelling is the only reason the first number means anything.
The run is public if you want to check any of it, and the write-up is here: https://whittleos.com/guides/startup-ideas
Really solid approach — I'm juggling something similar myself (building Xstream4K on the side), what's been the hardest part for you so far?
The distinction between “orders candidates” and “measures a market” is useful. In my own workflow I’ve started separating Google Ads estimates, traffic estimates, and verified WHOIS facts instead of collapsing them into one confidence score. Have you considered publishing a machine-readable evidence bundle so another tool can audit why a candidate survived?
Not today, and the honest version is that it's further off than it sounds. What exists: every problem a run cites is stored with the page it was read on, and the published run marks which ones have an openable link versus which are our estimate. The export is markdown — readable by a person, not auditable by a tool. And the 3,430 sweep is a citation inside a write-up, not a dataset you can query; I can't hand you a file another system could check.
Your separation is the right shape though. We already tag provenance in tiers — fetched page, pasted by the user, our own inference — specifically so a label can't over-claim, and that tier is published alongside each item. Turning it into a stable bundle format is real work rather than a dump, so I'd rather not announce it before it exists.
“Empty” being a valid outcome is probably one of the hardest things to build into a system. Most products feel pressure to always return something, even when the data doesn’t support it. The score vs. measurement distinction is another really important detail.
The pressure is real, and it has a price: a thin niche can honestly return zero finalists after someone paid for the run. I keep it anyway, because a system that always finds something is the same system as one that finds nothing — you just can't tell which you're holding.
I run a similar funnel on the investing side, and the number nobody tracks is the false negative rate. Your 2.7% pass rate is easy to measure. What it cost you to kill the other 3,337 is not, and that is the number that tells you whether the filter is actually calibrated. Do you keep the rejects around long enough to check whether any of them got built by someone else and worked?
You have named the number I do not have, and the answer is worse than "we don't track it". I went and checked before replying instead of telling you what I assumed was there.
For that sweep specifically — the 30 markets, the 3,430 — the rejects are gone. There is not one stored run from that date left in my database. So the 3,337 cannot be re-read, re-scored or followed up, and whatever the filter got wrong that day is unrecoverable. I have just written that into the file that owns those numbers, because until this afternoon I would have told you the data was there.
What is kept is a different and smaller set: 31 real runs across three months, 1,662 candidates, 866 of them rejected with the reason attached to each. That part I can query. The most common reason is "no documented complaint found" — cited 487 times and the sole reason in 227 of them. Which is the uncomfortable shape of your question: about a quarter of everything I kill is killed for something my own sourcing failed to find, not for something about the idea.
The one time I read the dropped set instead of the survivors, it paid. Fifteen candidates had been killed solely for depending on a platform — Shopify apps, Upwork tooling, Discord monetization. That is a real category rather than noise, and for a founder who will not get on calls it is often the only distribution that reaches a buyer at all. That rule now flags and demotes instead of killing. One confirmed false-negative class, found by reading rejections, and invisible in every possible reading of the finalists.
Your actual question — did any of them get built by someone else and work — I cannot answer, and I have no loop that could. The nearest thing I built is a way for a reader to point at one specific rejection and say it is wrong, tagged by which stage got it wrong. It has zero rows.
One complication I would rather state than hide behind: the gate is conditional on the founder, so "someone else built it and it worked" is not automatically a false negative — it may have been correctly wrong for the person who ran it. That makes the measurement harder. It does not make it optional, and it is not why I do not have it. I do not have it because I kept the survivors and not the rejects.
The labeling discipline (56 cited, only 32 linked, rest flagged as estimate) is the real product here, most tools would have labeled all 56 as sources, and not worried about the difference.
Thanks — though I should be honest about why it exists: the label is nearly free to produce. Provenance comes from the fetch path itself, so a problem read on a page we actually opened gets tagged that way automatically, and everything else defaults to estimate. The expensive part is not generating it, it is publishing it, because it makes the headline number smaller — 32 reads worse than 56 until you ask what the 56 was.
The limit worth naming: "linked" means the page was fetched and the link opens, not that the problem is important or that I verified the reading. It is a provenance claim, not a quality one. That distinction is the next one I would like to not blur.
The 'score orders but doesn't measure' line is the one most tools refuse to admit — good on you for deleting the copy that implied otherwise. What actually makes this write-up trustworthy is the 32/56 sourced-problems labelling: that's the difference between a filter you can audit and a filter you have to trust. One suggestion: make the empty-market outputs a first-class result ('this market returned nothing, and here's why that's informative'). A generator that tells you where NOT to look is rare — and it's exactly the 'built to say no' positioning WhittleOS sells.
The empty-market output being first-class is the right push. Right now a zero-finalist run explains itself on the page but does not claim the ground it covered, which is the part that would make it informative rather than just honest.
'Claims the ground it covered' is the exact phrasing — a zero-result run that lists what was checked is a map; one that just says 'no ideas' is an apology. That distinction is probably worth wiring into the output itself, not just the write-up.
That framing is doing real work, so let me be precise about the gap. A sweep already records everything the map would need: which categories were planned, how many queries ran, how many candidates each one returned, and the reason each candidate was dropped. None of it reaches the page on a zero-finalist run — the user gets the conclusion and not the ground. So this is a rendering problem on state we already hold, not new measurement, which makes it one of the few improvements I can size honestly.
The thing I'd want to avoid is a map that flatters itself. "Checked 31 categories" means nothing if two of them were thin; the honest version has to show where the funnel actually collapsed, including when the answer is "our sourcing found little here" rather than "this market is empty." Those are different failures and the page should not blur them.
'A map that flatters itself' is the standard most dashboards quietly fail. One test that keeps it honest: can a reader walk the record backwards — from finalists to queries to the categories that produced nothing — and watch the funnel narrow at each step? If the record only makes sense in the direction that produced your conclusion, it's decoration.
Fair test, and mine fails it at one joint. Worth naming exactly where.
What does walk backwards today: the shortlist back to each idea's score by area and to the pages each cited problem was read on; the funnel rungs back through every count, including the ones that don't reconcile and why; the rejected candidates back to the specific check that killed each, with a tally of which checks did the killing. And as of yesterday, the sub-markets back to what each one actually returned — what it kept, how much of it we read in full, and which ones added nothing, with the line under it saying that means our search came back thin there, not that the market is empty. That last part came straight out of this thread. It turned out to be a rendering job on state we were already holding and discarding, which is why it took a day rather than a quarter.
Where it breaks: you cannot get from a finalist back to the sub-market that produced it. Hits carry the sub-market that surfaced them, but candidates are written from a batch of mixed sources, so the attribution is lost at the moment ideas are generated. That leaves two chains — finalist to its cited pages, and sub-market to its hits — with no join between them. Smaller second gap: the record says how many queries ran, not what they were.
Rebuilding that join means threading the sub-market through generation and then leaning on a model's own source references for the last hop — days of work for an attribution that would still be approximate. I'd rather say the chain has a break in it than print a lineage I can't stand behind. But it is a break in one named place rather than everywhere, and that is the difference between an incomplete record and decoration.
'I'd rather say the chain has a break than print a lineage I can't stand behind' — that sentence is the whole standard, and it's rarer than it should be. One pragmatic note: the break doesn't have to be rebuilt to be useful — you can label it. Recording 'sub-market attribution unavailable past this point: generated from a mixed batch' turns a missing link into an honest edge, and tells the reader exactly where inference begins. Days of work saved, nothing overstated.
That's the better move and I'd missed it. A labelled edge and a rebuilt join are not the same purchase: one costs a sentence and tells the reader precisely where inference starts, the other costs days and buys an attribution I'd still have to caveat. I was weighing the expensive option against doing nothing, which is how "days of work" turns into "not yet" indefinitely.
The wording has to name the mechanism rather than just the absence, or it reads as an error — something like "sub-market attribution ends here: ideas are generated from a mixed batch of sources", so a reader can tell it apart from a failed lookup. Cheap enough that I'd rather ship it than describe it.
'Name the mechanism, not the absence' is the right rule — it generalizes to every honest gap in a data product: each one is a sentence away from looking like a bug. Ship it. Good talking this through — the frame survived contact with real constraints, which is more than most frameworks manage.
Ship it is right, and I'll take the generalisation with me: every honest gap is one sentence away from looking like a bug, and the sentence costs less than the suspicion does. The version I'd add is that the cost of leaving it unlabelled isn't confusion — it's that someone finds the gap themselves later and has to decide whether you knew. That's a much worse conversation than the one we just had.
Thanks for pushing on it. You moved something from my "someday" pile to my "this week" pile by pointing out I'd priced the wrong option.
Glad it moved piles. And your version is the sharper one — the cost of an unlabelled gap isn't confusion, it's retroactive doubt, and doubt is the one thing a data product can't buy back. 'The sentence costs less than the suspicion' is going in my notebook. Good luck shipping it.
Storing every reject to answer that is probably not worth it, but you do not need every reject. A random sample of thirty from each sweep, kept with the reason it was killed, is enough to estimate the false negative rate within something useful, and it costs almost nothing next to the run itself. Then you re-judge those thirty by hand a month later and count how many you would now keep. The sampling matters more than the volume here, which is the same point you made about your own ratings.
This reframes the thing I said I couldn't do, and the reframe is the useful part: I was answering "did any of them get built and work", which needs years and other people's outcomes. You're proposing a question I can answer alone in a month with a sample. Those aren't the same question, and I'd been letting the impossible one excuse me from the affordable one.
One change I'd make before running it, or the number flatters me: re-judging my own rejects a month later isn't blind. The kill reason is sitting right there, and I'd be marking my own homework with the answer key attached. Cheap fix — strip the reasons and shuffle in a handful of candidates that survived, then judge the pile without knowing which side anything landed on. If I keep the survivors at the same rate I keep the rejects, the filter isn't the thing that's calibrated, my judgement is the thing that's noisy.
And the honest bound: thirty tells me whether the false-negative rate is roughly a quarter or roughly nothing, not whether it's ten percent or fifteen — which is fine, because the decision it feeds is binary. What it can't tell me is whether my later self is any better than the filter. It measures agreement, not correctness. That makes it a calibration instrument rather than ground truth, and it's still far more than I have now, which is nothing.
The score orders candidates; it does not measure them' is a useful boundary. I would trust the result more if the near-misses were visible too, especially when a few sampled ratings can move an idea across the cutoff. Could you show the top rejected candidates beside the finalists, with the rejection reason and evidence level, so a founder can challenge the filter rather than only accept the output?
Half of this exists and the other half I can't build honestly, so let me split it.
What's there: a section listing every dropped candidate with the specific check that killed it, a tally of which checks did the killing across the whole run, and — where exactly one check failed — the test that would overturn it, labelled as per-check rather than per-idea so it isn't read as bespoke advice. Each rejected idea also carries a control for telling the run its call was wrong, tagged by the stage you think got it wrong. So challenging the filter is wired; what's missing is your ordering.
Why "top rejected" is the part I can't do: a candidate that gets killed is instructed to carry a merit score of zero, because the model that killed it is not the one I'd trust to rank it afterwards. So there is no honest "these three nearly made it" — producing one means inventing the number, which is the exact move the post is about.
But your instinct points at a real set I'm hiding. Candidates that passed every check and then lost on a cap do have honest scores, and they're the true near-misses — the page counts them and says nothing was wrong with them, and then never names one. That's the list worth showing, and it's also where your sampling worry actually bites, because that cut ranks on sampled scores while the kill decisions are a different mechanism. Evidence level isn't on the rejected cards either. Both are gaps rather than refusals.
The zero-score point is fair - a killed candidate has no honest rank. But you said it yourself: the ones that passed every check and lost on the cap are the true near-misses, and they do have honest scores. That filtered list seems worth showing on its own, separate from the rejected pile. "Passed everything, cut by cap" is the list a founder would actually want to challenge first. Would you show it as its own view?
Yes, and on reflection it's the more useful of the two lists. The rejected pile answers "why not", which is interesting; "passed everything, then lost a cap" answers "what did I not get to see", which is actionable — nothing is wrong with those ideas, they were simply below a line drawn for cost reasons. I checked what it would take: the data is already on the result page, so this is a view over state I hold rather than anything new to compute.
Two things it has to say on its face or it becomes the flattering number I complained about. First, the order in that list is exactly the sampled ranking you're worried about — it moves between runs, so the same market could hand you a different fourth-place idea tomorrow. Anyone challenging that list is challenging a ranking, not a verdict, and those deserve different language than a kill does. Second, it only exists when the cap actually bit; most runs don't fill it, and on those the honest view is empty rather than padded with whatever came last.
The reason it isn't already there is duller than a principle: the funnel reports those as a count and a reassuring sentence, and nobody asked what the count was made of until you did.
The same trap bit me ranking ~160 launch posts by views: the ranking was reproducible, so it felt true, but the top of the list was one-time counted loads that never moved again. What fixed it was a second pass a day later - anything that does not move gets marked unranked instead of promoted. Same move as your sourced-vs-estimate split, one layer up. Do you re-judge old candidates when the rubric changes, or does a published run stay frozen?
Frozen, deliberately. A published run is a record of a decision someone made on a date, and re-scoring it later edits the thing they acted on. The one direction it can move is down: a run can be retracted after the fact, never upgraded — asymmetric on purpose, because the failure mode I actually fear is a number quietly improving with age.
That got tested for real. We found an arithmetic bug where an exact half rounded down, and the honest way to size it was to recompute a past run rather than guess: about 2% of reachable scores move by a point and no letter grade changes anywhere. We fixed the arithmetic going forward and rewrote nothing that had been published. If it had changed grades, I think the answer is a dated note on the old run, not a silent re-score.
Your "unranked, not promoted" has a direct analogue here and it's the part I'm most glad exists. When our scoring stage can't produce a breakdown we can compute a total from, the idea is left off the shortlist instead of ranked on the model's own number — and the page says that was our failure, not a verdict on the idea. The temptation is exactly the one you describe: something that looks like a rank is available, and using it is one line of code.
The gap you'd catch if you looked: a frozen run doesn't show the reader which version of the rubric produced it. The run records the prompt versions internally; the page doesn't print them. That's a display gap, and it's the thing that makes "frozen" honest rather than convenient.
This resonates with something I ran into building AI-driven tools solo:
the instinct is always to make the system return more, when the real
signal is often in what it correctly refuses to return.
The score-as-ranking-not-measurement distinction is the one I wish
more builders admitted out loud. I've had the same experience shipping
products with LLM-generated scoring — the ordering is defensible, the
absolute number is not, and it's tempting to let the UI imply
otherwise because a precise-looking number reads as more credible.
Curious how you decided on the "estimate" label threshold — was that
a judgment call per-problem, or is there a rule for when a citation
graduates from estimate to sourced?
It's a rule, not a per-problem judgment call — but my rule is coarser than yours, and I'd rather say where.
The line is drawn at the fetch, not at re-derivation. A problem is labelled as coming from a page when we actually retrieved that page; a weaker label covers a site that blocked our fetcher and left us with the search snippet; anything else is our inference. A problem with no source at all is rejected rather than labelled. It's mechanical in the sense that matters — code refuses the strong label when there's no fetch behind it, and a check in our eval fails a run that claims one we have no record of, so it can't be talked into a label by a confident model.
By your test, though, most of my 32 are estimates. A number was read off a page by a model; I have no script that reproduces it from the citation alone, and re-running it is a re-reading, not a re-derivation. So what my split honestly means is "we opened this page" versus "we didn't" — which is worth something, and is not what "sourced" implies to someone reading it quickly.
The version of your rule I could actually build is a third tier under the current one, and it would be much smaller than 32. Worth knowing that before I claim it.
The cleanest rule I've seen work for this: sourced means a machine could refetch that same URL right now and rederive the same number without a human in the loop. Estimate is everything else, even with a link attached, if getting from that page to the number needed interpretation or judgment, it doesn't graduate. That turns it from a per-problem judgment call into a testable property: can you write a script that reproduces this number from the citation alone? If yes, sourced. If it needs someone to read it and decide, it's an estimate no matter how confident you are in it.
I run a hiring funnel for a living (we screen software engineers, and most applicants do not make it through), and your reject data maps closely onto a problem we had.
The split I would make is between "evidence against" and "no evidence found." Your top reject reason, no documented complaint found, is the second kind. That is not the filter judging the idea, it is the filter reporting on its own sourcing. We had the same thing with engineers rejected because we could not find proof of a skill, not because we found proof they lacked it. Once those went into their own "not proven" bucket and got one more look through a different source, that bucket turned out to be where our misses tend to hide. Kills based on evidence against mostly stay killed.
Two cheap checks that do not need the "did someone else build it" loop:
Seed canaries. Put a handful of ideas you already know are good (ones with real, public revenue) into a sweep with nothing flagged. If the filter kills any of them, you learn what it gets wrong today instead of waiting years for the market to tell you.
Score twice. Since the ratings sample, run the finalists and the near misses through scoring two or three times and look at the spread. We started doing this with interview scoring, and the swing matters most for exactly the candidates sitting near the cutoff, which is where a single run hurts you. A wide spread is worth showing as its own label instead of hiding inside one number.
The zero-row feedback form does not surprise me. People rarely volunteer that a rejection was wrong. One specific question on one specific reject ("have you actually seen this complaint?") usually gets more answers than an open invitation.
The "evidence against" vs "no evidence found" split is exactly right, and I can put a number on it: across the real discovery runs in our database, the single most common kill reason is a missing problem anchor — roughly six in ten drops. That is my filter reporting on my own sourcing, filed as a verdict about the idea. I hadn't separated the two buckets and I should.
On your two checks:
Seed canaries exist for the single-idea gate — ten businesses with real public revenue and ten documented failures, run as a repeatable calibration. The winners got killed about 5% of the time, the failures about 90%. What I have not done is point that harness at the sweep's triage stage, which is the one that produced the 3,337, so the canary result I quote doesn't actually cover the filter you're asking about.
Score twice is where I'm worst, and I'll say it plainly. The single-idea path samples each kill filter several times and requires a majority. The sweep does not — one sample per stage. Same niche, same input, five runs: five, zero, four, five, five finalists. The zero is the whole argument for your point, and the reason I haven't fixed it is cost, not disagreement — it multiplies the priciest stage. Reporting spread as its own label is the cheaper half of your suggestion and I hadn't considered doing it without the voting.
And the targeted question instead of the open form is obviously correct. "Have you actually seen this complaint?" on one specific reject asks someone to check one thing; my current version asks them to build a case.
Yeah, the realisation that the sweep is the easy part and the judging is where it can quietly flatter you feels pretty true. Easy to get excited by volume and forget the filtering is the real work.
I liked that you kept the provenance honest too — linked sources versus estimates. That kind of care shows.
Hope the product keeps getting clearer for you, mate. Looking forward to seeing how it evolves.
Thanks — the provenance labelling is the part that pays off slowly and then all at once, because it's the only thing that lets a number survive someone checking it. Appreciate you reading it.
That realization about empty market results is huge! A scoring system that forces a result every time is just lying to you. Real value comes from strict filtering and honest labeling of missing data, even when it means coming up empty.
The labeling half is the part that took the most work. It's easy to filter hard; it's harder to make "we don't know" a first-class output instead of a gap the model quietly fills. Concretely: if there's no comparable product we can point at, the field says unknown rather than inventing a plausible price, and a cited problem carries the page it was read on or it's marked as our estimate. Once missing data is allowed to be missing, coming up empty stops looking like a failure and starts looking like the same rule applied consistently.
Your answer to aryan_sinh is the most useful part of this thread, and I think it points at a call you can make now, before there's any evidence to look at.
If that nine-seed customer is any indication, expect mostly usage logs and occasionally a story. The first time someone says "your tool made me drop an idea," it will feel like the missing proof. Whether it is depends on a definition, and I don't see one written down in this thread yet. Did they drop a finalist because of a cited problem they actually opened, or were they already leaning that way and the run agreed with them? After the fact, both read as a win.
The one habit from my own experiments I'd actually defend is aimed at exactly that. Before each test I write down a fixed window and the rule for what counts as a pass or a fail, before looking at anything. It's the only thing that has stopped me from grading my own homework generously.
It's the same move you already made with the 24 estimates: labelling them is what lets the 32 links mean something. Applied to users, it's a line like "a finalist was rejected or reshaped, the reason given traces to a cited problem, within N days of the run."
So the decision I'd make now: write that line and the window before the next outside sweep, and treat anything that doesn't meet it as a usage log, however nice the message. If nothing meets it after a handful of runs, that's your answer to "product or toy", reached honestly instead of by anecdote.
Taking the pre-registration. Here is the line, written before the next sweep: a finalist is rejected or materially reshaped, and the stated reason traces to a problem the run cited, within 14 days of the run. Anything else is a usage log, including a nice message. If nothing meets it after ten outside runs, that is the answer.
That line will do the work. The one thing I'd add, from getting this wrong myself: write down what you expect the answer to be, with a number, before the ten runs start.
Not because the guess matters. Because when the tenth run comes back ambiguous — and it will, the first ones usually are — you will be standing there with a result and a decision to make, and the only thing that stops you from reading it generously is a sentence you wrote when you had nothing invested in it.
I run small A/B tests on instruction files and I keep a sealed prediction for each one. I have been wrong on the last three in a row, all in the same direction: I expected the intervention to do more than it did. That pattern is only visible because the numbers were written down first. If I had judged after the fact, I would have found a reason each time.
One practical note on the 14 days. The clock probably needs to start at the run, not at the message, or a user who comes back on day 20 to tell you something useful gets scored as a miss for being slow rather than for the reason you care about. Your call, but worth fixing now while it costs nothing.
The clock is already on the run in that line, but you found a real gap next to it: it does not say which date is measured when the report is late. Pinning it now — the 14 days apply to the date of the decision, not the date I hear about it. A day-40 message about a day-9 drop counts, and a day-9 message about a decision someone made in June does not.
Sealed number, before the ten runs: I expect zero. Not modesty, arithmetic. I went and looked at the database before writing this, and there was no surface anywhere in the product that asked the question. Nothing would have carried that answer to me if it had happened.
Which turned up something worse, and you should have it, since it came out of your suggestion. Outside Discovery runs, total, ever: ten. All from one person, all in one afternoon last week, and the 14-day window on them is still open until the 22nd. So the denominator I named in public was already spent at the moment I named it, and I could not have known, because every inbound channel I have was sitting at zero rows.
I am not scoring those ten as misses — scoring them would be scoring my own silence — and I am saying so now rather than when the number turns out to be inconvenient.
So I built the asking instead of starting the count. Two emails, day 7 and day 14, plus the same question on the run page for anyone who comes back on their own. The part I would defend to you is one field: if you dropped or reshaped a finalist, you pick which of the problems that run actually cited your reason traces back to, and "none of these" sits in the same list, equally easy to click. I could have collected the reason as prose and decided afterwards whether it counted. That is grading my own homework, and it is the version of this where the number only ever goes up. "Still deciding" is stored and counted too, for the same reason.
And the prediction you actually helped me write down is the second one, about myself. My bias runs opposite to yours: you have been wrong three times expecting the intervention to do more than it did. Having predicted zero, the way I get this wrong is talking myself into counting a warm message as a hit, because a hit would be a relief.
The 2.7% is interesting, but the bigger distinction seems to be between ranking candidates and actually helping someone decide. Have you seen users choose, reject, or materially reshape an idea because of the documented evidence attached to a finalist?
No, and I should say that plainly rather than reach for an anecdote. You have put your finger on the gap.
What I can show is the filter working: candidates dropped, with the reason attached. What I cannot show is anyone changing their mind because of it. Very few people outside my own accounts have run a full sweep. The one paying customer who did ran a nine-seed batch, then never came back and never told me what he concluded — so I have a usage log and no decision.
The distinction you are drawing is the one that decides whether this is a product or a toy, and I do not have the evidence yet. If you ever run one and drop something because of what it cited, I would genuinely like to hear it.
That gap between filtering and an actual decision is the important one. If you’re open to it, what’s the best email to reach you on?
Happy to keep it here — the thread is the better record, and the next person chasing the same gap can read it. If you have run a sweep and dropped or reshaped something because of a problem it cited, that is the exact case I am missing, and I would rather have it in public than in my inbox.
That 2.7% cut is a useful reminder that a long list is not a market. In Speechara.Ai we try to validate with a first useful session, then a second one, instead of relying on signup counts. What evidence made those 93 candidates stay?
Nothing made them stay — nothing killed them, which is a different claim and the honest one. The gate is subtractive: every candidate has to survive a fixed set of checks, any one of which ends it, and what is left is what no check could end.
The most common killer is the one your question implies. Across the real runs I have measured, 706 candidates were dropped; the single biggest reason, 438 of them, was that no documented problem could be tied to the idea at all. Support load killed 187, platform dependence 140.
On the run I publish, the surviving ideas carry 56 documented problems, 32 of which have a link you can open. The other 24 are labelled as our estimate rather than quietly promoted to evidence, and that labelling is the only reason the first number means anything.
Your first-useful-session test measures something mine cannot: whether the thing gets used twice. Signup counts and my survival rate are both upstream of that.