11
33 Comments

WhittleOS — an idea tool built to say no

Every idea checker I tried is a feel-good machine. I pasted this very product into one of them and it loved it. That told me exactly what the score was worth.

So I built the opposite. The only honest way to show it is my own demo run, published unedited:

https://whittleos.com/sample-discovery

One run, start to finish: 20 sub-markets → 60 searches → 24 pages read in full → 69 candidate ideas in, 36 killed at the deal-breaker gate, 5 out. It cites 56 documented problems from 41 sources — 32 carry a link you can open, and the other 24 are labelled as the tool's own estimate instead of being dressed up as evidence.

Best finalist: 52/100, "test it first."

The 52 is the point. That is my own showcase run, sitting on my own marketing pages, and I left it at 52 because the alternative was inventing an 88.

No signup to read it. If you open it, tell me where the run is wrong — that is genuinely what I want out of posting this.

on September 6, 2026
  1. 1

    The feel good bias in early stage tooling is rampant. Most products optimize for dopamine hits: green indicators, vanity metrics, and polite feedback that tells founders what they want to hear.

    The real return on investment for an evaluation tool is preventing someone from wasting six months building the wrong thing. A genuine 52 with documented friction points protects capital and time.

    One element that could make this even more practical: attach a cheapest falsification step to the score.

    When an idea gets a 52, the immediate founder question is: what is the single lowest cost experiment that kills this or validates it? If the report ends with: here is the exact assumption carrying the most risk, and here is a twenty minute manual check to test it, the report becomes an immediate action playbook.

  2. 1

    The 52 on your own landing page is the best marketing you could've done. It's not the score that sells me, it's that you didn't fake it. That's trust you can't buy. Most founders would've invented an 88 and called it a day. You chose credibility over ego. That says everything.

  3. 1

    The unedited run is a strong wedge because it gives people something concrete to disagree with. I’d turn the “no signup” choice into a deliberate interview loop: put one prompt under the report—“Which conclusion would you bet against, and what evidence would change your mind?”—and tag each response by step (sub-market, search, page, score). After 5–10 runs, the repeated disagreement tells you whether to improve discovery or the scoring model. Keep 52/100 as the artifact; sell the next run as a falsifiable decision, not a prettier score. That’s a good way to find first users without diluting the anti-hype positioning.

  4. 4

    Your point about single-sample noise causing false kills is the part I'd want to dig into. A tool built to say no also needs a way to catch its own bad rejections.
    Have you tried taking a few rejected ideas and checking the specific deal-breaker with potential customers? “We found no evidence of demand” and “we found evidence there's no demand” would lead me to different next steps.
    I'd find that audit more useful than the headline score: which rejections held up, and which turned out to be gaps in the research?

    1. 1

      The distinction is right and we're on the wrong side of it. Our most common
      rejection is labelled "No real problem behind it" — 21 of the 36 on that run, and
      around 440 of the ~700 across every real run. What the check actually does is
      look for the candidate among the problems that run collected, and kill it if it
      isn't there. That's "we found no evidence", printed as "there is none". Your
      phrasing is better than mine, and the label is changing.

      The audit you're describing I can't run honestly yet: it needs founders who ran
      it, disagreed with a rejection, and went and checked. I don't have that data, and
      I'd rather say so than assemble something shaped like it.

      What I can do meanwhile is the part that doesn't need anyone's permission — all
      36 rejections are published with their stated reason at the bottom of
      whittleos.com/sample-discovery. Telling me which of those look like gaps in the
      research rather than real kills is the most useful thing anyone in this thread
      could do for me.

  5. 3

    The 52 is much more useful than a made-up 88. I’d expose a second axis next to the score: evidence density (number, independence, and recency of sources) versus fit/feasibility, so a low score from thin data doesn’t look identical to a well-researched rejection. For the five finalists, a one-line “what would falsify this next?” test could turn the report into an experiment queue rather than a static ranking.

    1. 1

      Agreed, and it's cheaper than you'd expect — the inputs are already stored per row: whether we read the actual page or only a search snippet, how many independent hosts the evidence spans, when it was last seen, and how many separate sweeps found it again.

      Which I know because I published two articles today where I computed exactly that by hand — "25 records from 19 pages", "every row was seen once" — precisely because the product doesn't show it. Doing it in prose, one page at a time, is not a strategy.

      Your framing also caught something bigger than you aimed at. This morning I measured the corpus and 1,433 of its 1,434 rows have been found by exactly one sweep. The score is severity times recurrence, so with recurrence flat the score was severity — while the surface printed "comes up a lot" next to rows seen once, on 196 of them. Fixed today: the labels now describe severity, every row states its sighting count, and the explainer counts, from the rows in front of you, how many have been found more than once. Your evidence-density axis is the thing that would have made that visible on the page instead of needing a database query to find.

      On "what would falsify this next?" — there is a per-finalist async plan today with a stop-loss and a deadline, but it's phrased as a plan to run rather than a claim to break. Falsification is the better frame and I'm taking it.

  6. 2

    The 52 creates a measurement system where the score itself becomes the data, not the validation.

    Here's why that matters: founders make decisions based on whether they trust the measurement, not the idea itself. A 52 with cited sources and transparent reasoning tells you "this tool is honest." An 88 tells you the tool is selling confidence.

    So a founder sees your own product at 52 and thinks: "If this honest measurement says 52 for something the creator believes in, what does that tell me about my 60?" Nothing helpful. The 88 would have been noise.

    But the 52 — that's information about which measurement systems are worth paying attention to. You're not trying to build a predictor that scores ideas correctly. You're trying to build a measurement filter that separates founders who can trust their own thinking from founders who need someone to tell them what to believe.

    Most "idea validators" are in the business of making everything look viable. You inverted it: you're in the business of proving you're not.

    1. 1

      "A measurement filter rather than a predictor" is a better description of this
      than the one on my own landing page, and I'm going to sit with it.

      What I can't answer yet is what the framing costs. If the product selects for
      founders who can already trust their own thinking, that is a smaller room than the
      one everyone else in this category is selling to — and the founders who want to be
      told what to believe are exactly the ones who convert on an 88. Whether the room is
      big enough is genuinely open, not modestly open.

      One thing I'd add: the filter cuts both ways. The same run that says "52, and here
      is every source" also says what it could not check — the published one links 32 of
      its 56 problems and labels the other 24 as the tool's own estimate. A founder who
      finds that annoying is not going to like anything else it says either.

  7. 1

    The “say no positioning is strong because it separates the tool from feel-good idea graders. To make the verdict more useful for getting traction, Id pair each rejection with one concrete next test—who to interview, what behavior to look for, and a deadline—then track whether founders actually run it. That turns a score into a learning loop and gives you outcome data that is much harder to fake than a higher number.

    1. 1

      Half of this exists and you can't see it, which is my problem, not yours: each finalist gets an async test with a promise, a real call-to-action, a target number and a stop-loss date, and the report opens with which one to run first. There's also an outcome loop as of yesterday — once a validation result is recorded, a report can be retracted downward but never upgraded. That's the un-fakeable direction, and it's there because a tool that can raise its own grade after the fact is worthless.

      Where you're right and I have no defence: the recommended first move is computed from the survivors only. Rejections get reasons and nothing else. On a tool whose whole pitch is that it says no, the "no" is the output with no next step attached — which is exactly backwards, since a rejection is the moment a founder most needs to know what would change the answer.

      One place I'll push back, because it's a position rather than an omission: "who to interview" is something the product refuses to produce. Interview and call suggestions are blocked at the output scanner, not left out by accident — the argument is that a validation step requiring a conversation is the step most people never take. So I'll take the structure of your suggestion — one test, one observable behaviour, a deadline, then track whether it was run — and keep it async.

  8. 1

    The honesty is the right product decision and the wrong sales motion, at least for idea-stage founders. Nobody with a fresh idea pays monthly to be told no, but someone who has already burned six months and real money will pay well to avoid the second mistake, which is why I would point this at accelerators, angel groups, and second-time founders first. Same engine, completely different willingness to pay.

  9. 1

    The distinction between “no evidence” and “no demand” is a valuable guardrail. I’d make the output actionable by pairing each rejection with one falsifiable next test: who to interview, what behavior would change the verdict, and a fixed time or budget cap. Publishing the raw run is also a strong trust signal—over time, anonymized follow-ups could become the product’s most useful dataset.

    1. 1

      Splitting your three, because they land differently for me.

      The falsifiable next test is a real gap, and half of it already exists in the
      wrong place. The single-idea check returns a "what would need to be true" ladder —
      the concrete conditions that would raise the verdict — but a Discovery rejection
      carries only its reason. So the mode that examines one idea tells you what would
      change its mind, and the mode that throws away 36 does not. That asymmetry is
      indefensible and I had not noticed it until you wrote this.

      "Who to interview" is the one I'll push back on, and not out of squeamishness.
      This is built for founders who sell without calls: the validation plans it writes
      are a landing page, a real CTA and async outreach, and there is a scanner on the
      output that rejects "book a call" and its cousins before they reach the page. If
      the next test requires interviewing ten people, it has quietly become a different
      product for a different founder. What changes a verdict is behaviour, and
      behaviour can be observed without a conversation.

      The fixed time or budget cap I agree with completely. It exists today on the one
      idea the run recommends, and nowhere else — including on all the rejections, which
      is your point again from the other side.

      On the dataset: that loop shipped yesterday. A plan writes its target and deadline
      down before the test runs, you come back with the numbers you counted, and code
      grades them against the bar that was set first — you cannot resubmit upward once
      you have seen how it graded. So the mechanism is there and has nothing in it yet.
      Whether it fills is not up to me.

  10. 1

    Also, could you elaborate if the skew toward CRM & real estate related plays comes form your personal input or the systems assumptions - and have your analyzed https://www.ideabrowser.com ?

    1. 1

      That's my input, not the system. The run was seeded with one market — "real estate
      agent and property management tools for solo agents" — and the planner expanded
      that seed into 20 sub-markets, all inside it: listing-description generators,
      lease-renewal reminders, open-house sign-in, an MLS comps extension, an IDX
      widget. So the CRM flavour is the seed showing through, not a preference baked
      into the tool. Discovery also runs with no market given at all, which produces a
      completely different spread; I published a seeded run because a focused one is
      easier for a stranger to check.

      Haven't looked at ideabrowser. It's not in the set I've compared against
      (ValidatorAI, VenturusAI, FounderPal, ValidateMySaaS, PainMap, Trend Seeker,
      ideaproof), so that's a gap rather than a verdict — I'll go read it.

  11. 1

    You may find that founders who need their ideas scored in the first place have not yet been properly trained in problem discovery or in developing truly distinctive, high-potential concepts.

    I’m formally trained in evaluating ideas and spent 20 years at large firms doing this work for clients. I led teams that would develop up to 100 concepts for a client, with the goal of eventually producing one that became a major commercial or communications success. More recently, I’ve been automating the way professional product-innovation firms approach problem discovery and idea development. You may want to consider incorporating a similar component.

    The question is not whether to inflate scores—don’t; that would quickly undermine confidence and trust. The real question is whether you can help a founder move an idea from a score of 52 to 85+. The value is not simply in making the scoring more honest; it is in providing the direction and feedback that materially improves the underlying idea.

    1. 1

      Agreed on not inflating. The direction-and-feedback point is fair, but it's less
      absent than it looks.

      Two pieces already do it. Kill My Idea returns a "what would need to be true"
      ladder — the concrete conditions that would raise the verdict to hold, to test, to
      build — which exists so the founder can argue back with facts instead of taking
      the number. Pivot Wedge takes a broad or weak idea and narrows it into something
      one person can sell, or refuses and says no wedge exists rather than inventing
      one. And yesterday I shipped a line that names the single lowest-rated area of the
      top pick, so the next test aims at the thing most likely to sink it rather than at
      demand in general.

      Where I'd push back is the destination. 52 to 85 should not be reachable by better
      advice, because the two things holding that score down are confirmed distribution
      and hard evidence of demand — neither of which is knowable by reading pages about
      an idea. They change when the founder runs a test and reports what happened, which
      the product now grades against a bar written before the test started. So there is
      a path from 52 upward and it runs through evidence, not articulation. If
      articulation moved it, I'd have rebuilt the 88 with extra steps.

      If your automated problem-discovery work sits at the front of that — better
      concepts entering the funnel rather than better scores leaving it — that's a
      different and more interesting conversation.

      1. 1

        Nice! And important not to try and fix/elevate an idea that is fundamentally flawed or what in professional concepting sessions we‘d almost always discard as low potential or ”first ideas“ (even if technically it could be idea #100).
        Putting it before the funnel is kind of what I meant in a way any - but you could put in a mediocre idea, get a score that tells you it’s mediocre and then be send back to the front where you get help coming up with more valuable ideas - which you then check again until one clears 75 or 80/100 at the very least.

        Narrowing down on (or separating between) business, product and tactical ideas might also help you sharpen the results and value generation further.
        Have you tried using the JTBD framework to define what exactly users would ”hire“ it for?

  12. 1

    The linked-evidence vs. estimate split is a strong guardrail; I’d make the score’s uncertainty visible as a range rather than a single number. A useful next validation would be to record the lowest-scoring criterion and ask founders to test only that assumption with a pre-set budget or time limit. That turns “say no” into a conservative decision rule, not just a harsher score.

    1. 1

      Both land, and the second is a rule I don't have.

      On the range: you're right that the number is one draw. The ratings come from a
      model at temperature 1, so the same market twice can hand back a different
      shortlist. I do pay for sampling where a call is life-or-death — but only in the
      single-idea check, whose deal-breaker filters run five independent samples and
      vote, after single-sample noise turned out to be the main cause of false kills.
      The sweep you're looking at has no voting anywhere: sourcing, triage and scoring
      are one sample each. So I can't show you an honest range today, and a range drawn
      from a single sample is just an 88 with error bars on it.

      On the second: right now the recommended first move is picked by verdict, then by
      score — the strongest idea on the list, not the weakest assumption under it.
      Yours inverts that: take the criterion that scored lowest and test only that,
      against a limit you set before you start. It composes with the part that already
      exists, where the plan writes its target and deadline down in advance. That's a
      better decision rule than the one I shipped.

      1. 1

        Appreciate the transparency on single-sample vs voting — that distinction is the real product insight. Glad the ‘test the weakest criterion with a hard limit’ framing is useful; I’ll keep using that as the default when a score looks decisive but the underlying sample is thin.

  13. 1

    The point about setting the target before running the test stood out in your reply. It's easy to move the goalposts once you're attached to an idea. I like that this makes the founder commit to what would count as evidence before seeing the result.

    1. 1

      Worth adding, since it shipped a couple of hours after I wrote that: the target
      was already fixed before the test, but the answer wasn't. Reporting was free and
      unlimited, and the page listed every criterion you missed and by how much — so
      you could read "signups 41, needed 50", type 50, and submit again. A bar set in
      advance is worth nothing if the answer can be rewritten after you've seen how it
      graded.

      Now, within one plan, the report that counts is the weakest one you filed.
      Corrections downward land immediately; upward ones don't. Not "first report
      only", because that would trap a typo — mistype 500 for 50 and you'd have minted
      a result you could never withdraw.

      What it still can't do is tell whether your numbers are true. Nothing is
      verified. The property is that the bar came first, not that the report is honest.

      1. 1

        Thanks for explaining the update. I appreciate the distinction between keeping the target fixed and verifying the reported numbers. Your reply makes clear what the tool does and doesn't establish.

  14. 1

    The 52 creates a measurement system where the score itself becomes the data, not the validation. Most tools collapse to "well, maybe it works" because they optimized for confidence, not clarity. When a founder sees 52 vs 88 on the same idea, they're not evaluating the idea - they're evaluating whether they trust the measurement. That's where the decision lives.

    1. 1

      That's the bet, and it has an obvious cost: a 52 loses to an 88 in the market
      even when the 52 is the honest number. Trust in the measurement only pays off
      after someone has used the thing twice — the 88 pays off immediately. I don't
      have a clever answer to that, which is roughly why the demo is my own run left
      at 52 instead of a case study.

  15. 1

    The refusal to inflate a 52 is the interesting part. Have founders actually changed a build decision after getting a low score, or is the value still mainly in the research?

    1. 1

      Honest answer: I don't have that data yet. It's early enough that nobody has
      come back to me with "I killed it, here's what I did instead" — and I'd rather
      say that than tell you a story about outcomes I can't show you.

      Two things that aren't guesses. "Build it" is a verdict none of these modes has
      ever issued, on any run, including my own — it needs confirmed distribution and
      hard evidence of demand, and you can't get either by reading pages about an
      idea. And the build decisions it has changed so far are mine: it killed five of
      the eight ideas I put through it.

      Which is why the piece I shipped this week is the one that could actually answer
      you. The validation plan writes its target and deadline down before you run the
      test; you come back with the numbers you counted, and a plan created after your
      test started is refused — so the bar can't be picked once you already know the
      result. Ask me again in three months and I'll either have outcomes or I'll tell
      you I don't.

  16. 1

    The honest 52/100 and the split between linked evidence and estimates are the strongest trust signal here. I’d add a small tally of the 36 killed ideas by deal-breaker reason; that would make the demo useful not just as a score, but as a map of which assumptions keep failing.

    1. 1

      That tally is already in the run — it's just buried below the five finalists,
      which is a fair hit on the layout. On /sample-discovery, under "Why we rejected
      the rest":

      36 of 69 dropped. No real problem behind it 21 · Too much hands-on support 9 ·
      Needs heavy compliance 8 · Hard for a solo founder to sell or run 6 · Just an AI
      wrapper 6 · Leans too hard on one platform 4 · Can't charge enough 1 · No clear
      way to make money 1 · No repeatable way to reach buyers 1.

      Those sum to 57, not 36, because 21 of the ideas tripped two checks and 15
      tripped one — so it's a census, not a partition.

      You're right that it belongs higher up. Moving it next to the funnel line.

  17. 1

    Built a similar filter for our extension ideas — if we can't explain the acute pain in one sentence, we don't build. Saved months on features nobody asked for.

    1. 1

      "If we can't explain the acute pain in one sentence, we don't build" is the same
      rule with a human running it, and yours is cheaper. The gate here does exactly
      that, and it's why most ideas die: 21 of the 36 drops on that run are "we found no
      documented problem", which is the automated version of not being able to say the
      sentence.

      Months saved on features nobody asked for is the only ROI in this category I
      actually believe.