Here is the test, and why the number would have been a lie.
The setup: 8 companies with known endings, 5 failures and 3 successes, each reduced to the one-liner it could have been described by at founding, name stripped, run through the live gate under two founder profiles.
Only a KILL counts as catching something. The moment a hold or a pilot-first counts as a save, every tool in this category is accurate, including the ones that are merely agreeable.
Results on 2026-06-23. Solo lens: 4 of the 5 flops killed. Funded lens: 2 of 5. And the row I would have quietly dropped if I were selling this — the solo lens also killed 1 of the 3 successes.
That killed success is a payments company, and it is not an error I am hiding. For one person, it is the correct answer. For a funded team it is not, and the gate says so. The verdict is bound to the founder, which is the product.
Which is exactly why a single accuracy figure is incoherent here rather than just imprecise. Accuracy against what? The same one-liner has two correct answers depending on who is asking.
Three more reasons the number would have been marketing: the sample is small and hand-picked for fame, hindsight may leak through the anonymisation, and the weighted total is a single sampled pass that can drift a point between runs.
What I would ask of anything in this category, mine included: show me the rows where you were wrong. Not a case study. The table, with the misses still in it. A tool that only publishes wins has published nothing.
Every row, including the 3 that neither lens killed at all: https://whittleos.com/guides/do-startup-idea-validators-work