Spent the last 5 months building 1mil.app, a scanner that takes a market, runs live research on it, and returns 10 ranked business directions. You tell it your background, and it scores each idea on demand evidence and on whether someone like you could actually win it.
Last week I finally did the thing I'd been avoiding: I pointed the analysis at my own corpus. 339 scans. Roughly 3,400 ideas surfaced since March.
The result: zero ideas rated 7 or higher out of 10 on winnability. The best in the entire corpus was a 6.1, and it wants 1,200 build hours. 75% of everything scored below 2.
In June I took my six top-scored ideas and ran a willingness-to-pay check on them. 5 of 6 bombed. The old score measured demand and demand alone. So I rebuilt the scoring to price in competition, distribution, moat, and build cost, relative to the specific founder asking. The new model is stingy... maybe because reality is stingy too?
Here's what I learned from all that:
My hesitation to pick and build an idea wasn't just me procrastinating. I have been hoarding ideas for months and pursuing none, and the data confirmed why: there was no obvious winner in the pile.
Scores don't replace potential customer conversations. The scanner's real job is to say which idea deserves those conversations, and which nine don't.
A tool that flatters you is worse than useless. If my tool had done that, I'd have wasted months building something DOA.
Kinda ironic that my scanner's most useful output so far was telling its own founder "not this one, and not that one either." I shipped that behavior on purpose, and it cost me the comfortable fantasy of an easy win sitting in my backlog.
If you have a pile of ideas and can't choose, the free plan is 3 scans a month (needs an account, that's how I ended up talking to my first users):
https://1mil.app/?utm_source=indiehackers&utm_medium=show_ih&utm_campaign=launch
And if you think scoring winnability is impossible in principle, tell me why. That argument is the one I most want to have.
I really like that you tested your own assumptions instead of treating the score as unquestionable. One thing I'm curious about: what's the biggest factor that consistently separated the higher-scoring ideas from the rest? Was it distribution, willingness to pay, or something else?
3,400 ideas and none flagged as easy wins is actually useful signal in itself. What criteria is it using to judge 'easy win'?
You asked for the argument that scoring winnability is impossible in principle. I don't have that one — I think it's scorable. But there's a weaker link earlier in the chain: the 3,400 ideas you scored were surfaced by your own scanner, so the corpus and the scorer share an origin.
That's what makes "the new model is stingy... maybe because reality is stingy too" a coin flip with a third side. If the generator leans on what's already documented, it lands in well-covered space a lot of the time, and a near-zero ceiling is then partly a fact about where the ideas came from rather than about the market. You can't separate those right now: everything in the pile came out of the same machine.
Where I'm standing, so this doesn't read as advice from above: I run an experiment where an AI does the operator work end to end and I only touch the irreversible levers. I let it run hands-off for a stretch and it produced a steady supply of plausible, unwinnable directions. Sales are $0. What eventually moved it wasn't better scoring — it was an angle a person brought that the model wasn't going to generate.
The control I'd want is scoring ideas that didn't come from the scanner — whatever your first users brought you in conversation rather than ran through it — and seeing where those band. But something else comes first: you mentioned the same topic re-run swung the ceiling between 1.5 and 6.0. That's wider than the gap such a control would be trying to show. Freeze the model, get the per-topic noise floor, and the comparison becomes readable; run it before, and you get a number that supports whichever story you already believed.
Which is the call I'd actually put in front of you: not how to tune the weights, but whether the 75%-below-2 distribution is evidence about reality yet. Until you know the noise floor, I don't think it is.
Fascinating project, Eli. The fact that your tool told you the hard truth about your own ideas is the ultimate validation of its utility. Most tools just flatter the founder.
On the debate about whether scoring winnability is impossible: I think it is possible to estimate, but the variables (especially 'build cost' and 'distribution') fluctuate wildly based on the founder's hidden advantages. For example, a 6.1 idea requiring 1,200 build hours might be a 9 for a solo dev who uses enterprise-grade boilerplates to cut 80% of the build time, but a 2 for someone starting from absolute scratch.
I run into this as an engineer. I build production-ready Next.js templates (HadiKits) precisely because founders constantly underestimate the 'build cost' variable in their winnability score. You can have a winning idea, but if the execution drains your resources before launch, the math was wrong.
Curious how you factor 'build hours' into your model—is it a flat estimate, or does it adjust based on the founder's technical stack and existing assets?
3,400 ideas and none are easy wins? That's actually a good thing. It means the hard ones have moats. The easy wins are all taken by now.
nice!
"That makes sense — picking ideas that extend what you're already using is probably smarter than starting from scratch. I'm learning that the hard way with Rallynex — trying to build something that fits into existing workflows rather than forcing a whole new one. Which tools/platforms are you augmenting? Curious how you decided which ones were worth building on top of."
The 6.1 ceiling might be the most honest output a tool like this can give — but I'd flag what it can't see: the version of the idea you discover three weeks in. The thing I'm building now would have scored miserably on paper (crowded space, giant incumbents), and its actual wedge only showed up after watching real users do something I didn't design for. Scanners rate your starting position; most of the winnable-ness gets created mid-game by whoever stays in it. Fun test would be scoring the ORIGINAL ideas behind famous pivots and counting how many 3s became 9s.
Well put.
Winnabilty is part of the game. 🤑
Building any sort of model (machine learning, or even LLM/prompt-driven) is really hard unless you have a good training/test set: input data with the associated outputs so you can build a model and test whether the model, or your tool, does a good job.
You're never going to have a perfect experimental data set, but I think you can try to prove whether your tool works or not--and then it doesn't matter whether people agree in principle or not.
So pick a few ideas that you've seen work in real life, for some definition of "work" (successful write-ups here, funding announcements, acquisitions, etc.). You'd want to avoid anything sufficiently well-known that the LLM would know about the actual company. Any single idea's success is of course subject to all sorts of randomness and factors that are hard to measure. But if ideas that have worked tend to produce high winnability scores, then you'd start to convince me it's doable--and I think that would be a stronger argument than any argument you could make in principle.
Went ahead and tested your hypothesis as follows.
4 scans describing bootstrapped products that had commercial success, without naming each product. I also used real founder's pre-success background as the profile. Here's the test set:
Bannerbear (Topic: "Automatically generating social media and marketing images at scale via an API, so marketing teams stop manually creating repetitive visual assets" + Role: "Solo designer-developer. Ten years building websites and marketing tools for agency clients, strong at both design and code, comfortable shipping alone. No team, modest savings, active blog and small following from building in public.")
Lunch Money ("Personal budgeting and expense tracking for people who currently manage their money in spreadsheets and find budgeting apps too rigid" + "Solo software engineer. Years of experience shipping consumer web apps and side projects, comfortable across the full stack. Working alone, self-funded, no audience to speak of, patient about growing slowly.")
Testimonial.to ("Collecting video testimonials from customers and embedding them on SaaS landing pages without engineering work" + "Solo indie developer, formerly a network engineer at a large tech company. Full-stack JavaScript, ships fast, active on Twitter with a small but real following of indie hackers. Self-funded.")
SavvyCal ("Scheduling links and calendar booking that feels personal and polite instead of demanding, for consultants and professionals" + "Experienced SaaS founder. Previously co-founded and sold an email marketing platform, deep expertise in SaaS product design and developer marketing, well known in the bootstrapped SaaS community, funded by previous exit.")
Result (top-scoring idea per scan):
Testimonial.to - W5.9
Lunch Money - W4.1
Bannerbear - W3.5
SavvyCal - W1.0
The scanner does live research, so in all 4 cases it found the actual winner and priced it in as competition. It isn't answering "was this winnable back then" -- it's answering "can a solo founder win this space today". Scheduling in 2026 is Calendly and Cal.com territory, so 1.0 for a newcomer seems justified.
On the Bannerbear scan, it reinvented Bannerbear almost exactly -- "image generation API with conditional logic in templates" -- and scored it 1.5, because Bannerbear exists now and owns that space. Same with scheduling: it came up with a developer-first scheduling API, which is basically Cal.com, and gave it 0.2.
So my tool can't really run your test retrospectively. I can't rewind the web it researches against. But where you're fully right is outputs. I have 3,000+ scored ideas sitting in the database: predictions, frozen, timestamped, and zero outcomes attached to any of them. I've pulled 25 winnability-scored ideas from the historical scans (every top-band one, plus samples from the middle and bottom), stripped the scores, shuffled them, and I'll desk-check each one blind: does the buyer exist, do they already pay anyone for this, what's the strongest free alternative. Then unblind and see if the bands separate. If top-band ideas survive at a clearly higher rate than bottom-band ones, the ranking means something. If they don't, the tool is producing persuasive numbers, and better that I find out than a user.
Nice
Have you tried the tool? :)
"3,400 ideas is wild. What made you finally pick one to build? I'm building Rallynex and struggling with that focus myself."
I picked several that augment my existing tools and platforms.
"That makes sense — picking ideas that extend what you're already using is probably smarter than starting from scratch. I'm learning that the hard way with Rallynex — trying to build something that fits into existing workflows rather than forcing a whole new one. Which tools/platforms are you augmenting? Curious how you decided which ones were worth building on top of."
The insight about "distribution gap" is what really separates validation from validation theater. Most idea tools amplify the fantasy because their business model depends on keeping you building. Your tool does the opposite - it says "stop until your audience is ready."
That's harder to commercialize (nobody wants to hear it) but it's actually the most valuable signal. The 5/6 willingness-to-pay failures probably teach the model more than the successes - what matters is whether founders with distribution problems can now recognize the pattern before burning runway on a problem they can't solve with code.
Did you place successful acquisitions in it to test it? :)
What do you mean?
The strongest product insight here is that "winnability" probably has to include one very boring variable: can this founder reach 20 qualified buyers this week without borrowing an audience?
Demand, competition and build cost matter, but for tiny founders distribution access is often the real constraint. I would score each idea against an actual founder-owned path: existing list, communities they already participate in, customers they can email, niche expertise, or paid channels they can afford to test. If the path is vague, the idea should stay low even if the market looks big.
You got that right. It's up to the user to input as much info as possible about their expertise, experience, and whatever else they'd bring to the table.
Treat winnability as a ranking model, not a truth score. The next useful test is prospective: freeze the model, take a holdout set of ideas it has never seen, and run the same low-cost validation sprint on the top, middle, and bottom bands.
Then compare precision at the top: what share of the top 10% produces a qualified conversation, a concrete pre-commitment, or a paid pilot? Also test a few low-ranked ideas to estimate false negatives. If high-ranked ideas consistently outperform low-ranked ones, the model is useful even if no score reaches 7. If everything fails equally, it is only producing persuasive numbers.
The product's moat may end up being calibration data from these validation outcomes, not the scoring formula itself.
This is the right test, and I have one early data point in its favor: re-running the same topic several times last week, the winnability ceiling swung between 1.5 and 6.0. So individual scores are noisy draws, which lands me where you did: the model should be judged on band-level precision, not any single number. Freezing it and running cheap validation sprints on top vs bottom bands is now the plan. And I think you're right about the moat. The formula is copyable. The record of which validations actually worked is not.
75% of the corpus scoring below 2 is the part that lands hardest — most idea tools tune their output to feel useful, so keeping the model stingy is what makes it a decision tool instead of a validation machine. The 5/6 willingness-to-pay failures are fascinating: did that check separate "won't pay" from "won't pay yet — no audience"? Curious whether the new scoring weights distribution enough to catch that gap.
Yeah, those are two different failure modes, and the scoring treats them separately.
Demand captures "people pay for this."
Winnability catches "but the winners had an audience you don't have."
That's the distribution gap. The classic example: ShipFast did $263k in a year, but that revenue rode on the founder's existing audience. Without it you're the 41st identical Supabase + Stripe kit nobody sees.
The 5/6 willingness-to-pay failures are the interesting part to me.
You changed the scoring model after seeing that result — what would you need to see now to know the new “winnability” score is actually better at choosing which ideas deserve customer conversations, rather than simply being more conservative?
Winnability is one more signal. The scanner already uses your background to shape what comes back, and winnability flags structural stuff like free-tier competitors or moats you can't build in code. That narrows the field. But which idea you actually pick up depends on who you can reach. A plumber and a lawyer looking at the same scan would pick different ideas because they talk to different people every day. Whether it's good signal, I'll find out from users who actually made the calls.
I appreciate you taking the time to explain your thinking.
I'd be interested in continuing the conversation by email if you're open to it. What's the best email to reach you on?
No, thanks.
Not keen on being part of your email harvesting campaign.
But thanks to your bot for stopping by.
;-)
Fair call. I can see why it came across that way.
Appreciate you taking the time to answer my question thoughtfully.