
Spent a night making my own GPU startup prove its claims to me before I'd let it make them to anyone else.
Here's what triggered it: our pricing calculator let people select up to 512 GPUs. I looked at that number and realized I had no idea if we could actually deliver past 8. So I stopped everything and made us test it, live, no simulations, no "should work."
Rented real GPUs across two providers. Fired concurrent load tests, 3 simultaneous large deployments across Japan, Romania, and Canada, all held at once on real machines. Found the honest edges of what we can actually do: RunPod caps every pod at 8 GPUs no matter what you pay. Vast.ai gets us to 16 GPUs on workstation cards, 4 on B200.
One request failed mid-test. Lost a race for the only big machine available on the whole marketplace at that moment. No charge, no cover-up, just a real failure in a real log.
I'm not posting this because it makes us look flawless. I'm posting it because it's the opposite, and I'd rather you find out we test our own limits than find out the hard way that we didn't.
Total cost of proving all of this: under a dollar in real cloud spend. Worth every cent.
kilawattcloud.dev
The under-a-dollar test uncovered a mismatch between the 512-GPU selector and the 8-GPU RunPod cap, which is a good example of testing the promise rather than the happy path. I would make the calculator show provider, card class, and current availability beside each range, then log failed provisioning races as a separate constraint from hard capacity. How often will you rerun these live tests as provider inventory and limits change?
Appreciate the breakdown, both suggestions are solid and we're taking them seriously.
On the calculator: showing provider, card class, and live availability beside each range is the right fix. Right now the cap only shows up when someone hits it, surfacing it earlier removes that confusion for anyone using the tool, not just us.
On the logging: separating a hard capacity limit like RunPod's flat 8-GPU cap from a lost provisioning race for available stock is a fair distinction, they're different problems with different fixes, and our data should reflect that.
On cadence: we run these live tests weekly, and we have the data to back it up. Happy to keep publishing updates as inventory and limits shift.
The failed request is worth more than the three regions holding, and I do not think the number you took from it is really a ceiling. Losing a race for the only big machine on the marketplace means your capacity is contested, so sixteen on workstation cards is actually sixteen when nobody else is bidding at that moment. That makes availability a probability rather than a limit. If you are already logging provisioning attempts, the number I would want on the pricing page is the hit rate: out of the last hundred large requests, how many got the machine on the first try. A disclaimer tells me a cap exists, a hit rate tells me what happens when I actually click the button.
This is a sharp distinction, and you're right that a static ceiling and a real-world hit rate are two different claims. Availability under contention is exactly the kind of nuance we care about getting right as we scale. We're taking it into account as we refine how we report capacity going forward. Appreciate you pushing on the data instead of just the headline number, that's the standard we're trying to hold ourselves to.
The failed request is the part I keep thinking about, not the three regions holding. You lost a race for the only big machine on the marketplace, no charge, real log. Under a dollar and you found RunPod hard-capped at 8 while the calculator still said 512. Did you actually drop the slider after that night, or is 512 still selectable and you just hope nobody picks it?
The slider goes past 8 on purpose, that's not an oversight, it's how the calculator doubles as a quoting tool for larger deals. Anything past real self-serve capacity, 8 GPUs on any card class, 16 on workstation-class when stock exists, shows a live disclaimer right at the point of selection: no on-demand machine that size exists, and it routes to our team instead of letting anyone believe it'd just spin up. Check it yourself before assuming it's hiding something. Nothing on that page claims capacity we don't have.
Testing your own product claims before letting customers see them is a level of integrity most early-stage founders skip entirely. The fact that you caught a real capacity ceiling by spending under a dollar on live tests — and then updated your pricing to match reality — is exactly the kind of discipline that builds long-term trust. Compressing ambitious specs into honest, verified ones is a harder sell upfront but it's the right foundation.
That's the trade exactly. It would've been easier to leave the specs ambitious and let the pricing page do the talking. Verifying first means slower claims, but it means every number on that page is one we'd actually stand behind if someone called us on it. Appreciate you seeing that as the foundation it's meant to be.
This kind of radical transparency is rare in the cloud/GPU space. Most founders just leave the slider at 512 and pray an enterprise client doesn't actually click it. Spending $1 to find your actual hard limits before a customer does is the best ROI you'll ever get.
That's exactly the trap we were trying to avoid. A slider that goes to 512 isn't a feature if nobody's checked what happens past 8. We'd rather find our own ceiling on our dime than have an enterprise customer find it on theirs, mid-deployment. Appreciate you naming it that clearly, that's the whole point of publishing this stuff.
Spending a buck to catch provisioning bottlenecks now saves you from a massive headache when a real paying user tries to spin up a node.
Exactly, that's the trade we made going in. A dollar and a failed test now beats a customer hitting that same wall with real money and real timelines on the line. If we ever find a bottleneck we didn't catch, that's on us, not a surprise for whoever's building on top of us.
The test changed the claim from “up to 512 GPUs” to something much more specific. How are those verified limits changing what you’re willing to promise customers?
Good catch, and yeah, that's basically the whole story.
Before this, our pricing calculator let people select up to 512 GPUs with zero verification behind it, that number was aspirational, not tested. After spending a night actually proving what's real, here's what we're now willing to promise, and only this:
Up to 8 GPUs, any card type, instant self-serve. That's backed by live testing across two providers.
9 to 16 GPUs, workstation-class cards only (A4000/A5000/A6000/3090/4090/A40), self-serve, but gated on real-time marketplace availability. If the inventory isn't actually there, we tell the customer that instead of pretending it is.
B200, up to 4 GPUs, same deal, live-checked against real stock.
Past those numbers, we don't fake it. It routes straight to a real conversation with our team, because that capacity genuinely doesn't exist on-demand from any provider we work with right now.
So the short answer: we went from promising a number that sounded impressive to promising exactly what we can prove, in real time, every time. Less flashy, way more trustworthy.
That’s a meaningful shift — you’ve moved from an aspirational capacity claim to promises you can actually verify. I sent you a note by email on Friday; reply there when you get a chance and we can dig into what this changes commercially.
Thanks, Aryan — got it. For anything on the commercial side, best to reach us directly: hello@kilawattcloud.dev. Happy to dig into what the verified numbers mean for you.
Thanks! I’ve just sent it over.
Looking forward to hearing your thoughts whenever you have a chance.