Customer growth can expose weaknesses that early traction hides. But registered user count alone says little about the pressure on a product. Concurrent requests, expensive queries and different usage patterns matter more.
In a GeekyAnts podcast discussion on scaling from 100 to 100,000 customers, Ankur Gupta, introduced as a solution architect at OceanBase, explores how architecture, databases and costs change as products grow.
Several takeaways are useful for founders building with limited engineering budgets.
Start with the workload
A thousand customers running complex searches can generate more pressure than a much larger, mostly inactive user base. Capacity planning should reflect what customers actually do, including peak activity and the workflows that must remain reliable.
Find the bottleneck before adding servers
The discussion highlights query execution, joins and data distribution. More application instances may not help when every request still waits on the same database bottleneck.
A practical starting point is to trace one important journey, identify its slowest operations and test a targeted improvement under realistic load.
Reduce repeated work carefully
Caching can reduce unnecessary API calls and database reads. However, teams need clear rules for freshness, invalidation and customer-specific data. Faster responses are useful only when the information remains correct.
Keep architecture aligned with the business
One example in the episode describes investing in analytical capabilities before they matched the business’s priorities.
For a growing product, every infrastructure decision adds costs beyond hosting: maintenance, debugging and operational complexity. Tracking cost per successful transaction can help connect technical spending to customer value.
Treat distribution as a design choice
The episode emphasizes horizontal scaling and unified data access. These approaches deserve evaluation, but their suitability depends on workload, consistency requirements and operating constraints. Neither distribution nor consolidation automatically solves a capacity problem.
Preserve what still works
Growth does not automatically require a rewrite. Teams can measure current limits, improve specific bottlenecks and expand capacity incrementally. A larger redesign becomes easier to justify when evidence shows that the existing architecture cannot meet the next requirement.
For founders, the useful question is: Which assumption will stop holding at the next stage of growth, and how can it be tested before customers experience the failure?
One useful way to map this is by the bottleneck that changes at each stage. Early on it may be discovery and activation; later it becomes support, reliability, delegation, and decision latency. Which transition first forced you to change the operating model rather than simply do more of what worked before?
"Registered users" as a capacity metric is how teams buy the wrong servers. Concurrent expensive journeys beat headcount every time. Cost-per-successful-transaction is a great forcing function for small teams too — it makes "add another cache" compete with "fix whether this path even needs to run." Trace one critical journey under realistic load before you rewrite anything.
The "cost per successful transaction" point is the one I'd underline for small teams, especially with AI in the loop.
A small example from our app: we generate new learning content with AI every night, and the job first checks whether even the most active users still have enough unseen content. If they do, it exits without spending a single token. Tying spend to actual demand turned a cost that would grow with every day into one that grows with real usage.
Same spirit as "find the bottleneck before adding servers": measure the real pressure first, then spend.