A bunch of AI startups I spoke with recently have this perverse incentive of discouraging usage of their product.
They sell for a flat fee, but their own inference cost scales with usage.
Turns out on their most active users they had negative margins.
Changing their pricing to be a flat fee + token usage had their sales team struggling:
Closing clients without clear and predictable quotes is much harder.
The real fix for most: Get your own dedicated capacity, you can fix your unit economics and your incentives.
That is why I'm building Solheim: Reserve fixed capacity on a GPU for a flat monthly fee instead of per-token billing.
Predictable billing for predictable performance.
You trade elastic headroom for a ceiling, but capacity planning is much easier to solve than bad pricing -> Fix your product with engineering instead of financial acrobatics
The pricing tension is the strongest part here. When usage directly increases your own delivery cost, “predictable pricing” becomes an infrastructure problem, not just a packaging decision.