Solheim AI

Your Own Private Dedicated LLM

Visit Website
August 21, 2026 Virtual Private LLMs

A bunch of AI startups I spoke with recently have this perverse incentive of discouraging usage of their product.

They sell for a flat fee, but their own inference cost scales with usage.
Turns out on their most active users they had negative margins.

Changing their pricing to be a flat fee + token usage had their sales team struggling:
Closing clients without clear and predictable quotes is much harder.

The real fix for most: Get your own dedicated capacity, you can fix your unit economics and your incentives.

That is why I'm building Solheim: Reserve fixed capacity on a GPU for a flat monthly fee instead of per-token billing.
Predictable billing for predictable performance.

You trade elastic headroom for a ceiling, but capacity planning is much easier to solve than bad pricing -> Fix your product with engineering instead of financial acrobatics

2 Comments

  1. 1
    The negative margin on your most active users is the part worth sitting with, because it is not a pricing bug, it is a product design constraint. But I would separate two different fixes that get bundled together here. Dedicated capacity and flat pricing are not the same answer. Reserving a GPU for a flat monthly fee means you carry the idle cost yourself, so you have swapped a variable cost that spikes on heavy users for a fixed cost you pay whether anyone shows up or not. That is a real trade, not a free win. The other route is pooling variance. Charge flat, and bet that the average across a large base covers the tail. That only works if the tail is thin. If one user can burn 100x the median in an afternoon, the pool does not save you, it hides the loss until it is bigger. Disclosure: I am the founder of Piramyd, a flat $30/mo unlimited-token gateway for Claude Code, Codex and Cursor. I took the pooling route, so I have a stake in the answer. A question that decides which fix is right: on those users with negative margins, was it a steady above-average drain, or one runaway loop? Pooling handles the first. It does nothing about the second, and dedicated capacity does not either. It just raises the ceiling they blow through.
  2. 1

    The pricing tension is the strongest part here. When usage directly increases your own delivery cost, “predictable pricing” becomes an infrastructure problem, not just a packaging decision.

About

AI inference pricing is broken, per-billing token only makes sense for a small subset of companies. In traditional compute we have the concept of a "VPS" and for most parts that doesn't exist for GPUs. Now it does.