Quick context on why I'm doing this: I've been using Claude Code a lot, and like a lot of people, I kept hitting the same wall — not the token bill, but the usage window. You're mid-task, hit the ceiling, and wait hours for it to reset. Flat fee, but still capped in a way that doesn't match how I actually work.
At the same time I kept thinking about the EU angle — most inference still routes through US-jurisdiction infra, and that's getting harder to justify for a lot of teams, not for ideological reasons, just practical GDPR/AI Act ones.
So I'm building Solheim.ai: you get your own private LLM instance — think "rent the machine, not the tokens." One flat monthly fee, sized by how many requests you want to run at once (a simple slider), not by a token meter or a rolling usage window. No reset timer, no surprise bill on a busy week. Hosted entirely on EU infrastructure, running open-weight models (starting with Qwen3.6).
It's OpenAI-compatible, so it drops into Cline, Roo Code, VS Code, or your own backend without changing anything else in your setup.
Super open to feedback here, especially from anyone who's hit the Claude Code usage-window wall themselves, or anyone doing EU-only infra for compliance reasons — trying to figure out which of those two is the sharper pain point before I build further.
The "rent the machine, not the tokens" framing is the clearest way I've seen this pitched, usage-window resets are a genuinely weird UX pattern once you notice it: capped like a metered product, priced like a flat one, so you get all the friction of a limit with none of the predictability a flat fee is supposed to buy.
On your actual question, EU-hosted compliance is the sharper, more durable pain point of the two. The usage-window frustration is real but Anthropic could change that limit tomorrow and the complaint evaporates; a team that needs GDPR/AI-Act-compliant infra has a requirement that doesn't go away regardless of what any single vendor does with their rate limits. I'd build for the buyer who has no alternative path, not the one who's one policy change away from not needing you.
Since you're starting with open-weight models (Qwen3.6), the honest tradeoff to be upfront about is capability gap versus Claude/GPT-tier models for anything genuinely hard, worth knowing which segment of your target users cares more about "good enough and compliant" versus "best available," since that probably determines whether Qwen3.6 is a fine starting point or a ceiling that loses you the harder use cases.
Hi CodingPanda42, With Solheim you are explicitly testing which pain is sharper—usage-window limits or EU-only compliance—before building further. I’m building Build Before 2030: a manually reviewed public record where makers claim a Founding 100 number, set one measurable 30-day milestone, and later add proof of what shipped. The first cohort is free. Solheim feels like a strong fit. Interested? https://buildbefore2030.com
Those two pain points seem like they could lead to very different buyers even if the underlying product is identical.
What are you looking for as evidence that one is strong enough to become the primary reason to choose Solheim, rather than keeping both in the positioning?
Indeed, I think ultimately I just need to see what people will be using it more for.
There are of course high-profile cases of AI for coding blowing through budgets so there is a market there.
But my guess would be that inference is the bigger challenge in the big picture