Solheim
fixed fee with no usage or token limits
Last checked:
What is Solheim?
Solheim is an EU-hosted inference service that offers private LLM instances on a flat monthly fee, with no token metering or usage quotas. Users rent compute resources and set their own limits on concurrent requests and context window size, with an OpenAI-compatible API for easy integration into existing workflows.
Solheim pricing
Pricing model: Freemium
€15.00/month for 1 instance with 64k context on Qwen3.6-35B; pricing scales with instance count and context window size. Dedicated/enterprise tiers available by contact. No token billing regardless of usage level.
Solheim pros
- Flat monthly billing with no token metering or surprise costs—usage doesn't affect the invoice
- OpenAI-compatible API works directly with Cline, ZooCode, VS Code BYOK, and standard SDKs with no code changes
- EU-only infrastructure ensures GDPR and AI Act compliance, with no US hyperscaler exposure
- Configurable instance count and context window on-demand; beyond instance limit, requests queue instead of failing (no 429 errors)
Solheim cons
- Limited to open-weight models only (Qwen and DeepSeek); no proprietary or frontier models like GPT or Claude
- EU-only geographic availability; no option for US or other regions
- Requires email signup with no immediate API key provisioning; enterprise plans require contact/negotiation
Frequently asked questions about Solheim
How does billing work?
Fixed monthly fee based on instance count and context window size. All usage within those limits is included; no additional charges for tokens, requests, or overages.
What happens if I exceed my instance count?
Additional requests queue and execute in order without incurring extra charges or returning 429 errors—they simply wait for capacity.
Which models are available?
Open-weight models under permissive licenses: Qwen3.6-35B (128k context), Qwen3.8-27B (128k context), and DeepSeek V4 Flash (256k context). Users can request other open-weight models.
Is my data private?
Yes—each project gets its own VPL (endpoint and key), and all infrastructure is EU-based with no US hyperscaler involvement. Prompts are not used as training data.