Service
LLM Self-Hosting & On-Prem Inference
Consulting and architecture for self-hosted AI models — hardware class, model selection, hosting choice, procurement and capacity planning. Built on our own operational practice.
Self-hosted AI models became technically and economically within reach for the Mittelstand in 2026. What remains is the architecture question: which hardware, which model, which host — and whether operations are carried in-house or externally.
When self-hosting pays off
Four drivers:
- IP protection. Source code, design data, clinical data, or client communications must not leave your own network. External LLM APIs may process requests outside the EU, depending on provider, plan, and region.
- Compliance. Industry-specific regulation — CRA, NIS2, ISO 27001, MDR, HIPAA — restricts the transfer of sensitive content. Self-hosting establishes an auditable data-flow boundary.
- Volume economics. Above certain token loads, in-house operation becomes more economical than API pricing. The threshold depends on model, hardware class, and utilization, and is part of the consulting.
- Latency. Workloads that need low and predictable response times benefit from local inference without network latency.
Privacy by construction
All inference workloads run on-premises. No cloud GPU dependency. Authentication via the operating system's secure key store — no plaintext keys in the repository, no inheritance to foreign processes. Audit logs at the metadata level; prompt content is not persisted.
Our position: from our own practice
Creaminds operates its own LLM infrastructure for internal workloads — coding assistance, agents, the infrastructure of this website. What we use and have proven internally every day feeds directly into our consulting. The consulting goes beyond our own stack as soon as customer use cases bring different requirements.
What Creaminds delivers
- Workload analysis. Token volume, sensitivity classification, latency requirements, update frequency.
- Model selection. Recommendation from the current open-weight spectrum (Mistral, Llama, GLM, Qwen, and others), including fine-tuning options for your own domain.
- Hardware design. Selecting the right class — from workstation hardware for pilot projects through inference NPUs for high throughput to production servers for full model sizes.
- Capacity and procurement planning. Sizing against real peak load with a defined reserve, phased procurement, and lead-time strategy (see below).
- Hosting recommendation. Selecting suitable infrastructure providers from EU bare-metal, NPU-specialized providers, or hyperscaler EU regions — depending on risk and cost profile.
- Migration plan. Transition from your existing cloud LLM setup (Copilot, Claude, ChatGPT Enterprise) including pilot phase and parallel operation.
How we size and procure
Two mistakes dominate on-prem inference projects: sizing too tightly and procuring in the wrong order. Our approach addresses both.
- Size against the peak, not the average. AI coding tools generate burst load — agentic loops and tab completions come in waves. We size against peak load and build in a defined reserve that absorbs short-term bursts and organic user growth without short-notice re-procurement.
- Procure the pilot in parallel with production. Production hardware has months of lead time. Ordering sequentially loses that time. A quickly procurable pilot delivers immediate validation while the production hardware matures in parallel — the long lead time is off the critical path.
- The pilot stays in operation. After go-live, the pilot hardware is not discarded but becomes a permanent dev/staging environment, model evaluation platform, and cold standby for maintenance windows. This reuse planning is part of the business case from the start.
- Interconnect as a selection criterion. Large models that exceed a single GPU's memory need tensor parallelism — and therefore NVLink/NVSwitch rather than a PCIe-only topology. For the production class this is a selection criterion, not a nice-to-have.
- Single-node trade-off communicated openly. Starting with a single production node keeps CapEx minimal but introduces a single point of failure. We name it and mitigate it — through cold standby, announced maintenance windows, and vendor SLA — rather than hiding it.
A central economic lever of the platform is prefix-cache sharing: the shared context of a coding request (system prompt, codebase subgraph, tool definitions) is computed once and reused across developer sessions — rather than recomputed for every request. On modern GPUs this is the single largest cost lever.
Operating model
Two typical configurations, depending on the customer setup:
- Documented handover. Creaminds hands over the running stack fully documented to the customer's IT or an operations partner. Maintenance, updates, and SLA rest with the customer.
- Managed service. Creaminds takes over operation, updates, and monitoring as agreed. This variant requires a clearly defined service level and is calculated case by case.
Which model fits the individual case is decided in the customer conversation.
Hardware landscape 2026
Three classes, three use cases:
| Class | Examples | Suitable for |
|---|---|---|
| Workstation | NVIDIA DGX Spark, Apple Silicon Mac Studio | prototype, pilot, small teams |
| Inference NPU | Rebellions REBEL-Quad, FuriosaAI RNGD | high token throughput, low power budget |
| Production server | NVIDIA H200/B200, AMD MI300X | full model sizes, high parallelism |
Model landscape 2026
The consulting covers the open-weight spectrum — generalists, coding specialists, and domain-specific models. Selection and, where appropriate, fine-tuning follow the workload requirement.
Hosting geographies
The choice of hosting location is decided in the customer conversation. The deciding factors are compliance obligations, data classification, and workload size. Creaminds advises on EU bare-metal providers, NPU-specialized cloud providers, hyperscaler EU regions, and co-location models.
For whom
Mid-market companies that process sensitive data and want to deploy AI assistance productively without handing data to external providers. Typical triggers: rolling out AI coding tools across development teams, compliance requirements (GDPR, industry-specific confidentiality), unpredictable cloud API costs.
Approach
- Workload analysis and sensitivity mapping
- Model recommendation and hardware class
- Capacity sizing with reserve and procurement/lead-time strategy
- Host selection with contract-negotiation support
- Deployment plan and migration path
- Handover to customer IT, operations partner, or as a managed service