Sovereign on-premise vs hyperscalers: the case with data (cost, energy, performance and sovereignty)
Contents
Notation: amounts in euros (N €), decimals with a point; when a source quotes dollars it is marked “USD”. Data centred on Europe and Spain. The dollar symbol is not used (formula delimiter).
What this article covers
Synthesis article of the series (S2), and the heart of the proposal: the case with data for sovereign on-premise against the hyperscalers and the European cloud. Up to here, each track measured its own axis, cost per token (FinOps), goodput (benchmarking), energy and carbon; here the four are crossed, with a fourth dimension no US technical comparison puts up front: data sovereignty. The aim is to answer, with numbers and not with ideology, the question that underpins any investment in an AI platform: serve on your own iron, on a European cloud or on a hyperscaler? And to do it honestly: on-prem does not always win, and saying when it wins and when it does not is what makes the recommendation credible.
The framework: four axes, not one number
The mistake in almost every comparison is reducing the decision to the hourly cost of a GPU. The real decision crosses four axes, and only by seeing them together do you decide well:
| Axis | Question | Who measures it in the series |
|---|---|---|
| Cost (TCO) | how much does it cost to serve, all in? | FinOps (A2–A8) |
| Performance | does it meet the SLO, at what goodput? | Benchmarking (B2–B8) |
| Energy and carbon | how many watts and grams per token? | Energy (C2–C8) |
| Sovereignty | under which jurisdiction does the data live? | GDPR / EU AI Act / CSRD |
The first three are quantifiable and meet in the cost per token; the fourth is a constraint that can rule out an option however cheap it is. The synthesis consists of scoring each option on all four and deciding on the Pareto frontier, not on whichever axis happens to suit.
The cost axis: TCO and break-even
The real on-premise cost is amortised capex plus opex, and its cost per effective hour depends on utilisation:
$$\text{effective cost/GPU-hour (on-prem)} = \frac{\text{annual amortised capex} + \text{annual opex}}{8760 \times \text{utilisation}}$$This formula is the key to the whole debate: the cost per useful hour of on-prem rises as utilisation falls, because capex is paid whether the GPU is working or idle. The 2026 data:
| Data point | Value | Source |
|---|---|---|
| On-prem cost 8×H100 (floor, high util.) | ~2.83 USD/GPU-hour all-in | Spheron |
| Annual on-prem cost (floor) | ~237,000 USD/year | Spheron |
| AWS H100 (p5.48xlarge) | 4.10–6.88 USD/GPU-hour | Spheron |
| AWS 8-GPU on-demand annual (100 % util) | 287,000–482,000 USD/year | Spheron |
| Sovereign European cloud (Lyceum/Scaleway) | from 2–2.73 €/GPU-hour, zero-egress | Lyceum, Scaleway |
The break-even against AWS on-demand falls around 50–83 % utilisation depending on region and tariff; below ~70 % utilisation, the cloud wins on TCO, and above it, on-prem (Spheron). For very high-utilisation workloads, on-prem pays for itself in less than 4 months (Lenovo).
Three-year TCO: the full calculation for an 8×H100 node
Abstract numbers do not convince a committee; a three-year model with declared line items does. Take a sovereign 8×H100 node in Spain and compare it, for the same work, with AWS and with a European cloud. Declared assumptions: amortisation over 3 years, energy at a solar PPA of 32.5 €/MWh (with grid backup), PUE 1.3, and two utilisation scenarios (50 % and 80 %).
On-premise (own node), annual line items:
| Line item | Annual value | Note |
|---|---|---|
| Amortised capex (node ~270,000 € / 3 years) | ~90,000 € | 8×H100 server + network + storage |
| Energy (≈10.4 kW × PUE 1.3 × 8760 h) | ~3,850 € (at 32.5 €/MWh) | with solar PPA; at grid tariff, ~12–18 k € |
| Operation, cooling, maintenance | ~25,000 € | prorated staff, support, spares |
| Datacenter space (rack, connectivity) | ~12,000 € | colocation or own datacenter |
| Annual total | ~131,000 € | independent of utilisation |
At a fixed 131,000 €/year, the cost per token depends only on how many tokens you generate, that is, on utilisation:
| Utilisation | Useful GPU-hours/year | Cost/1M tokens |
|---|---|---|
| 50 % | ~35,000 | ~2.9 € |
| 65 % | ~45,500 | ~2.2 € |
| 80 % | ~56,000 | ~1.8 € (with cheap grid, ~1.1 €) |
Sovereign European cloud (Scaleway/Lyceum), pay per use: at ~2.2 €/GPU-hour with zero-egress, the cost per token is constant with utilisation (you only pay for what you use): ~1.5–2.2 €/1M tokens depending on model and batching, with no capex and no idle risk.
Hyperscaler (AWS p5), on-demand: at 4.10–6.88 USD/GPU-hour (≈3.8–6.4 €), the cost per token works out at ~2–3.5 €/1M tokens, and on top of that you have to add the egress, quite apart from the fact that for GDPR data the sovereignty axis rules it out already.
The reading of the model is the thesis of the whole of S2: at 50 % utilisation, on-prem (~2.9 €) does not beat the European cloud (~1.8 €); at 80 % with cheap energy (~1.1 €), it beats it comfortably. The crossover sits, as the literature says, around 65–70 %. Investing in on-prem is, fundamentally, a bet that you will sustain high utilisation, and that bet is won with scheduling, not with hardware.
The uncomfortable reality: the utilisation almost nobody reaches
Here is the honest figure missing from the “on-prem is always cheaper” speeches: most inference teams in production operate at 40–65 % GPU utilisation, because of traffic variability and the limits of batching; the 80–90 % assumption that makes on-prem attractive is rarely reached outside batch-only pipelines (Spheron).
This changes the naive conclusion: if your real utilisation is 50 %, on-prem is not cheaper than the cloud, because the capex you pay for the idle GPU eats you. That is why utilisation is not a detail, it is the variable that decides the cost axis, and it connects directly with the FinOps track (the idle of A2, the chargeback of A5) and with scheduling: raising utilisation is what makes on-prem pay. A badly used cluster of your own is more expensive than the cloud; a well scheduled one is far cheaper. The cost question is not “on-prem or cloud?”, it is “can I sustain high utilisation?”.
The hidden costs of the cloud: egress
The cloud has its own small print: egress costs (taking data out of the provider). In the hyperscalers, moving data out or between regions is billed, and in AI workloads with a lot of data movement (datasets, checkpoints, embeddings) it can be a significant line item that does not appear in the GPU-hour price. The advantage of the sovereign European cloud: most of them (Lyceum among others) have adopted the zero-egress model, they do not charge for moving data out or between regions (Lyceum). When comparing, the real hyperscaler cost is GPU-hour + egress + other charges, not just the GPU-hour; ignoring it artificially inflates the hyperscaler’s competitiveness.
An example of the order of magnitude: a platform moving 50 TB/month outbound (datasets, checkpoints, responses served to systems outside the provider) at a typical egress tariff of ~0.08–0.09 €/GB pays ~4,000–4,500 €/month, that is ~50,000 €/year on egress alone, a line item the size of a third of the cost of a node of your own, invisible in the GPU-hour price. On the European cloud with zero-egress that line item is zero; on-prem, internal traffic is not billed either. That is why a fair comparison must model egress according to the real data pattern: for workloads with a lot of outbound movement, it can invert the ranking between hyperscaler and European cloud. The cloud bill is not the GPU-hour; it is the GPU-hour plus everything you move.
To this is added contract and lock-in risk: the hyperscaler’s on-demand GPU tariffs can change, commitment discounts (reserved/savings plans) tie you for 1–3 years, and migrating out, because of egress and coupling to proprietary services, has a real exit cost. On-prem and the European cloud with standard APIs (Kubernetes, S3 compatible) reduce that coupling: the same manifest and the same vLLM run on your cluster or on Scaleway without rewriting. Operational sovereignty, being able to move the workload without rebuilding it, is a value that does not appear in the tariff but weighs on a three-year decision.
The performance axis: the provider does not decide, goodput does
One point that simplifies the synthesis: performance does not depend on the provider, it depends on the hardware and the configuration. An H100 gives the same goodput on your cluster, on Scaleway or on AWS, served with the same vLLM and the same config. What decides performance is the goodput under your SLO (track B), not who hosts the GPU. So, in a comparison at equal hardware, the performance axis neutralises itself: what changes between options is cost, energy and sovereignty. The exception: if a provider gives you access to newer hardware (B200, GB200) ahead of your on-prem buying cycle, then the cloud can win on performance per GPU, a real argument in favour of the cloud for staying at the hardware frontier without capex.
The energy axis: the European and Spanish advantage
Here on-prem (or cloud) in Spain or France has a structural advantage over a hyperscaler in a dirty region. Recalling the data from the energy track:
| Location | Grid carbon (gCO₂/kWh) | Price (indicative) |
|---|---|---|
| France (nuclear) | ~20–60 | low and stable |
| Spain (renewables + gas) | ~150–170 | low, volatile; solar PPA ~32.5 €/MWh |
| Germany | ~363 | high |
| Hyperscaler (region depends on provider) | depends; often not selectable | provider tariff |
The same workload in France emits ~9× less carbon per token than in Germany, and in Spain, with a solar PPA at 32.5 €/MWh (an all-time low), the electricity cost, 30–50 % of the TCO, is low and, under contract, predictable. A sovereign cluster in Spain or France controls where the energy is consumed and with what carbon; a hyperscaler gives you the region it gives you, often with no choice of grid intensity. For CSRD reporting, that selectability is a quantifiable advantage of on-prem and the European cloud.
In concrete numbers: the example 8×H100 node (~10.4 kW × PUE 1.3 ≈ 118,000 kWh/year) emits, depending on the grid, ~2.4 t CO₂/year in France (~20 gCO₂/kWh) against ~43 t CO₂/year in Germany (~363 gCO₂/kWh), the same machine, the same work, ~18× the difference in reportable footprint just by choosing the location. That decision, which a hyperscaler in an imposed region does not let you take, is exactly what on-prem and the sovereign European cloud put in your hands. The energy axis is not an environmental detail: it is cost (the price of the kWh), compliance (CSRD) and sovereignty (control of location) all at once.
The sovereignty axis: the one that does not depend on utilisation
And here is the axis that invalidates the cheapest option if the data is sensitive. US hyperscalers are subject to the US CLOUD Act: US authorities can demand data held by a US company even if it sits in a European datacenter. For data subject to the GDPR, that is a compliance risk. Sovereign European clouds operate under EU/EFTA jurisdiction, providing data residency and GDPR compliance, and are exempt from the US CLOUD Act (Lyceum · sovereign providers). On-prem of your own is the maximum degree of sovereignty: the data does not leave your cluster.
The key difference from the other axes: sovereignty does not depend on utilisation or on volume. However much a hyperscaler cheapens the GPU-hour, for GDPR data it is not an option, the jurisdiction risk is not offset by price. It links with the ENS × ISO 42001 × EU AI Act controls and the EU AI Act mapping: compliance is a hard constraint, not an axis to optimise.
The four instruments that turn sovereignty into a concrete constraint rather than a slogan:
| Instrument | What it requires | Implication for the architecture |
|---|---|---|
| US CLOUD Act | gives the US access to data held by US companies, wherever it is | a US hyperscaler does not guarantee jurisdictional residency even if the datacenter is in the EU |
| GDPR | residency and processing of personal data under EU law | requires an EU/EFTA provider or your own iron for personal data |
| EU AI Act | traceability, risk management and records for AI systems | favours the full control of the stack (logs, datasets, models) that on-prem gives |
| CSRD | verifiable reporting of environmental footprint | the selectable energy (clean grid, PPA) of on-prem and the European cloud is auditable |
The operational conclusion: for a European entity processing personal data or deploying high-risk AI, three of the four technical axes can favour the hyperscaler and it can still lose, because the fourth axis, sovereignty, acts as a prior filter. That is why S2 orders the decision like this: first the sovereignty filter (which rules out the hyperscaler for GDPR data), then the optimisation of cost, performance and energy among the options that pass the filter (sovereign on-prem and European cloud).
The scorecard: the three options scored
Crossing the four axes for the three realistic options of a European platform (order-of-magnitude, illustrative figures):
| Option | Cost/1M tok | Break-even | Energy/carbon | Sovereignty |
|---|---|---|---|---|
| Sovereign on-prem (ES/FR) | ~1.1 € (high util.) / ~3 € (low) | >65–70 % util. | controllable (clean grid, PPA) | total (EU) |
| Sovereign European cloud | ~1.5–2.2 € | no capex, pay per use | EU, zero-egress | high (EU) |
| Hyperscaler (US) | ~2–3.5 € + egress | no capex | imposed region | not EU (CLOUD Act) |
The reading of the scorecard: for data subject to the GDPR, the US hyperscaler is ruled out by the sovereignty axis, however competitive its tariff. The real decision comes down to sovereign on-prem vs sovereign European cloud, and there it is decided by utilisation and volume.
When each option wins
The honest recommendation, by scenario:
| Scenario | Winning option | Why |
|---|---|---|
| High, sustained volume (util. >65–70 %), GDPR data | Sovereign on-prem | lowest cost/token + total sovereignty |
| Variable or growing volume, GDPR data | Sovereign European cloud | sovereignty without capex/idle risk |
| Low volume or sporadic peak | European cloud (per use) | you do not amortise the capex |
| No sovereignty requirement, hardware frontier | Hyperscaler | access to new hardware without capex |
| Hybrid (base + peak) | On-prem + European cloud (burst) | cheap base of your own, sovereign elastic peak |
On-prem makes sense when there is very high, predictable utilisation (80 %+), strict sovereignty requirements, or a hyperscaler contract that works out expensive (Spheron). For sovereign platforms with a sustained base load, the winning pattern is usually the hybrid: sovereign on-prem for the high-utilisation base (where the cost per token is unbeatable) and sovereign European cloud for the peak and for growth (elastic, no capex, keeping EU jurisdiction). The best of both without giving up sovereignty.
Sizing the hybrid: how much on iron, how much on cloud
The hybrid is not “a bit of each”; it is sized with one figure: the base load percentile. The rule is to put on on-prem the load that is almost always present (the load that keeps the GPU at 75–85 %) and to send to the European cloud only the peaks that, if covered with iron, would leave GPUs idle most of the time. An example with a realistic traffic profile:
| Load band | % of the time | Where to serve | Why |
|---|---|---|---|
| Base (p0–p70) | always | on-prem (1 node 8×H100 at ~80 %) | minimum cost/token, high util. guaranteed |
| Middle (p70–p95) | daily peak hours | on-prem if it fits, otherwise European cloud | elasticity without idle capex |
| Peak (p95–p100) | sporadic | sovereign European cloud (burst) | absurd to buy iron for a rare peak |
With this split, the base amortises your own node at high utilisation (~1.1–1.8 €/1M tokens) and the peak is paid per use with no idle penalty (~1.5–2.2 €/1M tokens), all under EU jurisdiction. The expensive mistake is the opposite: sizing the on-prem for the peak, in which case the GPU spends most of its time idle at 30–40 %, the cost per token shoots up above 3 € and the cloud would have been cheaper. You size the iron for the base, not for the peak; the peak is exactly what the cloud does well. This principle connects with the capacity planning and the scheduling (Kueue/Volcano) of the series: the hybrid only works if the scheduler fills your own node before overflowing to the cloud.
Assumptions and sensitivity
The whole comparison hangs on assumptions that have to be declared, because moving them moves the conclusion:
| Assumption | If it rises | Effect |
|---|---|---|
| Utilisation | 50 % → 80 % | on-prem goes from losing to clearly winning |
| Energy price | expensive region → France/PPA | lowers on-prem TCO and carbon |
| Amortisation period | 24 → 36 months | lowers the on-prem cost per hour |
| Volume | < 2M tok/day → much more | crosses the break-even towards on-prem |
| Egress (hyperscaler) | low → high | makes the hyperscaler dearer than the European cloud |
The rule: no on-prem vs cloud comparison is valid without fixing these assumptions. One that says “on-prem is 3× cheaper” without declaring the assumed utilisation is propaganda; one that fixes utilisation, energy price, period and volume is a data point. The dossier must present the case with explicit assumptions and a sensitivity analysis, which is what makes it defensible before a committee that questions them.
Decision checklist
To take S2 from theory to decision, the questions that order the choice, in order:
- Is the data subject to the GDPR, or is the system high-risk under the EU AI Act? If so, the US hyperscaler is ruled out on sovereignty; you choose between on-prem and European cloud. If not, the hyperscaler enters the cost comparison.
- Can I sustain utilisation above 65–70 % on the base load? If so, on-prem wins on cost for that base. If not, the European cloud avoids paying capex for idle GPUs.
- Does the traffic profile have marked peaks? If so, hybrid: base on iron, peak on European cloud. Size the iron for the base, never for the peak.
- How much data do I take out of the provider each month? Model the egress; with a lot of movement, the zero-egress of the European cloud, or on-prem, win clearly.
- Which electricity grid and at what price? France or Spain with a PPA lower TCO and carbon; include it in the model and in the CSRD report.
- Have I fixed utilisation, energy, period and volume in writing? Without those four declared assumptions, the number is not defensible.
Whoever answers these six questions with data, not with intuition, has the case built. The series' recommendation for a sovereign European platform with a sustained base load is stable: sovereign on-prem for the high-utilisation base plus sovereign European cloud for the peak, with the hyperscaler reserved only for workloads with no sovereignty requirement where frontier hardware is needed without capex.
Limits and traps (data-driven)
- Unrealistic assumed utilisation. The 80–90 % that makes on-prem win rarely happens in production (40–65 % typical). Model your real utilisation, not the ideal one.
- Comparing only the GPU-hour. TCO includes energy, operation, cooling, egress (cloud) and capex (on-prem). Compare totals with the same assumptions.
- Ignoring sovereignty. For GDPR data, the sovereignty axis rules out the hyperscaler before cost does; it is not negotiable with price.
- Forgetting the hybrid. It is not “all on-prem or all cloud”; the base-plus-peak pattern usually dominates.
- Data in USD. US comparisons are in dollars and with dirty regions; convert them to euros and to your region’s grid (Spain/France) for your case.
The synthesis of S2, in one sentence: for sovereign European data, the decision is not on-prem vs cloud in the abstract, but sovereign on-prem (high utilisation) plus sovereign European cloud (peak) against a hyperscaler that the sovereignty axis rules out, and utilisation is the variable that splits the base between the first two. The rest of the series gives the numbers for each axis; this one crosses them into the recommendation. The next synthesis article (S3) sizes the investment; this one decides the architecture.
See also
- Cloud GPU: price comparison, commitment and sovereign neoclouds — the on-demand, spot and reserved prices of the European cloud providers that appear as an alternative in this analysis, with updated 2026 data.
- TCO of the on-premise GPU cluster: amortisation, energy and infrastructure — the full breakdown of on-premise TCO: server CAPEX, amortisation, energy, network and staff, with the spreadsheet that gives the real €/GPU-hour.
Sources
- Spheron · LLM Inference On-Premise vs GPU Cloud: 2026 Cost and Break-Even — https://www.spheron.network/blog/llm-inference-on-premise-vs-cloud/
- Lenovo Press · On-Premise vs Cloud: Generative AI TCO (2026) — https://lenovopress.lenovo.com/lp2368-on-premise-vs-cloud-generative-ai-total-cost-of-ownership-2026-edition
- Lyceum · EU Sovereign Inference Platform Comparison (2026) — https://lyceum.technology/magazine/eu-sovereign-inference-platform-comparison/
- Lyceum · Sovereign Cloud Providers 2026 — https://lyceum.technology/magazine/sovereign-cloud-providers-2026/
- Scaleway · H100 GPU instance (precio €, soberanía UE) — https://www.scaleway.com/en/h100/
- Nerd Level Tech · GPU Cloud TCO 2026: hidden fees, egress costs — https://nerdleveltech.com/gpu-cloud-comparison-2026-the-real-cost-of-ai-compute