Full TCO of an on-premise GPU cluster: from capex to the all-in €/GPU-hour and the break-even against cloud

Contents

Notation: amounts in N € or N USD (source denominated in dollars); decimal point; comma as thousands separator. The dollar sign is not used (it is the formula delimiter). Data centred on Europe/Spain. Generic example hardware: a cluster of N nodes, each with 4×H100 SXM5 80 GB.

TL;DR

A 4×H100 SXM5 node costs between 150,000 USD and 200,000 USD in total capex (GPUs + server + network + storage + prorated rack). Amortised over 3 years with European opex (energy at ~0.116 €/kWh industrial, PUE 1.54 average or 1.2 with liquid, 0.3 FTE of staff), the all-in cost ranges from 3.10 USD/GPU-hour (100 % utilisation) to 6.20 USD/GPU-hour (50 % utilisation). The break-even against AWS p5 on-demand (~6.88 USD/GPU-hour) is crossed at around 70 % utilisation; against a neocloud 3-year reserved (~1.49–2.10 USD/GPU-hour), on-prem never closes the gap in that scenario. Utilisation is the variable that decides the cost axis, not the price of the hardware.


The model: declared assumptions

Every calculation below starts from these assumptions. Changing any of them moves the conclusion; the sensitivity section quantifies by how much.

ParameterBase valueSensitivity range
Node4×H100 SXM5 80 GB (HGX baseboard)
ClusterN nodes (per-node model; scales linearly)1–32 nodes
Capex amortisation3 years (straight line)3–5 years
GPU utilisation70 %30 %–100 %
Energy price0.116 €/kWh (industrial Spain, Sept. 2025)0.06–0.20 €/kWh
PUE1.54 (global average, Uptime Institute 2025)1.15–1.80
Exchange rate1 USD = 0.93 € (reference Jun. 2026)

Energy source: GlobalPetrolPrices · Spain Business Electricity, Sept. 2025. PUE source: Uptime Institute Global Data Center Survey 2025 — the global average PUE has been stuck at 1.54 for the sixth consecutive year; hyperscalers 1.10–1.15; colocation/enterprise 1.58–1.80; facilities less than 5 years old, 1.48. PUE 1.2 is achievable with direct-to-chip liquid cooling.


Capex breakdown per 4×H100 SXM5 node

GPUs

ComponentUnit price (USD)QuantitySubtotal (USD)Source and date
H100 SXM5 80 GB (card)30,000–40,0004120,000–160,000GMI Cloud, Apr. 2026 · Introl, Apr. 2026

The range reflects market variability and volume discounts (5–15 % for orders above 50 units). H100 SXM5 cards require NVIDIA’s HGX baseboard; they are not sold loose for direct installation in standard servers.

Server / HGX baseboard

A complete 4×H100 SXM node uses NVIDIA’s HGX H100 4-GPU baseboard plus a compatible host server. Reference models: Supermicro SYS-421GU-TNXR (4U, dual Intel Xeon 4th Gen, HGX H100 4-GPU) and its Dell equivalent.

ComponentEstimated cost (USD)Note
Server chassis + CPU (2× Xeon) + RAM (512 GB DDR5) + redundant PSU18,000–25,000Based on a Supermicro SYS-821GE bare without GPUs at ~24,806 USD (xicomputer.com, Jun. 2026); scaled to the 4-GPU
HGX H100 4-GPU baseboardincluded in the GPU priceNVIDIA HGX platform; no separate public price
Inter-GPU NVLink (inside the node)included in the baseboard4 GPUs connected by NVLink 4.0 on the HGX baseboard

Marketing claim (no independent verification): Supermicro announces datacenter energy cost reductions of up to 40 % with liquid cooling in its HGX H100 servers (Supermicro press release).

InfiniBand NDR network

For a multi-node cluster with tensor parallelism across nodes, the GPU-to-GPU network is critical. NDR InfiniBand (400 Gb/s per port) is the de facto standard for HGX clusters.

ComponentEstimated cost per node (USD)Source / Note
NVIDIA Quantum-2 NDR 400G switch (64 ports, prorated across N nodes)2,000–4,000Switch ~35,000 USD (Introl, Apr. 2026); at 16 nodes, ~2,200 USD/node
InfiniBand NDR cables/transceivers (4 ports per node × ~1,000 USD/port)4,000Estimate based on ~1,000 USD per optical transceiver (Introl, Apr. 2026)
InfiniBand network (prorated per 4-GPU node)~6,000–8,000

For inference serving inside a single node (4 GPUs with NVLink), the inter-node network is less critical than for multi-node training. For disaggregated prefill-decode workloads spread across nodes, InfiniBand NDR is necessary.

NVMe storage

ComponentEstimated cost (USD)Note
Local NVMe (4 TB × 2 U.2/E1.S drives, working datasets and checkpoints)2,000–4,000~500–1,000 USD/TB enterprise NVMe, 2025
Shared object storage (NAS/MinIO, prorated per node)2,000–5,000Varies with the total capacity of the cluster
Total storage per node~4,000–9,000

Introl models 50 TB per GPU for effective operation in training clusters (Introl, Apr. 2026); for pure inference the requirement is significantly lower (model weights + logs).

Rack, PDU and datacenter connectivity

ComponentEstimated cost per node (USD/year)Source
Rack colocation (high density, 10–15 kW per node)5,000–12,000/yearEncoradvisors · Colocation Pricing 2026: high density 3,000–6,000 USD/month per rack; at 2 nodes per rack, ~1,500–3,000 USD/month per node = 18,000–36,000 USD/year in tier-1; lower in Spain
Rack PDU, electrical cabling (prorated)500–1,000 per node (amortised capex)Inside the colocation line item or in an own datacenter

Colocation in Spain/Europe is structurally cheaper than in US tier-1 markets (New York, Silicon Valley). For an own datacenter, replace it with the cost of your own space plus amortisation of the electrical and cooling infrastructure.

Capex summary per 4×H100 SXM5 node

Line itemRange (USD)Midpoint
GPUs (4× H100 SXM5)120,000–160,000140,000
Server chassis + CPU/RAM/PSU18,000–25,00021,500
InfiniBand NDR network (prorated)6,000–8,0007,000
NVMe + object storage4,000–9,0006,500
PDU/rack/other (capex)2,000–5,0003,500
Total capex per node150,000–207,000178,500

Sources: GMI Cloud (Apr. 2026), Introl (Apr. 2026), Spheron (Apr. 2026), xicomputer.com (Jun. 2026).


Opex breakdown per 4×H100 SXM5 node (annual)

Energy

A 4×H100 SXM5 node at full load draws approximately:

$$P_{\text{node}} = 4 \times 700\,\text{W (TDP H100 SXM5)} + 800\,\text{W (server)} \approx 3.6\,\text{kW (IT)}$$

The total datacenter power includes the cooling overhead, expressed by the PUE:

$$P_{\text{total}} = P_{\text{IT}} \times \text{PUE}$$ $$\text{annual energy cost} = P_{\text{IT}} \times \text{PUE} \times 8{,}760\,\text{h} \times \text{kWh price}$$

With the base values (PUE 1.54; 0.116 €/kWh):

$$\text{energy/year} = 3.6\,\text{kW} \times 1.54 \times 8{,}760\,\text{h} \times 0.116\,\text{EUR/kWh} \approx 5{,}475\,\text{EUR}$$

With a Spanish solar PPA (Q3 2025 reference price: ~34 €/MWh = 0.034 €/kWh according to PV Tech, Oct. 2025):

$$\text{energy/year (solar PPA)} = 3.6 \times 1.54 \times 8{,}760 \times 0.034 \approx 1{,}604\,\text{EUR}$$
Energy scenarioPrice (€/kWh)Energy cost/year per 4-GPU node
Spanish solar PPA (Q3 2025)0.034~1,604 €
Industrial Spain (Sept. 2025)0.116~5,475 €
European average (industrial tariff)0.160~7,550 €
Worst case (no PPA, high tariff)0.200~9,437 €

Staff / operations

Staff cost is the most variable line item with cluster size. For a small cluster (2–8 nodes), the rule of thumb is 0.3–0.5 FTE per cluster of GPU infrastructure support (Spheron, Apr. 2026).

Cluster sizeEstimated FTEFTE cost (€/year, Western Europe)Cost per 4-GPU node (€/year)
2–4 nodes0.3 FTE~120,00036,000–18,000
8–16 nodes0.5 FTE~120,0007,500
32+ nodes1–2 FTE~120,0003,750–7,500

Indicative salary reference: a GPU infrastructure engineer with CUDA, InfiniBand and Kubernetes knowledge in Western Europe, ~90,000–140,000 €/year fully loaded. Introl’s figures (Apr. 2026) in USD (~275,000 USD/year for the US) reflect the North American market, which is appreciably higher.

Maintenance and support

Line itemAnnual cost (% of hardware capex)Per 4-GPU node (midpoint)
Vendor maintenance / support5–10 % of capex~7,000–14,000 USD → ~6,500–13,000 €
GPU failure rate (~5 % annual) × replacement cost5 % × 4 GPUs × ~35,000 USD = ~7,000 USD expected~6,500 € (amortised as a provision)
Minor spares (cables, modules)~500–1,000 €

Introl quotes GPU failure rates of 2–3 % per year in small clusters; Google Research documented ~9 % annualised in Meta’s 16,384-GPU H100 cluster (Introl, Apr. 2026). 5 % is used here as a conservative intermediate value.

Depreciation (for accounting purposes)

Straight-line depreciation turns capex into an annual flow comparable to the cost of committed cloud:

$$\text{annual depreciation} = \frac{\text{node capex}}{\text{amortisation years}}$$
Node capex (USD)3-year amortisation (USD/year)5-year amortisation (USD/year)
150,000 (minimum)50,00030,000
178,500 (average)59,50035,700
207,000 (maximum)69,00041,400

H100 hardware depreciates fast: secondary-market analyses put the residual value at 20–40 % of the purchase price after 3 years (Introl, Apr. 2026). The arrival of Blackwell GB200/GB300 accelerates the perceived obsolescence.

Annual opex summary per 4×H100 SXM5 node (base scenario, 8-node cluster)

Line itemBase scenario (€/year)Range
Energy (PUE 1.54; 0.116 €/kWh)5,4751,604–9,437
Staff (0.5 FTE × 8 nodes, prorated)7,5003,750–36,000
Maintenance / support / failures9,0005,000–15,000
Rack colocation (Spain, high density)6,0003,000–15,000
Total opex per node~28,000~13,000–75,000

The extreme range reflects the difference between a well amortised own datacenter with a solar PPA and cheap energy (minimum opex) and tier-1 colocation with market tariffs and senior staff.


Deriving the all-in €/GPU-hour

Formula

$$\text{EUR/GPU-hour all-in} = \frac{\frac{\text{node capex}}{\text{years}} + \text{node annual opex}}{4\,\text{GPUs} \times 8{,}760\,\text{h} \times u}$$

where \(u\) is the average annual utilisation (0 to 1).

See the cost-per-token identity in cost per token and per request for the connection with throughput.

€/GPU-hour table by utilisation and scenario

Average capex (178,500 USD → ~166,000 €), 3-year amortisation → 55,300 €/year.

UtilisationOpex/year (base, €)Total cost/year (€)Useful GPU-hours/year€/GPU-hour
30 %28,00083,30010,5127.93
50 %28,00083,30017,5204.75
70 %28,00083,30024,5283.39
80 %28,00083,30028,0322.97
100 %28,00083,30035,0402.38

Low-opex scenario (solar PPA, own datacenter, large cluster): opex/year ~13,000 €.

UtilisationTotal cost/year (€)€/GPU-hour
50 %68,3003.90
70 %68,3002.78
80 %68,3002.43
100 %68,3001.95

High-opex scenario (market tariff, expensive colocation, small cluster): opex/year ~75,000 €.

UtilisationTotal cost/year (€)€/GPU-hour
50 %130,3007.44
70 %130,3005.31
80 %130,3004.65
100 %130,3003.72

From €/GPU-hour to €/1M tokens

The cost-per-token identity connects the hardware cost with the inference cost:

$$\text{EUR/1M tokens} = \frac{\text{EUR/GPU-hour} \times 10^6}{\text{throughput (tok/s)} \times 3{,}600}$$

For reference throughputs on H100 SXM5 with vLLM (see capacity planning for on-premise inference):

ModelTypical throughput (tok/s per GPU)Source
Llama-3 70B FP8, high batch~2,800Series B benchmarks
Llama-3 8B FP16, medium batch~9,000Series B benchmarks
Mixtral 8×7B, high batch~4,500Series B benchmarks

€/1M tokens table in the base scenario (€/GPU-hour 3.39 at 70 % utilisation):

ModelThroughput (tok/s)€/1M tokens
Llama-3 70B FP82,800~0.336
Llama-3 8B FP169,000~0.105
Mixtral 8×7B4,500~0.209

At 50 % utilisation (€/GPU-hour 4.75):

Model€/1M tokens
Llama-3 70B FP8~0.471
Llama-3 8B FP16~0.147

Occupancy (batching) multiplies the effective throughput and lowers the €/1M tokens without changing the hardware; it is analysed in GPU utilisation as a FinOps lever.


Break-even on-prem vs cloud

The break-even formula

Break-even happens when the total annual on-prem cost equals the annual cloud cost at the same utilisation:

$$\text{annual cloud cost} = \text{cloud GPU-hour price} \times 4\,\text{GPUs} \times 8{,}760\,\text{h} \times u$$ $$\text{break-even}: \quad \frac{\text{capex/year} + \text{opex/year}}{4 \times 8{,}760 \times u} = \text{cloud GPU-hour price}$$

Solving for the break-even utilisation:

$$u^* = \frac{\text{capex/year} + \text{opex/year}}{4 \times 8{,}760 \times \text{cloud GPU-hour price}}$$

Break-even table by cloud mode and on-prem scenario

Base on-prem scenario (capex/year 55,300 €, opex/year 28,000 €, total 83,300 €/year per 4-GPU node):

Cloud reference (price/GPU-hour)USD equiv.Break-even utilisation \(u^*\)Note
Neocloud on-demand (Lambda/Spheron ~2.90 USD)2.90 USD (~2.70 €)>100 %, on-prem does not competeNeocloud on-demand is cheaper even at full utilisation
Neocloud 3-year reserved (CoreWeave ~1.49–2.10 USD)~1.80 USD (~1.67 €)>100 %, impossibleNeocloud reserved beats on-prem in every scenario of this model
AWS p5 on-demand (6.88 USD/GPU-hour)6.88 USD (~6.40 €)~47 %Above 47 %, the average on-prem beats AWS on-demand
AWS p5 3-year reserved (~2.97 USD/GPU-hour)2.97 USD (~2.76 €)>100 %
GCP A3 on-demand (~10.98 USD/GPU-hour)10.98 USD (~10.21 €)~29 %Above 29 %, on-prem beats GCP on-demand
Azure ND H100 v5 on-demand (~12.29 USD/GPU-hour)12.29 USD (~11.43 €)~26 %

Low-opex scenario (total 68,300 €/year):

Cloud referenceBreak-even utilisation
AWS p5 on-demand (6.88 USD ≈ 6.40 €)~38 %
Neocloud on-demand (2.90 USD ≈ 2.70 €)~91 %
Neocloud 3-yr reserved (1.80 USD ≈ 1.67 €)>100 %

High-opex scenario (total 130,300 €/year):

Cloud referenceBreak-even utilisation
AWS p5 on-demand (6.88 USD ≈ 6.40 €)~72 %
GCP A3 on-demand (~10.21 €)~45 %
Azure on-demand (~11.43 €)~41 %

Reading the break-even table

  • Against neoclouds (on-demand or reserved), on-prem TCO does not reach break-even in any scenario of the base model. Neocloud reserved beats on-prem even at 100 % utilisation, because its hourly price is below the all-in cost of your own hardware. This is consistent with the analysis in cloud GPU: on-demand, reserved and spot.
  • Against on-demand hyperscalers (AWS, GCP, Azure), on-prem does have an achievable break-even: around 26–72 % utilisation depending on the scenario. At medium-high utilisation (>70 %), on-prem clearly beats AWS/GCP/Azure on-demand.
  • The variable that moves the break-even most is opex (staff above all), not hardware capex. A well sized cluster in cheap colocation with PPA energy can lower the threshold by 20 percentage points compared with the high scenario.
  • For GDPR data, the break-even against US hyperscalers is skewed: the sovereignty axis rules out US hyperscalers before cost does (see sovereign on-premise vs hyperscalers).
€/GPU-hourutilisation (%) →0306090100on-prem (fixed capex)AWS p5 OD (~6.40 €)GCP OD (~10.21 €)Azure OD (~11.43 €)neocloud OD (~2.70 €)≈47 % (AWS)≈29 % (GCP)

Sensitivity analysis

TCO vs utilisation

The all-in cost per GPU-hour varies inversely with utilisation because capex is fixed:

$$\frac{d(\text{EUR/GPU-hour})}{du} = -\frac{\text{capex/year} + \text{opex/year}}{4 \times 8{,}760 \times u^2} < 0$$

Going from 50 % to 80 % utilisation cuts the €/GPU-hour by \(\frac{4.75 - 2.97}{4.75} \approx 37\,\%\) in the base scenario. That 37 % reduction requires no hardware change, only more efficient scheduling (see GPU utilisation as a FinOps lever).

Utilisation€/GPU-hour (base scenario)Change vs 50 %
30 %7.93+67 %
50 %4.75reference
70 %3.39−29 %
80 %2.97−37 %
100 %2.38−50 %

TCO vs energy price

Energy price (€/kWh)Energy opex/year€/GPU-hour (70 % util.)Change vs base
0.034 (solar PPA)1,604 €3.00−12 %
0.116 (industrial ES, base)5,475 €3.39reference
0.160 (European average)7,550 €3.54+4 %
0.200 (high tariff)9,437 €3.67+8 %

Energy has a moderate impact on total TCO (8–12 % variation between the extremes), because hardware capex dominates. Over a very long amortisation (5 years) with a solar PPA, however, energy drops from 6 % to 1 % of total TCO and the differential is amplified. The price of energy matters more for the carbon footprint (CSRD) than for TCO when capex is dominant.

TCO vs PUE

PUECooling overheadEnergy/year (0.116 €/kWh)€/GPU-hour (70 % util.)
1.15 (liquid cooling, new facilities)+15 %2,166 €3.21
1.20 (liquid, modern datacenter)+20 %2,259 €3.23
1.48 (facilities <5 years, Uptime 2025)+48 %3,490 €3.33
1.54 (global average, Uptime 2025)+54 %3,627 €3.39
1.80 (legacy colocation)+80 %4,260 €3.47

The difference between PUE 1.15 (liquid) and 1.80 (legacy) is barely ~8 % of the €/GPU-hour at 70 % utilisation, because energy is only a fraction of TCO. PUE matters far more for the absolute energy cost and for CSRD reporting than for total TCO when hardware is the dominant component.

TCO vs amortisation years

AmortisationCapex/year (average node, USD)€/GPU-hour (70 % util., base opex scenario)
3 years59,500 USD (~55,300 €)3.39
4 years44,625 USD (~41,500 €)2.99
5 years35,700 USD (~33,200 €)2.72

Stretching the amortisation from 3 to 5 years lowers the €/GPU-hour by ~20 %, assuming the hardware remains competitive and the resale market supports the residual value. With the refresh cycle accelerated by Blackwell GB200/GB300, a 5-year amortisation carries a higher risk of technological obsolescence.

Sensitivity heat map (€/GPU-hour at 70 % utilisation, base scenario)

PUE 1.15PUE 1.54PUE 1.80
3-yr amort., solar PPA (0.034 €)2.722.742.76
3-yr amort., industrial (0.116 €)3.213.393.47
5-yr amort., industrial (0.116 €)2.542.722.80
3-yr amort., high tariff (0.200 €)3.443.673.78

Decision table: cost/control/sovereignty Pareto

The table below crosses the four dimensions with no implicit hierarchy; the ordinal reading depends on each organisation’s constraints.

Option€/GPU-hourInitial capexFull stack controlEU sovereigntyElasticityOperational risk
On-prem (util. >70 %, low opex)2.40–3.00high (150–207k USD/node)totaltotalnonehardware failure, idle
On-prem (util. <50 %, base opex)4.75–7.93hightotaltotalnonecapex with no return
Neocloud 3-year reserved (CoreWeave, Lambda)1.49–2.10 USDnonepartial (API)depends on the providerrigid contractminimal interruption
Neocloud on-demand (Lambda, Spheron)2.49–3.44 USDnonepartialdependstotalno interruption
AWS p5 on-demand6.88 USDnoneminimalNO (CLOUD Act)totalno interruption
AWS p5 3-year reserved~2.97 USDfinancial commitmentminimalNO (CLOUD Act)rigidno interruption
Sovereign EU cloud (Scaleway, Nebius EU)2.15–3.85 USDnonepartialyes (EU)totalno interruption
Hybrid on-prem base + EU cloud peak2.00–3.50 (weighted)mediumhighyes (EU)elastic peakoperational complexity

“EU sovereignty” column: US hyperscalers (AWS, GCP, Azure) are subject to the US CLOUD Act regardless of the datacenter region. Nebius has a Dutch legal entity; CoreWeave is a US company. See the full analysis in sovereign on-premise vs hyperscalers.

“Full stack control” column: on-prem lets you choose the driver version, the kernel, the NCCL configuration, MIG partitioning, and any system parameter. Cloud options offer control at container/pod level, with the hypervisor and firmware opaque.

The cost/sovereignty Pareto frontier for GDPR data excludes US hyperscalers, leaving: on-prem, sovereign EU cloud, and the hybrid. Among those three, the deciding variable is sustained utilisation and the predictability of traffic (see capacity planning for on-premise LLM inference).


Integration with the series FinOps model

The all-in €/GPU-hour of on-prem is the number that feeds the cost allocation pipeline of the series:

  1. Cost-per-token identity (cost per token and per request): engine throughput × €/GPU-hour → €/1M tokens.
  2. Chargeback and showback (chargeback and showback in GPU multitenancy): the all-in €/GPU-hour is the internal price charged to each tenant of the multi-tenant cluster.
  3. Utilisation as a lever (GPU utilisation as FinOps): raising utilisation from 50 % to 80 % cuts the €/GPU-hour by 37 % without changing the hardware, the highest ROI in on-prem FinOps.
  4. Capacity planning (capacity planning for on-premise LLM inference): the number of nodes to buy depends on the percentile of base load you want to cover in iron.
  5. Cloud comparison (cloud GPU: on-demand, reserved and spot): the all-in €/GPU-hour goes head to head with the cloud price in table A7 to compute the break-even.

Sources