Datacenter physical infrastructure (4/4): the UPS and the ride-through

Contents

We close the series with the component we have been promising in every article: the UPS, the uninterruptible power supply. It is the sprinter of the power chain. While the generator is the long-distance runner that sustains the facility for hours and the switching is the referee that decides the source, the UPS covers the most critical instant of all: the seconds that pass between the grid flickering and the generator taking over. Without it, none of the earlier pieces is any use, because the outage would already have brought the load down before the generator managed to start.

The ten-second bridge

Let us recall the sequence we saw in the first article. When the grid goes away, the generator has to start, reach its rated speed and voltage and accept the load; NFPA 110, for Level 1 and Type 10 systems, gives it up to 10 seconds for all of that, starting cold, with no preheating. During those 10 seconds, and the few more that the switching takes to transfer, the load cannot be left without power. What sustains it is the UPS.

That is its primary function: the ride-through, the bridge. But it is not the only one. An online UPS also acts as a permanent filter: it isolates the load from grid anomalies (voltage dips, overvoltages, transients, harmonics, frequency variations) and delivers a clean, regenerated waveform at all times. It conditions as much as it backs up. That is why the UPS is the sprinter and the guardian of the quality of the power that reaches the servers.

It is worth underlining that the vast majority of UPS “interventions” are not total grid outages but micro-incidents: voltage dips of a few cycles caused by a switching operation on the grid, the start-up of a large neighbouring load or a distant fault that the utility clears in milliseconds. Those disturbances, too brief for the generator to start, are precisely the ones the UPS absorbs without anyone noticing, and they are hundreds of times more frequent than a real blackout. The UPS therefore works every day even though the grid has never “gone away”: its value lies not only in the dramatic outage, but in the constant trickle of imperfections it filters out.

The consequence of this narrow role is counter-intuitive: the UPS has very little autonomy on purpose. Since it only has to cover the start-up of the generator, its batteries are sized for 5 to 15 minutes, not for hours. More would be expensive and useless: the long-distance runner is the generator, and asking the sprinter to run a marathon wastes money and space on batteries that will never be needed.

The topologies: VFI, VI and VFD

IEC 62040-3 classifies static UPS units by how much they isolate the output from anomalies at the input, and the three classes are worth knowing:

VFI (Voltage and Frequency Independent), known as double-conversion online, is the mission-critical standard. Power always passes through two conversions, a rectifier from AC to DC and an inverter from DC to AC, so that the output is entirely independent of whatever happens at the input: the load lives on a waveform built anew. It is the Tier III and IV topology.

VI (Voltage Independent), or line-interactive, regulates voltage variations with a transformer or an AVR, but lets the input frequency through. It is more efficient and cheaper, suitable for less demanding loads or edge environments.

VFD (Voltage and Frequency Dependent), or standby/offline, regulates nothing in normal operation: the load goes straight to the grid and the UPS only steps in when there is a failure. It is the one in a home computer, not in a datacenter.

There is also delta conversion, an online topology that routes a good part of the power directly from input to output and processes only the difference, gaining efficiency in steady state while keeping the behaviour of an online unit when a fault occurs.

The table summarises when to use each one:

TopologyIEC 62040-3 classIsolationTypical efficiencyUse
Double-conversion onlineVFITotal (voltage and frequency)97-98.1 %Mission critical, Tier III/IV
Delta conversionVFITotal under fault~98 %Mission critical with an efficiency focus
Line-interactiveVIVoltage onlyHighMedium loads, edge
Standby / offlineVFDNone in normal operationVery highNot suitable for a datacenter

The great historic reproach against double conversion was its efficiency: processing all the energy twice costs losses. But that argument has aged. Equipment from 2025 reaches figures that were once unthinkable: the Riello Multi Power2, with silicon carbide semiconductors, declares up to 98.1 % in double conversion; the Vertiv Liebert APM2, up to 97.5 %; and Schneider’s modular range, around 99 % with its high-efficiency modes. The double-conversion penalty, put at around 92 % in older documentation, is history today.

Eco-mode: the last point of efficiency

Even so, there is one more step. In eco-mode (or high-efficiency mode), the UPS feeds the load directly through the bypass, without going through the rectifier or the inverter in normal operation, and reserves double conversion for when the grid fails. Efficiency then rises to 98-99 %. In a 2N installation, where there is twice the equipment, that couple of efficiency points translates into a very real saving in energy, and in PUE.

The classic problem with eco-mode is the instant of the fault: on detecting the anomaly, the UPS has to transfer from bypass to double conversion, and that change takes between 1 and 16 milliseconds, during which the load is momentarily exposed to the faulty grid. It is the same server tolerance window we saw in the article on switching, and that is why the risk is real though bounded: a typical server rides out those milliseconds, but an unlucky transient, an overcurrent when voltage is restored after a dip, can trip protections. That is why mission-critical operators were historically reluctant. The manufacturers’ answer has been advanced eco-mode, multi-mode, predictive, with commercial names such as Schneider’s eConversion or Vertiv’s Dynamic Online, which keeps the Class 1 performance classification (the same as double conversion) while removing a good part of the transfer risk. Operating-cost pressure is pushing its adoption, previously unthinkable on critical loads.

The flywheel: backup without batteries

Not all ride-through is done with batteries. The flywheel UPS stores energy kinetically: a rotor spins at high speed, from 3,000 to 10,000 revolutions per minute, and its inertia delivers power during the interruption. Its autonomy is 10 to 30 seconds, which seems little until we remember that those are, exactly, the seconds the generator takes to start. To act as a bridge to the generator, it is more than enough.

Its advantages speak for themselves: with no batteries, there is no battery maintenance, no periodic replacements, no thermal management of a chemical bank and none of the associated risk; it takes up far less space and tolerates a wide temperature range, with an efficiency that reaches 98 %. Its limit is the short autonomy, so it is often combined with batteries in hybrid systems: the flywheel absorbs the short, frequent events (the most numerous ones) and the battery provides the autonomy for the few outages that drag on. This combination, as we shall see, fits the pulsing profile of AI loads particularly well.

That suitability for AI is no accident. Most grid quality events are extremely brief, voltage dips of a few cycles, and require no autonomy, only an instant response; a flywheel (or, along the same lines, a supercapacitor) handles them without spending a single battery cycle. Since batteries suffer under frequent cycling, discharging the flywheel for the short events and reserving the battery for the long outages extends the life of the chemical bank and reduces its maintenance. The Uptime Institute itself points out that manufacturers are researching flywheels, supercapacitors and advanced chemistries precisely in order to better tolerate the frequent, short, high-power discharges and recharges that AI loads impose. In the AI era, the chain’s sprinter also has to be a sprinter that recovers fast.

Batteries: lead gives way to lithium

Where there are batteries, the great transition of 2025-2026 is the one from lead to lithium. VRLA lead-acid batteries, the standard for decades, are giving ground to lithium iron phosphate (LFP) at a notable speed: lithium already takes around 40 % of datacenter backup installations, and up to 55 % among the hyperscalers.

The reasons are compelling. LFP lithium tolerates a wider thermal range, withstands many more cycles, of the order of 3,000 against the 200-500 of lead, and that extends its service life to 12-15 years against the 3-6 of VRLA. It takes up less space, recharges faster and, despite its higher initial cost, offers a ten-year total cost of ownership around 39 % lower. LFP chemistry is preferred precisely for its thermal stability and its resistance to thermal runaway, and it always comes accompanied by a battery management system (BMS) watching over it. As an intermediate option, thin plate pure lead (TPPL) persists, which recharges faster than classic VRLA and reaches long lifetimes, attractive in certain cost profiles and at low temperatures.

There is one figure that connects batteries directly with our world: in Schneider’s tests with real hyperscaler profiles, above 110 % load lead batteries have difficulty with high-frequency changes, while lithium handles them without trouble. Put another way, AI’s oscillating profile does not merely favour lithium: it almost demands it.

The move to lithium is not free of considerations. The energy density that makes it attractive is also what obliges it to be treated with respect: although LFP chemistry is much more stable than the NMC of electric vehicles and resists thermal runaway better, a lithium battery bank of several MWh concentrates a great deal of energy in a small space, which has led to specific fire detection and suppression regulations and to careful design of the room. The BMS is not an accessory: it watches every cell, balances the charge and disconnects on an anomaly, and its reliability is part of the reliability of the UPS. Against this, the appeal of thin plate pure lead (TPPL) as an intermediate step is that it keeps the operational and safety familiarity of lead while cutting back its worst defects. The choice between lithium and lead is, ultimately, not only about cost per kWh: it is about total cost, about space, about safety and about how it fits with the load profile.

Sizing: little time, unity power factor

Sizing a UPS has two peculiarities. The first, already mentioned, is the short autonomy: since only the start-up of the generator (some 30 seconds) plus its stabilisation (a couple of minutes) has to be covered, a minimum of about 5 minutes is enough, and 10-15 are recommended for margin. Battery autonomy is calculated, in essence, as the stored energy divided by the power demanded:

$$t_{\text{runtime}} = \frac{E_{\text{battery}}}{P_{\text{load}}}$$

The second peculiarity is the power factor. As we saw in the first article, apparent and active power are related by the power factor, and older UPS units, specified at 0.8, “stranded” up to a fifth of their rated capacity. Modern datacenter UPS units operate at unity power factor: they deliver all their kVA as useful kW. It is a quiet but important improvement when it comes to calculating how much real load a unit supports. And UPS sizing does not end with itself: the generator, upstream, has to be capable of feeding the UPS load plus the recharging of its batteries, another piece of coordination that the first article already anticipated.

An example lands the calculation. To sustain a load of 500 kW for 12 minutes, of the order of 100 kWh of useful energy has to be stored in the batteries (500 kW times 0.2 hours), plus the margins for ageing, depth of discharge and inverter efficiency, which in practice push that figure up considerably. It seems a lot for “only” 12 minutes, but remember that the short autonomy is deliberate: if the generator has not started in that time, the problem is no longer the UPS but the generator, and that is why both are tested together. The temptation to “buy more batteries just in case” is usually a bad deal: it costs more, takes up space and does not solve the scenario that really matters, which is getting the long-distance runner started.

The modular architecture

The way a UPS is built has changed as much as its chemistry. The modular UPS, with hot-swap power blocks, has displaced the monolithic one because it allows capacity and redundancy to be added incrementally, paying as you grow, and a module to be serviced or replaced without shutting the system down (Schneider calls it Live Swap). Internally it offers N+1 redundancy: if one module fails, the others take on its load without the system going down.

At whole-system level, the configurations reproduce the redundancy nomenclature we saw in the first article, applied to the UPS: from capacity (no redundancy) up to 2N (two independent systems, each capable of 100 %, equivalent to Tier IV availability), by way of parallel redundant (N+1), isolated redundant or catcher (a standby UPS that picks up the load if another fails) and distributed redundant. In the dual-path architecture, each path A and B has its own UPS, and the STS units we saw in the previous article transfer between their outputs for the loads that need it. All the pieces of the series fit together here.

It is worth understanding the nuance of each configuration, because it defines cost and behaviour under a fault. Parallel redundant (N+1) shares the load among several UPS units on a common bus; if one goes down, the others absorb its share, but they all serve the same load, so that a fault on the common bus affects them all. Isolated redundant or catcher keeps a standby UPS waiting that “catches” the load of any main system that fails, an efficient solution when there are several systems to protect with a single backup. 2N is the most robust and the most expensive: two systems that share absolutely nothing, each sized at 100 %, so that not even a bus fault can affect both. The choice, like all redundancy, is a balance between the cost of duplicating and the cost of unavailability, the same calculation that runs through this entire series.

The market follows this evolution. The recent launches, the new-generation Eaton 93PM with lithium, Schneider’s Galaxy VL and VXL ranges, the Riello Multi Power2 with silicon carbide, share three traits: modularity, lithium and efficiencies above 97 % in double conversion. The datacenter UPS market moves, according to analysts, between USD 4 billion and USD 9 billion in 2025 with high-single-digit annual growth, pushed precisely by AI demand.

The AI challenge: loads that pulse

And so we reach the challenge that is rewriting UPS design, and that sums up why this series matters for an inference factory. AI loads are not constant: they pulse. In a training cluster, thousands of GPUs draw power almost in unison, following the steps of the computation, and the power oscillates every one or two seconds. The Uptime Institute’s analysis from late 2025 puts numbers on it: in the worst case, the difference between the minimum and the maximum consumption can exceed 100 % at system level, the load doubles in milliseconds, and with the most recent hardware the swings reach 150 %, with overshoots of more than 10 % above specification. In GB200 NVL72 racks, sudden rises from 60-70 kW to more than 150 kW have been observed, above even the 132 kW maximum specification.

The UPS is the first piece of equipment to take that stress, and if it is not designed for it a load drop can occur. Schneider tested its Galaxy range with real hyperscaler profiles at 40 °C and made it withstand 1 to 100 % with no drops, up to 10 minutes of overload at 125 % and one minute at 150 %. Uptime, for its part, recommends specific mitigations: mixing AI loads with more stable loads and sharing generators to damp the fluctuations; choosing UPS configurations with more capacitance and greater redundancy, even N+2, to absorb the peaks; and leaning on the accelerator’s own power-smoothing tools, such as the power smoothing NVIDIA builds into Blackwell. And manufacturers are researching supercapacitors, advanced chemistries and flywheels precisely because they tolerate the frequent, short, high-power discharges and recharges that characterise AI better than a conventional battery does.

It is worth pausing on why the UPS is the first one affected. When thousands of GPUs go from their minimum draw to their maximum in a millisecond, someone has to cover that sudden demand instantly, and that someone is the UPS, because neither the grid nor the generator reacts that fast. If the UPS is not designed with enough capacitance and margin, its output voltage sags under the step and can drag the load down with it. That is why the Uptime Institute recommends, for AI loads, configurations with more redundancy, even N+2 instead of N+1, not so much for fault tolerance as for the capacity to absorb the peaks without repeatedly overloading the modules. It is a change of mindset: UPS redundancy stops being sized only against failure and starts being sized against the dynamics of the load as well. And upstream, the generator, which reacts more slowly still, is the link most stressed by these oscillations, which leads to recommendations such as mixing AI loads with stable loads and sharing generation to damp them.

Density adds another turn of the screw: with racks scaling from 100 kW towards 600 kW and the megawatt under discussion, traditional low-voltage UPS units are reaching their ceiling, and medium-voltage UPS architectures and centralised UPS systems with grid-scale batteries (BESS) are emerging. That convergence between the UPS and grid storage also opens up a new role: the UPS as a resource that provides services to the electricity grid, peak shaving, frequency regulation, reserve, blurring the boundary between datacenter backup and the grid itself. A battery bank that spends 99.9 % of its time waiting for an outage that almost never comes is an expensive, immobile asset; turning it into a resource that shaves consumption peaks or sells reserve to the grid while it is not needed for backup changes its economic equation. Manufacturers such as Vertiv already offer utility-grade storage systems integrated with the datacenter for this dual use. The challenge is that the UPS must never compromise its primary mission, being available for the outage, in order to provide services to the grid; the design has to guarantee that the energy needed for the ride-through is always held in reserve. It is worth not losing sight, however, of the systemic risks this introduces: during a minor grid fault in Virginia in 2025, several datacenters disconnected at the same time and caused a 1.5 GW drop, a reminder that, at this scale, datacenter physical infrastructure is already critical grid infrastructure.

Takeaways, and the close of the series

The UPS is the sprinter that covers the ten-second bridge until the generator starts, and at the same time the guardian that delivers clean power at all times. It is sized with little autonomy on purpose, it is built increasingly modular and with lithium, and it faces a new and formidable challenge: AI’s pulsing loads, which turn it into the first line of defence against oscillations that double the load in milliseconds.

And as in each of the previous articles, the conclusion returns to operations: a UPS is chosen well on its topology and its chemistry, but it earns its keep when, the outage having arrived, it transfers without a flicker. That demands testing the batteries, discharging them for real, measuring their actual autonomy, not trusting the nominal age; rehearsing the transfer coordinated with the generator; and watching every cell with the BMS. A battery bank that nobody has discharged under load is, like the generator that nobody has tested, an unverified promise. The whole power chain, from the utility feed to the chip, is only as strong as its least-tested link.

With this we close the physical infrastructure series. We began with the complete power chain, from the utility feed to the rack; we went on to the generators that sustain the facility when the grid goes away; we looked at the switching that arbitrates between sources; and we finish with the UPS that covers the critical instant. Four links of one and the same chain whose conclusion is always the same: in an inference factory, the GPUs are the expensive, visible asset, but it is the electrical infrastructure, silent, without glamour, measured in megawatts and milliseconds, that decides whether that asset works or shuts down. Designing it well, sizing it with slack for AI’s transients and, above all, testing it relentlessly, is what separates a factory that meets its SLA from one that learns the hard way on the day of the first outage.

See also

Sources