Datacenter physical infrastructure (3/4): the switching
Contents
We already have the power chain mapped out and the generators that sustain it when the grid goes away. What is missing is the referee: the component that, at every instant, decides which source (the grid, the generator, path A or path B) feeds each load, and that executes the change from one to the other without the compute noticing. That is the switching, and it is where the battle of the milliseconds is fought.
Because the problem of transferring load between sources is not deciding whether to do it, which is easy, but doing it fast enough that the server power supplies never notice the gap. And there the choice of technology is everything.
The referee of the chain
When a source fails or degrades, someone has to move the load to the alternative source without human intervention. That someone is a transfer switch, and there are two very different families of them that live in different places in the hierarchy.
At the building level, where the choice is between grid and generator, the electromechanical switch reigns: the ATS (Automatic Transfer Switch), or, in large installations, the transfer logic built into the generator paralleling switchgear itself. It is a transfer between sources that are hundreds of metres apart and that may take seconds to become available (the generator has to start), so here switching speed is measured in tens or hundreds of milliseconds.
Close to the critical load, by contrast, downstream of the UPS, between paths A and B, the solid-state switch governs: the STS (Static Transfer Switch). Here both sources are alive and available, and the aim is to transfer so fast that the load does not flinch. The boundary between “utility capacity” and “IT capacity” sits, precisely, around the PDUs and the rack distribution panels, and that is where the STS does its work.
The window that decides everything
To understand why there are two technologies, you have to know the figure that rules: how long a server power supply can go without power before shutting down. The answer, according to the ITIC/CBEMA curves and the IEC 62040-3 standard, is that most enterprise power supplies tolerate a total loss of between 10 and 20 milliseconds before dropping out. That is the window. The whole design of the switching consists of transferring within it, or of having something to cover the gap when you cannot.
That tolerance is not a manufacturer’s whim: the ITIC and CBEMA curves, which define the “envelope” of voltage disturbances an IT device must withstand without failing, capture the hold-up time provided by the capacitors inside the power supply. For a few milliseconds, those capacitors keep delivering energy even though the input has disappeared; past that margin, the internal voltage collapses and the server reboots. All the engineering of switching revolves around not exhausting that small cushion.
A grid cycle at 50 Hz lasts 20 ms; at 60 Hz, about 16.7 ms. A switch that takes a whole cycle is already at the limit. That is why the time of an STS is measured in fractions of a cycle: a quarter of a cycle is
$$t = \frac{1}{4f}$$that is, about 4.17 ms at 60 Hz. Inside the window, with margin. An electromechanical ATS, by contrast, takes much longer, and there lies the fundamental difference between the two.
The ATS: robust and slow
The ATS transfers by physically moving the connection. It uses electromechanical contactors or a motorised mechanism that opens one set of contacts while closing the other, like an electrically driven double-throw switch. That mechanism has a consequence: it is slow. A standard open-transition ATS switches in 60 to 200 milliseconds; a fast one, with a stored-energy mechanism, comes down to 50-100 ms. Eaton documents times of the order of 150 ms, some nine cycles. Any of those figures far exceeds the server’s tolerance window, and that is why an ATS, on its own, cannot feed critical IT load without a UPS behind it to cover the jump.
In exchange, the ATS is robust and cheap: it costs of the order of 60-70 % less than an equivalent STS, and its contactors with arc chutes withstand higher fault currents than the semiconductors of an STS of the same size. That is why it is the king of grid↔generator transfer at building level and of non-critical loads (lighting, HVAC, office IT), where a few hundred milliseconds of gap harm nobody. The reference standard for its construction and testing is UL 1008, which defines, among other things, the short-circuit withstand and closing ratings (WCR) according to the upstream protective device.
The ATS also supports several operating modes designed for different cases: standard open transition, delayed transition (with a programmed delay of half a second to three seconds, intended for motor loads, where reconnecting too fast and out of phase would damage the equipment) and closed transition, which we will see shortly. For an IT load, however, delayed transition is ruled out: that delay is an eternity for a server.
An important variant is the bypass-isolation ATS (maintenance bypass), which allows the switch to be isolated and maintained or tested while the load keeps being fed by an alternative route. In concurrently maintainable facilities, Tier III and IV, this capability is not optional: without it, maintaining the ATS would force an outage.
The STS: fast and electronic
The STS moves nothing. It uses thyristors (SCR), power semiconductors, as switching elements, and transfers the current from one source to another in a quarter of a cycle, 2 to 4 milliseconds, with no moving parts. Its conduction voltage drop is barely 1 or 2 volts and its efficiency is around 99.5-99.9 %. If the two sources are synchronised, it switches in make-before-break mode; if they are not, it does break-before-make in a brief interruption that, depending on the phase angle difference, stays around 4-10 ms, still inside the server’s window.
Its logic is that of a prudent referee: it continuously monitors the magnitude, frequency and phase angle of both sources, and transfers when the active one goes outside the configured thresholds (typically ±10 % of voltage or ±3 Hz), but only if the alternative is healthy. If both sources are out of tolerance, it stays where it is: transferring to a bad source is worse than holding on a degraded one. And after transferring, it waits a configurable delay (from 30 seconds to several minutes) before going back, so as not to enter a back-and-forth cycle.
The STS is listed under UL 1008S, which distinguishes it from the UL 1008 of the ATS, and it exists at several voltage levels: single-phase at 208 V for rack level, three-phase at 208 V for PDU level and three-phase at 480 V for the bus, as well as medium-voltage versions. Its flagship application is giving source redundancy to single-corded equipment: a server or a network device with only one power inlet can, thanks to an STS, receive power from both paths A and B without needing a second power supply. It is the way to put single-corded equipment into a 2N world.
Beyond transfer time, there are specifications an architect must look at when choosing an STS: the configurable retransfer delay (from 30 seconds to several minutes) to avoid cycling between sources; the fault capacity, which withstands between 10 and 50 kA for half a cycle or a cycle; and the integration with the management system via Modbus TCP, SNMP and dry contacts, so that switching is observable, something that, as we have repeated throughout this series, is not a luxury but part of reliability. The downside of the STS against the ATS is precisely its lower fault capacity: semiconductors tolerate less short-circuit current than contactors with arc chutes, which in sites with high fault current conditions the choice.
The high-end STS units of 2025-2026 have also refined details that matter in real installations: switching synchronised to the voltage zero crossing to achieve sub-cycle transfers (under 4 ms), true neutral switching in four-pole units, needed to transfer unbalanced loads safely, which are common in a datacenter, and redundant SCR controllers and drivers so that the STS itself does not introduce a single point of failure. The reliability of a static switch lies not only in the speed of its thyristors, but in the redundancy of the intelligence that decides when to fire them.
The table sums up the two worlds:
| ATS (electromechanical) | STS (solid state) | |
|---|---|---|
| Mechanism | Contactors / motorised | Thyristors (SCR) |
| Transfer time | 60-200 ms | 2-4 ms (¼ cycle) |
| Inside the server window? | No (requires UPS) | Yes |
| Typical location | Grid↔generator, building level | A/B paths, PDU/rack level |
| Fault capacity | High | Lower (limited by SCR) |
| Relative cost | Baseline | ~60-70 % more expensive |
| Standard | UL 1008 | UL 1008S |
Open or closed transition
There is a second dimension, orthogonal to the previous one, that defines how the change is made: whether the old connection is broken before or after the new one is made.
Open transition (break-before-make) disconnects from source one, pauses, and connects to source two. It is the natural and safe way when the two sources are not synchronised, because they must not touch each other, but it introduces the 60-200 ms gap of the ATS we already know. Hence it requires UPS backup for critical load.
Closed transition (make-before-break) does the opposite: it connects the second source before releasing the first, so that the load is never without power. The price is that, for an instant, the two sources are in parallel, and that is only admissible if they are synchronised and if the utility authorises it. The standards bound it: the United States electrical code (NEC 700.5(B)) limits that overlap to less than 100 milliseconds, and requires strict synchronisation conditions to close, with a voltage difference below 5 %, a frequency difference below 0.2 Hz and a phase angle difference of a few degrees. Closed transition is what allows you to test or return the load when the grid comes back without a single blink, a very valuable capability in mission-critical work. And its relative, soft loading or extended parallel transition, holds the parallel for a few seconds to ramp the load from one source to the other smoothly, ideal for the transfer tests that, as we insisted in the previous articles, have to be done without an outage.
The sequence of a power cut, step by step
It is worth stringing the pieces together in the order in which they act when, on an ordinary day, the grid goes away. At instant zero, the utility disappears. The IT load does not notice, because the UPS, which we will see in the last article, was already feeding it through its inverter (in double conversion) or transfers to its batteries in a few milliseconds: that is the ride-through, the bridge. Meanwhile, the control logic detects the grid failure and orders the generators to start. These take up to ten seconds to start and stabilise (NFPA 110, Type 10). When they are ready and synchronised, the paralleling switchgear, or the building ATS, transfers the load from the dead grid to the live generator; since that transfer happens upstream of the UPS, its gap of tens or hundreds of milliseconds is absorbed by the UPS itself without the load noticing. From there on, the generator sustains the facility.
When the grid comes back, the sequence reverses, and here closed transition shines: if the generator can synchronise with the grid, the system returns the load to the utility without a single blink, ramping it smoothly (soft loading) before shutting down the sets. Throughout this whole dance, the STS units of paths A and B have been doing their work in parallel, at load level, guaranteeing that every single-corded device always had a healthy source. Four components (UPS, generator, transfer switchgear and STS) choreographed so that the compute never knows that, out there, the lights went out.
The regulatory framework
Switching is heavily regulated because lives depend on it, not just bytes. The United States electrical code classifies systems by criticality: article 700 covers emergency systems (where an outage puts life at risk), 701 the legally required standby systems (which must come online in 60 seconds or less), 702 the optional ones, and 708 critical operations power systems. Each category imposes different requirements on the switch and on its listing. On that basis, UL 1008 governs transfer switches and UL 1008S the static ones, and closed transition is conditional on the interconnection agreement with the utility, because it means, even if for less than 100 ms, connecting your own generation in parallel with the grid. An architect who designs the switching without bearing this regulatory layer in mind is exposed to expensive surprises at the permitting stage.
Synchronising in order to close
Both closed transition and the generator paralleling we saw in the previous article depend on the same thing: synchronising before closing the breaker. The peaks and troughs of the waveforms of both sources must line up precisely, matching in voltage, frequency and, above all, phase angle, before they are joined, because closing out of synchronism causes currents and mechanical stresses capable of destroying the equipment. The guardian of that operation is the synchro-check relay (ANSI 25), a permissive device that only authorises the breaker to close when the three parameters are within limits. In large facilities, all this intelligence (synchronisation, load sharing, start sequence) lives in the paralleling switchgear governed by a master PLC, which is at once the brain of generation and that of large-scale switching.
The trend in these controls points to digitalisation and to grid-interactive functions: digital synchronisers with remote monitoring, predictive maintenance and grid-interactive capabilities (frequency regulation, reactive power control, black-start) that turn the datacenter’s generation and switching plant into an actor capable of talking to the electrical grid, not just of disconnecting from it. It is the same logic we saw emerging in the generators article with behind-the-meter generation: the boundary between the datacenter and the grid is blurring, and the switching plane is the point where that boundary is managed.
A/B architecture and the single point of failure
Let us put the pieces together in the dual-path architecture that runs through this whole series. In a 2N facility there are two complete, independent systems, A and B, each capable of supporting 100 % of the load. A dual-corded server connects to both and manages the transfer itself: here no STS is needed, because the redundancy is in the equipment itself. That is the ideal situation. The STS appears where that redundancy does not exist in the equipment: to give two sources to a single-corded load, or to transfer between two UPS outputs a set of loads that cannot have dual supply.
The mapping onto the Tier levels is illustrative. Tier II uses an ATS at the generator point and little else. Tier III has two independent paths but only one active at a time, and employs STS units at PDU or rack level to transfer between paths during maintenance. Tier IV keeps both paths active at once; dual cording eliminates the rack STS, but the STS remains present at transfer points between redundant UPS modules and for single-corded loads.
And here a design warning that connects with the philosophy of the previous articles: the switch itself can be a single point of failure. An STS that serves critical load and fails leaves that load with no alternative; a single UPS upstream of the STS is another weak point. That is why serious designs feed the STS from two independent UPS units, give the STS redundant controllers and drivers plus thermal monitoring, and, as a golden rule, specify a wrap-around maintenance bypass, with mechanical interlocking, for any STS or ATS that serves critical load. Replicating the sources and leaving the switch without redundancy is, once again, solving half the problem.
The maintenance bypass deserves emphasis because it is what makes possible the concurrent maintainability that defines Tier III and IV. Without it, testing or repairing a switch would mean leaving the load it serves unprotected; with it, the power is routed through an interlocked alternative path while the switch is isolated, tested and returned to service, all without a blink. There is a practical nuance in specifying it: if the two sources of the STS come from the same utility (a dual bus off the same feed), they are naturally synchronised and the STS can do make-before-break; if they are independent, you have to specify explicitly an STS with break-before-make capability, because it will have to transfer between sources that do not share phase. These details, invisible on a block diagram, are what separate a design that works from one that fails at the first real switching operation.
Towards solid state
Switching is living through its own revolution, pushed by AI. The corporate move says it all: Schneider Electric bought Vertiv’s transfer switch business (ASCO, a reference ATS manufacturer) for 1.25 billion USD, a sign of how strategic this link has become. The market for datacenter ATS and switchgear is projected to go from some 4 billion USD in 2025 to more than 10 billion in 2035.
The underlying technical novelty is the solid-state circuit breaker (SSCB), enabled by wide-bandgap semiconductors (silicon carbide and gallium nitride). It responds to faults in microseconds, isolating them so fast that the main bus voltage barely sags and the rest of the datacenter keeps computing without noticing. It is protection at the speed AI demands. The segment is attracting capital and innovation strongly: there are dedicated funding rounds to scale next-generation SSCBs for buildings and AI datacenters, and the major power semiconductor manufacturers flag it as one of the key pieces of the transition to 800 V distribution. It is a generational change: moving from protecting with mechanics, contacts that open, to protecting with electronics, which switch without arcing and without wear, finally aligning the speed of protection with that of the load.
Because that is the challenge that changes everything: AI loads swing brutally. An individual GPU oscillates 50-75 % of its power in milliseconds; at rack level, from 120 to 150 kW; and aggregated at gigawatt scale, system ramps exceed 1000 MW per second, stressing both switching capacity and grid stability. Added to the transition towards 800 V DC distribution, which NVIDIA proposes for its next architectures and which shifts the protection paradigm towards the DC world and SSCBs, the switching of an AI datacenter in 2026 looks less and less like that of five years ago. The referee has to be ever faster.
Another trend reshaping switching is integration with batteries and microgrids. As storage systems (BESS) enter the chain, to shave peaks, provide grid services or reduce generator hours, the switching plane stops being a simple “grid or generator” and starts orchestrating several simultaneous sources: grid, on-site generation, batteries and, where they exist, fuel cells. Grid-forming architectures with bidirectional flow can sustain the load without starting the generator for short events, leaving it only for extended outages or low battery state of charge. That multi-source orchestration, in milliseconds and with the intelligence to always pick the healthiest source, is the immediate future of the power chain’s referee. One reliability detail that is not minor: all this intelligence has to be protected from itself, giving the controllers redundancy so that the brain of the switching does not become, in its turn, a single point of failure.
Key takeaways
Switching is the referee that decides, instant by instant, where the power comes from. The master rule is the 10-20 ms window that server power supplies tolerate: the electromechanical ATS (60-200 ms) exceeds it and needs UPS backup, while the solid-state STS (2-4 ms) fits inside it and allows transfers without the load noticing. Closed transition avoids the gap at the price of synchronising and agreeing with the utility; the maintenance bypass guarantees that the switch itself is not a single point of failure; and the AI era, with its load swings and its jump to 800 V DC, is pushing this whole link towards solid state.
If you had to keep one idea only, it would be this: switching is not about choosing a source, but about doing it within a window of a few milliseconds, with the right technology at each point of the hierarchy, a robust ATS above and a fast STS below, without letting the referee itself become the link that fails. The rest are details, but details that are paid for dearly on the day of the first real switching operation.
One last component remains, the one that covers the gap the ATS cannot bridge and the one that first absorbs the stress of AI’s pulsing loads: the UPS. The fourth and final article closes the series with the sprinter of the power chain.
See also
- Datacenter physical infrastructure (1/4): the power chain
- Datacenter physical infrastructure (2/4): the generators
- Datacenter physical infrastructure (4/4): the UPS and the ride-through
Sources
- NFM Consulting, Static Transfer Switch vs ATS for Data Centers — https://nfmconsulting.com/knowledge/data-center-sts-vs-ats/
- Eaton, Automatic Transfer Switch Fundamentals — https://www.eaton.com/us/en-us/products/low-voltage-power-distribution-control-systems/automatic-transfer-switches/automatic-transfer-switch-fundamentals.html
- Eaton, Bypass-Isolation Contactor-Type Automatic Transfer Switches — https://www.eaton.com/us/en-us/catalog/low-voltage-power-distribution-controls-systems/bypass-isolation-contactor-type-automatic-transfer-switches.html
- LayerZero, Static Transfer Switches — https://www.layerzero.com/products/static-transfer-switches/
- ABB, Automatic Transfer Switching for Power Systems in Data Centers (UL 1008) — https://search.abb.com/library/Download.aspx?DocumentID=9AKK108468A1350&LanguageCode=en
- Consulting-Specifying Engineer, Understanding transfer switch operation — https://www.csemag.com/understanding-transfer-switch-operation/
- Power-Eng, Keep an Open Mind About Closed Transition — https://www.power-eng.com/om/keep-an-open-mind-about-closed-transition/
- ASCO / Schneider, NEC Requirements for Emergency Power Transfer Switching — https://cdn2.hubspot.net/hubfs/4333126/Files/Technical%20Information/ASCO/White%20Papers/ATS/ASCO%20national_electrical_code_requirements_for_emergency_power_transfer_switching_147492_0.pdf
- Generator Source, Paralleling Switchgear Explained — https://generatorsource.com/industries-served/data-centers/paralleling-switchgear-explained-how-we-power-hyperscale-data-center-growth/
- Socomec, Data Center Redundancy: Definition, Reliability, Best Practices — https://www.socomec.us/en-us/solutions/business/data-centers/data-center-redundancy-definition-reliability-best-practices
- Data Center Knowledge, Schneider Buys Vertiv’s Transfer Switch Business (ASCO) for 1.25B USD — https://www.datacenterknowledge.com/business/schneider-buys-vertiv-s-transfer-switch-business-for-1-25b
- NVIDIA Technical Blog, 800 V HVDC Architecture for AI Factories — https://developer.nvidia.com/blog/nvidia-800-v-hvdc-architecture-will-power-the-next-generation-of-ai-factories/
- arXiv, Operational Risks in Grid Integration of Large Data Center Loads — https://arxiv.org/pdf/2510.05437
- Uptime Institute, AI to trigger radical overhaul of data center electrification — https://intelligence.uptimeinstitute.com/resource/ai-trigger-radical-overhaul-data-center-electrification