We spent thirty years making data centre water colder. The next efficiency win comes from making it hotter than your bath…
Key Insights:
- AI factories need infrastructure designed for the compute coming next, not just the racks available today.
- 800V DC can reduce power losses as rack densities push beyond what conventional AC distribution can comfortably support.
- Running cooling loops at higher temperatures can dramatically increase free-cooling potential and reduce refrigeration energy.
- The next challenge is making power and cooling infrastructure flexible enough to support changing compute mixes, higher densities and future technologies.
The future of the AI factory might depend on water hotter than your bath…
That sounds like an odd place to start, but there’s a good engineering reason for it.
AI hardware is moving faster than the infrastructure being built around it, forcing engineers to rethink the design principles data centres have relied on for years.
It’s the focus of our latest podcast, Powering the AI Factory: The Future of Cooling and Energy Infrastructure, where our Head of R&D, Mohammad Royapoor, walks through the reference design we’ve developed and what it means for the power and cooling infrastructure needed for the next generation of AI.
This article expands on that conversation as part of The Centr, RED’s thought leadership hub.
The question is simple: if our infrastructure has to support the next generation of compute, what needs to change now?
Quite a lot, actually - starting with your bathwater…
What is an AI factory, and why does it change the brief?
A traditional data centre is built to store, move and serve information, with little dynamic compute performed as data passes through the facility. An AI factory does the opposite; it takes in power and data and turns them into intelligence - namely, tokens - at unprecedented scale.
Compute is no longer just one part of the workload - it’s the product, and it needs to run continuously.
That changes the definition of efficiency. A conventional facility might optimise for uptime and floor space, but an AI factory needs to maximise the amount of tokens it can produce from the power and water it consumes.
That means the infrastructure has to be looked at as a whole. Power distribution, cooling, networking and compute are all working towards the same output, so the efficiency of one system can't be considered in isolation.
A power system that's 98% efficient at distributing power is still a poor decision if it leaves the cooling system without enough headroom to keep GPUs running flat out. The same applies in reverse: an efficient cooling system doesn't count for much if the power infrastructure delivers power efficiently.
That has a commercial consequence, too. The output of an AI factory is tokens, so any power lost to an inefficient chiller or high-loss power distribution system is capacity that could have been used to produce more compute - and more revenue.
Designing for the rack that doesn't exist yet
Most infrastructure gets designed around the hardware that's available when the drawings are signed off, but we took a different approach with our reference design - starting with NVIDIA's Rubin Ultra.
When we started the work, Rubin Ultra was still in the early stages, and all we really knew was what NVIDIA had shared in its marketing material - it was expected to reach around 600kW per rack.
Not exactly a full spec sheet, but at least we had a number to design around…
Our thought process was simple: design the infrastructure to deal with 600kW, and the thermal and electrical demands of today's 230kW racks become much easier to accommodate. Design for today's hardware and retrofit later, and you push the difficult part of the problem into the future, when it costs more and causes more disruption to fix.
The design also isn't tied to NVIDIA. RED has carried out the same characterisation work at GPU, server and rack level for Intel and AMD hardware, so the infrastructure can absorb whichever manufacturer's roadmap is actually selected by the client (or all three at once).
That’s the benefit of designing from the ceiling down. Design for the ceiling and you can still operate below it….
Design around today's floor, and you've got nowhere to go when the next generation arrives.
400V AC has hit a wall
Most facilities running today are built on 400V alternating current, but its limits are starting to show.
Push more power down a 400V AC conductor and the current climbs with it - which means heavier cabling, more copper and more heat losses. At the power densities Rubin Ultra brings (on the way to a megawatt per rack), those losses become difficult to engineer away, and you’re basically left with two choices: accept them or change the architecture.
Then there’s the conversion overhead. A typical AC power train converts twice: AC to DC and back to AC, on top of ordinary step-down conversion. Every conversion is a thermodynamic toll booth - a little energy gets lost as heat each time, and none of it makes it to the compute load.
An 800V DC architecture skips the booth; no repeated conversion, no repeated loss. Our own figures put the savings at around 4-5%, and that saving is there every hour the facility operates, throughout its entire lifecycle.
But why 800V, rather than an alternative DC voltage? Go lower and the current rises, bringing its own headaches around conductors and protection. Go much higher - 1500V, for example - and arc-flash risk becomes harder to manage with current industry practices. 800V is roughly where the electrical upside is largest, and the safety problem is still solvable.
It's a trade-off, rather than a magic number. Other DC voltages can work, but 800V currently gives the right mix of electrical performance and manageable safety requirements.
This isn't some distant future-looking concept, either. NVIDIA is already pushing the industry towards 800V DC, and the direction of travel is pretty clear - the work now is figuring out how quickly, and how safely, the infrastructure can catch up.
The catch? The components don't fully exist yet
None of this is a simple swap, and the hard part isn't just the engineering - DC brings a different set of safety problems with it.
An AC fault has a get-out built into the physics. With AC, the current crosses zero 100 or 120 times a second. That zero crossing helps an arc self-extinguish, and it's a large part of why AC protection equipment has had decades to become well understood and widely deployed.
DC doesn't have that crossing. Once an arc forms, the current keeps flowing and the voltage keeps driving it. Safely breaking that fault requires more sophisticated protection technologies than those used in conventional AC systems: solid-state breakers and solid-state transformers are among the promising solutions, which are still moving from pilot projects and early demonstrators towards wider commercial adoption.
That's the uncomfortable part of moving fast: the science and the demand for 800V DC are way ahead of the supply chain that's meant to deliver it.
We’re not saying to rip out 400V AC and start again. Operators have already invested heavily in the infrastructure they have, and most of it still has years of useful life left. Instead, the plan is a staged retrofit, working backwards from a target end state:
- Rack-level rectification first. Keep the existing AC backbone and convert power locally at each rack.
- Upstream rectifiers and DC busbars next. Move the DC boundary further back into the distribution chain as the infrastructure is upgraded.
- Medium-voltage-to-DC as the end state. Bring in solid-state transformers and end-to-end DC distribution, removing the remaining AC conversion stages.
The important point is that these aren't arbitrary steps on a five-year roadmap. Each one is a decision about how much of the existing infrastructure can stay in place, how many stages are actually needed and when it makes financial sense to move to the next one.
Nobody, us included, knows exactly what compute demand or the supply chain will look like in five years. The architecture needs enough flexibility to deal with that uncertainty.
Two cooling regimes, one building
Densify a rack far enough and air simply can't move the heat quickly enough. That's why direct liquid cooling - with coolant travelling straight to the chip or through microchannels beside it - went from optional to essential in the space of two hardware generations.
But an AI factory rarely runs on liquid cooling alone - most facilities will still have networking equipment, ancillary loads and plenty of existing cloud-era hardware that run on air.
That gives every facility two genuinely different cooling problems running in parallel: a high-temperature loop delivering water straight to the chip, and a conventional loop running in the low-to-mid 20s for everything that’s still air-cooled.
Neither technology is particularly exotic, but the engineering headache comes from having to run both without deciding the compute mix years before anyone knows what that mix will actually be.
An operator might want a rack full of GPUs today, more conventional compute tomorrow and a very different balance further down the line. The cooling infrastructure has to cope with that without forcing a redesign every time the hardware changes.
That means designing one chilled-water system that can support both loops, with each running at the temperature it needs while sharing the site's redundancy. Get that right and the facility stays live and maintainable, without having to take large sections offline every time the workload changes.
Your bath is running at 40°C - the cooling loop is hotter.
A bath at 40°C is already warm, but push it to 45°C and you’d be jumping straight out the tub. That’s the temperature RED wants arriving consistently at the GPU.
To anyone used to conventional data centre cooling, sending 45°C water towards a GPU sounds completely backwards. The industry has spent decades installing chillers for precisely the opposite reason. Conventional cooling loops typically have a flow of around 20-25°C, with chillers often using mechanical refrigeration to pull the temperature down before sending it back into the facility.
It works, but keeping that water cold takes energy - and the cooling plant can become one of the biggest electrical loads in the data centre after the compute itself.
Direct liquid cooling changes the equation because the water no longer needs to be that cold. At 45°C, the GPUs are still comfortably within their required operating temperature, while the cooling loop is now warm enough for ambient air to do much more of the cooling without refrigeration.
Run the water through a dry cooler and, when outdoor conditions allow it, the heat can be rejected without mechanical refrigeration - literal free cooling.
That's the slightly counter-intuitive bit: in a facility with 45-55°C loops performing GPU cooling,you're not cooling the water down to a temperature the GPU needs; you're keeping it hot enough that the outside air can do the job for you.
Our own modelling shows what that's worth:
- At 20-30°C: In a Central European climate, a conventional cooling loop can run in free-cooling mode for around 40-60% of the year.
- At 45°C to chip: Raise the water temperature, and free cooling becomes possible for close to 99% of the annual cycle.
- More power for compute: With less demand from compressors, more of the power available from the grid connection can go to the GPUs. The same utility connection can therefore support more tokens.
- Lower PUE: Across the facility as a whole, careful optimisation of the power and cooling infrastructure could bring annualised PUE down to around 1.10 in a temperate climate, or around 1.17-1.18 in hotter climates.
Those figures are hard to imagine with the chilled-water systems the industry has relied on for the last decade, but once you run the cooling loop hotter the maths starts to change. Less power goes into refrigeration, more goes into producing tokens - and none of it depends on technology that doesn't exist yet.
What's still unsolved
Not everything is a solved problem waiting to be adopted.
As GPU densities keep climbing, even single-phase direct-to-chip cooling has a limit. At around 300 watts per square centimetre it can start to struggle to remove enough heat, which is where two-phase cooling starts to look interesting: instead of keeping the working fluid liquid, you use phase change and let it boil at the point of contact with the GPU.
That phase change can move a lot of heat, but getting the system to work reliably at scale is a much bigger engineering challenge. We’re already in conversation with OEMs developing two-phase cooling, but the technology also has to clear the environmental and regulatory hurdles around the fluids involved - including PFAS restrictions. Better heat transfer only gets you so far if the fluid can't be deployed at scale.
Then there’s the power side. We’ve been looking at high-temperature superconductors as a way to cut losses in future power systems, including cryogenic 800V DC. Our work with partners in North America is focused on characterising their performance and reliability, and the technology is showing real potential for ultra-dense AI campuses - but it is still moving from development into early deployment, with wider adoption likely a few years away.
Built for the neighbourhood, not just the client
It’s easy to look at an AI factory and see a huge facility taking power and water from the community around it - but the same building can also put something back.
The heat leaving the racks at 45°C has to go somewhere, and with the right infrastructure, it can be used by nearby residential or commercial buildings rather than simply dumped into the atmosphere.
The same applies to the power system. AI factories can store large amounts of energy on site and support the wider grid, rather than simply drawing from it - using stored battery power when local demand is high, or using on-site generation to help ease pressure on the grid. The same infrastructure built to support GPUs can, properly designed, help stabilise the network it sits on.
We’re leading work with CIBSE on a new TM65 technical manual specifically for data centres, giving designers a clearer way to measure embodied Carbon using real data and baseline figures for what typical, good and best practice looks like.
Alongside that, we’re working with a UK university to calculate the total embodied Carbon of the compute elements in data centres. The goal is to create a benchmark the wider industry can use, rather than leaving every project to set its own standard.
What comes next?
The next generation is already pushing further. Higher densities and higher power demands mean the infrastructure has to keep pace with hardware that can change before the building is even finished.
Build around today’s spec sheet, and the next hardware generation can force changes far beyond the rack. Design for something like Rubin Ultra, or even further ahead, and there’s room to run today's systems without having to redesign the infrastructure every time power and cooling demands increase.
Want the full conversation? Watch our Powering the AI Factory podcast episode for a closer look at RED’s reference design and the roadmap to 800V DC.
For more on how AI is reshaping data centre design, visit The Centr or read our piece Are You Ready For Rubin?