THOUGHT LEADERSHIP

ARE YOU READY FOR RUBIN?

07 AUGUST 26

15 MINUTE READ

Author

Key Insights:

  • NVIDIA’s Vera Rubin shows how quickly AI hardware is now moving beyond the assumptions built into current Data Centre design.
  • Operators need infrastructure that can adapt to cloud, AI, or a changing mix of both, rather than being locked into one workload model.
  • Higher rack densities are making power and cooling decisions central to AI Data Centre performance, efficiency and risk.
  • As AI factories monetise tokens, every watt lost to inefficient distribution, cooling or pumping power becomes a direct commercial issue.

The chip cycle is now faster than the build cycle

By the time a Data Centre is designed, built and commissioned, the chips it was designed around have already moved on. It's an old problem - there's a Silicon Valley joke that most new tech is outdated before you've even unboxed it - but it’s never felt this extreme before.

NVIDIA's Vera Rubin platform has just gone from 'coming soon' to being real, with early units already landing with a handful of customers. The gap between what today's facilities were designed for and what's actually arriving, has never felt this wide. 

This is a right-now problem and it's the subject of our latest podcast episode, Are you ready for Rubin?, where our CTO, Lee Prescott, unpacks what the move to next-generation AI infrastructure actually means for anyone designing, building or operating a Data Centre today. 

This article expands on the conversation - the numbers, the trade-offs, and the design thinking behind them. It's part of The Centr, RED's thought leadership hub, where we work through the problems reshaping Data Centre design rather than waiting for the market to settle them.

ARE YOU READY FOR RUBIN?

From early internet hotels to future token factories

To understand why Vera Rubin matters, it helps to see how far the industry has travelled. The modern Data Centre market kicked off in the late 1990s with "telehouses" and internet hotels - the beginnings of colocation, spreading rapidly across Europe. Then the dot-com bubble burst in 2000, and building largely stopped: the investment had arrived before the world was digital enough to use it.

When demand returned, it came through the cloud. The global giants like - Google, Microsoft, META, AWS, etc - drove a decade of steady, predictable growth, and colocation boomed on the back of it. Rack densities crept from one kilowatt to ten (which felt like a lot at the time) and cooling evolved from chilled water systems toward mass air free cooling as operators looked for efficiency gains.

Then, in late 2022, ChatGPT launched, and everything changed. Almost nobody saw it coming. The industry stopped talking about watts per square metre and started talking in kilowatts per rack: from 10kW to 30, to 50, to 130 in barely any time at all. Vera Rubin takes that to 230kW this year, with systems expected to move towards 600kW a year or so later, and megawatt-scale racks already on the horizon.

The problem - IT is moving faster than construction

Data Centres have always had to plan for change, but what's different now is the pace.

A typical build programme, from design through to procurement, construction, and commissioning, takes around two years even on an aggressive schedule - and with racks now refreshing almost annually, that's two full IT generations between design and the day it goes live.

Rubin is this year's version of that problem.

NVIDIA's rack-level power draw has jumped from around 130kW to roughly 230kW over the course of this year, and by the end of next year, that figure is expected to climb past 600kW. That’s two major jumps in less than two years.

That leaves operators dealing with two things at once:

  1. Understanding what Rubin actually needs in terms of connectivity, power and cooling - and getting that integrated into live projects this year,
  2. Planning for infrastructure that can absorb another, even bigger jump in 2027.

Most facilities being designed today aren't ready for either.

What does Rubin actually ask of a facility?

The headline is 230kW per rack, but the detail is where legacy facilities come unstuck.

Rubin-class systems change how resilience is delivered. Power redundancy moves into the rack through power shelves configured as 3+1 - “four to make three” - rather than relying on traditional rack-level redundancy models.

And the cooling requirements are just as specific: these racks run on water supplied at 45°C, accept only certified fluids, and depend on strict requirements around filtration and the cleanliness of the water passing through them. It turns water management into an ongoing operational requirement rather than a final commissioning check.

There's a genuine upside buried in those numbers, though. If the GPUs want 45°C water at the rack, the water entering the CDU can sit around 42°C - opening up much more potential for free cooling in many climates.

The catch is that the air-cooled part of the facility still needs conventional chilled water in the low-to-mid 20s, so the design challenge becomes running two different temperature regimes in one building - capturing the efficiency of the warm loop without compromising the cold one.

ARE YOU READY FOR RUBIN?

Build for cloud, or build for AI?

AI hasn't replaced the cloud Data Centre, it's added a second category alongside it: the AI factory, with its own power profile, cooling demands and economics. Most operators are still built for the first while working out how much of the second they actually need.

That’s why a lot of them are framing their next decision as a straight choice: build for cloud workloads, or commit fully to AI. It's an understandable question, but it assumes something that isn't true yet - that anyone actually knows what the eventual mix of workloads will look like.

Hyperscalers and colocation providers are still, largely, running cloud-era businesses that are now experimenting with AI on top. They don't yet know whether their customers will want a traditional cloud environment, a fully optimised AI factory, or something in between. Designing too specifically for one outcome now creates a different kind of risk later: infrastructure that can’t adapt as demand changes.

A simple way to think about this is Formula One. F1 teams don't build a car that's brilliant on one track and mediocre everywhere else - they design and engineer for adaptability across conditions, because they can't predict every race. 

Data Centre design now needs the same principle applied to power and cooling: infrastructure that performs well whether it ends up hosting cloud applications or AI training/inference clusters. That ability to handle both has become central to the design brief.

The power model is changing too - and the UK hasn't caught up

Cloud Data Centres were built on a simple rule: 100% of the power, available 100% of the time. Anything less failed the brief. 

AI factories don't need to work that way. A factory making a physical product doesn't stop the line because supply drops for a day - it runs on what it's got and keeps producing. 

Apply that logic to power, and the numbers look different. A supply that's only available 95% of the time, the kind that's traditionally been turned down as too unreliable, can become workable if the facility is designed to keep producing tokens even when power availability drops.

The impact is substantial. Take a one-gigawatt utility supply: designed as a cloud Data Centre, it gives you around 650-700MW of IT. Designed as an AI factory, using power that would otherwise sit reserved for rare peak events, the same supply gives you over 900MW. That's hundreds of megawatts more - all from the same grid connection.

That distinction matters more in a market like the UK, where grid capacity is already stretched thin. It means sites and supply the industry has written off as unusable are, in practice, still on the table - and most operators haven't caught up to that yet.

Combine that with the GPU shortage everyone's already navigating, and the operators willing to rethink what "enough power" actually means will have options the rest don't.

ARE YOU READY FOR RUBIN?

Why AI changes the cooling conversation

Rack density at this level creates a cooling load that air systems can’t physically remove fast enough - which is why liquid cooling has become an industry standard almost overnight. 

It also rules some things out, building a new facility around mass air free cooling might look efficient, but it leaves no way to liquid-cool GPUs and no way into the AI market for the life of the building.

But the more interesting part of the conversation is what's changing as a result.

PG25, propylene glycol at 25%, has become the standard coolant in most data centre deployments. It's what NVIDIA and other manufacturers have certified for their direct to chip cooling systems, and it's become the default almost without debate. But PG25 is thick and viscous, and in a facility where everything is pumped 24/7, that's a permanent pumping-energy tax. Every extra unit of pumping power is a unit not going toward compute.

RED's instinct is to question that kind of default. The obvious alternative, ethylene glycol, is thinner and cheaper to pump - the industry avoids it because it's toxic if it reaches a watercourse, and outdoors that's an important consideration. But inside a controlled, secure Data Centre environment, the calculation looks different. It needs testing against cold plate technology first, but if it qualifies, it's energy handed straight back to the GPUs.

PG25 isn’t the only area where things are changing. Next-generation systems are moving to 800V DC power distribution, roughly double today's voltage, which changes the risk profile of these buildings entirely. 

It's new territory for the industry - new safety rules, new design considerations, and skills most Data Centre engineers have never needed. Telecoms and the EV sector have been working with DC for years, though, and that's where those skills will come from.

As Lee puts it in the episode, it's time to stop thinking about data centres as buildings, and start treating them as machines: dense, specialised, and requiring a different level of engineering care throughout.

Tokens, power, and why every watt now has a price tag

Here's the reframe that changes how you should think about AI infrastructure economics: an AI factory doesn't sell compute or storage the way a cloud provider does. 

It sells tokens - the basic unit every AI output is built from, whatever the user actually typed. Strip away the interface and the underlying process is always the same: something goes in, GPUs work through it, tokens come out.

Tokens are the currency of the AI factory model, and that means power spent making tokens is revenue. Power lost to inefficiency - wasted in underutilised chillers or lost in electrical distribution - is money the business will never see again.

That single idea is why RED engineers don’t stop at the building itself, and where our role goes beyond that of a typical MEP consultant.

Rather than taking a client brief and treating whatever sits inside the racks as someone else's concern, RED's engineers get inside the IT equipment itself - understanding the commercial model at the GPU end and designing the power and cooling systems around that reality, not just the facility around it.

With tokens tied directly to revenue, getting this right is a commercial decision as much as a technical one.

Gigawatt scale - and the race to build it

The bigger these facilities become, the more the engineering challenge changes. Multi-gigawatt AI factories are being planned in Europe, four and five gigawatt facilities are underway in the Middle East, and in the US there's serious talk of ten-gigawatt-class campuses.

At that scale, the facility starts to become the grid's problem. AI workloads don't draw power smoothly the way cloud does - they spike, sometimes beyond what the connection was designed for, and grid operators are understandably nervous about that.

The answer is batteries. Supercapacitors at rack level smooth out the spikes where they start, and large-scale battery storage sits between the facility and the grid, absorbing fluctuations and giving operators more control over when power is used. Done well, it can even work both ways - allowing the data centre to support the grid by feeding power back when it’s needed.

Energy storage is becoming a bigger part of the conversation around resilient Data Centre design. We’ve looked at this in more detail in our piece on Designing Resilient, Future-proof Energy Systems for Modern Data Centres.

And then there's the simple challenge of building these facilities fast enough. Generators, transformers and cooling plant all come with long lead times, and every month a facility isn't making tokens is money lost. That's why the industry is turning to standardised designs, modular builds and offsite manufacturing - anything that gets a facility live before the hardware moves on again.

What comes next?

None of this stops with Vera Rubin. The step-up expected for next year's platform is bigger again, and the pattern is likely to repeat for the foreseeable future: rising density, rising power, and infrastructure that has to be designed for a target that keeps moving.

Lee's advice is blunt: don't design for today, or even for tomorrow. Leapfrog tomorrow, and go beyond it.

That's why RED becoming an NVIDIA Preferred Partner matters. Partnering with the company defining these platforms gives our engineers something a spec sheet can't: earlier visibility of where GPU technology is heading, direct access to NVIDIA's own engineers and specialists, and sight of the pipeline before it's public. 

We were already working this way - questioning defaults, designing around the IT rather than stopping at the racks - so partnering with the market leader was a natural way to stay ahead in a sector that was already moving fast.

That means clients get infrastructure designed for what's coming next, rather than starting from a public spec sheet once the rest of the market has caught up.

It's a move away from traditional Data Centre design toward fully integrated environments, where power, cooling, and compute are optimised together for performance, efficiency, and tokens per watt.

Looking further ahead, concepts such as superconducting cables cooled to -200°C that all but eliminate resistance losses, point to where high-density AI infrastructure could go next. It's early-stage and expensive today, but it's the kind of thinking the industry will need as systems continue to scale.

ARE YOU READY FOR RUBIN?

So - are we ready for Rubin?

Honestly, the industry is adapting quickly but readiness was never going to be about mastering one generation of hardware. Rubin is this year's checkpoint, not the finish line. 

The operators who'll be in the strongest position aren't the ones who over-optimise for a single chip generation, they're the ones building infrastructure flexible enough to absorb whatever comes after it.

The pace of change isn't slowing down, and flexibility, not a fixed spec sheet, is what readiness actually looks like now.

Want the full conversation? Watch the Are you ready for Rubin? episode of our podcast for our complete breakdown, including the market data behind the density curve and RED's approach to flexible cooling design.

And for more on how AI is reshaping Data Centre design, visit The Centr, or read our piece on how AI factories are changing the way we build data centres.

Explore a range of content brought to you by RED’s engineering experts.

LATEST FROM THE CENTRE