Green & Efficient HPC: Liquid Cooling Isn’t Just for Giants Anymore

Introduction

Liquid cooling has been part of high-performance computing from the beginning. Mainframes and systems more geared towards HPC, like the CDC-6600, ran hot and fast and needed liquid cooling, and Cray systems pushed those limits even further. The Cray-2 system, for example, used full immersion cooling, complete with a striking “waterfalls” implementation, to manage its thermal load. Seymour Cray, whose 100th birthday would have been this year, has been famously quoted about the need for “plumbing” in supercomputers.

For decades, liquid cooling remained the domain of elite, large-scale HPC systems. It was reserved for the most demanding workloads, the biggest budgets, the most advanced facilities, and where energy efficiency is a must.

A prime example today is Europe’s first exascale supercomputer, JUPITER, powered by the BullSequana XH3000 architecture at the Jülich Supercomputing Centre. Its first module, JEDI, topped the Green500 for energy efficiency, marking a new milestone in sustainable performance.

Green & Efficient HPC: Liquid Cooling Isn’t Just for Giants Anymore
The JUPITER Exascale Development Instrument at the Jülich Supercomputing Centre. Copyright: Forschungszentrum Jülich / Ralf-Uwe Limbach
Green & Efficient HPC: Liquid Cooling Isn’t Just for Giants Anymore

But as the Green500 list has turned energy efficiency into a competitive, and informative, benchmark at the high-end, the conversation has become much broader. Today, datacenters everywhere are grappling with the same challenges that once belonged to the largest HPC systems, such as demanding workloads, hot and fast systems, and issues with power, cooling, and energy efficiency.

Liquid cooling isn’t just for the upper echelon of HPC anymore, it’s crossed into the mainstream.

Why Liquid Cooling Is No Longer Optional

GPUs and memory-intensive nodes are the norm for a growing number of applications, and they are needed in large quantities. And there is no end in sight. The corresponding increase in rack densities and thermal loads lead to where we are today. HPC simulations, AI training, and real-time inference all push systems beyond air cooling’s limits.

To be sure, air cooling will continue to play a role in most environments. But the amount of heat that must be extracted and taken out of nodes, racks, and buildings is more than what air can handle alone. Liquid cooling addresses all of these, with superior heat transfer, higher density potential, and growing opportunities for heat reuse. It enables higher performance, smaller footprints, and better energy use.

Designing for Liquid Cooling Realities

But when you replace air with liquid, the whole equation changes. Pumps replace fans, effluent discharge becomes more complex, and leaks become major events. And unlike air, liquid coolants require storage and handling. Meanwhile, energy costs, environmental standards, and local ordinances are tightening. Datacenters must weigh not just peak performance but power usage effectiveness (PUE), water usage effectiveness (WUE), lifecycle costs, and compliance requirements.

It’s Not Just About Scale: Use Cases Across the Spectrum

While exascale systems like JUPITER show what’s possible at the high end, liquid cooling now delivers tangible benefits in far more modest deployments:

  • Rack-scale systems: Supporting GPU-heavy workloads without throttling.
  • Modular data centers: Where power and thermal design are tightly integrated.
  • Edge and remote locations: Where traditional cooing is limited due to space, power, or noise.
  • Retrofit and hybrid environments: Adding liquid components without full system overhauls.

The Challenge: Too Many Options, Too Many Variables

The Need to Think Holistically

Cooling isn’t a standalone fix. True efficiency requires system-level optimization and seeing the entire datacenter as a coordinated system, including:

  • Hardware synergy: Aligning CPUs, GPUs, DPUs, and accelerators for optimal energy use.
  • Energy-aware scheduling: Dynamically assigning workloads to match thermal and power limits.
  • Integrated infrastructure: Planning cooling, power, and compute together from day one.

Without a system-level approach, even the most advanced cooling systems will underperform or add unnecessary cost.

Too Many Technologies, Too Little Guidance

Liquid cooling solutions now include immersion, direct-to-chip, warm water, 2-phase, rear-door, in-row exchangers, direct liquid cooling (DLC), hybrid systems, and more.

Green & Efficient HPC: Liquid Cooling Isn’t Just for Giants Anymore
Green & Efficient HPC: Liquid Cooling Isn’t Just for Giants Anymore
Green & Efficient HPC: Liquid Cooling Isn’t Just for Giants Anymore
Green & Efficient HPC: Liquid Cooling Isn’t Just for Giants Anymore

Each option comes with trade-offs in performance, cost, and suitability to the specific datacenter and location. The level of complexity has risen so quickly that few organizations have sufficient expertise across all categories. Without guidance, it’s easy to overbuild, underperform, or miss efficiency gains.

Every Environment Is Different

No two data centers are alike. Effective cooling strategies depend on:

  • Legacy infrastructure: Existing HVAC, airflow, and equipment layout.
  • Geography and climate: Altitude, humidity, or seasonal temperature swings.
  • Utilities: Power density, water access, and reuse potential.
  • Regulations: Local codes, especially for government or research institutions.

Cooling should conform to the site, not the other way around.

Why Customization Matters

When there is no one-size-fits-all solution (and there isn’t with liquid cooling), success depends on matching the right approach to the right workload, environment, and infrastructure.

For example, colder climates may prioritize heat reuse. Arid regions must focus on water conservation. And high-memory nodes or edge deployments present different cooling demands than AI training clusters.

Workload types also influence design. AI inference workloads may benefit from different thermal tuning than simulation or model training. High-memory nodes may introduce bottlenecks that traditional thermal designs can’t address efficiently.

Tailoring the cooling solution to these realities is essential, not just for performance, but for cost and sustainability.

The Role of Co-Design

Co-design aligns compute, cooling, and infrastructure planning from the start, treating them as interdependent parts of the same system.

Co-design allows organizations to:

  • Right-size cooling and compute infrastructure: Avoid overbuilding while leaving room for growth.
  • Improve long-term efficiency: Reduce energy use and cooling overhead.
  • Accelerate deployment: Reduce delays and rework.
Green & Efficient HPC: Liquid Cooling Isn’t Just for Giants Anymore

With so many liquid cooling technologies and site-specific variables at play, few organizations have the full range of expertise required. Co-design brings the right experts to the table and aligns them towards your mission. From facilities and IT to HPC architects, vendor-agnostic technologists, and application engineers, the right teams comes together to help make smarter more sustainable choices.

Conclusion: Cooling Is Now Core to System Design

Once reserved for the most advanced HPC environments, liquid cooling is now widely needed. Whether you’re running an exascale supercomputer or a compact modular cluster, efficiency now depends on how holistically your system is designed, not just how it’s cooled. Co-design is what makes that level of integration possible.

If you’re looking to navigate the options or plan your next system, the experts at SourceCode can help.

Note: This article was first published in HPCwire on October 20, 2025. 

Related Blogs

Smart Backup Is Smart Business

Why Co-Design is the Future of Intelligent Infrastructure 

The Rise of Application-Defined Infrastructure: Why Workloads Should Drive System Design