Liquid cooling is set to reach 53% of AI chips in 2026, as Nvidia's Blackwell and Vera Rubin GPUs push past the heat levels air cooling can ever handle.
Why Liquid Cooling Is Becoming Non-Negotiable for AI Data Centers in 2026
Direct answer: Liquid cooling is projected to reach 53% penetration among AI chips in 2026, up from roughly 33% in 2025, because Nvidia's Blackwell and upcoming Vera Rubin platforms now push individual GPU power draw to 1,200 watts and full rack density beyond 100 kilowatts — levels that even sophisticated air cooling can no longer dissipate reliably. Google already runs liquid cooling on more than 80% of its AI servers, and AMD's Helios rack-scale platform is bringing liquid cooling to its MI450/MI500 lines in the second half of 2026. Liquid cooling has stopped being a specialized high-end option and become the default assumption for any AI-optimized data center built from here forward.
The Thermal Wall AI Chips Just Hit
For most of the history of data centers, air cooling was simply cooling — fans, chilled air, hot-and-cold-aisle containment, done. That assumption held for a remarkably long time, through generation after generation of denser CPU racks, because even a hot, fully loaded conventional server rack rarely broke much past 10 to 20 kilowatts. AI-optimized racks broke that pattern entirely and quickly. A single AI-optimized GPU, like Nvidia's Blackwell, can now draw up to 1,200 watts on its own, and a fully populated AI server rack running dozens of these chips can push rack density beyond 100 kilowatts, with immersion-cooled configurations enabling 150 kilowatts or more per rack. At that density, air simply cannot move heat away from the chip fast enough, regardless of how much airflow or how cold the intake air is. That physical ceiling, not a change in taste or fashion, is the entire reason liquid cooling is projected to reach 53% penetration among AI chips in 2026, up sharply from roughly 33% in 2025 — a meaningful share of new AI infrastructure now physically cannot run any other way.
This is a genuinely fast reversal of decades of received wisdom in data center design, where air cooling was assumed to be the cheaper, simpler, lower-maintenance default and liquid cooling was reserved for a narrow set of supercomputing or high-performance-computing use cases well outside the mainstream commercial data center industry. AI workloads didn't just nudge that assumption — they inverted it within the space of two or three hardware generations, to the point where, for the highest-density deployments built in 2026, air cooling is now the exotic, harder-to-justify option and liquid cooling is the default nobody seriously debates.
How Fast the Shift Is Actually Moving
The jump from roughly a third of AI chips running on liquid cooling in 2025 to a projected majority in 2026 is a genuinely fast swing for a category of physical infrastructure that usually changes slowly, constrained by construction timelines, supplier capacity, and long facility-refresh cycles. A few forces are compounding at once to make it move this quickly. Nvidia's rack-scale liquid cooling shipment volume is expected to double in 2026 alone, which says as much about supply catching up to demand as it does about demand itself. Google, already running liquid cooling across more than 80% of its AI servers, has effectively normalized the approach at hyperscaler scale, which puts competitive pressure on every other major cloud and AI infrastructure operator to match that reliability and density rather than fall behind on both cost-per-token and physical footprint. And AMD's entry into rack-scale liquid cooling through its Helios platform, rolling out across the MI450 and MI500 lines in the second half of 2026, means this is no longer a single-vendor story tied to Nvidia's roadmap alone — it's becoming the baseline expectation across the two dominant AI accelerator platforms simultaneously.
It's worth being honest that not every research estimate agrees on the exact number. Goldman Sachs has published a materially more aggressive trajectory — 15% liquid-cooled AI servers in 2024, rising to 54% in 2025, and 76% in 2026 — compared with TrendForce's 53% figure for that same year. Both point in the same direction and describe the same underlying shift; they simply don't agree on exactly how far along it already is, which is a genuinely open methodological question (covered in more depth in the Q&A section below) rather than a sign either figure is unreliable on its own terms.
Why Air Cooling Stopped Being Enough
It's worth being specific about why air cooling, which served the data center industry well for decades, hits a wall here rather than just gradually losing ground. Air has a fundamentally lower heat capacity and thermal conductivity than liquid — moving the same amount of heat away from a chip requires dramatically more airflow, more fan power, and more physical space around the chip than a liquid coolant loop needs to do the identical job. Liquid cooling systems can increase cooling efficiency by close to 25 times compared with air-based systems, which is not a marginal improvement at the edges; it's a difference in kind. A 140-kilowatt AI rack, a genuinely common density in a modern AI-optimized facility, simply cannot be adequately cooled by air alone at any reasonable fan speed or airflow volume — the chips inside it will throttle, or fail, well before the rack reaches anything close to its rated compute output. That's the blunt, physical reason "direct liquid cooling is mandatory" has become the standard language around platforms like Nvidia's Blackwell rather than "recommended" or "optional at higher densities."
This also explains why the framing conversation in the industry has shifted from cooling efficiency metrics that made sense for a lower-density world. Power Usage Effectiveness, or PUE, measures how much of a facility's total power draw goes to actual computing versus overhead like cooling — a genuinely useful metric when cooling was a comparatively fixed cost layered on top of whatever compute the facility happened to run. In an AI-dense facility, the more useful question has shifted toward something like tokens per watt: how much actual AI output a facility produces per unit of power consumed, factoring in that a facility with excellent PUE but old, air-cooled, throttled GPUs might still deliver meaningfully worse real-world AI throughput than a facility with slightly higher overhead but chips running at full, sustained performance because the cooling can actually keep up with them.
Direct-to-Chip and Immersion: The Two Approaches That Actually Work at This Density
Two genuinely different liquid cooling architectures have emerged as the real options once air cooling is off the table, and they solve the problem in different ways.
Direct-to-chip liquid cooling routes coolant through a cold plate mounted directly on top of the highest-heat components — the GPU and, increasingly, other hot-running chips on the board — while the rest of the server around it can still use conventional air handling for lower-heat components. It's the more incremental of the two approaches: it slots into a rack architecture that still resembles a conventional server rack, with the plumbing running to a cold plate rather than replacing the entire chassis design, which makes it the more common near-term choice for operators upgrading existing facilities or design lines rather than starting completely from scratch.
Immersion cooling goes considerably further: it submerges the entire server, or the entire board, directly in a dielectric (electrically non-conductive) fluid, eliminating fans entirely because there's no air moving through the chassis to begin with. Without fan noise, fan power draw, or fan failure as an ongoing maintenance concern, immersion cooling can support even higher rack densities — the 150-kilowatt-plus figures mentioned above are generally immersion-cooled configurations rather than direct-to-chip ones. The tradeoff is that immersion cooling asks for a more fundamental redesign of the server and the facility around it: different maintenance procedures, different serviceability assumptions (pulling a board out of a tank of fluid is a different physical operation than sliding a server out of an air-cooled rack), and generally a new-build-native approach rather than something layered onto an existing facility. In practice, most operators scaling up in 2026 are choosing direct-to-chip for the bulk of their fleet, with immersion reserved for the highest-density, most performance-critical deployments where the extra design commitment is worth the additional headroom it buys.
Nvidia's Vera Rubin and AMD's Helios: The Hardware Forcing the Issue
Nvidia's upcoming Vera Rubin platform is being described industry-wide as "fanless, all-liquid-cooled" — a genuinely different design philosophy from earlier GPU generations, which typically shipped with air cooling as the default and liquid cooling as a higher-end option for operators who wanted it. Removing the fanned option entirely, rather than offering it alongside a liquid-cooled variant, is Nvidia effectively telling the market that at Vera Rubin's power and density levels, there is no viable air-cooled configuration worth designing or supporting. That's a significant signal for any operator planning a facility refresh around next-generation Nvidia hardware: the cooling architecture decision isn't a choice to be made later, it's a prerequisite baked into the chip generation itself.
The supply chain behind Vera Rubin's cooling is worth naming specifically, because it shows how concentrated some parts of this build-out are even as overall liquid-cooling adoption broadens. Jentech has emerged as the sole heat-spreader supplier for the platform, while Cooler Master, AVC, Boyd, and Auras are the named cold-plate suppliers — a small, specific list of companies now sitting on the critical path for one of the industry's most anticipated chip platforms. AVC specifically is reported to begin shipping Vera Rubin cooling modules starting in the third quarter of 2026, which puts a concrete date on when this next wave of liquid-cooling-native hardware actually starts reaching data centers at volume rather than remaining a roadmap promise.
AMD's answer is Helios, a rack-scale platform rolling out liquid cooling across its MI450 and MI500 GPU lines in the second half of 2026. Helios matters beyond AMD's own market share because it confirms this shift isn't an Nvidia-specific hardware quirk — it's a property of AI-accelerator power density generally, showing up on both of the dominant accelerator platforms at essentially the same time. When two competing chipmakers converge on the same architectural conclusion within the same calendar year, that's a much stronger signal about where the entire category is headed than either company's roadmap would be in isolation.
Who's Already There: Hyperscaler Adoption in 2026
Google is the clearest existing proof point that liquid cooling works reliably at real operational scale, not just in a lab or a pilot deployment: more than 80% of Google's AI servers already run on liquid cooling as of this research. That's not a company dabbling in a promising new technology — it's a company that has already made liquid cooling the operational default across the large majority of its AI fleet, presumably because the alternative (running AI-dense racks on air cooling) stopped being a viable option at the density Google's own AI workloads require.
Amazon, Meta, and OpenAI are cited across this same research as adopters moving in the same direction, alongside Google, as the fanless, high-density new chip generations from Nvidia and AMD make air-cooled AI infrastructure a shrinking, eventually vestigial category rather than a mainstream option. Each has its own reason for moving fast: AWS runs a mix of internal AI training, its own custom accelerator lines, and third-party GPU capacity sold to cloud customers, all of which face the identical thermal ceiling as any other operator running current-generation accelerators at scale. Meta's push into ever-larger frontier model training runs depends on sustained, maximum-throughput GPU utilization over long training cycles — precisely the workload profile where a throttled, air-cooled rack costs the most in wasted training time. And OpenAI's compute needs, spread across its own infrastructure and multiple partner clouds, make it one of the more direct beneficiaries of liquid-cooling-native capacity coming online, since both model training and high-volume inference benefit from sustained rather than throttled GPU performance.
The practical effect of hyperscaler-scale adoption reaching this point is that liquid cooling stops being treated as exotic or specialized even by smaller operators and enterprises: once the largest, most sophisticated infrastructure operators in the world have normalized it across the majority of their fleet, it becomes the reference architecture everyone else designs toward, rather than an unusual choice that needs its own separate justification.
The Supply Chain Now Racing to Keep Up
A shift this fast in what "normal" data center infrastructure looks like creates real winners and real bottlenecks in the supply chain underneath it. The named cold-plate and heat-spreader suppliers behind Nvidia's Vera Rubin platform — Jentech, Cooler Master, AVC, Boyd, and Auras — are among the most directly positioned beneficiaries of this shift, since their components sit on the critical path for essentially every next-generation liquid-cooled AI server being built. Nvidia's own rack-scale liquid cooling shipment volume is expected to double in 2026, which is as much a statement about supplier capacity ramping as it is about customer demand; a doubling of shipments only happens if the suppliers behind those systems can actually build and deliver at that pace.
This also reframes how a data center operator or an enterprise buying capacity should think about vendor risk in 2026. A GPU allocation is no longer the only supply constraint worth tracking — cold plates, coolant distribution units, and the specialized fluid-handling components behind immersion systems are now genuine bottleneck candidates in their own right, concentrated among a relatively small number of named suppliers. An operator planning a large AI infrastructure build-out in 2026 or 2027 needs visibility into that secondary supply chain, not just chip allocation from Nvidia or AMD, because a cooling-component shortage can delay a facility just as effectively as a chip shortage can.
The Global Picture: Where the Reporting Is Strong, and Where It's Thin
Liquid cooling adoption is, at its core, a hardware and hyperscaler-driven story rather than a national-policy one, which shows up clearly in how unevenly it's been reported across regions.
United States
The US is genuinely the center of this story on both sides of the supply chain. It's home to the platform vendors setting the pace — Nvidia with Vera Rubin, AMD with Helios — and to the hyperscaler adopters putting liquid cooling into production at the largest scale: Google, AWS, Meta, and OpenAI. The US-based cold-plate supply chain, including Cooler Master, AVC, Boyd, and Auras, with AVC specifically beginning to ship Vera Rubin cooling modules in the third quarter of 2026, means the US sits at both the design and early-supply end of this shift simultaneously.
United Kingdom
No distinct, liquid-cooling-specific reporting on the UK turned up in this research. The closest adjacent context is that facilities inside the UK's AI Growth Zones are generally described as being engineered for liquid cooling from the outset — a sensible design choice for any new-build AI-optimized facility in 2026 — but that's a general architectural expectation rather than a UK-specific liquid-cooling story with its own numbers or named projects.
UAE / Dubai
No distinct regional-specific reporting on liquid cooling was found for the UAE in this research pass. Given the scale of Stargate UAE's compute ambitions, it would be reasonable to assume liquid cooling plays a significant role in its design, but that's an inference rather than a documented finding, and it's worth being explicit about that distinction rather than presenting it as a confirmed fact.
Australia
No distinct regional-specific reporting on liquid cooling adoption in Australia was found in this research pass. As with several other topics in this space, Australia's AI infrastructure coverage tends to center on grid and energy policy rather than facility-level cooling technology choices.
Germany
No distinct regional-specific reporting on liquid cooling adoption in Germany was found in this research pass, despite Germany's separately documented, large national data center investment activity (covered in more depth elsewhere). The absence of dedicated reporting here doesn't mean German facilities aren't adopting liquid cooling — new AI-optimized builds anywhere increasingly assume it by default — it simply means no distinct, named Germany-specific liquid-cooling story surfaced in this research.
Europe / France
The same applies to France: no distinct regional-specific reporting on liquid cooling adoption was found, separate from the country's broader, well-documented AI compute infrastructure investment. Given that Mistral's own compute buildout and the CampusAI joint venture both involve extremely high-density Nvidia GB300 deployments, it's reasonable to expect liquid cooling is integral to those builds, but again, that's a reasonable inference rather than a specific documented finding for this topic.
China
China is the one region outside the US with a specific, documented liquid-cooling data point: Chinese cloud service providers are cited as a driver of overseas AI project demand for liquid cooling, and Tier 2 data center operators more broadly are accelerating their own investment in the technology. That combination — Chinese CSPs' overseas projects and a broader domestic Tier 2 operator push — suggests liquid-cooling adoption in and around China's data center sector is moving in step with the global shift rather than lagging behind it, even though the available reporting doesn't break down adoption to the same granular percentage figures available for the US market.
Beyond Penetration Rates: What "Tokens per Watt" Actually Changes
The rise of liquid cooling isn't only a facilities story — it's changing how operators measure whether their infrastructure investment is actually paying off. Power Usage Effectiveness has been the industry-standard efficiency metric for a long time, and it still matters, but it was built for an era where the compute output side of the equation was comparatively stable across facilities running similar hardware. In an AI-dense world, that assumption breaks down: two facilities with identical PUE can produce meaningfully different amounts of actual AI output per watt if one is running chips at full sustained performance under liquid cooling and the other is running the same chips throttled under an inadequate air-cooling setup that can't keep pace with sustained load.
Tokens per watt — a measure of how much AI inference or training output a facility produces per unit of energy consumed — captures that difference in a way PUE alone can't. It reframes the cooling investment decision away from "how much overhead does cooling add to our power bill" and toward "how much of our GPUs' rated performance are we actually able to sustain, and for how long, before thermal limits force a slowdown." Framed that way, liquid cooling stops looking like a cost center layered on top of compute and starts looking like a direct multiplier on the useful output of every dollar already spent on the chips themselves — which is a large part of why operators are willing to absorb the higher upfront cost and design complexity of liquid cooling rather than sticking with cheaper, simpler air handling.
Retrofitting Versus Building New
One question every operator with existing air-cooled facilities eventually has to answer is whether to retrofit or start over, and the honest answer emerging across the industry in 2026 is that new, AI-optimized "AI factory" builds are treated as liquid-cooling-native from the design stage, rather than air-cooled facilities being converted after the fact. Retrofitting an existing facility for direct-to-chip cooling is technically possible for some rack configurations, particularly where the existing electrical and structural capacity can support it, but immersion cooling in particular tends to require a level of structural, plumbing, and layout change that's considerably more practical to design in from the start than to bolt onto a facility built around a completely different cooling assumption.
That distinction matters for how quickly the 53% (or, on Goldman Sachs' more aggressive estimate, 76%) 2026 penetration figure can keep climbing. It's not simply a matter of existing facilities gradually adding liquid cooling piecemeal — a meaningful share of the growth is coming from entirely new construction that was liquid-cooling-native from the first design review, which also explains why so much of this shift is concentrated among hyperscalers and dedicated AI infrastructure builders currently in an active construction phase, rather than showing up evenly across the entire existing global data center base.
What Adopting Liquid Cooling Actually Requires Operationally
Moving to liquid cooling isn't simply swapping one cooling unit for another — it changes a meaningful part of how a facility is run day to day. Coolant distribution units become a new piece of critical infrastructure sitting between a facility's chilled-water loop and the individual racks, and like any other critical infrastructure component, they need their own redundancy, monitoring, and maintenance plan; a coolant distribution unit failure in a liquid-cooled facility is a more acute event than a single failed fan in an air-cooled one, because it can affect an entire row of high-density racks at once rather than a single server. Leak detection is a genuinely new operational discipline air-cooled facilities never had to build — introducing liquid into a rack full of expensive, live electronics means leak sensors, drip trays, and clear incident-response procedures are baseline requirements before a facility goes live, not optional extras.
Staffing and training shift accordingly. Data center technicians who spent a career on airflow, filters, and fan replacement now need real fluency in plumbing, coolant chemistry (particularly for the dielectric fluids used in immersion systems), and the specific maintenance procedures each cooling vendor's hardware requires — direct-to-chip cold plates, coolant distribution units, and immersion tanks each come with their own service routines, and cross-training an existing air-cooling-trained workforce takes real time and investment. Serviceability itself changes too: pulling a board out of an immersion tank for a component swap is a physically different task from sliding a server out of an air-cooled rack, and facilities need to plan service bays, fluid-handling equipment, and technician procedures around that difference rather than assuming existing workflows simply carry over.
None of this is a reason to avoid liquid cooling — the physics covered earlier in this piece make it a requirement, not a preference, at current AI chip densities. It is a reason operators need to budget for the operational transition alongside the capital cost of the cooling hardware itself, and it's part of why so much of the industry's liquid-cooling growth is concentrated among hyperscalers and dedicated AI infrastructure builders who can absorb that transition cost across a large fleet, rather than a smaller enterprise operator converting a single existing facility piecemeal.
What This Means Going Forward
For any business whose product or growth plan depends on sustained access to AI inference or training capacity, liquid cooling's rapid rise is worth tracking as more than an infrastructure curiosity. Cooling architecture is now a real signal of how reliable and how fast a given compute provider's newest capacity actually is: a facility still running last-generation air-cooled hardware is very likely running at a lower sustained density, and possibly at throttled performance under heavy load, compared with a newer liquid-cooled facility running the same chip generation at full rated output. When evaluating cloud or AI infrastructure providers for a workload that genuinely needs sustained, high-density GPU performance, asking directly about facility cooling architecture is a reasonable, concrete diligence question, not an overly technical one to leave entirely to a vendor's marketing page.
It's also a reminder that the AI infrastructure layer underneath any AI-powered product is moving fast enough that architectural assumptions made even a year or two ago can already be behind the current hardware generation. Businesses building AI-powered features and agents should design with that pace of change in mind rather than assuming today's provider, region, or hardware generation will still be the best-available option in eighteen months — which is exactly the kind of forward-compatible thinking we build into AI agents and automation engagements, and into custom software development more broadly whenever a system's performance depends on infrastructure that's still evolving this quickly underneath it. Liquid cooling itself is a data-center-layer decision most software teams will never touch directly — but the pace at which it's becoming mandatory is a useful, concrete illustration of how fast the ground is still shifting under anything built on top of frontier AI compute.
There's a second, quieter implication worth naming: cooling architecture is becoming a genuine differentiator in how AI infrastructure providers market themselves, not just an engineering detail buried in a spec sheet. Expect procurement conversations in 2026 and 2027 to increasingly include direct questions about rack density, cooling architecture, and sustained-versus-peak performance guarantees, the same way businesses already ask about uptime SLAs or geographic redundancy — because at current AI chip densities, cooling architecture has become just as material to whether a workload actually gets the performance it's paying for.
Straight Answers on AI Data Center Liquid Cooling
What percentage of AI chips are expected to use liquid cooling in 2026?
Liquid cooling penetration among AI chips is projected to reach 53% in 2026, according to TrendForce, up from roughly 33% in 2025. That's a genuinely fast jump for a hardware and facilities category that typically moves slowly, and it reflects a real physical threshold rather than a shift in preference: a majority of new AI chip deployments in 2026 are running at power and density levels where air cooling simply isn't a viable option anymore. It's worth noting this figure sits alongside a more aggressive estimate from Goldman Sachs, which put 2026 liquid-cooled AI server penetration at 76% — both estimates describe the same underlying shift and point in the same direction, they just diverge on exactly how far along it already is, most likely due to differences in what's being measured (chips versus servers versus racks) and which segment of the market each estimate covers.
Why will liquid cooling dominate AI data centres in 2026?
Liquid cooling is set to dominate AI data centers in 2026 because the newest generation of AI accelerators has outgrown what air cooling can physically handle. Nvidia's Blackwell platform draws up to 1,200 watts per GPU, and fully populated AI racks now regularly exceed 100 kilowatts, a density air simply cannot dissipate reliably no matter how much airflow is applied. Liquid cooling can improve cooling efficiency by close to 25 times compared with air-based systems, which isn't a marginal gain — it's the difference between a chip running at full rated performance indefinitely and one throttling under sustained load. With Nvidia's Vera Rubin platform shipping as fanless and all-liquid-cooled, and AMD's Helios bringing the same approach to its MI450/MI500 lines in the second half of 2026, liquid cooling isn't one option among several anymore for high-end AI hardware — it's the only architecture the newest chips are actually designed to run on.
Why is liquid cooling no longer optional for 2026 AI workloads?
Liquid cooling stopped being optional once mainstream AI accelerators crossed the power and density thresholds where air cooling physically cannot keep pace, regardless of budget or engineering effort spent on the air-cooling side. A 140-kilowatt AI rack, an increasingly ordinary density in 2026, cannot be adequately cooled by air at any reasonable airflow volume — the chips inside will throttle or fail well before reaching their rated compute output. That's a hard physical ceiling, not a soft recommendation, which is why documentation around platforms like Nvidia's Blackwell now states plainly that direct liquid cooling is mandatory rather than optional at the high end. For any organization actually running 2026-generation AI hardware at meaningful scale, air cooling isn't a slower or cheaper alternative anymore — for a growing share of deployments, it simply isn't an alternative that works.
How much power does a single AI-optimized GPU like Nvidia's Blackwell now draw?
A single AI-optimized GPU like Nvidia's Blackwell can now draw up to 1,200 watts on its own — a striking figure compared to the far lower per-chip power draw of previous data center hardware generations, let alone consumer-grade chips. That number matters because power draw and heat output are directly linked: a chip pulling 1,200 watts is also generating roughly that much heat that has to be removed continuously for the chip to keep running at its rated performance. Multiply that across dozens of GPUs in a single rack, and the aggregate heat load explains directly why rack density has pushed past 100 kilowatts and why air cooling, which was never designed to handle heat loads anywhere near this concentrated, has become physically inadequate at the high end of the market. Upcoming platforms like Vera Rubin are expected to push per-chip power draw even further, which is precisely why Nvidia designed that platform as fanless and all-liquid-cooled from the outset rather than offering an air-cooled configuration at all.
How dense can a fully liquid-cooled AI server rack become (in kW)?
Fully populated AI server racks are now regularly pushing beyond 100 kilowatts of density, and immersion-cooled configurations specifically can enable 150 kilowatts or more per rack. To put that in context, a conventional, air-cooled enterprise server rack from a decade ago typically ran somewhere in the 5-to-15-kilowatt range — meaning today's highest-density AI racks are operating at roughly ten times the heat load per rack that air cooling was ever designed to manage comfortably. That jump is the direct result of packing more GPUs, each individually drawing thousands of watts less than Blackwell's 1,200-watt figure, into the same physical rack footprint to maximize compute density per square foot of expensive data center floor space. Immersion cooling's ability to push past 150 kilowatts per rack, well beyond what direct-to-chip cooling comfortably supports, is one of the main reasons operators reach for it specifically for their highest-density, most performance-critical deployments rather than defaulting to it everywhere.
What is direct-to-chip liquid cooling, and how does it differ from traditional air cooling?
Direct-to-chip liquid cooling routes a liquid coolant through a cold plate mounted directly onto the highest-heat components on a server board — principally the GPU — carrying heat away at the source rather than relying on air moving across the whole board to carry heat out of the chassis. Traditional air cooling depends on fans pushing chilled air across the entire server, picking up heat from every component along the way before it's expelled and recirculated through the facility's cooling system; that approach works fine at lower power densities but runs out of headroom fast once individual chips are generating over a thousand watts of heat each. Direct-to-chip cooling keeps much of the rest of the server's air-cooling architecture intact, upgrading only the highest-heat components to liquid cooling, which makes it the more incremental, easier-to-adopt option for operators upgrading an existing rack or facility design rather than starting a cooling architecture from scratch.
What is immersion cooling, and why does it eliminate the need for fans?
Immersion cooling submerges an entire server, or an entire circuit board, directly in a dielectric fluid — a liquid engineered to be electrically non-conductive so it can safely surround live electronic components without causing a short circuit. Because the fluid itself is in direct contact with every heat-generating component and continuously carries that heat away, there's no need to move air through the chassis at all, which means fans, and the noise, power draw, and mechanical failure risk that come with them, are eliminated entirely rather than just reduced. That fan elimination is also part of why immersion cooling can support the highest rack densities in the industry, upwards of 150 kilowatts per rack in some configurations — without fans and airflow clearance constraints dictating how tightly components can be packed, designers have considerably more freedom to maximize compute density within the same physical footprint, at the cost of a more fundamental redesign of how the server itself is built and serviced.
How much more efficient is liquid cooling than air-based cooling systems for AI data centers?
Liquid cooling can increase cooling efficiency by close to 25 times compared with air-based systems for AI data center workloads — a difference driven by the basic physics of how much heat a liquid can carry away relative to air moving at a comparable volume. That efficiency gap is what makes liquid cooling economically sensible despite its higher upfront installation complexity: at AI-scale power densities, the alternative to that efficiency gain isn't a cheaper version of adequate cooling, it's inadequate cooling that leaves expensive GPUs running below their rated performance. Framed against a metric like tokens per watt rather than raw cooling overhead, that roughly 25-times efficiency advantage is less about saving on a utility bill and more about actually being able to extract the compute performance the hardware was purchased to deliver in the first place, which is why operators building serious AI infrastructure in 2026 treat liquid cooling as a performance investment rather than purely a cost-control measure.
What percentage of Google's AI servers already use liquid cooling?
More than 80% of Google's AI servers already run on liquid cooling, based on this research — a figure that puts Google well ahead of the broader 2026 industry-wide estimates of 53% (TrendForce) to 76% (Goldman Sachs) penetration across all AI chips. That gap between Google's own fleet and the broader industry average is telling: Google has been operating AI infrastructure at serious scale for longer than most competitors, and its cooling architecture reflects lessons learned from running dense compute in production well before liquid cooling became the industry-wide default it's becoming in 2026. For any operator or enterprise trying to gauge where the rest of the industry is likely headed, Google's fleet-wide adoption level is a useful leading indicator — it shows what "fully liquid-cooled at hyperscale" actually looks like in practice, rather than as a future projection.
How did liquid-cooling adoption change between 2024 and 2026 according to Goldman Sachs' forecasts?
Goldman Sachs' figures show liquid-cooled AI servers rising from 15% in 2024, to 54% in 2025, to 76% in 2026 — a trajectory that describes a genuine majority-adoption tipping point occurring within a single calendar year, 2025, and then continuing to climb sharply afterward. That pace, roughly quintupling penetration in two years, is unusually fast for a physical infrastructure category, and it lines up with the same underlying driver covered throughout this piece: successive generations of AI accelerators crossing power and density thresholds where air cooling stops being viable, forcing rapid adoption rather than a slow, optional transition. Goldman Sachs' 76% figure for 2026 is notably higher than TrendForce's 53% figure for the same year, a discrepancy worth treating as a real, open methodological question rather than dismissing either number outright, since both come from credible research organizations tracking the same underlying market.
Why do these liquid-cooling penetration estimates differ between sources (e.g., TrendForce's 53% vs. Goldman Sachs' 76% for 2026)?
The honest answer is that this is a genuine discrepancy worth flagging rather than a confidently resolvable one based on the research available here. Estimates like these commonly diverge because of differences in exactly what's being measured — the share of AI chips shipped with liquid cooling support, the share of AI servers actually deployed with liquid cooling active, or the share of total data center power capacity running on liquid-cooled infrastructure, are all related but meaningfully different denominators that can produce different percentages from the same underlying market. They can also diverge based on which segment of the market a forecast weights most heavily; a forecast weighted toward hyperscaler and frontier-lab deployments, where adoption is highest, will naturally read higher than one that includes a broader base of smaller Tier 2 and enterprise data center operators still catching up. Both TrendForce's 53% and Goldman Sachs' 76% for 2026 point toward the same directional story — liquid cooling becoming the majority approach — even if the precise number depends on exactly how you draw the boundary around what's being counted.
What cooling system does Nvidia's Blackwell platform require, and can air cooling still work at all?
Direct liquid cooling is described as mandatory for Nvidia's Blackwell platform at its higher-density configurations, reflecting how far individual chip power draw — up to 1,200 watts per GPU — has outpaced what air cooling can dissipate at rack scale. Air cooling isn't entirely eliminated as a concept across every Blackwell deployment; lower-density configurations or individual components generating less heat can still use air in some contexts. But for the fully populated, high-density rack configurations that make Blackwell attractive as an AI compute platform in the first place, air cooling alone isn't a workable option, which is why "mandatory" rather than "recommended" is the accurate description of liquid cooling's role at the high end of this platform. That mandatory requirement is also precisely why Nvidia's next platform, Vera Rubin, dropped the air-cooled option entirely rather than continuing to offer it as an alternative.
What is Nvidia's Vera Rubin platform, and why is it described as 'fanless, all-liquid-cooled'?
Vera Rubin is Nvidia's next-generation AI compute platform, following Blackwell, and it's described industry-wide as fanless and all-liquid-cooled because Nvidia designed it without an air-cooled configuration option at all — a clear break from earlier chip generations, which typically offered air cooling as the default and liquid cooling as a premium upgrade. That design choice reflects Nvidia's own judgment that at Vera Rubin's expected power and density levels, an air-cooled variant wouldn't be a viable product worth engineering or supporting, not merely a less efficient one. For any data center operator planning a facility refresh around Vera Rubin, that fanless design removes the cooling-architecture decision entirely — there's no "which cooling option should we choose" conversation to have, because liquid cooling is simply built into what the platform is. AVC is reported to begin shipping Vera Rubin cooling modules starting in the third quarter of 2026, putting a concrete date on this next wave of hardware reaching facilities.
Who supplies heat spreaders and cold plates for Nvidia's Vera Rubin platform?
Jentech has emerged as the sole heat-spreader supplier for Nvidia's Vera Rubin platform, while Cooler Master, AVC, Boyd, and Auras are the named cold-plate suppliers. That's a notably concentrated supply chain for a platform expected to anchor a large share of next-generation AI infrastructure — a small number of named companies now sit on the critical path for cooling hardware across what will likely be a very large volume of AI servers once Vera Rubin ships at scale. AVC specifically is reported to begin shipping Vera Rubin cooling modules starting in the third quarter of 2026. For data center operators and enterprises planning AI infrastructure build-outs around Vera Rubin, that concentration is worth factoring into procurement risk planning directly — a supply disruption at any one of these named suppliers could meaningfully affect delivery timelines industry-wide, not just for a single customer's order.
What is AMD's Helios rack-scale platform, and when is it rolling out?
Helios is AMD's rack-scale liquid cooling platform, rolling out across its MI450 and MI500 GPU lines in the second half of 2026. It's AMD's direct answer to the same thermal challenge driving Nvidia's shift to Vera Rubin's fanless design — AMD's own AI accelerators have reached power and density levels where rack-scale liquid cooling, rather than an optional add-on, is necessary to run them at their intended performance. Helios matters as a signal beyond AMD's own product line specifically because it confirms this shift isn't limited to one chipmaker's particular engineering choices; it's a property of high-end AI accelerator power density generally, showing up on both of the industry's dominant accelerator platforms within roughly the same window. For enterprises and data center operators evaluating AMD as an alternative to Nvidia for AI infrastructure, Helios means that choice no longer comes with a different cooling-architecture tradeoff — both major platforms now assume liquid cooling as the baseline.
Why are Chinese cloud service providers cited as a driver of overseas liquid-cooling demand?
Chinese cloud service providers are cited as a driver of demand for liquid cooling specifically in overseas AI projects, alongside a broader trend of Tier 2 data center operators globally accelerating their own investment in the technology. That pattern makes sense against the backdrop of a genuinely global AI infrastructure buildout: Chinese CSPs expanding AI infrastructure outside China's own borders face the same physical constraints as every other operator building with current-generation, high-power-density chips, and liquid cooling is the only architecture that reliably supports those densities regardless of which company or country is doing the building. It's also a useful reminder that this shift isn't confined to the US hyperscalers and chipmakers that dominate most of the coverage — demand for liquid cooling infrastructure is a genuinely global phenomenon, showing up wherever large-scale, high-density AI compute is being deployed, including in overseas projects backed by Chinese cloud providers.
What is the difference between direct-to-chip cooling and full immersion cooling for AI servers?
Direct-to-chip cooling routes liquid coolant through a cold plate attached to the specific highest-heat components on a board, typically the GPU, while leaving the rest of the server's design and air-handling largely intact around it. Full immersion cooling submerges the entire server or board directly in a dielectric fluid, removing air handling and fans from the equation completely rather than upgrading just the hottest components. The practical tradeoff is complexity versus density: direct-to-chip is the more incremental option, easier to adopt into an existing rack or facility design and generally sufficient for current mainstream AI rack densities, while immersion cooling requires a more fundamental redesign of the server and facility but unlocks meaningfully higher densities, 150 kilowatts per rack or more, for operators willing to commit to that redesign. Most operators scaling AI infrastructure in 2026 are defaulting to direct-to-chip for the bulk of their fleet and reserving immersion for their highest-density, most performance-critical deployments.
Why is NVIDIA's rack-scale liquid cooling shipment volume expected to double in 2026?
Nvidia's rack-scale liquid cooling shipment volume is expected to double in 2026 because both demand and supply are scaling at the same time: more customers are deploying Nvidia's newest, liquid-cooling-native platforms at higher volumes, and the cooling supply chain behind those platforms — cold-plate and heat-spreader suppliers like Jentech, Cooler Master, AVC, Boyd, and Auras — has been ramping capacity to meet that demand. A doubling in a single year is a significant acceleration even relative to the broader liquid-cooling penetration trend across the industry (53% of AI chips in 2026, per TrendForce), and it reflects Nvidia's specific position at the center of the current AI infrastructure buildout: as the dominant supplier of the highest-end AI accelerators, its own rack-scale liquid cooling shipments are a direct proxy for how fast the newest, most power-dense generation of AI infrastructure is actually being deployed across the industry, rather than merely planned or announced.
How does 'tokens per watt' as a metric change how data centers evaluate cooling investment, compared to traditional PUE?
Power Usage Effectiveness measures how much of a facility's total power draw goes toward actual computing versus overhead like cooling — useful when the compute output side of that equation was relatively stable across similarly equipped facilities. Tokens per watt instead measures actual AI output, inference or training throughput, per unit of energy consumed, which captures something PUE alone can miss entirely: two facilities with identical PUE can produce very different amounts of real AI output if one is running its GPUs at full sustained performance under adequate liquid cooling and the other is running the same chips throttled because its cooling can't keep pace with sustained load. That shift reframes cooling from a line-item overhead cost into a direct performance lever — the question becomes less "how much does cooling add to our power bill" and more "how much of the compute we already paid for are we actually able to use," which is a more honest and more actionable way to evaluate whether a given facility's infrastructure investment is paying off.
What happens to a 140kW AI rack if it relies only on air cooling instead of direct liquid cooling?
A 140-kilowatt AI rack running on air cooling alone will not be able to sustain its full rated compute performance — the chips inside it will throttle their clock speeds and power draw to keep temperatures within safe operating limits well before the rack reaches the output it was purchased to deliver, and in more extreme or poorly managed cases, components can fail or shut down entirely under sustained heat stress. That throttling isn't a minor efficiency loss; at that density, it can mean a meaningful share of the rack's theoretical compute capacity is simply unavailable in practice, which defeats much of the purpose of buying that dense a configuration to begin with. This is the concrete, operational version of the abstract "air cooling can't keep up" claim running through this entire topic — it's not that air cooling performs somewhat worse at this density, it's that it functionally caps how much of the hardware's actual capability an operator can use at all.
Which cooling-hardware suppliers are best positioned to benefit from the shift to liquid cooling in 2026-2027?
The suppliers named specifically across current research as best positioned to benefit include Jentech, the sole heat-spreader supplier for Nvidia's Vera Rubin platform, and Cooler Master, AVC, Boyd, and Auras, the named cold-plate suppliers for the same platform, with AVC specifically set to begin shipping Vera Rubin cooling modules in the third quarter of 2026. These companies sit on the critical path for what's likely to be a very large volume of next-generation AI server deployments, given Vera Rubin's fanless, all-liquid-cooled design leaves no air-cooled alternative for operators to fall back on. Beyond the specific Nvidia-aligned suppliers, the broader category of coolant distribution units, dielectric fluids for immersion systems, and rack-level plumbing components stands to benefit from the same structural shift, even where individual company names are less widely reported than the Vera Rubin-specific supplier list.
Is retrofitting an existing air-cooled data center for liquid cooling feasible, or does it typically require new-build facilities?
Retrofitting is technically feasible for direct-to-chip cooling in some existing facilities, particularly where the underlying electrical and structural capacity can support the added plumbing and cooling infrastructure, but the broader pattern across the industry in 2026 is that new AI-optimized "AI factory" builds are designed as liquid-cooling-native from the start rather than air-cooled facilities being converted afterward. Immersion cooling in particular tends to demand a more fundamental redesign of layout, structural support, and maintenance procedures that's considerably easier to plan for in a new build than to retrofit into a facility designed around a completely different cooling assumption. That's a large part of why so much of the current liquid-cooling growth is concentrated among hyperscalers and dedicated AI infrastructure builders actively constructing new capacity, rather than showing up evenly as a retrofit wave across the existing global data center base — for many operators, waiting for a new facility is the more practical path to liquid cooling than converting an old one.


