Skip to content
Inside the Hyperscaler Custom Silicon Race: TPU, Trainium, Maia, and MTIA vs. Nvidia
Technology39 min read

Inside the Hyperscaler Custom Silicon Race: TPU, Trainium, Maia, and MTIA vs. Nvidia

Scult Team
39 min read

Google, Amazon, Microsoft, and Meta are all racing to build custom AI chips that cut inference costs and reduce their dependence on Nvidia's GPUs.

Inside the Hyperscaler Custom Silicon Race: TPU, Trainium, Maia, and MTIA vs. Nvidia

Direct answer: Google, Amazon, Microsoft, and Meta are each building their own custom AI chips — TPU, Trainium, Maia, and MTIA, respectively — and shipments of these custom ASICs are growing roughly 44.6% year-over-year in 2026, nearly triple the 16.1% growth rate of merchant Nvidia GPUs, on track to reach 27.8% of AI server shipments, the highest share since 2023. The driver is economic and strategic at once: custom silicon can cut inference costs by 50-67% compared with Nvidia GPUs, and it reduces dependence on a single supplier that commands roughly 86% of the AI accelerator market and around 73% gross margins. None of that makes hyperscalers independent of the broader supply chain, though — every one of these chips is still built on the same constrained TSMC 3nm manufacturing capacity and SK Hynix high-bandwidth memory that Nvidia itself depends on.

Why Nvidia's Dominance Created Its Own Competition

It's worth starting with why this race exists at all, because the logic is more interesting than "big companies want to make their own chips for prestige." Nvidia's position in AI accelerators by 2026 is close to a structural monopoly: roughly 86% market share and gross margins around 73%, figures that would be remarkable in almost any other hardware category. For a hyperscaler spending billions of dollars a year on AI infrastructure, that combination — one supplier, that dominant, at that margin — reads less like a vendor relationship and more like a tax on every unit of AI compute the business runs. When a single supplier's margin is that high, the arithmetic for building an internal alternative changes: even a chip that costs tens of millions of dollars to design in its first year can pay for itself quickly if it meaningfully undercuts what an equivalent volume of merchant GPUs would have cost.

That's the economic driver. The strategic driver sits alongside it: supply independence. When global GPU demand outstrips supply — as it has through most of the current AI buildout — the company with its own fabricated, road-mapped chip supply isn't purely at the mercy of another vendor's allocation decisions during a shortage. Google, Amazon, Microsoft, and Meta have each reached the same conclusion from slightly different angles: TPU (Google), Trainium (Amazon, alongside its Inferentia line), Maia (Microsoft), and MTIA (Meta) are all, at their core, bets that owning silicon designed specifically for each company's own workloads is worth the enormous cost and complexity of building and maintaining a chip program, rather than remaining fully dependent on one external supplier for the single most important input to the AI business.

The numbers on how fast this shift is happening are themselves notable. Custom AI ASIC shipments are growing at roughly 44.6% year-over-year in 2026 — nearly triple the 16.1% growth rate of merchant GPUs like Nvidia's — and ASIC-based AI servers are projected to reach 27.8% of total AI server shipments in 2026, the highest share the category has held since 2023. That's not a niche experiment; it's more than a quarter of the market shifting toward chips designed and controlled by the hyperscalers themselves rather than purchased as a merchant product from Nvidia.

Meet the Chips: A Field Guide to Hyperscaler Silicon

Each hyperscaler's chip program reflects a different history, a different primary workload, and a different level of program maturity. Understanding the differences matters more than treating "custom AI chip" as one undifferentiated category.

Google TPU v7 "Ironwood": The Most Mature Program by Far

Google's Tensor Processing Unit line is the oldest of the group by a wide margin — it launched in 2013, more than a decade before most of its current rivals started their own programs. By 2026, Google is on its seventh generation, known as "Ironwood," announced in April 2025. Ironwood specs read like a genuinely top-tier accelerator in its own right: 4,614 TFLOPS of FP8 performance, 192GB of high-bandwidth memory per chip, and the ability to scale into "superpods" of up to 9,216 chips working together on a single workload. Google is projected to ship roughly 4.3 million TPU units in 2026 — a scale that reflects both how central TPUs have become to Google's own AI products and how much capacity Google has built to support external cloud customers renting TPU access through Google Cloud.

That decade-plus head start shows up concretely in how the program runs today. Customer wait times for committed TPU capacity run around 2 to 3 months — a genuinely fast turnaround for accessing a specialized AI accelerator at scale — reflecting years of production experience, supply chain relationships, and software ecosystem maturity that newer programs simply haven't had time to build yet.

AWS Trainium3: Amazon's First 3nm Chip

Amazon's approach splits the workload into two purpose-built chip families: Trainium, aimed at training large models, and Inferentia, aimed specifically at inference. Trainium3, which became generally available around AWS's re:Invent conference, is notable as AWS's first chip built on a 3-nanometer manufacturing process — a meaningful jump in transistor density and efficiency over its predecessors. AWS has already proven the model at real scale: more than 500,000 Trainium2 chips were deployed as of late 2025, before Trainium3 even reached general availability, underscoring how much production AI workload AWS has already shifted onto its own silicon rather than exclusively renting Nvidia capacity for it.

The two-chip strategy — training and inference handled by architecturally distinct chips rather than one general-purpose accelerator doing both — reflects a cost-disruption logic: a chip optimized narrowly for one job can be more efficient at that job than a more flexible chip trying to do both training and inference reasonably well. AWS's public cost claims put Trainium's inference savings at up to 50% versus equivalent Nvidia GPU-based inference, a substantial number for any AWS customer running inference at real production volume.

Microsoft Maia 200: The Newest, Most Aggressive Entrant

Microsoft announced Maia 200 in January 2026, and the company has made some of the boldest public performance claims in the entire custom-silicon field: Microsoft says Maia 200 delivers roughly 3x the FP4 performance of AWS's Trainium3, and higher FP8 performance than Google's seventh-generation TPU. Those are Microsoft's own comparative claims rather than independently verified third-party benchmarks, but they signal how aggressively Microsoft is positioning Maia against the two more established programs it's chasing.

The gap between ambition and current deployment reality is visible in two numbers worth holding side by side: committed Maia capacity currently carries customer wait times of 18 to 24 months — many times longer than Google's 2-to-3-month TPU wait — and despite Maia's existence, roughly 70% of Azure's AI workloads still run on Nvidia hardware. Building a competitive chip and building the production capacity, software stack, and customer confidence to actually shift a majority of a hyperscaler's own workload onto it are two very different achievements, and Microsoft's numbers show it is still much further along on the first than the second.

Meta MTIA: Built for Meta, Not for Sale

Meta's MTIA (Meta Training and Inference Accelerator) stands apart from the other three programs in one specific way: it was purpose-built for Meta's own recommendation and ad-ranking systems — the algorithms that decide what shows up in a Facebook or Instagram feed and which ads get served — rather than as a general-purpose AI chip Meta sells or rents to outside customers the way Google and AWS both do through their cloud platforms.

That narrower purpose makes MTIA's existence, and Meta's parallel behavior, one of the more interesting data points in the whole custom-silicon story. Despite running its own MTIA program, Meta is reportedly pursuing a multi-billion-dollar deal with Google for TPU capacity specifically to run large language model inference — a workload MTIA wasn't originally designed around. That's a straightforward admission that a chip optimized for one job (recommendation ranking) doesn't automatically transfer well to a very different job (LLM inference), and that even a hyperscaler with its own multi-year chip program will still shop externally, including from a direct hyperscaler competitor, when a different chip is genuinely better suited to a specific new workload.

OpenAI's Broadcom Partnership: The Newest Entrant of All

OpenAI, not traditionally categorized as a hyperscaler but now operating at a scale that requires hyperscaler-level infrastructure thinking, is developing its own custom ASIC in partnership with Broadcom — a roughly $10 billion arrangement, currently in pre-production and targeting Q3 2026 for its first real deployment. The program is reportedly led by Richard Ho, an engineer who previously worked on Google's TPU program, which is a notable detail in its own right: OpenAI's approach to custom silicon is being shaped by someone who helped build the most mature program in the entire field.

Broadcom's role here extends well beyond OpenAI specifically — the company has built a dominant custom AI accelerator design-services business, holding more than 70% share of that specific market and a reported $73 billion in committed customer backlog. Broadcom doesn't compete as a chip brand the way Nvidia or the hyperscalers do; it's the design and manufacturing partner making several of these hyperscaler and AI-lab chip programs possible in the first place, which makes it one of the most structurally important — if less publicly visible — companies in this entire story.

The Economics: Why 50-67% Cheaper Inference Justifies the Cost

Designing a custom AI chip is not cheap. First-year design costs for a serious custom AI accelerator program run anywhere from $10 million to well over $100 million, before accounting for the manufacturing capacity, software stack, and specialized engineering talent needed to actually deploy it at scale. That's a serious capital commitment, and it only makes sense against the specific economics driving this entire race: 50% to 67% lower inference costs compared with running the equivalent workload on Nvidia GPUs.

The clearest illustrated example is Google's: eight TPU v5e chips running large language model inference cost around $11 per hour, representing a 65% to 67% reduction compared with equivalent Nvidia H100-based inference. AWS makes a similar claim for Trainium, citing inference savings of up to 50%. At the volume hyperscalers run inference — billions of queries across deployed products — even a 50% reduction in per-query compute cost compounds into an enormous absolute savings, easily justifying design costs that would be unthinkable for a smaller company running a fraction of the workload.

This is also where the concept of vertical integration becomes strategically central rather than just an engineering preference. Vertical integration, in this context, means a hyperscaler controlling the full stack from chip design through the software and cloud infrastructure that chip runs inside — rather than depending on a third party's chip roadmap, pricing, and allocation decisions for the most expensive input in its AI business. A hyperscaler that owns its chip design can optimize that chip precisely for its own dominant workloads (recommendation ranking for Meta, large-scale cloud rental for Google and AWS, enterprise AI services for Microsoft) in a way a general-purpose merchant GPU, designed to serve every customer's differing workload reasonably well, structurally cannot match for any one specific use case.

The Shared Bottleneck: Everyone Still Needs TSMC and SK Hynix

Here's the detail that keeps this entire race from being a clean story about hyperscalers escaping dependency on a single supplier: virtually every custom AI chip discussed here — TPU, Trainium, Maia, MTIA, and OpenAI's Broadcom-designed chip — is manufactured on TSMC's 3-nanometer process, the exact same advanced manufacturing node Nvidia's own chips depend on. TSMC's 3nm capacity is reportedly running at 100% utilization, with demand roughly three times available supply, and TSMC as a company manufactures around 92% of the world's most advanced AI chips across all customers combined. A disruption to TSMC's Taiwan-based production — from a natural disaster, a geopolitical event, or any other serious shock — has been estimated to potentially cost the global economy around $2.5 trillion annually, a figure that reflects just how concentrated the entire industry's manufacturing dependency has become, custom silicon included.

The memory side of the supply chain tells a similar story. SK Hynix holds roughly 62% share of the high-bandwidth memory (HBM) market — the specialized, tightly stacked memory every modern AI accelerator needs to feed data to its processing cores fast enough to be useful — and that concentration has contributed to a broader DRAM price surge as AI-related demand for memory has outpaced supply. A hyperscaler that spends billions building its own chip design team still needs HBM from the same small handful of suppliers Nvidia and every other chip maker in the industry is also drawing from, which means the custom-silicon race reduces one specific dependency (on Nvidia as a chip vendor) without eliminating the underlying manufacturing and materials dependencies the whole AI hardware industry shares.

That's a genuinely uncomfortable conclusion for an industry narrative built heavily around the idea of reducing dependency. A hyperscaler that spends years and hundreds of millions of dollars building its own chip design team can still find its entire program exposed to a single geopolitical event or logistics disruption on the other side of the world, in exactly the same way Nvidia's own supply is exposed today. Diversifying who designs a chip is a real and valuable form of risk reduction — it reduces exposure to one company's pricing power and allocation decisions during a shortage — but it is a narrower, more limited kind of diversification than reducing exposure to a single manufacturing region and a single memory supplier, which remains a layer of risk this entire industry, custom silicon included, has not found a way around.

Geography compounds the problem further. TSMC's most advanced manufacturing capacity sits overwhelmingly on the island of Taiwan, a concentration that predates the AI boom by decades but has become far higher-stakes now that trillions of dollars in hyperscaler AI investment sit downstream of it. Every custom chip program discussed in this piece — however differentiated its architecture, however independent its design team — ultimately converges on the same physical manufacturing footprint, which means the industry's actual geographic risk concentration hasn't meaningfully changed even as the number of distinct chip designers has multiplied from one dominant vendor to five or more competing programs.

The Maturity Gap: A Decade of Head Start Doesn't Disappear Overnight

Google's TPU program is genuinely in a different competitive position than the others simply because of when it started. Google began building TPUs in 2013; Microsoft's Maia program, by contrast, is generally understood to have started around 2019 — roughly six years later. Over a decade of production experience shows up in ways that are hard for a newer program to simply out-spend its way past: manufacturing yield improvements, supply chain relationships built over multiple chip generations, a mature software and compiler stack that makes the hardware usable for a wide range of workloads, and enough deployed capacity that customer wait times sit at 2 to 3 months rather than the 18-to-24-month wait Microsoft's Maia program currently carries for committed capacity.

That maturity gap is also visible in how much of each company's own workload actually runs on its custom chip today versus how much still runs on Nvidia. Azure — despite Maia's existence and Microsoft's bold performance claims for it — still runs roughly 70% of its AI workloads on Nvidia hardware. That's a useful corrective to any narrative that custom silicon is simply replacing Nvidia GPUs across the industry: even a company actively building and marketing its own competing chip is, by its own current deployment mix, still overwhelmingly reliant on the incumbent it's trying to reduce dependence on.

Complementary, Not Replacement: Why OpenAI Buys From Everyone

Perhaps the clearest evidence that custom ASICs and Nvidia GPUs are complementary tools rather than a straightforward replacement story is OpenAI's own procurement behavior. OpenAI is simultaneously developing its own Broadcom-partnered ASIC, maintaining a large-scale relationship with Nvidia for GPU capacity, and pursuing a separate deal with Cerebras — three different hardware strategies running in parallel rather than a sequential replacement of one by another. That pattern, highlighted directly in 2026 industry analysis, suggests the mature view of this race isn't "custom silicon versus Nvidia" but rather each type of chip finding the workloads and cost profiles it's best suited for, with sophisticated AI infrastructure buyers running a deliberately diversified hardware portfolio rather than betting everything on one supplier or one architecture.

Meta's own behavior reinforces the same point from a different angle. A company with its own MTIA program still finds it worthwhile to pursue billions of dollars in TPU capacity from Google — a direct hyperscaler competitor — specifically because MTIA wasn't built for the LLM-inference workload Meta now needs at scale. The practical lesson generalizing across this entire race: owning a custom chip program doesn't mean abandoning every other hardware option; it means having more leverage and more options across a genuinely diversified hardware strategy than a company fully dependent on a single external vendor would have.

The Hidden Cost: Software Fragmentation Across Five Different Chip Architectures

None of this comes for free on the software side, and it's a cost that gets far less attention than the chip announcements themselves. Nvidia's dominant market position was never just about hardware — its CUDA software platform, built up over roughly two decades, is what let developers write AI code once and reasonably expect it to run efficiently across generations of Nvidia hardware. Every custom chip discussed in this piece needs its own equivalent software and compiler stack, and none of them is a drop-in replacement for CUDA: code and models tuned for TPU don't automatically run efficiently on Trainium, Maia, or MTIA, and vice versa.

That fragmentation creates a real trade-off for any company deciding where to run its AI workloads. Chasing the lowest-cost chip for a given workload — TPU for one job, Trainium for another, Maia for a third — means maintaining engineering expertise across multiple, architecturally distinct software ecosystems, each with its own quirks, performance characteristics, and debugging tools. Standardizing on a single provider's silicon, by contrast, sacrifices some of the cost-optimization benefit in exchange for a simpler, more maintainable engineering setup. Large hyperscalers can absorb that complexity because they employ specialized teams for exactly this purpose; a smaller company evaluating where to run its own AI workloads has to weigh the same trade-off with far fewer engineers available to manage it. This is part of why Nvidia's ecosystem, whatever its cost premium, still commands the loyalty it does even from customers who could technically save money moving workloads to custom silicon: the software maturity and portability it offers has its own real value, separate from the raw price-per-chip comparison.

Open standards and cross-platform frameworks are the industry's partial answer to this fragmentation, letting some workloads move between chip architectures with less rework than a fully custom integration would require. But no framework fully erases the differences between five genuinely distinct hardware architectures optimized around different priorities, and any company weighing whether to chase custom-silicon cost savings needs to price in the engineering cost of that fragmentation alongside the sticker-price savings on the chip itself.

The Global Picture: A Story Concentrated in the United States, With Europe and China Watching From the Sidelines

United States

The hyperscaler custom-silicon race is, almost in its entirety, a US story. Google, Amazon, Microsoft, Meta, and OpenAI — every company discussed above — are American companies, and virtually all of the concrete 2026 developments in this space (Ironwood's launch, Trainium3's general availability, Maia 200's announcement, MTIA's deployment, OpenAI's Broadcom partnership) happened within the US corporate and technology ecosystem. That concentration mirrors the broader pattern in AI infrastructure spending generally, where the handful of companies with the balance sheets to fund multi-billion-dollar chip design programs are overwhelmingly headquartered in the United States.

United Kingdom

The UK's most relevant player is Graphcore, the British maker of the IPU (Intelligence Processing Unit), which has a long-standing strategic partnership with the Franco-German company SiPearl on AI and high-performance-computing silicon. This research pass found no fresh 2026-specific update on that partnership, so it's worth treating as useful background context on Europe's alternative chip ecosystem rather than a confirmed current development. Graphcore's IPU is also a fundamentally different kind of product than the hyperscaler chips discussed above — it's a chip Graphcore sells commercially to outside customers, closer in business model to Nvidia's merchant GPU approach than to an internal hyperscaler ASIC like TPU or Trainium that a single company builds primarily for its own use.

UAE and Dubai

No distinct regional-specific reporting on the custom-silicon race was found for the UAE or Dubai in this research pass. The region's more prominent AI infrastructure story in 2026 centers on dedicated power generation for large AI campus projects rather than chip design specifically, which is a different dimension of AI infrastructure investment than the topic covered here.

Australia

Similarly, no distinct regional-specific reporting on custom AI silicon was found for Australia. Australia's more prominent 2026 AI infrastructure story concerns data center grid connection and power policy rather than chip design.

Germany

SiPearl, the Franco-German company developing HPC and AI microprocessors under the European Processor Initiative, is headquartered partly in Germany. As with the UK's Graphcore partnership, this research pass found no fresh 2026-specific silicon milestone to report distinct from that ongoing background program.

Wider Europe and France

The same SiPearl and European Processor Initiative context applies across France and the rest of Europe more broadly. No confirmed new 2026 silicon-specific milestone distinct from Europe's broader sovereign-AI infrastructure investment turned up in this research pass. The practical read is that Europe's custom-silicon ambitions, embodied in programs like SiPearl, remain a longer-term sovereign-technology project rather than a competitor currently shipping chips at the volume or pace of the American hyperscalers' programs.

China

China's answer to this race looks structurally different from the American model. Huawei's Ascend accelerator line functions as China's domestic response, but it's better understood as a merchant-style accelerator — sold to many different customers across China's AI industry — than as an internal hyperscaler ASIC built by and for one company's own workloads the way TPU or Trainium are. That distinction matters: China's approach to reducing dependence on external chip suppliers (in its case, driven heavily by export restrictions rather than cost optimization) has produced something closer to a domestic Nvidia alternative than a domestic Google-TPU-style program, reflecting a different starting constraint and a different strategic goal than what's driving the American hyperscaler race.

What This Means Going Forward

The practical takeaway for anyone building AI-dependent products isn't that custom silicon is about to make Nvidia irrelevant — the data here says clearly it hasn't, even at the hyperscalers driving the trend. The takeaway is that AI infrastructure costs and performance characteristics are becoming genuinely differentiated across providers in a way they weren't a few years ago, when renting Nvidia GPU capacity from any major cloud provider was a fairly interchangeable choice. Where a workload runs, and which underlying chip it runs on, now has real, quantifiable cost implications — the kind of decision worth understanding rather than treating as an invisible implementation detail buried inside a cloud bill.

For a business evaluating where and how to build AI-driven products, that's a genuine reason to work with a technology partner who tracks this infrastructure layer as carefully as the application layer sitting on top of it. Our AI agents and automation work is built with an eye toward exactly this kind of infrastructure reality — designing systems that don't assume one vendor's pricing and availability will hold indefinitely. If you're weighing providers or trying to understand the trade-offs between different AI infrastructure options, our comparisons hub covers how we think through platform decisions like this more generally, and our glossary is a useful reference for terms like ASIC, HBM, and vertical integration that come up constantly in this space but rarely get defined plainly.

A few practical questions are worth asking of any AI infrastructure decision made against this backdrop. Is the workload latency-sensitive or cost-sensitive, since the two custom-silicon families in this piece optimize for different priorities and neither wins on every dimension at once. Is the team building on a given provider's silicon prepared to maintain that provider-specific expertise long-term, given the software fragmentation described above, or would standardizing on the more portable Nvidia ecosystem be worth the cost premium for a smaller team. And is the underlying business durable enough to weather a multi-year commitment to one hyperscaler's chip roadmap, given how differently mature these programs still are from one another. None of these questions has a universally correct answer — the right choice depends heavily on workload, team size, and growth trajectory — but treating chip architecture as a genuine strategic decision, rather than an invisible detail abstracted away by a cloud bill, is itself the mindset shift this whole race is forcing on anyone building seriously at scale.

Straight Answers on the Hyperscaler Chip Race

Why are hyperscalers investing billions in custom AI chips instead of just buying more Nvidia GPUs?

Because the economics and the strategic risk both point the same direction. Custom silicon can cut inference costs by 50% to 67% compared with running the same workload on Nvidia GPUs, and at the volume hyperscalers run inference — billions of queries across deployed products — that difference compounds into enormous savings that justify even a $10 million to $100 million-plus first-year design investment. On top of the cost case, there's a supply-independence motive: Nvidia holds roughly 86% of the AI accelerator market with around 73% gross margins, and depending entirely on one supplier for the single most important input to an AI business is a real concentration risk, especially during periods when GPU demand has outstripped available supply. Owning custom silicon doesn't remove every dependency — the manufacturing and memory supply chain is still shared — but it gives a hyperscaler more control and more negotiating leverage than depending on Nvidia alone.

What is Google's TPU v7 'Ironwood' and how does it compare to earlier TPU generations?

Ironwood is Google's seventh-generation Tensor Processing Unit, announced in April 2025, and it represents the most mature custom AI chip program of any hyperscaler discussed here — Google has been building TPUs since 2013. Ironwood delivers 4,614 TFLOPS of FP8 performance with 192GB of high-bandwidth memory per chip, and it can scale into "superpods" combining up to 9,216 chips working together on a single workload. Compared with earlier TPU generations, Ironwood reflects over a decade of iterative improvement in both the chip itself and the surrounding software and manufacturing ecosystem, which is part of why Google can offer relatively fast customer wait times — around 2 to 3 months — for committed TPU capacity, a turnaround newer custom-silicon programs haven't yet matched.

How many TPU chips is Google projected to ship in 2026?

Google is projected to ship approximately 4.3 million TPU units in 2026. That volume reflects both how deeply embedded TPUs are in Google's own AI products (search, workspace AI features, its Gemini model family) and how much external demand Google Cloud is serving from customers renting TPU capacity rather than Nvidia GPU capacity for their own AI workloads. It's a scale of production that took over a decade of iterative chip generations to reach, underscoring how much of a head start Google's program has over rivals that only began their own custom-silicon efforts in the last several years.

What is AWS Trainium3 and why is it significant as AWS's 'first 3nm chip'?

Trainium3 is Amazon's latest custom AI training chip, which reached general availability around AWS's re:Invent conference, and it's significant specifically as AWS's first chip built on a 3-nanometer manufacturing process — a meaningful step up in transistor density and power efficiency compared with the larger manufacturing nodes AWS's earlier Trainium and Inferentia chips used. Moving to 3nm puts AWS on the same advanced manufacturing node as Nvidia's latest GPUs and the other hyperscalers' newest chips, which matters competitively because manufacturing node is one of the clearest technical proxies for how efficient and performant a chip can be per unit of power consumed — a critical factor given how central power constraints have become to AI infrastructure economics.

Why does AWS split its custom silicon into two separate chips — Trainium for training and Inferentia for inference?

AWS's reasoning is a cost-disruption strategy built on workload specialization: training and inference are genuinely different computational jobs, with different optimal trade-offs between raw compute throughput, memory bandwidth, and power efficiency. A chip narrowly optimized for one job can outperform, on a cost-per-workload basis, a more flexible chip trying to handle both training and inference reasonably well. By building Trainium specifically for training and Inferentia specifically for inference, AWS can tune each chip's architecture toward its single job rather than compromising either chip's design to accommodate the other — a trade-off AWS's public cost claims suggest pays off, citing inference savings of up to 50% compared with equivalent Nvidia GPU-based inference.

How many Trainium2 chips has AWS already deployed as of late 2025?

AWS had already deployed more than 500,000 Trainium2 chips as of late 2025, before Trainium3 even reached general availability. That figure is a meaningful data point on its own: it shows AWS wasn't treating custom silicon as an experimental side project but had already shifted a substantial volume of real production AI workload onto its own chips at a scale most companies' entire AI infrastructure never approaches, well before the newer, more advanced Trainium3 generation arrived to build on that foundation.

That scale matters for how seriously the rest of the industry has to take AWS's custom-silicon program. Deploying half a million chips isn't a pilot or a proof-of-concept run — it requires manufacturing partnerships, data center power and cooling capacity, and software tooling mature enough to run real customer and internal workloads reliably at volume. It puts AWS's program in a similar league of production maturity to Google's TPU effort, even though AWS started later, and it's a large part of why AWS can credibly claim inference cost savings of up to 50% against Nvidia GPUs — those savings are backed by genuine production experience, not projected estimates from a small test deployment.

What is Microsoft's Maia 200 chip and what performance claims has Microsoft made for it versus Trainium3 and Google's TPU?

Maia 200 is Microsoft's latest custom AI accelerator, announced in January 2026 and deployed within Azure. Microsoft has made notably aggressive comparative performance claims for it: roughly 3x the FP4 performance of AWS's Trainium3, and higher FP8 performance than Google's seventh-generation TPU (Ironwood). Those figures come from Microsoft itself rather than independent third-party benchmarking, so they should be read as the company's own positioning rather than verified neutral comparisons — but the scale of the claims signals how aggressively Microsoft wants Maia perceived relative to the two more established programs it's chasing, even though Maia's actual production deployment (evidenced by 18-to-24-month capacity wait times and Azure's continued heavy reliance on Nvidia) is still considerably behind Google's and, in deployment volume, likely behind AWS's as well.

Why does Microsoft's Azure still run roughly 70% of its AI workloads on Nvidia despite building Maia?

This is one of the clearest illustrations that having a custom chip program and actually shifting the bulk of production workload onto it are two very different milestones. Building a competitive chip requires solving hardware design problems; shifting 70% of a hyperscaler's existing Nvidia-based workload onto a new architecture requires an entirely separate set of achievements — mature software tooling, proven reliability at scale, customer confidence, and enough manufactured supply to actually serve that much demand. Maia 200's 18-to-24-month wait times for committed capacity suggest Microsoft's production capacity simply hasn't caught up to its ambitions yet, and until it does, the pragmatic choice for both Microsoft's own workloads and its Azure customers is to keep running the majority of AI compute on the mature, proven Nvidia ecosystem while Maia's production capacity and software maturity continue to build.

What is Meta's MTIA chip built for, and why is it not commercialized externally?

MTIA — Meta's Training and Inference Accelerator — was purpose-built for Meta's own recommendation and ad-ranking systems, the algorithms that decide what content and ads appear in Facebook and Instagram feeds. That's a different design target than a general-purpose AI accelerator meant to handle a wide range of customer workloads, which is exactly why Meta hasn't commercialized MTIA externally the way Google and AWS rent out TPU and Trainium capacity through their cloud platforms — MTIA's architecture is tuned closely to Meta's specific, massive-scale recommendation workload rather than built as a flexible product for outside customers running arbitrary AI workloads.

Why is Meta reportedly pursuing a multi-billion-dollar TPU deal with Google despite having its own MTIA program?

Because MTIA was designed for recommendation and ad-ranking, not for large language model inference — a newer, different workload Meta now needs to run at scale as it builds out its own generative AI products. Rather than redesigning or repurposing MTIA for a job it wasn't built for, Meta is reportedly turning to Google's TPU program specifically for LLM inference capacity, even though Google is a direct competitor in several other markets. It's a pragmatic illustration of a broader pattern in this whole story: even a hyperscaler with a mature internal chip program will still buy externally, including from a rival, when a different chip is genuinely better suited to a specific workload than what it has built in-house.

What is OpenAI's own custom ASIC project with Broadcom, and when is it expected in production?

OpenAI is developing a custom AI accelerator in partnership with Broadcom, in a deal reportedly worth around $10 billion. The project is currently in pre-production and is targeting Q3 2026 for its first real deployment. It's reportedly led by Richard Ho, an engineer who previously worked on Google's TPU program — a detail that suggests OpenAI is deliberately drawing on the expertise behind the industry's most mature custom-silicon effort as it builds its own. The project sits alongside, rather than replacing, OpenAI's existing large-scale relationships with Nvidia and a separate deal with Cerebras, reflecting the same complementary-hardware-strategy pattern seen across the rest of this industry.

How large is Broadcom's custom AI accelerator design-services business, and how fast is it growing?

Broadcom holds more than 70% share of the custom AI accelerator design-services market — the business of helping other companies (hyperscalers, AI labs) actually design and bring their own chip programs to life — and carries a reported $73 billion in committed customer backlog. Those figures make Broadcom one of the most structurally important companies in the entire custom-silicon story, even though it doesn't sell a chip under its own brand the way Nvidia does. Several of the specific programs discussed throughout this piece, including OpenAI's, depend directly on Broadcom's design and manufacturing partnership to exist at all.

How much cheaper is inference on Google's TPU v5e compared with Nvidia H100 GPUs?

Running large language model inference on eight TPU v5e chips costs around $11 per hour, which represents a 65% to 67% cost reduction compared with running the equivalent inference workload on Nvidia H100-based GPUs. That's one of the more concrete, quantified cost comparisons in the entire custom-silicon conversation, and it illustrates directly why the economics of inference cost — not just training cost — are such a central driver of hyperscaler chip investment: a savings that large, applied across the enormous inference volume a major cloud provider runs, adds up to a substantial sum very quickly.

What cost savings does AWS claim for Trainium versus Nvidia GPU inference?

AWS has publicly claimed inference cost savings of up to 50% for Trainium compared with equivalent Nvidia GPU-based inference. While that's a somewhat smaller claimed reduction than Google's cited figures for TPU v5e, it's still a substantial saving at the scale AWS operates, and it reinforces the same underlying logic driving every hyperscaler's custom-silicon investment: even a partial reduction in per-query inference cost, multiplied across billions of queries running on AWS infrastructure, represents an enormous absolute cost saving that justifies the chip design investment many times over.

Why is Google's TPU program considered roughly six years more mature than Microsoft's Maia program?

Google began developing TPUs in 2013, while Microsoft's Maia program is generally understood to have started around 2019 — a gap of roughly six years. In an industry where chip design, manufacturing yield improvement, and software ecosystem maturity all compound with each successive generation, a six-year head start is substantial: Google has had multiple additional chip generations, multiple additional rounds of manufacturing process improvement, and years more time to build the compiler and software tooling that makes a chip actually usable at scale for a wide range of AI workloads. That maturity gap is exactly what shows up in the customer wait-time comparison — 2 to 3 months for TPU capacity versus 18 to 24 months for Maia.

What customer wait times exist for Google TPU capacity versus Microsoft Maia capacity in 2026?

Customers seeking committed Google TPU capacity currently face wait times of roughly 2 to 3 months, while customers seeking committed Microsoft Maia capacity face wait times of 18 to 24 months — roughly six to twelve times longer. That gap is one of the clearest, most concrete indicators of how far apart these two programs are in production maturity: Google has had over a decade to build the manufacturing capacity, supply chain relationships, and operational experience to serve demand quickly, while Microsoft's newer program is still scaling up production to meet the demand its own aggressive marketing and performance claims have generated.

Why does virtually every hyperscaler custom AI chip get manufactured on TSMC's 3nm process?

Because TSMC is, by a wide margin, the world's leading manufacturer of the most advanced chip nodes needed for cutting-edge AI accelerators, producing around 92% of the world's most advanced AI chips across all customers combined. Every major custom AI chip discussed in this piece — TPU, Trainium, Maia, MTIA, and OpenAI's Broadcom-designed chip — depends on that same 3-nanometer manufacturing capacity, which is reportedly running at 100% utilization with demand roughly three times available supply. That shared dependency means the custom-silicon race, for all its talk of independence from Nvidia, hasn't actually diversified the industry's underlying manufacturing risk — it has just redistributed who's designing the chips that all still need the same scarce manufacturing capacity to actually get built.

What would a disruption at TSMC mean for the hyperscaler custom-silicon race?

Given that TSMC manufactures roughly 92% of the world's most advanced AI chips, a serious disruption to its Taiwan-based production — whether from a natural disaster, a geopolitical event, or any other major shock — would affect essentially every player in this story simultaneously: Nvidia, Google's TPU program, AWS's Trainium, Microsoft's Maia, Meta's MTIA, and OpenAI's Broadcom-designed chip would all be competing for whatever alternative or reduced manufacturing capacity remained. Estimates have put the potential global economic cost of such a disruption at around $2.5 trillion annually, a figure that reflects how concentrated and singular this manufacturing dependency has become across an industry that otherwise talks a great deal about diversification and independence.

Why does SK Hynix's roughly 62% share of the HBM market create a bottleneck for custom AI chips as well as Nvidia GPUs?

Every modern AI accelerator — custom ASIC or merchant GPU alike — depends on high-bandwidth memory (HBM) to feed data to its processing cores fast enough to be useful, and SK Hynix controls roughly 62% of that specific memory market. That concentration means a hyperscaler's custom chip program, however independent its design and manufacturing partnership might be from Nvidia specifically, still competes for the same limited pool of HBM supply that Nvidia and every other AI chip maker also depends on — a dependency that has contributed to a broader DRAM price surge as AI-driven memory demand has outpaced available supply across the whole industry, not just for any single company's chips.

Are custom ASICs and Nvidia GPUs complementary or competitive at hyperscale, given OpenAI is pursuing both simultaneously?

The evidence points clearly toward complementary rather than strictly competitive. OpenAI is simultaneously developing its own Broadcom-partnered ASIC, maintaining a large-scale Nvidia GPU relationship, and pursuing a separate deal with Cerebras — three different hardware strategies running in parallel rather than one replacing another. Industry analysis frames this directly: ASICs and GPUs remain complementary at hyperscale, not mutually exclusive, because different chips suit different workloads, cost profiles, and availability constraints, and a sophisticated AI infrastructure buyer benefits from a deliberately diversified hardware portfolio rather than betting everything on a single architecture or supplier.

How much of total AI server shipments in 2026 are custom ASICs expected to represent?

Custom ASIC-based AI servers are projected to reach 27.8% of total AI server shipments in 2026 — the highest share the category has held since 2023. That's more than a quarter of the entire AI server market shifting toward chips designed and controlled by the hyperscalers building them, rather than merchant GPUs purchased from Nvidia, and it's a clear quantitative signal of how far this shift has already progressed rather than remaining a speculative future trend.

Put another way, roughly one in every four AI servers shipping worldwide in 2026 is built around a chip one specific hyperscaler designed primarily for its own workloads, rather than a general-purpose accelerator any customer can buy from Nvidia. That share is still a minority of the overall market — merchant GPUs remain the majority choice, and Azure's own 70%-Nvidia workload mix shows even chip-building hyperscalers lean heavily on merchant silicon today — but a share that large, growing at nearly triple the rate of merchant GPUs, is no longer a rounding error in any serious infrastructure planning conversation.

Why is custom ASIC shipment growth (44.6% YoY) nearly triple that of merchant GPUs (16.1% YoY) in 2026?

The growth-rate gap reflects custom silicon programs scaling up from a smaller base while benefiting from the specific cost and supply-independence motives driving hyperscaler investment discussed throughout this piece — 50% to 67% lower inference costs and reduced dependence on a single dominant supplier. Merchant GPU shipments are still growing at a healthy 16.1% annually, reflecting continued strong overall AI infrastructure demand, but custom ASICs are capturing a disproportionate share of that growth because hyperscalers have direct control over how quickly they scale their own chip production and how much of their own workload they shift onto it, in a way they don't have control over Nvidia's production allocation decisions.

What is the European Processor Initiative and what role does SiPearl play in it?

The European Processor Initiative is a European program aimed at developing high-performance-computing and AI microprocessor technology within Europe, reducing the continent's dependence on chip designs from the US and Asia. SiPearl, a Franco-German company headquartered partly in Germany, is the initiative's most prominent chip-development effort. This research pass found no confirmed new 2026-specific silicon milestone for SiPearl distinct from Europe's broader sovereign-AI infrastructure investment, so it's best understood as an ongoing, longer-term strategic program rather than a chip currently shipping at hyperscaler volume or pace.

How does Britain's Graphcore fit into the global custom-AI-silicon landscape relative to the US hyperscaler chips?

Graphcore, the British maker of the IPU (Intelligence Processing Unit), occupies a different position than TPU, Trainium, Maia, or MTIA: it's a commercial chip Graphcore sells to outside customers, closer in business model to Nvidia's merchant approach than to an internal hyperscaler ASIC built by one company primarily for its own workloads. Graphcore has a long-standing strategic partnership with SiPearl on AI and HPC silicon, though this research pass found no fresh 2026-specific update on that relationship. Relative to the American hyperscaler programs, which are backed by some of the largest technology balance sheets in the world and integrated directly into companies' own massive-scale cloud and product infrastructure, Graphcore represents a smaller, independent commercial alternative rather than a comparable-scale competitor.

Why do hyperscalers keep building their own chips even though first-year design costs run $10M-$100M or more?

Because the economics work out favorably at hyperscale volume. A custom chip program that costs $10 million to over $100 million in its first year is a serious investment, but set against 50% to 67% inference cost savings applied across the billions of queries a major hyperscaler's AI products run, that design cost can be recovered quickly and then continue generating savings for as long as the chip remains in production use. The math simply doesn't work the same way for a smaller company running a fraction of that inference volume — which is exactly why this race is concentrated among a handful of the largest technology companies in the world rather than being a broadly accessible strategy.

What does 'vertical integration' mean in the context of hyperscaler custom silicon, and why does it matter strategically?

Vertical integration here means a hyperscaler controlling the full stack from chip design through the cloud infrastructure and software that chip ultimately runs inside, rather than depending on an external vendor's chip roadmap, pricing, and supply-allocation decisions for the most critical input to its AI business. It matters strategically for two connected reasons: it lets a company optimize a chip precisely for its own dominant workloads in a way a general-purpose merchant chip can't match for any single use case, and it insulates the business, at least partially, from a single supplier's pricing power and allocation choices during periods of tight supply — exactly the situation Nvidia's roughly 86% market share and 73% gross margins created across the AI accelerator market.

Is Huawei's Ascend accelerator best understood as China's version of a hyperscaler custom chip, or as a merchant GPU competitor?

Huawei's Ascend line functions more like a merchant GPU competitor than an internal hyperscaler ASIC. Chips like TPU, Trainium, and Maia were built by and primarily for a single company's own workloads, with external rental as a secondary business; Ascend, by contrast, is sold across China's broader AI industry to many different customers, positioning it closer to a domestic alternative to Nvidia itself than to a domestic equivalent of Google's TPU program. That distinction reflects China's different underlying constraint — driven significantly by export restrictions on advanced foreign chips — and a correspondingly different strategic goal: building a broadly available domestic chip supply for an entire market, rather than one company optimizing silicon narrowly around its own specific, dominant workload.

Want results like this?

Keep reading