Microsoft introduces Maia 200 AI inference accelerator, built on 3nm process
Microsoft has announced Maia 200, a new AI inference accelerator aimed at improving the cost and speed of running models in Azure. The company says the chip focuses on inference, uses a 3-nanometer process and is being deployed initially in select US cloud regions, alongside a software toolkit for developers.
What Microsoft announced
Microsoft has introduced Maia 200, a next-generation AI accelerator designed primarily for inference workloads—the phase where trained AI models are deployed to serve real users in applications. The company described Maia 200 as an effort to improve the economics of AI token generation and to run large models faster and more efficiently on Azure infrastructure.

According to Microsoft’s announcement, Maia 200 is built on TSMC’s 3-nanometer process and is positioned as a significant step in the company’s first-party silicon roadmap. The chip is accompanied by a preview of a Maia software development kit (SDK) intended to help developers build and optimise workloads for Maia hardware.
Deployment plans and infrastructure focus
Microsoft said Maia 200 is already deployed in its US Central datacenter region near Des Moines, Iowa, with US West 3 near Phoenix, Arizona planned next, and additional regions to follow. The company framed Maia 200 as part of its broader heterogeneous AI infrastructure, where different types of chips and systems work together for varied AI tasks.
The emphasis on inference reflects a key industry challenge: scaling AI features to more users without exploding power and cloud-compute costs. If inference becomes cheaper per unit of output, AI assistants and business applications can offer richer features, longer contexts and more frequent usage while staying within budgets.
Why this matters in the AI chip race
Large cloud providers have been investing heavily in custom silicon to reduce dependence on a single GPU supplier, manage costs, and tailor hardware for their own AI stacks. Maia 200 adds to this trend by putting Microsoft more directly into the competitive landscape alongside other in-house accelerators used by hyperscalers.
What to watch next
- When Maia 200 becomes available across more Azure regions outside the US.
- How developers and enterprise customers experience performance and cost in real deployments.
- Updates to the Maia SDK and integration with common AI frameworks.
- Whether Microsoft expands its chip roadmap into additional model-serving or training use cases.
For the broader ecosystem, Maia 200’s impact will be measured not only in headline specifications, but in sustained availability, reliability at scale, and whether it lowers the per-request costs of serving large models in mainstream products.