Microsoft unveils Maia 200, a 3nm AI inference accelerator aimed at cheaper, faster token generation
Microsoft has introduced Maia 200, its next-generation in-house AI inference accelerator built on TSMC’s 3nm process. The company says the chip improves the economics of AI token generation and is being deployed in Azure data centres, alongside a software stack to help developers optimise models.
- Reporting desk
- Health India Network News Desk
- First published
Microsoft has unveiled Maia 200, its latest in-house AI accelerator designed primarily for inference workloads. In a detailed technical blog post, the company said Maia 200 is fabricated on TSMC’s 3nm process and is engineered to improve performance per rupee spent on AI token generation — a key cost driver as AI assistants and copilots scale to millions of users.

Microsoft highlighted architecture choices tuned for modern, low-precision inference, including FP8 and FP4 tensor cores, as well as a redesigned memory subsystem with high-bandwidth HBM3e and significant on-chip SRAM. The goal is to keep large models fed with data efficiently, boosting throughput without proportionally increasing power or total cost of ownership.
The company said Maia 200 is already deployed in its US Central datacentre region and will expand to additional regions over time. This suggests Microsoft is moving beyond experimentation into operational deployment, where chips must run reliably at scale, integrate with scheduling systems and support real-world model serving patterns.
Alongside hardware, Microsoft is previewing a Maia software development kit that includes tools intended to make model porting and optimisation easier. For AI infrastructure, the software layer often determines whether a chip delivers its promised performance in production, because compilers, kernels and observability can become bottlenecks even when raw FLOPS are high.
The announcement fits a broader hyperscaler trend: building custom silicon to reduce dependence on a single GPU supply chain, control costs and tailor systems to specific workloads. For Microsoft, inference efficiency is particularly strategic because it can directly reduce the per-query cost of AI features across Azure and productivity products.
For enterprises and developers in India consuming Azure AI services, the key implication is the possibility of lower serving costs and improved latency as custom accelerators are rolled out more widely. Actual benefits will depend on regional deployment timelines, model support, and how well Maia’s software stack performs across diverse workloads and frameworks.