News

Memory efficiency becomes AI’s constraint

Microsoft frames AI infrastructure around useful yield.

Jon Peddie

Microsoft is asking the AI infrastructure industry to focus on what its installed hardware actually produces. Rani Borkar, president of Azure Hardware Systems and Infrastructure, says memory shortages reflect a wider systems problem involving compute, software, networking, packaging, power, and cooling. For CIOs, ISVs, and silicon teams, the message is practical: capacity matters, though utilization determines the economic value of that capacity across increasingly expensive AI deployments.

The AI build-out continues to pull HBM and conventional memory into data center deployments at a rapid rate. Memory suppliers including Micron, Samsung, SK Hynix, CXMT, Winbond, and Nanya have expanded capacity, while broader semiconductor demand has pushed the World Semiconductor Trade Statistics forecast for 2026 global chip sales to $1.51 trillion.

Rani Borkar argues that capacity expansion addresses only part of the constraint. “Memory is viewed as a supply problem or a component problem,” she said at Semicon Taiwan. “In reality, it is a system problem.” Her argument focuses on useful yield: how much productive AI work an infrastructure investment produces after accounting for idle memory, stalled accelerators, data movement, scheduling delays, power limits, and software overhead.

Figure 1. Borkar describes two directions for future AI infrastructure.

Memory does not create value simply because it sits near a GPU or an accelerator. Models need data to arrive at the right time, in the right location, and at the right bandwidth. Software must schedule work effectively. Interconnects must sustain data movement. Power and cooling must support sustained operation. If one layer falls short, expensive compute and memory resources wait.

For silicon teams, that makes memory hierarchy, cache strategy, HBM placement, packaging, fabric bandwidth, and coherency part of a common architecture discussion. For ISVs, it elevates memory allocation, batching, runtime behavior, orchestration, compiler choices, and workload placement. An application with inefficient memory management can leave hardware underutilized despite strong component-level specifications.

AI infrastructure now turns nearly every system element into a potential constraint. More accelerators can expose a shortage of memory capacity or bandwidth. Additional memory can move the bottleneck into networking, power delivery, cooling, or the software stack. Faster links can reveal poor data locality or orchestration across distributed workloads.

Table 1. AI stresses every layer.

Peak TOPS, FLOPS, and memory bandwidth remain useful engineering measures. They describe the potential of individual components. CIOs need measures that capture delivered infrastructure value: model throughput, latency, queue time, accelerator utilization, memory utilization, energy per inference, and cost per useful output.

Microsoft’s framing recognizes that AI economics now depend on an end-to-end system. A higher-specification component does not automatically deliver proportionally higher model throughput. The organization that uses its hardware more effectively can create more capacity from the same installed base.

Microsoft’s silicon strategy

Microsoft has built Maia AI accelerators for Azure AI workloads and Cobalt CPUs for general-purpose and inference tasks. Borkar said Microsoft intends to develop successive generations of both product families. The company does not plan to sell those processors to external cloud or enterprise customers, enabling it to align the silicon with Azure infrastructure, networking, software, operating practices, and known workload patterns.

Figure 2. Two paths emerge.

Microsoft continues to deploy third-party processors, including CPUs from AMD and Intel, within Azure. This multi-supplier approach gives Azure access to merchant roadmaps, while its internal silicon teams optimize targeted workloads and system configurations.

Nvidia has expanded from GPUs into CPUs, networking, interconnects, rack-scale systems, and software. Microsoft begins with cloud operations, then shapes silicon around the requirements of its cloud platform. Both companies treat compute, memory, networking, and software as interdependent elements of AI infrastructure.

Borkar described full-stack integration as the important unit of value. Separating a component from the broader system can remove the context that gives it economic and technical value.

Capital makes utilization strategic

Microsoft expects 2026 capital spending of about $190 billion, over 60% higher than the prior year, according to the provided report. At that scale, utilization becomes both an architectural metric and a financial metric. A modest improvement in memory or accelerator utilization can unlock a meaningful amount of productive capacity without a matching increase in hardware purchases.

Borkar identifies two paths. The first improves current infrastructure through utilization, efficiency, and economics. The second changes the trajectory through new architectures, materials, and approaches to model construction. Microsoft’s “useful yield” framework connects both paths: It calls for immediate gains from the current stack while creating room for more fundamental changes to how AI systems use compute and memory.

Microsoft’s position moves the conversation beyond component acquisition. AI deployments still need more accelerators, memory, storage, networking, and data center capacity. System-level engineering determines how much intelligence those investments produce. Maia, Cobalt, Azure, and Microsoft’s third-party silicon ecosystem give the company a broad operating environment for testing that premise.

What do we think?

Borkar identifies utilization as the next economic challenge in AI infrastructure. The industry has directed enormous resources toward GPUs, HBM, and data center construction. Future gains will depend on extracting more useful work from those assets. Microsoft controls Azure operations, custom silicon, networking, and software, giving it a practical environment for full-stack optimization at hyperscale.

Inflection point. Microsoft’s argument may forecast an inflection point in AI infrastructure planning. The first phase of generative AI emphasized building more capacity as quickly as suppliers could deliver it. The next phase may prioritize intelligence produced per dollar, watt, byte, and square foot. If the industry adopts that measure, processor selection, memory architecture, networking, model design, and software optimization will carry equal weight in AI deployment decisions.

LIKE WHAT YOU’RE READING? INTRODUCE US TO YOUR FRIENDS AND COLLEAGUES.