The race to build more powerful artificial intelligence (AI) infrastructure is increasingly a systems problem. While accelerator performance remains important, expanding AI models and more complex workloads mean processors cannot be considered in isolation. Memory capacity, interconnect bandwidth, storage and power consumption can all become limiting factors as clusters grow.
This is becoming particularly significant as the AI infrastructure build-out accelerates. At the start of 2026, GlobalData estimated a global pipeline of large-scale data centre projects worth $2.306tn, while the growing adoption of generative AI is driving higher-density computing into more enterprise environments. It also forecasts that global data centre power consumption will triple between 2024 and 2030, at a compound annual growth rate of 21.1%.
In response, Huawei is adopting a system-level approach to the AI computing challenge. At HUAWEI CONNECT 2026 in Shanghai, the company unveiled its Peerium Computing Architecture and a new generation of infrastructure, including the Atlas 960E SuperPoD.
Why AI infrastructure is becoming a system-level challenge
Scaling an AI cluster by adding processors does not automatically increase usable computing power. As systems grow, processors must spend more time exchanging data. This raises memory requirements, and communication overheads can leave expensive compute capacity underused.
The challenge becomes more pronounced with frontier models. Huawei is designing its latest infrastructure with models of up to 10 trillion parameters in mind, while agentic AI adds another dimension. Agents can repeatedly interact with models, tools and data sources, generating more traffic between central processing units (CPUs), neural processing units (NPUs), memory and storage.
According to Huawei, in a traditional 100,000-NPU cluster, model workloads can use as little as 20% of available computing capacity because significant processing time is lost to data communication. The response is to focus on the architecture connecting those resources rather than treating each processor as an isolated source of performance.
Announced on 17 September, the Peerium Computing Architecture is designed to allow processors at the million-scale to operate collectively as a single computing system. It achieves this through nested parallelism, unified memory addressing and peer-to-peer interconnection. The architecture introduces Nested Bulk Synchronous Parallel, or Nested BSP, and moves away from the conventional master-slave arrangement in which certain processors or nodes coordinate others. The aim is to make computing resources more equal participants in the system and allow workloads to scale across a much larger pool of hardware.
Peerium and UnifiedBus connect compute, memory and storage
Central to this approach is UnifiedBus (UB), Huawei’s high-speed interconnect technology and a key building block of Peerium. Rather than using separate interconnect mechanisms for different hardware resources, UB uses a single protocol to connect CPUs, NPUs, memory, solid-state drives (SSDs), network interface cards and switches. It also provides unified memory addressing, enabling resources distributed across physical systems to be treated as part of a shared memory environment.
UnifiedBus can directly interconnect different computing and storage resources while reducing the protocol conversions normally required when data moves through a large cluster. That creates a common “data highway” across the infrastructure, designed to reduce communication delays and allow resources to be pooled more flexibly.
The Atlas 950 SuperPoD was the first generation of products based on the Peerium architecture. Its latest model, the Atlas 960E, builds on the same system-level philosophy by combining UnifiedBus with a new generation of optical interconnect technology. Huawei describes the Atlas 960E as the industry’s first SuperPoD based on near-packaged optics (NPO), operating as a single system that can scale to 4,096 NPUs via unified memory addressing and is designed to deliver 8 EFLOPS of compute performance at FP8 precision or 16 EFLOPS at FP4. The system incorporates 5,500 of Huawei’s Hi-ONE optical engines, each providing 7.2Tbit/s of transmission capacity and places optical components close to the processors they connect.
Using Hi-ONE means the Atlas 960E can avoid the approximately 48,000 conventional 800G optical modules otherwise required to connect its NPUs. Huawei says this reduces system power consumption by more than 550kW and doubles fault-free operating time, contributing to stated system availability of 99.8%. The Atlas 960E also uses a fully liquid-cooled design, another increasingly important consideration as rack densities rise.
With GlobalData expecting data centre electricity consumption to grow rapidly through the end of the decade, improvements in the power required to move data around AI systems could become almost as consequential as efficiency gains in the processors themselves.
Scaling from a SuperPoD towards million-processor AI clusters
Huawei’s longer-term ambition extends beyond a single SuperPoD. Multiple Atlas 960E systems can be interconnected via UnifiedBus or remote direct memory access over converged Ethernet (RoCE) to form larger SuperClusters. Huawei says a two-tier, four-plane Clos network architecture can connect up to 512,000 NPUs. Combined with a multi-rail topology, its architecture is designed to support up to one million NPUs.
The company is also extending the same principles beyond dedicated AI acceleration. According to Huawei, its upgraded TaiShan 950 SuperPoD applies UnifiedBus to general-purpose computing and supports up to 4,096 nodes with a unified memory pool of up to 256TB. Meanwhile, the OceanStor M900 context memory storage cluster is intended to provide petabyte-scale key-value cache capacity for AI inference.
That could become increasingly relevant as AI moves from model training towards high-volume inference and agentic workloads. These applications require rapid access to context and repeated exchanges between different types of computing resources. Pooling those resources through a common architecture can improve utilisation and make infrastructure more adaptable to changing workloads.
As AI infrastructure expands, the focus is shifting from the performance of individual processors to the efficiency of entire systems working together. Huawei’s Peerium architecture and Atlas 960E SuperPoD exemplify this shift and suggest that future progress in AI infrastructure may rely more on system-level design.