Huawei pulls Ascend 960DT forward by three quarters and launches the industry's first NPO-based SuperPoD
On September 17, 2026, at HUAWEI CONNECT 2026 in Shanghai, Huawei put two timetables on the table. The first concerns chips: the Ascend 960DT will be available in Q1 2027, three quarters ahead of the original roadmap; the Ascend 960PR moves a quarter earlier to Q3 2027; and the Ascend 970 and 980 follow in 2028 and 2029. Huawei rotating chairman David Wang framed it in his keynote as a one-generation-a-year cadence for the Ascend line. The second timetable concerns systems: the Atlas 960E SuperPoD, which Huawei calls the industry's first NPO-based SuperPoD. A chip pulled forward and a system rebuilt on a new optical foundation in the same event points not at single-chip performance but at the engineering discipline of turning tens of thousands of chips into one machine.
Start with the optical line, because it is the most concrete piece of the launch. Huawei names its next-generation NPO optical interconnect product Hi-ONE, the High-density Optical-interconnect-Node Engine. It is built on proprietary technology with a multi-physics design balancing optical, mechanical, electrical, electromagnetic and thermal performance, and reaches a transmission capacity of 7.2 Tbit/s per single engine. Huawei's claim is that this is the industry's first NPO product ready for mass production and the first NPO product with a built-in light source. Combined with UnifiedBus, its purpose is to make scaling SuperPoD interconnect systems far easier. Huawei recently submitted an implementation agreement on NPO to the Optical Internetworking Forum and says the response from industry partners has been widely positive.
That optical work directly shapes the Atlas 960E. A single Atlas 960E SuperPoD can scale up to 4,096 NPUs, delivering 8 EFLOPS of FP8 compute with up to 1 petabyte of HBM capacity. The detail that matters is that it uses 5,500 Hi-ONE units where connecting those NPUs the traditional way would require 48,000 800G optical modules. Huawei's conclusion is that this cuts power consumption by more than 550 kilowatts while doubling fault-free operating time, achieving 99.8% system availability. In other words, the value of NPO is not a faster single engine but an order-of-magnitude reduction in the number of interconnect parts, and part count directly determines the power, failure points and cabling complexity of a large cluster. At the tens-of-thousands-of-cards scale, all three are costs you can only discuss once they are solved.
Beyond the chip roadmap, Huawei filled in a relative weak spot: memory and context storage built for agents. The TaiShan 950 SuperPoD has been fully upgraded on UnifiedBus all-optical networking, supporting up to 4,096 nodes with a unified memory pool of up to 256 TB. Huawei offers three comparisons: for sandbox-intensive workloads, startup speed for 100,000 sandboxes is 30 times faster than traditional servers with sandbox density improved by a further 25%; and for vector search across 10 billion 1,000-dimensional vectors, search efficiency is twice that of traditional servers. Alongside it, OceanStor M900 is a context memory storage cluster providing multi-tier KV caching for agent-heavy and longer-context workloads, with one-hop direct access, a petabyte-scale KV cache at the L3.5 layer, and hybrid media plus an optimised retention algorithm that Huawei says extends SSD read/write lifespan 16-fold. What these figures share is that they all answer one question: as contexts grow longer and agents multiply, what holds the intermediate state.
What ties the systems together is the SuperCluster. It uses UnifiedBus to consolidate multiple interconnect protocols into a single unified protocol, significantly reducing protocol conversion overhead, which enables peer-to-peer interconnect between Ascend SuperPoDs, Kunpeng SuperPoDs and KV cache clusters, aimed at training and inference for 10-trillion-parameter models. Huawei also cites simulation results from its Markov Lab: a 100k-NPU cluster built from 4k-NPU SuperPoDs can deliver a 2.75x increase in MFU compared with a 100k-NPU cluster composed of 8-NPU servers. One counterpoint has to travel with all of this. China tech analyst Rui Ma noted on social media that Huawei had previously said its Atlas 960 SuperPoD would scale to 15,488 Ascend 960 chips, while this week's announcement referred to a system with 4,096 chips; her assessment was that the chip itself is coming way earlier, but the SuperPoD announced is much smaller than originally laid out. That gap is central to reading the launch: a pulled-forward roadmap and a shrunken system happened at the same time, and citing only half of it distorts the picture.
🤔 Frequently Asked Questions
When will the Ascend 960DT be available?
Q1 2027, three quarters ahead of Huawei's original roadmap, which had it in Q3 2027. The Ascend 960PR moves one quarter earlier to Q3 2027. Huawei says the Ascend line is now on a one-generation-a-year cadence, with the Ascend 970 and 980 following in 2028 and 2029.
What is the scale of the Atlas 960E SuperPoD?
A single one scales to 4,096 NPUs, delivers 8 EFLOPS of FP8 compute with up to 1 petabyte of HBM, and uses 5,500 Hi-ONE units in place of the 48,000 800G optical modules a conventional design would need, cutting power by more than 550 kilowatts with 99.8% system availability.
What is Hi-ONE?
Huawei's next-generation optical interconnect product based on near-packaged optics, short for High-density Optical-interconnect-Node Engine, with a transmission capacity of 7.2 Tbit/s per engine. Huawei calls it the industry's first NPO product ready for mass production and the first with a built-in light source, and says it has submitted an NPO implementation agreement to the Optical Internetworking Forum.
What should be flagged about this launch?
China tech analyst Rui Ma noted that Huawei had previously said its Atlas 960 SuperPoD would scale to 15,488 Ascend 960 chips, while this announcement described a system with 4,096, meaning the chip is dramatically earlier but the SuperPoD is far smaller than planned. Both halves belong in any citation of this launch.
🛠️ Recommended Tools
- Unit ConverterThis launch mixes EFLOPS for compute, Tbit/s for bandwidth, PB and TB for capacity, and kilowatts for power. Manual conversion across those units is exactly where the zeros go wrong.
- Text Difference CheckerHuawei had previously described a SuperPoD of 15,488 chips; this time it is 4,096, and the roadmap moved from Q3 2027 to Q1 2027. Diffing two versions of a spec sheet beats citing from memory.
- JSON FormatterKeeping the SuperPoD and cluster variants as structured data makes the comparison tractable: node count, compute, memory and bandwidth only line up when they are read as a set, which stops specs from being attributed to the wrong part.
Summary
On September 17, 2026, at HUAWEI CONNECT 2026 in Shanghai, Huawei announced several AI computing technologies and products. On chips: the Ascend 960DT will be available in Q1 2027, three quarters ahead of the original roadmap; the Ascend 960PR moves a quarter earlier to Q3 2027; and Huawei says the Ascend line is now on a one-generation-a-year cadence, with the Ascend 970 and 980 following in 2028 and 2029. On systems and optics: it launched the industry's first NPO-based SuperPoD, the Atlas 960E, scaling to 4,096 NPUs with 8 EFLOPS of FP8 and up to 1 petabyte of HBM, replacing the 48,000 800G optical modules a conventional design needs with 5,500 Hi-ONE engines, cutting power by more than 550 kilowatts at 99.8% availability; Hi-ONE delivers 7.2 Tbit/s per engine, and Huawei calls it the industry's first mass-production-ready and first built-in-light-source NPO product, with an NPO implementation agreement submitted to the OIF. On memory and storage: the TaiShan 950 SuperPoD is upgraded to up to 4,096 nodes with a unified memory pool of up to 256 TB, startup for 100,000 sandboxes is 30 times faster than traditional servers with sandbox density up 25%, and vector search across 10 billion 1,000-dimensional vectors is twice as efficient; OceanStor M900 provides a petabyte-scale KV cache at the L3.5 layer and is said to extend SSD read/write lifespan 16-fold. On clusters: SuperCluster consolidates multiple interconnect protocols onto UnifiedBus to support 10-trillion-parameter models, and Huawei cites Markov Lab simulation showing a 2.75x MFU gain for a 100k-NPU cluster built from 4k-NPU SuperPoDs versus 8-NPU servers. One counterpoint must travel with it: analyst Rui Ma noted the SuperPoD shrank from the 15,488 chips previously laid out to the 4,096 announced this week. All figures come from Huawei's official keynote page and AP and TechCrunch reporting, with nothing inferred.
Sources: Huawei: Advancing the Agentic World, Building a Solid Silicon Foundation (HUAWEI CONNECT 2026 keynote)
TechCrunch: Huawei plans Q1 2027 launch of new AI chip as it takes on Nvidia
AP: Huawei unveils new chip technologies as Chinese firm steps up the AI race with Nvidia