Huawei used its Connect conference to move the Ascend 960DT up by three quarters to the first quarter of 2027, pairing it with a 4,096-accelerator Atlas SuperPoD built on near-package optics. Per-chip, the part still trails Nvidia's next generation by a wide margin. That may matter less than it sounds, because the chips it actually competes against inside China are the ones US export rules still permit.
Key takeaways
- The Ascend 960DT arrives in Q1 2027 with up to 288GB of memory at 9.6 TB/s and 2 PFLOPS FP8 / 4 PFLOPS FP4 โ double the compute and capacity of the 950 series shipped earlier this year.
- A companion 960PR follows in Q3 2027 with 8 PFLOPS FP4 but only 192GB at 2.4 TB/s, splitting inference so prefill runs on one chip and decode on the other.
- Huawei says near-package optics let it collapse roughly 48,000 800Gbps pluggable transceivers into about 5,500 Hi-ONE modules, cutting 550 kilowatts and halving failure rates.
What the 960 series actually delivers
The 960DT โ DT for decode and training โ carries up to 288GB of memory running at 9.6 TB/s, roughly a 2.4x bandwidth jump over the 950 generation, plus 2.2 TB/s of chip-to-chip interconnect. Compute lands at 2 petaFLOPS in FP8 and 4 petaFLOPS in FP4. Huawei is also pushing its own HiF8 and HiF4 numeric formats alongside the standard FP and MX types, on the argument that a single format tuned for both precision and dynamic range can replace two conventional ones.
The 960PR, due two quarters later, inverts the trade. It doubles FP4 throughput to 8 petaFLOPS while dropping to 192GB at 2.4 TB/s, which is exactly the shape you want for the prefill stage of inference, where a prompt is chewed through in one compute-heavy pass. The memory-bound decode stage, where tokens actually stream out, stays on the 960DT. Nvidia had aimed its cancelled Rubin CPX at the same split, and Google's TPU pods have long leaned on the same logic.
How far behind Nvidia is it
The Register's read of the specs puts the 960DT in the neighbourhood of Nvidia's B300 family on memory and bandwidth while delivering about half the FP8 and a third the FP4 compute. Against Rubin, which promises 35 to 50 petaFLOPS of FP4 with 288GB of HBM4 and 22 TB/s, the gap is not close.
The catch is the one that keeps recurring in this market: Rubin and AMD's MI455X cannot be sold in China. Measured against the best accelerator Nvidia is currently permitted to ship there, the 960DT holds its own on dense floating-point work while offering considerably more memory, more bandwidth and support for much larger scale-up domains. For a Chinese lab choosing hardware for 2027, that is the comparison that binds.
Why the pod is the real product
Huawei's strongest card is networking, which is the business it has always been in. The Atlas 960E SuperPoD, slated for Q3 2027 and liquid-cooled, ties 4,096 accelerators into one unified memory domain and claims 8 exaFLOPS FP8 and 16 exaFLOPS FP4. It is built on near-package optics rather than pluggable transceivers โ Huawei says the switch replaced about 48,000 800Gbps pluggables with roughly 5,500 Hi-ONE engines, saved more than 550 kilowatts and cut failure rates in half, landing at a claimed 99.8% availability.
Against the Atlas 950, Huawei claims 2.3x training throughput and 2.5x inference throughput on a 10-trillion-parameter workload, with about 70% lower inference latency. Beyond a single pod, it says RoCE or its Lingqu fabric can chain up to 512,000 accelerators, and that a multi-rail topology puts million-NPU clusters within reach. That last figure is a paper configuration, not a deployment, and worth reading as roadmap rather than result.
What to watch before believing the roadmap
The practical question is not the slide deck but whether the silicon trains anything. DeepSeek reportedly hit trouble training on earlier Ascend parts and reverted to Nvidia hardware โ a software and reliability failure, not an arithmetic one, and the kind of problem a faster interconnect does not fix on its own. Roadmap slides from the event already sketch an Ascend 970 for 2028 at 14 petaFLOPS FP4 and 14.4 TB/s, and a 980 for 2029 at 28 petaFLOPS FP4 with 384GB, but three-year roadmaps in this industry are drafts.
Still, the direction is consistent with what is happening one layer up, where Chinese open-weight models have been taking share on developer routing platforms. A domestic training substrate that is merely adequate, available in volume and shipped nine months early is a different strategic object from one that wins benchmarks.
FAQ
When do the Ascend 960 chips ship?
Huawei says the 960DT arrives in the first quarter of 2027, three quarters ahead of its original schedule, with the 960PR following in the third quarter of 2027. The Atlas 960E SuperPoD that houses them is also targeted at Q3 2027. None of these are shipping today.
Is the Ascend 960DT competitive with Nvidia's Rubin?
Not on raw compute. Rubin is quoted at 35 to 50 petaFLOPS of FP4 against the 960DT's 4 petaFLOPS, though the two chips carry similar memory capacity. Huawei's counter-argument is scale โ 4,096 accelerators in one coherent domain โ rather than per-die performance.
What is near-package optics and why does it matter here?
Near-package optics moves the optical engine next to the chip package instead of using separate pluggable transceiver modules. Huawei says that cut roughly 48,000 pluggables down to about 5,500 Hi-ONE units, saving over 550 kilowatts and halving failure rates โ which is what makes a 4,096-chip scale-up domain practical to operate.






