AI Newsway

Huawei Pulled the Ascend 960 Forward Nine Months. Scale Is the Argument, Not the Die

The 960DT lands in Q1 2027 with 288GB and 4 PFLOPS of FP4, still far behind Rubin per chip โ€” but Huawei is selling 4,096-accelerator pods, and Nvidia cannot ship Rubin into China anyway

|5 min read0
AI Summary
Huawei advanced its Ascend 960DT accelerator by three quarters to Q1 2027, with 288GB of memory at 9.6 TB/s and 4 petaFLOPS of FP4 compute, plus a 960PR inference variant in Q3 2027. Per-chip it trails Nvidia's Rubin badly, but Rubin cannot be sold in China. Huawei's real pitch is the 4,096-accelerator Atlas 960E SuperPoD, built on near-package optics that replace roughly 48,000 pluggable transceivers.
Huawei's Ascend 960 roadmap leans on rack-scale systems, tying 4,096 accelerators into a single liquid-cooled Atlas SuperPoD rather than chasing Nvidia on per-chip compute
Huawei's Ascend 960 roadmap leans on rack-scale systems, tying 4,096 accelerators into a single liquid-cooled Atlas SuperPoD rather than chasing Nvidia on per-chip compute

Huawei used its Connect conference to move the Ascend 960DT up by three quarters to the first quarter of 2027, pairing it with a 4,096-accelerator Atlas SuperPoD built on near-package optics. Per-chip, the part still trails Nvidia's next generation by a wide margin. That may matter less than it sounds, because the chips it actually competes against inside China are the ones US export rules still permit.

Key takeaways

  • The Ascend 960DT arrives in Q1 2027 with up to 288GB of memory at 9.6 TB/s and 2 PFLOPS FP8 / 4 PFLOPS FP4 โ€” double the compute and capacity of the 950 series shipped earlier this year.
  • A companion 960PR follows in Q3 2027 with 8 PFLOPS FP4 but only 192GB at 2.4 TB/s, splitting inference so prefill runs on one chip and decode on the other.
  • Huawei says near-package optics let it collapse roughly 48,000 800Gbps pluggable transceivers into about 5,500 Hi-ONE modules, cutting 550 kilowatts and halving failure rates.

What the 960 series actually delivers

The 960DT โ€” DT for decode and training โ€” carries up to 288GB of memory running at 9.6 TB/s, roughly a 2.4x bandwidth jump over the 950 generation, plus 2.2 TB/s of chip-to-chip interconnect. Compute lands at 2 petaFLOPS in FP8 and 4 petaFLOPS in FP4. Huawei is also pushing its own HiF8 and HiF4 numeric formats alongside the standard FP and MX types, on the argument that a single format tuned for both precision and dynamic range can replace two conventional ones.

The 960PR, due two quarters later, inverts the trade. It doubles FP4 throughput to 8 petaFLOPS while dropping to 192GB at 2.4 TB/s, which is exactly the shape you want for the prefill stage of inference, where a prompt is chewed through in one compute-heavy pass. The memory-bound decode stage, where tokens actually stream out, stays on the 960DT. Nvidia had aimed its cancelled Rubin CPX at the same split, and Google's TPU pods have long leaned on the same logic.

How far behind Nvidia is it

The Register's read of the specs puts the 960DT in the neighbourhood of Nvidia's B300 family on memory and bandwidth while delivering about half the FP8 and a third the FP4 compute. Against Rubin, which promises 35 to 50 petaFLOPS of FP4 with 288GB of HBM4 and 22 TB/s, the gap is not close.

The catch is the one that keeps recurring in this market: Rubin and AMD's MI455X cannot be sold in China. Measured against the best accelerator Nvidia is currently permitted to ship there, the 960DT holds its own on dense floating-point work while offering considerably more memory, more bandwidth and support for much larger scale-up domains. For a Chinese lab choosing hardware for 2027, that is the comparison that binds.

Why the pod is the real product

Huawei's strongest card is networking, which is the business it has always been in. The Atlas 960E SuperPoD, slated for Q3 2027 and liquid-cooled, ties 4,096 accelerators into one unified memory domain and claims 8 exaFLOPS FP8 and 16 exaFLOPS FP4. It is built on near-package optics rather than pluggable transceivers โ€” Huawei says the switch replaced about 48,000 800Gbps pluggables with roughly 5,500 Hi-ONE engines, saved more than 550 kilowatts and cut failure rates in half, landing at a claimed 99.8% availability.

Against the Atlas 950, Huawei claims 2.3x training throughput and 2.5x inference throughput on a 10-trillion-parameter workload, with about 70% lower inference latency. Beyond a single pod, it says RoCE or its Lingqu fabric can chain up to 512,000 accelerators, and that a multi-rail topology puts million-NPU clusters within reach. That last figure is a paper configuration, not a deployment, and worth reading as roadmap rather than result.

What to watch before believing the roadmap

The practical question is not the slide deck but whether the silicon trains anything. DeepSeek reportedly hit trouble training on earlier Ascend parts and reverted to Nvidia hardware โ€” a software and reliability failure, not an arithmetic one, and the kind of problem a faster interconnect does not fix on its own. Roadmap slides from the event already sketch an Ascend 970 for 2028 at 14 petaFLOPS FP4 and 14.4 TB/s, and a 980 for 2029 at 28 petaFLOPS FP4 with 384GB, but three-year roadmaps in this industry are drafts.

Still, the direction is consistent with what is happening one layer up, where Chinese open-weight models have been taking share on developer routing platforms. A domestic training substrate that is merely adequate, available in volume and shipped nine months early is a different strategic object from one that wins benchmarks.

FAQ

When do the Ascend 960 chips ship?

Huawei says the 960DT arrives in the first quarter of 2027, three quarters ahead of its original schedule, with the 960PR following in the third quarter of 2027. The Atlas 960E SuperPoD that houses them is also targeted at Q3 2027. None of these are shipping today.

Is the Ascend 960DT competitive with Nvidia's Rubin?

Not on raw compute. Rubin is quoted at 35 to 50 petaFLOPS of FP4 against the 960DT's 4 petaFLOPS, though the two chips carry similar memory capacity. Huawei's counter-argument is scale โ€” 4,096 accelerators in one coherent domain โ€” rather than per-die performance.

What is near-package optics and why does it matter here?

Near-package optics moves the optical engine next to the chip package instead of using separate pluggable transceiver modules. Huawei says that cut roughly 48,000 pluggables down to about 5,500 Hi-ONE units, saving over 550 kilowatts and halving failure rates โ€” which is what makes a 4,096-chip scale-up domain practical to operate.

How do you feel about this article?

SJ

Discussion

Sign in to post
Loading...

Related articles

Amazon Triples Its Nvidia Order to 2 Million GPUs
Tech & Business

Amazon Triples Its Nvidia Order to 2 Million GPUs

Amazon is adding 2 million more Nvidia GPUs to AWS just five months after committing to 1 million, even as it scales its own Trainium and Graviton silicon.

Seung Jung22 days ago
Huang Defends Nvidia's 70% Growth Call: 'We Put In $1 and $100 Comes Back'
Tech & Business

Huang Defends Nvidia's 70% Growth Call: 'We Put In $1 and $100 Comes Back'

Jensen Huang reaffirmed Nvidia's 70% growth guidance, implying $680 billion in revenue, and answered circular-deal critics with '$1 in, $100 back.'

Seung Jung7 days ago
OpenAI Says Its Own Models Helped Tape Out Jalapeรฑo in Nine Months
Tech & Business

OpenAI Says Its Own Models Helped Tape Out Jalapeรฑo in Nine Months

AI-generated kernels beat OpenAI expert-written versions by up to 1.8x, as Jalapeรฑo posts its first InferenceX benchmark results.

Seung Jung23 days ago
Nvidia Nears $12.9 Billion Deal to Buy Hugging Face
Tech & Business

Nvidia Nears $12.9 Billion Deal to Buy Hugging Face

Nvidia has reportedly agreed to acquire Hugging Face for $12.9 billion, a multiple of roughly 80x revenue that buys the developer graph, not the income statement.

Seung Jung22 days ago
Nvidia Spent $27 Billion Without Filing a Single Merger Notice. The DOJ Wants to Know Why.
Tech & Business

Nvidia Spent $27 Billion Without Filing a Single Merger Notice. The DOJ Wants to Know Why.

Antitrust enforcers have opened their first real examination of the deal structure that has replaced the acquisition in AI: Nvidia has received a formal Justice...

Seung Jung5 days ago
Crusoe Raises $3.9B at a $30.9B Valuation and Bets on Data Centers You Can Truck In
Tech & Business

Crusoe Raises $3.9B at a $30.9B Valuation and Bets on Data Centers You Can Truck In

Data center developer Crusoe announced on Thursday the initial closing of a $3.9 billion Series F at a $30.9 billion post-money valuation. Atreides Management,...

Seung Jung5 hours ago