On September 17, 2026, at its Connect conference in Shanghai, Huawei moved the launch of its Ascend 960DT AI chip from the fourth quarter of 2027 to the first, three full quarters early. Two weeks before that, Bloomberg reported, citing unnamed sources, that DeepSeek plans to install at least 160,000 Ascend 950DT chips at a roughly one-gigawatt data centre in Ulanqab, Inner Mongolia, and to use them only for inference. Read together, the two stories say more than either does alone: one company is racing its roadmap forward, and the other is voting with real money, but casting only half a ballot. This piece does three things. It translates the jargon, works out what a "system-scale" strategy can and cannot fix, and then does the supply-chain arithmetic.

Key TakeawaysHuawei's strategy is to concede that its chip trails NVIDIA's and then to use interconnect, cabinets and scheduling to fuse tens of thousands of chips into one machine. That works on paper for inference and is unproven for training. What decides 2027 is not the PFLOPS on the keynote slide but three duller things: how much high-bandwidth memory China can make, how many developers will use Huawei's software, and when the first customer cases appear. The reported DeepSeek order is a live test of all three.

The Jargon, Translated: Chip, Bus and Cabinet Are Three Layers

The three names in the headlines sit at different levels, and reading them as one thing is how people get lost. The Ascend 960 is a chip, what Huawei calls an NPU and what everyone else would call a GPU-class accelerator. On the roadmap Huawei published at Connect 2025, it delivers 2 PFLOPS at FP8 with 288 GB of memory, 9.6 TB/s of memory bandwidth and 2.2 TB/s of interconnect. At this year's Connect Huawei split it in two: a PR version for prefill, the stage where the model reads your prompt, and a DT version for decode, the stage where it produces the answer one token at a time, with training as a second job. The DT is the one that moved to Q1 2027; the PR is set for Q3.

UnifiedBus, or Lingqu in Chinese, is an interconnect protocol: the language chips and cabinets use to talk to each other. It plays the role that NVLink plays inside an NVIDIA rack and that the network plays between racks. Huawei also ships a version that runs over ordinary Ethernet, so a customer need not replace its switches.

A SuperPoD is what you get when you use that language to join thousands of chips into one logical machine. The Atlas 950 SuperPoD is 8,192 Ascend 950DT chips with 8 EFLOPS of FP8 compute, due in the fourth quarter of this year; the Atlas 960 SuperPoD is 15,488 chips. Above that sits the SuperCluster, 64 SuperPoDs and more than 520,000 chips. The "one supercomputer" framing is Huawei's own. It means that, to the software, those chips look like one continuous pool of compute and memory rather than thousands of separate servers.

AI computing scales from an accelerator chip and HBM memory to an interconnected server board and a cabinet cluster.

So what actually separates "system scale" from "single-chip advantage"? Take a situation we have run into ourselves. You are serving a mixture-of-experts model and your user count doubles. On the single-chip path you swap in stronger cards, each one carrying more concurrent requests, and the cabinet count stays put. On the system path you add cabinets, but the new cards have to talk to the old ones fast enough, because a mixture-of-experts model routes every generated token through "experts" that live on different chips and hauls the conversation cache around with it. If the bandwidth is not there, the new compute gets eaten by the hauling.

The gap between the two paths, then, is not peak FLOPS. It is four things: bandwidth, how much data moves between chips each second; latency, how long each exchange waits; the software stack, whether the scheduler can cut the work into the right pieces; and availability, whether you can buy the hardware at all. In the Chinese market, the last one is often the only one that counts.

Why Huawei Chose This Road, and What Scale Can and Cannot Fix

Start with the constraint that will not move. Ascend's compute dies are made by SMIC, which has no EUV lithography, on a process a generation or two behind TSMC. That does not change before 2027. Huawei's answer is to say so plainly and then move the contest to a different arena: if it cannot win chip against chip, it will compete on who can join more chips into one machine.

There is one shipped, independently examined precedent for that road. In April 2025 SemiAnalysis compared Huawei's CloudMatrix 384 with NVIDIA's GB200 NVL72. Huawei's system uses 384 Ascend 910C chips against NVIDIA's 72 Blackwell GPUs and delivers about 1.7 times the total compute and about 3.6 times the total HBM. The price is power: roughly 560 kW against 145 kW, about four times as much. More chips, more optical transceivers and more electricity buy a system-level lead. Chinese electricity is cheap and plentiful, so the sum works in Inner Mongolia in a way it never would in Northern Virginia.

Which gaps can scale genuinely close? Three, in my view. Total throughput, especially in the decode stage of inference, where the bottleneck was always memory and interconnect rather than raw compute. Total memory, because a mixture-of-experts model with hundreds of billions of parameters can be spread across thousands of chips. And divisibility: inference requests are independent of each other, so when one chip fails the scheduler moves the request and the user barely notices.

Which gaps cannot be closed that way? First, the electricity and floor space per token. A weaker chip's efficiency gap does not shrink with scale; it multiplies. The bigger the cluster, the larger the extra power bill in absolute terms. Second, training. Training asks tens of thousands of chips to synchronise their gradients every few seconds, and if one chip or one fibre link stutters, the whole job waits. The Financial Times reported in August 2025 that DeepSeek, encouraged by regulators, tried to train its R2 model on Ascend and failed repeatedly, blaming unstable chips, slow interconnect and immature software, even with a Huawei engineering team on site. Training went back to NVIDIA; inference stayed on Ascend. That division of labour has not changed since.

Synchronized AI training chips face an interrupted link, while separate inference streams continue through other chips.

Third, software. CUDA is a moat NVIDIA dug over more than a decade with millions of developers. The numbers Huawei gave at Connect 2026 are 5,200-plus monthly active developers on CANN, its equivalent layer, 61% of them from outside Huawei, with support for more than 90 open-source projects including PyTorch, vLLM and Triton. That is real progress; before CANN was open-sourced in August 2025 there was no such number to report. But the gap to CUDA is not a percentage, it is orders of magnitude. A software ecosystem is not built with budget. It is built with time and with the number of people who have hit its bugs.

When you read this kind of news, I'd suggest keeping two columns. Verified: the Ascend 910C and CloudMatrix 384 are running in production; Ascend serves DeepSeek inference, and Huawei Cloud and SiliconFlow put that online in February 2025; DeepSeek confirmed in a WeChat comment last August that V3.1 uses the UE8M0 FP8 scaling format, built for "the next generation of domestic chips". Planned or reported: 950DT shipments this quarter, the 960DT next quarter, the 520,000-chip SuperCluster, the real yields of Huawei's in-house HBM, and DeepSeek's 160,000 chips. The first column is fact. The second is promise.

Three Things I Notice

Two timelines in the same season. The summer Huawei pulled the 960DT forward by three quarters, NVIDIA's Kyber rack for Rubin Ultra slipped to 2028, more than a year late, over midplane manufacturing problems, according to a SemiAnalysis analysis picked up by CNBC in July. I don't read that as Huawei catching up. I read it as evidence that rack-scale engineering is hard for everyone, NVIDIA included. And that happens to be the thing Huawei has done for thirty years. It is a telecoms equipment company; optical interconnect is its home ground, not an away game.

The HBM arithmetic. SemiAnalysis's production analysis estimates that SMIC's wafers could yield dies for more than a million Ascend chips, but CXMT's roughly two million HBM stacks a year are enough to finish only about 250,000 to 300,000 cards. Bloomberg's sources put 2026 output of the 950DT at 200,000 to 300,000. Two unrelated sources arrive at the same number. It follows that the reported DeepSeek order alone would consume more than half a year's production, which is why the timeline reads "partial capacity by late 2027 or early 2028". The keynote tells you how fast the chip is. The supply chain tells you how many you get.

Logic dies and unfinished AI accelerator packages await HBM memory stacks beside a completed chip package.

Half a ballot. DeepSeek's 160,000 chips are reportedly for inference only, even though Huawei markets the DT variant for training as well. That is the honest version of the "training still bottlenecked, inference substituted first" story. I don't think it is a consolation prize. As agent applications spread, the number of tokens spent on inference is exploding, and inference's share of total compute will keep rising, so winning inference alone is winning the larger half. But it also means new models are still born on NVIDIA hardware. Domestic substitution is replacing the running, not the making.

What It Means for the Supply Chain, the Clouds and the Buyers

Demand has already been pushed toward domestic chips by policy. Last September the Cyberspace Administration told ByteDance, Alibaba and others to stop buying NVIDIA's China-specific RTX Pro 6000D; Jensen Huang has since said publicly that NVIDIA's share of China's data-centre market has fallen to zero; and the H200 received a US export licence in February, but CNBC reported that month that it had produced no revenue, with Huang only announcing orders and restarted production in March. Supply, in turn, is pinned by HBM. For the cloud providers caught in the middle, the selection question is no longer "which is faster" but "which one can deliver next year, and whose software team will sit with mine for three months".

A word from our own experience. When we build AI agents and knowledge systems for business clients, not one of them has ever asked about the FLOPS of the underlying chip. They ask what a million tokens cost, whether latency holds at peak hours, and whether the vendor will still exist next year. A domestic chip serving inference becomes an option for a buyer the moment it brings the price down and puts a service-level agreement in the contract. If all it has is a keynote, it isn't one. We argued in our pieces on China's open-source model strategy and Xiaomi's MiMo that openness at the model layer already spares buyers from picking a side. The chip layer isn't there yet, but on the inference side it is getting closer.

So a suggestion for the next few quarters: watch three things, not the roadmap. Whether the Atlas 950 SuperPoD really ships in Q4. Who the first customers are, and what models they run at what scale. And whether CXMT's HBM output doubles or triples the way Huawei's schedule requires. If any one of those falls through, the Ascend 980 slide doesn't matter. Supply, yield, software maturity and customer cases decide outcomes far more often than the cadence of launch events.

A Question to Leave Open

I'm not going to call a winner here, because the evidence doesn't support a verdict for either side. I'd rather leave a question worth arguing over. When the single chip is no longer the only ring, will the measure of an AI chip company shift from "peak FLOPS" to "can you reliably run a hundred thousand chips as one machine"? If the answer is yes, then NVIDIA's Kyber delay and Huawei's 960DT pull-forward are two rounds of the same fight, not two different fights.

There is a second question behind it. Four times the electricity for 1.7 times the compute is a sum that fails in North America and works in Inner Mongolia. If power is the new scarce resource, does the compute race end up as an energy race? Argue with me in the comments. And if you are choosing an inference platform for your own business, talk to us.

Further reading, by search term: Huawei Connect 2026 keynote; SemiAnalysis CloudMatrix 384; SemiAnalysis Huawei Ascend production ramp; Bloomberg DeepSeek Ascend 950DT Ulanqab; NVIDIA Kyber rack 2028.