← Thinking Thinking

Yunqi 2026: Alibaba’s Chip-Model-Cloud Full Stack

Yunqi 2026 put three releases on one stage: Zhenwu V900 (3× per generation, 216GB, 2027Q1), Qwen’s zero-human RSI loop (33 rounds, AA 40→45), and a cloud rebuilt for agents (AI storage cost −69%, a 20GW target). This piece reads each in turn — how real the “3×” is against Ascend 960 and NVIDIA B300, where RSI’s boundaries sit, and why the unit of competition has moved from the single chip to the whole system.

2026-09-23Thinking44 min read
Yunqi 2026: Alibaba’s Chip-Model-Cloud Full Stack

On the morning of September 22, at the main forum of the Yunqi Conference in Hangzhou, Alibaba CEO Wu Yongming offered two judgments and one metaphor.

Judgment one: the total quantity of machine thinking will reach more than 1,000 times that of humans, a ratio that sits below 3% today. His reference point is the power revolution: machines already perform 99.9% of physical work, and machine thinking is expected to take the same share; from today to 1,000 times the human level, the headroom is still tens of thousands of times. Judgment two: the defining product of the machine-intelligence era has not yet appeared. The metaphor: today’s AI coding is roughly where the electric light stood in 1882. When Edison lit the first commercial power station on Pearl Street, people used it to replace kerosene lamps; air conditioning, refrigerators, and washing machines did not arrive for decades.

All three statements point at the same thing: thinking is becoming a commodity that can be produced at scale, and the industry is still at the stage of “the power station has just been built.” For a full-stack vendor, the next questions are concrete: how to build the station, how to lay the grid, and what will generate the power.

Alibaba’s answer is a chain of interlocking links: models are getting stronger, and have begun to reverse-design the chips that run them; chips and systems are getting stronger, and are being used by the models to adapt themselves; and the cloud is rebuilding itself around the load profile of agents. At this conference, that chain was broken into concrete numbers for the first time: a 3× generational jump in chip performance, a 69% cut in AI storage cost, 33 rounds of model iteration with zero human intervention, and 500,000 cards in a single cluster.

As one of the few Chinese vendors staking heavy bets on self-developed AI chips, large models, and AI cloud at the same time, how Alibaba organizes these three pieces carries reference value for the domestic compute ecosystem. This article takes the three in order, chips, models, and cloud, and answers four questions: how solid is the Zhenwu V900’s “3×”? Set it beside Huawei’s Ascend 960 and NVIDIA’s B300, and what route has each chosen? What did Qwen’s “self-iteration” change, and how credible is it? And what does a cloud “rebuilt for agents” actually change?

Chips: 3× Performance, 216 GB, and a Spec Sheet with No Absolute Numbers

The Zhenwu V900 is T-Head’s third-generation training-and-inference AI chip. Officially, its single-chip performance is 3 times that of the previous-generation Zhenwu M890; it carries 216 GB of memory and 1,200 GB/s of inter-die interconnect bandwidth; it natively supports FP8 and FP4 low-precision compute; and it covers both high-precision training and low-precision and ultra-low-precision inference. Volume production is set for the first quarter of 2027, about two quarters ahead of the original plan, and the company calls it the most powerful self-developed AI chip in China by compute performance.

Put the roadmap together and the cadence carries more information than any single spec: the 810E in 2024 (96 GB of memory, 700 GB/s interconnect), the M890 in the second quarter of 2026 (144 GB, 800 GB/s, 3× the 810E), the V900 in the first quarter of 2027 (216 GB, 1,200 GB/s, 3× the M890), and the J900 planned for the third quarter of 2028. One generation a year, each generation tripling; the official multiples compound to roughly 9 times from the 810E to the V900. That cadence has now run for three generations. Commercially, the Zhenwu series had shipped 560,000 units cumulatively as of April, and served more than 650 customers as of June. By IDC’s 2025 figures, T-Head ranks second among domestic cloud AI accelerator vendors, with a share of about 16%.

en-fig1 · The Zhenwu staircase: 3× per generation, ~9× in three years
en-fig1 · The Zhenwu staircase: 3× per generation, ~9× in three years

But the launch left a blank from beginning to end: no absolute compute figure was ever given. TFLOPS per card, PFLOPS at a given precision: disclosures of that kind were uniformly absent. The “3×” also came without qualifying conditions: which precision, dense or sparse, per chip or per system. For readers used to reading a compute table, this spec sheet is missing a column.

A missing column can be read another way: once the V900 and the Ascend 960 meet in the same delivery window, the outside world will naturally fill them into one and the same table. What can be anchored horizontally today is capacity and interconnect: 216 GB of memory plus 1,200 GB/s of inter-die bandwidth puts it roughly at the level of Google’s previous-generation TPU v7, slightly above it (per Weijin Research).

500,000 Cards: Folding Cross-Server Communication into Memory Semantics

Beyond the single-card numbers, the real protagonist of the V900 launch is the system. Once a cluster passes ten thousand cards, communication between chips eats a large share of training time, more than 40% by Huawei’s public estimate; a faster card alone spins idle if the system cannot keep up. T-Head’s answer is the ICN Switch interconnect chip: inside a supernode (thousands of cards wired into one logical computer over high-speed interconnect), it gives card-to-card links native memory semantics and unified memory addressing, achieving full-bandwidth high-speed interconnect across a thousand cards, with a unified memory pool in the hundreds of terabytes. In plain terms, higher layers of software no longer have to care which server the data sits on, and a thousand-plus V900s look more like one “super-chip.” Paired with a new-generation supernode server (holding the V900, the ICN Switch, the Panmai 920 smart NIC, and the Zhenyue 510 storage controller), a single cluster can scale to 500,000 cards.

On the same afternoon, Alibaba Cloud launched the Panjiu AL64, the world’s first supernode server to put SNPO — Scale-Up near-package optics (NPO) for tightly coupled multi-card interconnect domains — into a product: 64 cards per chassis, non-blocking scale-out to 1,024 cards, and optical modules moved from the rack panel to sit beside the compute chip. The driver is that copper no longer reaches far enough inside ultra-large Scale-Up domains; SNPO is led and opened by Alibaba Cloud together with the Open Data Center Committee (ODCC), and is not tied to any single vendor’s compute chips or switching system. Whether that openness pays off is still open: the spec leaves interfaces for multiple vendors, and how fast it is adopted depends on how many in the ecosystem actually follow. Lightelligence debuted 3.2T/6.4T SNPO silicon photonics chips at the same event.

The generation now on sale has not been idle: the Lingjun supernode based on the M890 has already opened service in Ulanqab, will expand to Zhongwei in Ningxia, has run models of more than 2 trillion parameters such as Qwen3.8-Max and Kimi K3, and will add service nodes in the fourth quarter of this year.

Three Ways to Assemble It: V900, Ascend 960, and B300

Over the past month, three companies have laid out their next step in AI infrastructure: Huawei launched the Ascend 960 supernode at its Connect conference on September 17, NVIDIA’s B300 is already in volume production and shipping, and Alibaba showed the V900 at Yunqi. Put all three on one table and the first thing that strikes the eye is convergence: Huawei says “expand the unit of competition from a single chip to the whole system,” Alibaba says “like one super-chip,” and NVIDIA’s rack-level design is called NVL72. All three are pushing “one computer” forward.

Start with the single cards:

Zhenwu V900 Ascend 960DT NVIDIA B300
Positioning Training and inference, volume in Q1 2027 Training, ready in Q1 2027 Training and inference, in volume
Memory 216 GB 288 GB HBM 288 GB HBM3e
Memory bandwidth Undisclosed 9.6 TB/s About 8 TB/s
Scale-Up interconnect 1,200 GB/s (inter-die) 2.2 TB/s 1.8 TB/s (NVLink)
Low-precision compute Undisclosed (3× M890) FP8 2P / FP4 4P FP8 about 7.5P / FP4 about 15P
Power Undisclosed Undisclosed About 1,400 W

Three points deserve a close read. On capacity, 216 GB against 288 GB puts the V900 one notch lower, and per-card capacity directly shapes how trillion-parameter models get partitioned. On interconnect bandwidth, the Ascend 960’s 2.2 TB/s is the highest of the three, above the B300’s NVLink, and behind it sits Huawei’s bet on connectivity: replacing 48,000 pluggable optical modules with 5,500 Hi-ONE near-package optical engines, cutting more than 550 kW of power. And in the compute row, the V900 is the only empty cell.

en-fig2 · Three spec sheets, one blank cell: V900 / 960DT / B300 side by side
en-fig2 · Three spec sheets, one blank cell: V900 / 960DT / B300 side by side

Now the systems. Huawei’s Ascend 960 supernode runs 4,096 cards per domain and delivers 8 EFLOPS of FP8 compute plus 1 PB of memory, with round-trip latency inside the domain as low as 2 microseconds; networking multiple supernodes together reaches a maximum cluster of 512,000 cards, extensible toward 1 million with multi-rail topology. On Alibaba’s side, the V900 supernode is characterized by “full bandwidth across a thousand cards plus unified addressing,” caps clusters at 500,000 cards, and uses SNPO to bring optics into the Scale-Up domain. NVIDIA’s strategy remains rack-first: NVL72 puts 72 GPUs in one rack, all linked by copper, with racks networked over InfiniBand or Spectrum-X.

The similarity between the two Chinese routes deserves its own note: both push optical interconnect to near-package, both publicly take cluster ceilings past 500,000 cards, and both chose the phrase “unified memory addressing.” Per the Supernode Definition and Practice White Paper (led by Peng Cheng Laboratory and the Global Computing Consortium, with 38 participating organizations, published in July 2026), this is exactly the core feature that sets a supernode apart from a traditional cluster: several thousand cards look like one machine to the software. The difference lies in how each uses optics: Huawei swaps pluggable modules for in-house NPO optical engines in pursuit of low power, while Alibaba uses the open SNPO spec to decouple the supply chain and leave interfaces for multiple vendors. NVIDIA’s answer is, for now, “make the rack smaller and the bandwidth inside it bigger,” leaving optics between racks.

en-fig3 · Three “computers”: 500,000 cards, 512,000 cards, 72 GPUs
en-fig3 · Three “computers”: 500,000 cards, 512,000 cards, 72 GPUs

Finally the timetables, which put all three to the same test: the 960DT and the V900 both point to the first quarter of 2027, while the B300 is in volume production and on sale with its next generation ramping. That will be the first window in which one set of public data can support a three-way measured comparison: whether they ship on time, whether the software stacks are ready, and whether supernodes can deliver their theoretical compute on real training jobs will all be settled then. One caveat: the table above mixes three tiers of confidence. The Ascend 960DT figures are Huawei-official, the B300 figures are media-compiled (not officially disclosed), and the V900 has no absolute-value disclosure at all. The vendors also disclose precision (dense vs sparse) inconsistently, so the columns are juxtaposed only as an order-of-magnitude reference.

The Supporting Cast: CPUs, NICs, and Storage Controllers

The chip story has a second layer: the V900 is only one piece of T-Head’s full-stack matrix. The Yitian server CPU, whose roadmap was published for the first time at the same event, points to one judgment: in the agent era, the CPU is back on the critical path. Task planning, tool calling, code sandboxes, and retrieval orchestration all depend on single-core performance and memory bandwidth. To that end, the 2027 Yitian 730 uses a fully in-house microarchitecture, with single-core performance up to 1.4 times that of the Yitian 710, while the Yitian 720 supports large-scale concurrency with 192 cores; further out, the Yitian 750 will interconnect directly with Zhenwu chips over the ICN bus. With the Panmai 920 NIC, the Zhenyue 510 storage controller, and the ICN interconnect chip, T-Head has now filled in the data center’s core chips. On the software side, the underlying software stack open-sourced in July 2026 covers more than 96% of the mainstream frameworks, with upstream versions taking no more than a week on average to complete adaptation, the link where domestic chips have most often been stuck.

Models: Self-Iteration Starts Reaching into Hardware

RSI: Handing the Experiment Loop to the Model Itself

The core move of Recursive Self-Improvement (RSI) is to let a model find its own problems, design its own experiments, and improve itself. In Wu’s framing, this is the concrete technical path toward superintelligence (ASI); a year ago Alibaba’s judgment was that “AGI is just the starting point.” At Yunqi, Alibaba gave engineering detail for the first time: with zero human participation, Qwen3.8-Max built its own training pipeline, constructed its own training data, designed experiments, and located defects, iterating for more than one month and completing 33 effective rounds (about one a day on average); the result was a rise on the Artificial Analysis Intelligence Index from 40 to 45, crossing into the score range occupied by top overseas models, which the company says leads all current domestic models.

Unpacked, what gets automated is the research loop itself: hypothesis, experiment, localization, correction. This used to be the working rhythm of researchers; now the model runs 33 laps on its own. That is more worth recording than a point gain, because what changes is the way capability grows: iteration speed is no longer bounded by human working hours.

It is worth asking: why now? At least three things came together. The model crossed a capability threshold, able to execute over long horizons, call tools, and correct from failure feedback. The supporting toolchain was already in place, so the model did not have to invent its own experiment infrastructure. And measurement was in place, so the effect of each round can be read out immediately rather than waiting on a human judgment. Remove any one of the three and autonomous iteration stays a demo.

Precisely for that reason, two qualifications accompany any reading. The starting point and the evaluation system are still human-built: the pretrained base, the evaluation sets, and the initial version of the training framework are all human assets, and what the model changes is the loop after that. Nor is 45 a generational gap; its significance is that it turns “autonomous research” into a process observable on a monthly cadence for the first time. And all these numbers come from Alibaba’s own account, with no third-party replication yet.

The Model Starts Rewriting Its Own Supply Chain

The spillover from RSI is more concrete than the score. Among the engineering results disclosed at the conference, the model worked on both ends of the supply chain at once.

On the software side, it first completed adaptation to T-Head’s new GPU, which it had never touched before: building the inference framework for the next-generation model Qwen3.8-Flash-Next from scratch, cutting long-context first-token latency by 47% and per-output-token latency by 60%, and raising per-instance throughput in everyday chat by 96%. Most of the adaptation of a new model to the vendor’s own new chip was done by the model itself. On the hardware side, given a real chip module spec, it ran autonomously for more than 60 hours and made more than 10,000 EDA tool invocations, completing the full flow from functional architecture to verification to physical implementation, and reducing the module’s physical area by 42%. Among the various claims about “AI participating in chip design,” this is a relatively hard engineering result: the constraints came from a real spec, and the output is a physical implementation that can actually be built.

The training pipeline was not left out either: on the cloud side, the PAI asynchronous Agentic RL framework strings trajectory sampling, reward accumulation, model training, and parameter updates into one pipeline, while CPFS on the storage side guarantees the data supply. The common thread across the three: the model has begun remaking its own toolchain and hardware chain, adding a new source to the next round of capability gains. Once those gains compound, the curve gets steeper than sheer scale expansion alone.

en-fig4 · The model rewrites the model, and the chip
en-fig4 · The model rewrites the model, and the chip

5–10 Trillion Parameters, and a One-Tenth-of-a-Yuan Cost Line

The scale line has not stopped. Qwen4 is training on a new architecture, and the follow-on Qwen4.5 and Qwen5 are planned to extend total parameters to 5 trillion to 10 trillion. On the competitive coordinates compiled by the industry, the frontier threshold currently sits around 3 trillion, and 10 trillion is the first target publicly called out by a domestic vendor. Multimodality is advancing in step: the omni-modal Qwen3.8-Omni, image model 3.1, speech 3.1, the video generator Wan3.0 (first on both the Artificial Analysis text-to-video and video-editing leaderboards), and a batch of new world models, with the next-generation video generator due in November.

The other line is cost. Qwen3.8-Flash open-sourced the next-generation architecture early: by reworking the attention mechanism and the model’s internal information-passing mechanism, training cost falls by about 90%, inference reaches frontier performance with a few billion active parameters, and cache pricing is pressed to one-tenth of a yuan (RMB 0.1) per million tokens. The open-source pool keeps growing: more than 460 open-source models, more than 3 billion cumulative downloads, and more than 300,000 derivative models.

Put the two lines together and it gets interesting: pushing parameters toward 10 trillion on one side, pressing the price per million tokens to a tenth of a yuan on the other. Scale and efficiency are two faces of the same problem, and if thinking is to become something metered like electricity, both have to happen at once.

Cloud: A Ledger Rebuilt for Agents

From AI Native to Context Engine

The framework Alibaba Cloud CTO Li Feifei gave is this: for enterprise agents to reach production at scale, the shortfalls cluster in four pieces of infrastructure: comprehensive and continuously updated context; an execution environment that supports autonomous operation and continuous evolution; trusted collaboration that folds in identity, permissions, and audit; and the ability to move from static rules to real-time sensing, dynamic reasoning, and autonomous action. Against that demand, Alibaba Cloud pushes its stack from AI Native Cloud to Agent Native Cloud and then to the Context Engine, built on three words: Model, Harness, Context.

The generational difference here has a concrete referent: AI Native solved “serving the model,” request in and result out, with requests unrelated to one another. Agent workloads are the opposite: tasks run for days, need checkpoints, retries, and isolation along the way, and context must be continuously updated and carried across sessions. A stateless architecture has no ready answer to any of those needs. What the Context Engine sets out to take back is exactly this infrastructure work of “stateful execution.”

Productization comes down to two things. The first is AgentCore, which standardizes and manages the infrastructure an agent needs to run: long-task support, failure retries, checkpoint recovery, asynchronous execution, plus security isolation, identity and permissions, and full-link audit. The company says it can raise task completion rate by 99% and lower total cost of ownership by 70% (vendor-disclosed, not independently verified). The price of managed service is migration cost: once execution, state, and memory are standardized inside the platform, the application is bound one layer deeper to the cloud. The second is a set of runtime environments including Agent Sandbox, Agentic OS, and Agentic Computer. The concrete problem this launch set aims to solve: when an agent’s task goes from “answer a question” to “run continuously for days,” the cloud cannot leave every application to implement its own checkpoints, retries, and sandboxes, just as nobody in earlier years should have had to implement database transactions themselves.

The Moving Bill for Context: A New Kind of Statement

On the infrastructure side the numbers are harder. The new-generation CPFS pulls parallel file storage to hundreds of TB/s of throughput, hundreds of millions of IOPS, and 100 PiB per file system; the three effect figures the company gave are: average model startup time down 50%, peak compute utilization up 30%, and AI storage cost down 69%. On the inference side, Tair KVCM unifies orchestration across in-memory pools, caches, and remote shared storage, reaching a 99% effective cache hit rate and cutting per-token cost by 50%; in customer production environments, KV CacheStore expands the cache coverage window by 900%, cuts first-token latency by 54%, and lifts throughput by 20%.

The mechanism behind these numbers is not complicated: agent-era workloads carry ever-longer contexts and ever more state, and the KV cache has become a third cost center sitting between compute and storage. Whoever moves and stores it more cheaply can quote a lower inference price. Storage, caching, and scheduling, once inconspicuous line items on a cloud bill, are becoming the main battlefield that decides the economics of inference.

Widen the view further: Alibaba Cloud targets operating datacenter capacity above 20 GW by 2032, and states plainly that customer demand far outruns current supply and that supply-chain shortages limit how fast it can expand; in the fourth quarter of this year it will first add supernode service nodes. According to analyst estimates cited by Wallstreetcn, at USD 12 billion to 15 billion of annual revenue per GW, 20 GW maps to a trillion-RMB revenue roadmap. That is only linear extrapolation; whether it lands requires all three of power, chip supply, and demand.

en-fig5 · The cloud ledger: cuts start with the moving fee
en-fig5 · The cloud ledger: cuts start with the moving fee

Devices: Handing Intelligence to the User

The device layer is the touchpoint of a different path: Qwen AI hardware showed three products at once, AI glasses N1 / N1 Pro and AI clip-on earbuds, with reservations opening that day and sales on October 13. On the office side, Qwen Office launched “enterprise context” and the QwenNote A2 recording card, reaching more than 30 million users in its first month; on the mobile side, Qwen Intelligence offers handset makers a full-stack AI solution. On desktop, the open-source Qwen-27B already runs smoothly on local hardware. None of them is a heavyweight on its own, but together they show one thing: cloud model capability is seeping into concrete user scenarios through hardware and office software. Their shared weakness is just as clear: the hardware is only starting out, and experience and ecosystem will decide how far it goes.

Summary and Judgment

Compress the Yunqi 2026 launch list into one sentence: Alibaba has changed the unit of competition for compute from a single chip to a whole system, and folded the system’s three parts (chips, models, cloud) into one feedback loop.

The unit of competition is already the system. Over the past month, Huawei, Alibaba, and NVIDIA have used nearly identical language in the same direction: one computer, unified memory, supernode. A single-card gap no longer decides the outcome on its own; the degrees of freedom in system design (optical interconnect, memory semantics, open standards) have become the new variables. The two Chinese vendors push optics into the Scale-Up domain and take cluster ceilings past 500,000 cards, while NVIDIA stays rack-first, and the three assemblies are three different answer sheets.

What is more real than the language is the numbers. Mutual transformation has begun producing measurable results: the model cut chip module area by 42%, raised its own inference throughput on the new chip by 96%, and the cloud cut AI storage cost by 69%. But all these figures come from vendor accounts and lack independent replication, and the disclosure strategy is itself part of the competition: giving only multiples and not absolute values preserves the room to “say it at delivery time.”

The hardest test is on the timetable: the first quarter of 2027 is the first validation window. By then the V900 and the Ascend 960DT will be ready in the same quarter, the B300 will be on sale with its next generation ramping, and the three system designs will face their first same-stage test. Three nearer observation points: the execution of Alibaba Cloud’s supernode expansion this fourth quarter, the next round of RSI iteration numbers, and the training progress of Qwen4.


Disclosure: This article is based on public talks and official press releases from the 2026 Yunqi Conference (September 22–24, Hangzhou), along with reporting and interpretation from STAR Market Daily, Cailian Press, Shanghai Securities News, Southern Metropolis Daily, Qudong China, Weijin Research, and other outlets, cross-checked before writing. Some figures are vendor-disclosed and not independently verified. Nothing here constitutes investment advice. Figures are current as of September 23, 2026.