
On September 17, Huawei Connect 2026 opened in Shanghai, and the Ascend 960 supernode took the stage: 4,096 cards drawn into a 2μs round-trip latency domain. The next day, at the same event’s UnifiedBus (Lingqu) Technology Forum, something less conspicuous happened: the UnifiedBus Base Specification 2.1 was officially released. 586 pages.
The document’s timeline deserves a first look. Its file properties show creation in the early hours of September 15, with last edits later that same day. By September 16 at the latest it had appeared on the UnifiedBus community site, corroborated by a long-form deep read that surfaced within a day or two. On the 18th, Huawei officially released it at the conference’s UnifiedBus Technology Forum, with detailed interpretation. The day before was the Ascend 960 supernode keynote; the day after came the protocol text that will keep supernodes of this kind running for years. Product and specification have entered a same-week cadence, and that is no coincidence: the 960 is about the system, UnifiedBus is about the protocol, and the former is currently the latter’s largest buyer.
Huawei’s official summary of 2.1 comes in four sentences: new 224G/256G generational rates, a new reliable reroute mechanism, higher system bandwidth efficiency, and lower communication latency under multipathing. Behind those four sentences sit 12 categories of revision on top of 2.0.1. This article does one thing: take apart those 586 pages to see the big signals hidden in a small version. One is a technology signal: the compute interconnect is starting to set its own rate cadence. The other is an industry signal: for the first time, the UnifiedBus ecosystem has a timeline for stepping outside Huawei’s walls.
A Small Version, Big Meaning
Let us be clear about what 2.1 is not. It does not touch the architecture. Everything 2.0 established stays exactly as it was: the six-layer protocol stack (physical, data link, network, transport, transaction, function), Load/Store semantics, credit-based flow control, multipathing. Vendors building on 2.0 do not have to tear anything up, and the baseline line in the revision record states it plainly: an update based on UB 2.0.1. The architectural major-version upgrade is not part of this release; the specification says nothing about future versions, and the industry expects it to be called UB 3.0.
So what did 2.1 change? The revision record lists 12 categories: three new rate tiers; the RS(256,240) FEC code family; UB Gearbox mode; dynamic packet-rate control; reliable reroute; three fused-write transactions; two queue transactions; the CFG7 slim transaction-layer packet; memory-management RAS features; access authentication and TEE optimizations; the UB neighbor discovery protocol; plus editorial corrections. Boiled down, it is three things: add capability, squeeze efficiency, harden operations. In the words of Cheng Chuanning, a UnifiedBus protocol and chip architect, speaking at the forum, 2.1 keeps evolving on top of 2.0’s full openness, with reliability further strengthened, to give the industry a more mature and more stable technology foundation.
The weight of that statement depends on the audience. Connecting and running long-term are two stages of supernode commercialization. When UB 1.0 shipped with the Atlas 900 supernode in March 2025, the industry validated the first: bus semantics could turn a cabinet of cards into one computer. By this forum, commercial deployments had passed 1,000 (the figure from Zhu Zhaosheng, chief strategy officer for Huawei Computing), and the question had shifted to the second: running workloads at scale, with high reliability and high efficiency, over the long haul. Every revision in 2.1 points at that question. The sections below take it apart along four lines: physical layer, reliability, transaction layer, and scale-and-operations.

Three Rate Tiers: Setting Its Own Tempo
The physical layer is where 2.1’s changes concentrate, and where the information value is highest. The specification divides it into two modes: PHYMode-2 treads the Ethernet rate ladder, taking the compatibility route; PHYMode-1 lives in UnifiedBus’s custom-rate domain, serving the high-speed custom tiers. 2.1 adds one tier to each of the two modes, plus one further custom rate: three new rates in total.
The first tier, 212.5Gbps, rides PHYMode-2’s existing ladder. That mode’s original rate sequence was 2.578125, 25.78125, 53.125, and 106.25Gbps, all aligned to the Ethernet ecosystem, with electrical characteristics referencing the IEEE 802.3 family; the newest tier, 212.5Gbps, references 802.3dj, the 200G-per-lane standard. What it means is compatibility: reusing the Ethernet world’s mature SerDes, cable, and optical-module supply chain, so vendors moving onto the UnifiedBus stack do not have to replace their physical-layer supply chain. In linear-optics scenarios without retiming, 212.5Gbps references OIF CEI-224G-LINEAR-PAM4; the specification also holds out an optional high-insertion-loss figure for DAC cable scenarios: a maximum of 42dB (bump to bump, at Nyquist frequency), physical headroom for long-reach copper interconnect.
The second tier, 224.0Gbps, is Custom rate 4, running on PHYMode-1 with electrical characteristics referencing OIF-CEI-06, aligned to the 224G linear-optics ecosystem. The Ascend 960 supernode’s Hi-ONE optical interconnect engine (7.2T per engine) belongs precisely to this generation: the spec text was finalized on September 15, Hi-ONE took the stage on September 17. The rhythm of protocol paving the way for product is plain to see.
The third tier, 256.0Gbps, also runs on PHYMode-1. Ethernet has no such tier, and its alignment target is not on the network side: this rate is a door into the supernode left open for PCIe devices. UnifiedBus’s ecosystem positioning is open access and diverse components, and PCIe is the component ecosystem with the largest installed base today. 256.0Gbps appears on no Ethernet roadmap; it is a rate the compute interconnect defined for itself.
Two supporting mechanisms come with them. The new FEC code family, RS(256,240,T=8), corrects 8 symbols (GF(2^8)) and sits alongside the existing RS(128,120) family, so error protection at the high-rate tiers upgrades together with the rates. More interesting is the UB Gearbox: a low-speed link on one side and a high-speed link on the other, connecting two UBPUs (UB Processing Units, the devices on the bus that run the UnifiedBus protocol) that work at different rates; the rate ratio is 1:2 and the lane ratio is 2:1, and it interleaves and de-interleaves at bit granularity, oblivious to the protocol, requiring only that both sides run PAM4. What it buys is investment protection: 100G-era gear joining a 200G-era network interoperates with a Gearbox in between, no rip-and-replace required. The other side of the bargain deserves attention too: the Gearbox is a real component hung on the link, and clock recovery plus bit interleaving add latency, power, and materials cost; the spec also defines only the single 1:2 rate ratio, with PAM4 required on both sides. Investment protection has limits; mixing rates across more than two generations still takes re-planning.

Put the three rates side by side and a conclusion surfaces: the compute interconnect is forming a tempo of its own. 212.5G aligned with Ethernet is industry compatibility; 224G aligned with optical physics is product-first; 256G aligned with PCIe is ecosystem ambition. Compatibility with industry holds at the standard rates, but the tempo no longer follows the network interconnect’s generational clock. This is the sharpest judgment in 2.1. The specification does not say it outright—put the three rate tables side by side and that is what they mean.
Three-Level Reliability: When a Link Drops, the Pointer Resumes
2.0’s reliability came in two levels: link-layer hop-by-hop retransmission for occasional bit errors, and transport-layer end-to-end retransmission for packet loss and reordering. The missing level was the network layer: what happens when a whole link goes down. 2.1 adds reliable reroute, and the mechanism is worth a close read, because it is the protocol-side foundation under the supernode availability claims.
The setup: when two UBPUs have two or more reachable links between them, UBFM (the UB Fabric Manager, the system management software responsible for topology and port configuration) configures Port1 as the backup port for Port0 after power-on. When a failure occurs, the port on the receiving side detects the link break in the RX direction (the physical-layer LinkUp signal drops to zero). The upper processing unit, on receiving the report, looks up an internal table entry for the CNA address of the failed link’s far end (a short network address, the network-layer house number that UnifiedBus assigns to each device) and reads the receive pointer of the broken port (RcvPtr_Packet: a pointer the receiving side maintains per packet, marking where the start of the current or next packet to be received sits in the peer’s Retry Buffer). A reliable-reroute LastAck message then goes out over the backup path. On receipt, the sending side resumes from the position the pointer marks in the Retry Buffer, out through the backup port. The specification sets two time limits for the whole process: the LastAck should complete within 500ms, and a sender that has not received it within 1s treats the case as a timeout. Retry Buffer space polluted by the broken link is released only after the resumption is confirmed complete.
There is also a plan for scenarios without a direct backup path: when two UBPUs share only a single link but each can reach a third UBPU, the LastAck and the rerouted data can be forwarded through the third. The specification’s example is the triangle topology of UBPU0 talking to UBPU1 via UBPU2. The cost is one extra hop; what it buys is freedom in topology design.

Why this matters becomes clear when read alongside our Ascend 960 piece. The supernode swaps 48,000 optical modules for 5,500 optical engines; as integration rises, the criticality of a single component rises with it: one bad unit affects more links. Huawei claims 99.8% availability for the system and a doubled mean time between failures, and the support for that claim does not stop at hardware redundancy; the protocol side must be able to move traffic off a failed link without loss. Network-layer reroute, plus the memory-management RAS that 2.1 introduces in the same release (covered below), forms the protocol half of that claim. Hardware makes links fail less; the protocol keeps service alive when they do.
Efficiency and Latency: Packing the Completion Ack into the Packet
The transaction layer is 2.1’s most technically substantial part, and it is where the last two of the official four sentences (bandwidth efficiency, communication latency) mainly land.
The first group is the fused-write transactions: three new write transaction types (Write_with_atomic_store_add, Write_with_remote_sync, and Write_with_remote_sync_with_be). Long names, one mechanism: fold synchronization semantics into the write transaction itself. Traditionally, data goes over and the completion acknowledgment makes its own round trip; fused writes fold the completion signal (and even atomic accumulate and response semantics) into the same write packet, so in a multipath, out-of-order environment the data arrives one-way and the confirmation arrives with it, with no standalone ACK round trip. The specification’s appendix examples give a quantitative reference: a small-IO request transfer in Push mode needs only 0.5 RTT; a large-IO response transfer using Send plus Write_with_immediate is also 0.5 RTT. A control group sits in the same example table: for the same large-IO response transfer, plain Send plus Write takes 1.5 RTT—the Write_with_immediate variant saves the entire round trip. That 0.5 RTT is a product of fused protocol semantics; it does not mean the physical transfer became half-length. The data still takes the full path; what is saved is protocol behavior: the confirmation no longer occupies a round trip of its own. The cost of fusion lands on the semantic side: data, synchronization signals, and accumulate operations ride in one packet, and if the packet is lost the whole group is replayed, with execution exceptions caught by the transaction acknowledgment (TAACK) return path. The complexity has not disappeared; it has moved from the application layer into the protocol stack.
The design throws in a by-product, too. In the Write_with_atomic_store_add example, the target side receives the transaction packet and atomically accumulates statistics into a designated statistics address space, with exception information carried back in the response. Control-plane traffic such as task counts and scheduling feedback can ride along with the data traffic instead of building a separate lane.
The second group is the queue transactions: Atomic_store_with_return_status and Atomic_sync. The former hands off a task and brings queue status back with it: whether the queue is full, how deep it runs, whether it is congested, all known to the initiator the moment delivery completes. The latter performs efficient synchronization in the cache-coherent space. UnifiedBus’s existing transaction primitives cover the Send, read, write, and atomic families, but queue semantics had no entry point of their own; these two fill that cell and give pipelined task distribution a protocol-level primitive. The cost is that queue depth and congestion state become hardware state the target side must maintain in real time and return with each delivery, adding one status lookup to the delivery path. A delivery that carries its own queue status hands the scheduler exactly the input dynamic decisions need; the direction here is unmistakable.
The third group slims the packet header. A new CFG7 configuration supports a slim transaction-layer packet format: a network-layer header based on 24-bit CNA addresses connects straight to the transaction layer, the transport layer can be configured into bypass mode (the specification’s phrasing: the transaction layer can call network-layer services directly, reducing protocol overhead), and per-hop extension headers and check fields fall away. Paired with an upgrade to packet-rate control, in which the sender inserts idle blocks (No_Operation Blocks, or NOBs) to ease the receiver’s packet-rate processing load, 2.1 adds a dynamic mode on top of 2.0’s static-only approach: the insertion rate is computed from average packet length, data unit length, and receiver burst capacity, so a mix of short and long packets no longer floods the receiver with bursts. The cost of the saved fields is just as clear: flows that bypass the transport layer no longer enjoy its end-to-end retransmission and checks; error detection and recovery fall back on link-layer hop-by-hop retransmission and the transaction layer itself. The slim packet suits traffic with deterministic paths where upper layers can cover reliability; it is not a universal format. Idle blocks, for their part, are a clear bandwidth tax: the more conservative the insertion, the lower the effective bandwidth, and the point of dynamic mode is to collect it at exactly the right rate.
Together the four changes form a single intent: every bit of on-wire overhead has to earn its fare. The bottleneck in intra-supernode communication is often not peak bandwidth but the share of payload that protocol overhead eats. What 2.1 does at the transaction layer is bend that share back the other way.
Scale and Operations: 24-Bit Addresses and a List of International Standards
The address space is where 2.1’s pragmatism shows most clearly. The network address CNA expands from 16 to 24 bits, and the addressable space grows from 65,536 to 16,777,216, a 256-fold expansion. The driver is arithmetic: a 4,096-card supernode unit, multiple endpoints per card, and expansion toward SuperCluster mean the 16-bit headroom starts to run tight right around this generation. The specification does not switch aggressively: the 16-bit format and the IP format are both kept, 24-bit is an addition rather than a mandate, and existing designs keep running.
Operations-side enhancements come in three parts. Memory-management RAS: when UMMU (the UB Memory Management Unit, the component that performs address mapping and access-permission checks inside a device) hits an exception in address translation, token validation, or permission checks, it reports it as a Class A error through the event queue, making faults on the memory path observable. Access authentication aligns with industry standards: DMTF SPDM (DSP0274) plus the RFC 9334 remote-attestation architecture, with the UBPU as attester, the verification server as verifier, and UBFM as relying party, supporting passport mode and background-check mode, so device identity and trust state gain a standardized measurement vocabulary. The TEE (trusted execution environment) extension tightens the boundary of the trusted computing base (TCB): the specification defines the TCB as the UBPU hardware root of trust, the hardware security module, and TEE-related compute units, with general-purpose operating systems not on the list. Add ULDP, a neighbor discovery protocol based on IEEE 802.1AB LLDP: devices identify each other automatically after power-on.
The normative references are even more telling when pulled out on their own. The list includes IEEE 802.3dj, 802.1AX-2020, RFC 2131, and OIF CEI-224G, spanning four categories: Ethernet physical layer, link aggregation, address assignment, and electrical interfaces. For a compute interconnect specification led by a Chinese vendor, the reference list reads almost like a collection of international standards. This is not posturing; it is an engineering judgment: reuse mature underlying standards wherever possible, and spend home-grown effort on the semantic and transaction layers, where there is genuinely no ready-made answer to borrow.
Ecosystem Realities: 1,000+ Deployments, 52 Identifiers, and a 2028 Timeline
Beyond the protocol, the September 18 forum also produced a set of ecosystem numbers worth reading on their own.
First, the official scorecard. In his remarks, Zhu Zhaosheng said that after a year of development, more than 1,000 Huawei supernodes are in commercial use; 30+ organizations have taken part in supernode definition and practice; and 7 technical documents, covering base technical specifications, component reference designs, OS management, and more, are fully open. He likened compute-interconnect communication to Mandarin: vital to the entire compute infrastructure, with UnifiedBus staying equal, neutral, and fully open in its technical specification, supporting upstream and downstream partners.
Then, the registry. In the snapshot as of September 18, the vendor identifier registry on the UnifiedBus community site lists exactly one vendor: Huawei. The device and module identifier count stands at 52, all of them Ascend 950 PR series; module identifiers, zero. On one side, 1,000+ deployments and 30+ organizations from the forum; on the other, one vendor and 52 device records in the registry. The contrast need not be bad news; it shows the ecosystem is still at the stage of partner evaluation, reference designs, and IP project launches, and that actual chip-level registrations have yet to arrive. But it marks the real position: paper openness walks ahead, and silicon arrival needs a timeline.
The timeline is here, and it is the most substantial piece of news from the forum. Wang Zhen, deputy general manager of Netforward, unveiled a UnifiedBus-based product roadmap: the UB IP project launched in Q2 2026; a three-party interoperability test comes at the end of the year; and two product lines, IO Die and Switch, are expected to complete launch and mass production in 2028. It is the first chip-level production timeline given by a third-party company since UnifiedBus opened. Released in the same window: Volume 1 of the UnifiedBus Compatibility Test Specification 2.0, covering test cases for the physical, data link, and network layers, with the community site saying UnifiedBus compatibility testing capability is now in place. Protocol, testing, third-party silicon: the full set came together at one forum.

The application side showed signal, too. Ning Yunxiao, a core developer of vLLM Ascend, presented a KV Pool practice built on UnifiedBus memory semantics: pooling, sharing, and uniformly accessing memory resources across devices and tiers to expand usable KV Cache capacity and cut explicit data movement and I/O overhead. It is the first time UnifiedBus Load/Store semantics have been publicly validated by a third-party developer on a critical inference-side workload, and it lands right on the growth curve of long-context and agent workloads. UBSIM, a simulation platform from a Zhejiang University team, supports cross-layer co-simulation from parallelism scheme to network topology to compute allocation, giving partners a tool to evaluate UnifiedBus architectures without buying hardware.
The Open Boundary: Free, No Patent Suits, but No Forks
The other half of the ecosystem is the license terms. The license accompanying 2.1 upgrades to V2.0 and continues the same structure: a free copyright license, non-exclusive, non-sublicensable, non-transferable, granted only for implementing the UnifiedBus specification to develop compliant products. The prohibitions can be checked word for word: no revision, alteration, or modification of the UnifiedBus specification or derivative works by any means, and no excerpting or quoting of any content from the specification to develop other standards. As the specification’s text states: Huawei may update this specification as it deems necessary or appropriate, and you agree that Huawei’s updates to this specification do not require prior notice to you.
Read Zhu Zhaosheng’s ‘equal, neutral, fully open’ alongside this no-fork clause, and the true shape of UnifiedBus’s openness comes into focus. Three models line up for comparison: NVLink is closed and proprietary, unavailable to companies outside the alliance; UALink and UEC are consortium-governed, members sitting at one table to write the rules; UnifiedBus is unilaterally open: the specification is free to read, free to implement, free for commercial use, but control of the text sits 100% with Huawei. Ecosystem members can build chips (Netforward), build complete systems (Talkweb and others have built supernodes to 2.0 on their own), and write upper-layer software (openEuler components are already open source), but they cannot alter the protocol itself, nor move its content into another standard.
It is a rational but not generous choice. The upside is speed: 2.0 to 2.1 took only a year, and the roadmap is not hostage to alliance politics; the price is a ceiling on how deep ecosystem members can participate, with the scale ambitions of the buyers unable to reach the specification layer. The more likely form of the UnifiedBus ecosystem is a deep domestic collaboration circle, not an international standards body. The two ecosystems each earn on their own terms; different models, no ranking implied.
Summary and Judgments
This article has taken apart the 586 pages of the UnifiedBus Base Specification 2.1: the physical-layer tempo of three new rates and the Gearbox, the three-level reliability completed by network-layer reroute, the transaction-layer upgrades of fused writes and queue transactions, the 24-bit addresses and the international-standards reference list, and the ecosystem timeline that took shape at the forum for the first time. Three judgments close it out.
First, the compute interconnect’s rate cadence is starting to stand independent of Ethernet. 212.5G on the Ethernet ladder is the foundation of industry compatibility; 224G on optical physics is the runway for Huawei’s own products; 256G toward PCIe is the door of ecosystem ambition. Each tier has its own alignment target, which says the rate table is self-authored, not copied from the networking world. Going forward, one test of a compute interconnect roadmap’s maturity will be whether its rate generations keep evolving independently.
Second, 2.1 is the protocol precondition for the 960 generation to run sustainably. The specification was finalized on September 15 and released on the 18th; the 960 supernode took the stage on the 17th. The same-week cadence shows Huawei’s specs and products are meshed in lockstep. Reroute and memory RAS directly support the 99.8% availability claim, and the 24-bit address leaves addressing headroom for scaling beyond 4,096 cards. When reading the supernode story, the protocol version line and the product version line should be read side by side.
Third, the UnifiedBus ecosystem is moving from paper openness toward silicon, but the ceiling is in plain sight. 1,000+ commercial deployments, Netforward’s 2026 interop and 2028 production, and the compatibility test spec in place are the first set of proof points. A registry that still holds exactly one vendor, Huawei, marks the ceiling. The no-fork license model makes it a deep domestic collaboration ecosystem rather than an international-standards one, and whether Netforward’s production lands in 2028 is the first touchstone for this judgment.
Four things to watch going forward: UB 3.0’s timeline and architectural direction, and whether it touches Load/Store semantics; the results of Netforward’s year-end three-party interoperability test; when a second vendor appears in the registry; and whether KV Pool enters the vLLM mainline. When any one of the four lands, UnifiedBus will be worth a return visit.
Sources
- UnifiedBus Base Specification 2.1 (UB-Base-Specification-2.1.0-zh-clean.pdf, 586 pages, finalized 2026-09-15, unifiedbus.com): rate table 3-1, Data_Rate_Support_2, the FEC code family, §3.5 UB Gearbox, §5.3.9 reliable reroute, the chapter 7 transaction-layer opcode table, appendix I table I-3, and the revision record (except where noted, every mechanism detail and number in this article comes from the specification text)
- unifiedbus.com: homepage announcement cards (the official four-sentence summary); launch announcements for the UnifiedBus Base Specification 2.1 and UnifiedBus Compatibility Test Specification 2.0 Volume 1 (2026-09-18); the UnifiedBus Technology Forum news page (published 2026-09-20: 1,000+ deployments, 30+ organizations, the Netforward timeline, KV Pool, UBSIM); the vendor-identifier and device-identifier pages of the registry
- Huawei Connect 2026 day-one release materials (the Ascend 960 press release, keynotes, and HiSilicon Optoelectronics disclosures; itemized verification in this site’s Ascend 960 article): 4,096 cards, 5,500 Hi-ONE optical engines, 2μs RTT, 99.8% availability, and a doubled mean time between failures
- WeChat account Jikuibu, ‘Huawei Releases UnifiedBus Base Specification 2.1’ (2026-09-16/17): a long-form deep read of the entire specification; a reference for the industry intent behind the three rate tiers and for ecosystem positioning (where it diverges from the specification, such as the 42dB insertion loss and pointer-field spelling, this article has corrected against the specification text)
- Related articles on this site: Same Node, Double the Compute: How Ascend 960 Rebuilds the Chip, the Supernode, and the Cluster, Huawei’s One Network: The HC 2026 Networking Panorama, and Deep Analysis of the Lingqu Protocol
