← Thinking Thinking

Five Networks for the AI Factory: Spectrum-X Multi-Plane and the 512K-GPU Scale

Spectrum-X multi-plane splits the AI factory network into five dedicated planes: scale-up, scale-out, scale-across, plus newly announced scale-in and AI…

2026-08-27Thinking16 min read

Hot Chips 2026 specialist · AI factory networking

The overview set judgment two: open interconnect is the second united front: UCIe, UALink, unified Ethernet, NVLink Fusion; everyone outside NVIDIA's own garden is standardizing. This piece enters the main battlefield: what those standards run on is a factory network being redesigned. Day 2's three networking talks (Spectrum-X multi-plane, BlueField-4, Thor Ultra), each on its own a product launch; put on one timeline, we see an elevation: the network moves from a tool that connects compute to the AI factory's own skeleton and operating system.

Five networks for the AI factory: which traffic each plane serves
Five networks for the AI factory: which traffic each plane serves

1. Spectrum-X Multi-Plane: Five Dedicated Networks

The overview covers this in brief; this section dissects how the five networks divide labor.

One table first (fig1): scale-up handles inter-chip proximity bandwidth, still NVLink (NVL72 = NVLink spine + switch trays; see the AgentX piece, not rewritten here); scale-out handles cluster east-west; scale-across handles multi-campus, multi-factory interconnection (the concept is covered in our scale-across piece); scale-in is new this year, bound to BlueField-4 (§3); AI context serves context traffic for agent workloads.

Mechanism one: how multi-plane isolates. The general-fabric approach is QoS prioritization on one network: training synchronization flows, KV-cache reads, storage checkpoints, and management traffic share the same switch fabric, fighting for priority through schedulers. The multi-plane approach is physical separation: five traffic classes each run their own plane; tail latencies don't pollute each other, and one plane's failure domain doesn't spread to the other four. The unit of isolation upgrades from scheduling policy to physical topology: that is the substantive meaning of "dedicated."

Mechanism two: scale-in and AI context are the two new faces. Scale-in's release name binds directly to BlueField-4 ("BlueField-4 Spectrum-X scale-in network"), and the official slide pairs it with the "AI factory scale expansion 64×" claim: the leap from 8K to 512K GPU (fig2) takes it as one of its preconditions. AI context's very name writes agent workloads into the network layer: the volume and access patterns of context and KV-class traffic justify a dedicated network. The load-bearing claim is cited with official attribution: a single general-purpose fabric cannot serve all five traffic classes: that is NVIDIA's thesis (same source as the overview), not escalated to our judgment here; what we add is the mechanism: the five traffic classes differ in latency sensitivity by orders of magnitude, and mixed-network scheduling costs diverge at 512K scale.

From 8K to 512K: a 64× scale leap
From 8K to 512K: a 64× scale leap

2. CPO Volume Production: The Optical Interconnect Watershed

The one line "CPO in volume production" in the overview, with metrics here: micro-ring modulators in mass production, laser cost −4×, MTBI +10×. For our corning-glassbridge and cpo-vs-npo pieces, this is the increment landing: those pieces debated routes; this conference delivers the production ledger: micro-ring production solves yield and cost curves, and −4× lasers and +10× MTBI directly change deployment economics. The micro-ring modulators, now in production, land inside Spectrum-X CPO switch modules — the deployment-economics gains cash in at the switching layer. The optical-interconnect basics are not rewritten (the RNG and Optical Shuffle pieces are the foundation).

3. BlueField-4: From DPU to AI Factory OS

Three facts are set in the overview (co-design, Astra, factory-OS positioning); this section dissects the mechanism of the positioning leap.

The depth of the co-design first: a member of the Vera Rubin seven-chip system, "BlueField-4 Spectrum-X scale-in network" is its release name: the NIC and the network plane ship together; scale-in's data path lives on BlueField-4. Astra's 7Tb/s aggregation is its throughput face. ServeTheHome's on-site sentence, on record: "a completely different class of device from Thor Ultra": between an 800GbE NIC and a factory OS lies an order-of-magnitude difference in system positioning (see §4).

The load-bearing mechanism for judgment three sits here. Our MRC piece (June 1) set the thesis: push intelligence from the switch to the NIC. This year's increment advances along the same line by two steps: at the protocol layer, MRC evolves to MRC++ (data-path details await full-text reading); at the positioning layer, the DPU narrative leaps from "offloader" (doing chores for the CPU) to "AI factory operating system" (the factory's data-plane and IO orchestration rights). Two steps, one direction: what the NIC holds has been elevated to orchestration rights.

The orchestration-rights division (cross-link to the Intel piece): this piece's BlueField "factory OS" means data-plane and IO orchestration; the Intel piece's "orchestration layer" means control plane and workload orchestration (where agent runtimes land). The two lines meet at the rack; each concept holds its side.

4. Broadcom Thor Ultra: 800GbE and MRC++

The lead sentence is in the overview; this section adds the ledger behind it.

The Broadcom-OpenAI 10GW collaboration is the backdrop: self-designed accelerators plus Ethernet networking, ramping from H2 2026 through 2029; the single-cluster 100,000+ chip claim comes from that context. Lineage on record: the predecessor Tomahawk Ultra (July 2025, die-to-die interconnect); Thor Ultra extends that line from switch silicon to the NIC. The ConnectX-8 comparison data and the MRC++ data-path details have not yet been published.

5. Judgments and What to Watch

Judgment one: traffic layering determines network layering; the cost of five networks is stated first, then broken. State the cost: five networks mean five operations stacks, five procurement paths, five failure domains; the fragmentation critique has real weight. The break is at scale: general fabrics rely on QoS to prioritize within one network, and scheduling costs rise with scale; at 512K GPUs, the isolation benefit of physical separation starts to outweigh five operations stacks. Both ledgers go into the checkpoints (the CPO cost curve is the overview's variable two).

Judgment two: CPO volume production is the optical-interconnect watershed. The line from demo to deployment is production metrics: micro-ring production, −4× lasers, +10× MTBI move our earlier route debates onto deployment economics. How fast the cost curve falls determines the speed at which optics penetrates from the switch side toward the host side.

Judgment three: the network becomes an OS — BlueField-4's positioning change is a system-level signal. The DPU's leap from offloader to factory OS (§3) takes orchestration rights away from host CPUs and switches; MRC's evolution to MRC++ is the protocol-face evidence of the same leap. Network devices begin defining how the factory runs — the operating-system version of "elevated to skeleton."

What to watch: MRC++ data-path details; Thor Ultra vs ConnectX-8 comparisons; the CPO cost curve and MTBI (overview variable two); scale-in and AI-context field materials; the 10GW ramp cadence.

Forward commitment: when the MRC++ full read lands and the CPO cost curve updates, we will come back and settle the account.

Three networking talks, positioned: NIC, DPU, and multi-plane switching
Three networking talks, positioned: NIC, DPU, and multi-plane switching

Further reading: Hot Chips 2026 Overview · Four Routes Against the Memory Wall · Intel's Triple Launch · Cloud Silicon Generation Two