Supernodes, Rubin racks, UnifiedBus, CPO, memory, and power are becoming one systems problem.
2026-09-0519 key events
Topic direction
Why watch
The real competition is shifting from single-chip speed to rack, network, memory, power, packaging, protocol, software stack, and deployment economics.
Current read
The next infrastructure boundary will be decided by the coupled system: accelerator, NIC, switch, protocol, memory locality, runtime, power, storage, and operations.
Looped Transformers Hit the Frontline: Why OpenAI, Zhipu, and Academia Converged on the Sam…
Recurrent depth turns model depth from an architectural constant into a runtime variable. This piece traces the route's seven-year lineage and engineering boundaries: Huginn made the case, Nanbeige proved it i…
1 published post
1 published postThe next infrastructure boundary will be decided by the coupled system: accelerator, NIC, switch, protocol, memory locality, runtime, power, storage, and operations./thinking/looped-transformer-recurrent-depth/
Product
MetaRoCE in Depth: RDMA Designed for Loss, an Open Standard for Million-GPU Scale
MetaRoCE is Meta's clean-sheet RDMA transport protocol, opened through OCP. This analysis works through six questions: why it appeared, what it actually is, why Meta opened it, how the protocol works inside, w…
1 published post
1 published postThe next infrastructure boundary will be decided by the coupled system: accelerator, NIC, switch, protocol, memory locality, runtime, power, storage, and operations./thinking/metaroce-nic-intelligence-open-transport/
Market
The Token Factory Ledger: Inference Economics in SiliconFlow's Prospectus
Revenue of RMB 55.3M in 2025 (+653%), gross margin of -24%, net loss of RMB 345M, valuation of RMB 7.74B — SiliconFlow took its “Token Factory” story to the Hong Kong Stock Exchange. This article opens its led…
2 published posts
2 published postsThe next infrastructure boundary will be decided by the coupled system: accelerator, NIC, switch, protocol, memory locality, runtime, power, storage, and operations./thinking/token-factory-siliconflow/
Technology
Agents Pull the CPU Back to the Table: Four Questions from Hot Chips 2026
Six Hot Chips 2026 CPU talks distilled into four questions: the serial bottleneck (Vera's wide front end vs IBM's 5.7GHz), the memory power budget (SOCAMM2's 30-40W full load vs 100W-class RDIMM), three packag…
1 published post
1 published postThe next infrastructure boundary will be decided by the coupled system: accelerator, NIC, switch, protocol, memory locality, runtime, power, storage, and operations./thinking/hotchips-2026-agent-cpu/
Product
Five Networks for the AI Factory: Spectrum-X Multi-Plane and the 512K-GPU Scale
Spectrum-X multi-plane splits the AI factory network into five dedicated planes, scaling 64× to 512K GPUs via 8 planes by 4 rails. CPO reaches volume production: 4× fewer lasers, 5× lower power, 10× better MTB…
5 published posts
5 published postsThe next infrastructure boundary will be decided by the coupled system: accelerator, NIC, switch, protocol, memory locality, runtime, power, storage, and operations./thinking/hotchips-2026-ai-factory-network/
Technology
OpenAI Turns In Its Scorecard: An Analysis of the First Measured Results for Its In-House J…
At Hot Chips, OpenAI delivered Jalapeño's first measured results: 1.5-1.9× the per-watt throughput of Blackwell systems across three models, 1.7-3.6× lower latency. A reckoning against our June estimates: proc…
1 published post
1 published postThe next infrastructure boundary will be decided by the coupled system: accelerator, NIC, switch, protocol, memory locality, runtime, power, storage, and operations./thinking/jalapeno-first-results-audit/
Technology
Dissecting SemiAnalysis AgentX: The First Big Exam for Agentic-Era Inference
SemiAnalysis spent $3M capturing real Claude Code traces, open-sourced 393 sessions, and measured agentic inference across 1,000+ chips. This piece dissects AgentX 1.0 layer by layer: the dataset pipeline, cac…
2 published posts
2 published postsThe next infrastructure boundary will be decided by the coupled system: accelerator, NIC, switch, protocol, memory locality, runtime, power, storage, and operations./thinking/agentx-1-cuda-moat/
Research
Protocol Stacks at 1.6T Ports: The Kernel Stack's Last Mile and Four Ways Out
As ports head to 1.6Tbps, the protocol-stack CPU tax becomes a cost line the size of bandwidth itself. Seven papers, four routes: Presto's full TCP state machine in a switch pipeline (Best Student Paper, 16 co…
7 published posts
7 published postsThe next infrastructure boundary will be decided by the coupled system: accelerator, NIC, switch, protocol, memory locality, runtime, power, storage, and operations./thinking/sigcomm2026-terabit-host-stack/
Market
The Toll Booth, Bought Out: When Intelligence Flows Through Stripe's Pipes
Stripe acquires OpenRouter for $7.5 billion: a 5.8x valuation jump in 82 days. The routing layer proved middle-layer value can be captured, then the payments layer bought it whole. Closing chapter of the toll-…
1 published post
1 published postThe next infrastructure boundary will be decided by the coupled system: accelerator, NIC, switch, protocol, memory locality, runtime, power, storage, and operations./thinking/stripe-openrouter-token-toll/
Research
A Generational Upgrade Without a New Engine: GLM-5.3 and the Second Half of Post-Training
GLM-5.3 ships on the exact same 743B base as GLM-5.2, with every gain from post-training: Terminal-Bench up six-fold, coding near Fable 5, CyberGym past Mythos 5 as the best open-weight score. We break down th…
1 published post
1 published postThe next infrastructure boundary will be decided by the coupled system: accelerator, NIC, switch, protocol, memory locality, runtime, power, storage, and operations./thinking/glm53-post-training-scaling/
Product
Agent Storage Paradigm Reassessment: After FMS 2026, Projections Became Products
About a month later, the four-stage framework is validated and revised with FMS 2026 products. Stages 3 and 4 happen simultaneously. Bus dimension deep-dive: three new paths from CPU-dominated to GPU+DPU-domin…
2 published posts
2 published postsThe next infrastructure boundary will be decided by the coupled system: accelerator, NIC, switch, protocol, memory locality, runtime, power, storage, and operations./thinking/agent-storage-paradigm-reassessment/
Research
KV Cache Scheduling Engineering: After Compressing to 7%
3 published posts
3 published postsThe next infrastructure boundary will be decided by the coupled system: accelerator, NIC, switch, protocol, memory locality, runtime, power, storage, and operations./thinking/kv-cache-scheduling-7-percent/
Company
HPE turns AI factory into stack integration
HPE Discover ties AI factory, network, agent workload, and strategy into one systems bet.
4 linked posts
The HPE series connects agent workloads, AI factory, networking, and strategy. That makes HPE a useful marker for whether enterprise AI infrastructure becomes an integrated stack rather than a GPU procurement project.Prediction: enterprise AI infrastructure vendors will compete on integration proof, not only benchmark numbers./thinking/hpe-discover-2026-strategy/
Market
MaaS exposes inference economics
MaaS articles make the cost stack visible: KV cache, batching, routing, utilization, and token distribution.
2 linked posts
The MaaS articles expose a layered cost structure: model serving, KV cache, batching, utilization, token distribution, and customer acquisition all interact.Prediction: inference winners will need cost observability and routing intelligence as much as model quality./thinking/maas-inference-tech-stack/
Technology
KV Cache rewrites the memory hierarchy
Inference forces infrastructure to treat memory, SSD, and scheduling as one hierarchy.
3 linked posts
Long-context and agentic workloads make cache size, memory bandwidth, storage tiering, and scheduler policy visible to product economics.Prediction: memory/storage vendors will become direct participants in inference architecture decisions./thinking/kv-cache-deep-dive/
operation
Colossus shows compute is not battle power
SpaceX Colossus rental signals that raw GPU count does not equal usable AI capacity.
1 linked posts
The Colossus signal shows that AI infrastructure value depends on turning capacity into usable service. Idle, mismatched, or hard-to-integrate compute becomes a cost problem.Prediction: utilization and workload fit will become first-class due diligence metrics for AI infra./thinking/spacex-colossus-compute-reality/
Technology
Optics debate moves into AI fabric timing
CPO/NPO debate shows optical interconnect is now an investment and deployment timing question.
2 linked posts
The CPO/NPO debate and Optical Shuffle supplement both point to the same question: when does copper stop scaling economically for AI fabrics?Prediction: optical adoption will be uneven and tied to cluster scale, packaging risk, and operational confidence./thinking/cpo-vs-npo-semianalysis-optics-debate/
Technology
Topology and protocol become architecture
MRC, ZCube, RailFly, and Lingqu connect transport, topology, NICs, switches, and AI workload shape.
5 linked posts
The MRC, ZCube, RailFly, Lingqu, and RNG articles collectively show the same shift: topology, transport, and endpoint intelligence are moving into AI architecture decisions.Prediction: the next cluster design cycle will evaluate NIC, switch, routing, and protocol as one architecture bundle./thinking/mrc-protocol-when-nic-becomes-the-brain/
Technology
Power, packaging, and memory become the hidden bottleneck
800V power, 3D stacking, Rubin BOM, and CPU comeback point to non-GPU constraints.
5 linked posts
The 800V, 3D stacking, Rubin BOM, CPU comeback, and supernode articles all push attention toward power, packaging, host processing, memory, and rack cost.Prediction: the next infrastructure advantage will come from solving cross-layer bottlenecks rather than only buying faster GPUs./thinking/analog-chip-800v-datacenter-power-revolution/
Open questions
Q01TrendWill infrastructure advantage shift from accelerators to coupled rack-level systems?