← Observe Developing · Medium

AI Infrastructure Rebuild

Supernodes, Rubin racks, UnifiedBus, CPO, memory, and power are becoming one systems problem.

2026-09-0519 key events
Topic direction
Why watch

The real competition is shifting from single-chip speed to rack, network, memory, power, packaging, protocol, software stack, and deployment economics.

Current read

The next infrastructure boundary will be decided by the coupled system: accelerator, NIC, switch, protocol, memory locality, runtime, power, storage, and operations.

Last updated · 2026-09-05 Key events · 19 Entities · NVIDIA / AMD / Broadcom / Cisco / Arista / Huawei / HPE / SpaceX
AIInfrastructureSupernodeRubinUnifiedBusCPONetworkingInferenceDatacenterMemoryPower
Key events
Timeline axis Latest Earlier

Scroll right for earlier events

Reverse event list

Product

Looped Transformers Hit the Frontline: Why OpenAI, Zhipu, and Academia Converged on the Sam…

Recurrent depth turns model depth from an architectural constant into a runtime variable. This piece traces the route's seven-year lineage and engineering boundaries: Huginn made the case, Nanbeige proved it i…

1 published post 1 published post The next infrastructure boundary will be decided by the coupled system: accelerator, NIC, switch, protocol, memory locality, runtime, power, storage, and operations. /thinking/looped-transformer-recurrent-depth/
Product

MetaRoCE in Depth: RDMA Designed for Loss, an Open Standard for Million-GPU Scale

MetaRoCE is Meta's clean-sheet RDMA transport protocol, opened through OCP. This analysis works through six questions: why it appeared, what it actually is, why Meta opened it, how the protocol works inside, w…

1 published post 1 published post The next infrastructure boundary will be decided by the coupled system: accelerator, NIC, switch, protocol, memory locality, runtime, power, storage, and operations. /thinking/metaroce-nic-intelligence-open-transport/
Market

The Token Factory Ledger: Inference Economics in SiliconFlow's Prospectus

Revenue of RMB 55.3M in 2025 (+653%), gross margin of -24%, net loss of RMB 345M, valuation of RMB 7.74B — SiliconFlow took its “Token Factory” story to the Hong Kong Stock Exchange. This article opens its led…

2 published posts 2 published posts The next infrastructure boundary will be decided by the coupled system: accelerator, NIC, switch, protocol, memory locality, runtime, power, storage, and operations. /thinking/token-factory-siliconflow/
Technology

Agents Pull the CPU Back to the Table: Four Questions from Hot Chips 2026

Six Hot Chips 2026 CPU talks distilled into four questions: the serial bottleneck (Vera's wide front end vs IBM's 5.7GHz), the memory power budget (SOCAMM2's 30-40W full load vs 100W-class RDIMM), three packag…

1 published post 1 published post The next infrastructure boundary will be decided by the coupled system: accelerator, NIC, switch, protocol, memory locality, runtime, power, storage, and operations. /thinking/hotchips-2026-agent-cpu/
Product

Five Networks for the AI Factory: Spectrum-X Multi-Plane and the 512K-GPU Scale

Spectrum-X multi-plane splits the AI factory network into five dedicated planes, scaling 64× to 512K GPUs via 8 planes by 4 rails. CPO reaches volume production: 4× fewer lasers, 5× lower power, 10× better MTB…

5 published posts 5 published posts The next infrastructure boundary will be decided by the coupled system: accelerator, NIC, switch, protocol, memory locality, runtime, power, storage, and operations. /thinking/hotchips-2026-ai-factory-network/
Technology

OpenAI Turns In Its Scorecard: An Analysis of the First Measured Results for Its In-House J…

At Hot Chips, OpenAI delivered Jalapeño's first measured results: 1.5-1.9× the per-watt throughput of Blackwell systems across three models, 1.7-3.6× lower latency. A reckoning against our June estimates: proc…

1 published post 1 published post The next infrastructure boundary will be decided by the coupled system: accelerator, NIC, switch, protocol, memory locality, runtime, power, storage, and operations. /thinking/jalapeno-first-results-audit/
Technology

Dissecting SemiAnalysis AgentX: The First Big Exam for Agentic-Era Inference

SemiAnalysis spent $3M capturing real Claude Code traces, open-sourced 393 sessions, and measured agentic inference across 1,000+ chips. This piece dissects AgentX 1.0 layer by layer: the dataset pipeline, cac…

2 published posts 2 published posts The next infrastructure boundary will be decided by the coupled system: accelerator, NIC, switch, protocol, memory locality, runtime, power, storage, and operations. /thinking/agentx-1-cuda-moat/
Research

Protocol Stacks at 1.6T Ports: The Kernel Stack's Last Mile and Four Ways Out

As ports head to 1.6Tbps, the protocol-stack CPU tax becomes a cost line the size of bandwidth itself. Seven papers, four routes: Presto's full TCP state machine in a switch pipeline (Best Student Paper, 16 co…

7 published posts 7 published posts The next infrastructure boundary will be decided by the coupled system: accelerator, NIC, switch, protocol, memory locality, runtime, power, storage, and operations. /thinking/sigcomm2026-terabit-host-stack/
Market

The Toll Booth, Bought Out: When Intelligence Flows Through Stripe's Pipes

Stripe acquires OpenRouter for $7.5 billion: a 5.8x valuation jump in 82 days. The routing layer proved middle-layer value can be captured, then the payments layer bought it whole. Closing chapter of the toll-…

1 published post 1 published post The next infrastructure boundary will be decided by the coupled system: accelerator, NIC, switch, protocol, memory locality, runtime, power, storage, and operations. /thinking/stripe-openrouter-token-toll/
Research

A Generational Upgrade Without a New Engine: GLM-5.3 and the Second Half of Post-Training

GLM-5.3 ships on the exact same 743B base as GLM-5.2, with every gain from post-training: Terminal-Bench up six-fold, coding near Fable 5, CyberGym past Mythos 5 as the best open-weight score. We break down th…

1 published post 1 published post The next infrastructure boundary will be decided by the coupled system: accelerator, NIC, switch, protocol, memory locality, runtime, power, storage, and operations. /thinking/glm53-post-training-scaling/
Product

Agent Storage Paradigm Reassessment: After FMS 2026, Projections Became Products

About a month later, the four-stage framework is validated and revised with FMS 2026 products. Stages 3 and 4 happen simultaneously. Bus dimension deep-dive: three new paths from CPU-dominated to GPU+DPU-domin…

2 published posts 2 published posts The next infrastructure boundary will be decided by the coupled system: accelerator, NIC, switch, protocol, memory locality, runtime, power, storage, and operations. /thinking/agent-storage-paradigm-reassessment/
Research

KV Cache Scheduling Engineering: After Compressing to 7%

V4 Flash CSA+HCA architecture-level hot-cold tiering, strategy benchmarks, end-to-end 8×H100 case study, delta encoding outlook. FMS 2026 hardware validation added.

3 published posts 3 published posts The next infrastructure boundary will be decided by the coupled system: accelerator, NIC, switch, protocol, memory locality, runtime, power, storage, and operations. /thinking/kv-cache-scheduling-7-percent/
Company

HPE turns AI factory into stack integration

HPE Discover ties AI factory, network, agent workload, and strategy into one systems bet.

4 linked posts The HPE series connects agent workloads, AI factory, networking, and strategy. That makes HPE a useful marker for whether enterprise AI infrastructure becomes an integrated stack rather than a GPU procurement project. Prediction: enterprise AI infrastructure vendors will compete on integration proof, not only benchmark numbers. /thinking/hpe-discover-2026-strategy/
Market

MaaS exposes inference economics

MaaS articles make the cost stack visible: KV cache, batching, routing, utilization, and token distribution.

2 linked posts The MaaS articles expose a layered cost structure: model serving, KV cache, batching, utilization, token distribution, and customer acquisition all interact. Prediction: inference winners will need cost observability and routing intelligence as much as model quality. /thinking/maas-inference-tech-stack/
Technology

KV Cache rewrites the memory hierarchy

Inference forces infrastructure to treat memory, SSD, and scheduling as one hierarchy.

3 linked posts Long-context and agentic workloads make cache size, memory bandwidth, storage tiering, and scheduler policy visible to product economics. Prediction: memory/storage vendors will become direct participants in inference architecture decisions. /thinking/kv-cache-deep-dive/
operation

Colossus shows compute is not battle power

SpaceX Colossus rental signals that raw GPU count does not equal usable AI capacity.

1 linked posts The Colossus signal shows that AI infrastructure value depends on turning capacity into usable service. Idle, mismatched, or hard-to-integrate compute becomes a cost problem. Prediction: utilization and workload fit will become first-class due diligence metrics for AI infra. /thinking/spacex-colossus-compute-reality/
Technology

Optics debate moves into AI fabric timing

CPO/NPO debate shows optical interconnect is now an investment and deployment timing question.

2 linked posts The CPO/NPO debate and Optical Shuffle supplement both point to the same question: when does copper stop scaling economically for AI fabrics? Prediction: optical adoption will be uneven and tied to cluster scale, packaging risk, and operational confidence. /thinking/cpo-vs-npo-semianalysis-optics-debate/
Technology

Topology and protocol become architecture

MRC, ZCube, RailFly, and Lingqu connect transport, topology, NICs, switches, and AI workload shape.

5 linked posts The MRC, ZCube, RailFly, Lingqu, and RNG articles collectively show the same shift: topology, transport, and endpoint intelligence are moving into AI architecture decisions. Prediction: the next cluster design cycle will evaluate NIC, switch, routing, and protocol as one architecture bundle. /thinking/mrc-protocol-when-nic-becomes-the-brain/
Technology

Power, packaging, and memory become the hidden bottleneck

800V power, 3D stacking, Rubin BOM, and CPU comeback point to non-GPU constraints.

5 linked posts The 800V, 3D stacking, Rubin BOM, CPU comeback, and supernode articles all push attention toward power, packaging, host processing, memory, and rack cost. Prediction: the next infrastructure advantage will come from solving cross-layer bottlenecks rather than only buying faster GPUs. /thinking/analog-chip-800v-datacenter-power-revolution/

Open questions

  • Q01 Trend Will infrastructure advantage shift from accelerators to coupled rack-level systems?
  • Q02 Possible path Which layer captures margin first: model, serving stack, memory, network, power, or cloud operations?
  • Q03 Constraint Do NICs, switches, and protocols become part of the accelerator architecture once memory locality dominates?
  • Q04 Validation Will enterprise AI factories be integrated products, or remain custom assembly projects?

Linked Thinking

Core essays 94

September 20262 posts
August 202625 posts
2026-08-31AI 基础设施 The Token Factory Ledger: Inference Economics in SiliconFlow's Prospectus Revenue of RMB 55.3M in 2025 (+653%), gross margin of -24%, net loss of… Read 2026-08-31AI 基础设施 2026Q2 NVIDIA Earnings: From Selling Chips to Selling AI Factories Revenue of $96.2B, guidance above $100B for the first time. The evidence… Read 2026-08-30AI 基础设施 Agents Pull the CPU Back to the Table: Four Questions from Hot Chips 2026 Six Hot Chips 2026 CPU talks distilled into four questions: the serial… Read 2026-08-27AI 基础设施 Five Networks for the AI Factory: Spectrum-X Multi-Plane and the 512K-GPU Scale Spectrum-X multi-plane splits the AI factory network into five dedicated… Read 2026-08-27AI 基础设施 Hyperscaler Silicon, Generation Two: Three Routes to the Training–Inference Split Three clouds all brought second-generation custom silicon: MTIA 400's dual… Read 2026-08-27AI 基础设施 Intel's Triple Launch: 18A Full In-House, 350W Air-Cooled Inference, UCIe Open Packaging Three Intel talks form a complete deployment-cost logic chain: Diamond… Read 2026-08-27AI 基础设施 Four Routes Against the Memory Wall: Hot Chips 2026 Delivers the Answer from the Silicon Side Micron sizes the wall with four bills: about 90% of package area to… Read 2026-08-27AI 基础设施 Hot Chips 2026 Overview: Per-Token Economics Is Forcing Decode Specialization Twenty-seven talks across two days, collected into one map and inventoried… Read 2026-08-26AI 基础设施 OpenAI Turns In Its Scorecard: An Analysis of the First Measured Results for Its In-House Jalapeño Chip At Hot Chips, OpenAI delivered Jalapeño's first measured results: 1.5-1.9×… Read 2026-08-25AI 基础设施 Dissecting SemiAnalysis AgentX: The First Big Exam for Agentic-Era Inference SemiAnalysis spent $3M capturing real Claude Code traces, open-sourced 393… Read 2026-08-25AI From Solo Bet to Consensus: Groq 3 LPX Full Production and Low-Latency Inference's Eight Months Eight months ago, folding an inference-chip company into its platform was… Read 2026-08-22SIGCOMM Protocol Stacks at 1.6T Ports: The Kernel Stack's Last Mile and Four Ways Out As ports head to 1.6Tbps, the protocol-stack CPU tax becomes a cost line… Read 2026-08-22SIGCOMM Collectives Become a Runtime: SIGCOMM 2026's Collective-Communication Turning Point Three production realities — MoE explosion, superpod heterogeneity,… Read 2026-08-22SIGCOMM Scale-Up Is the New Datacenter Network: SIGCOMM 2026 Re-runs 2015-2020 NVL72 and CloudMatrix384 production volumes turned the intra-rack… Read 2026-08-22SIGCOMM KV Cache Becomes a Network Citizen: How SIGCOMM 2026 Rewired LLM Inference as a Networking Problem After PD disaggregation, KV cache became explicit network payload. Six… Read 2026-08-22SIGCOMM SIGCOMM 2026 Close Reading: The Networking Community Moves Its Coordinate System onto AI Infrastructure A session-by-session count of 110 SIGCOMM 2026 papers: AI-workload papers… Read 2026-08-22AI The SubQ Three-Month Audit: The Technology Is Real, the Narrative Ran Ahead On May 5, Subquadratic released a model, SubQ, claiming a 12-million-token… Read 2026-08-11存储 Agent Storage Paradigm Reassessment: After FMS 2026, Projections Became Products About a month later, the four-stage framework is validated and revised… Read 2026-08-11存储 FMS 2026: The Storage Hierarchy Is Being Networked AI is transforming the storage hierarchy from discrete layers into a… Read 2026-08-08AI推理 KV Cache Scheduling Engineering: After Compressing to 7% V4 Flash CSA+HCA architecture-level hot-cold tiering, strategy benchmarks,… Read 2026-08-08AI The Efficiency Limit of Model Memory: From MLA to CSA+HCA First understand Tang et al. three-axis taxonomy, then test against… Read 2026-08-08AI推理 Inference Pricing Teardown: The Physical Cost of a Token Reverse-engineering inference pricing from V4 Flash/Pro dual-model… Read 2026-08-07AI The Model Is the Computer: Why AMD Bought Taalas Taalas etches model weights directly into mask ROM silicon, eliminating… Read 2026-08-06AI Two Paths: Kimi K3 vs DeepSeek V4 Architecture Divergence K3 and DSV4 pursue different extremes from MLA. Intelligence: occasional… Read 2026-08-06AI 基础设施 Physical Constraints in the AI Supply Chain: Four Bottlenecks and the Reshaping Landscape Nine major CSPs 2026 capex reaches 886.7 billion USD, 1.5x global chip… Read
July 202620 posts
2026-07-31AI 基础设施 NVLink's Moat: The Battle for Open Scale-Up Interconnect In a 72-GPU MoE model, every token-routing all-to-all step waits on the… Read 2026-07-31AI 基础设施 Routing and Congestion in 100K-GPU Clusters: When Physics Pushes Back A 100K-GPU cluster network isn't just a bigger network — it's a network… Read 2026-07-29AMD AMD Helios Supernode Teardown: Can Ethernet-Based Scale-Up Crack NVLink's Moat? AMD's Helios rack-scale AI platform is the first credible challenge to… Read 2026-07-28EDA Server Sizing for EDA Engineers: A Hardware Guide That 80% of Buyers Get Wrong Specing an EDA cluster with an AI-cluster mindset is the most common… Read 2026-07-23AI When Alphabet Burns $5.9B in a Quarter: The AI Infrastructure Capex Paradox Google Cloud revenue grew 82% YoY. Alphabet total revenue grew 24%. Then… Read 2026-07-22AI WAIC Revisited: Engineering Choices, Technology Delivery, and Enterprise Decisions in Supernode Design Whether an enterprise should purchase a supernode depends on four things:… Read 2026-07-21AI infrastructure When AI Agents Reinvent the File System: From "Everything Is a File" to "Everything Is Context" Agent workloads are rewriting the foundational assumptions of storage… Read 2026-07-18AI WAIC 2026 Field Notes: Year One of Supernodes, Training on Domestic Silicon, and the Third Path A live examination of China's AI industry after three years of… Read 2026-07-17存储 When Storage Becomes Agent Working Memory: Three Years of FMS Trend Migration and Technology Roadmap Agent storage paradigm analysis: KV Cache cross-tiering, Agent random IOPS… Read 2026-07-16AI Scale-Across: AI Clusters Are Growing Across Cities GPU clusters have already exceeded the power supply limits of a single… Read 2026-07-15AI Silicon Photonics' Three-Year Reshuffle: Component Bottlenecks, Route Divergence, and Supply-Demand Variables AI data center optical interconnect is transitioning from pluggable to… Read 2026-07-14semiconductor When the Best Earnings Meet the Worst Crash: A Structural Anatomy of the Semiconductor Selloff and Memory Market Forecasts SK Hynix plunged 15.37% in its worst single-day drop ever, erasing $1.3… Read 2026-07-12AI KV Cache as Infrastructure: When the Cache Layer Decouples from the Inference Engine Reasoning system design from workload physics. Six cluster challenges,… Read 2026-07-11AI GPU Is Becoming Oil: When Compute Turns from Scarce Resource to Tradable Commodity Ornn launched a GPU spot market, Nvidia lost $1T in market cap in two… Read 2026-07-09AI The Chinese Market Panorama Under Gartner's $2.6T AI Spending Framework Gartner May 2026: global AI spending $2.59T across eight layers. China… Read 2026-07-06AI The $2.60 Trillion AI Infrastructure Stack: A Layer-by-Layer Breakdown of Gartner's 8 Segments Gartner forecasts $2.60 trillion in global AI spending for 2026, with… Read 2026-07-05AI Where the 100x Comes From: Decomposing the Three-Layer Multiplication of AI Hardware-Software Co-Design SemiAnalysis founder Dylan Patel's 100x framework: AI efficiency gains… Read 2026-07-04ODCC The Infrastructure Generation Gap: AI Data Centers at ODCC 2026 In eight years, power density has increased 15-25x. Six technology… Read 2026-07-03semiconductor τ Scaling V2: From Theoretical Framework to Production Evidence Deep read of He Tingbo's τ scaling paper V2. 381 mass-produced chips,… Read 2026-07-01AI $3.1 Billion for Data Infrastructure, Not AI: Schneider’s Acquisition of Cognite and the Value Thesis for Vertical AI Schneider spent $3.1B not on an AI model but on data infrastructure. This… Read
June 202631 posts
2026-06-30AI The Glass Bridge: How Corning Uses Glass to Remove CPO's Last Production Hurdle 9μm fiber core vs 0.5μm PIC waveguide — CPO's biggest production… Read 2026-06-30AI When Tokens Stop Being Costs and Start Being Capital: The Economics of Tokenmaxxing 2.0 Empirical validation of compounding correctness is turning inference from… Read 2026-06-28AI Two Cracks in the Export Control Wall: Apple Courts CXMT, GLM-5.2 Rivals Mythos Apple lobbies to buy chips from blacklisted Chinese memory maker CXMT; WSJ… Read 2026-06-27AI Inside Jalapeño: What Happens When an AI Company Builds Its Own Heart # Inside Jalapeño: What Happens When an AI Company Builds Its Own Heart On… Read 2026-06-26LineShine LineShine Addendum: New Details Confirmed by The Next Platform's Deep Dive **Sources:** Read 2026-06-25AI4S From ISC 2026 to AI4S: HPC Is Shifting from a "Peak-FLOPS Machine" into a "Scientific Validation Engine" **Declaration:** This article is written based on publicly available… Read 2026-06-24HPC LineShine Tops TOP500: 2 EFLOPS Pure-CPU (6/29 Update) ISC 2026 confirms LineShine #1 in both TOP500 and HPCG. Top 500… Read 2026-06-21AI Inside the Inference Engine From Prefill to CUDA Kernel — a Millisecond-Level Breakdown of an… Read 2026-06-21AI Anatomy of an Inference Bill 85% of your AI bill is infrastructure tax; only 15% creates value Read 2026-06-21AI The Three Blind Spots of AI Observability When your monitoring says "all green" while your AI systems quietly burn… Read 2026-06-18HPE The $14 Billion Bet: HPE Discover 2026 Strategic Overview A decade of divestiture, then all-in on networking and AI factories. HPE's… Read 2026-06-18HPE Network as the AI Control Plane: HPE's Networking Gamble QFX six-tier coverage from training to inference, GreenLake Intelligence… Read 2026-06-18HPE Decoding the HPE AI Factory: Compute, Storage, Software, and the Cray Integration Experiment DL 394 Gen 12, Alletra MPX 10000 MCP-native storage, GreenLake full-stack… Read 2026-06-18HPE When AI Agents Become Workloads: HPE's Agent Infrastructure Blueprint Zero-code registration, three-tier identity, NVIDIA sandbox, MCP-native… Read 2026-06-16MaaS MaaS Inference Tech Stack: How Six Levers Cut Cost by 96% Technical dissection of DeepSeek's 96% per-token cost reduction. Six… Read 2026-06-16MaaS The Token Distribution Era: MaaS Service Models and Business Anatomy China's MaaS break-even line is 5-7 yuan/million tokens, with mainstream… Read 2026-06-15AI The Tyranny of Memory: How KV Cache Is Reshaping Every Layer of AI Inference In the 1M-context era, KV Cache is the defining bottleneck of inference… Read 2026-06-14AI When SSD Becomes Memory: How AI Inference Is Rewriting the Storage Hierarchy Three facts, sitting side by side in the first half of 2026, create a… Read 2026-06-13AI基础设施 Compute Power Is Not Fighting Power: What the SpaceX Colossus Lease Tells Us About AI Infrastructure Reality 220,000 GPUs built and leased out within a year. The SpaceX Colossus lease… Read 2026-06-12CPO One Report Wiped Out Optical Stocks: Is CPO Actually Dead? On June 9, 2026, SemiAnalysis sent a research note to institutional… Read 2026-06-12存储 RAMageddon: The Memory Famine and Storage Supercycle in AI Data Centers AI inference will rewrite the storage hierarchy. HBM margins exceed GPUs,… Read 2026-06-11AI Text Diffusion vs. Autoregressive: The Paradigm War DiffusionGemma at 1,107 tok/s, Mercury in commercial deployment, Dream 7B… Read 2026-06-08AI基础设施 Making K8s Understand Super-Nodes: openFuyao and the Lingqu Cloud-Layer Breakout The Lingqu cloud layer wraps hardware capabilities into K8s-native… Read 2026-06-08AI基础设施 Heart of the Super-Node: How the Lingqu Service Layer Weaves 8,192 Cards Together The Lingqu service layer answers how 8,192 cards cooperate — UBS Engine… Read 2026-06-08AI基础设施 Making Linux Understand Super-Nodes: Technical Anatomy of the Lingqu Kernel Layer Lingqu (UnifiedBus) super-nodes require systemic changes to the Linux… Read 2026-06-07AI Breaking the Transceiver Bottleneck: How Optical Shuffle Reshapes AI Cluster Economics Panduit engineer Castro proposes Optical Shuffle at IEEE 802.3, cutting… Read 2026-06-05AI A New Direction for Data Center Networking: What RNG Opens Up In April 2026, AWS switched all new non-GPU datacenters to a flat topology… Read 2026-06-04AI NVIDIA Rubin Respins: Is AMD GPU Competitiveness for Real? Fubon Research reveals Rubin was respun due to MI450 pressure. Full… Read 2026-06-02AI From CLOS to ZCube: Network Topology Evolution for AI Computing Clusters From Charles Clos's non-blocking telephone switching network in 1953 to… Read 2026-06-02AI MRC: When the NIC Becomes the Brain of the Network OpenAI, together with NVIDIA, AMD, Broadcom, Arista, and Cisco, overturned… Read 2026-06-01NVIDIA The PC, Reinvented: Computex 2026 and NVIDIA's Infrastructure Ambitions At GTC Taipei, Jensen Huang unveiled RTX Spark—33 years of technology… Read
May 202616 posts
2026-05-31AI Data Center The $27 Billion Hidden Thread: AI Data Center Power Distribution Revolution and the New Analog Chip Landscape As rack power consumption surges from 15kW to 1.5MW, power distribution… Read 2026-05-30AI Deep Analysis of the Lingqu Protocol The Interconnect Bet Behind China's AI Compute Breakout Read 2026-05-28AI The CPU Is Back: How Agentic AI Is Rewriting the Server Processor Landscape Three things happened almost simultaneously in May 2026. AMD's Venice… Read 2026-05-27semiconductor From CoWoS to Tau Scaling: 3D Stacked Chips, Technology Evolution, and the Route Fork On May 25, 2026, Huawei's He Tingbo unveiled "Tau Scaling" at ISCAS 2026,… Read 2026-05-27AI The $7.8M AI Rack: What Morgan Stanley's Rubin Teardown Reveals About Value Chain Restructuring Morgan Stanley's Howard Kao team published a comprehensive BOM (Bill of… Read 2026-05-26AI Two Revolutions, One Network: The Co-Evolution of AI Training Cluster Topology and Protocol AI training networks at 100K GPU scale are undergoing simultaneous… Read 2026-05-26AI Starting from ZCube: Network Topology Design for PD Disaggregated Inference As inference replaces training as the primary battlefield for AI… Read 2026-05-25thinking DeepSeek V4 + Ascend: Full-Stack Validation of Domestic AI Inference KADC 2026 Series Analysis · Part 4 · End-to-End Validation / Domestic AI… Read 2026-05-25thinking Ascend Supernode Architecture Leap: From Training-First to Agent-First KADC 2026 Series Analysis · Article 1 · AI Infra / Hardware Architecture… Read 2026-05-25AI RailFly: Network Topology Design for Prefill-Decode Disaggregated Inference Using ZCube as a baseline, we analyze the core value and limitations of… Read 2026-05-24AI PDC Disaggregated Serving for DeepSeek V4-Pro: From Compute Principles to Deployment Configs Targeting NVIDIA B300 and Huawei Ascend 950 Supernode, this article… Read 2026-05-23AI Ascend 950: Huawei's Third Path Calibrated with architect report: B300 36PFLOPS dense FP8, SLA-safe EP… Read 2026-05-21AI Broadcom Tomahawk 6 vs NVIDIA Networking Chips: A Full-Stack Benchmark from Silicon to AI Factory A full-stack comparison from switch chips to optical packaging, from… Read 2026-05-21AI NVIDIA Q1 FY2027 Earnings Deep Analysis: Technical Signals and Strategic Games Behind $81.6B Agentic AI demand has gone parabolic, but NVIDIA faces a triple squeeze:… Read 2026-05-20AI The Death of CPX and the Birth of LPU: A Paradigm Shift in AI Inference Architecture The Death of CPX and the Birth of LPU: A Paradigm Shift in AI Inference… Read 2026-05-20AI The Hook: Why Were NVL72's Copper Cables Sentenced to Death Within Two Years? This article is based on publicly available information as of May 19,… Read

Cross-topic links 10

August 20263 posts
July 20262 posts
June 20263 posts
May 20262 posts