locsic.com
Thinking
Long-form notes for reading the direction of technical change.
Thinking is a professional column: fewer quick posts, more durable essays. It keeps observations, source trails, assumptions, and evolving viewpoints visible over time.
Essays
30 essays shownJuly 2026
30 essaysNVLink's Moat: The Battle for Open Scale-Up Interconnect
In a 72-GPU MoE model, every token-routing all-to-all step waits on the slowest hop. Microsecond latency is catastrophic; sub-microsecond is the cure. As clusters scale from 8 to…
Routing and Congestion in 100K-GPU Clusters: When Physics Pushes Back
A 100K-GPU cluster network isn't just a bigger network — it's a network governed by different physics. This article systematically breaks down eight core challenges in routing…
AMD Helios Supernode Teardown: Can Ethernet-Based Scale-Up Crack NVLink's Moat?
AMD's Helios rack-scale AI platform is the first credible challenge to NVIDIA at the supernode scale. 72 MI455X GPUs, UALoE Ethernet scale-up, a 2 GW Anthropic deal, and a deep…
MCP Protocol 2026-07-28 Major Revision: From Tool-Calling Protocol to Infrastructure
On July 28, 2026, MCP published its largest specification revision since inception. Statelessness, MRTR, Tasks extension, and formal extension system collectively push MCP from a…
Server Sizing for EDA Engineers: A Hardware Guide That 80% of Buyers Get Wrong
Specing an EDA cluster with an AI-cluster mindset is the most common mistake in EDA infrastructure planning. This guide derives CPU, memory, storage, and network requirements…
The Open-Weight War: When the People Selling Walls Want to Close Open Source
Kimi K3 did not just ignite a model performance race—it split the AI industry.
When Agents Enter the Organization: How Enterprise Context OS Reconstructs Enterprise Information Infrastructure
Context is the third enterprise resource after compute and data. Enterprise Context OS = Context Store + Context Compiler + Agent Runtime.
When Agents Enter the Organization: How Context Becomes Enterprise Infrastructure's Next Frontier
As agent populations surge, context becomes the new bottleneck resource. Starting from the read-write cost inversion, the article derives the cognitive object model, Cognitive…
The Future of AI Native Organizations: A Topology Rewrite
Starting from the NBER paradox, this article defines AI Native organizations and dissects their operating model across seven dimensions: collaboration, communication,…
When Alphabet Burns $5.9B in a Quarter: The AI Infrastructure Capex Paradox
Google Cloud revenue grew 82% YoY. Alphabet total revenue grew 24%. Then free cash flow turned negative $5.9B. AI infrastructure capex is growing faster than Alphabet cash…
WAIC Revisited: Engineering Choices, Technology Delivery, and Enterprise Decisions in Supernode Design
Whether an enterprise should purchase a supernode depends on four things: whether your primary model is limited by communication bottlenecks, whether your data center can support…
When AI Agents Reinvent the File System: From "Everything Is a File" to "Everything Is Context"
Agent workloads are rewriting the foundational assumptions of storage architecture. From POSIX file systems to cognitive file systems, from KV Cache to Agent state management…
WAIC 2026 Field Notes: Year One of Supernodes, Training on Domestic Silicon, and the Third Path
A live examination of China's AI industry after three years of gear-shifting. Four deep dives: supernode economics, domestic chip training crossing 0-to-1, Oriental Chip's third…
When Storage Becomes Agent Working Memory: Three Years of FMS Trend Migration and Technology Roadmap
Agent storage paradigm analysis: KV Cache cross-tiering, Agent random IOPS pressure, four-stage evolution framework. FMS 2026 has validated core predictions; framework revised.
Scale-Across: AI Clusters Are Growing Across Cities
GPU clusters have already exceeded the power supply limits of a single site. The next step—not building bigger, but connecting farther.
Silicon Photonics' Three-Year Reshuffle: Component Bottlenecks, Route Divergence, and Supply-Demand Variables
AI data center optical interconnect is transitioning from pluggable to NPO/CPO. Who's at the bottleneck, where are the opportunities, and where are the supply gaps?
When the Best Earnings Meet the Worst Crash: A Structural Anatomy of the Semiconductor Selloff and Memory Market Forecasts
SK Hynix plunged 15.37% in its worst single-day drop ever, erasing $1.3 trillion from chip stocks. A five-dimensional structural anatomy of the semiconductor crash—ADR arbitrage,…
Decoding Anthropic's Loop Engineering Guide: Four Loop Types and Their Boundaries
Full analysis of Anthropic's official Loop Engineering guide. Four loop types, SKILL.md verification encoding, seven token levers, four code quality principles.
KV Cache as Infrastructure: When the Cache Layer Decouples from the Inference Engine
Reasoning system design from workload physics. Six cluster challenges, post-CXL hardware reasoning, AFD bandwidth economics, KV Memory Node architecture.
The Productization of Agent Toolchain: When Loop Engineering's Six Building Blocks Become a Market
From Claude Cowork to ChatGPT Work, from MCP to Agent Gateway, loop engineering's six building blocks are crystallizing into a five-layer product market.
GPU Is Becoming Oil: When Compute Turns from Scarce Resource to Tradable Commodity
Ornn launched a GPU spot market, Nvidia lost $1T in market cap in two months, and Micron tripled. Compute is turning from scarce resource to tradable commodity.
When the Loop Becomes the Unit of Engineering: The Paradigm Shift from Prompt to Context to Loop
Boris Cherny said he no longer writes prompts—he writes loops. As Anthropic and OpenAI converge on the same loop primitives, loop engineering is moving from concept to…
Inside Microsoft's ResearchStudio: Can AI Automate the First and Last Mile of Research?
An engineering manifesto on skill engineering, a deep teardown of Microsoft Research's AI research system, and an epistemological question about how expertise is transmitted.
The Chinese Market Panorama Under Gartner's $2.6T AI Spending Framework
Gartner May 2026: global AI spending $2.59T across eight layers. China holds 15-20% but with 70%+ in infrastructure vs 54% global average. A four-dimensional breakdown—compute,…
The $2.60 Trillion AI Infrastructure Stack: A Layer-by-Layer Breakdown of Gartner's 8 Segments
Gartner forecasts $2.60 trillion in global AI spending for 2026, with infrastructure capturing 55%. This article walks through every segment: definitions, scale, leading players,…
When the AI Bill Catches the Payroll: The Token Cost Paradox
Anthropic spends 4x payroll on compute. Uber burned its AI budget in four months. The cheaper tokens get, the more enterprises spend. The e-commerce disruption of retail is…
Where the 100x Comes From: Decomposing the Three-Layer Multiplication of AI Hardware-Software Co-Design
SemiAnalysis founder Dylan Patel's 100x framework: AI efficiency gains come from the multiplicative effect of co-designing model architecture, kernel optimization, and chip design.
The Infrastructure Generation Gap: AI Data Centers at ODCC 2026
In eight years, power density has increased 15-25x. Six technology directions—power, cooling, UEC, scale-up, in-network computing, token economics, NPO—coupled and evolving…
τ Scaling V2: From Theoretical Framework to Production Evidence
Deep read of He Tingbo's τ scaling paper V2. 381 mass-produced chips, Kirin 2026 LogicFolding measured data, three-layer τ reduction AI architecture (UB + Hi-ONE + 3D Folding)…
$3.1 Billion for Data Infrastructure, Not AI: Schneider’s Acquisition of Cognite and the Value Thesis for Vertical AI
Schneider spent $3.1B not on an AI model but on data infrastructure. This article examines what holds lasting value in vertical AI scenarios.