← Thinking Thinking

From Solo Bet to Consensus: Groq 3 LPX Full Production and Low-Latency Inference's Eight Months

Eight months ago, folding an inference-chip company into its platform was NVIDIA's solo bet; eight months later it became six companies' default…

2026-08-25Thinking35 min read

2026-08-25 · CPX/LPU series, part two · Preceded by "The Death of CPX and the Birth of LPU"

On December 24, 2025, NVIDIA announced it would pay roughly $20 billion in cash for access to Groq's technology licenses and core team. Jensen Huang clarified that same evening: "We are not acquiring Groq as a company." On August 24, 2026, at Stanford's Memorial Auditorium during Hot Chips 2026, NVIDIA announced that the Groq 3 LPX inference accelerator is in full production. From deal to production: exactly eight months.

On the same day, NVIDIA released three numbers: 3,400 output tokens per second, 30× higher throughput per megawatt, and 35× lower cost per token. Also on the same day, according to US market-close reports, NVIDIA's stock closed down for a seventh consecutive session — its longest losing streak since September 2022.

The good news of mass production and the market's doubt landed on the same calendar page. The coincidence doesn't need to be forced into a causal story. Low-latency inference happens to sit between the two: it is NVIDIA's bet on premium token economics, and it is also one of the test strips the market uses to measure returns on AI infrastructure.

Eight months ago, folding an inference-chip company into its platform was NVIDIA's solo bet. Eight months later, that bet has become the industry's default architecture. This article answers four questions: how hard are the numbers, what lies between production and availability, how did our May predictions fare, and why did six competitors converge on the same answer within half a year.

从 CPX 之死到 LPX 量产的八个月时间线
从 CPX 之死到 LPX 量产的八个月时间线

That journey, compressed: Rubin CPX unveiled in September 2025 as a long-context inference GPU; the $20B Groq transaction on December 24; CPX erased from the roadmap and LPX debuted at GTC in March 2026; Rubin pilot production and first deliveries in June and July; full production announced August 24.

One: The Quality of Three Numbers

The conclusion first: these three numbers are not on the same credibility level, and the gaps are wide.

3,400 tokens/second is a benchmark-environment record. The figure comes from Artificial Analysis benchmarking, running Gemma 4 31B (an open-source agentic model with 31 billion parameters) at a 100K-token context, and NVIDIA calls it the fastest performance ever recorded for that model. Worth noting: the benchmarked hardware came from an NVIDIA-supplied environment, with no independently procured third-party run yet. For reference: human reading speed is roughly 5–7 tokens per second; OpenAI's Ultrafast tier, in limited preview since August 13, delivers 750 output tokens per second on the full GPT-5.6 Sol model; and Cerebras's CS-4, announced August 18, hits 4,400 tokens per second on GPT-OSS-120B (120 billion parameters, about four times the size of Gemma 4 31B) — a number Artificial Analysis measured in a third-party environment. NVIDIA also claims LPX responds up to 4× faster than "the nearest alternative platform" for latency-sensitive workloads, without specifying the baseline. Put together, these numbers expose a problem before they settle anything: model sizes span roughly 15× from 8B to 120B, so directly comparing token speeds is comparing framings, not speeds.

30× and 35× are vendor-measured progress. Up to 30× higher throughput per megawatt and up to 35× lower cost per million tokens, versus the previous-generation GB300 NVL72, on SemiAnalysis's AgentX agentic-coding trajectories (with real context growth, tool calls, and sub-agent spawning preserved), running DeepSeek V4 Pro. Two footnotes sit in NVIDIA's own blog post: first, the tests were run by NVIDIA itself and are "currently pending SemiAnalysis review"; second, the results "don't yet reflect Vera CPU performance for tool calling." In other words: early self-measured data, review incomplete, part of the gain not yet counted in. There is also a quiet metric migration: at GTC in March, the 35× referred to projected throughput per megawatt; in August, the 35× refers to cost, while the throughput figure at the same metric is 30×. The same brand number changed its unit of measure. Replacing projections with measurements is progress, but the like-for-like figure shrank about 15% — a detour there is no need to take.

15× is the most solid number of the three. OpenRouter's data shows agentic tasks consume 15× more tokens than simple chat: in a research task, the agent queries databases, retrieves documents, invokes sub-agents, runs code and verifies results, accumulating context across hundreds of steps. This number speaks to demand-side volume growth; by itself it does not prove that low latency is a necessity. The latency argument is a separate chain: an agent's task completion time equals per-step latency multiplied by step count, and when steps number in the hundreds, per-step slowness is amplified linearly by the chain. That argument is qualitative — but so far no one has produced a quantitative rebuttal.

Ranked by credibility: 15× (industry data) > 30×/35× (vendor-measured, pending third-party review) > 3,400 t/s (benchmark environment, pending production verification). The ranking doesn't change the directional judgment, but it sets how much respect the word "production" deserves.

速度参照系:口径决定数字成色
速度参照系:口径决定数字成色

The chart above makes the point mechanically: every bar carries its qualifiers, and the taller the bar, the heavier the footnote.

Two: Between Production and Availability, Four Dependencies

A production announcement covers capacity alone. From production line to user-invokable API, Groq 3 LPX has to clear four dependencies.

The first is form. LPX is an extension of the Vera Rubin platform: one rack holds 32 compute trays, each with 8 LP30 chips — 256 chips working in concert over ultra-high-bandwidth interconnect, deployed alongside the NVL72 host rack. The NVL72 handles training, prefill (the input-understanding phase of inference), and context; LPX specializes in decode (the token-by-token generation phase). Public materials don't say whether this is a technical requirement of co-deployment or a commercial bundle — the two characterizations differ considerably, and that gap is real.

异构分工:一次推理在两种机架间的旅程
异构分工:一次推理在两种机架间的旅程

The second is supply chain. The LP30 is fabbed on Samsung 4nm — which is actually an advantage: NVIDIA's GPUs crowd TSMC's CoWoS advanced-packaging capacity, and the Samsung line sidesteps that jam. Foxconn exclusively builds the compute trays. The risk sits in the third link: every LPX rack needs 32 FPGAs for fabric expansion logic (link scheduling, load balancing, protocol conversion). We flagged this as the architecture's soft spot in May, and no public data on its actual overhead exists yet. The fourth link is on the GPU side: HBM4 supply for Rubin GPUs remains tight; if GPUs run short, the whole heterogeneous pipeline limps.

The third is customer rollout. Only one launch customer is named: Nebius, deploying to its Token Factory production inference platform, online by end of 2026. Nebius CTO Danila Shtan's statement carries the key design: "through the same API developers are already using, with no migration to a new stack." Low latency sold as a premium API tier, no application rewrite required. The smartest arrangement in the business model. But supply-chain reports say 6,000 LPX racks ship in 2026; the named customer list has exactly one entry, deployment scale undisclosed, and the remaining racks' destinations (presumably hyperscalers bundling with NVL72) are not public. The gap between 6,000 racks and one named customer is the biggest open question on the demand side.

The fourth is the most interesting: the echo of the deal structure. The footer of NVIDIA's press release reads: "Groq and LPU are used under license from Groq, Inc." The $20B transaction bought a perpetual, non-exclusive license to the patent portfolio plus the core team; Groq the company still operates independently — and the same press release says Groq plans to be among LPX's earliest adopters, right after Nebius. The licensee buying the licensor's mass-produced product. That detail describes the transaction's power structure more precisely than any strategic declaration. Incidentally, Senators Warren and Blumenthal have already sent a letter questioning whether the license-and-acquihire structure was designed to dodge merger review, with a response deadline of April 3, 2026 for Huang.

At this pace, 2026 is only the opening act. The real watershed is LP40 in 2027: that generation natively supports NVLink, and GPU–LPU interconnect upgrades from "racks cooperating" to "architectural integration." Most LPX customers are voting for that generation.

Three: The Ledger — What Our May Article Got Right

On May 20 we published "The Death of CPX and the Birth of LPU," the first analysis of Groq 3 LPX. Time to settle the ledger. Ledger discipline: state the judging event first, then check the old text — so that vendor plans we merely relayed don't get booked as our predictions.

"Genuine predictions, delivered." First, the timeline: we judged LPX would ship in Q3 (originally H2); the judging event is the production announcement, which landed August 24 — inside Q3, about five weeks ahead of quarter-end. Second, positioning: we argued the LPU would serve as decode's second, leaving the GPU in place; NVIDIA senior director Dion Harris verified it nearly word for word: "This is not about replacing the GPU, but about matching the right processor to each part of the workload." Third, and most valuable: the disaggregation had to happen. In May it was one company's roadmap; eight months later it is industry consensus.

"Relayed plans, no credit." The 6,000 racks, Foxconn exclusivity, Samsung foundry — all were vendor plans relayed from supply-chain reports in May; the August 24 announcement added no new information on those three. Booking them as "prediction hits" would be self-deception.

"Watch items, still open." The FPGA scheduling overhead, the failover logic of GPU–LPU heterogeneous systems, and the real pricing of low-latency tokens: three questions raised in May, still without public answers today. Pricing above all: Harris says clouds can offer premium tiers for latency-sensitive customers, but the tier price itself remains blank. Until that number appears, the commercial loop of the low-latency business cannot be validated.

"What we didn't foresee." The speed of convergence among competitors: Cerebras, AMD, and OpenAI all moved to the same disaggregation pattern within half a year (next section). And SpaceXAI announced its first-generation Starmind agent satellite will be based on an "optimized" Vera Rubin NVL72 — the official wording goes only as far as "planned"; the prototype timeline (2027, per social-media reports) is unconfirmed officially. We record it as a plan, without embellishment.

The biggest lesson of this ledger: architectural judgments outlast timetable judgments. Timetables slip (HBM4 qualification and liquid-cooling engineering both caused slips); structural judgments, once sound, are only a matter of time. When writing predictions, bet on structure, not on calendars.

Four: Six Players Converge — One Answer, Two Bets

Zoom out. The most important event in low-latency inference this half-year is that everyone arrived at the same answer: abandon single-chunk universal compute, split prefill from decode, give each the right hardware.

Player Route Headline number Status
NVIDIA Groq 3 LPX LPU chip array + on-chip SRAM (Samsung 4nm) 3,400 t/s @31B (benchmark env) Full production; Nebius first, online late 2026
Cerebras CS-4 Wafer-scale 3×WSE-3T concentrated SRAM 4,400 t/s @120B (AA third-party) Q3 2026 ramp; AWS/AMD prefill offload
AMD × Taalas Weights hardwired into chip (mask ROM) 16,960 t/s @8B (vendor-measured) Acquisition announced Aug 6; integrating with Instinct
OpenAI Ultrafast Buys Cerebras capacity, resells 750 t/s @GPT-5.6 Sol Limited preview since Aug 13; first customers incl. Jane Street, Rogo
Huawei Ascend 950DT High-bandwidth HBM route (HiZQ 2.0, 4TB/s) Awaiting tests On Huawei Cloud in August
Yuanchuanwei LPU+ China's direct LPU counterpart Tape-out H1 2027

The three on-chip storage routes differ in how they trade flexibility, capacity, and speed. Groq builds a distributed SRAM pool from 256 small chips — flexibility. Cerebras concentrates SRAM on a dinner-plate-sized wafer — capacity and razor-thin on-die latency. Taalas is the most radical: trained weights hardwired into mask ROM, one transistor storing one 4-bit weight and performing its multiply, a ~200W chip delivering nearly 17,000 tokens/second — at the cost of a chip that serves exactly one model and needs a respin whenever the model updates. AMD's logic in buying it is clear: Instinct GPUs for prefill, Taalas for decode — the same division of labor NVIDIA built.

But the division of labor converged; the memory media did not. Huawei's Ascend 950DT does decode on in-house HBM (HiZQ 2.0, 4TB/s), skipping SRAM's speed extremes in exchange for ecosystem compatibility and continuity of the existing software stack. This is a serious alternative: the SRAM route bets that "decode's latency sensitivity is worth everything"; the HBM route bets that "decode's capacity needs will eventually outweigh its latency sensitivity." The longer the context and the larger the model, the more persuasive the latter becomes. Which bet is right gets settled in 2027.

六家收敛格局:分工收敛,介质分两路
六家收敛格局:分工收敛,介质分两路

OpenAI's position is unlike the others. It doesn't build chips; it joins the pattern with buying power: Cerebras's largest customer (a multi-year purchase agreement signed in April) and also its patron (~$1 billion in data-center development funding). Customer as patron, with an equity position that behaves like futures: Cerebras's S-1 discloses that OpenAI holds a near-zero-strike warrant over ~33.4 million Class N shares (total exercise cost of roughly $334), vesting in tranches as OpenAI buys capacity, up to ~10% of the company if fully earned, alongside a $1 billion loan at 6% interest. Nearly free to exercise, provided the purchases get made. One calendar coincidence: the master agreement and the warrant are both dated December 24, 2025 — the same day as the NVIDIA-Groq transaction. Both of low-latency inference's ecosystem routes took shape on the same Christmas Eve. The Ultrafast tier has been in limited preview in the OpenAI API since August 13, and early customers like Jane Street sketch the real buyer profile for low latency: in finance, fast and slow convert directly into money.

The hardest public demand-side evidence comes from the only one of the six that is a public company whose main business is low-latency inference. Cerebras went public in May; its August 12 Q2 report showed total revenue of $180 million (+74% year over year, slightly below expectations) and remaining performance obligations (RPO) of $25.4 billion — the bulk of it OpenAI's multi-year contract, and, as the company said on the earnings call, that figure does not yet include potential backlog from AWS or other hyperscalers. Orders are a harder leading indicator than marketing; order concentration is itself the risk. The market's anxiety is concentrated on the returns side.

Finally, NVIDIA's moat. On public benchmark numbers alone, LPX did not take the absolute speed record (3,400 vs 4,400 — different framings, for reference only). LPX's weight is elsewhere: it is one part of the Vera Rubin platform's co-design (NVIDIA's framing: seven chips, five purpose-built racks). What customers buy is a workshop of the AI factory. Cerebras and peers are selling faster engines; NVIDIA is selling the whole car. That difference sets the level of competition.

Five: 2027 Is the Verification Year

Lay the known milestones on a timeline: Nebius Token Factory online by end of 2026 — the first production-environment data point; Cerebras CS-4 widening supply in Q3 2026; LP40 in H1 2027 with native NVLink integration; and farther out, the 2028 Feynman architecture stacking LPU units into the GPU main die — the endgame we pointed to in May.

Three watchpoints belong on the calendar. First, third-party measurements after Nebius goes live and the actual pricing of low-latency tokens — together they decide whether this is an engineering marvel or a business loop. Second, the customer list as CS-4 supply widens, and the pace at which $25.4 billion in RPO converts. Third, LP40's official specs, and how deep the NVLink integration goes.

As for that losing streak: seven down days reflect macro sentiment on AI capex; JPMorgan's warning on off-balance-sheet credit and the 15% server price increases starting 2027 point at the same problem — enormous investment, unproven returns. Low-latency inference happens to be the first test on the returns side: if latency-sensitive customers will pay a premium for fast tokens, AI infrastructure gains a new layer of revenue quality; if not, the fastest chips are just more expensive inventory.

Eight months ago, low-latency inference was one company's bet. Now it is six companies' roadmaps, a $25.4 billion order book, and a delivery curve still being drawn. Turning a bet into consensus took eight months; turning consensus into business will take longer.

Further reading: The Death of CPX and the Birth of LPU: AI Inference's Paradigm Shift (part one of this series) · NVIDIA Rubin Respin and AMD MI450 · The next piece, "Four Assaults on the Memory Wall," returns to the root cause of all this: bandwidth, capacity, and cost cannot all be optimal in one memory medium.