← All shows

State_of_AI_with_Nathan_Benaich_State_of_AI_Compute_Index_June_2026

Published Jun 29, 2026 · Duration 13:09 · Language en · 8 highlights

Summary

本期节目介绍了Air Street Press与Zeta Alpha合作发布的《AI算力指数》2026年6月第五次更新(V6),核心结论是开放研究文献中的AI算力使用在2025年短暂停顿后于2026年强势反弹,预计达到49,339次芯片引用,同比增长10.7%,略高于2024年峰值。NVIDIA依然是绝对主导,占约91%的引用份额,挑战者如AMD、华为昇腾和苹果虽有增长但规模仍小,其中苹果反超AMD成为最受引用的非NVIDIA、非谷歌加速器,主要反映本地推理而非前沿训练。NVIDIA内部已变成产品周期更替的故事:A100趋于停滞、Hopper(H100/H200)成为活跃的装机基座、Blackwell仍主要停留在管线阶段。本次更新最大的变化是指数不再只追踪论文引用,还加入了集群基础设施视角,显示算力已从研究采购转向工业化部署,xAI Colossus等大型私有集群的规模远超国家级超算。Grace Blackwell管线中已部署的不足4%,约81%仍处于“已宣布”状态,约束已从GPU数量转向电力、土地、冷却、互联、债务和承购等环节。需求侧首次以“吉瓦”衡量,前沿实验室不再绑定单一芯片供应商,而是组建涵盖NVIDIA、TPU、Trainium、定制芯片和主权云的算力组合。总体而言,开放文献已复苏,NVIDIA仍是默认选项,挑战者故事真实但仍弱小,而下一波算力浪潮在纸面上远大于已落地的规模。

Highlights

  1. Last year's update asked whether 2025 was the first real slowdown in open AI compute citations after six years of growth. With final normalized 2025 counts now in hand, and 2026 projected from counts through June 1st, the answer looks clearer. 2025 was a pause, not a rollover.

    去年的更新提出了一个问题:在连续六年增长之后,2025年是否是开放AI算力引用首次真正的放缓。如今有了最终归一化的2025年数据,以及基于截至6月1日数据对2026年的预测,答案变得更清晰:2025年只是一次暂停,而非全面回落。

    Resolves the central question of last year's report
  2. A lot of companies, labs, and governments are trying to challenge Nvidia. In the open research literature, they're still mostly not succeeding. There is movement in AMD, Huawei, and Apple, but NVIDIA remains the language researchers use when they describe their compute.

    众多公司、实验室和政府都在试图挑战NVIDIA。但在开放研究文献中,它们大多仍未成功。AMD、华为和苹果有所动作,但NVIDIA依然是研究者描述其算力时所使用的“语言”。

    Strong framing of NVIDIA's entrenched dominance
  3. Apple moved from 741 to 998, overtaking AMD to become the most cited non-NVIDIA, non-Google accelerator in the open literature. That likely reflects the spread of local inference and developer workflows rather than frontier training.

    苹果从741上升到998,超越AMD成为开放文献中引用最多的非NVIDIA、非谷歌加速器。这很可能反映的是本地推理和开发者工作流的普及,而非前沿训练。

    Surprising rise of Apple as top challenger
  4. Groke is now the most cited startup chip, up 49% to 264, and in December, NVIDIA acquired it. The fastest way to dent NVIDIA's share turned out to be getting bought by NVIDIA.

    Groq现在是引用最多的初创芯片,上升49%达到264,而在12月,NVIDIA收购了它。结果证明,撼动NVIDIA份额最快的方式,竟是被NVIDIA收购。

    Ironic and memorable insight about startup chips
  5. The biggest change in V6 is that the compute index is no longer just a paper citation tracker. The cluster charts now show how much of the AI buildout has already moved from research procurement into industrial infrastructure.

    V6最大的变化在于,算力指数不再只是一个论文引用追踪器。集群图表现在展示了有多少AI建设已经从研究采购转向了工业化基础设施。

    Marks a fundamental shift in the report's methodology
  6. XAI Colossus 1 is the largest tracked hopper deployment at 200,000 GPUs and by itself is almost five times the entire tracked A100 set.

    xAI Colossus 1是追踪到的最大Hopper部署,拥有20万块GPU,仅它一个就几乎是整个追踪A100总量的五倍。

    Striking scale comparison of a single private cluster
  7. The index tracks 78,136 deployed GB200 slash GB300 GPUs, 322,440 installing and 1.66 million announced. In other words, less than 4% of the tracked Grace Blackwell pipeline is deployed. Roughly 81% is still announced.

    该指数追踪到78,136块已部署的GB200/GB300 GPU,322,440块正在安装,166万块已宣布。换言之,追踪到的Grace Blackwell管线中已部署的不足4%,约81%仍处于“已宣布”阶段。

    Reveals how much of the next wave is still vaporware
  8. This is why GPU count is becoming an incomplete question. The constraint has moved outward to power, land, cooling, interconnect, permitting, debt, and offtake. A GPU order is not a cluster. A cluster is not always available capacity.

    这正是为什么GPU数量正在成为一个不完整的问题。约束已经向外转移到电力、土地、冷却、互联、审批、债务和承购。一笔GPU订单不等于一个集群,一个集群也不总是可用的算力。

    Reframes the real bottleneck of the AI buildout
Full transcript

You're listening to the State of AI Compute Index June 2026 on Air Street Press. Today we released the fifth update of the State of AI Report Compute Index in collaboration with Zeta Alpha. You'll now find updated counts as of June 2026 for AI research papers using NVIDIA, TPUs, Apple, Huawei, AMD, ASICs, FPGAs, and AI semi-startups. We've also expanded the infrastructure side of the index. A 100...

Hopper, standalone Blackwell, Grace Blackwell, and a new demand side view of Frontier Lab contracted compute in gigawatts. Each chart can now be downloaded, shared, and embedded. A few notes up front. The 2026 citation figures used real counts through June 1st, 2026, plus a volume. Adjusted projection for the rest of the year. Year over year deltas are calculated against final normalized 2025 counts.

not last year's mid-year 2025 projection. GPU count charts show NVIDIA data center GPUs by owner slash operator, split into deployed, installing, and announced. Grace Hopper and Grace Blackwell parts are counted by GPU dies, so one NVL 72 rack equals 72 GPUs. Tenants are not double counted against the operators whose clusters they use. The breather was short.

Last year's update asked whether 2025 was the first real slowdown in open AI compute citations after six years of growth. With final normalized 2025 counts now in hand, and 2026 projected from counts through June 1st, the answer looks clearer. 2025 was a pause, not a rollover. Across the tracked accelerator categories, the 2026 projection reaches 49,339 chip citation counts.

up 10.7% year over year, and just above the 2024 peak. NVIDIA remains the default. Its chips appear in 44,715 of those counts, up 10.9% year over year, and about 91% of the track total. That does not mean frontier labs have suddenly become more transparent. The largest model developers still publish less of their best work, and many papers built on managed APIs or shared cloud services do not specify the underlying silicon at all. But the open literature has not stopped reflecting hardware diffusion. The 2025 dip looks more like publication cycle timing, API abstraction, and a quiet year between hardware waves than a collapse in compute usage. The more interesting finding is the relative lack of upheaval. A lot of companies, labs, and governments are trying to challenge Nvidia.

In the open research literature, they're still mostly not succeeding. There is movement in AMD, Huawei, and Apple, but NVIDIA remains the language researchers use when they describe their compute. Indeed, AMD citations nearly doubled from 251 in 2025 to 472 in 2026. Huawei Ascend 910 rose 56% from 137 to 213.

Apple moved from 741 to 998, overtaking AMD to become the most cited non-NVIDIA, non-Google accelerator in the open literature. ConsumerMax now outside the leading Silicon Challenger. That likely reflects the spread of local inference and developer workflows rather than frontier training. TPS, meanwhile, declined 7% in the open paper data, despite Google's obvious importance to frontier AI compute.

That tension is a useful reminder that paper citations are a leading indicator for some forms of adoption, but a poor measure of private, API-mediated, or closed lab usage. NVIDIA is still king, but the kingdom is changing shape. The NVIDIA chart is now a product cycle chart. A 100 citations are basically flat at 15,327 of 0.4% year over year. The A100 is no longer where the slope is.

200 citations have more than doubled to 9,823. Up 111% year-over-year as the 2024 and 2025 Hopper build-out finally works its way into papers. Blackwell Family mentions tractors B100 slash B200 slash B300 in the Zeta Alpha data are still small at 685, but are up roughly 3.4 times plus 239% year-over-year. The older stack continues to drain out.

V100 citations are down 34%, RTX 3090 down 23%, P100 down 36%, and K80 is now barely visible. The RTX 4090 is still useful in the academic longtail, up 9% to 6,557 citations, while the new 50 series cards are appearing quickly from a small base. Startup Silicon is no longer one story.

The individual startup chip chart is more fragmented than the headline category suggests. Groke is now the most cited startup chip, up 49% to 264, and in December, NVIDIA acquired it. The fastest way to dent NVIDIA's share turned out to be getting bought by NVIDIA. Cerebras is second at 242, up 8%. Sambanova rose to 52, Cambricon to 33, while Graphcore fell to 40, and Habana to 24.

That is not yet a market share story. Startup chip citations remain tiny next to Nvidia, and papers using startup silicon still often include authors from the chip company itself. But it does show the category splitting into different jobs. Groke is showing up around low latency inference, cerebras around large scale, wafer scale systems, and the older acquired or de-emphasized platforms are fading from view. The right conclusion is not that startup chips are breaking CUDA, they aren't.

It's that the category is specializing into niches, inference latency, wafer scale, while NVIDIA keeps the general case. Hopper is the installed base. The biggest change in V6 is that the compute index is no longer just a paper citation tracker. The cluster charts now show how much of the AI buildout has already moved from research procurement into industrial infrastructure. Across the tracked hopper systems, we count 460,904 deployed H100, H200, and GH200 GPUs plus 1,328 installing. That is more than 11 times the tracked A100 total of 41,208. The A100 chart now reads like a legacy fleet chart. Hopper is the live installed base. XAI Colossus 1 is the largest tracked hopper deployment at 200,000 GPUs and by itself is almost five times the entire tracked A100 set.

Tesla Cortex follows at roughly 66,000 H100 equivalent GPUs, then Meta's Gen AI clusters at 49,152, a CoreWeave H200 cluster at 42,000, Voltage Park at 24,000, and Germany's Jupiter booster at 23,536 GH200. There's a geopolitical point hiding in the table. National HPC systems are exact.

and visible, which makes them easy to count. Private fleets are estimated, harder to observe, and much larger. Europe now has serious machines in Jupiter, Alps, Isambard AI, Leonardo, Mernostrum 5, and Jean Zay. But the largest private AI clusters are operating at a scale national supercomputing programs mostly do not match. Blackwell is mostly pipeline. Standalone B200 slash B300 deployments in the index Total 16,024 GPUs, led by Deutsche Telekom's Munich Industrial AI Cloud at 10,000, SoftBank's DGX Superpod at 4,000, E2E Networks in India at 1,024, and SK Telekom's high-end cluster at around 1,000. IRN's 50,000 B300 order sets in announced. Grace Blackwell is where the real pipeline sits. The index tracks 78,136 deployed GB200 slash GB300 GPUs,

322,440 installing and 1.66 million announced. In other words, less than 4% of the tracked Grace Blackwell pipeline is deployed. Roughly 81% is still announced. The largest announced and installing programs now look more like sovereign or hyperscale industrial projects than normal data center procurement. Human in Saudi Arabia at up to 600,000 GB 300, Stargate Abilene at a 450,000 GPU target, South Korea's national 260,000 Blackwell program, and scales 200,000 GB 300 commitment to Microsoft, XAI Colossus 2 at roughly 110,000 installing, Argonne Solstice and Equinox at 110,000 combined, and Stargate Norway at 100,000 announced. This is why GPU count is becoming an incomplete question. The constraint has moved outward to power, land, cooling, interconnect, permitting, debt, and offtake.

A GPU order is not a cluster. A cluster is not always available capacity. And contracted capacity is not the same as a model training run. The demand side is now measured in gigawatts. The new frontier lab contracted compute chart looks at the other side of the market, not who owns a cluster, but which labs have contracted capacity in disclosed gigawatt terms. For gigawatt disclose deals, open AI is at 10 gigawatts of NVIDIA systems.

Anthropic is its six gigawatts total, split between one gigawatt of NVIDIA and five gigawatts of non-NVIDIA capacity. That non-NVIDIA figure is understated because several major commitments, including open AIs, AMD and Broadcom deals and Anthropics AWS training commitment, are not disclosed in gigawatts and are therefore not charted. Still, the direction is clear. Frontier labs are no longer choosing a single chip vendor. They're assembling compute portfolios.

NVIDIA for the broadest software ecosystem and fastest path to scale, TPUs, and Tranium for strategic supply and cost control, custom silicon for future leverage, and Neoclouds or sovereign projects when hyperscaler capacity is not enough. CUDA is still king, remains true in the papers. But in frontier lab procurement, the question is becoming broader. Who can turn contracted gigawatts into reliable, liquid-cooled, networked, usable intelligence infrastructure?

Looking ahead. Several known unknowns will shape the next Compute Index update. How quickly announced Grace Blackwell projects become deployed clusters rather than press releases. Whether OpenAI's Stargate program, Humane, South Korea's National Blackwell program, and NScale's Microsoft commitments hit their published timelines. Whether AMDMI300 slash MI350, Huawei Ascend, TP Use, and Tranium show up more visibly in open paper metadata.

or remain hidden behind closed lab and cloud abstraction layers. Whether Blackwell appears in the literature in late 2026, the way Hopper appears in the 2026 data now. Whether national AI factories can close any of the gap with private frontier lab infrastructure. The headline from v6 is simple. The open literature has rebounded. Nvidia remains the default. The Challenger story is real, but still small. Hopper is now the installed base.

And the next compute wave is much bigger on paper than it is in the ground. See the live charts here, www.stateof.ai slash compute. A few notes. We take the view that usage of chips in AI research papers by early adopters is a leading indicator of broader industry usage, but not a complete measure of closed lab or API-mediated compute. 2026 citation figures are real counts through June 1st, 2026 plus volume-adjusted full-year forecasts from Zeta Alpha's open-source AI paper index. Year-over-year comparisons use final normalized 2025 counts as the baseline. GPU count charts use public disclosures, operator materials, EuroHPC, top 500, semi-analysis, the next platform, data center dynamics, company filings, and NVIDIA materials. Large private company figures are best available estimates.

National HPC figures are exact where published. Announced capacity is labeled separately from deployed and installing capacity. We do not invent cloud instance counts where vendors do not disclose them.

Delete this episode?

This removes the episode page and its saved audio from this library.