Dwarkesh Podcast - Why smarter AI models could drive up compute prices 10x
Summary
本期讨论了未来几年顶尖AI实验室可能面临的算力格局:以Anthropic为例,收入连续多年增长约十倍,而行业算力供给每年仅增长约三倍。作者认为,要维持这种差距,实验室必须提高利润率、推高算力价格,或把更多算力从训练转向推理,但实验室并不愿因推理挤占训练资源而放慢前沿模型研发。随着模型能力增强,同样的算力能够创造更高收入,作者估算若一块H100能承载人类水平的软件工程师,其年价值可能超过25万美元,远高于当前租赁价格。算力因而可能成为更昂贵的稀缺资源,最强实验室凭借更高的变现效率获得竞价优势,而更节省算力的模型也能收取更高溢价。与此同时,低价值的AI应用可能被自动化科研等高价值用途挤出,因为后者更愿意为有限的token和硬件支付高价。作者也审视了这种稀缺论是否会重蹈马尔萨斯式预测的覆辙,但指出摩尔定律、晶圆厂建设、ASML光刻机供应和先进制程晶圆分配都难以显著加速,甚至维持每年三倍增长都不容易。长期看机器人制造可能再次让算力变便宜,但在当前“奇点前”阶段,模型业务强大的规模经济很可能进一步集中智能产业的财富与权力。
Highlights
-
For the last three consecutive years, Anthropic's revenue has 10x'd year over year. And it's likely to do so again this year, so they ended last year with 9 billion in revenue. I think they'll probably end this year with somewhere between 100 billion to 150 billion dollars in rev ...
过去连续三年,Anthropic的收入每年都增长了十倍,而且今年很可能再次如此;他们去年以90亿美元收入收官。我认为他们今年年底的收入可能达到1000亿至1500亿美元。
Dwarkesh Patel A startling growth premise -
If a true human-level software engineer could run on an H100 equivalent, then at today's prices for software engineers, that H100 should rent for over 250k a year. That's over 15x the current spot price for an H100, and this is not even accounting for the fact that your AI can wo ...
如果一名真正达到人类水平的软件工程师可以运行在一块H100等效算力上,那么按今天软件工程师的价格计算,这块H100每年的租金应超过25万美元。这是H100当前现货价格的15倍以上,而且还没有考虑AI能够在夜间和周末持续工作。
Dwarkesh Patel A vivid valuation of intelligent compute -
If it costs $20 an hour to rent an H100, then it would be extremely stupid to use a weaker, less efficient model because it's going to burn more tokens on your expensive compute to get the exact same result. Basically, if you have a model that can get the same result by using les ...
如果租用一块H100每小时要20美元,那么使用更弱、更低效的模型将极不明智,因为它会在昂贵的算力上消耗更多token,却只得到完全相同的结果。基本上,如果一个模型能以更少算力取得同样结果,那么从某种意义上说,你就创造了更多算力。
Dwarkesh Patel Efficiency becomes synthetic supply -
I don't see how any of the three elements that constitute that 3x can be much accelerated. So 1.4x of that is coming from Moore's law. Far from increasing it, I think it'll be a miracle if we can just keep it going for a few more years. 1.2x is coming from building new fabs.
我看不出构成三倍增长的三个要素中,有哪一个能够大幅加速。其中约1.4倍来自摩尔定律;别说继续提高,我认为如果还能维持几年就已经是奇迹。另约1.2倍来自建设新的晶圆厂。
Dwarkesh Patel Physical bottlenecks challenge exponential demand -
The fact that Anthropic revenue has been 10x'ing year over year, whereas their compute has only been 3x'ing year over year, I think illustrates how strong the economies of scale are in the model business. When you train a model, you just had to spend this one-time cost to learn a ...
Anthropic的收入每年增长十倍,而算力每年只增长三倍,我认为这说明模型业务的规模经济有多么强大。训练模型时,你只需支付一次性成本来学习各种技能,随后这些技能就能由所有用户共享。
Dwarkesh Patel The central insight on AI scale economics
Full transcript
Dwarkesh PatelToday I want to talk about what the compute situation for the labs will look like over the next few years. For the last three consecutive years, Anthropics revenue has 10x to year over year. And it's likely to do so again this year, so they ended last year with 9 billion in revenue. I think they'll probably end this year with somewhere between 100 billion to 150 billion dollars in revenue. Now for this trend to continue, Anthropics would need to make 1 trillion dollars in revenue by the end of next year.
Dwarkesh PatelOf course, there's no deep reason why this has to be true. It's a very wild conclusion, and it's ultimately a question of AI capabilities. Does AI get that useful by the end of next year? But suppose the trend does continue. Well, I want to think through what happens in that world. Now, the other big trend in AI is that lab compute only three Xs year over year. For a lab to keep 10 Xs revenue year over year while compute only three Xs, one of the following three things needs to happen or some combination of the three needs to happen.
Dwarkesh PatelOne, lab margins have to increase. Two, the price of compute has to increase. Or three, the percentage of compute that labs spend on inference rather than training has to increase. My understanding is that basically all three of these things are already happening. With regards to the margins, Anthropics inference margins reportedly went from 40% to the middle of last year to upwards of 80% now available.
Dwarkesh PatelWith regards to compute, the spot prices for compute are more than 40% higher than they were in the February trough that we had earlier this year. And with regards to the share of compute that goes to trading versus inference, in 2024, according to Epoch, OpenAI was spending just a quarter of its compute on inference and that number is likely closer to 50% if not higher now. Now, Lazer preferred not to do this final thing of increasing the share of compute they spent on inference. The way the labs see the world, the whole point of inference revenue, is to help convince investors to give you more money in order to train the next bigger, better model. And if you're spending most of your compute on inference, you're basically declaring that AI progress has stalled, and you're just now in the business of being a cloud provider. Now, this is a less compelling business than building AGI. And so the labs do not want to be in this business, nor do they think they are in this world. They think that within a year, they'll have built models that make the current ones look extremely shitty. But they need to invest.
Dwarkesh Patela lot of their compute, the majority of their compute, into doing the training and experiments that are necessary to build the next model. So that leaves only two options for how you can get out of this gap between the fact that lab compute only increases 3x year over year, but revenue increases 10x. Either the lab's margins have to increase so that they get this surplus, or the price of compute has to increase so that everybody in the stack below the lab gets a surplus.
Dwarkesh PatelIt's not clear to me which world we end up in. Do we end up in a world where we go from 80% margins for some of the top models to greater than 90% margins if the lab margin effect dominates? Well that would require the leading model to be so far ahead of the competition because the nature of margins, why they exist in a market economy is that the thing you are serving is so much better than what somebody else could go get and replace you on the market. But it's just really wild for me to consider that the margins for something like intelligence will be greater than 90% and they don't get competitive at that level. So that leaves only one other possibility of this escape valve between these two trends, which is that the price of compute has to increase. As I mentioned, this is already starting to happen.
Dwarkesh PatelAnd the effect is even stronger when you look at the tranche of compute that the frontier labs actually need to accumulate. Because they can't just go out and buy a spot instance. They need to make sure that they get enough scale to get really good efficiency and flexibility. And also that they have the kind of compute that lends itself to the security they need for their own weights and for their customer's information. So I think a relevant case study here is to look at the computers that Google and Anthropic are renting from SpaceX.
Dwarkesh PatelGoogle, for example, is paying $900 million a month for 110,000 GPUs that are a blend of GB200s and GB300s. The price that Google is paying here is 2x the spot price per hour for those GPUs. And that spot price itself is more than 40% higher than it would have been in February. I want to emphasize the key conclusion here, that as AI models get smarter, they'll be better able to monetize the same amount of compute.
Dwarkesh PatelIf a true human-level software engineer could run on an H100 equivalent, then at today's prices for software engineers, that H100 should rent for over 250k a year. That's over 15x the current spot price for an H100, and this is not even accounting for the fact that your AI can work nights and weekends.
Dwarkesh PatelOf course, you might expect that if we had 10 million extra software engineers suddenly appear in the economy, the marginal value of a software engineer would decrease and thus the revenue that that H100 would be able to generate would not be 15x higher than it is right now. But I actually don't know if this is true. If we applied this argument to people instead of AIs, then this would be the classic lump of labor fallacy.
Dwarkesh PatelFor example, economists generally believe that high school immigration does not decrease wages in the long run because of how innovation and specialization increase the value of labor. Maybe this labor supply shock will be so big and so fast that we can't count on this general heuristic anymore. But if you believe what standard economics says, then the marginal value of labor and thus the marginal value of compute should stay astonishingly high. So let's think about what changes in such a world.
Dwarkesh PatelWell one of the things that would happen is that as the top labs get better and better at monetizing compute and the cost of compute increases it becomes harder for anybody else to compete against them because they have to bid for this resource against somebody who is basically able to make better use of it. Another thing that will happen and I think this is actually the most interesting implication of this whole thought exercise is that if you can train the best most efficient model then you'll be able to charge much higher margins than you can today.
Dwarkesh PatelThis is the Alkin Allen Effect on Economics and what it's basically saying is that if it costs $20 an hour to rent an H100, then it would be extremely stupid to use a weaker, less efficient model because it's going to burn more tokens on your expensive compute to get the exact same result. So labs will be able to charge a much larger premium if they can train a model that better economizes this scarce input. Basically, if you have a model that can get the same result by using less compute, then you've in some sense created more compute and the value of compute is going to increase. Another thing that will happen is that a lot of current popular applications of AI will probably get priced out. The reason AI is relatively cheap right now is that AI just can't do a lot of things that top humans can do, but this at some point will no longer be the case. And at that point, Google or Anthropic or OpenAI will be willing to pay more for the tokens to automate AI research than you or I will be willing to pay to make more AI slop talk.
Dwarkesh PatelI'm a bit worried that this kind of analysis honestly pattern matches a lot onto the ways that people in the past have been wrong about scarcity. I'm, for example, thinking of the famous Simon Ehrlich bet. Paul Ehrlich was this famous doomer about population growth, and he made this bet that a basket of commodities would increase in price rather than decrease in the decade preceding 1990. And this is a very famous bet because it's supposed to illustrate how Ehrlich's Malthusian worldview was wrong.
Dwarkesh Pateland how he did not anticipate the way in which market signals and human ingenuity can find better ways to economize scarce inputs. I'm guessing that the analogy to this bet is probably wrong. Other analysis has shown that if that bet had remained in a different decade, Erlich might well have won. But more generally, I think the supply of compute is much less elastic and much less capable of absorbing large demand shocks.
Dwarkesh Pateland much less capable of being accommodated by using different substitutes than the extraction of different metals. To illustrate why I think this 3x in compute capacity year over year is hard to budge or potentially even sustain is that I don't see how any of the three elements that constitute that 3x can be much accelerated. So 1.4x of that is coming from Moore's law. Far from increasing it, I think it'll be a miracle if we can just keep it going for a few more years. 1.2x is coming from building new fabs.
Dwarkesh PatelThis process is ultimately going to be bottlenecked up to 2030 and potentially even beyond by just building new ASML UV machines. Dylan, when he was on the podcast a few months ago, talked about this in great detail. And 1.8x comes from the fact that AI is absorbing a lot of wafer allocation that was previously going to smartphones and PCs.
Dwarkesh PatelThis is probably gonna hit a wall by the end of next year, when at the leading edge, N3 nodes at TF2MC, AI will have gone from 60% to 86%. At some point, you have just absorbed all leading edge weight for capacity for AI, and you can't keep increasing this number. So I don't know how we get, even to continue to do 3X compute scaling year over year for the next few years, much less go beyond that. At the end of the month, I go through the time on our tradition of closing my books. I started by opening Mercury, which is my banking platform.
Dwarkesh Patelto make sure that all my transactions are properly categorized. Auto-categorization rules handle the predictable stuff pretty well. But I'm constantly working with new contractors, you know, tutors and researchers and videographers, and I'm also trying new tools. Manually categorizing all of these transactions would add a couple of hours of overhead every single month. So instead of going through them one by one, I have Command, which is Mercury's built-in AI, to get a stab at all of them at once. Command proposes a category for each transaction and provides its rationale.
Dwarkesh PatelI just review, I fix anything that's off, and I approve. And once all this work is done in Mercury, it syncs everything with QuickBooks. And Command's judgment calls are genuinely good. It does the obvious things like looking at the vendor, but it also investigates who on my team made the purchase and looks at notes and memos to build up as much context as possible. This is just one of the ways you can use Command to automate the backend of your business. To learn more, go to mercury.com slash command.
Dwarkesh PatelMercury is a fintech company, not an FDIC insured bank. Banking services provided through Choice Financial Group and ColumnNA members FDIC. AI-generated responses and suggested actions may vary and are not guaranteed. Now, I want to clarify that at some point in the future, compute will get cheap again. At some point, we'll just have robots that can convert shores of silica sand and mines of copper into new computer chips. And then the price of compute is basically the raw inputs and the tools required to do this processing.
Dwarkesh PatelI'm just talking about this current pre-singularity regime where AI compute merely 3x's year over year, which is not enough to offset how much more valuable AI is becoming over time. By the way, the fact that anthropic revenue has been 10x in year over year, whereas their compute has only been 3x in year over year, I think illustrates how strong the economies of scale are in the model business. And logically, this makes sense. When you train a model, you just had to spend this one-time cost to learn all these different skills that then get to be shared across all your users.
Dwarkesh PatelThis is very unlike human labor, where each instance has to be retrained from scratch. I wish we didn't live in a world with such strong economies of scale for intelligence, because I'm worried about power concentration, but it seems we do. Okay, this was an aeration of a blog post that I also released on my website at dworkesh.com. Check it out for other posts or to be notified when I release a post in the future. Otherwise, I'll see you for the next full episode.