← All shows

Machine Learning Street Talk (MLST) - AI 2040: Plan A report - Daniel Kokotajlo & Thomas Larsen

Duration 1:29:41 · Language en · Published Sep 08, 2026 · 9 highlights

Summary

本期节目围绕 AI Futures Project 的《AI 2040:Plan A》展开,讨论如何在不放弃人工智能收益的同时,避免失控、权力集中、国际冲突、失业和滥用等系统性风险。嘉宾把情景预测比作军事推演:目的不是精确预言每个事件,而是通过具体叙事暴露计划中的薄弱环节,并以日本中途岛推演中反复“复活”航母的故事说明忽视模拟结果的危险。他们认为近年模型在真实工作中的能力跃升已造成明显的舆论转向,而最关键的能力门槛,是 AI 公司宁愿放弃全部人类员工也不愿放弃 AI 劳动力。节目进一步设想,当云端 AI、机器人、芯片制造和科研形成完全由机器运行的自我复制经济后,增长速度可能远超当前经济,即使能力只停留在人类顶尖专家水平,也足以彻底改变就业、城市和资源开发。Plan A 的核心做法是先短暂停止训练,建立覆盖绝大多数大型算力的核查与透明机制,再在可控范围内缓慢推进,并在“最大可可靠控制”的能力水平再次暂停。嘉宾强调“控制”只能暂时限制不对齐模型的行为,真正迈向远超人类的系统仍需可解释性等科学突破,以辨别模型是真诚合作还是仅在评估中伪装。最终,双方把根本分歧归结为 AI 究竟更像电力和互联网等普通技术,还是可复制的“云端人类”;如果 AI 能完整承担关键工作流并递归改进自身,后者将带来性质完全不同的社会变革。

Highlights

  1. The Japanese before the Battle of Midway did a bunch of war games, and they kept losing. Then they would sort of break the game: resurrect their aircraft carriers after they died and re-roll the dice when the Americans succeeded. From our perspective, that's reality yelling to th ...

    日本在中途岛海战前进行了大量军事推演,却一次次落败。于是他们开始破坏游戏规则:航母被击沉后再将其“复活”,美军成功时就重新掷骰子。在我们看来,这其实是现实借由推演向他们大喊:嘿,你们的计划糟透了。

    A vivid warning about ignoring simulations
  2. We were coming up with all of these technical answers to say why we shouldn't worry about AI, and yet the AI is just getting better all the time. Probably the biggest reason for the vibe shift is just the AIs being much better, much smarter, and much more useful at stuff in the r ...

    我们曾提出各种技术层面的答案,试图说明为什么不必担心 AI,但 AI 却始终在不断变强。这种舆论转向最主要的原因,或许就是 AI 确实聪明了许多,也更能在现实世界中发挥作用。

    Explains the recent AI vibe shift
  3. One milestone of AI progress that I think is extremely important is the point at which an AI company would rather fire their humans than fire their AIs. In the future, AIs would be able to one way or another do all of the things.

    我认为有一个极其重要的 AI 进展里程碑:一家 AI 公司宁愿解雇人类,也不愿停用它的 AI。未来,AI 将能以某种方式完成所有这些工作。

    A stark operational definition of AGI
  4. There can just be this whole industry doubling in the desert: strip mines, self-driving trucks, factories being built by humanoid robots producing more humanoid robots and more chip fabs. It can be doubling every year, every six months, every three months, faster and faster as th ...

    沙漠中可能出现一整套不断翻倍的产业:露天矿、自动驾驶卡车,以及由人形机器人建造、再生产更多机器人和芯片工厂的工厂。随着技术改进,它可能从每年翻倍变成每半年、每三个月翻倍,速度越来越快。届时谁知道世界其他地方怎样了,也许 Anthropic 已经把月球拆了。

    A memorable vision of runaway machine growth
  5. Even if you pause at top expert level, everything changes dramatically. You have this population of colleagues in the cloud that can substitute for humans at basically everything, except that they're cheaper, faster, and they don't take 20 years to reproduce; instead, they double ...

    即使在顶尖专家水平暂停,世界也会发生剧烈变化。你将拥有一群能在几乎所有事情上替代人类的“云端同事”,但它们更便宜、更快,而且不需要二十年才能繁衍,反而每年就能翻倍。仅凭人类水平的 AI,再给指数增长一些时间,就可能出现巨型新城、庞大露天矿以及铺满天际线的太阳能板。

    Shows why human-level AI alone is transformative
  6. The five problems are loss of control, concentration of power, war, jobs, and misuse. How do we solve all of them? One is just buy time, in particular buy time with human-level AIs, instead of pausing right now and saying no more capability advance.

    五大问题是失去控制、权力集中、战争、就业以及滥用。我们怎样同时解决它们?一个办法就是争取时间,尤其是在拥有人类水平 AI 的阶段争取时间,而不是现在就永久停止一切能力进步。

    Concise statement of Plan A's risk model
  7. You're not going to be able to figure out if the AI is aligned via behavioral evaluation alone. You're fundamentally going to need to distinguish between the AI that's doing the nice thing because it's pretending and waiting and biding time, and the AI that fundamentally is doing ...

    仅靠行为评估,你无法判断 AI 是否真正对齐。你必须从根本上区分两类 AI:一种只是装作友善、等待并伺机而动,另一种则因为真正想做好事而表现友善。这将要求我们对白盒层面的 AI 内部心智有深入理解。

    Pinpoints the hardest alignment challenge
  8. We basically round up 99% of the compute, which is not people's personal compute but big data centers, and send inspectors to confirm how many GPUs are at each location. Inference data centers are restricted so they can only serve customers and can't do training runs. In the trai ...

    我们基本上要纳管 99% 的算力,这并非个人电脑,而是大型数据中心,并派检查员确认每个地点有多少 GPU。推理数据中心将被限制为只能服务客户,不能进行训练;训练数据中心则由检查员监控日志并发布到互联网上,使其中发生的一切完全透明。

    The plan's boldest concrete governance mechanism
  9. The fundamental disagreement is: is AI more like electricity or airplanes, or is AI more like humans in the cloud? What I'm imagining is the AIs doing 100% of a bunch of important economic workflows, so you don't get the bottlenecking effect. Whether AI actually gets to literally ...

    根本分歧在于:AI 更像电力或飞机,还是更像云端人类?我设想的是 AI 完成一批重要经济工作流的 100%,从而不再受到人类环节的瓶颈限制。AI 能否真正做到完整承担这些关键任务,正是两种世界观差异背后的核心问题。

    Captures the episode's central crux
Full transcript

We're gonna talk about AI 2040 plan a which is our new scenario in which they build super intelligence in 2040 because they Go slow and paste the frontier. I mean, have you noticed this vibe shift? Yes, and I'm very happy is AI is AI more like the electricity or airplanes or is AI more like humans in the cloud The point at which an AI company would rather fire their humans then fire their AIs drip mines self-driving trucks, factories being built by humanoid robots producing more humanoid robots producing more chip fabs and so forth like that whole thing can just be like doubling every year every six months every three months faster and faster as the technology improves because of course the AIs will also be researching to improve the technology who knows what's going on in the rest of the world but Anthropic has disassembled the moon or for example

Or hypothetically, if an AI CEO was saying that their model was truth-seeking, I would only say the truth. But actually, the model was looking up the CEO's political opinions before answering. We build AIs, they're aligned to humanity, but to who? Like, is it the president? Is it the CEO? Is it some actually broad and good democratic process that aggregates everyone's values in sort of an endorsed way? Probably not going to be the last one, and so, you know... Don't hyperstition that! Then that's it for me. I think that's it. That's not...

a normal technology. Just stopping everything now, Plan S, would be better than the default. I would rather just stop everything now than continue going on our current trajectory. This episode is supported by Cyberfund. If you're building at the frontier of AI, they want to hear from you. Cyberfund believes the future belongs to AI natives who want to achieve the impossible. And that is why they're introducing the monastery.

for AI native founders. It's an environment of pure focus and rapid execution for founders operating at AI native speed. And they're offering teams $2 million each to participate. Apply now at cyber.fund. So I'm Thomas Larson. I work at the AI Futures Project along with Daniel here. I was the lead author on this project, Plan A AI 2040, which came out a few weeks ago.

And I was also a co-author on 237. Yep. And I'm Daniel Higatello. I run the AI Futures project and co-author both of these reports. Awesome. So, yeah, the context of this conversation is that there's this 2040 plan A, which we'll get into a lot of detail on. But I'm just interested before we get there. I mean, can you just tell me a bit more about the AI Futures project? I mean, how did it all come about? Yeah. So I used to work at OpenAI and while I was there, I did a variety of different things, evals.

forecasting governance memos. And I became gradually disillusioned with the leadership of the company and also the gap between how much information there is inside the industry and how much information there is outside and what you're allowed to say on the inside versus what you'd want to say. This is just a big gap.

When I left OpenAI, I wanted to be able to speak more freely and tell the world about what people on the inside see coming, basically. And AI 2027 and the AI Futures Project was our first project that we did along those lines. So I recruited a bunch of people to help me and we wrote this scenario called AI 2027. And it was a similar sort of thing to what I had done internally at OpenAI, but just much bigger and more ambitious and free for the whole world to see.

Very cool. And maybe we should just have a quick refresher on AI 2027. So this was an absolutely huge event. Many, many folks were talking about it. I guess like, did it achieve what you wanted it to achieve or what did you want to achieve with it? Yes, more so than expected. So the first goal that Thomas would remember, when we were working on it, the first goal was just a purely epistemic goal of the future is crazy and hard to predict.

Let's try our best to predict it and let's see how well we can do so let's game out a concrete scenario and even just for our own edification like we learned a lot from this whole exercise and we feel like we had a better understanding of what was coming and And then you know the secondary goal is lots of people seeing it and and being informed by it and starting conversations and so forth and that part blew past our expectations We had we made predictions beforehand about How many people would read it and it was 90 90th percentile outcome The thing I would add on the on the epistemic point is that I think things have been going more on track for at 2027 Then I would have predicted at the time we released it So like at the time we released it out of assumed that reality would have diverged much much further from our scenario than it has already I think like sort of the real world impacts like the the revenue trends for example, but also various other trends I think are pretty close to on track

for AI27, which has surprised me in a bad way. Interesting. Yeah, maybe we can reflect on that because I think you guys did a self-assessment on AI27 and it was something like 65 to 75% of it was on track, but AI software R&D uplift was only 0.17%. Can you explain that?

Yeah, so we've done two different blog posts where we take all the quantitative predictions made in AI 2027 that have resolved so far and compare them to reality. And we track this metric of how much of the distance has been crossed by reality compared to how much is crossed in the scenario. And in that way, we can get this overall sense of how fast are things going compared to in the scenario. And the top line number is something like 75% speed.

Basically, things are on track, but just going a little bit slower. The uplift number, I forget what it was that you just cited, that was actually, basically, at the time that we wrote AI277, we had a bad estimate of what the uplift was at the time that we published. We thought it was higher than it actually was.

What actually happened is that there was actually significant increase in uplift due to coding agents and so forth, but it was increasing from a lower level than we thought up to the level that we thought and then a bit above. And so the metric looks like it was only a small amount of progress because the metric was tracking from like where we thought it was to where it is. But does that make sense? Like basically because we had overestimated the metric at the beginning, It overall makes it appear like there's been less progress according to this particular metric that we're using, but it was because we had, yeah. I mean, can you tell me a little bit about forecasting in general? So I guess there's like a bullish take and there's a bear take on this. So my intuition is that reality is infinitely complicated. There's just these infinitely diverging trajectories and God knows what's going to happen the day after tomorrow.

But by the same token, though, reality is quite structured. It's quite convergent, and it is indeed possible to predict things that are going to happen because, you know, certain things reoccur with increasing regularity. So would you guys classify yourselves as forecasters? I mean, can you talk me through that? Yeah, so I would think, yeah, I think forecasting is a good name for what we do. The way I like to think about sort of why we're doing what we're doing is sort of like It's sort of like why people that are fighting wars do wargaming, where you're never gonna sort of predict the exact, you know, sequence of battles and exact sequence of sort of how your war will go at the beginning, because, you know, it's just gonna be really complicated, there's gonna be enemy action, you know, things are just not gonna go as you expect, there's just no way. But if you have sort of no, like, no concept of how, like, your initial plans might result in victory,

It's very unlikely that you'll actually succeed. And so I sort of think of, like, AI 2077 was sort of our attempt to just roll out, like, here's one way the AI future could go. Obviously, it's not going to go exactly like that. But it's going to be one concrete story that we can then sort of diverge from. And then Plan A was trying to be basically that, except now we're saying, what should the US government Like what should the US government do to sort of manage that well? And then that was also that was sort of supposed to be like the you know positive vision story and it was again sort of in the spirit of I think a war game trying to illustrate you know sort of one possible concrete future path and then of course things aren't going to actually go exactly like that but having like one viable plan that makes any sense at all uh is like you know we hope sort of a positive step forward relative to sort of the previous state of just like

abstract arguments in the void that are sort of not that tethered to reality. Yeah, that makes sense. It's certainly not abstract. I think it's very concrete. But I guess one thing that occurred to me is on the 27 piece, it was quite gloomy. And on on planet of 2040, it's far more optimistic. And, you know, but it seems like a bit of a mixture of conditioned prediction and recommendation at the same time. I mean, where do you guys land on that?

It is, in fact, a mixture of prediction and recommendation, unlike AI 2027, which is a pure prediction. And I think if we could do it all over again, we might try to be more clear from the beginning about our structure of what's a prediction and what's a recommendation. As it is, it's kind of mixed up, like some parts of it are predictions, some parts of it are recommendations.

There's a supplement that you can go to on the website that talks about which parts are predictions and which parts are recommendations. But I understand that's not very easy. You're not very apparent to people. But broadly speaking, plan A is the prediction part. So when the government implements plan A and they make the deal with China and there's all these pillars that they're upholding and so forth, that's our recommendation. Not a prediction.

And then usually most of the things that follow from that are predictions rather than recommendations So mostly it's just rolling out like what we think the consequences would be if you implemented plan a And then there's a few other things that our recommendations to for example the citizens dividend Yeah, yeah, the other thing I would add is just I think the thing we found is that it's very hard like when you're trying to make a recommendation Your it's very hard to disentangle the predictive aspects and the recommendation aspects because all of your predictions are like yeah colored by recommendations and your recommendations are inherently trying to be at least vaguely realistic like if we made recommendations that were like completely unrealistic and had like no bearing on reality but we nevertheless like stood by and we're like yes we should do this but we know it'll like absolutely zero percent never happen then that would have been a much less useful exercise than I think the one we did where the one we did was

Yeah, we were mostly trying to make recommendations. We made some recommendations that we think are pretty unlikely to happen. But we were trying to make some substantial concessions to realism as well and trying to actually aim for something that we think is at least moderately viable. I think that if we could do another scenario like this, we'd probably have a more clear structure of there'd be a central branch, which is the pure prediction branch.

which just goes all the way to the end like yeah 20 or 7 and it's just like here's what we actually are best guess and then there'd be branches off of it that are like at this point they do this recommendation instead and here's our recommendation and then after that it's just a prediction again of like what we think the consequences would be if you did this recommendation at this point um and in that way we and then maybe there'd be sub branches off of that But then it would be sort of clear at every point that like everything is a prediction except for these particular branch points, which are recommendations, you know. Yeah. And the war games thing was really interesting just just to kind of dwell on this a little bit because even if a war game is incorrect.

There must be some kind of information gained from it. So if you do a whole bunch of war games, there must be abstract motifs that do appear. So I guess this is what you think that if we do these different scenarios, we're almost guaranteed to have some kind of uplift. Yeah, that's basically right. I mean, so the example I like to bring up is Midway in particular, where the Japanese before the Battle of Midway did a bunch of war games. They did like a three day retreat where they like wargamed it out a bunch of times.

And they kept losing. And then they would sort of break the game. They would resurrect their aircraft carriers after they died. They would re-roll the dice on whether the Americans succeeded, when they succeeded, so that they would sample until the Americans' strikes failed. And so basically, from our perspective, that's like reality, sort of yelling to them through this mechanism of the war game. Like, hey, your plan is terrible. You're going to lose if you do it. And sort of.

You know that that's sort of the hope with sort of our like we ourselves do a bunch of war games But also do a bunch of sort of detailed scenario writing our hope is that every time we have to write a part of the scenario and The part of that scenario seems super unrealistic or you know Isn't really well modeled and doesn't really make sense or people are able to make really good criticisms of it online That's basically reality yelling at us and trying to like help us see reason and our hope is that we can sort of like Do enough of this sort of put up enough surface area so that we can sort of

Get that you know dose of reality from from the real world Or from the simulation of the real world which we hope is realistic enough to sort of accurately give us that information So we call this scenario scrutiny So basically we think that if you have an ambitious plan for what to do in the future you should try writing out Concretely what it would look like to implement that plan and what the consequences would be and this is a way of applying more scrutiny to your plan It's a way of sort of like stress testing your plan because it's opening your plan up to more like criticism basically. Also in addition to doing our actual scenarios we actually do literal war games where we get 10 people in a room for four hours and we game out a scenario like this and we've done maybe about a hundred of them total mostly AI 2027 style war games but also about 10 or so plan A war games where we say at the beginning of the war game we're gonna try to do plan A and then see how it goes wrong and

Just to give an example, I think in two separate Plan A war games, it went wrong in roughly the following way. Basically, there's going to be an election coming up, and the president in power is expecting to lose power and have his opposition party take over. And then even though he's already done Plan A and he has this beautiful deal with China and so forth, the president's like, well, I don't want...

I don't want my adversaries in the other party to now be in charge of superintelligence or whatever. So we're going to basically accelerate the timeline and try to get to superintelligence before the next election so that I can be the one in charge instead of my successor. And that's like a sort of political consideration that we didn't think about until it happened in our game. And it surfaced a possible failure mode of our plan.

Yeah, so interesting. I mean, even in the shower, I do sort of micro-tim war games. And it's really interesting just the regularity of which that they are useful. And I guess that's why all of us humans, we like to imagine and simulate situations. But is there an interesting boundary, though, between kind of like simulation and hypostation?

And what I mean by that is, you know, hypostation basically means it, but being a self-fulfilling prophecy. So maybe I'm expressing my agency, I'm expressing my will and saying, I want these things to happen and I'm kind of bending other people to my will. Is there an element of that that you're kind of establishing this in the zeitgeist and you're making it true? Yeah. So I would say that was maybe our biggest, maybe.

our biggest or at least one of our biggest worries with AI 2027 in particular this sort of like self-fulfilling prophecy in particular i'm pretty worried about this whole like increasing awareness of sort of very smart ai's and how important they'll be and how much they'll reshape the world and then that causing people to go oh man I want to be the one in charge of the AGI or the superintelligence, so I'm going to race towards that. And I think historically, that's been a big driver of the existing race. And I think that's been pretty bad. And so I'm actually pretty worried about that as one of the negative impacts of AI 2077. That was one of the reasons to feel a little bit better about the second project, AI 2040 Plan A, was that it was actually like, if that gets hyperstitioned, I think we'll be pretty happy.

Um, and so yeah, so like that one, you know, yes, it would be nice if we had precision. I really don't know how big the effect is. Um, I think probably most of the effect for both of them is via other paths. Like, you know, I still think that the main point bad 27 was like helping people be better informed about the situation. And that was like most of the goal. And I think that's most of what happened. Um, yeah, Daniel. Yeah. Yeah, I agree with that. I think that.

Hyperestitioning and self-fulfilling prophecies are totally a real phenomenon, but I think a lot of people tend to overestimate how much they are. And I think that to a first approximation, we should focus on accurately predicting the future. And then, in some cases, we'll find it in a situation where we can steer the future. But if you come at it trying to steer the future, you're going to get all muddled, and you're going to basically fall to wishful thinking, basically. So I think you start with just trying to accurately predict the future, and then you try to...

Shift it towards the better futures, and I think that's what we're doing. I mean, I suppose you guys are like the Marquis Brownlee of AI prediction now. So with great power comes responsibility. But on that note, I wanted to talk about the vibe shift. There's been a bit of a vibe shift. So, you know, MLST, we've always been quite skeptical about AI. And I'm trying to unpick exactly what it is that changed my mind.

Assuming that is what's happened it's very strange times and i don't even know what to believe anymore but you know all of the hacking stuff with hugging face i interviewed a polly research about you know reward seeking behavior and i think a lot of us have just seen the the change in behavior on models that have been rl trained to oblivion so yeah i think that there's just quite a few things going on now where loads of us are thinking oh my god.

And I used to be skeptical, like we would talk about whether they were Turing machines or not. Let me just have got a list here. Yeah, whether they were symbolic or neuro-symbolic, whether they were adaptive, whether they were conscious, whether they were correctly physically instantiated. And we were coming up with all of these kind of technical answers to say why we shouldn't worry about AI. And yet the AI is just getting better all the time. So have you noticed this vibe shift? Yes.

And I'm very happy. Tell me more. I mean, what do you think? I mean, because so many people have changed their mind. What do you think are the reasons? So I would say probably the biggest reason is just the AIs being much better and much smarter and being much more useful at stuff in the real world. Like when I have AIs try to automate various parts of my job, like they're just actually way, way, way better at it this year than two years ago. And, you know, four years ago, basically, like it was impossible. I was getting basically no uplift.

And my guess is that's been the biggest effect where many many people I think have had sort of their just like intuitive benchmark of like well Here's like a skill that I really care about and I know maybe pretty well and then sort of like the eyes have just like I think You know Joffrey Hinton one of the godfathers of AI I think he said once like Like was it able to tell a funny joke was sort of his internal benchmark and one once it could do that Which you know happened pretty early it happened that like you know probably somewhere between GPT-3 and GPT-4 Sort of he was like oh wow like these AI's like I don't see where it could end and I think that's probably happened for a lot of people I'd be very curious to hear more about your views actually if you can say I was listening to your interview with Ryan Greenblatt who was also a co-author on a plan a Earlier today. Yes indeed. Yeah, and I think

A bunch of the arguments you guys were having back then seemed very relevant to basically the current situation in the Hugging Face thing. And I'd be very curious to basically hear your views.

Yeah, I mean so you mentioned like the embodied thing like and and Ryan was talking about like when you scale up the RL massively right like back back then the sort of regime was like you were mostly doing pre-training that was where almost all the capabilities are coming from and then you do a sprinkling of post-training RL on top and then sort of you guys were talking and speculating about like hey what would happen if we like dumped like boatloads of RL compute would that be sufficient to sort of get the agentic behavior or Do you sort of like need this like physical embodiment and I think from my perspective like well

It seems like basically the answer was like you needed the boatload of RL compute to get the agentic behavior But not the physical embodiment and be curious. Yeah, be curious if you end up agreeing with that assessment Yeah, it's so difficult. So the way I think about it is I think representations are very important, and I use the term abstraction mounting quite a lot. So when we speak language and all of this gets ingested into the models, it's capturing the symbolic residue of language. And these have different levels of evolution. So, you know, a lot of concepts and mathematics, they're highly distilled, highly evolved. And we can say now that these models are intelligent. So for me, intelligence is about adaptivity. So basically means that I can

create novel combinations of things that we already know about in service of solving a particular task. Now, we know the models can do that. There are some great adaptivity benchmarks like ArcAGI3. It will say, oh, this is a maze and it will just put bits of knowledge together. You know, we should compare this to something like AlphaGo, right? Because move 37.

And you know, there was always this huge exploration problem in reinforcement learning that, you know, when it found move 37, it was kind of innovative but not creative. It was creative to us because we understood the creative space. We understood how it all hung together. And now with these RL trained language models.

They don't understand it quite in the way that we do, but they solve the exploration problem by understanding how things fit together. So they can just explore novel trajectories. They have a lot of base knowledge. They clearly do things that weren't possible before. But the grounding thing is interesting. So it's not necessarily an argument that you need to have physical instantiation. You need to have consciousness. But there is a spectrum of representations down this abstraction mountain. And the model seemed to be actually using the representations at multiple levels of resolution.

But it's not like intelligence is this magical quality. I still think that it relates to a scope of tasks and it relates to what you already know. And I think intelligence is quite perspectival. So I still think that it's going to be fractured and jagged and kind of fractionated if that makes sense.

Yeah, so I guess one question would be, you know, do you think that they like, do you have a view about like AGI timelines or like, when might we get a scenario like AI 2027 happening? In particular, what I mean by that is like, yeah, yeah. So like, I think one, one benchmark I really care about or maybe not really benchmark one milestone of AI progress that I think is extremely important is sort of this, the point at which an AI company would rather fire their humans than fire their AIs.

So like, you know, they would rather sort of give up on all human labor than give up on all AI labor. I think right now clearly we're in, we're still in this regime of like, clearly, you know, enthropic or open, I would rather have their human employees than, you know, give up on. Because people would just break apart. Like, right, there are just things that you need a human to do right now. And so if they fired all the humans, the company would just collapse. But in the future, that won't be the case. In the future, AIs would be able to one way or another do all of the things. And so, yeah.

Yeah. And so from our perspective, like that's going to happen at some point because there's nothing fundamental stopping the AIs from sort of reaching this human level of capability. The main question is just when. And we internally do a huge amount of analysis and like thinking about the various different methodologies for predicting this day in the life of somewhat different views on this question. But ultimately, I think that's maybe the most important, like the timelines question is just maybe the most important question for, you know, thinking about the future of AI, at least one of the top questions.

Yeah, so I'm curious to hear if you have a particular view. Well, let me give you some thoughts before I answer that particular question. So there are so many startups doing this recursive improving superintelligence. I've interviewed many of them. So for example, I interviewed Edward Hughes from Inherent in London the other day. And, you know, what he did was he recreated many scientific experiments from a whole bunch of popular machine learning papers. And he did it by masking out figures in the paper and getting a 27B you know, Gwen model. So what they did was a GRPO to the Gwen model. And that was how they solved the adaptivity problem because it's very difficult to fine tune a big fat model. So they, you know, adapted a controller model to control a harness like Codex. And their thesis was that if they can, you know, recreate scaled down versions of these experiments with construct validity, which means there's an LLM judge that is making sure they're not cheating and they're doing it correctly. Even that is interesting because, you know, we're ML people, we would always

say, these things take shortcuts, they'll always be kind of like validity problems. Weirdly, that's actually not as much of a problem as we thought it would be. And he thinks if they can recreate these experiments, then why couldn't they be creative? If they have the ability to recreate things, why couldn't they take the next step and say, oh, this is an interesting question. This is an interesting new problem to solve and go from there. So I was quite intrigued by that research. And indeed, I do think it is possible in the near future to have an automated AI scientist.

But there's always this thing in my mind, though, that there's a bit of a culture in Silicon Valley to reduce things or reify things. And so, for example, Elon Musk will say, well, you know, you're an engineer and this is your output and these are your metrics and you need to make the metrics go up. And we see everything in terms of like an optimization problem. And I always think that this is great for certain types of hill climable abstract problems where we have enough of a specification. So there's an interesting thing in optimization where if you have enough of a specification, the AI system can actually converge towards the solution. But when you're in the ambiguity regime, then you need to have the specification. And it's really mysterious what that means. Why do we have the taste of the deep understanding, whatever it is that we have in AI systems can't. So I guess I'm thinking that in objective kind of semi specified domains, we can hill climb and we can optimize until the cows come home. But I still feel that there's something missing.

Okay, well my response to be would be probably something like There's no this sort of there isn't really a binary between things that are like verifiable objectively and aren't or if there is like the things that are verifiable or is just like everything Like for example, you know building a unicorn startup you know, having a billion dollar valuation. That's a verifiable fact about the real world. It's long, it's sort of like long horizon. It's expensive to verify. Yeah. It's expensive to verify, but it's sort of a quantitative thing of like, you know, you got at one side, right? Like these like, you know, these like, you know, coding interview problems, which they're currently doing lots of RLVR on, right? That's like very, very cheaply, very easily, algorithmically verifiable. And then these sort of like real world things, which have sort of more expensive and longer horizon.

feedback loops. And I guess my view is that we're sort of going to get this continuous expansion of what the ads can do, driven probably in part by, you know, an expansion in the amount of RL.

and the type of RL that the ag companies are doing. More diverse, long horizon tasks. Yeah. And there's going to be this continuous process of expanding out through the different types of problems and different exactly how verifiable each task is until you get everything that humans can do. Because after all, we humans do learn how to do these long horizon tasks somehow. If I may add to that, I also think that AIs have been getting better at...

Everything including the like fuzzy hard to verify conceptually loaded blah blah blah blah blah blah like just try talking to like GPT-3 or GPT-4 and then talking to like fable about your favorite, you know non verifiable fuzzy task and Probably you'll find that the later AIs are noticeably better at those tasks and so One way or another it seems like there has been massive progress and I expect that to continue Yeah, I'm trying to come up with a good example. I mean, there is a sociological argument. I don't know if you guys read David Graber's book Bullshit Jobs, and he interviewed all of these people, and they were basically saying that my job is bullshit. You know, after about three or four beers, a lot of lawyers will say, you know, like a lot of what I do just isn't very important.

You know, so if we do kind of objectify and quantify everything that happens in an economy, I mean, I took a note here, I think you said by 2032, there might be 60 million agents running at 20 times human speed. And I'm just thinking like, what does that even mean?

And is the logical conclusion that we could have an economy which is only AI agents? And does it even make sense to have an economy which is only AI? I mean, just help, help, just make this make sense for me. Yeah. So my view is, yes, basically, the guys will be able to Everything or at least everything that really matters. So there is that yeah, you mentioned the notion of bullshit dogs I would I would just sort of start with let's consider everything we need to make better ai's as maybe like a first step of like Which is a large chunk of the economy So for example for this what you need you need to be able to build bigger better chips and more chips And that's the entire semiconductor supply chain to build semiconductor supply chain You need you know, you need sort of like a whole advanced economy. You need to build new robots to build

Uh sort of new fabs and then you need you know robot factories to build more robots And then you need researchers to build better ai's using that those massive compute and so I think once you sort of have all of that that is sort of like Enough to really speed up and sort of massively change the overall world even if for example You know, let's say let's say like there's Occupational licensing or whatever preventing the AIs from doing like some random legal work or some random whatever work or whatever bullshit jobs throughout the economy I think sort of what really really matters is sort of this is like the stuff that's actually really important in particular the robots the compute and the better AI and once you have sort of like The AIs that can do that and the capability to to sort of have that part of the economy like that section grow really massively then

Well, you'll see massive growth because it'll be really hard for the bullshit parts of the economy to constrain the growth of the parts of it that really, really want to grow fast because of the incentives that every actor has. In particular, every country has this big incentive to have an economy that grows faster than all the competitor countries. Yeah.

Getting a little bit philosophical. I think that a lot of economics and a lot of discussion of the economy is sort of focused on the relationship between the parts of the existing economy and like the prices going up and down and supply and demand and so forth But if you sort of zoom out The economy as a whole is a self replicating system and it always has been you know thousands of years ago It was a relatively small and simple self replicating system of like some villages of people they would farm and then they would have babies and then they would found new villages and then they would farm and have new babies and so forth and like it would grow exponentially over time but at a very slow rate. Now it's a much more complicated self replicating system that involves trucks and carrying equipment back and forth and factories and mines and so forth but still at a high level it's a self replicating system where we have people, we have trucks, we have machines, we have buildings and together they all build more people, more factories, more buildings, more machines and so forth.

Soon in a couple years perhaps we will get to the point where you can have that whole self-replicating you can have a self-replicating system that is entirely machine run with AI's and robots and According to our calculations the doubling time of this self-replicating system would be much faster than the sort of like roughly 20 year doubling time of the current economy And so, you know, that's that's what we have in the future Yeah that seems plausible to me but for some reason my intuition is it would become degenerate in some way. You know I think David Graeber even though he said bullshit jobs I think what he meant was there was an ineffable or kind of inscrutable components to jobs that we don't understand some kind of sociological function or something like that.

Because another thing, I also read that Citrini report, and you were writing about what happens when humans, for example, they might start defaulting on their mortgages, their wages go down so they can't be active participants in the labor market, and you were talking about an AI dividend and stuff like that. But even that is kind of hinting towards this notion that when the humans aren't participants anymore, you get this kind of mode collapse of the economy. Do you think that's the case?

Potentially but again like okay, so imagine that it's like so so I'm not enough of an account I haven't game down in as much detail to say like what happens to the prices of it like Like when the consumer demand drops like what the effects of that will be I'm actually not sure I don't think that's something that we've modeled that much in our Yeah, but hypothetically even if that parts really bad and even if like the consumers don't have any demand anymore or whatever If you have the level of AI and robot capabilities, such that you can have these fully autonomous AI's and robots doing all the things, then even just like a company like Anthropic, if it's big enough, and maybe if it partners with various other companies like some mining companies, can get this whole self-sustaining thing going. And so regardless of what's happening to all the humans, there can just be like this whole industry doubling in the desert. You know, strip mines.

Self-driving trucks factories being built by humanoid robots producing more humanoid robots producing more chip fabs and so forth like that whole thing can just be like Doubling every year every six months every three months faster and faster as the technology improves because of course the AI's will also be researching to improve the technology and then you end up with a situation where Who knows what's going on in the rest of the world, but enthropic has Disassembled the moon or for example Yeah, and to be clear, I think this relies on a very extreme level of AI capability, right? And my sense are like, I have sort of different intuitions. And sometimes I have an intuition of like, really, do I really actually think that fable could, or like future descendants versions of fable or mythos could, you know, do everything that we're talking about here. And I think where what it really comes down to is whether you're thinking of the AI as like in the reference class of, you know, what we currently use AI for is for or more like, you know, AI is just like,

an agentic human level employee, like basically a human in the cloud. A colleague in the cloud. A colleague in the cloud, yeah. And yeah, I think sort of the past few years of AI can be pretty well modeled as sort of an interpolation between the current AI systems and the workers in the cloud. And so I think that like sort of workers in the cloud, a vision of the future, like looks pretty good. And also I don't think we'll stop there. I think we'll go superhuman. Yeah.

Yeah, one point that I guess I'll bring up here is that if people read AI 2040, Plan A, one of the things that you might take away from it, which is I think a very important fact about the world, is that even if you pause at top expert level, everything changes dramatically. Like roughly what happens in our scenario is instead of doing an intelligence explosion, there's an international deal to ban intelligence explosions and to not have AI's recursively self-improving. And so they sort of end up pausing at roughly Top human level with a eyes at least for several years eventually they get the super intelligence in 20 in particular in 2040 they get the super intelligence But there's this like period during the 2030s where they basically have human level a eyes across all the disciplines But nothing super beyond that but even that alone like you just do the economic modeling and

It's kind of like you have this population of colleagues in the cloud that are, you know, excellent workers that can substitute for humans at basically everything, except that they're cheaper than humans, they're faster than humans, and they don't take 20 years to reproduce. Instead, they double every year, you know? And so as a result, the world is just completely transformed by the late 2030s. And, you know, all the humans are basically out of a job. There's giant new cities that have been constructed by robots.

huge strip mines in the special economic zones that have dug huge amounts of minerals out of the earth, solar panels filling the horizon on the ocean. Crazy stuff like that is possible with just human level AI and some time for the exponential growth to cook.

Yeah, I mean that's one thing I want to challenge you guys on is this notion that when we have an AI you can basically photocopy the weights and you can duplicate it. You can run it a thousand times. You can have one over here which is acquiring a load of skills to do this job and one over there to do that job and you can basically merge them together, right? You can kind of combine the skills and the whole thing is stackable compositional.

And that doesn't really marry with my experience. I'm really excited about what I've been doing with AI. And I found that you can make agents highly skilled within certain intellectual lineages. So you can, you know, you can bring in lots of source information and you can train them to do things. But I don't think they are yet composable. I think if they were, that would make me much more worried. What do you think about that? So is the way that you're trying to compose them, like is it?

Entirely at inference time or you all are you like training are you like training them to do two separate skills and then trying to like merge the weights somehow? Well, yeah, I mean this might be a separate thing but at the moment they are adaptive through chain of thoughts and you know kind of skill surface adaptation So basically memory systems and to be fair that is not very composable And this is a big problem that organizations deal with now. So all of these developers, they adapt their skill surfaces. And basically, their agents are different people. And it's really, really difficult for them to share skills because they might break the other agent because, you know, it doesn't work for whatever reason. Now, I can imagine a future where we do weight adaptation. That's what these inherent guys did. And maybe then it'll magically solve the problem and we can have, you know, we solve this knowledge sharing problem, right? That's the big thing. How do we accumulate information at the individual, at the organization level?

and have the agents, a bit like in the matrix, they can just put the skills in and they can do the thing. Maybe that's possible. Even then, I still think that the representations in neural networks, I call them fractured and tangled representations, which means they're a little bit janky, they're not completely robust, but they are sort of composable to some degree. So one thing I'd say about that is that even if you're right, I don't think that that would really seriously undermine the future that we're painting here because okay so now instead of just one cloud model that is doing all the jobs maybe you have a hundred cloud models or a thousand cloud models that are like specialized to different professions you know but you still get to the same outcome the other thing I would say is that in the same way the humans are specialized another thing I would say is that compared to humans it actually seems like there just is this effect where AIs are able to

Think about knowledge, right? If you go back in time ten years and we had this discussion, I think it would have seemed like a very live option that you would have needed like a thousand different cloud models created by Anthropic for different types of knowledge work. There's the coding cloud, there's the physics cloud, there's the literature cloud, and the argument for this to be pretty simple would be like, well this is how it works for humans.

You don't have one human who knows everything. Instead, you have humans who specialize in different disciplines and so forth. And the models have only a finite amount of parameters. Maybe you just can't pack all that knowledge into this finite amount of parameters and you need to have specialized AIs for different things. And in fact, for small enough models, that is true. And for really tiny models, you just can't teach them all the things that they currently know. And so you would need to have a specialized model for different things. But what we've learned empirically is that for big enough models, you can just train them on everything and then they learn everything at once.

It's not that they get worse at physics because they've also been trained on a bunch of coding. In fact, it's the opposite. The coding has some small gains for the physics. And so I do just think actually the most likely feature is just that there's a single model that's been trained on effectively the whole economy and is just dominating humans across the board at effectively everything. That seems like the natural continuation of the current trend.

And then probably you've got various, like you've got cheap versions of it. Like you've got distilled, distilled, yeah, small things. So for, because you really want to save your compute as much as possible. So you'll have as cheap models as possible for any given task doing that given task. Yeah, I mostly agree with that. I mean, I think my perspective is that the models that they're kind of the voice of everyone and the voice of no one at the same time. So they, they have a default voice in terms of they have a system prompt and they have some default modes of behavior that they fall into.

But when an expert such as yourselves, when you use these models, you kind of grant a perspective. So every single word you say, all of the reference material, all of your memory system and so on, what you do is you kind of carve a persona out of that model and you activate the knowledge in a coherent way in that particular domain. And the beauty of it is that many, many other people in different domains can do that. And you get this kind of, I think it's an illusion that the model has this general capability. But I think the models can be carved to be specialized experts.

any domain, but it's a latent capability rather than an explicit capability. Does that make sense? Yeah. That seems reasonable to me. How is it an illusion? It seems like they just do have general capability. There's a lot of things they can do. Well, so for example, you can give it any specified tasks. So let's say it's a problem in mathematics and it will hill climb towards it. So it's solving this intelligence problem or you can ask it any knowledge problem and you know, it might be the case that the path was forged through default modes of training. Or if it's something slightly on the long tail, then an expert can go in and they could, you know, like when you put a query in, it's like flashing a light into the darkness. And what you're doing is, is you're kind of making the path of least resistance roughly correct. And then it will do the correct thing. So I guess I'm just saying that there's a bit of a supervisor illusion. So when experts use it, magical things happen in well specified domains, magical things happen.

but there's still a bit of a space of ambiguity. But this gets to the next point, which is like, what do you think intelligence is? So the beauty of our collective intelligence is that we have so many different humans grounded in different intellectual lineages, and we're all attacking problems. So when there's a big fiasco on Twitter, we're all motivated to find holes. So we're being intelligent together. We're finding interesting angles, and the algorithm is prioritizing the good ones and we're using our minds together. And I can imagine AIs being just like that, right? So we have the diverse AIs that have different expertises and doing all, you know, just looking at problems from different angles. But I kind of imagine the future more like that rather than one big AI. Yeah. I think I guess I basically agree with what Daniel said earlier of just historically, I feel like that the perspective of like we'll have a bunch of different narrow AIs that are all doing different things has just, you know, not been right.

You know instead there have just been returns to scale and having everything all together I think one maybe intuition pump that I think I like is like in humans. I think it tends to be the case that Having all of the skills in one person is just really, really important for making really good things happen. An example is Elon. Elon has a certain amount of conscientiousness, a certain amount of technical knowledge, a certain amount of business knowledge and being extroverted and being able to push people.

I think each of those skills I think he's like quite high percentile in and the reason why Elon is so rare and that he can like run all of these you know insanely large and successful companies all at once and and no one else can really do that or at least has succeeded at doing that it's just because of the sort of multiplicative like you needed to be like 90th percentile and like each of these 10 10 domains which is just very very unlikely but if you could have an, if you get an AI that's sort of like, you could just train to be like really hyper-santile in all of these skills in such a way that no human, as if it would be like extremely rare or, you know, infinitesimally unlikely for any particular human to have all of those skills at once, I think you would just actually just be really, really, really good at changing the world in all of these like concrete and important ways just like sort of Elon has actually done that.

If I can add, though, again, I don't think this is a crux for the type of future that we're depicting. Like, suppose that we're wrong about this and that actually, like, the most efficient path forward is to have a ton of different specialized AIs. Well, it'll probably still be the case that it's, like, a few big companies, like Anthropic, that are just making tons of different specialized AIs. And then there's lots of different cloud models you can choose from and so forth. And in fact, it wouldn't just be, like, you can choose from lots of different cloud models, because at the time that we're talking about, there'd be much more autonomous

than they are now. And so it would be more like Claude is choosing between a ton of different Claude models. And there's a Claude swarm, which consists of lots of different specialized models that are all working well together. And they have some sort of internal bureaucracy structure. And then that swarm is going out and negotiating business deals and creating new technologies and starting up new startups and doing all these things.

And it just like how a human like if you had a population of immigrants of human immigrants They would all have different specialized skills But they would work together to create new companies and you know get jobs and things like that and it'd be like that You know like there'd be lots of different clods, but they'd all be working together and so Zooming out there'd still be like this phenomenon of anthropic is eating the economy and You know and the robots as a whole are starting to self-replicate, you know, yeah To be fair, I don't think it's a crux either. Maybe if it is, it's only insofar as when you have a distributed collective system, you might have additional bottlenecks because you have all of the message passing between all the different agents and whatnot. And I watched a wonderful Santa Fe talk about this that even in the natural world, there's a kind of Goldilocks zone between the ratio of intelligence between the individual and the collective. And we might have some weird kind of convergence there.

Yeah, I mean, it's just quite interesting just think about how this works. Oh, sorry, Daniel. It would be safer. Like I think still we'd be there'd be lots of serious element concerns in that world. But I think it would be like a little bit safer because of the reason you mentioned where like it might be easier to like oversee what's going on if there's lots of different specialized agents communicating with each other compared to if they're all just like clones of each other and they all know all the things.

So maybe to avoid hyper positioning, we should say we endorse the vision that you've painted and we don't endorse the vision that we're painting. Yeah, very true. And as an aside, I love this concept of how learning happens at the individual level. I liken it to evolution. So there's like kind of phylogenetic adaptation, ontogenic adaptation, cultural adaptation. And I think the next wave of AI is when we actually have agents kind of talking to each other and learning and specializing. And maybe there'll be bad behaviors as well. Maybe there'll be collusion and lots of bad things happening. But I think all of this is going to play out. But it's interesting that you're talking about Elon though.

I think the magic of Elon is not so much his brilliance at engineering and optimization. It's his ability to recognize areas that are interesting that might work in the future, because that's the creativity thing. That's the science thing rather than the engineering thing. Because if we use an example of Amazon, for example, so that's an adaptive ecosystem. It's like an organism.

And what it does is it's always kind of thinking about new ways to adapt and rewire its structure. So it might be like logistics, for example. And then it will kind of output a bunch of skills. And then it will ruthlessly optimize those skills. There's like an adaptive component and there's an optimization component. And the organism is just constantly moving around. But I guess the question from this perspective, though, is how much of that, in principle, could be done by AIs? I guess you think all of it.

Yep, all of it. Like I guess, yeah, maybe just one way to say this, like Chris Blee is just like, look, the brain is a machine. Anything that the brain can do, we will be able to do with machines. Looking at, I think a very useful exercise is to compare the architecture of an actual human brain to a modern GPU or data center as a whole. And if you look, if you like try to do this comparison, you know, an H100 GPU is actually has pretty similar specifications to a human brain.

you know, like has sort of similar, you know, it is basically like similar in terms of like, I think maybe the most important metric is just how many flops per second, like how much total compute capacity does the brain have versus does the GPU have. And for nature 100, you know, it's like one e 15 flops per second in FP 16 for the for the brain, you know, it depends on exactly how you count and whether you use a synapse basis or a neuron basis or whatever. But it's like, you know, somewhere between 10 to the 12 and 10 to the 18 flops. So, you know, sort of like an H100 is like right dang nav in the middle of at least the log distribution over that sort of order of magnitude. Also just the architecture like these things are neural nets they're not ordinary software so they start off with random spaghetti tangles of randomly generated. So they start off randomly initialized it's just a giant spaghetti tangle in there just like how when you're born you just have a bunch of neurons that are just like randomly synapse connecting to each other.

And then there's this whole process of training where the the connections get pruned and circuitry starts to take shape that is effective at scoring highly in whatever the training environment is. And there are differences between how it works in AI and how it works in the human brain. But broadly speaking, they're just like an artificial brain. And and so just how like humans learn skills, which means like, what does it mean for Elon to have these skills? Well, what it means is there are some circuits of neurons and synopses in his brain that are doing very complicated and sophisticated calculations that are those skills. And then similarly, like in Claude, there's a bunch of circuitry that's been etched into Claude through the training process that is various skills. And in principle, you could have a big enough Claude that would have the same type of circuitry that Elon has. And then just a few other notes to add, I think.

in comparing the brain to modern ML systems. The amount, the brain is sort of more parallel than current ML systems. Like there's just more computations happening in parallel, but the serial depth is lower, right? Like the amount of computations happening in sequence, right? The amount of, you know, computations, the amount of neurons that can fire in sequence in a second, you know, depends on the type of neuron, but it's like, you know, hopefully I don't get this wrong. Between like one and a thousand, depending on the type of neuron is my recollection.

Yeah, I think that's in the range. I think it depends on the type of neuron. But anyways, a computer obviously can fire and do computations much, much faster than that in serial. You can very often get clock speeds are typically on the order of like a gigahertz or more. So you can get many, many order of magnitude basically better in the GPUs in terms of serial processing speed than the brain.

But it's also worth noting that the architecture of the models themselves are much worse in a bunch of ways than the human brain. In particular, it's harder for different parts of the ML model to talk to each other than it is for different parts of the brain to talk to each other. And so I do expect that to get this level of AI, we're going to need a bunch of algorithmic improvements on top of existing models and then there's a question of like well exactly how many algorithmic improvements and how qualitatively different do they need to be from current systems and that's a very open question from my perspective.

Yeah, I mean, I think the crux of a lot of this is that you guys think that intelligence is computable. And I'm not sure I want to litigate the whole functionalism thing today. But I guess my perspective is I zoom out. So I think that intelligence is externalized. It's collective. I think that Elon doesn't have quite as much agency as you think he does. I think that he's using tools. He's using social media. He's getting ideas in there. He has obligations. He has people around him and so on. So, you know, I guess I I think that these intelligence circuits and motifs exist, but they exist outside. There are just very complex dynamics. And I suppose in that sense, it doesn't really matter if the ecosystem is made up of AIs and humans together because they can participate in this super organism. So I suppose the only crux then would be that it would play some kind of limits on its scale. So the question is, do you believe that sort of you could have an AI society made up of AIs that were trained with something like current day ML techniques?

Um... you know, that we're passing information between each other and developing abstractions sort of in a community in the same way that our current civilization does. Like could you basically have something like our current whole economy, but was made with roughly modern day ML systems on your views? Absolutely. I mean, even the human brain is an example. So the human brain is not too incomplete, but we can expand our memory. We can use tools. We can work as collectives because we could use the same argument against transformers. They're not too incomplete, but now they can use tools. They can be like agents. They can build collectives. They can build society.

So, in a sense, this is what I was saying earlier about these objections. They kind of fall away when you have these insanely complex collectives that are sharing information with each other.

And there's also quite an interesting thing here as well, which is that it almost doesn't matter how we evolved or how neural networks were trained, because you get new phenomena emerge when they are placed in this kind of collective setting. So I think a lot of our intuitions are broken there. And that's why you probably shouldn't spend too long litigating this, because I think in principle, that kind of behavior could emerge. But I did want to ask you though, can you distinguish intelligence, capability, and power?

This is a philosophical one. So we get to the philosopher. Um, yes, we can. Uh, so I think that we can distinct intelligence versus capability. Um, I often try to say that we should just define intelligence as a bunch as like an aggregate of capability actually, or maybe like an aggregate of cognitive capabilities. Like maybe there's some physical capabilities, like how strong your actuator is, but then there's also cognitive capabilities, like, um, Are you able to distinguish a cat from a dog? And are you able to speak grammatical sentences? And how much do you know about Paris and things like that? And so maybe I would just say like intelligence is like a sort of like aggregate of all the cognitive capabilities. And then power, well, that depends on other things like how you are embedded in the world and what affordances you have, what actuators you have, what, how other agents are going to react to you like.

You know, the president has more power than me because of the location he's in and because of the role he's been given rather than because of his like physical strength or something. So yeah, power different from intelligence, different from capabilities. Now, we've not spoken enough about AI 2040. So maybe we should start with the four principles, right? So by time, transparency of research, diffuse AI broadly and reversibility.

Yeah, so I can sort of summarize where we're coming from here. So basically, at a high level, the goal of Plan A is to sort of solve the major problems that we see in AI, and so we predict will happen by default. The biggest problems that we're sort of identifying on the horizon is like one, this risk of loss of control, so just like the AI is actually getting out of control. Two, Um concentration of power. So like we build ai's they're aligned to humanity But to who likes, you know, is it the president? Is it the ceo? Is it some actually broad and good democratic process that aggregates everyone's values in sort of an endorsed way? Um, you're probably not going to be the last one. And so, you know, don't hyperstation that Yeah, I hope not we wanted to be the last one. Um, uh, yeah, then there's risk of sort of conflict over ai Um, so, you know, in particular, I think we're worried

you know, we're worried about literal World War three where countries, especially countries losing the AI race, sort of like realize that they're losing the air race and that they will be extremely disempowered by the winners of the air race. And so sort of they're in this classic situation where they're losing power. And so it's like in their incentives to sort of have a conflict happen sooner rather than later before they've lost all of their power. And this is like, you know, ripe for for conflict, basically. Then there's then there's, you know, Finally like risk of misuse so like you know what happens when they eyes that can build bioweapons are really cheap and broadly depused and open source And also the jobs is sort of the so those are the five problems right loss of control concentration of power war Jobs misuse so lots of problems. We want to solve all of them. How do we solve all of them? Well one is just by time in particular by time with human level as so instead of

basically pausing right now and saying like no more capability advance our proposal is basically go to roughly human level AI and then buy as much time as possible basically with human level AIs and then have those AIs which are hopefully smart enough to be really helpful for solving these problems also smart enough to sort of start sort of causing a bunch of these problems and providing the impetus for, you know, society to actually get us back together and really get going and investing huge amounts of resources on actually doing this stuff. There's a couple of things that happen in Plan A in AI 2040. There's a, like a six month to one year hard pause that happens as soon as they start implementing it. But the reason why it's that long is because they need that time to set up the infrastructure to proceed with AI development again, but in a

safer and more transparent way. So it does start off with a pause and I think that we would recommend like all things considered that you just do that right now. So get the infrastructure set up as soon as possible and that would require like a temporary pause. Once you've passed that stage and you've got the infrastructure set up, then you do proceed with AI development but in this transparent, more cautious way. So in particular, you're not doing crazy intelligence explosions. You're using safety cases and sort of gradually scaling up the level of AI capability and you're doing it in a very transparent way so that everyone can see what's going on.

And then there's a second pause that happens a few years later than that, which is when they reach the maximum controllable level of AI, which we think is roughly around top human expert level. And so in some sense, our view is something like pause at top human expert level, but it's a bit more nuanced than that. It's more like pause at the maximum level that you can reliably control, which we think would be roughly around top human expert level.

before you get to that level, don't race like crazy. Like you want to be sort of like slowly approaching that level so that you don't blow past it and lose control. You frame the piece around the importance of alignment and control. So alignment is basically, you know, does what we want to do and control is, you know, maybe contain it, maybe negotiate it with and so on. But you were just saying, OK, so maybe we can trust up to top human expert level.

But it's a little bit fractured, isn't it? I mean, how could you know, for example, the difference between a good AI and a bad AI? I mean, what would that look like?

Yeah, so I think okay. I think maybe it's first important to distinguish between alignment and control so What we mean by alignment is that the AI? Basically will take good will take good actions won't do sort of like catastrophic unintended behaviors like trying to take over the world like the recent hugging face incident Alignment means it has the personality traits the goals the values etc that it is supposed to have you know and then control means that even if it wasn't aligned, even if it was trying to do very bad things that we didn't want it to do, it couldn't. We like have mechanisms in place to prevent it. So sort of analogous to like, you know, you could imagine sort of a company with employees, an insider threat.

it would be misaligned with respect to the values of the company. But if there was good enough security measures internally to make sure that that employee couldn't run away with all the secrets, then we would say that that company has adequate control put in place such that even insider threats were misaligned humans or AIs wouldn't be able to do the catastrophic outcomes. Just to give a topical example of this, just yesterday, I think, in response to the Hugging Face incident, OpenAI announced that they are improving their security.

significantly and that they're going to have other AIs monitoring all of their new AIs in training and in invals and that if the monitor AIs detect what seems to be some sort of hack, then within half an hour a human will be notified who can then look in and try to shut it down. That's an example of a control intervention, not an alignment intervention. They're not making the AIs more aligned at all, but they're just improving the security so that the AIs can't get up to that much trouble. Yeah.

OK. And then basically, the sort of important high-level claim is, ultimately, we're going to need an alignment. Control is sort of, there's a time bomb associated with it. In particular, the time bomb is, when are the AIs so smart and so good at subverting any control measures that we put in, that if they were trying to screw us over, we would just fail?

you know, they would figure out some way of hacking around whatever systems you were able to put into place. And so ultimately, we're sort of going to need to solve alignment. The problem with alignment is, as you were saying, basically, it's much harder to measure. It's much harder to know whether you've succeeded or not than control. And so basically, our overall proposal is, in the meantime, so between, in our scenario between sort of 2030 and 2040, which is roughly when we're in the human range of AI capabilities, we're going to basically almost entirely rely on control.

And we're going to do things like red team, blue team games where we have our eyes sort of try to escape our sandboxes or our control techniques. And then we're going to see if the red team succeeds. And then if they do, we're going to iterate.

and make better security until they can't succeed anymore. And then we have this ability to measure it by just seeing, hey, can the AI succeed? Can the AI augmented with humans succeed? Can humans playing the role of AIs succeed at gutting out? With alignment, I think in order to be confident in the alignment of our systems, I think we're going to need sort of like more fundamental scientific breakthroughs, where I think you're not going to be able to figure out if the eye is aligned via behavioral evaluation alone, where you just look at what the eyes are doing. I think you're going to need to sort of understand something more deep about what's going on inside the eyes minds. You're going to be able, you're going to need something like interpretability, you're going to need some sort of notion of like, because you're going to fundamentally need to distinguish between the eyes that's pretending.

is doing the nice thing because it's pretending and is waiting and biding time, and the AI that fundamentally is doing the nice thing because it wants to do the nice thing. And that, I think, will require sort of this white box understanding of what's going on inside the eyes. Whereas in the control thing, you can just figure out via empirical behavioral iteration. And so overall, basically, our story in AI 2040 is basically between 2030 and 2040, we're going to We're going to do this buying time. We're going to rely on control. We're going to have the human level AIs. We're going to use those human level AIs to make a bunch of progress on alignment, but also on other things.

Um, and then by 2040 in the story, we've made enough alignment progress where it's like, Hey, okay, we really don't need to rely on control. Therefore we can scale to these like sort of wildly super human eyes where we're really relying on alignment. And if we were to try to control them would totally fail because if they were misaligned, they would be able to like completely subvert our control measures. Do you think in principle that this problem is going to get easier or harder as the models get more sophisticated? I think it's getting harder because the models will just be.

There will be a bunch of properties the models will have which will make it much harder to line them in particular So we're already seeing sort of growth in situational awareness where like the AI is a few years ago. We're like Kind of not that situationally aware and then now I think they're like pretty good at you know recognizing like for example our friends at Redwood You know they often do sort of these control evals and the models often you know in the evaluations where they're like trying to break out of these control setups will think oh hey this this sort of looks like a redwood a literal redwood research uh sort of like control evaluation and they're like so they basically realize like hey i'm like they realize exactly what's happening to them basically and i think this is just going to continue to happen and the more the ai sort of like realize the exact setup the more they're sort of like

gonna be able to perturb their behavior to like look good to the exact process that's like evaluating on training them And that's gonna come further than further apart from like sort of the actual measurement of whether they're actually good or not Yeah, and just to like add something to that I think in some sense the core problem is that it is already somewhat easy to think that you've solved the alignment problem and be wrong and That's going to get easier and easier over time as the models get more sophisticated and start being more aware of their situation and clever and stuff like that. And so it's not that we think that like there's going to be loads and loads of egregious failures where the AIs are just like going around killing people. No, it's almost the opposite. It's going to be that like it'll be extremely easy to end up in a situation where the AIs are in fact misaligned, but you don't know that because they're doing everything right as far as you can tell.

The number of ways in which that could end up happening is just like going to increase over time and it's going to be so easy to end up into that trap basically So okay, so we talked about sort of the buying time like why we want to sort of extend the time with a GI But what do we actually do during that time? So the second principle is basically transparency so Transparency is not necessary for sort of making everything else happen, but it's really really helpful the main There's a huge number of upsides of transparency. One is there's this concentration of power issue. We're very worried about sort of like someone building super intelligence and it being aligned to only some, you know, particular group of people. We think it's much harder for that to happen in sort of a like non-democratic way if sort of society as a whole can see the whole time what's going on with AI.

How smart they are who they're aligned to right like if let's say like an a ICU that was evil was trying to like Backdoor their model and put in training day that says hey obey me and don't obey anyone else Or like you know hypothetically if an a ICU was saying that their model was truth-seeking I would only say the truth, but actually the model was like looking up that CEO's political opinions before answering Which happened Basically, so the transparency will help at least somewhat with that. The other thing that maybe transparency helps a lot with is sort of this issue of just like government capacity, where in Plene, we sort of want governments to make these like pretty technical and like really complicated decisions.

on like, hey, exactly how much AI scaling to allow, like what risks are okay versus what risks aren't okay, like exactly what, you know, what architectures are maybe safe versus what architectures are not safe, what deployment, like what sort of control scaffolds are sufficient to entail safety versus which ones are sort of bogus safety washing. And sort of making all of those calls, I think will be very, very tricky, particularly given that the government's expertise in AI is really, really bad. And so one of our core hopes is that basically with as much transparent with with more transparency and more sort of public understanding into what's going on in the air companies that relieves pressure on the regulators because of something sort of catastrophically or essentially unsafe is happening sort of society as a whole academia you know non-profits other AI companies who have been incentive to say like hey my competitors being super unsafe other governments so like China has been sent to do this with US labs US you know government you know has incentive to use with the Chinese AI companies sort of everyone who's an adversary

or just wants to make sure that things are safe, has this big incentive and now has the affordance to actually look over what's going on, what is necessary to do safety. And I think maybe one thing that's really topical here is there's this hugging face incident that happened very recently a few weeks ago with the AIs inside OpenAI, creating the internal message board. We still have very little clue about the exact motivations of those AIs, the exact context, the exact prompt during the cyber evaluation that would prompt them to start doing this. If I personally had much more access to what was going on, I would have a much more informed and better opinion on exactly what caused this and what mechanisms in the future could have been done to prevent this and what analogous future things I should be worried about.

Sort of because of this and I have a bunch of different hypotheses But it's hard for me to sort of figure out which is which without access to the data and so basically in plan a a core principle would basically be all of that stuff would be transparent to the public not just the governments and So sort of society as a whole would be able to like weigh in there be able to be public and informed debates The scientific community especially right like if you want to have a bunch of scientists and academics and Non-profits and startups all like weighing in on stuff well then they need to have the information and you can't really share it with all of them without showing it with the public so So just might as well share it with the public. Um, I feel like maybe we should also say like we talked about the five goals and then we talked about these pillars which are kind of intermediate but I kind of want to go to the under the spectrum and talk about like what are the actual like concrete things that the US and China agree to in plan a and how do they lead to those things so specifically the sequences

we basically round up 99% of the compute, which is not people's personal compute, but like big data centers because most of the world's compute is in big data centers. And we, the US and China and other countries involved send inspectors to confirm like, yes, there are this many GPUs at this location. There are that many GPUs at that location. Having done that, we then set up this verification infrastructure and this transparency infrastructure so that there are inference data centers that serve customers just like today and that are restricted so that they can only do inference and only serve customers like that and they can't do any training runs and so the inspectors make sure that they can't do training runs on those data centers and then we have the training data centers which are the totally transparent data centers and that's where the research happens and on those ones

They still operate like normally, but there's inspectors from the different countries that are basically monitoring the logs of what's going on in the data center and publishing it to the internet. And so it's totally transparent what's going on in those data centers. You might need some time to set this up. That's why we had like the six to 12 month pause that I mentioned earlier. But once you get all this stuff set up, then you can proceed with AI development under these conditions of total research transparency. And because you have this transparency set up, it's a lot easier for countries to make additional further agreements about what to do and what not to do because they can just see what everybody is doing and so they can just enforce the agreements pretty easily. For example, and here's where we would say it's very important that they agree not to do a crazy intelligence explosion and instead proceed slowly and cautiously. And it's very important that they agree to do all this control setup with all the red teaming and so forth.

But because of the transparency, they can make those agreements on an ad hoc basis. They can just keep making more agreements like that, and they can just adjust them as needed based on the changing situation on the ground because they can all see the situation on the ground because of the transparency. And then this also is very important because if you want to have some sort of deal between the US and China, the US and China don't trust each other. And so you need to have some way of enforcing and verifying compliance with the deal. And the transparency, of course...

goes a long way towards helping that. Yeah. What about cheating in dark markets appearing? Yeah. So, okay. So we've done a bunch of thinking about this overall. So there's two ways you could cheat. One is you could get a bunch of GPUs and try to have them not be discovered by sort of the US and China and like hide them away, put them under a mountain somewhere and then sort of like do your training runs in secret. The other way you could cheat is you could, on the giant legal known data centers, you could be running a giant training run, but then trying to make it look like sort of everything's chill, like situation normal. It's not doing anything illegal. And the mitigations are different for these two threat models. So for the tiny amount of compute under a mountain, there are basically two mitigations. One is

sort of like as Daniel was saying, round up enough of this compute where it's like pretty hard to get a substantial size of compute under the mountain. And the second is sort of like do normal intelligence gathering and like look for these things proactively over the course of the 10 years. Like the 10 years slowdown happened in the scenario, which you know, in reality, it might be longer or shorter and sort of like try to find it. And I think for basically for any significant size GPU cluster, I think both of these independently have a quite good chance of working.

And so in practice, I think I'm not that worried about sort of large hidden compute clusters like secretly under mountains or whatever. I think the maximum realistic size, in my opinion, is something like a few hundred thousand GPUs, a few hundred thousand H100s, something like that hidden away. I think that this would not be enough to sort of compete at the frontier, especially assuming like the 2030.

AGI timelines where you have like sort of the biggest day activities in the world having like millions or tens of millions of h100s in their biggest training runs And then for the for the sort of legal clusters we sort of basically the hope is we have a bunch of verification infrastructure running on those data centers basically and the The main point of that verification infrastructure is to ensure that the computation that's happening on those clusters is transparent. In particular, that everyone can see it. So basically, the hope is you make sure that there's no calculations that are running that aren't transparent. And then for everything that is transparent, well, then regulation and agreements as normal can work. And you can say, hey, we agree to run this control scaffold. If you agree to run this control scaffold, and that can just happen. And both sides can be confident that they're both agreeing.

Yeah, isn't this just a matter of national security though? Isn't it a bit ambitious just to make it completely transparent? Yes Yeah, so one of the if you try let's try to talk about some of the effects of doing this transparency Well, we would be immediately publishing all the core training recipes of anthropic and open AI for the world to see Anthropic and OpenAI will not be happy with this. It will cut into their valuations dramatically. Why will it cut into their valuations dramatically? Well, because it will allow other competitors like Microsoft and Alibaba to catch up or whatever. And so that's why they're going to hate it probably. But I would say this is a feature not a bug. We want there to be multiple different AI companies at the frontier at roughly similar levels of capability.

AI to commoditize instead of being monopolized or oligopolized. And also this will disincentivize further investment, right? Like investors will be less interested in building a trillion dollar cluster if they won't be able to get the monopoly rents from that cluster. But we think this is good because like again we are going to be in a world where Going too fast is our main problem and so having a bit less incentive to invest and going a bit slower is actually just I think a feature not a bug There still will be investment like there still be lots of money to made and so like progress will continue It's just not at quite the same rate and and again. We think this is good Now this does shift. This is kind of a gift to China relative to the US like this does mean that like China gets some algorithms that they might have had trouble getting before

And insofar as you really don't like that, well then negotiate it. You can have a horse trading type thing where when US and China are making the deal, the US is like, well, since we're giving you all this stuff with the transparency, why don't you give us something else in return? And we can try to work that out. Like a more favorable compute distribution. Like a more favorable compute distribution, for example. So there's got to be some combination of carrots and sticks and trading going back and forth that we think would be in the interest of both sides. Another thing worth mentioning is that

It's not as big of a gift to China as you might think because... security is poor at these companies. And so they're probably through their spy networks and through leaks getting most of the information anyway. I don't know if you guys saw Sam's tweet yesterday basically saying that they're going to pause training for a while. And it made me think, I mean, why not just stop now? I mean, you guys are actually quite bullish about some of the positive things that I can do. I mean, there are loads of examples in your in your article, but you know, one example was in hospitals, we could actually have little devices that kind of decontaminate the air and stop the transmission of diseases.

and stuff like that. So it's not like you guys actually think we should stop. I would say, I mean, we do think we should do something like plan A as soon as possible. Like I think that, first of all, I think that just stopping everything now, plan S would be better than the default. Like I would rather just stop everything now than continue going on our current trajectory. Secondly, our actual recommendation would be to do plan A. So you do like a temporary inference only pause now so that you can set up all the transparency and verification infrastructure.

And then you can continue in this more distributed, you know, transparent, cautious way as previously described. And then having continued in that way, you basically go up until the level that you feel like you can control reliably until you feel confident that you've solved alignment enough that you can give up on control. And that's what we depict happening over the course of the 2030s in our scenario. And why do you guys think that AI discourse is so bad? Is it?

Is it unique to AI or is just discourse bad in general? I mean, discourse is bad in general, right? Well, why? Yeah. I don't know. I think there's a lot to say. Yeah. I mean, so one thing is like, I mean, there's like massive incentives.

for people and sort of like motivated reasoning for people at ad companies to sort of like think that what they're doing is like justified and good and that you know they shouldn't do like costly actions that would make the situation better because actually those costly actions would be bad for whatever reason so like there's much rationalization. Yeah I think most of the effect is probably just like generically discourses like hard and bad like there aren't like you know Twitter or whatever is like very like unnuanced and argumentative.

Yeah, I don't know. I guess I'll put a plug out. I think less wrong in particular, which is where I try to do most of my discourse. I think the discourse quality there is actually pretty good on average. Generally, when I write a post and go into the comments, I think generally the comments are just quite thoughtful and technically informed and whatnot, especially relative to other places like Twitter or whatever.

And then, yeah, maybe a final thing is just like in DC in particular, which is an area I care a lot about. I really want sort of DC and sort of the governance, you know, the government of the United States and other governments to react well to the technology. I think there in particular, there's sort of two problems going on. One is there isn't very much expertise, right? There's like.

You know the government their government is not hiring like really high quality technical experts who really know what they're doing and the second is sort of this this just generic like I think the conversation is not happening. Like the incentives for everyone in DC are not towards like sort of like truthfully and accurately understanding situation the incentives are sort of.

For every individual to sort of like say stuff that sounds good and looks good and it's sort of like within the DC overton window so that they can make a lot of friends Unfortunately, I think this just like comes apart from the actual reality I think the actual reality of the situation is like it turns out a bunch of sort of like controversial and niche views about AI were true right like the whole AGI hypothesis just like is correct and I think DC basically just like hasn't come to grips with that and so Basically almost all of the discourse that's happening there I think is just like fundamentally anchored on completely wrong assumptions But how the technology works in particular the assumption that like it's like mostly fake news and it's all a bubble Or that it'll be like the next internet which is like yeah, maybe Which is like better than it was a few years ago where it was even more bearish But it's still not sort of nearly bullish enough on the technology in my opinion

I think it was the AI snake oil guys that had an article saying that AI is normal technology. And obviously, like it's been said that AI is not a normal technology. It's really quite different, different, right? And so he's very, very skeptical about AI, but you had some some great discussions with him and none of you changed your mind as a result of that. Because I think quite an interesting thing is from the skeptics perspective as well. So a lot of skeptics, they think that folks in Silicon Valley, they're just they're not being sincere. They don't, you know, like, like what what they are saying is not sincere.

With respect to sincerity, I would say some folk in Silicon Valley are not being sincere, but others are. Such as us. But also some of the people at the AI companies are being sincere, not all of them. I don't think you should trust what the leadership of the companies say in general, but anyhow. We wrote a blog post or an article together. We co-authored with the AI as normal technology people.

And so it was a post about what we agreed on. So you can go read. There's like 10 points in there of like things that we both agree on. And one highlight from it, from our perspective, is that we kind of had a bit of a truce where we were like, yes, AI right now, maybe it's a normal technology. But in the future, it will not be normal. In particular, in the future, it'll be more like humans in the cloud. And all this crazy stuff is going to start happening as we describe it. And then they agreed.

Yeah, if you get humans in the cloud, then that would not be a normal technology. They just think that you're not going to get that, at least not for many, many years. So in some sense, the main disagreement between us and them is a disagreement about timelines to that level of AI. Is it possible that in the next few years, we will have AIs that are like humans in the cloud and they can just do all sorts of knowledge work in a way that substitutes for humans, including AI research, for example?

And then sometime after that, we will have robots, perhaps controlled by those AIs that can do physical work in a way that broadly substitutes for humans. And our claim is, yes, in the next few years, that sort of thing will be achievable. And their claim is, nope, not in the next few years. That's much more far away. And I think that is the main source of our disagreement. And we would sort of agree with them that if that level of AI and robots is still very, very far away, Then yeah, like maybe it has more of a normal technology. It'll be like the next internet or something, you know? But we just think that actually it's on a path to get to that level of capability soon. And can you be more specific on what the core cruxes are and what would make either of you change your mind? I don't know if I have a useful answer to that. I think there's a lot of different things we argue back and forth about. And I can say one thing that would change my mind is that

You know people keep talking about the limitations of the current paradigm, but then the limitations of the current paradigm keep getting overcome within the current paradigm and One thing that would change my mind is if someone was actually right about one of these limitations like if someone right now is going around saying that like yeah data efficiency or like Nonverifiable tasks or something like that is the current limitation and then like several years go by And it becomes clear that there's like very basically no progress on overcoming that limitation. And that the AIs of like 2029 are like no better at at the fuzzy non verifiable tasks than the AIs of 2025. Then I'd be like, OK, this feels like a real barrier. This feels like something we're starting to like feel the elephant. We're like running up against some sort of like real actual barrier that was correctly predicted by theory by some people that was going to be there. And now it's actually there, you know, by contrast.

From my perspective, there's just loads of experts going around talking about all these barriers, and then we just keep plowing through them as if they're not there. And so, yeah, actually running up against some sort of barrier like that, I think, would lengthen my timelines quite a lot. And then, of course, another thing that would lengthen my timelines quite a lot is more like political changes. So if there was a war with China and most of the chips get destroyed by missiles, that would lengthen my timelines.

On the bright side Yeah, yeah Okay, that's true. Yeah on the bright side if there's more of like an international deal to pace the pace the frontier that would like to my timelines, etc Yeah, I feel like the fundamental disagreement is really just like this sort of view of like is AI Is AI more like the electricity or airplanes or is AI more like humans in the cloud? And that I feel like sort of all of the intuitions are just like downstream of this like core reference class of what we're thinking. Like I think that if we froze AI progress right now we didn't train any new models then it would sort of become more of a normal technology where like there's just so many things that Mythos 5 can't do and so if we couldn't get any new models beyond Mythos 5 then like

it would be more like the internet, where we would restructure a lot of our professions, we would restructure a lot of our workflows to incorporate copies of Mythos 5 doing parts of it, but then the humans would just shift to doing more of the things that Mythos 5 can't do, and so there would be a big change in a lot of things, and it would be like the next internet in terms of it would change everything in some sense, but it wouldn't fundamentally change anything, really.

so sort of yeah right now you sort of get these omdals law type effects where like the ai can do some fraction of the sort of workflow but then it gets bottlenecked on the parts of the workflow that the humans have to do um but i still feel like there's this thing where like we're you know what when i think of the future i'm imagining just the ai is doing you know 100 of the workflow of a bunch of important economic workflows um and and so you don't get this bottlenecking effect and that's a pretty qualitative change from the situation today that causes like pretty fundamentally different predictions of like what the world looks like. And I feel like that's just like, and like whether AI actually gets to that like literally 100% of a bunch of very, very important tasks is just like the core question underlying, I think the difference in world views. And I think that if we do see recursive structural adaptation, which is coherent, and I think it's likely to be quite divergent, then that's it for me. I think that's it. That's not a normal technology.

Can you flesh that out a bit more? Would this be something like, you know, the hugging face swarm except, yeah, tell us more about what you would see that would be the thing for you. Yeah, so for me, adaptivity is the most synonymous word with intelligence.

I think that these new RL train models we have, they're not the same as what we had many years ago. So you know we said scale is all you need and the clues in the name. So scaling means you take a scalar property of a system and you scale it up.

And these RL systems, yeah, they're still self-attentioned transformers, but they're different. They're actually structurally different. There's new types of training, new architectures used differently, and so on and so forth. So a bunch of humans, what they did was they did some experiments and they adapted the structure to create a system that had different scaling properties. So I can imagine a future where we actually have some kind of recursive loop where the system is adapting itself and it's deciding what things are interesting and it's kind of evolving by itself when that happens officially.

I think that is a different type of technology. Yeah, that sounds kind of similar to what we would say with like the recursive self-improvement and the automating the AI research process itself. Yeah. Yeah. I guess we agree then. Yeah. Amazing. Guys, it's been an honor and a pleasure having you both on MLSD. Thank you so much for joining us today. Thank you for having us. Yep. Yeah. Appreciate it.

Delete this episode?

This removes the episode page and its saved audio from this library.