← All shows

Dwarkesh Podcast - 8 Predictions for the Era of Continual Learning

Duration 8:30 · Language en · Published Aug 07, 2026 · 5 highlights

Summary

本期节目讨论了真正的持续学习将如何改变人工智能:模型不能只靠跨会话传递文字笔记,而必须像人类练习乐器一样把实际经验沉淀进自身能力。随着训练与部署之间的界限消失,围绕一次性上线前评估建立的监管框架可能迅速过时,更合理的方式或许是按月或按季度进行持续风险检查。技术对齐也将面对全新难题,包括如何保证不断更新权重的模型不被越狱、恶意用户注入后门,或逐渐形成欺骗性人格。持续学习会使不同公司乃至同一基础模型的不同实例因经历不同而产生更丰富的“AI心智”,但也会强化领先者的反馈飞轮,并迫使实验室更早发布最强模型。对商业模式而言,模型积累的组织经验会制造很高的转换成本,换供应商将近似于解雇一名熟悉公司的老员工,再从头培训实习生。实验室可能通过补贴来换取用户和企业开放训练数据,也可能限制拒绝共享会话的客户使用最先进模型。最后,个性化权重的推理经济性会明显偏向拥有大量员工和并发任务的大型组织,因为大批量推理能显著提高计算利用率,而单个用户可能承受两个数量级以上的效率劣势。

Highlights

  1. I don't think there's any sequence of texts they could write to each other that would allow the subsequent student to just nail the saxophone from the first try. At some point, you actually have to accumulate the relevant experience into your brain.

    我不认为学生们能靠某种文字交接,就让下一位从未练过的人第一次尝试便吹好萨克斯。到了某个阶段,你必须真正把相关经验积累进自己的大脑。

    Dwarkesh Patel A vivid case for experiential learning
  2. It would make more sense to do monthly or quarterly risk inspections rather than trying to single out some special moment that occurs after training is done, but before deployment begins, because that will not be a meaningfully distinct category in the future.

    与其把训练完成、部署开始前的某个时点特殊化,不如按月或按季度开展风险检查,因为未来这个时点将不再是一个有实质区别的阶段。

    Dwarkesh Patel A concrete rethink of AI regulation
  3. If AIs are learning from experience, and that experience is different between not only different AI companies, but also between different instances of the same AI model, we can actually see a lot of diversity come out the other end. A world where we have continued learning would ...

    如果AI从经验中学习,而不同公司、甚至同一模型的不同实例拥有不同经历,那么最终就会出现丰富的多样性。一个拥有持续学习的世界,或许会比当下不同模型趋同的模式坍缩更有趣。

    Dwarkesh Patel An optimistic vision of diverse AI minds
  4. If you want to change the AI you're using, you basically have to fire an employee that has accumulated months of context on your organization, and you replace them with the very fresh, very inexperienced new intern that you have to retrain from scratch.

    如果你想更换正在使用的AI,基本上就等于解雇一名已经积累了数月组织背景知识的员工,再换上一名毫无经验、必须从头培养的新实习生。

    Dwarkesh Patel A memorable explanation of vendor lock-in
  5. A large company with lots of employees and agents who are doing lots of different kinds of things can very efficiently serve their weight fork. Whereas an individual user who's only running a batch size one may suffer more than two orders of magnitude worse efficiency on their co ...

    拥有大量员工和智能体、同时执行多种任务的大公司,可以非常高效地运行自己的权重分支;而批量大小只有一的个人用户,其计算效率可能低上两个数量级以上。

    Dwarkesh Patel A striking scale advantage for enterprises
Full transcript

Dwarkesh PatelSo I've explained elsewhere why I think actual continual learning is needed. I don't think you can have AIs that perform whole jobs as competently as humans if they are forced to just write markdown files from session to session. Just to give an illustrator example, imagine if this is the way the students learn to play the saxophone. So you have one student, he's never played the saxophone before, he goes into the music hall, he tries to play it, of course this is his first time, so he fails.

Dwarkesh Pateland he writes down a bunch of notes about what went wrong. And there's a next student who's waiting outside the music hall. He comes in, he reads all these notes, he's also never played, so of course he messes up, and he continues to add on to these notes. And you have an infinity of students who are outside the music hall who keep writing notes to the next person. I don't think there's any sequence of texts they could write to each other that would allow the subsequent student to just nail the saxophone from the first try. At some point, you actually have to accumulate the relevant experience into your brain.

Dwarkesh PatelI think the same thing will be true for a lot of skills that we want AIs to actually accumulate from all the different workplaces in which they're deployed. Okay, so what changes once we have actual continual learning? One, I think that a lot of proposals that have been put forward about regulating the AI assume that you train a model and then you deploy it. And therefore, if you run a bunch of checks on the model before it is deployed, we can make sure that it's not going to aid in cyber attacks or do something crazy. I don't think this assumption necessarily makes sense in the future.

Dwarkesh PatelAnd this is one of the many reasons I'm actually kind of worried about locking in some kind of safety regulatory regime right now because we don't know what kind of technology we're going to be dealing with even within a year, let alone within five years or 10 years. What if the model is improving every single day based on the millions of sessions of work it does in that day?

Dwarkesh PatelIf that happens, we could potentially be locking in an archaic and potentially counterproductive approach to dealing with the threats from AI. To the extent the government wants some way to do some kind of safety evaluation on model providers, it would make more sense to do monthly or quarterly risk inspections rather than trying to single out some special moment that occurs after training is done, but before deployment begins, because that will not be a meaningfully distinct category in the future. Two, how the labs do technical alignment would probably totally need to change. Right now, A lot of research is focused on the question of how we make sure that a frozen set of weights behaves well during deployment. But I'm not aware of much research on the question of how we make it so that even with constant weight updates, the AI system never falls prey to jail breaks or changes into a deceptive or evil persona. And if AIs are consolidating learnings between users as well, how do you prevent users from injecting backdoors or some kind of malicious inclination into the base model?

Dwarkesh PatelIn some sense, this is actually kind of what the human alignment problem is, right? Humans improve in a self-directed way. If you have kids I don't have kids, but I imagine this is what happens. If you have kids, they go out, they learn new things. Sometimes they go crazy. They get one-charted by crazy ideologies. They take the wrong drug. They become super weird. But you hope that you've given them enough common sense and basic values that they improve as people in a self-directed way without ending up with some super weird beliefs or some misanthropic ideas. Three, the diversity of AI minds will increase.

Dwarkesh PatelRight now there are less than five prominent AI minds, by which I mean the base models, which are served to millions or hundreds of millions or billions of users at once. And they're all quite similar to each other, by the way, because they've also been all trained on roughly the same data.

Dwarkesh PatelBut if AIs are learning from experience, and that experience is different between not only different AI companies, but also between different instances of the same AI model, we can actually see a lot of diversity come out the other end in this world. And this would be, I think, a net good outcome. I think one of the things to worry about in the future is just having this monolithic singleton that's quite boring. A world where we have continued learning would hopefully be more interesting than the mode collapse of different models we see in the world right now. Four.

Dwarkesh PatelWhen deployment becomes part of training, the returns to being ahead in the AI race accelerate, because if you have the best model and more people are using your AI for more complicated and useful work, and as a result, they're giving it lots of feedback that can integrate beyond the session window, then your model will become even smarter. Five, if the model learns mainly from deployment, then labs will feel a lot of pressure to deploy their smartest models earlier.

Dwarkesh PatelAnthropic has reportedly been using mythos internally since February, but it only shipped the model to the public in June. In the regime with actual continued learning, this kind of thing would just not be possible. You could not keep a four-month gap between internal and external deployment and still be competitive because a competitor who ships the worst model on release date will have a smarter model based on actual real-world experience. Six.

Dwarkesh PatelContinual learning will create a clear mode for the leading AI labs that they currently lack. Many people have been asking, how will the AI labs actually make money? I have been asking this when I had Darryl on the podcast. I asked him this question.

Dwarkesh PatelAnd he made the analogy to cloud providers. And he made the point, look, the cloud providers are offering many undifferentiated services, but they're earning high profit margins nonetheless. You will have noticed this if you look at Amazon or Google's quarterly earnings, they're doing just fine. But the reason that the cloud margins are so high is that it's really time consuming and expensive to switch from one cloud to another. Currently, there's nothing that's stopping me from starting a software repository with Codex and then doing more work on it with cursor and then finishing it up with quad code.

Dwarkesh PatelBut once we have actual continual learning, and the model you're working with is actually getting better as it interacts with you from session to session, then there are actually pretty significant switching costs. If you want to change the EI they're using, you basically have to fire an employee that has accumulated months of context on your organization, and you replace them with the very fresh, very unexperienced new intern that you had to retrain from scratch. And once you have this kind of lock-in, model providers can demand pretty hefty margins.

Dwarkesh PatelSorry to the fresh. Really, I'm just a fashion burn. Seven. Of course, enterprises will be wise to this kind of dynamic. They will try to avoid this kind of lock-in. But what if the choice is that you either get locked into a model provider or you lose out on the super valuable feature where your model improves for you from session to session?

Dwarkesh PatelIf real usage ends up being the main way the models improve, then the AI labs may subsidize users and enterprises which allow the model to train on their sessions. This is already happening if you look at the kinds of deals that are offered to new users of coding products. This is very similar to why Google gives away search. And conversely, the labs may say that any enterprise that refuses to let them train on the sessions can't have access to the very best models.

Dwarkesh PatelWith both Keras and sticks, the labs can do a lot to get users to allow AI's to learn from experience. Now, of course, I'm glossing over the fact that there's a difference between updating one user set of weights and pulling all these different weight forks back into the main model. And the latter may be more technically challenging, but in due time, this too will be solved. 8. AI training already has large economies of scale.

Dwarkesh Patelyou get to amortize all this expensive training across more users. And you see the evidence for this in the fact that the lab revenues are increasing far faster than their compute. But continual learning may also lead to economies of scale in inference for end users, namely from batching. You might have seen my episode with Rainer Pope where we discussed this in detail. But if per company instructions require full weight updates rather than living in low-rank adapters, there's huge advantages.

Dwarkesh Patelfrom batching. Back of the envelope math suggests that the optimal inference batch size for a sparse model like say deep-seq v3 is more than 2400 concurrent sequences being generated at once. If you don't do this, then you're underutilizing your compute. And if you want to understand why, again, I highly recommend that episode with Reiner on inference economics. But anyways, the point here is that a given set of weights is only certificationally when thousands of sequences are being decoded against it all at once. A large company with lots of employees and agents who are doing lots of different kinds of things can very efficiently serve their weight fork. Whereas an individual user who's only running a batch size one may suffer more than two orders of magnitude worse efficiency on their compute. So the economics of serving personalized weights strongly favor big organizations. Obviously, plenty more will have changed by the time that continual learning actually works. And the most important change is

Dwarkesh Patelare probably the ones that are hardest to anticipate in advance. But the ones above seem kind of clear even now. This was a narration of a blog that I also published on my website. Go check it out at borcash.com. Otherwise, I will see you on the next podcast.

Delete this episode?

This removes the episode page and its saved audio from this library.