← All shows

Machine Learning Street Talk (MLST) - How Replication Could Teach Machines What Good Science Looks Like — Edward Hughes

Duration 2:01:48 · Language en · Published Sep 11, 2026 · 12 highlights

Summary

本期访谈围绕如何让AI科学家从回答既定问题,迈向主动提出问题并进行开放式创造展开。Ed认为AI本身并无独立意义,真正重要的是它与科学家、工具和组织生态结合后,能否加速跨学科知识发现。节目将创造力理解为在保留意义所必需的约束时,有选择地打破旧约束,并强调创新只有经过个体、知识领域和社会评价共同作用,才会成为被承认的创造。双方进一步讨论开放式探索,主张复杂科学问题无法由单一全局奖励函数预先定义,而应依靠局部目标、好奇心和不断修正的不完美世界模型前进。Ed介绍了Inherent的Replica任务集与Faraday智能体:一个经过强化学习训练的小模型把前沿编码模型当作工具,在跨领域论文复现实验中反而超过了该工具模型及其他前沿智能体。其关键不只是重画论文图表,而是学习严谨实验、动态监督、逐轮信用分配以及识别不可复现结果,并把深度复现视为通往创新的课程起点。访谈最后把递归自我改进提升到公司层面,设想多个人类与多个智能体共同积累文化和能力,并提出组织需要像电气化时代重构工厂一样,发明一种超越传统OKR、适合开放式探索的新结构。

Chapters

  1. 开放式AI科学家的创造力 0:00–1:02:53

    本节介绍了Ed从理论物理学、DeepMind研究到创办Inherent的经历,以及他构建横跨科学领域、与人类共同工作的AI科学家系统的愿景。对话深入区分创新与创造力,提出创造力来自约束满足、适度打破既有约束、社会认可与文化积累,并借AlphaGo第37手、文艺复兴、音乐和生物进化等案例展开说明。双方还讨论了开放式探索的核心:不预设单一全局目标,而以局部好奇心、欠明确问题和事后评价推动发现,同时强调AI成果必须既新颖又能被人类理解和学习。

  2. 法拉第智能体与科研创新 1:02:53–2:01:48

    本节先从科学理论不可避免的先验偏置谈起,讨论模型架构是否需要不断还原到更基础、可压缩的形式。随后重点介绍科研智能体 Faraday:它通过复现论文中被删除的图表来学习实验设计、批判性判断和严谨验证,并借助小模型监督前沿编程模型,在跨领域测试中取得更好表现。对话还深入探讨了逐轮信用分配、避免数据泄漏与科研作弊、神经模型和符号工具的结合,以及从论文复现逐步走向原创发现的路径。最后,嘉宾展望了人类与多智能体协作的“递归型公司”,认为组织结构也需要围绕 AI 科研和开放式创新重新设计。

Highlights

  1. In 2016, towards the end of my PhD, I saw the AlphaGo match. I remember watching that, like so many other people, being fascinated by Move 37. And I became convinced that the future was going to involve agents that could really aid humans in making discovery and accelerate the ra ...

    2016年,在博士即将结束时,我看了AlphaGo的比赛。和许多人一样,我被第37手深深吸引,并由此相信,未来将出现真正帮助人类进行发现、加速科学进步的智能体。

    The origin story behind his AI-scientist mission
  2. Most of the biggest paradigm-shifting discoveries happen when you have knowledge in one area that gets transported to a different area, and then that unlocks some unexpected connections, some unexpected advance. We think we're building a horizontal intelligence layer for all of s ...

    大多数真正改变范式的发现,都发生在某个领域的知识被迁移到另一个领域时,由此解锁意想不到的联系和进展。我们认为自己正在为整个科学体系构建一个横向智能层。

    A compelling thesis for cross-disciplinary discovery
  3. I don't think that Move 37 was creative. I think that Move 37 was innovative without being creative. In my mind, creativity requires another step, which is to recognize that the thing you have done is creative.

    我不认为第37手具有创造力;它是创新的,却并非创造性的。在我看来,创造力还需要多一步:认识到自己所做之事是有创造性的。

    A provocative distinction between innovation and creativity
  4. All that evolution requires is that individuals survive and reproduce. If you really believe that creativity is satisficing, then the default crutch that we reach to in machine learning, which is optimization, is the wrong thing to reach for to build creative agents. The way that ...

    进化所要求的不过是个体能够生存并繁衍。如果创造力的本质是满足约束而非追求最优,那么机器学习默认依赖的优化,就不是构建创造型智能体的正确工具。打开新创造空间的方法,是保留部分现有约束,同时打破其中一些。

    Challenges optimization as the foundation of creativity
  5. An exaptation is an adaptation that was giving the organism some advantage, which then finds a use somewhere else, an unexpected second use. My favorite one as an ML person, of course, is GPUs: developed for gaming, and it just so happens that that's exactly what you need in orde ...

    所谓“外适应”,是指一种原本给生物带来优势的适应,后来在别处获得了意想不到的第二种用途。作为机器学习研究者,我最喜欢的例子当然是GPU:它原本为游戏开发,却恰好成为训练现代神经网络所需要的工具。

    A memorable example of discovery through repurposing
  6. It's that very act of copying that is creative. In order to transmit an idea, whether that's a physical idea or a more advanced idea, for example a cultural norm or a technology, you need to recreate what someone else had in their brain. And that requires an act of creativity on ...

    恰恰是复制这一行为本身具有创造性。要传递一个想法,无论是身体动作、文化规范还是技术,你都必须在自己的头脑中重建他人头脑中的内容,而这本身就要求创造力。

    Reframes imitation as an inherently creative act
  7. Most evaluations in AI at the moment are built in foresight: somebody dreams up a capability they would like the AI system to have, and then they develop some environment and some reward function in advance. If we really want agents that are creative or that behave like scientist ...

    目前大多数AI评测都是预先设计的:人们先设想希望系统具备的能力,再提前构造环境与奖励函数。如果我们真的想要具有创造力或像科学家一样行动的智能体,就必须把评估方式倒过来,不再事先规定它应该做什么,而是在事后审视它实际做出了什么。

    A concrete proposal to evaluate open-ended intelligence
  8. Let's suppose that you have a perfect world model, and you're now out in the world trying to make discoveries. What you'll quickly find out is that whatever you do, all that happens is what you expect. If you want to make a discovery, then you have to have some imperfection in yo ...

    假设你拥有一个完美的世界模型,并试图在世界中进行发现,你很快会发现无论做什么,发生的都只是你已经预料到的事情。因此,如果想要有所发现,你的世界模型就必须存在某种不完美。

    A sharp paradox linking ignorance to discovery
  9. We talk about the idea of an open-ended system, to an observer, having to produce artifacts that are both novel and learnable. There's no benefit in a system producing some incredible discovery that just cannot be parsed by humans. The most advanced AI science systems also need t ...

    我们认为,对观察者而言,一个开放式系统必须产出既新颖又可学习的成果。如果系统做出了惊人的发现,却完全无法被人类理解,那就没有益处。最先进的AI科学系统还必须能够教育人类,或把发现翻译成人类语言。

    Defines human legibility as a limit on scientific progress
  10. What we find is that our Faraday agent is able to perform better than the frontier model. It's able to perform better both than the coding agent that it's using as a tool, but it's also performing better than other frontier coding agents like Claude. We ran a prompt optimization ...

    我们发现Faraday智能体的表现优于前沿模型:它不仅超过了自己作为工具调用的编码智能体,也超过了Claude等其他前沿编码智能体。我们还对GPT-5.5 Codex运行了提示词优化循环,虽然性能略有提升,但Faraday仍保持了相当大的优势。

    The episode's headline empirical result
  11. For the first few months, we had these agents proactively reaching out to us and trying to help us with stuff, and to be perfectly honest with you, it was quite annoying. But about two or three months ago, I think we reached a phase transition where the agents were aware enough o ...

    最初几个月,这些智能体会主动联系我们并试图帮忙,老实说相当烦人。但大约两三个月前,我们似乎跨过了一个相变点:智能体对公司的上下文了解得足够多,也拥有了足够的操作能力,于是它们的主动行为开始真正有用。

    A candid story of agents crossing a usefulness threshold
  12. When the electric dynamo was invented, factory owners were able to replace their big steam-powered turbines with electric dynamos, and this gave a small productivity boost. What unlocked the really extraordinary productivity gains was when people reconfigured the whole factory, p ...

    电动机刚发明时,工厂主只是用它替换大型蒸汽轮机,因此只获得了小幅生产率提升。真正惊人的提升来自人们重新设计整座工厂,在各工位配置独立电动机,并由此发明生产线。如今的问题是:我们如何从头重塑AI研究这座“工厂”,让智能体处于中心?

    A powerful analogy for AI-native organizational redesign
Full transcript

How do we start to build AI scientist systems that go beyond simply answering questions that we pose and start to ask the kinds of questions that lead to open-ended creative discovery? You spoke about Move 37 and I think you and I would agree that that was definitely creative. I don't think that Move 37 was creative. I think that Move 37 was innovative.

without being creative. And what we find is that our Faraday agent is able to perform better than the Frontier model. So it's but it's also performing better than other Frontier coding agents like Claude, for instance, and had these agents proactively reaching out to us and trying to help us with stuff. And to be perfectly honest with you, it was quite annoying. They just had no idea how to help us. But at some point about Two or three months ago, I think we reached a phase transition. What's the equivalent of OKRs for the age of recursive self-improvement? Indeed. And I think that is the next era. It's the era of collective intelligence rather than the era of individual intelligence. This episode is supported by Cyberfund. If you're building at the frontier of AI, they want to hear from you.

Cyberfund believes the future belongs to AI natives who want to achieve the impossible. And that is why they're introducing the monastery for AI native founders. It's an environment of pure focus and rapid execution for founders operating at AI native speed. And they're offering teams $2 million each to participate. Apply now at cyber.fund.

Just a quick piece on inheritance. So you were at Google DeepMind before. How much funding have you got? How many employees? What's the valuation? Yes. So we've raised $50 million from index ventures and radical ventures at a post money valuation of $225 million. Currently, we're a team of 11 people. And presumably you think that AI is the most important thing in the next five years, just generally. No, I wouldn't say that.

I think that AI by itself in some sense is meaningless. Really, it's what can AI enable in interaction with the wider ecosystem? And that for us, we're most interested in broader scientific progress. Ed, it's amazing to have you on MLST. Welcome. Pleasure to be here. Tell us about yourself.

So yes, my background originally I was a theoretical physicist as my PhD and I was studying string theory in particular the scattering of particles but using string theory as a calculation mechanism and one of the central.

theses of the work I was doing was an idea called dualities in physics. Duality is a mathematical map, if you like, between two different theories. And it can take you in some quite unexpected places. So usually it's useful if you have calculations in one theory, which are very hard. You can map across this duality and calculations in another theory that are much easier. And so I was using some of these dualities in order to take calculations that were very difficult and map them into a different geometric space in order to do them. And I think that taught me a lot about how to think about transforming problems, something we'll probably come back to when we talk about creativity. But towards the end of my PhD, I became convinced that working in scattering amplitudes was not the

the most profitable thing I could do in terms of making an impact on the sum total of human knowledge. Funnily enough, many of the people that I was citing at the time of my PhD, people like Jared Kaplan and Jeffrey Pennington, also made a similar decision around about the same time. So in 2016, towards the end of my PhD, I saw the Alpha Gay Match. I remember watching that, like so many other people being fascinated by move 37.

And I became convinced that the future was going to involve agents that could really aid humans in making discovery and accelerate the rate of scientific progress. And so I joined DeepMind in 2017, really animated by that question, how do you build an agent that can itself make discoveries? And I started by working out on that in the context of reinforcement learning, particularly multi-agent reinforcement learning. The reason I was so interested in multi-agent reinforcement learning was because human technology seems to have been created not so much by one individual, but by the sum total of human culture. And I like to think of cultural evolution as the fastest intelligence generating process in the universe. And so I became, I suppose, obsessed by this idea of really trying to distill cultural evolution into agents.

Now, at some point on my journey, the foundation model started to arise. I started to work on building much larger models, so I built a model called Adaptive Agents, which was all about using meta-reinforcement learning across a very, very large space of tasks in order to build what was at that time. I think the largest trained from scratch RL agent, 500 million parameters, so it's pathetic by modern standards, but at the time it was rather large.

And then I worked on world models. So I suppose my claim to fame there was that I wrote from by hand, effectively, before coding agents, the final implementation of the GD1 model in a single file. So I think that was the last time I really wrote a fully artisanal human written implementation of an algorithm.

And then from there, I led a team called AI Scientist and what we were trying to do at DeepMind was apply coding agents to the problem of doing scientific research. And this was in parallel to many of the developments at Sakana and elsewhere.

And so I led that team with Lewis Kersh for just over a year. Lewis had been Jürgen Schmidhuber's PhD student and had come to DeepMind with a very similar interest to mine. But in the middle of last year, we'd become convinced that we needed to build an entirely new company to take this idea into its most ambitious form. And that was really for three reasons. The first reason was that we came to believe that developing an agent capable of discovery wasn't just about a single agent. It was about the way that agent interacted with the entire.

ecosystem around it of scientists. And so what we wanted to do was some fairly radical, organizational design transformations. And to do that in a large established company is much more difficult. That's the kind of thing we're building at inherent. We call it living within the experiment. That's the first reason. The second reason is that we started to think about the way that the major paradigm shifting discoveries come about in science.

And at least from my reading, and also from my personal experience of having shifted areas a few times, I observed that most of the biggest paradigm shifting discoveries happen when you have knowledge in one area that gets transported to a different area, and then that unlocks some unexpected connections, some unexpected advance.

And so that required us to really take a horizontal view of science. And so what we think of we're building at inherent is that we're building a horizontal intelligence layer for all of science. And the third reason that we wanted to do this was really an infrastructural reason. So it turns out that if you want to both reinvent the organization and you want to build this horizontal AI scientist technology. What that implies is that you want to give the AI scientist all the same affordances and context as humans. That's very tricky to do in an established company because you don't want your agent running riot on YouTube if you're Google. But in a new company you can start building the infrastructure from the ground up, not so much

just around sandboxing, but also around permissioning so that you can have agents that really have very equivalent environments in which to operate as humans do on a day to day basis. And so that's really a founding tenet of absolutely everything that we do. I think that creativity has a lot to do with respecting constraints. And I use the word constraints because that feels like the most abstract form of knowledge because we have a privileged form of cognitive knowledge and there's cultural knowledge. And I even think that constraints in the physical world can be thought of as some form of knowledge. So the most abstract possible way to describe creativity is respecting constraints. And what we want to do at different levels is accumulate knowledge, which means we need to find or discover these constraints. And I love using the analogy of a maze. So when we are discovering the shape of a problem, what we're doing is we're kind of discovering the walls.

in the maze. In our evolution, it happens at multiple levels. So there's this kind of DNA phylogenetic evolution in the course of our individual lifetimes. We have this ontogenic evolution. I'm overloading Lamarck a little bit here because he was talking about it in terms of heritability as well. But you see the analogy with AI. So we adapt the weights and then our individual agents get experience and they adapt their skill surface and their memory systems.

And then it feels like there's a third wave, which is the fastest form of knowledge accumulation, which is cultural accumulation. So in the future, the agents themselves will be talking with each other, collaborating or maybe colluding and building up this latent cultural knowledge. So, I mean, do you think, you know, just before we kick off, is that a reasonable model to understand how AI is going to progress? Yes, I think there's a lot of that I agree with. Maybe I'll unpack it backwards. So I think the one wrinkle I'd add to that at the cultural level is I don't believe that there will be some separate culture for agents and a separate culture of humans. In fact, I think the super intelligence in so far as it will exist will be a combination of humans and agents interacting in very deep and very complex ways. And that knowledge will be that the

the better that we're able to interconnect agents and humans and leverage their complementarities, the faster we'll be able to drive this accumulation process of knowledge. Now, I think that absolutely that there is this distinction between if you like the slow weight updates, which are maybe more akin to the evolutionary process for DNA and then the fast in context updates. I suppose one of the things I'm very interested in at a an architectural level is whether this analogy will still hold in five years time, for instance. Now with a lot of pieces, if you like, that could sit in between those two, epigenetic effects, for example, are one. And what's the analogy for epigenetic effects in an AI system? So one of the pieces in our

New paper that we talk about is this idea of a weaker coding agent using a stronger coding agent as a tool. And then what we do is we actually update the weights of the weaker coding agent because that has the benefit of generalization more so than just updating the context. But because the weaker coding agent is using the stronger coding agent, that weaker coding agent is of course injecting things into the context of the stronger coding agent. So now you have to ask yourself, well, is this weights or is this context? And the answer is, of course, it's both. And perhaps we've got the opportunity to have a much more intricate spectrum between these two things than we currently have. And that's one of the themes that we're in particular investigating because in the

cobbled together nature of our current AI stack. What happens is you get these abstractions that become very sticky, rightly so because they work well. But because of the burden of knowledge for humans to understand the frontier, we just accept a very large number of these abstractions because it's just too complicated for us to be examining all of them in combination.

And the promise, I think, of AI scientist systems, agents that really understand this horizontal, is that they might be able to weaken multiple constraints at once and thereby enhance the ability for creativity. Another loop that you opened earlier was, you know, you spoke about Move 37 and I think you and I would agree that that was definitely creative.

It feels like there are some limitations to its creativity. So I would call it a form of concrete creativity. So it doesn't understand in the sense that Margaret Bowden would speak about in terms of understanding how it hangs together in the context of the system and what is possible and counterfactuals and whatnot. But it's still a form of concrete understanding. But another interesting angle there as well is you were talking about the difference between possibly human knowledge and AI knowledge, because I was speaking with Tom McGrathick.

good fire, he did interpretability on AlphaZero. And his idea is very much that these things are learning the space of human concepts and beyond. And we could actually mine those representations as a new form of science. So these things are discovering interesting things that perhaps we would discover but haven't discovered yet. And we could actually use this as a laboratory for discovering interesting new knowledge. Well, interestingly, I don't think that move 37 was creative.

I think that move 37 was innovative without being creative. So the way that I think about innovation is innovation is the process of taking unknown unknowns and making them into known knowns. And in order to get an innovation, you can't generate an innovation if you sort of already knew what it was you were looking for. You have to have something that's unexpected, but then it also becomes valuable. So what's the difference then between innovation and creativity? Well, in my mind, creativity requires another step, which is to recognize that the thing you have done is creative. And who was it who recognized that move 37 was a remarkable move? It wasn't the it wasn't AlphaGo. It was the commentators, for example, who if you watch the famous footage say, oh, that must have been a clicker. Was it a click miss? It wasn't, of course, it was it was exactly the right move.

And so that is a kind of meta, meta cognition. So you have to sort of understand that you didn't know something and now you are, now you're updating your own knowledge as a function of that. And it's, it's interestingly discussed by Mahali, Dick St. Mahali. He wrote a lovely book called The Psychology of Creativity and Invention. And he's also the person behind flow. So many of your listeners will already know a concept from him.

But in this book, he interviews a very large number of different creative people from the different disciplines. And he comes up with an ontology of what creativity is. And he says creativity has got three components. There is the creative individual. That's the bit that we always focus on. But that's really in some sense, the tip of the iceberg. The second piece is the domain.

And the domain is a set of symbolic rules, if you like, to which the creative person is adding or perhaps breaking one of the rules and then thereby expanding the space of possibility. But the third piece, which is perhaps the most forgotten one, is the field. And the field are the set of other individuals who are going to decide whether the creative person's contribution gets admitted into the domain.

And he gives this lovely example of Florence in the Renaissance. So we're talking 15th century. And in the 15th century in Florence, there was an enormous flowering of creativity, whether it was architecture, science, even the way that society itself was structured.

And the question is, what was it that led to that influence? Now, you might say perhaps it was just an expansion, the number of creative individuals. Was there some mutation of the DNA, some new educational system? It seems quite unlikely that there could be a mutation in the DNA. And so far as we know, there was no great change in education. Well, then you have to ask, okay, was it the domain?

And I think at least in part it was the domain because at that time, many building techniques that had been lost to antiquity, which were in fact known to the Greeks and Romans, were being rediscovered via archaeological means, via analysis of the building structures that people were uncovering. But it couldn't just have been the domain because much of this rediscovery was happening in Rome.

And Rome didn't have the same flowering of creativity as Florence. So the third thing you need is the field. And what Florence had that the other Italian cities didn't, were lots of very rich families. The Medici is the most famous among them, but I didn't believe it was the richest. And they were rich from the wool trade and then also from becoming financiers as well. And they had this idea of making Florence the most beautiful and most cultured city, and that was in some sense to weave a protective cloak around the city at a time when there were many city-states, there was quite a lot of conflict, and they believed in this idea of beauty as in some sense a kind of psychological defence. And as a result, there were a very large number of creative

uh, constructions that were admitted into the domain and it became a competition between the artisans of the day. And so I think it's instructive to think about how that might play out with AI scientist systems. And in particular, I think it's, um, it becomes much more interesting when these AI scientist systems start to be able to do, uh, things which are generalizable. Um, so what do I mean by that? Well, of course move 37 remarkable innovation.

But it doesn't really tell you how to do innovation in other domains. We didn't immediately see a line from move 37 to discovering a new material, for example. And even if you think within a single organization, the line between move 37 and say alpha fold wasn't a particularly direct line, it's not like the AlphaGo agent or indeed the AlphaGo training techniques really informed alpha fold.

in a very direct way. But in principle, if you had had a generalizable discovery engine, then it could make a discovery about the weather and then figure out, oh, there's some part of that discovery. Perhaps it's the architecture of the neural network that was used in those to make that model. I wonder whether that applies to protein design.

And those kinds of connections, I think, generalizable connections are going to be what leads to a large acceleration in the rate at which we can make discoveries. Yes. I mean, because you were discussing how we recognize creativity and the social component is extremely vexed because it's very tempting to think there's some degree of social proof in creativity. And indeed, perhaps there is. I mean, there's a famous example of a urinal with a bit of masking tape on and everyone just decided that it was creative.

I tried to think about it abstractly. So for me, something is creative when it becomes a mode in the state space to a certain extent. So that clearly became a social mode. I mean, what's difficult about.

us as individuals and cultural learning is the introduction of agency and the fact that we could have done differently. But I suppose going all the way down to physical creativity, you know, evolution isn't an agent. It's not doing planning. But there are still these canalized modes. And the way I think about it is it's a bit like the system has discovered an interesting new subspace. And that subspace is being used in a myriad of situations.

So we would call that discovery creative. And perhaps even with AlphaZero, maybe if we enumerated many possible game trajectories, and if we saw something that looked like a category. So this particular type of pattern was being rediscovered, reused in many different situations. We could immediately look at it and just draw a boundary around that category. But maybe AlphaZero would kind of have competence without comprehension. So if it was using this thing in many different situations, maybe then we would call it creative.

Yes, I think I want to come back to the idea of relationship between creativity and constraints for a moment. So if we look right back at evolution itself, I think evolution quite clearly is creative. It certainly generated this enormous amount of diversity in the natural world.

And the way it's done so is exactly by satisfying. Satisfying is just a posh word for saying satisfying constraints. So why do I say it's satisfying rather than optimizing? Many people might think it is an evolution trying to optimize for the best, the best of an individual with the most adaptive traits. Well, actually all that evolution requires is that individuals survive and reproduce. And once you've done that, There's not a lot else that you can do. Now, perhaps you could say you can do second-order survival and reproduction. That is true. You probably care about your children surviving and reproducing as well. So it's not quite as simple as that. But even at that second order, that's still a constraint satisfaction problem. And there's two interesting implications of this. The first one is, if you really believe that creativity is satisfying, then the default

crutch that we reach to in machine learning, which is optimization is the wrong thing to reach for to build creative agents. And secondly, if you believe that constraints, satisfying is important for creativity, then the way that you open up new creative spaces is that you take your existing constraints and you break some of them. Now, why do I say break some of them?

Clearly, if you break all the constraints, then there's no meaning left. The way that we construct meaning is, and indeed the way that we construct laws of the universe, is that we rely on things being repeatable.

This is the so-called principle of induction, which is not something that you can prove, but it's something that we just observe. The laws of the universe seem to stay the same from moment to moment. So you can't break every single constraint, otherwise we'd live in a world of white noise. But breaking some of the constraints is very useful because at least some of the constraints arise because of our existing theories. We don't have access to the universe. We only have access to the universe.

through our observations and measurements of it. And in order to make those observations and measurements, we do two things. We have tools that allow us to make those measurements. And then we have our own neural apparatus, which allows us to make interpretations of those. And so by relaxing or breaking some of those constraints about the interpretation or about the tools, where they then able to access new insights about the universe. And so that's where creativity really arises. Yes. You've opened so many loops there. I don't know how to close them off in order, but I'll try my best. The thing that you just said is very interesting, which is this very vexed issue of coherence, right? I mean, atonal harmony is the great example of this. And I often argue with my co-author on this article that we wrote, whether that is breaking the constraints or inverting them or just respecting them in some other way.

because a lot of people talk about knowledge being quite situated and what they're meaning in that case is that it's only coherent if you if you respect the constraints and sometimes it's not possible to break the constraints.

And another thing you spoke about, and this is also related to your 2024 ICML paper, which was open-ended. This is, I think, necessary or required for air. Yes. Beautiful paper, by the way. And I said to Tim Ruck-Taschel at the time that I felt that was actually a definition of creativity rather than open-endedness. Actually, I think they're basically the same thing. And the reason I think they're the same thing is that intelligence is basically about optimization. Right? So intelligence is like, you know, I I'm trying to find the shape of the maze and I don't know it's full shape yet and I'm trying to fill it in and I can go in that direction. Creativity, as you were saying before, it's about discovering new questions, new problems, new mazes. And as Kenneth said in his book, Why Greatness Cannot Be Planned, there's a weird paradox there that when you optimize towards something, it's really, really difficult for you to find something interesting and creative because you've got the blinkers on.

Yes, well, I think let me give two examples that pertain to that description. So one of them is rather beautiful concept in evolutionary biology of X adaptation. So we know, of course, that adaptations.

persist across evolutionary time because they give some advantage to the individual, which allows them to be selected for perhaps they're better able to escape from predators, perhaps they're better able to find food, for example. So what do I mean by an exaptation? Well, an exaptation is an adaptation that was giving the organism some advantage, which then finds a use somewhere else, an unexpected second use.

And indeed, we see this in biological evolution, but we also see this all the time in the famous discoveries of science. So whether it's Alexander Fleming and Penicillin by leaving the Petri dish out, whether it's the invention of the microwave during a radar testing, where I think the individual in question left the chocolate bar in his pocket so that it melted. My favorite one as an ML person, of course, is GPUs. GPUs.

developed for gaming. And it just so happens that that's exactly what you need in order to, well, first of all, optimize the training of convolutional networks, but then now, of course, adapted for the optimization of all modern neural networks. So that's one example. The second example I want to give is completely different, is coming back to the musical example. So as you know, I have this sort of moonlighting career as a semi-professional musician.

And I love that example you gave of atonal music. There's perhaps an even sharper one, which is the famous Tristan chord. You can go and look this up on Wikipedia if you don't know about it. But it's the chord that Wagner used right at the very start of his opera Tristan and his older. And it's the very first chord in the piece. The piece starts with three individual notes of melody and then this chord.

And it's seen as a very creative chord. And it's really interesting to inspect why that is. Now the chord that he uses is not new. This chord, you can go and see this chord back hundreds of years before the same chord was used. But the thing that was new is the context, is how it was situated to use your term.

The context is that the chord is very harmonically ambiguous. You're not at a point where you've yet established the key of the piece. And so as a listener, you immediately question, okay, well, what is this chord saying? And in general, up until that point in musical history, harmony had in some sense been used as an accompaniment to melody.

But at this point Wagner is questioning that and he's asking the question well what if rather than using harmony as accompaniment I use harmony as communication directly and so the thing he's trying to communicate in this chord is.

exactly that ambiguity what what is going to happen suspense perhaps this this sense of confusion or impending chaos but also a slight sense of hope as well there are many things that are happening in that chord and this really prefigures a lot of musical developments in the 20th and indeed 21st century where harmony is used.

for color, it's used for emotion. And in fact, we're all intimately familiar with this because this is used to incredibly great success in film music. you can immediately identify just by the very first couple of seconds of a chord or a harmonic sound world at the start of a scene, even before you've heard, you know, 30 seconds of melody, that this is going to be a rather chilling scene or rather hopeful scene or a love scene, for example. And so I think that is a sort of microcosm of the point you are making, which is that you really creativity has to be judged by standing on the shoulders of giants has to be judge situated in the place that is currently in the cannon yes again absolutely fascinating and.

The way I interpret this is you're pointing to let's say if I edit a video or I make some music or something like that. You're saying in principle it could be quite ambiguous and then it'll be interpreted using the constraints of observers now the observer thing is very important because in your 2024 paper you're talking about the perspective of an observer whether a stream of an events produced by an open-ended system is novel and learnable and there's a kind of a virtuous complexity gradient that we can climb.

I still think that. coherence is a binary property. So when artists create things, they usually have a set of constraints that guides its creation. It could be an intention, you know, like when Michelangelo was painting the Sistine Chapel, there was lots of cultural constraints and what he intended to do at the time. But the observer relative thing is interesting because let's say, you know, you're a very clever person and you write some mathematics and you show it to someone and they don't understand it. It looks like slop to them because they can't recognize the constraints that guided the process. But it's not

slot because I think it's objectively coherent. They just don't understand it yet. But still, when you have this kind of cultural transmission, this is a great form of new adaptivity because it will be reimagined and reinterpreted in a different context. Yes. And I think that actually a lot of creativity does arise from that underspecification. I think it's one of the rather wonderful features of humans is that we can't really transmit our ideas to each other. We have this very high noise, very narrow bottleneck channel, which is our...

description of things in words to try and communicate an incredibly high dimensional state space in our brains and in some ways in our bodies as well. Athletes, for example, come to mind. And I want to just come back for a moment to the book by David Deutsch, the beginning of Infinity. Right there, the beginning of Infinity. What was his definition of science? Well, okay, let me do the definition of science and then I'll come back to this point also of replication. So he says that science is a search for good explanations about the universe. And he's very precise about what he means by a good explanation. He says, a good explanation is one that is hard to vary. So let me give you an example. Let's suppose that you say that the sun rises every morning because it is pulled on a chariot by the gods.

And let's suppose that over time the sun is rising later and later every morning. For example, that happens in the northern hemisphere as we go from summer towards winter. Now there are a number of different ways of varying that explanation. Maybe the gods are getting more tired. They're sleeping in so the sun is rising later. Perhaps the gods are angry and that's the reason why. Now let's suppose that one day the sun doesn't rise at all.

for, say, an hour. And maybe this is something we'd explain as an eclipse. But it can now be explained as the gods either inflicting wrath or the gods giving you a chance to sleep in. Perhaps the gods are inflicting their favor upon the world. Now, suppose instead that you try to explain that the diurnal cycle, 24-hour a day cycle of day and night, by the fact that the earth is spinning on its axis.

And now, let's suppose that you have to explain that the sun is rising later every day. Well, the most natural thing to say then is, okay, well, perhaps the Earth is spinning a bit slower to make the sunrise later, but then hang on.

process, the diurnal cycle is still 24 hours. So how can it be spinning slower but also have a cycle of the same period? So now you see you're forced into a more creative space and you're forced into maybe suggesting that the Earth is not only spinning but it's tilted on its axis and it is orbiting the Sun.

And now, let's suppose that you have the solar eclipse phenomenon. Well, that's pretty odd because you can't just sort of spin the earth into the place where the sun disappears and then spin it back again. That seems like that would require an enormous feat of celestial engineering, if you like. So you then have to posit some other body, in this case the moon, that comes between the sun and the earth. But that body has to come between the sun and the earth at a particular time. And with the regularity that It is consistent with the rest of the theory. So that's what he means by good explanation. Let me just come on to this other point, though, about how he thinks about cultural transmission. And this is really buried quite late in the book. And I think it's a really beautiful account of how creativity arises. So he talks about two mysteries. The first mystery is a very prosaic one. If I stick my hand in the air like this,

then I can ask you to copy that, and you will be able to copy that. And that is an incredibly cognitively difficult thing to do. Why? Because you have got a visual cue of me sticking my hand up, but you haven't got any of my proprioception. You certainly don't have any information about my muscles or indeed what I did with my neural secretary in order to do that. And you've got to reproduce that within yourself. That's one mystery. Mystery two.

is the mystery that around somewhere between about 10,000 and 4,000 years ago, there was this real explosion in the creativity of humans, at least as measured by the archaeological record of the density of different types of technology. Now, that's not to say that there wasn't creativity before. We know cave paintings go back a lot further than that. But certainly in terms of the sheer variety and accumulation of these technologies, something special that seems to happen around that time period.

Now, even though that time period is a few thousand years, it's certainly not long enough for biological evolution to have done much. So something, arguably, the biological prerequisites must have already been present in our brain. So how was it that biological prerequisites were present in our brain, but they weren't being used for anything? What was it they had adapted for? And he rather beautifully solves both problems at once.

And his solution is that it's that very active copying that is creative. In order to transmit an idea, whether that's a physical idea or a more advanced idea, for example, a cultural norm or a technology, you need to recreate what someone else had in their brain. And that requires an act of creativity on your part. And what changes in that six thousand year period is not really much about the individuals. It's something about the field.

Suddenly the leaders and the societies of the time come to value people who accept that creativity, not just for copying, but for doing new things. And so this is a lot of the reason why.

at inherent, we're starting with the idea of replication and thinking of that. Yes. And we will probably get into your paper just in a short while. But to push back on the copying thing, so the canonical example of bad shallow replication is a photocopy or let's say a forger.

So a forger can just make the right colors and the brushstrokes and so on. But all of the inner structure, the abstract structure, the intentions, the motivation, the constraints are absent. And I should bring in Michael Tomasello, because we're interviewing him in a few weeks. And he said, human cumulative culture depends on shared intentionality, teaching, normativity, and ratcheting, not just copying.

This is really interesting because I think what you're saying is that we can do this kind of imitation learning, but what we actually need to do is recreate, because as Kenny of Stanley said, it's not about where you end up, it's about how you got there. So the challenge is to recreate the path which led there. And I think, you know, let's say ancient humans, they painted on the inside of caves and stuff like that. And what made it learnable and cognisable was the fact that we have the same physiology. You know, we have the same structure of the brain, same affordances and whatnot. So maybe that was an easier problem to recreate that abstract structure than say an artificial intelligence, which is learning on more surface level data. In some ways, although I think in some ways that it's almost harder in artificial intelligence because there aren't as many constraints. So what do I mean by that? I think

The wonderful thing about humans copying each other is that we don't have access to most of the information. It's a very, very partially observed setting. So I can't see your neurons, but I can take the very small number of bits of information you give me and reconstruct at least some of what your intention is. And it's that.

bottleneck, and the fact that we have also constrained physiology, that means that copying, in my view, actually begets many of these other downstream facets that Michael Thomas said I talked about.

I don't claim that, in fact, it has to be just unidirectional. I think actually very likely there was some sort of ratchet copying is part of the mix. There's wonderful work by people like Cecilia Hayes, for example, who talk about the same equipment for social learning being actually what you need for a social learning and the interaction between the two being very important. But to come back to your idea of the photocopy, quite clearly, the photocopy is an anti-pattern. Now, we can't Fortunately, we can't photocopy humans and we can't photocopy human ideas. And it's that which has led to creativity. Unfortunately, in the case of AI, we actually can photocopy the weights of a model. You can't do that with a closed weights model, but you can with an open weights model. And that lack of constraint actually makes it harder to arrive at creativity. So I think of a lot of what my job is and I think increasingly to some extent as AI becomes

more spread in society, this will become a more common role for humans to play is as a constraint engineer. What is it that we need as the interfaces between these systems that will promote novelty and creativity? Or in other words, how do we need to regularize away from purely generating facsimiles into generating that's much more complex series of social and technological interconnections that lead to some of these things like invention and normativity and shared intentionality between humans and machines. This brings me on to another thing as well. So a lot has been spoken about functionalism, for example, which is that, you know, we're building machines and we say that if they have the same abstract functions, then essentially they're the same as.

as us doing the same thing with our physical instantiation. But when you look at evolution... Interesting kind of Prometheus moments happen so you know that there's the emergence of language and this copying cultural accumulation that you just spoke to and something fascinating is happening now which is that we are training these foundation models we're doing some RL post tuning and then they exhibit different forms of intelligence and agency in different configuration so they weren't trained to do this but we can now create a society of agents.

And they just have this kind of phenomenon that wasn't part of their evolution. I mean, another example of this is I could take a herd of lions, for example. Every individual lion is not too incomplete. It's not intelligent in the way that we are. But you could imagine a configuration of lions that had more intelligence and more capability than all of them as individuals.

And don't you think this almost goes against the path dependence idea because, you know, now we're seeing a phase change, like an emergence of new capability, new intelligence, new agency, that none of the individuals were evolved to exhibit? Well, I think that there are latent capabilities in the ways these models are trained. I mean, if you I sometimes like to think about this through the lens of something I call the strong Morovec paradox.

So Moravex Paradox is this idea that things that seem complicated for humans to do, like playing chess and go, turn out to actually be relatively easy for AI, or at least we kind of figure out how to get AI to do those things earlier. Things that seem quite easy for a human to do, like making a cup of coffee, is still sort of way out of the realm of possibility for modern robotics to do that reliably in a new kitchen, for example.

And I think it's a strong version of that, which says actually the things which are right at the tip of our cultural evolutionary tree, right at the tip of knowledge, things like solving protein folding or weather prediction or materials design.

are going to be the first things that we figure out how to use AI for really effectively. Things very early on in the evolutionary tree, like the origin of life, symbiogenesis, for example, or autocatalytic reactions, I think that's almost going to be the latest thing that we figure out. And so how have we got to these systems which do exhibit emergence and I think you're right they do well we just trained on all of human cultural knowledge and it turns out that once you've encoded that in the internet then you do get this measure of generalization and indeed you that that's not just generalization of knowledge it's also generalization of the ability to do things and I think that the same trick can be played as we start to build

larger and larger spaces of environments in which to train these agents. And it's no surprise that the place that is working most effectively is coding agents using command line interfaces because it's relatively easy to synthetically generate a very large number of these different environments. If I may come on to one more point, which is around evaluation, and where's the kind of boundary of this, if you like, where is the frontier at the moment?

And I increasingly believe the frontier is in how we evaluate. So most evaluations in AI at the moment are built in foresight. Somebody dreams up a capability they would like the AI system to have. And then they develop some environment and some reward function, which they code in advance. And that's typically what we call a verifiable reward. So it's something that If the agent produces a behavior or an output that's desired, there is a fixed procedure that can run in order to validate whether that works or not. Now, that is very, very good at generating agents that can fulfill the kinds of goals that a human might want to set. But it's not very good at training agents to come up with their own goals, to ask questions rather than answer them.

And so if we really want agents that are creative or that behave like scientists, we have to flip around evaluation so that we're not presupposing in foresight what it is that we expect them to do. We're instead looking in hindsight at what they've done. And then we are judging it either as a human or as an individual agent or as a set of agents. And that's much more like how we would judge something like a PhD.

It would be patently absurd for a PhD advisor to come in, say to their PhD student, on their first day. I've written down a set of three questions. And after four years, I'm going to ask you these three questions, and you're going to tell me the answers. And if you get them right, I will give you a PhD. Rather that we have a system whereby after four years, there's a Viva and the PhD student presents their work, and that's then evaluated by a group of their peers. Yes, indeed. I suppose another interesting question is, I mean, you know, folks like Kenneth Stanley and Jeff Clean have long spoken about this, in addition to your work as well.

There's something interesting about open-endedness, which, roughly speaking, is rather than trying to solve known problems, you almost flip it on its head and it becomes about the discovery of problems. And maybe we can use the word question here as something analogous. But it feels to me that when you are able to ask a question, It seems all but solvable. It's just a matter of computation. It feels like being able to ask a question means that you already have one step in that epistemic phylogeny. And then it becomes almost like a search problem from there. Would you agree with that? Well, I think that questions exist at varying different levels of specification.

On at them at the most concrete level, if you like, there are questions where when you ask them, you can specify a procedure for knowing whether the answer is right or wrong. So if you think about formal maths with the lean prover, for example, that allows you to specify a conjecture and also the lean solver will compile a an attempted proof. And if that compiles, then you know that under the assumptions and the existing theorems that are within that setting, this thing is true according to the system you have set up. And so that's the kind of deepest level of specification. There's also questions that are very, very underspecified. So one example that of course we all care about is how should we solve climate change?

Now, I can't specify, nobody can specify a procedure. If someone came up with a proposal and said, this is how we should solve it, it's the following five steps. And in fact, even if someone came up with a procedure which exactly specified everybody in the world should do for the next 10 years, still it would be impossible to decide a priori how to evaluate the quality of that procedure. And so I think that The way that open-endedness sees the world is that you can't come up with these concrete problems in advance. And there's then a couple of things that you can do. One thing you can do is you say, okay, we're going to hop around between different sorts of concrete problems. And that's the kind of thing that map elites from Jeff and others does very well.

And another thing that you can do is you can say, okay, we're going to create a curriculum of under specification. And that is much more understudied. Part of the reason it's more understudied is that before language models, it wasn't really clear how you would even tackle a curriculum of under specification. But just in recent years, we've had work like Omni and Omni Epic from Jeff's group.

which start to use language models as these models of interestingness. And suddenly that allows us to flip from questions which have to have a precise specification in code, for instance, to questions which can be really quite underspecified. And that's exactly the kinds of direction that we're taking of building this curriculum around the specification is inherent.

Yes, I spoke with Jeff about that. That was Jenny, I think. Yes, indeed. Jenny's wonderful. And even that, the way I kind of think about that is, you know, by Kenneth, he said they have these fractured entangled representations, and we can actually come up with systems to kind of use the fact that they are better at discriminating than generating. So we can almost come up with these loops to sort of iteratively discriminate to produce better generators so that we can evolve in different directions. But I just want to do a quick definitional thing, which is there's a bit of a vexed issue of what open-endedness is. And to me, roughly speaking, it's when you don't know where you're going.

But you just gave the example of climate change, which actually seems like we do know where we want to go. We just want the global temperatures to go down. But that feels like open-ended, because the sheer space of complexity just getting there is very large. So there's this canonical version of open-endedness, which is that the goal space is unknown. And then there's this domain of intelligence, where we are allowed to come up with intermediate sub-problems. But that space could potentially be very complex as well.

Yes. So I think that the statement that I agree with in the way that Ken and Joel Lehman phrase Open-endedness is that you cannot have a single global goal. Now, that doesn't necessarily mean that you can't have local goals or indeed partial goals that contribute to that. So in the case of climate change, you came up with one plausible goal, which is make global temperatures go down.

Clearly, I can set that up as a straw man because then you'd have to specify, but go down by how much. But also, it's not even clear that even if you were to satisfy that on average, would that even be what we wanted? Perhaps if you satisfied that on average by making some part of the world far, far colder, that would not be what we want to achieve. So you start to realize that for these very complicated underspecified problems, there isn't really a single.

a single reward function that you can specify in advance. And that was exactly what Jimmy Sekretan and Ken and others showed in the Pick Breeder experiment that actually if you want to arrive at these creative outputs from a system that has some kind of representational constraint, then it's much better to follow your local curiosity.

Now following your local curiosity is itself following a goal. There's nothing wrong with local goals. It's not a sort of free-for-all and it's not incoherent. Importantly, the people in Pickbreeder were not all drunk. And I claim that if you in fact had got people doing Pickbreeder and all they'd be doing was just clicking the screen at random, you would not have found this interesting behavior. They were in fact doing something which...

has got some internal coherence to it. But importantly, it's not guided by a global goal. And that's the distinction that I think is most important. Yeah, I agree with that. I mean, I think Kenneth would say that if the goal is complex or ambitious, it's likely to be deceptive, which basically means it's underspecified because it does sound sometimes like he's saying there's no planning and no goals at all. But you know, local goals are well understood. And also he's not saying it's like a thousand monkey, monkey's experiment where all the monkeys are going in random directions.

all of those agents are following, he calls it their own path of interestingness. But what that means is they're respecting their own constraints, and their constraints could actually be very deep and structured. They could be domain experts. And the deceptive point is, I think, a really deep one. And it comes back to world modeling, actually, in my view. So if you want an agent to be able to make a scientific discovery, it has to operate in a space where the goal is deceptive. Why is that the case? Well, if you think about an agent that's got a world model, and a world model just to remind people what that means, that means that it's a forward model, action conditioned forward model of the world. So if I take this action in this state, what will happen as I roll that out? And arguably that's also what

really any scientific theory does. It tells you, okay, as a function of this state of the world, when I introduce this perturbation, what happens to the state of the world? Now, let's suppose that you have a perfect world model, and you're now out in the world, you're trying to make discoveries. Well, what you'll quickly find out is that whatever you do, all that happens is what you expect.

And having a perfect world model is what would allow you to know whether the goal was good, right? So on the flip side, if you want to make a discovery, then you have to have some imperfection in your world model. And as a result of that, the goal has to be deceptive because you have to get to some point where you think, okay, I thought that the way out of this maze was over here.

But now I realize I've been laboring under a misapprehension about this local goal that I've picked, and that what seemed to be moving towards the light, what seemed to be a good idea, is now not really working out for me. And then you update your world model and you realize, OK, there is some light source that somebody has put there adversarily in order to confuse you, for instance, to take your maze analogy.

And so I think that there's something very deeply connected between the idea of open-endedness and also the idea of building world models. In particular, what I think of now in the realm of science as experimental world models. So a world model for what will happen if you carry out some new experiment. Yes. And I suppose in both cases, we're talking about an epistemic gap. So, you know, if it's deceptive, there's an epistemic gap, but there's also a more virtuous kind of gap, which is actually respecting a lot of structures we already have.

But when we look at the unit distance disproof on the recent GPT model, what we find at the moment is that the models are kind of navigating spaghetti space. So it's incredibly verbose. But even if the opposite were true, even if they were using very high level abstractions that mathematicians use, would that necessarily be better? Do you think that there are, because what we're talking about here is we're shining a flashlight into ideas space?

And we could have a very high aspect ratio and we could just traverse the spaghetti or we could do what we do and we could just traverse the very high level of abstractions. Which one of those two extremes is better? Honestly, I've got no idea. And I think it's a fascinating topic for future discussion and research.

It's almost like asking which language is better, either which human language or which programming language. And the answer is, well, none of them. But it's certainly the case that in certain languages, certain things are more compressible and in certain languages, other things are more compressible. And arguably, there is some interaction between language and culture that leads to different kinds of creativity. And that's why it's so important that we preserve.

different languages and we preserve that diversity. So that I expect that there will be a period of time where we will have very different ways of AI solving problems to humans. I hope that that will persist, in fact, because I think that will lead us to different kinds of creativity from different kinds of constraints being broken. But then of course, what you need if you're going to have that kind of approach as you need a translation layer. And so this is why when we talk about the definition of open-endedness in the paper that I wrote with Michael at ICML a couple of years ago, we talk about the idea of an open-ended system to an observer having to produce artifacts that are both novel and learnable. And it's really that learnable piece, which is to this translation point. There's no

benefit in a system producing some incredible discovery that just cannot be parsed by humans. And what's really interesting is how few people work on pushing the boundaries of Go. Now Go is a two player zero sum game. There is a Nash equilibrium, which means there is a perfect way of playing Go. And we're fairly sure we haven't found that yet.

Running an alpha go like algorithm with more and more compute you're going to get better and better at go but we're already so far beyond human go playing capability that it's not interesting to humans because it's not learnable and I think that in some ways is going to put an interesting friction on the rate at which we can make discoveries that's going to.

necessitate the most advanced AI science systems also being able to educate or translate into human language. Yeah, and a great example of that was the Kepler's conjecture so when Thomas hails you know he did hundreds of thousands of dynamic programming problems and. The annals of mathematics couldn't verify whether it actually solved the problem or not but it was a it was not very nice from from a sort of it wasn't very intellectually satisfying but it feels like there is a step towards crystallization.

right so i think many times we do some initial adaptation we prove that something is possible and then we crystallize it down we find the abstractions and maybe that's the kind of a i we need so we need to start in the in the bigger space and then we crystallize and recreate legible abstractions so i suppose i'm saying in the case of go do you think that's even conceivable do you think it is compressible in a way that would be legible to us. It's a very good question and part of me.

thinks there are things which are very hard to compress. Now, we've been on a very good philosophical run with the philosophy of reductionism. It served us incredibly well for 400 years, or arguably goes back even further to that to William of Occam and the idea of Occam's razor, take the simplest possible explanation if you have nothing else to distinguish between the explanations. But there is a school of thought that believes that reductionism may not be the be all and end all. And actually this comes, one illustration of this comes back to my theoretical physics roots and the problem called naturalness. So in the standard model, there are a large number of dimensionless constants, not a very huge number, but enough to wonder what values these should be tuned to. And dimensionless constants are important.

because in some sense they are physically meaningful. If you have a dimension attached to your constant, then by rescaling what you mean by a meter, you also rescale the value of the constant, and therefore the exact value you attribute to it is not something that you need to worry too much about. But dimensionless constants, they can't do that. And so for the last 50 years or so, physicists have been arguing about whether the values of these constants are themselves meaningful. And part of the problem here is that the values of these constants in order to end up in the universe we live in have to be tuned very, very, very precisely. So to many, many decimal places. And another problem is that some of these constants end up being very close to salient numbers, things like one, for instance. And so then we have this interesting

reductionism fuels question of, well, if this number is one followed by 16 decimal places, followed by a few other numbers, one followed by zeroes for 16 decimal places, followed by a few other numbers, five, two, seven, I don't know the exact ones, then is it a problem? Should we be looking for a theory that explains how that number comes to be different from one? So a reductionist would say, of course, this is pointing towards something more fundamental physics that's out there.

But somebody who's not a reductionist, perhaps someone who believes the anthropic principle that we're in this just so universe, this universe which is perfectly attuned to human life, and the explanation of that is that we existed it, then you wouldn't worry about this. So I think the same thing is really true about machine learning. If we have very complicated thought patterns, or if we have very large models which resist interpretability, does that mean that we're missing a trick and we should try and compress these things or is it the case that perhaps nature just doesn't compress? And I think the jury's out. And what do you think on the whole kind of real patterns thing? Do you think there is some natural convergence towards the types of knowledge these systems will find? Yeah, well, to some extent, form follows function. And in so far as we have trained these models on

Data the form of the models is Reflects the data and reflects reality in some ways is teaching us about reality And in fact yesterday I was listening to your interview with John jumper and I thought he put this very succinctly and beautifully when he was talking about the advances of Alpha Fold 2 Alpha Fold 1, where as if I remember rightly they used no more data than Alpha Fold 1, but in some sense they were just more in tune with reality in Alpha Fold 2. The architecture had been optimized to really represent that particular problem, not all problems, but that particular problem to a better degree. And so I think we are...

we are learning to make these models of the universe that really do reflect something deep about the underlying structure. But do I think that that... eventually we'll end up in a kind of pure empiricist paradise where we're bringing absolutely no biases to the table. I actually tend to think that's impossible. And in fact, this again comes back to David Deutsch. He says that all of science is theory laden. It's necessarily theory laden. And one way of seeing that is to come back to the idea of a world model. So what guides us in...

the experiments we do, is exactly the model of the world we have. We can't just setting up, there is no perfect instrument to set up that measures everything about the universe. So even in choosing the instrument with which you measure things, that is necessarily theory laden. And so you can't ever get to this sort of empiricist paradise. And so as a result of that, I do believe there will always be opportunities to uncover Perhaps some bias that we hadn't seen about the way we're measuring things. And then that itself will unlock or remove the constraint of how we were designing the systems. And then that will unlock another level of improvements. And so I think this is a false dichotomy to talk about, okay, should we replace the transformer architecture or build on top of the transformer architecture? I mean,

Probably some version of a transformer architecture will continue to work for some problems. There are probably other architectures that work perhaps more generally. But when we look back, much like we might look back at the Wright Brothers aeroplane and see echoes of that in a Boeing 747, for instance, we will look back at the transformer and see echoes of that in whatever it is that are our most powerful AIs in 20, 30, 40 years. Ed, we should move on to talking about your papers. So tell me all about it.

So in this work we were interested in giving a scientist agents the capability of replicating research papers and I'll come back to why but let me explain what I mean by that first. So paper replication is the process of.

taking a research paper and then redoing the original experiments that led to those results. And in some sense, it's a public good. It's something that scientists should be doing because it gives us these firmer foundations on which to sit and it reveals perhaps the tacit knowledge, the pieces that weren't captured in the original paper. It can also be the jumping off point for new creative explorations because a paper can't possibly be a perfect facsimile of of the research that was done, often you'll discover some wrinkle on the original method, which then sparks a whole new investigation. And in the setting that we had, we were really interested in whether an AI agent could not just replicate a paper, but do it in a scaled down version. There were two reasons for this. Firstly, just pragmatism, we wanted to be able to do many, many replications and generate data for the agent to learn from.

But secondly, to some extent, the ability to quickly validate or falsify a direction of investigation is a really valuable skill. And in my experience, it has been perhaps the determining skill of whether someone's a good researcher or truly excellent researcher. So the reason why we wanted to developer agents that were good at replication is that exactly as I said earlier, we believe replication is the first step on a curriculum of under specification towards innovation. So the very same skills that allow an agent to replicate a paper, making good decisions about what experiments to do, critiquing the way that it's gone about the experimental process, gathering information that's maybe tacit or that it doesn't know.

are the same skills that it would be necessary for it to design and implement its own experiments and therefore advance the frontier. So what we did in the paper is we developed a task space, which we called replica. So there are some classic ones by the big hitters of the field. And there are some much more recent ones, including ones in the area of open-endedness. And for each of these papers, we have a language model redact a figure from that paper in such a way that it's gone from the PDF and it's irreversibly removed. And then we task the agent with, given this paper with the figure redacted, use the description in the paper to recreate the original figure. And we give the agent instructions that it's not allowed to go and access the original paper figure. And we give it

one hour and we give it a one seventh slice of an h200 GPU using something called MIG which is the NVIDIA multi-instance GPU slicing protocol and the way that we score the capability of this agent is that we have another coding agent a frontier coding agent as a judge and that frontier coding agent has got the instructions for the original agent plus it's got a bunch of guidelines in order to detect cheating and we validate that judge against human taste if you like so the ultimate arbiter is whether this is a kind of human.

the humans think this is a good replication. And we collect data to demonstrate that the judge agrees to the human. And so that's the replica task space. The other thing is we develop an agent. We call that agent Faraday. And we do something a bit unusual in developing Faraday, which is that we take a small model, in this case, a Quen 3.627B parameter model. And we have that user frontier model as a tool, so a frontier coding agent as a tool with something we call catcoding agent as tool. And we post-train that 27B parameter model using many, many rollouts on replica. And because there are 300 tasks, we can get some diversity of data from that. And so we're really training a capability, a general capability at replication or a general scientific intuition.

Okay, we train on 242 and we test on 68 held-out tasks. And the held-out tasks are deliberately not in the same area of AI. So we train on classic ML papers, if you like, core capabilities, things like post-training, CNN architectures, open-endedness, LSTMs are some of the older ones. And then we test on AI for science papers. So quite different using AI to make models of the world.

And what we find is that our Faraday agent is able to perform better than the Frontier model. So it's able to perform better both than the coding agent that it's using as a tool. So it's clearly kind of instructing that it's squeezing more capabilities out of that, that coding agent. But it's also performing better than other Frontier coding agents like Claude, for instance, and also than the Frontier open weights models like GLM 5.2, which is recently released.

Now, you might be wondering, OK, well, did we just prompt GPT 5.5 codecs really badly? And that was a very reasonable question to ask. And in fact, we asked that question as well. And we ran a prompt optimization loop.

on GPT 5.5 Codex. It's effectively just an Andre Carpathie auto research loop where we say, okay, you can see the rubric judge score, and you can do many, many iterations on the prompt, and it builds up this very complicated prompt, which effectively just shouts at GPT 5.5 and says, hey, don't do all these cheating things, be a rigorous scientist, merely make sure that you try and iterate and do many, don't stop after 10 minutes.

and it accumulates this large prompt and it does improve the performance of Jupyter 5.5 codecs on these tasks by a tiny amount, but we still have a quite a sizable advantage.

What's interesting is that changing the weights of this small model and allowing that small model to really intervene during the rollout and instruct codecs in different ways and to check on what's happening during the run of codecs and to think about, OK, as a function of what's happening during the run, should I stop it or continue it?

is buying you quite a lot of advantage. So in terms of then the next steps for this, really there are a few. So one obvious direction is scaling up. Just across the three main conferences last year, this is ICML.

iClear and NeurIPS. It was something like 12,000 papers accepted. So even if we just say, okay, we want to just take a couple of years worth of papers, we could increase the size of our tasks set by towards some magnitude. And that, of course, would enable us to train a larger model as a coding agent that hopefully get an even stronger improvement. But the more interesting piece is, okay, we started with replication. How do we go towards innovation?

And there's one, there's various ways of thinking about this, but one I like to use is if you imagine that you were really great at replicating papers, right? And that for any given paper, you could do a really high fidelity replication. Well, now if you're also able to imagine a paper that doesn't exist, maybe you take an existing paper and you kind of imagine a change to the figure. And that's a very minimal kind of innovation.

But you could imagine something like, let's take the original transformer paper, and let's say, actually, I'm going to demonstrate the same results, but they're going to be twice as sample efficient. And then you just modify the paper, and you say, OK, this is what you've got to replicate. Then your replication agent is going to go gangbusters trying to replicate this idea, which is actually a completely new result.

And interestingly, I had a chat with Cleon Jones, co-author of the Transformer paper about exactly this kind of... Yeah, very fascinating. And he was talking about, okay, what was the process for the original Transformer paper? And really what they were trying to do is effectively replicate some existing results that were being done with...

recurrent neural networks, but without recurrent neural networks. So his constraint that he imposed was, well, let's just use convolutional nets. So they weren't actually interested in attention mechanisms whatsoever to start with. So they were replacing the RNNs with ConvNets, and then they were doing a machine translation task and trying to get good behavior, and hopefully at least as good, maybe a bit better behavior, but better performance than the existing tasks. And gradually they accumulated this sort of cobbled together pieces, ConvNets were one of them. At some point, a friend of his came over and said, look, I've built this attention mechanism based on the work of Dima Badan now a few years before. And it's kind of sitting around in this part of the Google code base, would you just throw this in? You know, I just kind of want to see how it does. And so he threw that in and it helped. And it was then later that they ablated everything else. And eventually they removed the ConvNets that they'd originally put in. And they found out that

nothing mattered apart from this attention mechanism. And that's when he came up with this title and it was Leon who came up with the title attention is all you need. And so, but what's interesting about that is that the thing they were kind of trying to do initially was just a replication. And then it was a series of steps to impose different constraints to the papers that had been imposed before. And so you can now start to see how you could take a good replication agent and actually use it potentially in in collaboration with humans is what we intend to really accelerate the rate of innovation. Yeah, and I think I buy it. So you're saying we start with and we should say it's deep replication. It's not shallow replication. So you have this LLM judge.

And it's not just saying, is the figure the same? It's saying, is it in the spirit of the paper? Is it showing understanding and all the rest of it? And then maybe we should explain the GRPO thing. And also, this is a really interesting model that you've discovered, because I've long been thinking about this. How can we actually attractively build adaptivity into these systems? And so as you were explaining, you've got this Gwen model. And by the way, the new 3.627B Gwen model, it's amazing. The guys at Two for Labs in Switzerland, they were using it for their arc.

the three harness and they said it's dramatically better than the last version apparently. So you're doing this adaptation with GRPO on that Gwen model and you're using that as a supervisor for the coding agent. And I guess the rationale there is that the coding agents, they have the latent capability. So it's like what you prompt it, you know, it's like, what's the magic word? If you can give them the right guidance, then you've got that big capability. Exactly. So I think of these coding agents, coding agent models as they're pretty good engineers.

Now, they do write slot code, so they're not brilliant engineers. There's some taste problems there. But if you kind of give them a goal, then they will go after it, and then to a large extent will succeed at that. But they're not great at asking questions, so they're not great scientists. And so really, we're building that scientific layer. And you're right to mention the new Quen model that's just come out. We're actually about to test that one as our next...

the next model we're going to post drain on top of. So actually the results we've got in this paper are themselves already behind the curve and we should be able to get even stronger results with the newest model. You asked that you said something else. I suppose more broadly. What do you think is being learned here? So you said you've got this data set and I think you use Gemini to remove a bunch of the figures and now the purpose is to recreate the figures in the right way showing deep understanding.

What exactly is the model learning? Is it learning some kind of abstract process of how to recreate these things in general? Let me give you a couple of examples of the kinds of things that Faraday learns. One of the papers that we had the agent replicate figures from was a paper called Voyager. It's an interesting paper because it's about building a skill acquisition library in a crafting game from a couple of years ago. It's very germane to open-endedness.

And but in this, there's a particular figure in this paper which is demonstrating how the library of skills is acquired over time. And the best competing run of Claude and Codex on this was, it turns out it was Claude.

Claude did was it hand coded a library, so it kind of simplified the skill acquisition by having a predefined skills that needs to be acquired. And part of the purpose, part of the point of Voyager is that in fact the system itself needs to be coding up those skills and then reusing them. And so what Faraday does instead is he maintains much more faithfully the idea behind the paper, which is, OK, can you not only use the skills in the library, but also construct the library on the fly? So that's just one example of the kind of rigor and faithfulness. As another example, I'll give you from one of the test tests, the AI for science test, and it's a paper called, No, I'm not a material scientist, so forgive me if I get this wrong. But one of the things that it was trying to do, I think, is predict some of the intratomic

potentials. So this is a figure where there's a model, a generative model that's trying to do this and the figure is both looking at the... has error bars on runs of this generative model. And it has parts where it looks at the behavior of this generative model under various physical conditions. And with the best competing run here is the codex run. And what that does is it runs one seed, so it can't actually provide any error bars. And it also omits these much more detailed subtle ablations of the model under different conditions. And Faraday adheres much more tightly to the specification of what a rigorous scientist would do, displaying both the error bars and also doing this deeper analysis of the different conditions. What you're seeing here is, I think, something that we would start to call good behavior from, say, an intern or a research scientist at the start of their career, which is...

being forensic in analysis, being rigorous in the way that you go about doing science, and really starting to ask the right questions to gain the maximum information you can about the setting rather than stopping at what might seem like a sort of surface level claim. I suppose another thing, as you said in the paper, scientific research is incredibly lossy.

And, you know, so you don't really give all of the details. So I guess I'm surprised that it's even possible to replicate most papers using this method. I mean, were you surprised by that? Yes. So this is a lacuna, if you like, in the paper. Now, not all papers will replicate perfectly. And the papers we've chosen, we deliberately chose papers which were well known and highly cited.

And the main reason for that was that we just wanted things that were going to be interpretable to us and also interpretable to humans in our expert network who were helping us to ascertain that the strengths and weaknesses of different agents. Now, as a result of that, because these are highly cited and well-known papers, they are, I think, at least I believe, much more likely to be replicable because if they weren't replicable then probably we would have discovered by all the people trying to build on top of them. And so in some ways we dodged the bullet of how do we figure out whether this paper is replicable or not. We somewhat addressed that

by virtue of having a judge, which as you say is much more interested in the process of replication than it is in the output. So visual fidelity is just one of a number of different pieces in our rubric. That rubric also includes things like experimental integrity and claim reproduction, whether the overall claim is reproduced.

As we scale the task space, I think we are going to come into this thorny issue of how do we build judges that are able to reward the agent for figuring out that a paper is not in fact replicable at all? And how do we avoid good hearting, this metric, perhaps for either cheating behavior or for simply not trying?

on papers that aren't replicable. Arguably, if a paper is not replicable, you should try even harder to figure out, okay, what is it that doesn't work out? Because then that in itself is innovation. Yes, and because you spoke about in the paper the kind of trade-off between using some kind of hill-climable scalar reward function versus using LLM as a judge with a load of criteria. And I mean, I suppose, what would cheating look like? I mean, as an example, I was intrigued by this, so I just downloaded a random ML paper and I just cut out the figures and I told Codex to recreate the figures. I was expecting it just to cheat immediately and do an incredibly good job and just to find the paper. It actually did a terrible job.

But yeah, I mean, there must be so many forms of data leakage, right? I was assuming even the tables of results and some of the description around that would allow it to shortcut and basically just recreate the figure, even if it wasn't there. And I was surprised that that didn't really happen in, you know, for me. Yes. So of course, there are a lot of clues in the rest of the paper and part of the instructions that we give both the model and the judge is that the agents shouldn't shortcut and merely grab results from elsewhere on the paper. And we definitely see early in training examples of just that behavior. Of course, it gets punished by the judge. And when we were sort of tuning our training procedure, that was one of the first things that we had to figure out. How do we stop it from just determining where the points should be getting a very good score on visual fidelity and that dominating the training.

I think that there are other... more subtle forms of cheating, which have to do with, for example, stacking the odds in the favor of the method that you want to work. So whether that's things like running on 20 environments and showing the results on just the one environment that sends to work, or doing optimal stopping, for example, say, once you have the result then you just cut the experiment at that point, which of course then means that all of your statistical tests don't apply in the way that they were designed to.

These kinds of cheating behaviors, I think, are more subtle. We haven't done yet, this is so hot off the press, we haven't done a full forensic analysis of everything, although we have sent some of the examples of replications to paper authors for their inspection and they've had very kindly had a look for us and haven't found examples of cheating, at least in those ones. I expect that we have still got some cheating going on and that this is going to continue to be a problem, I think eventually it will end up in a gray area. I think in the end, we're going to have to figure out what the norms are around this. And it actually brings us closer towards that point of how do we believe humans will use this? Because at the moment, we're building this as sort of foundational technology, but our intention is that

the capabilities of Faraday 2, Faraday 3, Faraday 4, etc. will be in collaboration with humans. And then to some extent, it will be around what norms do humans develop about using these technologies? And how do we go about evaluating and reviewing the outputs for things that we think are normatively good or bad in research itself?

as we were saying before so construct validity is very important so that's roughly you know is it actually following the abstract thinking process that the scientists were going through that they're not short cutting but another interesting thing is that you deliberately amortize the results in some way so what you do is you have a limited wall clock time and you're saying to the agent well if you can't do the full thing you might need to do a smaller version of this thing you know and prove that out is that lossy in any way I mean do you think that some scientific results only really materialize at a certain scale, and there is no simpler version of it. Yes, for sure. And there is definitely papers in our dataset replica task space, where we do see that there's no sensible scale down, or at least no sensible scale down is found. Now, let me give you an example. The AlphaGo paper is a fantastic example, very difficult on a 1 7th mix slice of a H 200 GPU and with one hour to train AlphaGo.

So one experiment that we do to assess the capability of Faraday more generally is that we, in fact, do some evaluation where we scale up the resources that Faraday has given. And so this is something that's completely out of distribution from training. But we deliberately pick papers where you can make we believe, and we sort of hand assess this if you like, that you should be able to do a replication of the figure or of the paper with eight B300 GPUs, which is a quite a sizable amount of compute, and with eight hours. Now, the eight hours we chose for rather prosaic reason, which is that that's roughly the amount of time that our model can go to before it exhausts the 256K context limit. So we didn't do any compaction. But what we find is that

the model not only is pretty good at generalizing to this setting, it also does quite considerably better than the Claude in this setting and arguably actually the advantage over Claude is larger than the advantage over Claude was on the one hour task. So there's something about this kind of scientific rigor that's really paying off more when the space of possibilities that you have to explore is larger. Now one challenge is how would we kind of continue to scale this up and really I think we have to hope that training on relatively small scale things teaches us or teaches the model the same capabilities as one would need to do large-scale experiments. I think we have hoped that's true.

because that is literally how it works for humans. You do not get your new employee at Google or OpenAI or Anthropic to immediately go and train the next version of GPT or Claude or Gemini because they will waste resources. They first need to learn how this works at small scale. And then it turns out you can develop those intuitions and generalize them up. And one of the core things you do in this dataset generation is you decompose papers into a list of tasks, essentially.

And I guess you're prompting a language model to do that. How have you done that? So we take the paper and we ask Gemini to identify figures which are plots. So we're not doing tables at the moment. That was just for simplicity to give it a single surface. And then we have Gemini redact the figures that uses a Unix utility so that you now have a separate what we call gold figure, which is supplied to the judge in order to determine, OK, how well has this replication been done? And you have the PDF with the redacted figure. And we do that for every single figure that's a plot in the main text of the paper. We focus on main text. Again, somewhat to stack the odds in our favor of getting things that are replicable so that we can, for the moment, dodge this question of, OK, how was this result replicable or not?

And what we find is that for most papers, there's one or two figures that work for that. So for some papers, there's up to 13 figures that work. And then we assemble that all into a data set. In order to get the reward function, what we found worked really well is generating a per task.

judge rubric. So a rubric is if you're like a mark scheme it's like the kind of thing you would give to an examiner who is looking at your work at school and it tells you okay you should reward the agent for scientific reggae, you should reward the agent for visual fidelity, you should reward the agent for claim reproduction and so on and so forth.

And what we found is that by having a per task rubric, so by having an intermediate stage where we adapt the rubric to the particular task and then use that consistently for that task for the entirety of training, we're able to achieve two things, better agreement with human raters and also much less noise. Which brings me onto something that you inquired about earlier, the question of GRPA. So one of the key achievements in the paper was that we got GRPO to work and that wasn't without its difficulties. We went through a period we called the RL crisis where just nothing worked and I know from talking to people at other companies they've had their RL crises and I expect that we'll have RL crises in the future. And part of the difficulty here is exactly because we're doing RL

On non verifiable tasks. So you have non verifiable tasks. You have an LLM as a reward and LLM is a stochastic generative model. So it's got inherent noise. And so now you're going to have to deal with the fact that from rollout to rollout, the same, you know, the same kinds of behavior can be judged differently.

And so in order to... We had another problem also which is that these are long horizon tasks. So we have these tasks last at least an hour in the kind of final stage of training. And we are interested in...

multi-turn behavior. So there's many different things that our 27B parameter model can do. It can use any kind of Unix utility that it has on the system. It can use Codex as a tool. It can interact with the internet. It can download things from the internet into the container. So it's really got this, you know, almost equivalent to a human kind of action space. And so that combination multi-turn one hour time period and then noisy rewards tends to mean that GRPO Goes well for a while and then collapses. So we did a couple of things. We did many things and we distilled the deads a couple that worked. One thing is very basic modification, which is that we just score the rollout multiple times with the same judge and we take the average of that. So that's relatively standard piece. The other piece I think is quite new, which is that we do per turn credit assignment and we do that in a slightly intricate way. So we have the judge in addition to producing this.

Rollout level score say okay for every turn of the agent during this rollout how much weight would you attribute to that and this is now a distribution there's positive numbers they will sum to one and we then normalize that weight so that we're not.

changing the overall distribution according to the number of tokens in each turn. So we wouldn't want these kind of weights if you had a very long turn to really magnify this rollout compared to all the other rollouts. So we do this normalization step. And then we use these weights to adjust the advantages during GRPO. So what we're really saying is that when you're upweighting or downweighting, the behaviors. We want to do that on a turn level doing GRPO rather than on a rollout level. And if we look at the way that the weights work, we do a little bit of interpretability on this. What you find is that the judge

ends up assigning more weight to turns that are in the middle of the rollout, because this is kind of the load bearing stuff. It's very... If you were to anthropomorphize this, it's very relatable as a human. You know, you kind of start your work. You know, the first kind of bit is a bit routine. You're just trying to get into the swing of things. At some point, you get into flow, and you're really making the important decisions. And then towards the deadline, then hopefully, if we're going to meet the deadline, it's just kind of crossing the T's and dotting the I's.

The other thing is that there's more weight assigned to terms where the 27B model is prompting the Codex model. And that's because decisions about what you ask the Codex models do are very important. That's the kind of really load bearing stuff. And so that we found that this combination of these two pieces combined with our rubric judge really enabled us to get stable training. And actually, in the end, we stopped training just because we wanted to put a paper up. We didn't stop training because we were in a collapse regime. Yeah. I mean, the way I insured that is so by going turn based.

what you're doing is you're putting these gradient updates in where there is signal and you're not where there is noise because you know what one school of thought is oh it should be based on an entire rollout but does that then mean intuitively that you kind of have black holes in some parts of the state action space so then you're just kind of relying on Gwen's default behavior and you're not updating those parts of the trajectory.

I think it's not that we don't tend to put... Well, the judge doesn't tend to put zero weight on parts of the trajectory. It could in principle, but we find it's more just sort of a change in the distribution of weighting. So what that means is that there are parts of Quen's behavior that we're doing a lot to change, and then parts of Quen's behavior that we're doing a little bit to change on each step. And so it turns out that what that...

does is it buys a stability because there are some things that Quinn is actually pretty good at doing. If you ask it to read a PDF, for example, it's good at doing that. You can do that straight away. If you're doing uniform credit assignment, then you're updating, you're saying, okay, great, you read the PDF, you're doing that every single rollout. You really don't need to do that. What you really need to kind of up weight are the pieces where it was actually genuinely something different and interesting that led to the better performance of this rollout versus the other rollouts in the group and so that's what that's what this is achieving. So a lot of folks at the moment like Gary Marcus they're claiming victory for neurosymbolic AI and they're pointing to all of the insane harness engineering that's going on you know there was that primal intellect harness that came out the other day and.

I don't know what to believe anymore. So you've got an interesting way because you're using GRPO and RL and it's actually very innovative. I think it's amazing. But what a lot of other people would have done is they would have just adapted, they would have come up with a harness and they would make the harness do library learning and skill transfer. And you see what I mean? I mean, did you consider that as an option? Yes, absolutely. And in some ways, this paper is exactly a reaction to that.

We very deliberately are not constructing harnesses in this work. And my co-founder, Lewis Kirsch, did a very interesting analysis of AI scientist works that are based on harnesses and the capabilities of the base models. And in that analysis, which he presented at a conference workshop a few months ago, he found that around about three months after you've built the harness, then the base model can do the thing the harness could do.

And so we wanted to explore a different paradigm that may also be true for our coding agent as to paradigm we'll see. But we at least wanted to kind of assess something different. And one reason why you might want to use our paradigm rather than the harness is that history teaches us, I think, that when you have these capabilities in the weights of a model, they're more flexible and generalizable than when you have them hard coded into a harness.

When you have them in a harness, however, it's not that all harnesses are bad. When you have them in a harness, a harness is likely more sample efficient than having them in the weights. And so it depends what trade-off you want. So if you were doing something like Alpha Revolve, where you have a very specific problem, how do we do four by four complex matrix multiplication more efficiently, then building a harness may well be the best thing you can do. Neurosymbolic AI to solve specific problems.

And this is what you get in all of these wonderful papers, things like Alpha Revolve, also things like the Darwin girdle machine or hyperagents from Jenny Zhang. They're all doing harness engineering and they're great at solving specific problems. But what we've observed is that this doesn't tend to generalize. And what we wanted to do is build a system which you can then apply to a completely different problem, quite a difficult long horizon problem, which is to replicate a paper in a completely different area of ML research. And so we believe that doing that by changing the weights for model is going to work better. Now, I think that these are, these in fact could be combined.

And there's a wonderful paper. I think it's called Ivo Tune by one of our research fellows, Anja Serena. She wrote this last year. And what she does is she has harness engineering. And then she has RL on top of that. And so you can think of this by analogy with AlphaGo. AlphaGo had search, which in some ways is this kind of symbolic piece. And there was also the neural piece of RL distilling this into the weights. And so one of the areas that we're very excited to look at next is how could you use harnesses at training time to boost the performance within a rollout of the agent and then distill that back into the weights so that you get the best of both worlds so you still get the the flexibility and generalizability of having a neural model that operates at test time. Yeah and quicker side it was good that you preemptively cited you again.

Just to prevent any turbulence downstream. I'm only joking. I've always been a neurosymbolic guy and I've always felt that there's something very powerful about symbolic constraints. They're incredibly powerful. And like one school of thought in AI is that the AGI that we build, the intelligence would have to be symbolic. And what we're starting to see now is that the symbolic stuff is important, but it can actually be recrystallized back into the model.

Right. So it's useful as a tool for generating data. We can bring it back into the model. The exception is, I think, that sometimes we need to crystallize specific skills, as you were just saying, that clearly are better if they're in symbolic land. But if we're talking about pure creativity, do you think it's strictly better that eventually we move those representations back into the model? No, I think it's a combination. I think there will be things that sit in the model and things that sit outside the model.

and I don't have a strong prior on what those things will be. I'm not sure it's possible to have a strong prior about what those things will be. I think that where we sit now, we can really much more clearly see how this kind of system might work and might lift itself up by its own bootstraps than we could before.

And it's good that you mentioned Jurgen. Of course, he saw this right back in 1988, I believe it was his master's thesis. And then of course, too much work after that. So I see the kind of construction of the symbolic pieces as something that may itself be done by the AI scientist system. So you could imagine an AI scientist system building a special purpose model, which might be itself a neural model.

it might be a skill, it might be some combination of a skill, a harness and a neural model that it can then use as a tool. So this paradigm of coding agent as tool is the tip of very large iceberg, which ends up with having a large number of different agents, a large number of different skills, a large number of different neural models, a large number of different symbolic pieces of equipment that are very sample efficient, all interacting together and also all interacting with humans. And what might the evolution of this system be? So at the moment, it's a single coding agent.

But I could imagine there could be a swarm of agents. We could potentially use a larger model to do fine-tuning. I mean, at the moment, it's quite exciting for folks at home because it feels like I could do this. It's actually a really powerful thing. But if you want to build the really, really powerful version of this, would you take, let's say, a 200 billion parameter model and do the same thing? So I think there's one dimension that we're very interested in which is scale. And we would like to see whether we can get scaling laws from this kind of approach. Now, of course, part of the philosophy is that we'd have a smaller model controlling a larger model. So there's some limit on how big you would want to make the smaller model controlling the larger model if you were to do that scaling. But another dimension is, as you say, to scale the number of agents that are in interaction with each other and in interaction with humans. So in fact, we're very interested to hear from people who might

be interested to work with us in seeing how they could use this model in their work. At the moment, we don't think that the model is general or reliable enough for a release, an arbitrary release to the world. But if there are people with creative ideas of how this could accelerate their work, that's something that we want to discover. Because of course, that will inform the kinds of interfaces that we need to build between the agents and the humans and between agents and other agents.

However, perhaps the most interesting and unusual thing is how we plan to use this within Inherent as a company. So our mission, as I said, is to recursively self-improve to discover new knowledge. And we think of recursive self-improvement quite differently from most other organizations doing this. We think of this as a phenomenon at a company level.

And so what that means is we're continually trying to close loops and put agents at the very heart of everything we do. And that starts with giving agents all the same affordances and context as humans. And it also means that the way that we as humans work is that we proactively adopt agents as quickly as possible. So already we're starting to use the Faraday model internally.

further kinds of replication that we might want to do to advance our research. And also we're trying to learn from the way that we're using the model in order to accelerate both the construction of the model and the future research success of the company.

In addition to having, of course, that neural model, we're accumulating a huge amount of context, whether that's the code that we write, whether that's also context about how we run the company of various types. And what we've found is that once you get to a certain amount of context, and once you give the agents a certain number of affordances, you really reach this Rubicon moment. So for the first few months as we as we ran the company, we had these agents proactively reaching out to us and trying to help us with stuff. And to be perfectly honest with you, it was quite annoying. They just had no idea how to help us. But at some point about two or three months ago, I think we reached a phase transition where the agents were aware enough of the company context and they had enough affordances to do useful stuff in the company that the things they were doing proactively.

became genuinely useful. And of course now you have a lever that you can pull to scale the number of agents and then to scale the rate at which you can do useful stuff in the company. And so that's what we mean by the recursive company. The company itself as a whole is self-improving as a function of the interactions of agents and humans rather than building a single agent that somehow will magically do this in its own isolated box.

Yeah i mean i've experienced the same things i i've i've got a skill surface and memory system and it's just because there was a phase change where it's now incredibly coherent for the types of things i do so i have a very differential experience probably to most people using a i because it's so good for me the problem is my my skill surface and memory system is spaghetti.

So it's very supervised. It's very specialized to me. I can't really, it would be useless in a large organization. You've built a general purpose system, which presumably just could be the future of how we use agents that could in principle ingest trajectories from how everyone in the organization is using AI. So there's this virtuous cycle where everyone has the new version of the small agent driving the bigger agent. And this is almost the dream of open-endedness, isn't it? Because now organizations can evolve their own agent.

systems that are coherent for them. Indeed. And I think that is the next era. It's the era of collective intelligence rather than the era of individual intelligence. The way that most people use agents at the moment is one to one, right? Sometimes it's one to many, you know, I might have multiple agents running at once. But at Inherent, we think that that is a somewhat impoverished way of using agents. In fact, If human society only operated by one-to-one relationships, then that would really slow down the rate at which we could do cultural evolution. So we really see the problem of using agents and developing collective intelligence as many to many. How can we build surfaces that enable many humans to be collaborating with many agents and intervening in a very organic way at different points in the cycle?

And I think that extends also to the physical world. So it's rather remarkable that the way that most of us do our work in an office looks the same as it did in the late 1980s. People turn up and they sit down at individual computer screens. And this is while we now have models which can go off independently and do a huge amount of work and can be scaled to very large numbers.

And I think it can't possibly be the right optimum and there's organizations that have done very small experiments in the grand scheme of things like valve famously has a desks on wheels and that's a bit they attribute a large part of their success to this idea. What would it mean to really develop the next generation organization that is evolving itself but not just doing it.

in the digital space or even just in the interface between many humans and many ages, but also in the physical space of the laboratory or the office itself. One thing that fascinates me is what the topology of this would look like.

Ken of Stanley is always talking about committee meetings and objectives, you know, like the tyranny of objectives. I mean, I could imagine a blended approach where everyone has their own agent, which is adaptively learning like this. And then maybe they choose to share certain streams of data within domains to and shared agents in the organization. And maybe at the organization level that there's there's a big agent. And there's all that you can just imagine that there are potential pitfalls here, you know, because all sorts of bad behaviors and good behaviors might might emerge. How do you see that coming up?

Yeah, well, I think for us, it's all about what we call living within the experiment. So we have to do a lot of experimentation. And I think that it's really an unknown unknown. And coming in with any particular worldview about the hierarchy or the structure, it's likely to not exactly pan out in that way. Now we've got priors and we're trying various things. But we haven't solved it yet. I want to give you a historical analogy to this, which I think is quite instructive, and it's to go back all the way to the Industrial Revolution and the invention of the electric dynamo, and that happened in the 1890s. What happened when the electric dynamo was invented is that factory owners were able to

replace their big steam powered turbines with electric dynamos. And they were more efficient, and this gave you a kind of small productivity boost. The problem was that the factory was configured for this single source of power, the big steam powered turbine. And what that meant is the factory had all of these systems of complicated ratchets and pulleys that then powered them at different machines in the factory.

That meant that it was very inefficient. If there was a power failure, then the entire factory had to shut down. No one could make any progress. It was also massively unsafe because you had to build quite tall and narrow factories to accommodate all these kind of pulleys and ratchets. They tended to be very dark, very difficult to operate, to kind of work as a human in these places.

What unlocked the really extraordinary productivity gains and also unlocked all sorts of products that people would have been inconceivable in the existing factories of the day was when people reconfigured the whole factory and that reconfiguration meant.

putting individual dynamos at individual workstations and then inventing the production line. And then that has various benefits. First of all, you can get rid of all the pulleys and ratchets. Second of all, if you have one dynamo that fails when the rest of the production line can continue, so you don't have a single point of failure. But thirdly and perhaps most importantly is it just improved the quality of life and the safety of the factories because now you could arrange them horizontally. You could put skylights in so there was natural light and you had a much safer and better working environment. And I think that analogy holds now with what we're trying to do at Inherent. How do we reinvent the factory for AI research from the ground up to put AI agents at the center?

Yeah, I was interviewing a Cesar Hidalgo who wrote a book called The Laws of Knowledge and he was citing an example from Jeff Bezos and he said he wasn't worried about Barnes and Noble competing with Amazon when they started selling books because they have the wrong structure. You know, they would need to do structural adaptation. And this is part of the reason why Kenneth talks about, you know, diversity preservation rather than just diversity because you actually need to keep multiple options open to adapt to, you know, to change your structure.

And is there something organizations need to wrestle with? Because if you think about it, there's there's something good about organizations having a clear purpose and coherence. But by the same token, there might be some optimal configuration. I think a lot of people just intuitively think that some kind of decentralization is good. Because when they discover a new strategy, shouldn't they be able to adapt themselves to start doing that instead? Yeah, I think Certainly the evolution of organizational design is important, I think new technologies tend to be get new forms which can make better use of those. I push back against the idea that there's an optimal configuration because of course the technology is always changing and that will change the organization. One advantage of building a new organization is that you can do many more experiments, you can move much quicker.

precise example of that, one of the most famous inventions at Google from an organizational point of view was the OKR objectives and key results. And that's really powered a lot of the success of the company. And I spent almost nine years in the company and made very good use of those. It's a goal directed process, which identifies the goals upfront and then works towards those goals as measured by the key results, a measurable quantitative metric.

But of course, to some extent, that goes against the philosophy of open-endedness, at least if you make this time period of those goals too large. If you take a more open-ended view, you would want those goals to be able to adapt and change within that time period so that you can take different stepping stones, perhaps ones that were unexpected. And so one thing that we're now trying to figure out at Inherent is Can we invent the next organizational paradigm, the one that is based not around optimization like OKRs, but around open-endedness? What's the equivalent of OKRs for the age of recursive self-improvement? Yeah, of course, you were at Google DeepMind before. I mean, we don't need to go into too many details here. But I mean, do you imagine a future that this is basically a revolution and the incumbents won't be able to adapt fast enough or

Do you think that, because it's interesting, isn't it? All of these large companies that they're implementing AI agents, but you're making a new company, which means you can create the structure de novo. So you can adapt to meet the situation. So do you think in five years time, we're just going to see entirely new companies that are doing it differently or do you think these old guys can adapt? Look, I wouldn't have started a new company unless I thought that we had a chance of doing something significantly different and that would you know, be really revolutionary in terms of the speed at which we're able to create new capabilities. I think that the existing companies have verticals that they are already going to be incredibly successful in and will continue to be successful in. But I do think that there is an emerging market of AI science and of scientific discovery.

And we know that growth is powered by innovation. And we also know that ideas are getting harder to find. And there's a wonderful paper with exactly that title. Whether you measure that by research or productivity, I believe there will be this new market of AI assisted, AI accelerated R&D. And I think that'll be hugely beneficial to the world because as we know, innovation powers growth.

But ideas are getting harder to find and there's a great paper. Nicholas Bloom wrote this wonderful paper a few years ago with exactly that title. Whether you measure this by research or productivity, whether you measure this by the number of inflation-adjusted billion dollars that are needed to develop a new drug, even if you measure this by the age of the Nobel Prize winner at the age that they make their discovery. All of those are going in the wrong direction. And I think it's...

Sort of obvious why the reason for this is that we have what's called a burden of knowledge with the victims of our success as a species we're accumulating so much knowledge that the.

time and effort it takes to get to the frontier of any given domain is is just so large that you now can't have individuals who know enough across domains. And so this is the promise of building a horizontal layer of intelligence across all of science. And the way we see ourselves fitting in is Can we be that intelligence layer that can power the next generation of autonomous labs, the next generation of R&D organizations that can power even the next generation of company construction to solve really, really difficult problems? And that's a somewhat different kind of market to the market for coding agents. It's a different market to the market for, say, the use of AI for existing corporates. It's the different market to the market for chatbots. And so I do think that there is an opportunity to come in and

and define what is meant by the culture for that market, what's meant by the interfaces between humans and agents in that market, and how do we really power the next generation of growth for humanity? This has been absolutely fantastic. Thank you so much for joining us today. It's been a lovely, lovely to chat to you. Thank you very much for having me.

Delete this episode?

This removes the episode page and its saved audio from this library.