← All shows

Machine_Learning_Street_Talk_MLST_Why_a_Nation_Can’t_Outsource_Its

Published Jul 13, 2026 · Duration 55:53 · Language en · 6 highlights

Summary

本期节目采访了英国前沿实验室 Cosine 的创始人 Alistair,主题围绕英国首个主权大模型(sovereign LLM)的构建。他解释了公司如何借助政府主权 AI 部门的支持,获得布里斯托 Isambard 超算集群的算力配额,从而具备了训练大模型的雄心。面对「别人用数十亿美元、你如何用数百万美元做到」的质疑,他指出关键在于 Cosine 采用授权模型权重、由客户自行部署的商业模式,因此无需为推理承担巨额数据中心成本,而 Anthropic 等公司近期的压力主要来自推理而非训练。他深入探讨了模型架构,认为总参数量与活跃参数量至关重要,并援引博客分析推测 Sonnet 约为 1.3–1.5 万亿参数、Opus 约 1.5–1.8 万亿。在强化学习方面,他强调「信用归因」(credit attribution)的重要性,用老师只给 B 却不指出问题的类比说明当前 RL 对所有 token 平均加权的低效,主张让模型学会形成可复用的抽象。他还讨论了对抗「代码垃圾」(slop)的方法,如运行时验证、漏洞利用验证(exploit validation)和端到端测试,以及 Swarm 多智能体编排(曾用 253 个子智能体一次性完成机械表编译器项目)。最后他坦言美国的出口管制与 Fable 被封反而成了机遇,虽感觉自己像「二等公民」,但誓言英国团队别无选择、必须把主权模型做成。

Highlights

  1. How can you do in millions what they are doing with billions? Yeah, no, it's a very fair question and one that I probably get more than anything else.

    你们怎么可能用数百万美元做到别人用数十亿美元才做到的事?是的,这是一个非常合理的问题,也是我被问得最多的一个问题。

    Frames the central provocation of the whole interview
  2. I think that one of the biggest reasons you've seen people like anthropic struggle recently and the reason they've signed the deal as they have with like the colossus cluster and so on, is because of inference and not because of training. So you do need significantly less resourc ...

    我认为你最近看到像 Anthropic 这样的公司陷入困境、以及他们签下 Colossus 集群这类协议的最大原因之一,是推理而非训练。所以如果你不打算做推理这一块,你需要的资源会显著更少。

    Counterintuitive claim that inference, not training, drives the compute crunch
  3. the likelihood is, and again, no one really knows outside of anthropic, that something like Sonnet is in the, I believe, like 1.3 to 1.5 trillion total prams, MOE, and has probably 100 plus billion active. And then an opus is probably the 1.5 to 1.8 range, and probably has 150 to ...

    很可能——当然除了 Anthropic 内部没人真正知道——像 Sonnet 这样的模型总参数大约在 1.3 到 1.5 万亿之间,是 MOE 架构,活跃参数可能超过 1000 亿。而 Opus 大概在 1.5 到 1.8 万亿的区间,活跃参数可能有 1500 到 1800 亿。

    Rare public speculation on the secret sizes of frontier closed models
  4. the teacher just gives you right. Okay, that's a B. Thank you so much. And you're like I don't know what made it a B... it would be far easier if the teacher just circle the sentence and be like this is rubbish, you don't say this, right? And that is fundamentally the principle w ...

    老师只给你一个'好,这是 B',然后你却完全不知道是什么让它得了 B……如果老师直接圈出那句话说'这里写得很糟,不该这么写',那会容易得多。而这从根本上就是我们试图在整个强化学习中引入的原则。

    Memorable analogy for why credit attribution makes RL far more efficient
  5. we had what we call exploit validation, which is literally like, okay, the swarm comes up with its list of things. And before it makes it to you, we spin up the application in a virtual machine in a runtime... And we tell the agent, well, you've seen the source code, so if it's v ...

    我们有一个叫做'漏洞利用验证'的东西,字面意思就是:智能体集群列出一堆可能的问题,在它们到你手上之前,我们会在虚拟机的运行时里真正把应用跑起来……然后我们告诉智能体,你已经看过源代码了,如果它真有漏洞,你应该能想办法攻进去;如果攻不进去,那就把它从列表里去掉,因为它显然是误报。

    Concrete, clever technique to eliminate false positives via real runtime proof
  6. people coming to us saying, OK, we now realize what you're doing is really important. How can we be involved?... The involvement and the urgency, importantly, from those companies, from government, from just like citizens as well, has gone through the roof... but also thank you, ...

    人们来找我们说,好吧,我们现在意识到你们做的事真的很重要,我们怎么才能参与进来?……更重要的是,来自这些公司、来自政府、乃至普通公民的参与度和紧迫感都暴涨了……不过也要谢谢你,唐纳德·特朗普。

    Wry admission that US export controls became an unexpected gift for UK sovereign AI
Full transcript

I was horrified as many were when Fable suddenly got banned. We have obtained the mandate to build the UK's first sort of sovereign LLM. How can you do in millions what they are doing with billions? There is a huge amount of room for error or a huge amount of wiggle room. It was not something that was on my bingo card in January and numbers like Tentril being knocked around. We were compressing most of the internet at that point. Really funny up to Claude Delicio production database. It was a...

Oh, you're right to point that out. Agentic harnesses are getting less important over time. A model can probably do with a bash only, basically, any task these days. Also, thank you, Donald Trump. Yeah, for the first time, I feel like a second-class citizen, because they are going faster than I am. And I really hate that. That bore was my blood more than anyone else. And we are going to do everything we can to pull this off. We have no choice but to make it happen.

Alistair, it's great to meet you, mate. We are here in London, where it's customary to say hello, Giza. Hello, Giza. So we're in London, are we? We are in Hoxton right now, so we're in Shortwich. We're about half a mile away from where Cosine started in my apartment, which was in Hoxton Square, just over there. So we haven't come very far that we have expanded a fair bit since then. And what is Cosine? Cosine is a frontier lab based here in the UK prior to about Three months ago we built Best in Class coding agents specifically for highly regulated and high-side environments, so think things like financial services, insurance, defence and so on. But more recently we have obtained the mandate to build the UK's first sort of sovereign LLM, which is a much more ambitious vision and something which is on a scale much larger than we've done before but it's very exciting to be working on it.

So, tell me about that Sovereign AI piece. I should say, by the way, so I read the article about you in the telegraph. And I was horrified, as many were, when Fable suddenly got banned. Yes. Because of this export control. And now everyone suddenly is thinking about Sovereign AI. So, tell me the story. Well, it ties into a bunch of different things. It ties into the backstory of cosine.

One of the reasons we're fortunate to be able to do Sovereign AI is that we have a lot of expertise around model training, model building, all of the infer algorithms, data, people that you need to do that kind of thing. And we've been doing that for some time. So we already had all of those things in the organization. And then probably, I'd say, nine, 10 weeks ago, we were inducted into the...

government's sovereign AI unit, or backed by them, I should say, that is something that they have put out to increase the number of sovereign AI initiatives, companies that are being built in the UK. And what that looks like in practice for us is an allocation of compute on the Ismpard Supercomputer cluster out in Bristol. And that has allowed us to, well, honestly, it's one of the things that's unlocked our ability to even have the ambition to do something like this, right?

Fundamentally, one of the biggest blockers for like a startup of our size or smaller, to be honest, on being able to approach work like this is a huge part of it is the compute. Like if you raised, you know, 50 to 100 million bucks, like a good chunk of that would go on compute on doing a project like this and to have an allocation come from the survey unit is huge because it genuinely does enable it. And we still use some private compute on the side, but fundamentally all of it will be done on his and Bard which is, which is super cool. And to be honest with you, it was not something that was on my bingo cards in January at the beginning of the year when we started out. I mean we still do our, I'd say, our conventional business of like coding agents and the models that we already have built, but we've been able to take that vision and really take it to the extreme in a way that we wouldn't have been able to otherwise. So I'm not being funny but the million dollar question is.

Well, you're actually more than that. I'll give you more than that. So, you know, folks over there in the US, they've got probably on the order of hundreds of billions. You've got Mistrah, which is on the order of, you know, let's say, 14. Single to double digit billions. Something like that. How can you do in millions what they are doing with billions? Yeah, no, it's a very fair question and one that I probably get more than anything else.

We at CoSign are not an inference company, and that sort of ties into the kinds of deployments and the way that we sell our product. So if your viewers are going to give a bit of background, given the fact that we predominantly deploy into highly secure, high-sided environments, most of the time, mainly all of the time, we are not hosting the model ourselves.

A customer isn't hitting, you know, cosine slash, you know, API slash V1 and then hitting a chat completions endpoint or something like that from us. They are either taking the model weights that we give to them, deploying them on their own GPUs. We have a lot of that. That is like the most air gap, the most secure deployment we do.

or they are renting GPUs in some hyperscaler cloud they're already a part of so whether it be like Azure or AWS or whatever and then they'll run the model there. What that means in practice for cosine is that we license the technology that we built. We don't actually make a margin on tokens or anything like that and all of this ties into your question meaning we don't have to spend a lot of the money that the Americans are having to spend on data centers for inference purposes. Now, that's not to say that you don't also need a huge amount of compete for training, obviously you do. And a huge amount of the infrastructure they have in the US will also be used for training, but I think that one of the biggest reasons you've seen people like anthropic struggle recently and the reason they've signed the deal as they have with like the colossus cluster and so on.

is because of inference and not because of training. So you do need significantly less resource if you're not going to do the inference bit and we're fortunate in the way that we sell the product it means we don't really have to. And then on the other side of that, we are taking some interesting research approaches in terms of how you pull something like this off and we can talk more about that in a minute I'm sure in terms of how we are architecting the model of how we're training it, some algorithmic stuff, all of that's to say that we do have A credible shot at pulling off the full run, including like the continued pre-training, the mid-training, the post-training, all of those bits, but to be completely transparent with you there is a huge amount of room for error or a huge amount of wiggle room. There are obvious...

places where we have had to make trade-off decisions. That extends like the scope of RL. I'd always like to do larger RL runs, more generations, more inference time compute during the RL process to get more variety. We can't do as much of that as we would like to if we had like 10 times more compute, for instance, right? So there are trade-offs, but fundamentally I think given the way that we've scoped the project, and we can get into more detail, I'm sure, but given the way we've scope the project, I think it is viable in that very narrow scope. We have some of the largest companies in the UK all feeding use cases and their desires for what they want the model to be able to do directly into us so that we can train a model that's really good for them. I think that is...

potentially a feedback loop that hasn't really been explored as much in the space. Obviously you don't want to say bad things about some of these companies, but yeah, the models from Mistral, from Co here and so on. They're not competitive. Even the Chinese models, arguably, they've only really started getting competitive the day before yesterday, so you know, just a chance or so, yeah.

GLM 5.2, so the vibes are good, or they won on the ARC challenge, it didn't do very well, but maybe that was just a red herring. I can talk about that in a minute, but yes. Yeah, it's very cool. But you know, like, so is it because you were, you were almost employing rights because they weren't trying to make it better. Like, how can we make models that are as good as those front-end models? Crudely, I think there are a couple of key things. Well, maybe three key things. I think one is architecture slash raw model size.

I think the second is active parameter count, and the third is data. I believe, and correct me if I'm right, I believe the largest model that Mistral have made to date is the 675B Mistral 3 large, right? Sparsamoe, very, very similar to Deep Seek's architecture, if not the same, I think.

As a result, fundamentally, that model exists in the way that it does because it fits a use case that they have seen. It probably fits a GPU deployment profile that they have seen in the enterprises they're trying to sell to in France or Europe. And as a result, pragmatically, they're like, right, this is probably the biggest that we can get away with given what certain companies have access to. As a result, you're obviously going to cap out how far you can go in terms of model performance.

There was a very interesting analysis done in a blog post, which I can't remember the name of, but I can send you post, because I just found it very interesting, of a breakdown of the probable sizes of models like Sonnet and Opus. Did you see that? It was so cool the way that it was done, through Vertex and figuring out, okay, given what we know about open-weight models and latency times and stuff like that.

What I saw was that they came up with a bunch of questions and they could infer based on the general knowledge. But that was a little bit sketchy, wasn't it? I've seen a more empirical one. I'll send it to you. That blog post alone played a large part in our decisions. Architecturally, we took for the sovereign model.

One of the things that was very clear from that is that the likelihood is, and again, no one really knows outside of anthropic, that something like Sonnet is in the, I believe, like 1.3 to 1.5 trillion.

total prams, MOE, and has probably 100 plus billion active, right? And then an opus is probably the 1.5 to 1.8 range, and probably has 150 to 180 billion active, depending on the D type we're talking about, whether it's FP8 or FP4 or whatever. What about Fable?

that it wasn't up for long enough for me to try to, because I basically, what I wanted to do is give Fabel that blog post and then point it and then point it and be like, right, do the analysis on yourself and tell me how big you are. Never got that far though, because I don't think it was up for long enough. It was, it's a ridiculous uplift though, isn't it? It's, I thought I'm sure you saw on X, like numbers like 10-trill being knocked around. I don't know if that's true or not. I have no idea. I think it's obviously bigger than the other ones and obviously has way more active premises.

but I couldn't speculate, I genuinely don't know. But yeah, I think that the net of that article was that, you know, Opus was in the region of 1.5 to 1.8 with around 150 billion active. Obviously, there's a lot of algorithmic and data work that goes into it, but I think if you don't at least match that architecture, then you're already...

going to struggle to sort of reach that ceiling. I think an example of this, and it's way more nuanced than this fairly basic argument I'm going to make. But if you look at like a deep-seek V4 Pro, 1.6 trillion total, can't remember the exact number, but it's going to be in the region of 30 to 50 billion active. And my view is that The reason for these architectural decisions is largely for Inference Ability of the Model. Like it gets like sure great that you have those top line 1.6 trillion parameters but like if no one can run it because they need two nodes of B300's just to be able to like fit it into memory and actually you know run it at decent TPS then like how many people can actually take advantage of that? I think that is one of the reasons that whether it be like

Chinese open source or European or American open source has not reached close source performances like there is that pragmatic question of okay well the labs have a huge number of GPUs and they have enough inbound demand to make sure those GPUs are utilized to a level where they're not that worried about having them up whereas you know the open source community doesn't really have the same argument. If you're running it yourself, if you're gonna run a model of that scale on your own hardware and it's not really being utilized that much by your organization, you're gonna worry about how much money you're spending on just those GPUs being idle. So I think that on that first point, architecturally, I think that overall parameter count is obviously important. I think active parameter count is also incredibly important. And then the last bit is data.

I think that obviously the labs have some element of edge in data, both because they are able to procure so much from the brokers who sell it, and also because they have internal data functions which are very mature at this point. I'm not saying the other labs don't have that, but they're definitely not at the same scale, both in terms of spends or maturity. Particularly, I think to an extent, and this is definitely not true but I think to an extent like the pre-training corpora, is there going to be that much difference once you're in like the 30 trillion token range? Okay, you kind of all have roughly the same stuff. We're compressing most of the internet at that point. Mid-training similar story. I think post-training has been and remains one of the most interesting areas for these labs. I think that is having worked with some of the labs on post-training data because obviously we're...

we're very good at coding. It's been interesting to see how they've been procuring even the formats that they've been using, like the raw data and the different use cases that they're interested about when they're putting requests out for, we want this, we want that, we want the other things. And I think that is one of the key areas they're differentiating. Having really good post-trained data and also just being able to run RL at like ridiculous scale. Quick pause. Agents are getting smarter every day. But even the smartest agents get stuck without the right context.

and the right tools. That is where notion comes in. With the recent launch of custom agents, notion became the collaborative AI workspace where teams and agents work side by side. And now, their new development platform is turning that workspace into infrastructure developers can build on.

Now this is exactly how I run MLST. The whole show lives in notion. My guests, the publishing calendar, the commercial side, everything is in there. But what's changed is that it's now agentic. I just talk to my agent. It can be Claude or any agentic harness. And then it then talks to notion via the MCP or the CLI. And it's just done. And then I can access it on my phone. It's an absolute game changer. We'll get to the oral bits. I know you've got an opinion that, you know.

our training is really, really important. But I mean, there's a few things on what you just said. I mean, first of all, the M O E thing, is it just a trade-off for inference speed, or do you think there's actually some beneficial advantage in terms of like factorisation having an M O E? I think at least like my opinion is, if we could have trillian parameter dense models we would, it would just be strictly better. I think it would just be better. Yeah, I think, but I think like, again, that is taking into the logical extreme of...

Could you be what would you need to inference that what would you need to actually run a parameter in a model of that of that size? I think we're like an interesting and interesting example of this is And and it's not really a fair fight, but if you take at small scale something like a GPT OS S120B, right?

That is an MOE. I believe it's five billion active. I can't remember it was a while ago, but it was something like that. And then you take a like Devstral 2-123B, which is, I believe, like a similar architecture to Lama 70B, so it's fully dense. I don't know if you've used them back to back before. Devstral feels so much better than the GPCSS model. And that they are architecturally quite different, particularly in the attention mechanism.

But fundamentally, I think that a huge part of that just comes from the fact that one, you have 120 billion active parameters, particularly the other, you have five. And it is a very, we've deployed both customers, like last year. And the difference was night and day in terms of how it felt and what they got out of the coding agent when they were running it. Oh, interesting. Yeah. Because on the other stuff, so you know, you were talking about like the data, the pre-training, the algorithmic stuff. I mean, maybe the algorithmic stuff is kind of converged only because we now have this basin of attraction where there's

kernel optimizers and entire ecosystems around this and maybe it's kind of converged. But the data thing is interesting, right? Because there must be an insane amount of engineering. And you know with LLMs, it's a little bit like what's the magic word? If you frame the question in the right way, it has that representational friction and it does interesting things. So I'm guessing anthropic do a whole bunch of date, data, curation, and pruning. But another thing that anthropic do is They have Claude code, they have the ecosystem. And it actually is coming in all day every day. Yeah, exactly. And those trajectories, because you know, you've got this big thing, which is that it's not about where you end up, it's about how you got there. And I think software engineering is the, it's not about running code. It's actually like just creating process. Yeah, creating mental abstractions, doing experiments, refining those abstractions, showing them with the team. So, so this process, like so iteratively, I'm running code, I'm testing things, I'm refining my abstractions.

Anthropic have access to so much of that data. How much of an advantage is that? Obviously it's a huge advantage. I don't know what their terms of service say. I'm not saying that they train on all that data. I don't know whether they do or not. Certainly one of the interesting things that having trajectories means and something that like obviously internally we...

you know collect our own trajectories of our own use of cosine obviously we don't have anyone else's because of the way we deploy it but for cosine employees using cosine at least we do still get a fair number of trajectories nothing like the order of magnitude that anthropic get but it the most useful thing for trajectories for us is seeing how users prompt models. So like there's one thing when you're putting together our L data sets and you can come up with these like beautifully formed like problem statements that get fed into the model that are very well structured and very well formed and they're very clear about what the expected outcome is.

The reality is that users don't prompt models in that way at all. They're like, FU, it doesn't work, why the hell, are you dumb? Why have you done this, that and the other thing? And the reality is, is that's the distribution that the model's going to be exposed to in the real world, is you need training data that looks like that. And I think that I would speculate there's a lot of the alpha, that anthropic get out of their trajectories they get back, is obviously, obviously they're helpful. I don't think that they can honestly know It's not like there is a grader attributed to that trajectory, like, it's not like an RL thing where you know that this trajectory was good, like, absolutely. I'm sure that you can do some judging stuff and you can see how the user's responded to be like, OK, well, at this point, the user was probably satisfied that the job was done. But it's also fundamentally seeing like, what is the user saying during the conversation? And how are they saying it is incredibly useful? And even our scale, we have

There's an argument to be made about how represented it is for the entire development ecosystem. That's a different question, but we do have an idea of generally speaking across months and months and months of usage.

do engineers interact with these products and how do they respond and all these things. And that's very useful to make things more realistic at training time. Yeah, 100%. Yeah, because I mean, you've spoken about slop online. It's one of my pet topics as well. So, you know, vibe coding works in the sense that quite often it does the thing that you tell it to do. And, you know, the test will pass. But actually, you're building a spaghetti monster and you're throwing, you know, more bad spaghetti after good spaghetti. And instead of doing a one line fix that it's supposed to do it will give you like 200 extra lines.

you know, there's the whole understanding, debt thing, we can talk about that. But, you know, so what we actually want to happen and neural networks do do this to a certain extent, they learn statistical inferences that represent some kind of abstract structure and that helps them to generalize. And obviously what we want them to do is to learn problems in the abstract so that they can generalize to new novel problems that they've never seen before. So, you know, the idea is we capture the thought process and then we capture that generalization if we do it well enough. That's the rough idea. I mean, how do you think about slop?

So, Slop is also one of my pet hates. When you are, obviously, any of your viewers who I'm sure use Claude code or open code or whatever agentic coding harness they like, see it nauseam is like, both the combination of model vibe and also slot problems. And you hit the nail on the head, I think, in your question, which is when you said that, okay, the test pass, right, okay, is it technically functionally correct? Sure.

right but to what cost is the big is the big question and I think fundamentally I mean like when you think about how these models are trained to do software engineering and this is the realisation that we sort of had when we when we built outposts and I'm sure you've seen the blogs about how we did it within that realm normally And again, we don't have that much insight over how this works in the big labs, but normally when you're training a model to be better at software engineering, you have some kind of software engineering problem.

you have a problem statement you have some kind of test often a unit test but not always depending on the task type that needs that is either failing in the in the prior state and passing in the after state and then you know you give the agent the problem it goes and does its trajectory it's roll out and then at the end you run the unit test and if it passes then okay I'm like you got it right great you get a reward and then the the waits update Obviously, the problem with that is a fewfold, but fundamentally, it could have come up with the most insane way of doing something, whether it be in terms of commands that it ran that could have been unsafe in terms of code that's absolutely crap compared to what it should have actually done. All of those things get reinforced whether you like it or not when you give that reward based on purely correctness. There are a number of things that we have done and are continuing to do.

in the RL process to try to ameliorate this. And in terms of slop specifically, there are a couple of key things. One is that, like, correctness gates... Everything else like if you get the problem wrong regardless of whether you did an elegant way like you don't get rewarded But beyond that we do have other rewards that target the exact things that we've been talking about and also in some cases But not all we do have reference implementations for these things like if if you are using a poor request from a permissively licensed open source repo, as like some seed data, you do have the original patch, the human mate, and you can actually do some level of like, okay, let's compare what the agent wrote, let's compare what the human wrote, and does this seem reasonable to be, you know, 500 lines longer, probably not, maybe we shouldn't give the full reward for this. So there's stuff you can do on the pure reward level, but we also

doing stuff on an algorithmic level, which has been looked at more and more in the space, which is to do with credit attribution in trajectories. One of the big problems with our relays, it stands. I think is that, and I am by far not the first person to say this thing, Andre Caparthi said this like over a year ago, but fundamentally this notion of, you have a rollout of maybe 256,000 tokens, right, in some extreme cases.

And that culminated in like a 1 or a 0, depending on what the model did. And what we're saying at the moment in many cases is, okay, all of those tokens are equally weighted in getting us to that answer, which when you think about it is insane, because that's clearly not true. And in so many cases, the...

that there will be small decisions or important decisions in a trajectory that were sort of, you know, forks in the road that could have resulted actually a bad outcome, whether the model decided to get on the right one and ended that. And what we and others in the space, I guess, are trying to do right now, is, okay, if you can find out what those ranges of, like, a high entropy tokens or places where you know a decision was made, finding that is half the problem and then figuring out once you know that this is an important thing that happened determining whether it was good or bad relative to the final outcome is a different story. But if you can do that, A, your RRL gets significantly more efficient because you're not relying, I mean the analogy I always come up with is like, okay, say you were doing your English A level and you were practicing and you'd written like an essay for your teacher and you'd written a 2,500 word essay.

On a question and the teacher just gives you right. Okay, that's a B Thank you so much. And you're like I don't what made it a B is like I'm not gonna tell you what made it a B is B And then what you're gonna have to do is you're gonna have to write hundreds of essays and you'll get an A on some we'll get a B on others get a C on others and eventually you are gonna be like okay, so when I do this I tend to get an A more so I think this is probably a good thing to reinforce but it would be far easier if the teacher just sent you know circle the sentence and be like this is rubbish you don't say this right and that is fundamentally the principle we're trying to bring into like RL across the board because you get so much more performance you get more out of the flops that you have and also you're teaching the model to learn the things that are actually important and not just the filler right

I know, I mean, the great thing about machine learning is that it just generalizes low down the abstraction mountains. So from very superficial statistical generalizations, the bad thing about machine learning is it generalizes to exactly. So yeah, I completely agree with you know, we have a huge problem with benchmarks and machine learning. So, you know, we are obsessed with Passette 1, Passette 5 accuracy. And we don't seem to care about reliability, about, you know, consistency, about security, about abstraction forming. That's clearly the most important thing. I mean, Francois Cholet, he did the art challenge.

unfortunately they were brute forceable. Now we've got this new version, which is so difficult to brute force, you have to form abstractions to get any kind of good performance on it. So you're saying there's a new form of RRL, perhaps different from the deep seek type of RRL, which is rather than just being rewarded for getting the right answer, you're actually forcing it to form reusable abstractions and go higher up the mountain. Yes, in short, you've explained that far better than I did, but yes, that is essentially what we're trying to get to. One of the nice things is that I think that essentially what we're talking about here is credit attribution within a trajectory. That ports quite nicely to a bunch of different RL algorithms that are invoked at the moment. You can use it with a GRPO. You can use it with GSPO and all the different flavors of the algorithm, but fundamentally,

having a rigorous and importantly unopinionated way of pointing at ranges of work that an agent has done and being like, this is good, this is bad and so on. And for the advantage that's being calculated across those different trajectories that you've made, for that advantage to be attributed to those ranges, disproportionately to the rest of the tokens and the trajectory, means that your weight updates will be more targeted at making those characteristics either appear more frequently or less frequently. Do we still have an epistemic problem because you know that the one problem with machine learning is it doesn't really have like the notion of true and false so yeah we can do feedback from code execution feedback we can do a whole bunch of abstract lenses on actual processes that engineers are doing but don't we still have this gap though that we don't really know whether it was correct or not. It's one of the heart yes it's one of the hardest things particularly.

I mean, one of the, obviously, one of the reasons that coding has taken off as a use case is because you have verifiable rewards in some guys. And I think one of the things that, I mean, I've just said is that we are trying to bring some of the more fluffy, taste-related things and make them verifiable, but on a more floating scale. I think for, obviously, for things like law, for things like other non-verifiable domains, That is way, way harder. That is way harder. I think that is one of the reasons, one of the key reasons that we haven't seen like the same revolution in other industries as we've seen in coding or maths or physics, because you can't just statically compile some law and see whether you get a one or a zero. And I don't know what the answer to that is. I really don't. Maybe there are new ways of codifying those domains.

or maybe you just bring everything in distribution, which I think is what's happening these days. I think if you just make everything in distribution and you target every use case and you have some level of whether it be human judge or good LLM judge that has been either trained or well prompted by human, you get close. But I still think that essentially in many cases you need to enumerate a lot of these problem sets at train time. Make sure the model's good at them because otherwise you'll The dream of generalisation across the board, I think, has already been sort of shown to not really exist that much. I know, but we're in such an interesting time because every new model comes out and, you know, Fable was so much better, but we want to have systems that do more with less, which is what you're saying. So we want them to acquire these abstractions. And it's a really, really weird situation, right? Because do you think we'll ever get to a point where we can remove the human from the loop? Because right now I think the basis of AI psychosis is that you have

very, very talented humans. And they know how to ask the question and they have taste and they go in the right direction. And there's this virtuous co-creation cycle. It's very, very good. So, you know, we're now in the realm of it's not vibe coding. It's a genetic engineering. It's very exciting. But do you think it could ever be done without humans? Yes. I think it can be done without humans. I don't think we're...

anywhere near there yet. What would that look? Do you mean in well-specified? You know, I'm not talking about style trends, but you know, like there is that anthropic, they built a sea compiler. That does a well-specified problem. Yeah. I give you like a novel application and I can only vaguely specify it. We're going to need a human for a long time, aren't we? If you want it to be good and maintainable and actually in the style of something that a senior engineer would write, For now, yes, you definitely need a human there. Probably post-hoc to be like, okay, here's the mountain of stuff that you need to change. And all the design decisions that like in your chain of thought, you thought were good, but actually weren't because of, you know, real reasons. But I do think that I do think that we will get there. I think it's going to be through a combination of model improvements, creating

our environments and problems that really look like the kind of things you're talking about and also harness engineering I think a combination of all those three things yes and we'll get to harness engineering but okay but in the meantime we've got the spaghetti monster mitigation strategy and and I think one of the big problems is code review Yes, because you know like the the AI psychosis is manifested as understanding that so increasingly I become less aware of what's going on and that's actually really bad for me maintaining my competence and for me being able to evolve the software going forward so that's the right questions yeah exactly so so how can we do this because now we are generating ridiculous amounts of code right is this a case of let's use more AI to do the review or do we still need humans in the review I think

that I think we need more runtime validation of what AI is producing. And what that looks like is, I think code review will evolve somewhat. I think it will be more like proof that the thing that it says it's doing is actually doing that thing. Obviously, like AI reading git diffs is not useful. I think it can catch things, and I've seen it catch things in the past.

But one analogy, and it's not actually for code review, but one analogy that I can tell you that's really good when it comes to, say, our cybersecurity scanning products that we have, which I guess is an allegous to code review, because it's reading basically a whole code base using a swarm. But one of the key things that we saw with that was that it would go through that process, and it would pick up so many things across large code bases, but this could be a problem, this could be a problem, this could be a problem, and I think one of the things that we see in that and in AI code review is like sure like if you just look at that code in isolation and that function definition it can look quite dodgy but in reality the code path is never hit or there is like this other function that's called first that mutates this variable that then means that that doesn't do what you think it does or there's an environment variable that happens at runtime that means that this doesn't happen all of these things that you know it's just ask analysis you can't do it

And one of the best ways that we mitigated that problem in that product was that we had what we call exploit validation, which is literally like, okay, the agents, the swarm comes up with its list of things. And before it makes it to you, every single one, so we spin up the application in a virtual machine in a runtime, right? And in as close to production way as possible.

And we tell the agent, right, well, you've seen the source code. So if it's vulnerable, you should be able to figure out how to get through it, right? You should be able to like craft your, your horrible zip file to exploit this thing that you think exists. And if you can't, then you just take it out the list because it's clearly a false positive. And what we have been working on is applying the same logic.

but to just PRs, right? And we're not necessarily looking for cyber vulnerabilities, but we're looking for, okay, you have allegedly built out this feature, you have this new screen that has this table in it or this form in it that does this. And when you click on this button, it should result in a new entry in the DB and it should show up in the, all the stuff you'd expect, right? And instead of like, obviously you can look at the diff and for the most part, you can get a lot of mileage out of that, but also just like, show me it's doing that.

prove that it's done that in some reasonable way before it even makes it to me. Because otherwise you end up in this situation, I was just talking to a customer earlier today, where they were saying, when they first started adopting like agentic coding tools, they were in this spot where either they would end up with this enormous backlog of code review, right? And no one likes code review, let's face it, no one actually enjoys that.

or you get people being like, after a minute, it looks good to me. Merge, thank you so much. And you just get this yellow merging into like your main branch. And that's bad as well. So I think there is going to be way less cognitive burden if in however you're doing a review, you can see the code, but you can also see canonical proof.

that at least on the happy path, the thing that it's saying it's doing is actually happening. And the other half, I think at least in the present data, this is also really comprehensive end-to-end testing of everything you build.

and that is something that really sucks to have to build out, but once you have it, it saves you from so many problems because I'm sure you've seen, I've seen in like personal projects and so on, you vibe code for an afternoon, you built 10 new features, then all of a sudden the other five you have before stop working and like, why is this happened? It's like, oh, well, okay, the abstraction you had, I've just messed with it, and now it doesn't work for that thing.

So yeah, it's a combination of like defensive stuff like the end-to-end stuff and also just proactively lifting mental burden from people but being like look here's like either a screen recording or some screenshots or whatever of me showing you that this is what I think it is. I know I know it's really funny up to Claude Delicior production database. It always say oh you're right to point that out.

I'm so sorry for doing that. You can do about it by the way, but yeah. I mean, this is another alignment problem though, right? Because, you know, if we are accumulating understanding debt, the functional descriptions themselves are going to suffer from that because we don't understand the functional description anymore. And it's not just functional descriptions, you know, like there are intents, there are, you know, there's behavior, there's all of these different levels of describing a system, user stories and stuff like that. And unfortunately, these are different views of the blind elephant, right? They don't necessarily have friction with reality.

And you see the problem here like we're just... We're just losing touch with what it's supposed to be doing. In many cases, actually, Claude and the models, they're writing the functional tests, and then they're kind of hacking their own. Oh, it didn't pass. OK, I'll just change the test. Now, it passes. Great. Here we go. You see it all the time. I know. So, you know, it's a very difficult problem, but one that we need to fix. And, you know, for me, I think a lot of it is to do with scoping and constraints, you know, to at least cut down the subs problem. But we should move on. And what are your thoughts on agentic engineering? So you guys are going to agentic harness. If I understand correctly, a couple

a lot of years ago. You actually forced everyone to start using that because you really want to optimize the hell out of it. And you were talking about the RL piece. So maybe there's some co-evolution with the agentech harness and the RL. What's important in this agentech harness? I think that so, broadly, I think that agentech harnesses are getting less important over time. Oh, interesting. Why? The models are just getting so good. Oh. I think that...

I think that you can get, like, the proof point is that, like, a model can probably do with a bash only, basically, any task these days, more slowly and with more tokens, but it can probably still do it. That's not to say that agentech harnesses aren't important, but I think, like, over time, where's the value coming from? Where's it coming from? It's coming from the model and not from the harness. I think in the way that we built ours and we have been building agentech harnesses for a very long time, like, the first agentic model that we had was a fine tuned GPT-4 turbo model that we trained in January of 2024 and The coding agent harnesses didn't exist at that point Claude code didn't exist none of this stuff existed and We had to build one out and we

Did that symbiotically with the design of the model which is something we still do today because you do get way more performance out of like tightly coupling the two fundamentally harness engineering is still something we care a lot about these days we actually care more about efficiency than anything else in a world where token costs and tokenomics which is a word I have for the first time today or forward tokenomics is becoming increasingly important to particularly enterprises who we sell to We want our harness to use as few tokens as possible for stop. That's what we're trying to do Yeah, we'll get to that in a second because if the models can do epistemic quantification, you could in principle allocate a budget of tokens to get certain things done. But even before we get there, I mean, I want to push back on the harness engineering because one thing I have found is subagents are a game changer. Yes, absolutely. And that is because a lot of problems are too complicated for an LLM to do in a single path. So I'm sure you've had a similar experience.

As a problem becomes more specified, as you reduce the ambiguity, the entropy goes down, the models get better and the models are better when they have less rot in their context. So what happens is the engine is after doing a bunch of engineering.

What they find is that they kind of decompose problems into agentic subtasks. And then they have a fresh agent and the agent has a clear specification and you do this one thing. Yeah. Exactly. But, you know, so what you're doing logically as an engineer is you're factorizing a problem into smaller sub-problems. You're getting agents to orchestrate and you're not rotting the context in the main one. I mean, how do you see that evolving over time? Because now it's quite a manual process, but do you think, I mean, you've got this swam thing maybe. Oh, thank you for bringing that up because that was exactly what I was going to want.

to me is sub-agent orchestration taken to the logical extreme. We're kind of lucky to be able to do it because I think that obviously Cloud Code has, I believe, workflows, and Codex has sub-agents as well. A swarm is what it sounds like on the tin. A swarm is genuinely like a real, a lot of sub-agents running at the same time in a hierarchy way. And because cosine isn't trying to serve hundreds of millions of people a day, we are able to serve a swarm-like feature.

Whereas I think if I'm Thropic had a swarm for like opus, I think they would have even colossus one would run out of run out of tokens. Swarm does exactly what you've just outlined automatically, basically. So there is a there is a video on my on my Twitter and also my LinkedIn of me taking our Lumen out post model, which is post trained Kimmy K2.6. So definitely not an opus or a mythos or anything like that. And I asked it, okay.

um i want you to build me a mechanical watch compiler was the use case that i came up with i am swissed by birth so i have a reason to do this and i basically asked and actually someone did this exact same task with fable i don't know if you saw it on twitter when fable was out but essentially like i want you to build me a sdk in python for me to be able to like specify mechanical watches in code because i don't know herology but i want to be able to do it anyway i want it to be physically congruent on to to use some kind of physics engine. And then I also want like a 3D viewer to be able to like see the thing running. And it all needs to be possible in real life. You can't have things like intersecting which other wouldn't be possible and so on. And that is something that out of the box, Kimmy cannot do. There's no way, not even close, like it would be terrible. In fact, even Gemini 3.5 and O person 5.5 can't really do it. But as soon as you put them in a swarm and for what?

Swarm looks like for cosine is you have one orchestrator at the very top it breaks down a problem into subproblems for basically product managers or whatever you want to call them. We call them sub planners, but they own verticals of this. So like within that task, you would have had a sub planner to do the SDK. You would have had a sub planner to do the 3D view. You'd have a sub planner to write the documentation and so on. And then those sub planners could then dedicate to workers. And then they have like a flat layer of as many workers as they like. And for that problem, we use 253.

Subagents which I think is more than you tend to see in a cloud code session and so on I think I think you'd probably hit you usage limit pretty quickly that way But when you do that it is possible and you can do that entire project in one shot. And I am contradicting myself quite badly because I've just said harnesses don't matter but in that respect they do obviously matter. I'm pretty excited about that but you know it raises the question that when first of all when you start to have loads and loads of agents you have more understanding debt and less interactivity because for me the lack of interactivity is part and parcel of the understanding debt. And sometimes you want to interject you want to say you've gone slightly wrong there I want to change what this agent's doing.

And what many folks have found when they build these agent systems is that the agents interfere with each other, they've kind of overrided some, they're going to deadlock. How are you dealing with all that? So it's a hard problem and we experience all those things. In terms of, you can, so one of the key things that we did is we made you...

we gave the ability to interject to like an agent on the lowest level so say you had a worker that was like two levels down from the top one you can actually like talk to that one which is an important thing and with regards to like other problems in terms of like trading on each other's toes you can put right locks on file, so that you only one agent can edit it at a time. You also can provide context to agents when they are using it. Say like an agent is reading a file that it just read. We have stuff in the harness which says, okay, another agent has just edited this file, so don't be surprised if you see it slightly differently as to how you sort of last time. All these things help. They are not like a panacea, but they certainly help. And they make sure the agent is less surprised when they're like, oh, where did that come from? That wasn't in my last edit.

But fundamentally, it is, it sort of comes back to the point I made earlier where, okay, yes, you will run this thing and it will provide you a huge amount of value very quickly, but you are still going to have to come through it afterwards and be like, actually, the reality is, I don't like this abstraction, you've done, I don't like the way you've done this. There is going to have to be some sweeping afterwards, I think. Yeah, what are your thoughts on memory? Very hard to get right.

Okay, tell me more. Very hard. Oh, so... Well, it's a similar thing actually to the RL because memory is not about the destination. It's about how you got there. Yes. Memory is very hard to get right. We've tried a bunch of different approaches.

fundamentally I think every approach to memory that exists right now is a bit of a hack right it's like a tool and it's it's in many cases like a vector DB or like an embedded version of like some tip bit of knowledge but they're very hard for the agents to know when to query it's also fundamentally quite hard for the agent to know whether something was useful enough to write to memory and It's also difficult to keep these things up to date. We've had many situations where an agent's been doing a trajectory and when...

It's been doing something that was genuinely the right thing to do. It's like used its memory and it's like the memory's been old and out of day. And then the agent's like, oh, well, the memory says you should do it this way and it changes its tack and there you're like an engine engineer. You're like, no, please, don't do that. There are things that we're looking at internally with regard to like continued, continual learning and stuff like that to try to avoid memory being a tool and for it to just something that for it to be something that's just in the latent space with the model.

That is also very hard, but I think it is a more intuitive and elegant solution to it just being a tool. It's also a tool that's very hard to get right during RL because it is a huge...

surface area for gunnery in terms of reward hacking in terms of leakage in terms of it's being able to clear something query something from the future that it shouldn't have access to yet despite all the guard rails that you can put in place it's just hard yeah exactly and in a sense this this is another area for AI psychosis because I've written a memory sea line and I would almost argue that now you don't even need you know vector databases and so on you can just have an inverted index you know just because the models are so good at asking in different directions And so, yeah, there is a huge problem. It needs to know to retrieve. That's a big one. But it actually works. But it only works for me because it creates this fractionated spaghetti mess again. So there's another spaghetti mess in the memory seal. But it works really, really well. But it doesn't work very well at the organizational level because my spaghetti monster doesn't play with, you know, with John spaghetti monster. So if we could solve that problem and you're drawing an interesting picture as well of how we can actually optimize the different layers of the sandwich together. Yeah. I think that

As soon as someone gets hit right, you'll just know when you're using it immediately. I haven't seen a single implementation, whether it be in Claude code, to be honest, whether it be ours, or chat GPs, or any of them, where I'm truly like, oh no, this isn't just a hack. This isn't just a rag. This is the one remnant of rag that still exists, really, in the more traditional sense. And I am certain that there is a better way out there somewhere.

Oh, definitely. But I think another thing you've said is specialization, not generalization. I mean, for me, a lot of agentic engineering is like emergent specialization. So it's like, let's take a big intelligence to crystallize a small intelligence to do the particular thing we're doing. But as we're nearly out of time, final question is synthetic data generation. So what are you guys doing about that? Tons. I don't know how much of it I can get through in five minutes, but I'll do my best. So there are a number of areas that we have done.

synthetic data generation in in the past we're doing a lot more of it in the more forward-looking sense for like the sovereign model because the sovereign model can't just be good at software engineering it has to be useful across the board and that and that that means that we have to get good at that synthetic data generation particularly in the RRL realm in things that aren't just coding The Bren and Butter though is coding, and the way that we have done this in the past is with a very cool and sophisticated pipeline that we have built out over the course of like a year and a half now. But that, one of the big cold start problems in coding RL, particularly if you're using like open source repositories as a sort of seed data, and even if you're using closed source, that actually doesn't matter.

Nearly all of the PRs which commits or whatever you want to refer to them as don't have a built-in grader, right? The ones that do are often bug fixes. The classic example, I guess, if you're not enough to be in the space is like a sweet bench style problem, you have a git issue, you have a PR that fixed it, and then because it's open source, you have some sort of regression test that was added, and then that's your seed data. The real world doesn't look like that, unfortunately, and software engineering in the broad doesn't look like that, meaning that...

to do RL well, you still need to fundamentally be able to do tasks that aren't bug fixes, which is the vast majority of what engineers do. And you still need to be able to tell whether like the agent got it right or not, broadly speaking. And what we do at cosine is we take like real work that was done. We're not, we're not magicking up, made up problems for the problem for the model to solve because there is fundamentally, if the model can come up with a problem, it can probably solve it. But we take real problems that were solved, whether it be feature work, refactoring, whatever it is. And we, what we're synthesizing is we're synthesizing ground truths or greatest, all ways of measuring, whether that thing has been done. And it is a bit of a minefield because

particularly within RL obviously it's not supervised so fundamentally we need to be able to test these things in a way that isn't too tightly coupled to like the original implementation like there are many ways to skin a cat as we know and what that looks like in practice is you need a implementation agnostic enough way of testing it that's still rigorous enough to check functional correctness it is a very fine line to tread and we have done a lot of work around it We have a long pipeline built on it. We have custom post train models that live inside that pipeline that essentially have gotten good because we had to do a lot of manual labeling in places and stuff like that. But what it has allowed us to do is have a autonomous pipeline where this is particularly important for enterprises.

We can point it essentially any programming language, any type of task, any stack, anything like that, and say, OK, I want RL data for this problem set, and we can get a good chunk of it. And that is why for our outpost model, when we were coming up with the languages that we wanted to get the model good at, things like Java, 4 trans, C++, I could go on. We use that pipeline to gather the RL data for these things. And in many cases, and in some programming languages, there aren't even test suites.

That's where it gets really hard. That's where you have to get a bit, a bit inventive in terms of like measuring whether the agent has gotten something right or not. Like, I think it, yeah, it's things like very log and says very log where you have to actually run, I think, what's called a synthesizer in your environment to be able to tell whether like this chip actually works or not. Yeah, exactly. Fortunately, I don't run this pipeline. A track called Bend does and he knows far more about that than I do. But yeah, that is broadly.

within software engineering, how we've gotten very good at it, and it is something that we are generalizing out into other use cases that aren't just software engineering as well. Very cool. So in closing, you might argue that the US government have handed you a commercial advantage here because now it's more important than ever to build software and AI. You guys have this model coming out, you know, towards the end of this year.

Do you think you're going to be able to do it? But also, are you still at risk from a supply chain point of view? Because you know, like so much hardware is controlled by America. How's this going to pan out for you guys? Not even I think that if the hardware is already in the UK, I don't know how much there is they can do about that. All of the infrastructure, the model is going to be trained on already exists and is up in the UK because we're doing it very soon. In fact, experimentation already is happening upstairs.

In terms of... Has it been a bit of a gift? What's happened recently? Given our positioning, absolutely. Yes, without a doubt. It has been probably the busiest I have ever been since founding the company. In terms of people coming to us saying, OK, we now realize what you're doing is really important. How can we be involved? The consortium of companies that you read out is growing by the day. And the involvement and the urgency, importantly, from those companies, from government, from just like citizens as well.

has gone through the roof, and I feel very fortunate to an extent lucky that obviously we were well positioned to take advantage of this early. Also, we did part ourselves in that position, but also thank you, Donald Trump. How did we not see this coming then? Because it was such an oh shit moment for so many people. We did, though. I think a lot of people that we certainly did, but it was always fobbed off as like, I'm sure I can. I suppose that could happen. But like...

I personally didn't expect to happen as soon as it did. I am also freshly surprised by the 5.6 news that I'm sure you've seen as well, where that's going to be rolled out. And there is even for me, like a huge part, maybe like, oh man, that really sucks. I wanted to try that model that I don't know where I'll be able to now. And that might just be the existence for us now, unless we and others do work to get that level of performance out in some other way. And that's like our job now, I guess.

Yeah, I mean, for the first time, I feel like a second-class citizen, because those folks over there in America, they've got better AI than I do. Yes. There you are going faster than I am, and I really hate that. And believe me, that boils my blood more than anyone else, and we are going to do everything we can to pull this off. You mentioned how you're going to do this, like we are just going to, we're just going to...

make it happen. We have no choice but to make it happen. Please do. Yes, we are going to do everything we can to make it happen. On behalf of everyone in the UK, please do. We'll do our best, thank you. So it's been a pleasure. Thank you so much for having me.

Delete this episode?

This removes the episode page and its saved audio from this library.