← All shows

The Pragmatic Engineer - Building Codex with Tibo Sottiaux

Duration 1:13:21 · Language en · Published Sep 09, 2026 · 7 highlights

Summary

本期访谈回顾了 Tebo 从应用数学、Google 与 DeepMind 的研究工具工作,走到 OpenAI 并领导 Codex 团队的经历。Codex 最初源于帮助 OpenAI 研究人员更快编写内部 Python 基础设施的模型与小型代理,随后与 A3 项目合并并发展成面向公众的产品。团队选择用 Rust 构建核心代理,是为了以清晰的产品边界换取长期的正确性、安全性、效率与扩展能力,而非只追求早期开发速度。Codex CLI 开源且允许接入其他模型,体现了团队希望靠模型与产品本身取胜、而不是依赖锁定用户的理念,同时也带来了低质量贡献和功能被提前复制等成本。Tebo 将代理框架描述为略微领先于模型的“拐杖”:框架先用提示、护栏和工具补足能力,模型训练成熟后这些补丁便可逐渐缩减。随着代理能够深度验证依赖关系、自动发现逻辑与安全问题,代码审查的重点正从逐行确认正确性转向讨论产品意图、系统契约和关键不变量。代理也显著降低了依赖升级、维护和重新架构的成本,但良好抽象与模块边界反而更加重要,因为大量代理可以在极短时间内同时修改系统。对于希望成为优秀 AI 时代构建者的工程师,Tebo 强调深度好奇心、快速理解系统、持续追问“为什么”,以及真正理解用户社群与清晰表达意图。

Highlights

  1. Although it was the most fun I've had on solving hard technical challenges, I learned a lot from not having product market fit, not having the right users, not having the right feedback loop. There's a lesson there that I carry with me: always deeply think about the impact that y ...

    尽管那是我解决高难度技术挑战时最快乐的一段经历,但我从缺乏产品市场契合、没有找对用户和反馈闭环中学到了很多。我一直带着这个教训:要始终深入思考自己产生的影响,也要思考所参与项目的整体重要性。

    A candid lesson from a technically excellent failure
  2. The decisions early on are really turned out to be quite important, as long as you don't sacrifice too much of the velocity. We had very prolific and amazing Rust developers, and then you get a lot of validation as well at compile time. Primarily we were focused on correctness an ...

    只要不过度牺牲开发速度,早期决策后来往往会显得非常重要。团队里有非常高产且出色的 Rust 开发者,而编译期也能提供大量验证。我们首要关注的是正确性,同时也重视效率。

    Explains the counterintuitive Rust bet
  3. If we were going to be successful, open source itself would change, and the role of code itself would change. Being part of that community seemed important instead of divorced from it. We didn't have all the answers, so let's just make this a level playing field and encourage a l ...

    如果我们能够成功,开源本身会改变,代码所扮演的角色也会改变。成为这个社群的一部分,而不是与之脱节,显得非常重要。我们并没有所有答案,所以不如创造一个公平的环境,鼓励现阶段的大量试验与探索。

    A principled case for open-sourcing an AI agent
  4. You set it up with a couple of crutches so that it can actually do the thing to a level of reliability and in a way that is efficient and also with the behavior that you expect as a user. Over time, what we see is the system, the developer message shrinks, and then the harness al ...

    你先给模型配上几根“拐杖”,让它能够以足够可靠、高效且符合用户预期的方式完成任务。随着时间推移,我们看到系统提示和开发者消息会缩短,整个代理框架也会随之缩减。

    A memorable model of how harnesses evolve
  5. When we benchmark them, they're superhuman in code review. This is not just true for correctness; this is also true for security, where they're capable of reasoning across very, very complex things. Really what we see is this discussion around the intent that takes place around t ...

    在基准测试中,它们的代码审查能力已经超过人类。这不只体现在正确性上,也体现在安全性上,它们能够跨越非常复杂的关系进行推理。我们真正看到的是,围绕拉取请求的讨论正在转向意图:你究竟想做什么,而这件事本身是否值得去做?

    Redefines the human role in code review
  6. Maintenance is really like a tax that you pay over time just to keep things running. A large part of maintenance just comes for free. The cost of mistakes is going down, but the good old rules of software engineering, of having good abstractions, really help.

    维护就像为了让系统持续运行而长期缴纳的一种税。如今很大一部分维护几乎可以免费完成,犯错的成本也在下降;但软件工程那些关于良好抽象的老原则依然非常有帮助。

    Captures how AI changes software economics
  7. There are two things that are important: deep, deep curiosity for how things work, and an ability to train yourself to understand things very quickly. If you can't explain what you're trying to achieve, if you can't explain your intent, if you don't have a tie to a community, if ...

    有两件事很重要:对事物如何运作抱有极深的好奇心,以及训练自己快速理解事物的能力。如果你无法解释自己想实现什么、无法讲清意图、与某个社群没有联系,也缺乏判断力,那么要做出卓越工作就会困难得多。

    Timeless career advice for AI-era engineers
Full transcript

Codex is one of the most popular AI coding harnesses today, but how did it all start? Many of you will know today's guest Tebo from his generous and pretty fricking Codex usage resets. He was also there when Codex as a product started and has led the broader Codex team since. Today we cover how Codex started and why it was built in Rust and made OpenSorners. How code reviews are changing inside the Codex team and OpenAI. What it means when maintenance and re-architecting are getting ridiculously cheap.

the merge of codecs into chat GPC looked like and the many underappreciated engineering challenges of this project. If you want to understand how Teams inside of OpenAI plan, review and ship software, this episode is for you. This episode is presented by TurboBuffer, a ridiculously scalable fast and cheap hybrid search engine built on top of object storage by an engineering team that are really going to like after spending time with them. TurboBuffer is the tool that companies like Entropic, Notion, Cognition and Harvey all use to connect their AI products to massive amounts of unstructured data.

When I've talked with engineers who use TurboBuffer, the theme that always comes up is reliability and performance at scale. The reasons for this have everything to do with TurboBuffer's architecture. TurboBuffer uses only object storage for state and MVME SSDs with memory cache for compute. Data in TurboBuffer is organized into namespaces. You can think of a namespace as a database table or a search index or an S3 prefix depending on the world you come from.

When a namespace is not being queried, it stays on cheap object storage with no associated compute cost. When a namespace is active, TurboPuffer pulls it up into hot-catching tiers so queries are very fast. This design fundamentally makes it effortless to scale to hundreds of millions of namespaces. If you're building a multi-tenant AI product, every user and their agent can have their own dedicated search index without any overhead. And each namespace can hold hundreds of millions of documents without any special configuration.

you can scale TurboPuffer virtually without limit, and the performance, reliability, and operating model all stay the same. If you need to connect AI to lots of data, TurboPuffer should be your first choice. Check it out at www.turbo puffer.com. Tibo, welcome to the podcast. So good to have you here. Thank you for having me. It's so good to see you again. It's good to do this last time with it in person, now we're doing our video. First, I wanted to ask you, how did you get into tech? When did you first know that You want to work with computers? It's a good question. It was a long, long time ago. My parents actually decided to move out of Brussels where I was born and just thought it was great to just buy a small house and refurbish it, but it was in the middle of a village with not much going on. I think there was like roughly 200 people living there.

not many that I felt like I wanted to talk to or, you know, could make friends with it. And so I kind of got stuck. This is like very early, like eight, eight years old. I kind of got stuck as like, you know, computers and, you know, it's like early days of like the, for me, the internet. And, you know, that was my way to learn about things. And so just the rest is just like, you know, came from that. So like I owe it to my parents to have moved into the middle of nowhere and then had no choice but to get interested in computers. Once you finished high school, you went on and went to university, right? Actually studying it properly. Yes. I studied mathematics, applied mathematics at university. I went there quite early. And so I graduated early as well. I thought for a long time that I would actually not make it and that I would drop out.

It's like small companies and small consulting business like while I was studying. I was working for banks. I was working for, I was very interested in supply chain and applied mathematics problems. And I was sort of like selling that and learning a lot through that. Eventually I ended up in the startup world in Belgium. I did that for a little while and then moved to London to work initially at Google and then DeepMind and then now moved.

To be here at open the eyes like this is California. I love the California weather. We can talk about that It's been very good right after University you started you've you've founded a startup right you had the startup bug in you or the entrepreneur entrepreneurial bug Yeah, so this the startup was all about pharmaceutical supply chain looking at the supply chain for clinical trials and try to optimize and decide, hey, should you produce more medicine? Where should you send it? Where should dispatch it? How do you avoid waste? And through that, making clinical trials more efficient. And this was using traditional non-ML techniques, more optimization, solving Monte Carlo simulations, these kinds of things. Stochastic multi-stage optimization problem, really.

And we also applied it on steel industry and we applied it to electrical grid as well in Europe. It was like anything that sort of had the shape of like an optimization problem, we sort of like get interested in. And you know, to this day, like this, this, this company still exists. And I think they do some of the most interesting work still, but it's changing a lot, you know, with, with modern AI for sure. But it's interesting because you kind of said like, oh yeah, that wasn't ML. It was just a traditional stuff. And then you go into like Monte Carlo simulation and optimization and this algorithm.

I get a sense that you kind of just went deep, right? That it was like, okay, like here's a problem space. Like how can I use mathematics stuff that I learned stuff that I didn't learn to just go deeper and deeper? Do I sense that correctly? Yeah, that's that's why I was obsessed with applied mathematics is just really this idea of you have theoretical mathematics or you have theoretical science and physics and like there you just you do it because there's something to be discovered and something beautiful about it. And it's all about patterns and pushing the frontier. But you don't necessarily always know how you're going to apply it. And then there is the real world, right? There's all these cool problems that just lie around. And I was very interested in seeing how can I make the world better? And so how do I apply sophisticated mathematics to just optimize the world around me? And that was a lot of the thesis behind that startup. Yeah. And then after a startup, you ended up at Google.

at Google London. It was in 2015, and I remember in 2015, Google was a really, really competitive place to get into, like, maybe as competitive as OpenAI today in terms of the industry or terms of prestige. You worked on maps initially, and then you moved over to DeepMind. Can you talk a little bit about what you worked on and then why did you move on from a already really interesting space that you securely loved, you know, like the optimization logistics and all these things? Yes. I didn't start on Google Maps. I started on...

a project that was meant to make the web faster and meant to make websites faster, especially on mobile. At the time, Google was trying to see the transition from desktop to mobile and more and more traffic going to mobile phones. So we wanted to get ahead of that, so funded a number of initiatives and projects. I was working on one of them. This was really, really fun because it was a small group.

within actually the ads organization, it was meant to offset the loss for the ad revenue loss because of this shift of traffic to mobile and worked on it for roughly two years. And then it was canceled. And although it was the most fun I've had on solving hard technical challenges, I learned a lot from not having product market fit, not having the right users, not having the right feedback loop.

not trusting your product manager when they say the project is going well when in fact it's not going well at all and then you know one day it's just like this VP flew in from California and then it was just like oh yeah it's like you know we're canceling this project you know unfortunately you only have you know hundreds of users and this is clearly not Google scale and then it's unbelievable but people were surprised and I think.

There's a lesson there that I carry with me, of course, is just always question, always go to, always deeply think about the impact that you're having, but also the importance of the overall project that you're contributing. And then I moved into Google Maps. Google Maps was super fun, worked on reviews. And then after roughly a year, I couldn't ignore DeepMind. It was just, it was this special place, headquartered in London.

So many great things were happening. This was really the early days with rumblings of things like AlphaGo. And they just seemed to be doing extraordinary things and just really tackling the very, very hardest problems that you can tackle. And with my background, I was obviously drawn to that. I started there. I worked on a lot of the research infrastructure, research tooling. This is a theme that I carried on for almost a decade. And this is very much also how I approach things is, how can I build tooling and products that help make others more efficient and bring a lot of utility to them. Initially, I was doing this for research and then like over time, you know, I got like into thinking about things in a much more like more general and general and general way, you know, eventually, like, you know, ending up where I'm now. And a fun story that you recently shared on X as well is how you were part of the team that built this internal Google bot that was, you know, if you want to say similar to chat GPT.

but a year before chat GPT. Can you talk about that? That is a new story. I haven't heard it before. This was part of DeepMind. There were multiple efforts as well. There was Brain as well that was separate at the time. They had their own efforts on large language models, but it was definitely something that was being explored. It was not the main thrust of DeepMind. DeepMind was very much worried and busy thinking about grand challenges and games and thinking about RL, not in the language sense.

And so there was this group that was pushing on large language models and thinking about what if large text corpuses are everything? What if you just pushed language to its maximum and you just scaled language models? Would that be enough to get to general intelligence? That was a hard debate at the time. And then one group decided to just really push on that. And then it felt really natural as I was building, tooling.

you know with others for research is like you know obviously you're like you know what can we do with this model like how do we present it you know to the researcher like how can they sort of like you know debug the inputs outputs and eventually you sort of like end up with you know like a chat system so we built that internally we had a lot of fun initially the models were like you know kind of like almost like a little bit absurd like you know not very coherent not super useful but it was a lot of fun to sort of like think her with them that caught up like you know like wildfires like you know this application is just sort of like you know everyone was kind of like sharing little conversations within DeepMind it felt more than like a research project or like a research a project for researchers and so then there was this desire over time to like launch it as an external product but DeepMind was just like not set up you know there was like the the right way to launch products at Google there was like you know the whole machinery of like you know how you do that um you know the whole like

the less production stock. You're obviously very, very optimized over the years to do things well, but also very, very hard as an environment to truly innovate. And then I want to ask what made you look around or maybe even consider open AI, but I feel you partially answered this question. Just putting myself back into your shoes. It's 2024 or 2023. You're inside of Google who are publishing amazing papers, doing really good research. You're doing super fun stuff, pushing the limits of what's been done before.

Yeah, it's inside a company where you already moved. You know, for people who are feeling kind of comfortable or good about where they are right now, which I imagine you must have been, like, what made you still explore? All right, like, what else might be there? Yeah, I was very comfortable. It's a good place. But really, I had a desire to, you know, meet great people but also join a mission that I truly believed in and that, you know, I felt like the people were through.

to the mission and care deeply about impacting the world in a very, in a deeply positive way, but also in a direct way, not, not being like, oh yeah, it's just like, you know, we just do this work over here. And then it's like, it's the job of someone else to figure out, you know, how to, how to make this useful. It's like, I wanted to join a group where, you know, like all the parameters were sort of like considered together where, you know, research and product were like really co-designing.

opening, I was just like crushing it. I thought Chatchapiti was like, you know, taking off as a guy was, I met a couple of people from open AI and then I was like, wait, what, you know, you only have like 20 people working on Chatchapiti? Like that is, that is an insanely small number that must be like extremely empowering. Like, you know, how does that work? How do you manage to maintain, you know, a product with that level of scale and with that level of autonomy with, you know, only only 20 engineers?

And then, you know, I said, kind of dug and dug and dug. And so it was just an amazing group of people, amazing mission, you know, super talented, super driven. And like it was, it was like drew me in. And then I joined pre-reasoning efforts immediately, like typical opening eye fashions. Like I joined, it was like, oh, yeah, you know, like there's this thing going on, like, you know, we're going to launch reasoning models, like, you know, it's like a new paradigm. And then, you know, start sprinting on that. And like, you know, like a month later, like the company launched 01.

I want one preview and now it's exhilarating to be part of I wanted to be part of like a place that moves fast cares about impact would be in tune with the world and you know just really listen and sort of like that's also to me that you know what I've carried with me like when when building codex when building products it's like having a community listen to the community just really focus on like a really intense feedback loop and then building something that is just like you know you just really want to care about it and like, you know, care about the utility of it that it provides to the world. And then, of course, you start to work pretty quickly on codex. So you joined in 2024. Can you take us back? What the thinking back there when you joined was about AI or LMS and code. I know there was this ASWE effort back then. We talked about it in the deep dive as well that we did in the pragmatic engineer, the autonomous software engineer. A3, yeah, that's what it was.

pronounced internally. We don't have the nice read I've heard anymore. It's Codex. But really for me, I started building infrastructure for research. A lot of what I did before was large scale data storage, analysis, and then tools to understand training runs. I did a lot of different things over my years, but it was always about building for others and making them faster.

and just really caring about, you know, fundamentally doing that well. And then through tooling and infrastructure making new things possible. And so when I joined OpenAI, it was like with the same idea. And then with the read with one preview and like, you know, some of some of the later models, it was very clear that we had to use the models themselves to help us go faster. And so I just really got obsessed with this idea of what were the limitations? How were we going to use those models for research itself?

So we got together with other folks in research. We started training models. We started building little agents. Those were truly the precursor to Codex. And this was like, we were training internal models to be very proficient on the Python code base of OpenAI. And then very proficient with having like good taste in architecture, good taste in code style. It was a Python only.

And then the idea was like, you know, we would sort of like use that to build infrastructure very quickly and, you know, help researchers code faster as well. And then, you know, and then we would move faster. And then over time, when you just kind of like pushed out and simplified to its core, you're making a lot of, you know, we found like where you could make a lot of progress very quickly and then learn very quickly. And then Greg and Sam are, you know, people with immensely supportive and also Greg was very adamant that, you know, we would we would not just focus on ourselves, but we would also focus on benefiting the world. And so he just sort of encouraged that we would be thinking about this not just as a tool for OpenAI itself, but also as something that we would actually make it through a product. And this is when we merged this research effort with this A3 effort, and we started building one thing. And then that led to a sprint, which was like the initial cloud codecs that we launched, which

didn't really have PNF because it was like a little bit too high friction. And we also launched the Codex CLI and we continued to push, but there was always this idea of, hey, how do we get models to really help here? You mentioned that first you started to build this model to train on the Python code and actually help build in for better. But then you made this interesting decision where for Codex, you built it in Rust.

And at the time, the model was not on distribution for Rust, right? It wasn't as good as Rust and it was in Python or TypeScript. Why did you make that kind of a decision? Was it kind of like, did you expect that it'll catch up or you figure that performance is more important? Because it was very counterintuitive. Most of the other harnesses built were actually not built in Rust. They were built on distribution on TypeScript or Python or something else. Yes. From first principles, like we very early on we were thinking about the product interface and the agent as different things. So it was very important to build the core of the agent in a way that was robust, that was secure as well, that was engineered for efficiency and skill. And having worked through projects over the years that go from

Hey, this is a fun thing to like, hey, we need to scale this to the scale of like the largest data center. The decisions early on are like really turned out to be quite important, as long as you don't sacrifice too much of the velocity. And so it's like, it's a, it's a tradeoff, but we had very prolific and amazing Rust developers. Our internal models were not bad at Rust. And then you get a lot of validation as well at compile time. It's like, you know, statically verified and all these things.

And that is great for agents too. So turns out, you know, it was quite clear that, you know, Rust as a language would actually be quite good for agents fairly quickly if we decided to put some effort into it. But primarily we were focused on correctness and we were focused on efficiency as well. Interesting. So you're saying, you know, it's worth, in your case, it was worth thinking ahead of where you want this thing to be. And for example, things like a language choice, obviously with agents, you can rewrite a bunch of stuff and easier than in the past, but it's still like you can save yourself reworking by putting in the right scaffolding or the baseline of what you're building on, right? I think we could have been successful if we had been in a TypeScript or maybe even Python and then it would have been fun and then we would have rewritten it at some point. But having a very clean separation between the agent itself, which can exist irrespective of the product,

It was a very important principle. And if you write everything in the same code base, in the same language, just like, inevitably, you're going to be a little bit sloppy and you're going to intertwine things more than you should. And then it's going to prevent further innovation after that. And so that was very important. Like the Rust boundary, in a sense, was very useful for that. One interesting decision that you made, which is unique across all of the major labs is...

having this built-in open source, right? The CLI is open source, the SDK and the AppStream are all open source. When and why did you decide that? It's not a given, especially, you know, there used to be jokes about open AI having things closed, but this is actually the opposite, where like this is open, whereas like some competitors would ship closed source harnesses, which again, I think it's very easy to understand why you want something closed source. Why did you want an open source? There was something really...

cool about the idea of having the code open source because fundamentally what you're billing is you're billing a coding agent. And so we were sort of like thinking about, well, if you have that, you know, you're obviously going to point it at itself. And, you know, maybe, you know, you can build a community of, you know, contributors that use it to improve it. And then, you know, you can learn a lot from that. Also, it felt at the time is like, you know, very clear to us that if we were going to be successful, open source itself would change, and the role of code itself would change. And so being part of that community seemed important instead of divorced from it. I think it's hard to solve problems if you don't witness them yourself. And then the other thing was just, it still feels like early, but it was very early at the time. It felt like we would have some ideas for how to solve things well.

And we were co-designing these with the training and the research. And it's all about expressing the capabilities of the model and the most flexible in the best way. But also, we didn't have all the answers. And being very open about, hey, this is what a good harness looks like. This is how we think about it. We did a couple of very technical deep dives and blog posts. And we talked about it a lot. And we thought, hey, it's just like, The world is vast out there. There's like, you know, crazy smart people is like, you know, we're we're going to get inspired by other open source project as well. And so let's just make this a level playing field and so like encourage a lot of tinkering and exploration at this stage. Now, this has been now, you know, like a year later, a year and a half later, which is a very long time in right now in this time frame. But looking back or taking the experience.

What are the benefits you've seen the kind of engineering benefits the engineering teams benefits from being open source and just honestly what are things that are kind of hard being open source, right? Like there must be downsides like just trying to get an honest take on both sides Yeah, they're definitely downsides it comes at a cost, right? The benefits are it's almost something to build in the open. It's awesome to have like small a small repo as well. Whenever we hire someone and they join the Codex team, it's like they've seen the repo before. They've looked at PRs. Allboarding is done. Yeah, it's done. In onboarding, it's just like you use Codex to look at the repo with you and you ask some questions, but it's not a secret issue that you can get productive right away. We get a lot of good contributions, although we get a tsunami of random stuff as well. Obviously, you and everyone else.

Right. Open source is changing. I think this is one of the examples. That's right. And then to me, it just, and to a lot of the team, it just brings a lot of energy to just be part of the community and like be directly contributing, not just saying that we care about the community, but actually doing things that, you know, you can see it's costing us effort, right? We don't have to do it. The downsides are, you know, it's separate from the rest of our code. So, you know, sometimes we have to draw like artificial boundaries and, you know, work across multiple repos.

When we're working on something particularly exciting and you know we're building it in the open then you know at times we find that you know others copy it you know before we have the time to release it and it's like it's just a little bit sad but also it's like it's part of the game you know it's it's like you're building in the open and it's like you know that's that's sort of like the contract that you signed is like you know you can copy it we have a very permissive license as well.

But it does sting a little bit when you're working on something and you're like, you know, and then the third thing is just like everyone else is like, you know, we are overwhelmed with, you know, random contributions and, you know, we have to deal with that additional tax. But then that pushes us to, you know, also like try and solve for it, right, which I think is good. And on top of the open source, one thing that surprised me about Codex and I didn't even know about it until recently, it's not tied to the open AI models. You can.

use other models with codecs. You know, like putting myself in a vendor's shoe, it might not be very obvious because again, all the other vendors I look at when they do a CLI, it's kind of use it with our models. Again, what made you decide to be this permissive about, you know, using or allowing to use your harness with other models? It felt quite natural if you are part of this.

community and building an excellent coding harness is like, why would you couple it to your model? That felt like, so like quite disappointing to make that decision. So it didn't feel right. And in general, it's like, I think, you know, it's like, I kind of tried to make decisions that I'm like, yes, you know, it's just like, I can just sort of like, explain it, you know, it is correct. It's the same reasoning with, you know, it is open source in the first place. It would have been trivial for anyone to fork it and then add support for another thing. But then you're just encouraging people to just go and use that fork and then not suddenly you have overhead. And the only reason you have a fork is because you wanted to change 10 lines of code to add support for another model provider. That feels very silly. So why not just support it in the first place? The other thing is we benefit a lot from being able to just give optionality. So maybe today you

You love using OpenAI models, and you're super productive with them. But tomorrow there's a new model that comes out. You want to try that. Why force you to go and completely change your setup just to try a new model? And then we benefit from the feedback that we didn't get, which is maybe there's something that you liked about that model. Maybe it actually didn't work well. But it's sort of like being nice to our users and to the community feels like the right thing to do here.

And then, you know, we also, we will also try like other models, right? So, and, you know, we tried them in the same harness and, you know, it's just all good. And then this is also often like, this optionality is very important to companies that we work with. This is something that, you know, we absolutely lean into. This last point, I think, you know, as any serious company, you want to have optionality and you want to use a tool that gives you that optionality. But I kind of appreciate it because it feels to me like it's kind of honest, like, look, like, it forces the whole company to compete the best in everywhere in the model layer and the harness layer with open source with with chooseable models. And it kind of like doesn't doesn't allow you to like kick back and say like, all right, we're done. We can we can we can hang back for a little bit for now. Yeah, I want us to win users by having, you know, the best models, the most efficient models, the best product. And then, you know, if we do all of these things, it's like, we're going to have a good time. If we sort of like force you to use the product because, you know, this one thing is just like, then I don't think that will attract

I love the idea of winning based on merit, not based on lock-in. And this is a perfect time to mention our season sponsor, Entire, who also play by the same rules. Like it or not, Git is becoming a bottleneck for modern, agent-heavy software development.

Devs are creating more code with agents, these agents are pushing more code, many Devs are running more parallel agents, these are pushing even more code. GitHub is clearly struggling to keep up and has frequent outages. So, what's the solution? Entire was founded by GitHub's last CEO Thomas Dumka, and he rebuilt Git hosting for the Agenda care from scratch. Entire was built to be very fast, and to have your repos regionally close to you to reduce latency, allowing for fleets of agents to push in parallel.

Some numbers they published. Entire can handle 418 pushes per second. That's up to 89 times faster than every competitor on the market. When GitHub is down, you can still keep working and you don't even need to migrate early from GitHub. You just sign up to Entire and the platform mirrors your repo. And one more neat thing. Have you ever wondered what prompt resulted in this specific code being generated? I find that the prompt and conversation with the agent carries more information than the peer itself, at least for me.

Entire captures all the prompt history with your agent, right in the repo, easy to check back. And has a pretty innovative UI to show all of this. If you're looking for Git hosting that works even when GitHub is down, head to entire.io slash pragmatic, install the CLI, and mirror your repo with a click, I've already done it. Oh, and did I mention that it works with any agent and its open source? I'd also like to mention our season sponsor, Antizizis. Teebo talked about how the experience of the software you use should feel delightful.

Delightful includes no annoying bugs. But when you're using agents to write your code, how do you avoid shipping bugs? Reviewing every line of code is becoming a challenge with the amount of code that agents generate, which is why Anticisis goes well beyond code review. Anticisis runs your whole system in a hostile simulation. This simulation includes both targeted testing and fast testing. By running this simulation, it finds every bug before your users do. And because the simulation is fully deterministic, doesn't only find bugs, it gives you a perfect reproduction of every issue, which makes it much easier to fix issues. The first thing I thought when I heard about Anticisys is that automated bug discovery and fully deterministic testing sounds like science fiction, but it's actually hardcore engineering under the hood. Jane Street, Flying.io and the Etsy the Community Ship Agent written code with full confidence because they know it's been verified by Anticisys. To see more case studies and details, head to Anticisys.com slash pragmatic.

And with this, let's get back to Tebow, and why competition between tools is great. Yeah, and I think as an engineer, I always see that whenever there's competition, as someone who's using tools, it's always amazing. I remember when Microsoft had, with JetBrains, with the IDU wars, and then there's the clouds battling with each other with all the features. And now, of course, we have the harnesses, we have the models. And as a user, it's great, because now we have more choice. They just develop faster.

I guess our voice gets heard a bit better, so it's great to hear. Speaking of the harness, can you tell me how it works today in the sense of like, when I start a codex task, does it run always on my machine? Does it choose the cloud? Does it use a sandbox? And how do I control this or know this or how much should I know about this as an engineer? Yes. So by default, it runs in, it runs sandboxed.

everything that if there is a command that should run with additional permissions outside of the sandbox, it will ask you as a user for permission. But everything, every tool execution happens within the sandbox by default, and it runs entirely on the machine, your local machine. And this has been the case for more than a year now.

But it is something that is evolving and shifting where like you can select to run this in the cloud, which then runs in like a managed VM, where it's the same VM that you get through chargeability work. And you can sort of like inspect it, but like it runs in a Cata container, it's like a secure environment. And so everything runs inside of that VM and it doesn't run on your machine. And then the only thing that happens in your machine is like the, your input and then the streaming back of the output. And so that obviously then is much nicer on your CPU and your machine and you can scale up much, much, much more. And this is just a step. It's going to be much more seamless in the future to use cloud machines and then maybe have a combination of partial execution on your laptop, partial execution on cloud machines.

And really the thing that we're thinking about that is very natural is, as models just get better and more capable, they can leverage so much more compute and many more resources than are available on your local machine. And so it would be a constraint at some point to just limit execution on your local machine. But one thing that is great about it running locally, and I think the reason I love it when it runs locally, of course, is the pain because If I'm doing some work, it's like, you know, I have several agents. It's eating CPU. If I want to close my laptop, I cannot kind of leave it like half open, right? When I was in one of the offices of an AI company, I had a little half open and they're like, are you running agents? I'm like, yeah, I have one. But the reason, the reason I do it because I have my local tools. I have my local Postgres database. I have my, my this, this and that. How are you thinking about the cloud is amazing, but it doesn't have this setup or it's just a pain to set it up.

Are you thinking or are you experimenting with making these setups? And I'm kind of reminded of a topic that we talked about pre-A, which is cloud development environments in like 2022, 23. They're hot and then we talked about AI more. I think outside of large tech companies, like cloud dev boxes really never took off because there's a very big upfront cost. And then you need to pay like a maintenance cost as well and you just don't benefit from it as like a a solo developer or like a small team with the level of capabilities that we have in agents now is like the setup almost is free, right? So like this setup cost and this maintenance cost is like, if your agent is capable of doing it, you know, you should just do it for you. So for example, if you're saying like, hey, you know, I have like, I have my local SQLite or I have a local server and MCPs and whatnot, it's like, how hard is it to actually configure exactly the same setup and keep it in sync?

on a cloud dev box. Well, maybe it's not that hard if the model just does it for you. And so I think we're going to see a resurgence of, you know, fully cloud orchestrated machines, which then frees you from your laptop, right? It's like one thing that we've, you know, a ton of success with your work is like it's just available on your mobile. I started my day just dictating a bunch of tasks into it next to the coffee and it just does it. It has access to my calendar. It has access to my email.

It has access to Slack and it's just so awesome to just be able to walk around and get stuff done without having to carry my laptop everywhere. And I think it's the same thing, we shipped Codex remote where execution is still happening on your laptop, but it would be wonderful if you didn't have to keep your laptop open. Can you tell me a bit on how in the past, how did you improve Codex? Because I remember when I first used Codex, this was one of the early versions.

You know, like you could talk to it, it did stuff. But for example, I said, like, all right, make this change. And it did that change. And I had unit tests and it didn't run it. And then later, a few months later, I don't know exactly when it just started to run it automatically. Were these things, did you improve the, you know, the script that runs, you know, the instructions? I'm not sure how exactly you call the, you know, the bootstrapping script or whatever that is.

Is it improving the model? As a dev, how can I imagine you making each version better between the harness and then between the model and what's the connection between the two? Yeah, this is a good question. So the harness, in a sense, is always a little bit ahead of the model. Oh, really? How so? What I mean by that is that you have the model, it's capable of certain things, but then you...

you set it up with like a couple of crutches so that it can actually do the thing to a level of reliability and in a way that is like efficient and also with the behavior that you expect as a user. And so that's the role of the harness, right? It's like, you know, provide guardrails, like safety, make it more efficient, make it more like steerable, controllable. And then the harness usually is also responsible for, you know, what we call like the, developer message, which is sort of like infected in the context at the start of each turn. And so that affects obviously like the purpose that I do to affect like the behavior of the agent throughout the turn. A lot of what you have is like the result of the harness and the model. Initially, like maybe you're like, oh, it doesn't run tests. So, you know, you have to remind it to run tests. And then, you know, we train a battle model that is

just capable of better reflecting on what is it that you really want when you ask for something, and then you don't actually have to tell it anymore. So over time, what we see is the system, the developer message shrinks, and then the harness also shrinks. Inside of the Codex team, do you have specific goals? Do you say like, all right, right now the Codex as a harness and model combined is not very good at this?

it's kind of doing silly mistakes here. Or how can I imagine how as the engineering team, how you're working on the next, you know, version of, of codex is, because the thing that I don't really get as a, as a dev is like, okay, there's a model, which to me is this magical thing, which will get better. Of course, I'm sure you have some feedback channels, but you also have the harness, which is the tools that you're building. Like that's probably what the team is responsible for. How do you set even your goals, right? Like in traditional software, you'd be like, we will build this feature and you build that feature because you know how to do it, but it feels a bit more fuzzy to me.

this development process? Yeah, it is. And it's why we co-design most things. And it's a process where it's a collaboration between research and the engineering team, like primarily building the core agent harness. It's always a question of like, OK, we see today that we are very good at this, but we're not very good at this. And we have a desire to do another thing because it would be a very cool products feature.

So like, look at it as like, okay, should this be like a harness change or should this be a model change? And if it's a model change, like, how soon can we have it? Can we have it in a month? Can we have it in, you know, three months, six months? And we sort of like work through that. And then depending on, you know, how soon we can just fix it in the model at which level of training, then we might decide to not even do something in the harness at all.

and not just wait for the model to solve it. It's agents all the way, right? So we use agents to analyze a lot of the feedback, to come up with themes, to just help us have these conversations and decide on priorities. But we analyze it across all of coding. We analyze it across all of non-the other domains, like finance, comps, marketing, all the things where our users are using these agents nowadays.

And there's subcategories within those. And then we roughly know how well we perform. And then we're always pushing the frontier. And there's a thing that is interesting is as we make, as our pre-training model gets better, as we make the overall model better, the whole thing lifts up. But then there are sometimes things that we pay a little bit more attention to. You mentioned you analyze agents all the way. Can we talk about the software development lifecycle on Codex in the sense of whenever a new engineer joins a team, any team, it's like, okay, how are things done here? And you know, back pre AI, it would have been you joined a company like Uber or Google, and they would tell you like, cool, the way it works is we have an idea or the PM has an idea. We make a plan, we get together, we do some estimations, we break up the work, we code the work, we do tests, we do code reviews, we release, we do feature flags, and then, you know, we were on call. That's how it used to be. When someone joins the Codex team, you know, they,

because they've been contributing to the open source part. But what do you tell them? How do things get done here? If they're like a total newbie. I introduce them to great people. And then the thing that they hear the most about when they have a question is like, have you asked Codex? And Codex is just by default, I'd open the eyes plugged into everything. So it has access to Slack, it has access to all the documents, access to all the code.

still surprises new starters that you can basically ask at anything and it will very often just like come up with like a really good response. And so the easiest way to understand the state of a project or who's working on something or why decision was made is like critics knows about it all internally. So you just use all of that. We do a lot of work in you know for that reason we do a lot of work in public channels.

We open up documents with fairly broad permissions and so that everyone has access to this information as well, and so that your agents can go through things and reason through things. And then we have a couple of other things that are just really very helpful for team productivity and team collaboration that we haven't released yet, but are going to come, like some of it at Dev Day, although that just makes you very grounded and in tune with the rest of the team.

And allows you to just very, very quickly understand the state of things and produce things yourself. The general recommendation is just like, hey, care about the user, care about the coherence of the product, care about the models and where they're going. If you're doing something and you're building like this 10,000 lines of code crutch to work around the model flaws that you're probably doing the wrong thing. So we have a set of principles, but it's just really.

sort of like a team culture and ethos at this point. And you know, it's just very much sort of like carries on, you know, when people join, it's just like through the rest of the team, just like, you know, sort of like teaching their roots. And then when I have an idea, I think it's a good idea. I talk it through with Codex, maybe I talk with it with some of my colleagues, like here's a cool new feature I'm going to build as my first contribution or first major contribution to Codex. How do I go about that? Obviously, I code it down with Codex. I obviously test it and make sure that it works.

From there on, what's the process? Do you still have the concept of code review or AI code review of verification, of rolling out, of verifying, of stage rolled out, you know, the thing is because Codex itself, it goes out to millions of people, like I just crossed a big 20 million active user mark. But if it's chat GPT, then it also goes out to like even a lot bigger number of people. Yes. But it's surprisingly like a similar process.

whether you ship on Codex or Chatchapiti, even though Chatchapiti goes out to a billion active users and growing, you can ship a PR. You can make a change and get it shipped the next day or even the same day, and it just goes out to a billion users and it's fine. We just really instill a sense of ownership and care. So people are very empowered to make changes, even large changes. The general thing that is being asked is sort of like...

evidence that it's going to be well received, evidence that is like a worthy addition, evidence that, you know, it's like it is worth maintaining over time, but also like the cost of maintenance is like just really, as you know, gotten done significantly as well. So we think about these things slightly differently than you know, say like two years ago or three years ago. The other thing as well is like, you know, we automate as much as possible. So like a lot of like the process of like code review and deploys and you know, catching regressions is like, you know, all of that is like pretty much automated.

And so you get to just focus on just really the idea and how it's going to help our users. And you care about the coherence of it all and the overall power of the agent and making things better. And we have a long, long list of things that we aspire to do and haven't gotten to yet. And then there's the north star direction, which is a delightful, simple to use personal AGI.

that knows everything about you, that it needs to know, has access to the right resources, can take sometimes risky actions on your behalf, but then you get the push notification and then you can verify that. And it's a thing that you deeply understand as a user, but also it knows about your schedule, it knows about your goals, it can be proactive, and it should be extremely natural. It should be something that you can control through.

natural language, voice, maybe it should understand your emotions. If it has a camera, it should be the most natural thing on earth. It should not be a thing with 10 different buttons and configurations. AGI should be simple to use. You mentioned briefly the review, the code review, but I wanted to go back to it. You worked at Google on a product used by hundreds of millions, which is Google Maps, and Google is very well known for their culture of...

a very strict code reviews. They have, I think, two layers of code reviews. There's a language correctness review, and they've taken, I think they've really perfected it across industry for a long time, and they do believe that it works and they use it. How do you think that part is changing specifically the human review? Because for a very long time, until maybe a year or two ago, I would have said your code review has all these benefits, knowledge sharing, the second pair of eyes, removing the bus factor, because now someone else understands them when that person is out, that they can jump in.

Conversations are happening about architecture, not just a code. But now there's a lot more code. And what was the value of code review? In what cases? And so on your team, because you guys are so ahead of this, where do you see humans still being or developers being involved in the review stage valuable? And where is it fine? Did you find it fine to hand it off to an agent?

Yeah, the role of code review is changing. One of the early projects that I did on codex was like working with research on developing a code review model that was going to be to a level where it can spot mistakes in logic and reasoning to a degree where it would require humans like, you know, multiple, potentially multiple hours to catch the same level of mistake because It requires like really digging like, you know, three, four levels deep into like the dependencies and like, you know, just understand that maybe the documentation actually was wrong and like the implementation of like the third party dependencies like different from what you expected. And so therefore your invariants are not upheld. And these things is just like, you know, unless you're an expert in that library, you wouldn't know. And therefore you have a bug. And so we developed like these code review models and, you know, we released them. And now they're like the same level of like capability and like ability to

spot these mistakes by doing like, you know, deep verification are like just part of the mainline models. Like when we benchmark them, it's like they're like super human in code review. And this is not just true for correctness. This is also true for security, for example, where they're capable of like reasoning across like, you know, very, very complex things. And then, you know, coming up with like, Hey, you know, you have a critical security vulnerability here, which is now mandatory across like all of open AI pull requests, like we block pull requests from merging if, you know, we flag them with like a security issue.

And this is like all automatic. And the role of code review now is like, I think it was always about correctness. It was always about, you know, ensuring that things work, but it was also sort of like a little ritual for information exchange and, you know, bringing people on the same page and like, you know, encouraging like a discussion, which ideally would have happened before, but sometimes it just only happens like around the code because once it merges, it's just actually runs in production, it's doing stuff. And then you have to maintain it. So there's like this social aspect to it as well. And so I think.

all of it is changing. The correctness, the security, I think that will be automated. Really what we see and I see is there's this discussion around the intent that takes place around the pull request. What are you even trying to do? And is that a right thing to attempt to do? I think you can have that discussion outside of the pull request. It doesn't have to be around code.

So maybe this helps crystallize where a discussion needs to happen versus where we did it because maybe we didn't have the type of tool that we have right now. Yeah, I think this is going to change. And it was like a forcing function because you have to have that discussion where it's good to have that discussion before you merge it and it becomes production code. But I think there are other ways to have these discussions and design things together and make sure that the intent is good. And then the code doesn't matter as much.

And it's interesting because when I think back of all my code reviews, like of course I have like memories where like it was great. We had a good discussion or I learned something really interesting, but a bunch of times, honestly, it was such a pain in the ass. Like I was trying to get my stuff. You're paying. Hey, could you read my code and like, no, right now I'm busy. No, I really need this to unblock me. And then you context switch. And then I feel it's always been like.

good and bad right so i feel whatever we do there will be always upsides and downsides but now they're just moving so i guess one upside is as an engineer you might have to not give your attention to just kind of basic stuff that doesn't need your input per se. Yes it saves time and progressively what we're going to see is also.

like you have an agreement on, you know, the box and the overall contract of what it's supposed to do. And then, you know, what is inside the box, as long as you have like strict guarantees in terms of resource utilization, data access, security, these kinds of things. It's like, what happens inside the box is, you know, it could be literally anything. It's like, don't really need to care. And like really what you need to agree on is like, what does the box actually do and what are the invariants that must be satisfied. And I think that is then worthy, you know, having like a really good conversation on, maybe assisted by your favorite agent. But then once you have that and you have that understanding, it's just like changing anything within the box doesn't require for discussion, and it just really preserves your attention. The cost of maintenance has gone down. Maintenance is always such a hot topic whenever we build something inside of all these companies like Google, Uber, or even startups. Building was the fun part, but then maintenance was the painful, and that's when we learned, okay, we're building it, et cetera.

Inside of Codex and OpenAI, what do you see maintenance becoming cheaper changing in terms of instead of what you're building, what the ambition is, the I guess custom tooling, those kind of things? Maintenance is really like sort of like a tax that you pay over time just to keep things running and it's always been necessary. It will continue to be necessary, but where I think it changes is like a lot of it is just going to be automated.

you know, it's like, okay, you have you have this third party dependencies, like you need to upgrade the version. Remember, it's like, oh, yeah. You can fully automate this, you know, if you have good change log and, you know, and the code is well documented and like, you know, and the model can just like reason through it. It's like, you know, I can just like blast through your code base to it in a couple of hours. And, you know, previously you would have like sort of punted on it because it's not the most fun thing to do. But it's actually really important.

for your business. So it's really important for your project, especially for security vulnerabilities. You want to stay up to date. You want to apply all these patches. I think that's just going to be fully automated. So a large part of maintenance, it just comes for free.

I think it's awesome to also think about before when you wanted to just completely react, you have to do like a new architecture because you're trying to make space for like a new different kind of trade-offs or you have a new understanding of like the workload or you're trying to fit a new feature. And suddenly you realize like your current system is just very limiting and you need to completely re-architecture it. That was like a really, really costly endeavor. So sometimes like multiple years. And I think this is also like super, super accelerated now.

So like the cost of mistakes, you know, I would say like, you know, is going down. But then at the same time, the good old rules, I would say, of software engineer of like, you know, having good abstractions, like really help, like, you know, is going back to this, like having the box with invariance, like, you know, if you sort of like draw the right shape, you're going to be able to change things much more quickly within the box and like not affect the rest of the services or the rest of your infrastructure. And I think it's important. It's important to design for very quick iteration and change.

I remember when I talked with Peter Steinberger, that was before he joined OpenAI, but about open claw and how he thinks about it. Like, you know, he told me that he doesn't read the code, but he kept thinking about like, I could see that he's holding the architecture in his head, and he was telling me how he rearchitects a lot. And he thinks about how to make it modular, how to allow 100 contributors to each build their thing without stepping on each other's toes. So I'm hearing what you're saying that this.

This care this this planning this this structuring has become maybe just a lot more important to Like which which was which was something back in a day You know it was like the architect or the staff engineer or experienced folks were doing this thing and other Engineers around them were kind of building this you know smaller parts but sounds like now all engineers need to Be aware of when you're building your software right and plan for it Yeah. And the GPT models are getting better and better at this as well, thinking about long-term maintenance and good architecture. And this is a natural next step, right? It's not just about code quality in the sense of, oh, is this code clean within this file? But is architecture actually correct to reduce maintenance burden over time and make space for future product or future extensions or changes? And just really this act of engineering over time.

kind of like something that models are starting to become capable of thinking about very well. I think it's just kind of fascinating to understand that the software that we're building is just going through the life cycle much, much faster, right? You know, before you had, you know, you were, you were scaling it, you were starting it, you know, maybe as like a small team of, you know, yourself, maybe a couple of engineers. And then you would add engineers like slowly. And then, you know, maybe after a year, you know, it's like, if it's very, very successful, you would have 50 engineers on it or like a hundred engineers on it.

you would have time to see it coming. You would have time to see like, you know, the humans on board and you can think about the documentation, all of that stuff. But now it's just sort of like that explosion of like, you know, suddenly you have like a hundred agents contributing to this thing is like, you know, that can happen like, you know, in a weekend. And so, you know, you're just going into it at, you know, major, major speeds compared to before. Okay. But how do you and the folks at Open AI like deal with it? Does it not mess with your mind?

Like, you know what I mean in the sense of like you've been in this business for quite some time now, like decades or well over. And there was a pace that we kind of got used to. And obviously it's now a lot faster. But how do you get your head around the fact that, A, it's faster, B, the stuff that you've been doing a year ago, right now you're not doing because now the model is good at it.

How do you reconcile that? Because I'm sure there's stuff that you've been really good at related to software that now you can hand off to the agent. Do you not get a little bit of sting? We talked about it stinging for your features to be implemented open source, but it can also sting that I've been really good at refactoring or right now it might be architecture, but maybe the model will be really good at that. And now I'm like, okay, damn. I'm glad, but also it would have been nice for me to do that. Yeah, I think there's a craft aspect to it, which Occasionally, I still, you know, pull up an editor and like write some code and it's just like, it feels nice. And it's sort of like, I have fond memories of like late nights sitting in Vim and, you know, just like... Cranking it out. You know, drinking Coke Zero and, yeah. Just not having to think about anything else other than like the problem in front of me. But really...

I think it's all about being in the flow and solving problems. And what I find is folks here and also everyone I talk to is just adapting very quickly. And I think if you have a mindset where it's all about code as a tool to solve problems, and you can solve so many more problems, it's like before you wanted to benchmark something and you weren't quite sure where you were going to net at.

You can just do it. It's going to take you no more than 30 seconds to launch something in the background and get proper numbers and be able to do a better trade-off. It should make you a better engineer if you just really care about the outcome and the system working well. And so what it allows us to do at OpenAI, it allows us to run our inference much more efficiently. It allows us to get much more effective compute and deploy that to the world.

But everyone's just very focused on that and solving important problems at the speed that was not possible before. And I haven't yet encountered someone who's like, oh, that's not good. That's not fun. Do I understand correctly that it sounds like if you have ambitious problems, if you have way more problems than what you can solve today or tomorrow or the next week, sounds like this is not really a problem. Because when you get more efficient somewhere, you keep going. Which is a lot of startups, right? Like startups are always way more ambitious than.

That's what they're able to do. We're not out of problems for sure, right? And I don't think we will be for a while. We have a long, long road ahead of us in terms of like, mathematical breakthroughs, scientific breakthroughs, you know, making the world a better place, like just really building for humans and solving the most important problems that everyone is facing and just doing it in a deeply human way. That's what we're here for.

Also just going back to coding and like you know these late nights is like I think there's like it's also like maybe like a Glamorous version of it just like I also had very a lot of late nights where I was trying to refactor something And you know, it's just like it would be like three hours deep into the refactor and then realize like actually this is a dead end And I must restart from scratch and it was like very frustrating and so it's like there was like There are like these very very fun times that there's also the time where it's like It doesn't compile in yours. Why is it not compiling yet? I'm sure you had the time where you go, you go later, it's now super late. You need to go to bed because you need to get some sleep and then you can't really sleep. And you have this thing where you have some tasks that is halfway and it upsets you. Sometimes I remember dreaming about the code as well. And I guess one thing I don't really have these days when I'm working on my software for my business.

Is I don't really have something that is halfway because I can just tell it do this and then I can leave it as a state where it's kind of like, you know done either finish or it's either working or it's I have proof that it's failed but it's interesting because you know everything's sped up, right? Yeah, I may be like I do have like what a lot of people do and I do myself is like, you know, I have like sometimes like bigger questions that I'm asking myself like and I you know from conversations I've had during the day or like I haven't yet, you know just like had the time to just look into it and so

I will send off Codex to just look at it overnight. And then I'm very excited to then wake up and look at the results. And so it's always like an exciting morning. Well, I feel there's an arch to doing long running tasks. And of course, you can use the slash goal, which will go and run. That's also something that was recently added like a few months ago, right? The slash goal command to Codex. Yeah. And back to maybe like the harness is a crutch, right? Is a slash goal was like...

necessary to allow, not to keep the model on track, on a singular goal for a very long period of time. It allows the model to literally run for days or weeks if it's a really, really hard problem. But with the new generation of models, what we're seeing is, you don't need a slash goal anymore. You don't need a harness around it. You can just tell the model, hey, go and work for a week, and it will actually do it.

Speaking of hard problems and the fact that you're not out of them one of the interesting things that you shipped from the outside it I would say it it was you know as an engineer was moderately interesting is the what you call the merge which is Codex appeared inside of chat gpc and the reason I say that's as engineers it kind of moderately interesting because we've been using codex like yeah it's there you can now open it in the chat gpc upgrade like I just went there and I just immediately went to codex because I don't I don't really use chadgbt in the app per se. But I talked with folks at OpenAI and people in your team, and they were telling me there was a lot of preparation going on, a lot of engineering challenges. Can you give a sense of how big this project was, what you needed to do, and why was it difficult to pull off, and how did Codex and other tools help you get it done in ways that would have been hard before?

Since you've launched the merge the the numbers that you keep sharing of how many people use codex It's like it's going up way faster than before. So I assume there's a big scale problem. You've solved here a lot of things we're Challenging with the merges first of all completely different stacks charge dbt is like fully Managed cloud-based like, you know, you run everything on our on our systems. We store things like traditional way of like building things, built for scale, built for efficiency, codecs fully local. And so the mergers just really like how do you get the same benefits and the same capabilities from this local coding agent and then build a product around it and build it in a way where it can benefit like a much, much broader pool of people, which is just also why, you know,

all of us joined OpenEI to benefit this very, very broad population across the world. And so it was a very exciting journey of figuring out how do we build a cloud version of this that, in essence, is capable of very, very much the same things, but is also built in a way where we can serve it through tens and hundreds of millions of users in a way that is still efficient so that we can include it all the way into the plus plan.

ChargeGPT work is essentially running the full codex harness together with a cloud computer. It's a very powerful machine, actually. People have picked up on it and showed what you can do. If you are creative with the prompt, you can get ChargeGPT to that.

train another model in there. Wow. There are some pretty wild things. You can get it to install Blender and do 3D modeling. It's very permissive. It has internet access. It's a powerful machine. And then Codex just works on it. And this is what we ship through. So a lot of system challenges. The team did it very quickly. Obviously, Codex helped to make it more efficient.

look at and build a lot of the infrastructure, and then help resolve a lot of the little differences as well that had been occurring between Codex and ChatchaBT, like merging plugins, architecture, merging library, and really, really working towards a unified system, which is really the goal. You shouldn't feel like you can do something in Codex that you can't do in ChatchaBT or vice versa. What we're trying to build is one unified product.

that gives you access to the same intelligence but in the way that you want to use it. And so it was very fun as well because Codex throughout the whole journey also acted as a journalist to sort of like document all the steps and the debates and the discussions that the teams were having. And it was very animated debate, you know, of how we should do it and how we should name the thing and, you know, when to introduce it in one way and like what to merge into what. There were like many different permutations considered. And so there's like a very fun journalistic element to it, where we have a full recounting that could exit over time. And yeah, it's kind of become known as well as the toggle arc of OpenAI, where we introduced the work toggle, which there was also a lot of debate around whether this was the right thing. And then we kind of grew to just really like it. But over time, we're going to merge things further.

headed into the direction of like full unification. And, you know, we kind of view this as like a temporary state where, you know, you have like, you have better or stronger capabilities when you're in work mode. But over time, we're bringing this, you know, all all the way through like, you know, everyone that uses chance. And how do you personally use codecs? Like, what's your, what's your working setup in terms of agents, in terms of tasks, in terms of what you what you manage with it. And related to this, I asked Peter Steinberger what I should ask about you. And he said like, you need to ask him how do you deal with the fact that you're involved with all these projects? Your calendar is like Tetris, but usually you show up pretty cheerful. My calendar is fine. And it's just, I am capable of doing so many more things nowadays because I have the technology like Codex and

I actually shifted a lot of my work on mobile using Charge of Duty Work. Whenever I have something that I want to take note of, I just fire it off. I use dictation a lot. Whenever I have a question, instead of writing it down to look into later or delegating to someone, I just fire it off in Charge of Duty Work and I get a report. It has a whole bunch of custom skills.

custom instructions where it's now like very, very tailored to like, you know, produce the kinds of reports and slide decks and code explorations, you know, in the style that I can consume effectively. And so every time I'm like between meetings or like, you know, you'll kind of like see me like, you know, I was just like dictating through my phone. As I said before, it's just like, we do a lot of work in public channels. We have like a lot in Slack, we have a lot in Notion and Google Docs as well. And so there's pretty much like, there's no question really that I feel I cannot ask that Codex will be able to do at least a first pass of thinking through, whether it is public sentiment on a feature, looking at production logs for how much usage we have on a certain thing, making a list of things that we should deprecate because they're not getting traction, understanding what a certain team is up to. Any question I have, I can get an answer to within 30 minutes. And so that's how I use it. I use it for everything. It's like my personal...

uh agent in like all the ways and then oftentimes on on weekends as well I do some like code explorations or like I built some prototypes and I have fun like sort of like imagining the future of the product in some ways and I do that with others on on on the teams it's not always the same team and it's just like in one day I can build things that I sort of like I had it in my system right it's like it's like I woke up one day I was just like we should explore what it means to build this and then I can just sort of express all of that and get like something in front of people in a day so that they can think through it and criticize it and hopefully get inspired by it. It's like by no means, you know, we need to ship it, but it's more like, okay, I flush it out of my system and then, you know, I go on and like, you know, do other things. So it's just like so, I know it's such a magical time and it's like so empowering. And as closing, what would your advice be for a software engineer, someone who builds software who would want to

get the skill set and the experience to have the opportunity to work at a place like the Codex team, like OpenAI or like an AI startup. So like, you know, just become this really great builder with these tools. Because the question that comes up is often like, should I start with the theory? How important are the basics? Should I just get really good at using the tools? Yeah, I think there are two things that are important is deep, deep curiosity for how things work.

and an ability to like, you know, train yourself to understand things very quickly. And so it is the case that things will continue to change, but people that do externally well at OpenEI are like, you know, people that just sort of like are able to like grok a system quickly and like, you know, also dive into like a new code base and sort of like, you know, make sense of it. But obviously, like all of that is helped with agents nowadays, right? So.

just like there's so much information that you need to absorb and that you know being able to understand and reason through it. I know a lot of that is asking good questions really about you know how do things work and just like going into like the five whys which I think you know you can just kind of keep digging and digging and digging and you know you're learning very very fast through that. The other thing is being in tune with the community or you know the people that you're trying to solve a problem for. It's like not everything is like solving a direct problem sometimes you're solving a problem that will be useful you know to like another group of people.

in the pursuit of solving a problem for humans. But just being crisp about the taste or the needs or the requirements and being able to think clearly and exercising through this clarity of thought feels really important to me. If you can't explain what you're trying to achieve, if you can't explain your intent, if you don't have a tie to a community, if you don't have the taste, it's going to be much harder to do great work. That was some tea. Well, thanks so much for this conversation. This was awesome. Thanks for having me.

I've always wanted to get together with Tebow and I'm glad that we finally made it happen. I appreciated how Tebow talked about not just the upsides of open source but also the downsides. Most notably how competitors can copy features you are just working on in the open right now and then ship it right before release and just how much this things. Plus you get a lot of low quality contributions that you still need to somehow deal with. Another interesting one was Tebow saying how the hardness is always a step ahead of the model.

From the inside, the Codex team see their job as building clutches for the model with the harness, the tools, and the setup instruction. And then the next version of the model will be trained to need fewer of these clutches.

I'll be honest, as a dev this sounds a little demotivating that the stuff I build in the next version of the model it'll just know and we can get rid of it. Plus I do suspect that it's not just about building these clutches but also building tools that models will use and it's not like the next version of the model will reinvent an mcp protocol or scales or plugins, at least I hope not. I also enjoyed hearing what the merge, merging chat gpc and codecs look like from the inside.

it was merging a previously fully local coding agent, Codex, into a managed cloud-based stack and doing it efficient enough so that it can be included in OpenAX $20 per month plan when $20 is not all that much in terms of compute purchase. It was pretty amusing to hear how Codex itself acted as a journalist of the whole project as it was present in all the Slack conversations and all the documents and so it could capture all the important debates and decisions. I'm not gonna lie, this part felt a little bit of a big brother feel to it.

where the AI is always watching, but it could well become the new normal in startups in the future. I have not yet decided how I feel about this. And finally, I appreciate Tebow's advice for engineers to succeed, be curious, understatt symptoms quickly, and be in tune with the group you are building for. It's reassuring to hear from Tebow as well how much the fundamentals still matter. Do check out the show notes below for deep dives on how codecs, clock code, and cursor were built, and other related topics. If you like what you heard, Please hit a rating on a podcast player that you're using. It means a lot to me and to the show. Thanks and I'll see you in the next one.

Delete this episode?

This removes the episode page and its saved audio from this library.