In short
Podcast Episode Summary: OpenAI Codex Team - From Coding Autocomplete to Asynchronous Autonomous Agents
Overview In this episode of Training Data, hosts Sonya Huang and Lauren Reeder engage with Hanson Wang and Alexander Embiricos from OpenAI's Codex team. They delve into the advancements of Codex, OpenAI’s AI coding agent capable of functioning independently for up to 30 minutes, generating complete pull requests from simple task descriptions. Key topics include the evolution of AI from coding autocomplete tools to autonomous agents, the training required for real-world software engineering tasks, and the implications for the future of coding.
Key Takeaways
Evolution of Codex
- Transition from Autocomplete to Autonomous Agents:
- The original OpenAI Codex focused on competitive programming and line-by-line code completion.
- The latest Codex model (Codex1) is designed for day-to-day enterprise development tasks, emphasizing real-world application rather than competitive coding.
- Codex’s Functionality:
- Codex operates in its own environment, allowing it to work independently and produce pull requests based on the tasks assigned.
- It aims to transition developers from “pairing” with AI to “delegating” tasks to autonomous agents.
Technical Challenges
- Long-Running Inference:
- One of the main technical challenges involves optimizing the AI's ability to handle tasks over extended periods (up to 30 minutes).
- The team focuses on training the model to understand the preferences and styles of professional software engineers to produce mergeable and quality code.
- Creating Realistic Training Environments:
- Emphasis on realistic coding environments during training to mirror real-world complexity and workflows.
User Experience and Interaction
- Delegation Mindset:
- The Codex model encourages a shift in mindset for developers to utilize it effectively, promoting experimentation and parallel task execution.
- Developers are encouraged to engage Codex by submitting multiple variations of tasks for optimal results.
- Code Review Dynamics:
- Future coding processes may shift towards increased code review responsibilities for humans as agents take on more coding tasks.
- The ability of Codex to cite its outputs enhances the review process, making it easier for developers to verify and trust AI-generated code.
Future Vision
- Predictions for Software Development:
- The future may see the majority of code being generated by AI agents working autonomously in their own environments.
- Integration of Codex and similar agents into more tools and platforms, allowing seamless interaction across development environments.
- Potential Applications Beyond Coding:
- Codex could serve as a powerful tool for non-engineering roles, such as project management, highlighting its versatility and broadening its user base.
Cultural References
- The Culture: A sci-fi series by Iain Banks that presents an optimistic view of AI.
- The Bitter Lesson: A paper by Rich Sutton discussing the significance of scale in AI development.
Conclusion The discussion emphasizes the transformative potential of AI in software engineering, particularly through tools like OpenAI's Codex. As technology evolves, so too will the roles and workflows of developers, moving towards an era where AI agents significantly enhance productivity and coding capabilities.
---
Recommended Readings
- The Culture by Iain Banks
- Works by Richard Sutton on reinforcement learning and AI principles.
Favorite AI Applications
- ChatGPT for various tasks
- Waymo for robotics applications
---
This summary encapsulates the key themes and insights from the podcast episode, serving as a concise reference for stakeholders interested in the advancements of AI in coding and software development.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00In my opinion, the easier it is to write software, then the more software we can have right now, we think of like, I bet you if we look, pull up our phones, well, you folks are investors. But if you're not an investor, I bet you if you pull up your phone, most of the apps on it are apps that are built by large teams for millions of users. And there's very few apps that are built like just for us and the specific thing that we need. And so I think as if it comes more and more practical to build like bespoke software for people or teams, we'll end up having higher and higher demand software.
0:35for.
0:43Welcome to training data. Today we're joined by Hansen Lang and Alexander Embarecos from OpenAI's Codex team for a fascinating look at the future of software development. Codex is OpenAI's series of AI coding tools that helps developers delegate tasks to cloud and local coding agents. Unlike the original OpenAI Codex, which was developed in 2021 to autocomplete lines of code, the latest evolution of Codex can complete entire tasks for you autonomously in the background. The key difference between O3 and Codex is that while O3 is great at competitive programming, Codex has been RL -Tune to be great at day -to -day Enterprise Development Tasks.
1:21Alexandra Hansen share more about the backstory for Codex and the brother paradigm shift from snappy auto -complete to longer running background agents. Plus they share their surprising vision for how developers will interact with AI in the future as sync and async experiences merge. Hint, it might look more like TikTok than your current IDE. Thank you guys for joining us. It's wonderful to have you here. Hey, thanks for having us. Chris, be here. We'd love to hear a little bit more about what you guys work on. Tell us about the Codex team in your story. Well, yeah, Hanson, I'm one of the researchers that helped train the Codex One model.
1:56And I'm Alex the product lead. I think for me, the name Codex is is such a great callback to the original Codex model. That was like kind of like a hot moment for me when it first came out because I think GPT -3 was really cool. But then Codex was like the first moment where I felt like it's like wow this can really do something that is going to change the world. And it's like actually kind of like how I got into the whole like startup space. One of the first couple demos I did was using codex to do data analysis. I think it's actually funny story. I was here for as part of Sequoia's Arc program.
2:28That's how I met Lauren. When the demos we did would be actually used open at codex to do data analysis. That's how I started in the start space. I think I was time went on as the later versions of GPT came out. became super clear that using AI for agent use cases, it was gonna be the future. And so I joined the company to work on on agent coding efforts. Yeah, and this was like, a per standard open AI style where we like the naming to be as easy to follow as possible. This is the codex of like, I think it was 2021. Yeah, this is pretty much it, right? Exactly, yeah. So it was actually like the model power and GitHub co -pilot.
3:06And then recently, as we were working on this product, which we'll talk about, we thought, This is a super fun brand, also a very app name, code, codex, codex, execution. We decided to sort of resuscitate the brand and keep using it. You said resuscitate. It was codex dormant for a while, and then you all resuscitated it for a new year. We have to use the brand recently. Okay, really cool. Can you tell us a little bit about codex, the agent, and what it does? Yeah, I think so. Basically, codex is a coding agent that has its own container and its own terminal, kind of like fully in the cloud. you give it a task and it comes back to you with a PR in this sort of like one -shot style.
3:42And we actually experimented with a lot different form factors kind of along the way, but kind of in the end decided to settle on this one. Yeah, so like, you know, we've been working on a bunch of agents and we've been working on a bunch of coding products as well. And basically, in our mind, codex is like this thought experiment for how would it work to code with AI, but where we sort of put all our effort into thinking about what would that feel like if the AI is working on its own computer, independently from you. So you're delegating to it rather than pairing with it. And so some of the things that we're really proud of with this codecs launch are thinking about the compute environment and how do we set it up so that the agent can actually work on its own, but be productive, and creating the model which has to talk more about.
4:22Basically that isn't just good at writing code that looks good or is functional, but also is really good at writing code that is useful for professional software engineers and like, mergeable ideally without even touching their own computer. So what is the difference between codex and codex CLI? Yeah, we've definitely gotten some questions about that. I promise this is all going to make even more sense over time. So basically codex for us is like our brand for like agent code. And we have this vision of like, you know, we're going to have this agent and mostly the agent will work on its own computer.
4:53But it's all shall be able to meet you in like any of the tools that you use wherever you work. Be that your terminal or your IDE or your your management tool. Codex CLI is basically codex in your terminal. CLI stands for command line interface. So it's in your terminal, you can work with codex. That's your environment. And then codex or codex in chat .pt is basically codex working on its own computer. Today, those are just distinct things. As a brief aside, one of my favorite things about working you open AI is how willing we are to cut scope and just launch things quickly. But over time, we'll actually bring those things closer together.
5:25So you can really think of it as just like codex and it can be in JGPT or it can be in your CLI. Very cool. And so what did you have to do differently for the model to make it useful beyond just writing the next line of code? Yeah, so I think one of the most interesting progressions, so if you go back to the O1, the first reasoning model that we launched, we highlighted how good it is at math and even coding competitions. As of now, I used to be a competitive code and it's better than me, a competitive coding. It's better than most, almost all people at OpenAI at that. But I think one of the things that we saw was that, you know, despite being good at these programming competitions, it wasn't actually that good at producing mergeable code.
6:06And so, like, we even highlighted it in the blog post with models like O3, like, the code that it generates often, you know, like, isn't quite to the taste or style that a, you know, professional software engineer would expect. So a lot of the effort that we spent on training this model was aligning the model to basically the taste or the preferences of professional software engineers and that something that took a lot of, I guess, specialized training. Yeah, I have this very product -e analogy that I like, which is like, if you take our like, re -seagmodels, which are great at coding, they're great at coding, but it's kind of like this really pro -cretionist competitive programmer like College Grad, who doesn't have many years of job experience being a professional software engineer at like, on a team.
6:48Right? And so a lot of the work we did to go from like, O3 to like, Codex1 was actually like, the equivalent of those first few years of job experience, where it's like, hey, what is a good PR description look like? You know, PR titles, how do you read the style of the code base and then make sure your code is in the same style? How do you test well? How do you show that you tested well? Stuff like that. What's typically the aha moment for when someone uses codex? Yeah, I think one of the things we have in the onboarding is like find and fix a bug in the code base. I think that's one of the areas where codex really shines is like specifically like bug fixing.
7:20just because it can actually like independently try not just to see if you know something looks a bit off but it can actually go and then like verify that okay like I can try and reproduce a particular issue and so I think like even you know like leading up to the codex launch there were a couple bugs where you know like we were sitting there kind of like wondering what's going on and on it honestly like sometimes the easiest thing to do is just like paste in a description of the issue into codex and we were surprised how frequently and we could have actually end up with a usable fix. Yeah, like fun story here.
7:53Hopefully this doesn't give away too much, but at 1 a .m. the night before launch, the morning of launch, at 1 a .m. we were looking at a bug with like an animation, a lot of animation, and you know, this is the kind of thing like, okay, I guess we could cut it from launch scope and be okay to launch without it. But we really wanted to get in, and we just couldn't figure this out. And so an engineer ended up like describing what the bug was and putting it into codex and actually like a fun pro tip for anyone who's like using codex is that, If there's a really hard task, it can be useful to ask codex to take multiple cracks at it.
8:22So they pasted that description and ran it four times. Like, hey, there's this bug. We can't figure out what's going on. And three of those rollouts did not work. And then one of the four, which is the fix to the bug that we were stuck on for hours at 1am before launch. And so, landed the fix, deployed the code. And the animation was in for launch. That's awesome. Maybe tell us more about how you all are using it internally at OpenAI. like is every engineer, is every researcher using codex now in their workflows? Yeah, and actually, can I give you the other kind of magic moment? Oh, yeah, please do.
8:51So one of the interesting things about codex is that it's a very different form factor from maybe what people are used to, like a lot of the AI products that people are used to, especially in software. Maybe like GitHub, Kupa was the first really good one. There are really things that kind of work with you in flow, and you're just kind of seamlessly going back and forth. You're kind of pairing, right, and it's flavors on pairing. And we think that's awesome. And the codex CLI is a tool that you can use in that way. But for codex, we really want it to push this idea of you're delegating. Because in the future, we imagine that actually the vast majority of coding is actually going to be done independently from the human working on their computer who can only do one thing at a time.
9:32And so it'll be done by agents working on their own computer. And so that is a very different thing to delegate to an agent than it is to pair with sort of an AI model that's like in your tooling. And so you have to kind of use it differently. And so when we actually were working on an alpha before launch, we would just give this agent to people and be like, hey, like just use this however you want. And we noticed that many, many of the people trying to use our alpha of codex, we're just like not really finding it super useful. And then we're like, huh, that's interesting. Let's look at how people like add open AI are using like internal tooling like codex.
10:08And we realized there was like a big difference which is the mindset of using it. The mindset that works really well for codex is like kind of this like abundance mindset and like, hey, let's try anything. Let's try anything even multiple times and like see what works, it saves me time. And so we've kind of shifted the way that we even on where people into the product to try to create this aha moment which is running many tasks in parallel. So like for us, if we see someone like trying it out and like they've run like 20 tasks in like a or an hour, that's amazing and like we they're probably gonna like they've understood Basically how to use the tool fascinating.
10:41How does that change the role of the human when you have to review all of this code Like if two of the three work then want you to yeah, I think we put a lot of focus on also like making the outputs easy for people to review So like one of the things that we're proud of is like We haven't seen this in too many other tools. It's like the ability for the model to cite its own work. So not just like the files that are changed, but also even like the terminal outputs. So like if it ran a test and you know, for some reason the test wouldn't work, it actually like tells you that and it tells you like, here's the exact kind of like terminal command I ran, here's the output.
11:13Makes it much easier to verify the outputs. But it is like a great point. I think we're shifting to a world where like a lot of the time that we spend normally coding, a lot of that's going to shift to actually reviewing the code. Do you need humans to review the code? Because I think of code as one of those things where it compiles or it doesn't, once it compiles, you can go and check if it does the thing it was supposed to do. Do you even need humans to do the code review? I think, yeah, I mean, for the foreseeable future at least, I do see that to be the case. I mean, I think a lot of it's also just building trust with the early users.
11:48I think people really need to have a feeling for what things are working well, but things are not. And I think there's always just some external context about what makes this code correct that might be beyond what you initially provided as context. Yeah, if you think of what a developer does, and this is obviously oversimplifying, but there's like, do, you call that ideation, maybe then there's design, like, okay, what are we actually doing? And then like, planning, how are we going to do it? Then there's implementing, and then validating, you know, testing those changes, and that, you know, that's basically a loop, and that that small loop of like, implementing, and then testing is what code X is right at right now, although we can talk about how you can use it for planning too.
12:34And then there's actually deploying the code, and then maybe maintaining the code, writing documentation, et cetera. And so like, you know, I forget the exact exact, but I feel like that I remember recently is like engineers spend like maybe like 35 % of their time coding. It's not actually the majority of even what engineers do. And so, you know, the future that we're trying to build towards is that is one where, you know, if you're a software developer or even like in any profession, all the work that is like easily automatable, that's usually the grungier type of work, you're not doing. You're delegating that.
13:02And then the work that is more interesting because maybe it's ambiguous or maybe because it's really hard, that's the work that you're driving. So we're trying to build towards that work, that world. And I think we have to get there iteratively. So for example, right now, if you're a human in your right code, another human's going to review that code. And so we're not going to come in and just try to change that. And we're like, OK, let's plug into that. So the way the product works right now is like, you, the developer, are being accelerated by the tool. You ask for some code to be written. You decide if it's good and you want to push it out to your team, and then your team can review it.
13:34And then over time, we'll basically expand what we can do. So it'll help more and more with like planning, maybe even designing, maybe even thinking about what to do and response the things that are happening in your app or at work. And then we'll push to like make review easier and easier as handsome as describing. Yeah, and do you think I see a future where you have, you know, like multiple agents collaborating together. So you have codex, the codex agent writes the code and then maybe like the operator agents, the one that's testing it and all of the things that all the different agents that we've been working on at the company can kind of like come together.
14:04That's awesome. Have you seen people, now that you can delegate doing writing code, people beyond engineering teams start to use codex? And it was beginning to the world of vibe coding. You guys are helping us bring us further down that hole. Yeah, this is actually super funny. We were, so the answer is yes, but I'll tell you a story. We were working on our launch blog post with Lindsay here. And we were talking about like what quotes to quote from customers. And we had a customer that wanted to say, yeah, like we on the engineering team love this and also it's like a power tool for PMs And I remember looking at that quote and be like this is a really cool quote Because I'm up for it on the product team and I use it to just like avoid having to bug an engineer about things or to answer questions But I remember looking at that quote and be like do we want that in the launch blog post because the target audience for what we're building is like Specifically professional software engineers not vibe coders So I think we ended up not including that exact line, but I think over time like as As we have agents that can help us code, I would expect more and more people to be able to contribute to code bases.
15:06I think a number of professional software developers goes up or down over time. This is just my opinion, but I think it goes way up. I think a - Not vibes co -thers, professional software developers. Yeah, I think so. But yeah, in my opinion, the easier it is to write software, then the more software we can have right now, we think of like, I bet you if we pull up our phones, well, you folks are investors, but if you're not an investor, I bet you if you pull up your phone, and most of the apps on it are apps that are built by large teams for millions of users. And there's very few apps that are built like just for us and the specific thing that we need.
15:38And so I think as it becomes more and more practical to build like bespoke software for people or teams, we'll end up having higher and higher demand software. Yeah, as I think about how I use it, I think it just really is a multiplicative factor right now rather than any sort of replacement, just like especially looking at the patterns of our internal power users. There's a really dramatic difference in the top users of Codex doing 10 plus PRs every day. It's just really such a multiplicative factor that I can't see a world in which it's lowering the bar to creating software so much. That's it.
16:15I think this is a really important question and to be completely honest, we don't know. This is something that we as a company pay a lot of attention to. I want to talk a little bit about what's happening under the hood on the technology side. So you mentioned that the model itself, one of the things that makes it different from competitive programming is you've made it more, be good at the things that a professional software developer would do. Is that the biggest difference on the model side? Or should we think that it as a close cousin of 03? Yeah, so it's definitely the same models, 03, with additional reinforcement fine tuning.
16:47But that said, yeah, I think part of it is kind of like these more qualitative aspects of what makes a good software engineer versus simply like a good, let's say like, coder, you know, like style, even like how it writes comments. That's, I think, that's like one of the things that people have noticed with other models. And then on top of that, I also want to highlight one of the big challenges was like making good environments for the agent to kind of learn in. And so if you think about like real world software repositories, it's like so varied and complicated. I think about how much dev ops has to go into setting up for repository and that's something we're kind of like learning the hard way with the our environment setups.
17:29But should we talk about the multi repo I was showing you yesterday? Yeah, I was showing Hanson the repo for the startup that you know, open AI acquired and so we joined and so we were looking at that repo together, thinking about it for you as an environment and Hanson's like so So, where are the unit tests? You know, because the agent uses unit tests to verify it. I know, this is a real startup that has no unit. I mean, I was the same. So, I can't complain. So, yeah, you have all these really messy environments. So, yeah, we have to, over the course of training, we have to, basically, generate these really realistic environments for the agent to learn from.
18:10And I think one of the reasons that we're able to make such an end -to -end product work is that we have the same environment that we used during training and the same, like, basically this containerization infrastructure that we're using to serve in production. So our users are, like, we're running our own computer environments. When users use codecs, they're running in the exact same environments that we're using to training. So you don't have the agent saying, but it works on my machine. Exactly. Yeah. Okay. Okay. I think these are also the longest running agents I've seen out of OpenAI, deep research, maybe it was the previous one that was longest running.
18:49And my understanding is, you know, codecs can, you know, something's been 30 minutes on different tasks. Are there any kind of surprising challenges and things you've encountered just getting inference time to scale up on, you know, query for so long? Maybe I'll start with the product side and then there's many on the moment side. But on the product side, actually, the thing that I think the most about is user intent. It's like actually, if you imagine someone using auto -complete in their IDE, it's not super hard necessarily. I mean, obviously it's difficult, but it's not super hard to predict what are they trying to do right now for the next microsecond.
19:26But for doing a task that takes 30 minutes, it's actually fairly difficult to help a user describe the task. They may not even know exactly what they want for 30 minutes worth of work. So something that we spent a while debating, and it's still a thing we debate, is what is the right granularity of a task for someone to give to Codex? And how can we make it easy so that Codex can be really flexible where you can use it for one line changes? You can use it for big refactors that you know exactly what you want, or larger features where you know what you want? Or maybe can you use Codex when you don't know exactly what you want?
20:00And so maybe you should ask codex for a plan and then you can have it codex to just tasks and then do those tasks afterwards. So that's still a topic of the bait and iteration for us. Yeah, I think that's actually a good pro tip for using it. It's actually really good at coming up with its own plans and then sometimes it's really tedious to specify everything you up. You want upfront, and that's kind of like one of the unique challenges about working. If you wanted to work for an hour at a time that you kind of do have to specify a lot upfront, which means that you have to spend like, I don't know, like 10, 20 minutes coming up with that.
20:33But if you use actually like the ask mode to first like, you know, generate like a high level plan of what you want to do, and then you can like iterate on that with the model before you, you know, send it off for an hour. It really is like working with an intern. Yeah. What about on the model side? Anything that's surprising in terms of model behavior as it starts to run for so long? Yeah, I think our models have gone a lot better at kind of like sticking kind of like on task as it, especially with these longer rollouts, I will say there are cases where, you know, even though there is a limit to the model's patience, even though it's quite high.
21:07So it can be frustrating sometimes, you know, it's like, it goes off for like 30 minutes and then, you know, this is a case that we're working to get better at where it's like, you know, it's kind of like just like a human it comes back to you, it's like, sorry, I don't, this is too much. I don't have enough time to do this, actually, like that's one of the things it says. It's just like a term. So they're very human -like, yeah. In any case. Yeah. I'm curious how you think about the right interaction patterns and how they evolve and how the suite of products around this evolve over time. We have codex.
21:35We have codex CLI. What else do you think is out there in the design space for engineering and building products? Yeah. So the codex, as we launched it, is really just like, you know, it's a research preview. It's a thought experiment. A useful one, but it's still very early. And what we're most proud of with Codex is the model and the beginning of this foundation for compute environments. And the UI we shipped is one that we iterated towards and there's some fun stories there. But it's definitely not the final form factor. And for those listening, basically the UI we shipped is an interface in chatch .pt where you can submit a task and ask Codex to either answer your question or write code and then you kind of have this like, I mean, that looks a little bit like a 2D list of things that you can go look at merging.
22:24Really, I think for, so we built that to really lean hard into this idea of an asynchronous agent that you delegate to. But what we want to build towards is a setup where you don't have to think about whether you're delegating or whether you're pairing with an agent. And it should just feel like working with a teammate and where that teammate is like ubiquitously present and all the tools you work with. So you should be able to pull up any tool that you're working in, be it your terminal, your IDE, your issue management tool, maybe your learning tool, your errors, you know, it shows you errors, and just ask for help.
23:00Maybe even, Codex has already taken a look before you even got there, and it has like an opinion there. And you can be able to ask something, be it a short question or a long question, it'll just like appropriately decide how much time to spend before answering you, and just like help you land those changes. So basically, we want to kind of blend this idea of like pairing and delegation, but the first thing we shaped was just like the purest thought experiment. The other thing I'll add to this is like one of the unique things about working at OpenAI's that we are the makers of Chatch Bt, which is sort of the, you know, the AI's the most people use.
Read the full transcript
23:35And so we don't actually see a future where as you go about your day, you're deciding whether to use the Codex agent or I don't know, you're like shopping agent or taxi ordering agent. By the way, I'm just naming random things here. Or you're like marketing agent. Actually the way we think this should work is you should just have one assistant that you talk to and you can ask it anything about anything and it can just do the things you need. That's Chatch .pt that will become our assistant. And then if you're a power user of a certain type of tool, so let's say you're a software would develop or you spend a lot of time in certain functional tools, then you can go into that tool and have a spoke interface with buttons, with lists that you can use to efficiently go about your day.
24:18Do you think we'll still use IDs? Yeah, for sure, but they'll evolve. Right now, they're very focused on writing code and as Hansen was saying, probably agents will be writing more and more code and so it's gonna become, there'll be a shift in emphasis towards landing code or reviewing code or validating them or maybe even a shift in emphasis towards planning bigger arcs. Yeah, I think we're already seeing a lot of people on the team. They kind of like, first thing in the morning, they come in, they make coffee and then they kick off a few tasks just to kind of get a starting point. And then they come back after their breakfast and they look at the tasks where the PR has got generated, then they'll take those and the ID is kind of like the place where you take you know, it's not, it's maybe we'll get you like 80 % of the way there hopefully or even more.
25:03But then there's always this last mile where you go in and really fine tune Based on kind of like your own vibes How do you see the broader market evolving like within opening? I you have so many different strategies here and As you think about async task as you think about some of the things that you mentioned moving into chat gpt We're seeing a lot an explosion of other tools and specialized models Are you obviously are biased when I'm curious what your readers of the broader market? Yeah, it's a crazy time to be a developer right now like there are just so many new tools that are just so helpful.
25:38Like a fun story, recently I was in the airplane, and there was no Wi -Fi, and I had thought that I was gonna maybe write some code and like build a thing, and there was no Wi -Fi, and I was like, you know what's good? Like it's just not worth my time to like even try to write code anymore. Whereas, you know, the startup that I was working on like many years ago, like part of the genesis of that startup was like me writing some code without Wi -Fi in an airplane. And I just wouldn't even do that anymore, because like the market, it's just like, it's just changed so much. And I think we're going to see an equivalent shift in an equivalent amount of time.
26:06So in the next few years, coding will look completely different. I think right now, most of the tools that people find the most value from are tools that work really closely with you. In your development environment, like, baske pairing. And I think the shift that we're going to see, but we have to figure out how this will happen. But the shift that we're going to see is that actually the majority of code will be written by agents. And those agents won't be working in your environment where you can do one thing at a time, but they'll be working in their own environments. And they won't just be triggered by you, like thinking of specific tasks, but they'll be connected into the tools you use doing work there.
26:41And so I think, well, she basically that shift towards agents. I think we're going to have to figure out a lot about code review as you were asking you about. Personally, I don't exactly know how that's going to work, but I do know that even already at OpenAI, we're seeing much more code is merged by agents, but actually also even more code is generated by agents as folks are like, you know, like, say, kicking off tasks four times to like choose their favorite implementation. And so it's like not 100 % clear how we should even like manage all this code that is being written. Some things that I will say though, in case it's useful to the audience, is that there are definitely things you can do to your code base to make it more addressable for agents.
27:21This isn't necessarily a particularly novel, but, you know, obviously using like typed languages is really helpful. Another thing that's very helpful is having smaller modules that are like better tested. Like we joke about my test at all, yeah. Yeah, I'm a test. Like we joke about my startups. Refollow, but I bet you we would have written it differently if we were writing it today. And even there's small things like the code name for this project is WAN. This is the code name for codecs. It's like WHAM. And when we named it, we were very intentional in doing so because we knew we would have code in the server, for the website, in various other places.
27:56And we wanted it to be really easy for the agent to search for what WAM related code and find it. And so we named the project WAM and we grepped the code base first to figure out how it was there. Like if we would have called it something like code or codex or agent, you can imagine it would have been really hard for the agent to use codex. And now the agent's going to be confused. Well, so in the code, this is kind of my point, right? Like intentional design. Like in the code, we use the term WAM, like a lot. because that's actually much easier for the agent to find. Obviously, if we didn't use a word like that, the agent could still find its way, but it would have to spend much more time to find the right files.
28:34It is cool that a lot of the things that actually make the code base easier for humans to also tends to make these agents like good tests, for example. Writing good docs is another good example, where now I think there's even more of an incentive to do that, because not only does it make your life easier, it makes the agents life easier. Okay, sorry to be the annoying VC, But cloud coding tools are also like, I think, agent coding experiences from others. Do you think the, I'm curious how you think your experience is compared today? And then do you think the market is probably going to converge towards the same vision of what, you know, syncing and async coding look like, and in that version of future, what do you think opening eye wins on?
29:11I think we're going to see a little bit of everything, right? Like even in what you mentioned, like there's like tools that are working on your computer, there's tools that are working on their own computer. Like as I mentioned, like I think we're going to see the majority of work being written. where the agent has its own computer, but it will still be really important for us to invest in accelerating developers who are doing work on their own computer too. So ideally, we get the best of both worlds there, but most work is done in agent compute. I think the way I see it as well is like, I think one of the hardest part of software engineering really is like taking all the context from the world and like encoding it in these requirements, these like design docs.
29:46And then the implementation, like I think as we alluded to earlier, is like not actually like that much of the life cycles is spent on that physical coding. And so I think where chat GPT shines is like, it is this assistant that has memories now. It has access to a lot of different connectors to all the different tools you use. We have operator, deep research that have all these different capabilities. And so I think the vision where that all comes together is where a tool like Codex can really shine once it has access to all that knowledge. It's able to make use of that. And I think with that, it should be able to do a much more effective job at just the coding part?
30:23Yeah, imagine hiring a software engineer. And the only thing that that software engineer can do is take a task from you and produce a PR. Or it has these very well -defined features, and it can exactly do those things. And then you ask for a random thing, oh, hey, the team is getting together. Do you mind also, I don't know, getting a meeting room and leading a brainstorming? It just be so frustrating if you hire a teammate, and they refuse to do that kind of work. And so similarly, I think it's really, we're building towards a future where agents that you're working with are a little bit more generalized.
30:57To reference, Hans was talking about operator and deep research. If you think operator has a web browser, deep research has a different flavor over web browser, codex has a terminal. Really, your teammate has pretty similar tools, like a human teammate. And so the goal for us eventually is to pick places where we want to really invest in a specific audience to make rapid progress. So obviously we're doing that with coding, with codex, or like GPD 4 .1, where we generate specific e -vails for that audience and then made a better model for them for developers. But then over time, generalize these things into simple things that everyone can use.
31:33So I think, again, with OpenAI and ChatGit. I feel like that's a place where the products we've able to look very different from something that's very only specifically for coding. What do you think of the primary UI that developers use to interact with Codex, like do you think it'll be a chance at PT, the CLI, the ID, all above? Yeah, I think a mix of all the above, I think we just kind of want to meet developers where they are in that moment. So it might not even be like in the editor or in the terminal, it might be like on Slack. Like, you know, someone messages you like, hey, like there's a bug and you're just like, hey, like, go fix it.
32:09I'll give you my like fun, future UI that is like not at all serious. But maybe the future of working with agents, if you're a startup founder in the future, and you have a team of you and a couple of founders in many agents, actually looks like TikTok. Maybe you have vertical feed, and it's basically an agent has produced video that you can watch with an idea like, hey, a customer wrote in with this request, I think we should fix it. And then you swipe right to say like yeah, let's let's fix this. Let's do this You swipe left to say no, we did it. I didn't say this was gonna make a lot of sense And then you press and hold to provide feedback So you feel like yes, like do it, but you know make sure the font is in italic And so basically you have all these agents who are like subscribe to information at your company or on your team and They're proactively coming up with ideas and doing them and then giving you updates and you're kind of just curating the work that is being done.
33:09And they show you the previews of what the world could look like. Yeah, obviously that's a half joke though. I think that'll be the arms -like working with agents. And then there's definitely going to be really important for people to be able to go do the work themselves and on pair with agents in. I get that it's a half joke, but it's a really cool visual, because I think everyone agrees conceptually with this idea of collaborating and reviewing all the different changes that an agent makes. is gonna look very different from how we code today, but like nobody's actually given me a visual of what that might look like, so that's a really cool idea.
33:39I love it. Awesome. Should we wrap with the lighting round? Let's do it. Okay, recommended piece of content are reading for AI fans. For me, that's like immediate. That's like the culture by Ian Banks. You read it. Yeah, it's amazing. Yeah. It is a science fiction series, started being written in the 80s, and it is unusually positive in its view of how a future space -faring human and non -human race could look. And there's a lot of questioning about what is the purpose and meaning of life when we have AGI. Yeah, I think for me it's anything by Richard Sutton. I think that was my introduction to reinforcement learning.
34:19And I think it's kind of a joke here that we read the bitter lesson every single day. That's kind of the philosophy of OpenAI. I think even with Codex, we give it a terminal and it literally uses POSIX tools. That's the most bitter lesson way of working with the computer. And your favorite AI apps? Gotta be chat DVD. Not chat if you see it coming in. They were so boring. They were so boring. I don't mean. OK, either it could be a new feature that you guys have released other than Konex or something outside of OpenAI. OK, so I guess I don't, it's funny. I don't really think of AI apps. But I do like it when my life gets easier.
34:59So, you know, some things that I like are like when you're using AI, but it's kind of invisible. So, like, just I'm in product. So, I often like file bugs, and like linear has a really elegant integration when you file a bug from a Slack conversation. It just generates the bug from the Slack conversation. But they never say AI anywhere. Just like, you actually kind of don't even notice that it's using AI. Oh wait, I came up with an answer for favorite AI app, Waymo. Oh, there we go. Yeah, I think for me, like, co -pilot I've definitely been the thing that keeps delivering value every single day for me.
35:36OK, robotics, bullish, bearish. Bullish? Yeah. Which new application or application category do you think will break out in 2025? Other than coding? Yeah, I mean, I think when you had Yusa and Josh on, It's kind of the same answer, but 2025 is definitely the year of agents. I think we're going to see agents take off in a lot of different categories. Yeah, I have to agree with that. What's every agent so you're most excited about? Aside from coding agents. That's a good question. Well, I mean, so my take would be like, you know, if we, I know this meant to be rapid fire, but like kind of the way we think of agents is you have reasoning models, right?
36:13And then you give those reasoning models like access to tools of the trade. and then you figure out how to train that agent to do the specific function. So it's not just about writing, it's about journalism, it's not just about coding, it's about software engineering. So that's kind of what we're doing. And in my mind, the reason I'm so excited about agents this year is because we now have a few agents' ships from OpenAI and other companies are shipping agents too. And so we're starting to see what kind of the shape this is and starting to identify the primitives. And so specifically what I'm excited about is like as we bring this together and you come up with like an agent that you don't have to provision like separately for every single function, but it's an agent with a computer that has a browser and has a terminal and it can do like multiple things without you having to like exactly specify like you are my coding agent or something.
36:58Really cool. Thank you so much for joining us. Congratulations on what you've felt at Codex and thank you for giving us a preview of how you think the coding market will evolve and also giving us a peek into how you know long running async agents and experiences will play out. Really appreciate it. Thank you. Thanks for having us. Thank you.
From the publisher
Hanson Wang and Alexander Embiricos from OpenAI's Codex team discuss their latest AI coding agent that works independently in its own environment for up to 30 minutes, generating full pull requests from simple task descriptions. They explain how they trained the model beyond competitive programming to match real-world software engineering needs, the shift from pairing with AI to delegating to autonomous agents, and their vision for a future where the majority of code is written by agents working on their own computers. The conversation covers the technical challenges of long-running inference, the importance of creating realistic training environments, and how developers are already using Codex to fix bugs and implement features at OpenAI.
Hosted by Sonya Huang and Lauren Reeder, Sequoia Capital
Mentioned in this episode:
The Culture: Sci-Fi series by Iain Banks portraying an optimistic view of AI
The Bitter Lesson: Influential paper by Rich Sutton on the importance of scale as a strategic unlock for AI.




