ChatGPT Codex: The Missing Manual

16 May 2025

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

```markdown

Latent Space

The AI Engineer Podcast - Episode Summary

Episode Title

ChatGPT Codex: The Missing Manual

Episode Description In this episode, the podcast delves into ChatGPT Codex, the first cloud-hosted Autonomous Software Engineer (A-SWE) developed by OpenAI. Hosts Alessio and Wix talk with core developers Josh Ma and Alexander Embiricos about Codex's origin story and its future roadmap.

Key Concepts Discussed

  • Introduction to ChatGPT Codex
  • Launch of ChatGPT Codex and its implications for AI-driven software engineering.
  • Personal Journeys into AI Development
  • Backgrounds of Josh Ma and Alexander Embiricos in AI and software development.
  • Evolution of Codex and AI Agents
  • Transition from traditional coding tools to AI-assisted development.
  • The significance and potential of AI agents in software engineering.
  • Understanding the Form Factor of Codex
  • Discussion on the interface and interaction patterns users will experience with Codex.
  • Building a Software Engineering Agent
  • Technical overview of what constitutes an "agent" in software development.
  • Best Practices for Using AI Agents
  • Strategies developers can employ to maximize productivity when using AI tools.
  • Emphasis on code structure and modularity to facilitate better AI interactions.
  • Navigating Human and AI Collaboration
  • How humans and AI can work together effectively in software development.
  • The importance of user feedback in refining AI capabilities.
  • Future of AI in Software Development
  • Predictions on how AI will further integrate into everyday coding tasks.
  • Planning and Decision-Making in AI Development
  • The role of strategic planning in developing AI tools like Codex.
  • User, Developer, and Model Dynamics
  • Examination of how users interact with AI models and the developers' role in shaping these interactions.
  • Iterative Deployment and Future Improvements
  • The importance of refining AI tools through iterative user feedback and updates.

Key Takeaways

  • Collaboration with AI: The need for a paradigm shift in how developers interact with coding tools, moving from traditional methods to AI-assisted processes.
  • User-Centric Design: The significance of incorporating user feedback into the development of AI tools to better meet developer needs.
  • Best Practices:
  • Install linters and formatters to enhance AI performance.
  • Utilize commit hooks to create a more structured development environment.
  • Make codebases modular and well-documented to facilitate AI understanding and interaction.
  • Future Roadmap: The development of Codex will continue to evolve, focusing on multimodal inputs, enhanced user interfaces, and greater integration with existing development tools.

Conclusion The episode concludes with an invitation for feedback from the developer community, emphasizing OpenAI’s commitment to iteratively improve Codex based on user experiences and needs. Both guests express excitement about the future of AI in coding and its potential to revolutionize software engineering.

Links

  • Follow Josh Ma on [GitHub](https://github.com/joshma)
  • Follow Alexander Embiricos on [x.com](https://x.com/embirico)
  • Explore more about Codex and AI development at [Latent Space](https://latent.space)

```

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:06Hey, everyone. Welcome to the Latent Space podcast. This is Alessio partner and CTO at Decibel, and I'm joined by my co-host, Wix, founder of Small.ai. Hello, hello. Calling in from Singapore here, but we are in the remote studio because the OpenAI team keeps shipping, and today they just live-streamed and released ChatGPT Codex. Welcome to Josh, who I think we've talked about, we've met while you were at Airplane, right? Yeah, yeah. I've been building DevTools for a bit now, and I have to talk to you when I'm building DevTools. I mean, you have now seen me complain a lot when things happen. So I don't know if it's a good or bad thing.

0:51It's a gift, man. Feedback is a gift. Thank you. Alexander, we're new to each other, but you've been leading a lot of the codex testing and demos and stuff. Yeah. Hey, I'm Alexander. I'm on the product team here. Awesome. So, yeah, we're going to just assume that everyone's watched the live stream. He also released a blog post with a bunch of like test demo videos. Basically a bunch of like, it's very, it's very interesting. I noticed in the demo videos, it was like individual engineers sitting by themselves, very lonely. And then they're just talking to their AI friends, coding with them. I don't know if that's the vibe you want to give off, but like that's how I came across.

1:31Yeah, man, we those, those videos, we were going for like maximum authentic, just like engineers about how it helps them. Yeah, I'll take the feedback. But no, I mean, it's true. I mean, sometimes, you know, on-call is a lonely job. Like, mobile engineer is a lonely job. Like, there's not that many of those. So, yeah, totally. But anyway, so what did you guys individually do? Maybe we can kind of start there. How did you get pulled into the project? And we'll start from there. Yeah, maybe I can go first because then we have a fun story about how we started working together. okay so actually before working at OpenAI I was working on a native macOS software called Multi which is like about like it was kind of like a pair programming tool but we thought of ourselves as working on like like human to human collaboration and then basically as Chatapiti and stuff came around we started thinking about like oh what if instead of a human pair programming with a human it was like a human pair programming with an AI so I'll skip this whole journey but that was this whole journey and then we all ended up joining OpenAI.

2:33And I was mostly working on desktop software. And then we shipped reasoning models. And, you know, I'm sure you guys were, like, head of the curve in terms of understanding the value of reasoning models. But for me, like, it's kind of, like, starts off as better chat, but then when you can give it tools, you can actually, like, make it an agent, right? Like, an agent is, like, a reasoning model with, like, tools and environment, guardrails, and then maybe, like, training on, like, specific tasks. So anyways, we got super interested in that, and we were just, like, starting to think about, like, okay, how do we bring reasoning models into desktop.

3:06And at the same time here at OpenAI, there was a lot of experiments going on with giving these reasoning models access to terminals. I wasn't working on those first experiments, to be clear. But that was the first true, wow, I really feel the AGI moment that I had. It was actually while I was talking to David Kay, a designer who was working on this thing called Scientist. And he showed me this demo of it updating itself. And like nowadays, like, I don't know if any one of us would be like the most impressed to like change the background color. Modifying its own code? Yeah. And then, you know, it was like they had a hot reloading setup.

3:41So I was just like mind blown at the time. And I know it's still a super cool demo. And so we kind of were experimenting with a bunch of these. And I sort of joined one of the teams that was like tinkering with this. and we kind of realized hey, it's just super valuable to figure out how to give a reasoning model access to a terminal and then now we have to figure out how to make that a useful product and how to make it safe you can't just let it go loose on your local system, but that's where people were initially trying to use it. So a lot of those learnings ended up becoming the Codex CLI which shipped recently a lot of the work there the thinking that I'm most proud of is enabling things like full auto mode.

4:26When you do that, we actually increase the amount of sandboxing so that's still safe for you. We were working on these types of things and then we started realizing we want to let the model think for longer. We want to have a bigger model. We want to let the model do more things safely without having to do any approvals. And so we thought maybe we should give the model its own computer. the agent in its own computer. And then at the same time, we were also experimenting with putting the CLI in our CI so it could automatically fix tests. We did this crazy hack to get it to automatically fix linear tickets in our issue tracker.

5:05And so then we ended up creating this project that is Codex, which is basically really the concept of giving the agent access to a computer. Actually, I realize that. I don't know if you were asking. Well, I personally did, but anyways, I told the story. I hope that's okay. Sure. No, I mean, you weave your personal story into the larger narrative anyway, but yeah, and I'm sure Josh has a part two. Yeah, yeah, so my story's somewhat different. I've had OpenAI for two months here, and it's been one of the most fun, chaotic two months of my life, but maybe I'll start back at the company I had founded a few years back called Airplane.

5:45We were building an internal tool platform. The idea is to let you build internal tools, but really lean into developers. and make that really easy. And it sounds unrelated, but in many ways, the similar themes started coming up. What's the right form factor for doing local development? How do you deploy tooling to the cloud? How do you run code in the cloud? How do you compose all these primitives of storage and compute and UI to let developers build software really quickly? I like to joke that we were just, I don't know, two years too early. Towards the end, we were playing around with GPT 3.5 and trying to really make, it was really cool.

6:24It could actually build a React view really quickly, right? And I think if we had kept going on it, maybe it would have turned into some of the AI builders that you see today. But that company ended up getting acquired by Airtable, where I ran some of the AI engineering teams there. And for me personally, towards the beginning of this year, I saw the progress we were making in software, agentic software development. And for me, it was a bit like my own moon landing kind of moment that I suspected was about to happen, right? Whether or not I was involved in the next two years, I think we are going to build an agentic software engineer.

7:05And so I talked to my friend over here. I was like, hey, are you guys working on something like this? And, you know, you can see a wide-eyed look. He's like, I'm not allowed to tell you anything, but maybe you can talk to the team. and so very fortunately this was right when Alex and folks were spinning up things and I remember actually in our interview we riffed on the form factor should it be CLI the issues with that, waiting for it to finish and not being able to interrupt all the time wanting to run it 4 times, 10 times in parallel and at that point I said maybe it should be both and we sort of are going for that right now but yeah I'll just say I was very excited And so I'm very excited to just be pushing this forward.

7:46And I think Codex is still really early, excited to share it with the world, but there's a lot more to build. Yeah. I'll say it was a very fun conversation when we first met because he came in. I've never had this happen before. It's like, here's exactly like kind of the change that I see in the world and therefore the type of product that I want to build. I know you can't confirm if you're working on it, but just so you know, this is the only thing I want to work on. And then I was like, I asked just a few open-ended questions and we immediately got into some of the core debates around the form factor of the tool.

8:17And I was like, okay, this is awesome. I think a DevTools person can spot another DevTools person like that. Yeah, blink twice if you're working on this. But for what it's worth, early iPhone team at Apple was the same because iPhone team members did not know if they were on the same team. They're not allowed to tell each other. So they had to triangulate. Wow. It's like a two-piece. And talking about form factor, so you mentioned the CLI, which you already released, and I think there's other, you know, Cloud Code, AIDR, a bunch of other tools out there. Should people think of Codex in ChagipD as like a hosted Codex CLI?

8:54Like, are there big differences between the two? Let's talk about that. Yeah, go for it. Yeah, I think of it as, I think that's the short of it, right? Allowing you to run codex agents in OpenAI's cloud. But I think that the form factor, like, it's a lot more than just where the computer runs, right? It's how does this bind to the UI? How does this scale out over time? How do you manage caching and permissioning? And how do you do the collaboration story? and so let me know if you disagree but I think the form factor is the core of it. It's been honestly a really fun journey. The other day or maybe last night in the AM, Josh was sleeping because he had to do the live stream.

9:46I didn't have to. A bunch of us were looking back at the dock where we planned what we were going to ship and we were like man, we had a lot of scope creep. and effectively all that scope creep was kind of like incrementally made sense because we kept leaning further and further into this idea that this is like not just like a model that's good at coding but rather this is an agent that is good at like independent software engineering work and like the more we lent into that the more like things started to feel really special so you know uh we could i'm gonna just label and then set aside like the entire conversation around like the compute platform that Josh has like, you know, been leading.

10:27But let's just take the model, for example. You know, we don't just want it to be good at code and like we don't just want it to like solve like say like Sweebench tasks. You know, Sweebench is an eval for those who don't know that's like has a certain way of like functionally grading outputs. Because if you look at a lot of like Sweebench passing like outputs from like an agent, they're not really like PRs that you would merge because like the code style might be like different. Like it works, but the code style is different. So, you know, we spent a lot of time, like, making sure that our model is, like, graded at adhering to instructions, graded at inferring code styles so that you don't have to tell it.

11:02But let's say that you got, then, a PR. That it was, like, the code style was good. It followed your instructions well. It still might be really hard to merge if you have this, like, enormous description, like, just, like, model of, like, how it thought about building it. And, you know, you probably need to pull it onto your computer to, like, test the change and validate that it works. and maybe that's okay if you're just running one change but in a future world that we imagine where actually maybe the majority of code is actually being written by agents that we're delegating to, doing tasks in parallel, it becomes critically important that you can actually integrate those changes easily as the human developer.

11:38For instance, some of the other stuff we started to train was PR descriptions. Let's really nail this idea of a good, concise PR description that highlights the relevant things. So our model will actually write like a nice short PR description with like a PR title that like adheres to your like, you know, repo format. We have a way to prompt that more if you want with agents.md. And then in the PR description, it'll actually cite like relevant code that it found along the way or relevant code in its PR. So you can like mouse over and just see it. You know, and perhaps my favorite thing is actually the way we handle testing.

12:11So the model will attempt to test its change. and then it will tell you in this like really nice way with just like a checkbox kind of thing whether or not those tests passed and again it will cite if the test passed like a deterministic reference to the log so you can like read it and be like okay I know that this test passed right or if the test failed it'll be like hey this didn't work I feel like you need to install like PMPM or whatever and you can like read the log and like see what it is so those are some of the things that I think I've lost track of the original question but anyways Those are some of the things that we've been really leaning into as we build this software engineer agent in the cloud.

12:47I think also just it feels very different. You can look at the features, but I think for me the feeling is it takes a leap of faith the first few times. You're like, I'm not really sure if this is going to work. And it goes off for like 30 minutes. But then it comes back and it's like, wow, this agent went out, wrote a bunch of code, wrote scripts to help code mod its own changes, tested this, and it really went through the full end-to-end thinking about the change it wants to make and I you know had no faith that at the start that it was gonna be able to successfully do it and after using a bit you're like wow like it actually you know pull through and so that kind of long-running independence is something that's like hard to really like summarize you have to really try it but it finally feels very different and yeah that feels special it yeah i used it i opened up br for it a few minutes ago i was in the lucky first 25 of people to get the roll up uh yeah it's very nice you know it kind of shortcut it because it couldn't figure out how to run rspec in rails and so i just checked the syntax of the ruby file and it was like looks good to me but i think it doesn't have the agents that md yet so i think once i set that up they'll be good um no just it just don't use ruby man like python oh yeah That is skill issues.

14:06Once it's good enough to migrate the whole thing, that'll do that. I mean, it is funny that there is, just briefly on the note of like, don't use Ruby or not, there's like a bunch of things that I think teams can do to like make better use of AI agents. Oh, please. Stop using Ruby. Number two. But yeah, if you could list some things out, that's like best practices. You know, I noted from the live stream that they mentioned pro users install linters and formatters so that basically these are in the loop verifiers that the agent can kind of use, right? So, which, like, you know, it turns out to be dev best practices as well, but now the agents can also auto-use it.

14:41Commit hooks have always been a tricky thing for humans because I've been on teams that were like, no, everyone, everything has to have a commit hook. And then I've also been on teams that were like, no, like, this thing gets in the way of committing, so let's rip everything out. But actually for agents, it's actually really good to have commit hooks. Yeah, I mean, you took the words out of my mouth. I think the three I was going to say would be one, agents.md. We put a lot of effort into making sure the agent understand this hierarchy of instructions. You can put them in subdirectories and it'll understand which ones take precedence over which others.

15:16So over time, we also have O3 and Fora writing our agents.md files for us. I love the tips. You actually open sourced the prompt descriptions here. Yeah. Anything to highlight? I mean, I think I would start simple and not try to overdo it. And a simple agents on IMD will get you a long way rather than no agents envy. And then it's more of like you learn over time, right? What we would really like to do is auto-generate this at some point for you based on the PRs you create and the feedback you give. But we figured we ship faster rather than later. And you mentioned you have 03, 04 writing HSMD for you as well?

16:00Yeah, like, I'll, like, give it my entire directory, right, and just say, like, hey, produce an HSMD. Well, actually, these days I'm using code 1 to do it because it can traverse, codex 1, sorry, to traverse, like, your directory tree and generate those things for you. So, yeah, I would recommend slowly, gradually investing in HSMD. And then, you know, you took the words out of my mouth, like, getting very basic linting, formatting up that's your really big wins because it's similar to how if you open a new project in VS Code, right? You get some out-of-the-box checking. The agent's starting. As a human, you're starting without that advantage.

16:40And so this is trying to give that back to the... And yeah, I don't know. Do you have anything else? Yeah, so one analogy there. And then actually I have just some thoughts we've observed of even using other coding agents, just any coding agent, how to prepare for that. But the analogy I kind of like is if you start with a base reasoning model, actually, you basically have this really precocious, incredibly intelligent, incredibly knowledgeable and weirdly spikily intelligent college grad. But we all know if you hire that person and ask them to do software engineering work independently, there's just a lot of practices that they're not going to know about.

17:22and so kind of a lot of what we've done with with Codex 1 is basically give it its first few years of job experience and like that's effectively what the training is so that it just kind of knows more of these things and like if you think about it like a PR description is a classic example of that like writing a good PR description right and possibly knowing what not to put in it actually right and then so that's what you get there so now you have this like weirdly knowledgeable eligible, spikily intelligent college grad with a few years of job experience. And then every time you kick off a task, it's kind of like their first day at your company.

17:57Right. And so agents.md is basically a method for you to kind of compress that like test time exploration that it has to do so it can like know more. And as Josh said, obviously we want to like right now it's like research previews. So you have to like update it yourself. But there's like a lot of ideas we have for how to make that automatic. So that's just a fun analogy. Yeah. Maybe the last one I'll say is make your code base discoverable. It's the equivalent of maintaining good engineering practices for new hires that you make, letting them understand your code base faster. A lot of my prompts start with, I'm working in this subdirectory.

18:32Here's what I'd like to accomplish. Can you please do it for me? And so giving that guidance at scoping helps. Yeah. Okay, I'll give you three things for generally. First, language choice. I was hanging out with a friend the other day who's a bit of a latecomer to AI. And he was like, oh yeah, I want to try building like an agent's product. Like, should I build it in JavaScript? And I was like, you're still using JavaScript? Like, no wonder. Like, you just like, you know, you use at least TypeScript, like give it some types. So, I mean, I think that's a basic one. I don't think anyone listening to us now needs to be told this.

19:07Another one is like, you just make your code modular, right? The more modular and testable it is, the better. But you don't even have to write the test. Like an agent can write the test, but you kind of need to design the architecture to be modular. I saw this presentation recently by someone here who was like, they weren't vibe coding, it was a professional software engineer, but using tools like Codex to build a new system. And they got to build a system from scratch and there was kind of this graph of their commit velocity. And then their system had some traction, so then it was like, okay, now we're going to port it into the monolith that is the overall Chatshapd code base that has seen ridiculous hyper growth.

19:42and so maybe it's not like the most architecturally pre-planned and like their commit rate, the same engineer, same tooling, actually the AI tooling continues to improve, their commit rate just like plummets, right? And so I think the other thing is just like, yeah, architecture, like good architecture is like even more important than ever and like I guess the fun thing is like for now, that's a thing that humans are really good at. So like, you know, kind of good, you know, important for the software engineers to do their job. I don't know, just don't look at my code base. Yeah, well, definitely not in mind.

20:12But the last thing is just kind of a fun story, which is the codename, the internal codename for our project is Wham, like W-H-A-M. And we chose it. Actually, I was working with our research lead, and he was like, hey, make sure you grep the code base before you choose the codename. So we, like, searched the code base, and the string Wham was, like, only present, like, in a few larger strings and never present as its own string. And that means that whenever we prompt, we can be very efficient. We can just say in WAM, right? And then WAM code that is like, you know, for our web code base or our server code base or in like our shared types or anywhere else is like really efficient for the agent to find, right?

20:55Whereas let's pretend like alternatively we would have called our product like ChatGPT code, you know, not saying we didn't consider that. then it would be super hard for the agent to figure out where we wanted to direct it to. And so we'd probably have to provide like more relative like folder paths. So there's a lot of this stuff like as you start to think ahead, like, oh, I'm going to have an agent. It's going to be using terminal to grep. Then, you know, you can start like naming things intentionally. Would you name would you start naming things less for humans readability and more for agent readability?

21:26Like what's kind of the trade off in your mind? Yeah, I it's interesting because I definitely had. different priors coming in and open AI. Like, I currently believe that the systems are actually very convergent. Like, there's a lot of, maybe it's because as long as you see humans and AI writing it, like maybe there's a world where it's only AI is maintaining a code base and the assumptions change. But the moment you start to break that fourth wall and a human is coming in, doing code review, deploying the code, like, it has human fingerprints all over it, right? And so how humans communicate to AI, where to make the change.

22:00how humans communicate a bug that needs to be done or communicate business requirements, right? All those things aren't going to go away immediately. And so I think the whole system still feels actually very human. I think there's a cooler answer I could say that it's like, oh, no, it's like this alien thing. It's completely different. But I don't know. I think it's like these are started off as large language models. There's a lot rooted in human communication. Yeah. By the way, if there's somewhere you want to take this, you should actually cut us off because I realize we're just kind of monologuing between each other.

22:29No, no, I think this also ties to the agents.md, right? It's like, why is it called agents.md and not readme.md? There's kind of like, I guess in your mind, some fundamental difference with how the agents and the human consumes the information. So I'm curious if you think that's at the class naming level. It's just at the instruction level. Like, where does it break down? Yeah. Go over. Okay, so this is like a few options for this naming, right, which we considered. So you could go for reading, readme.md. You could go for contributors.md, right? You could go for codexagent.md and then maybe codexcli.md is these two separate files, right?

23:06But that are sort of branded. There's also cursor rules, windsurf rules. Everyone has rules. Yeah, and then you could go for agents.md, right? And so there are a few trade-offs here. I guess one is openness and one is specificity, I suppose. and so when we thought about it you know we thought about well probably there's things that you want to tell an agent that you don't need to tell a contributor and similarly there's things you want to tell contributors to like really like help them set up in your repo or whatever that you don't need to tell the agent the agent can just figure that out so we were like okay maybe this is going to be different and like you know the agent's going to read your readme anyways so like maybe agents.md ends up being like the stuff that you need to tell the agent that it's not like like, automatically figuring it out from the readme.

23:52So we kind of made that decision. Then we considered, okay, there are, like, different form factors of agents, right? Like, the most special thing about what we are building and shipping is, like, it's just, like, an out-of-the-box way to use a cloud-based agent that can do many tasks in parallel and can, like, think for a long time and can use a lot of tools safely, right? And so we thought, like, well, you know, how fundamentally different is the set of instructions that you want to give that from an agent that you're working with more collaboratively on your computer? We had a good amount of debates about that, to be completely honest.

24:24And then we ended up concluding, like, actually, those sets of instructions aren't different enough that we need to namespace this file. If there is something you need to namespace, you could probably just, like, say it in plain language within the file. Then the last thing we consider is, like, well, okay, how different do we think the instructions you have to give, like, our agent are to the instructions you might give to an agent running on a different model or built by a different company? And we just think it kind of sucks if you have to, like, create, like, all these different agents or whatever.

24:49you know, it's like part of why we made the Codex CLI open source is like a lot of problems, like safety issues that you need to figure out for how to deploy these things safely and no one should have to figure these out like more than once. So that's why we went for like a non-branded name. And I have one specific example of why you know, readme and agent.imd are different. For agents, I don't think you really have to tell code style. It looks at your code base and just writes code that's consistent to that. Whereas like a human's not going to take its time, sorry, their time to go through the code base and, you know, follow all the conventions, right?

25:27So that's just one example, like, you know, at the end of the day, like, there are differences between how these two kinds of developers approach it.

Read the full transcript

25:38Cool. I think that's a really good set of advice. I think you just gave us our episode title. Like we're just going to call it best practices for using chat GPT codecs. And, you know, I mean, I think people are going to want best practices. So I noticed something that's very interesting, right? Like I think there's always two versions in terms of building agents. One, which is you try to be more controlling. You try to make it more deterministic. And then the other, you try to just prompt it and trust the model. And I think your approach is very much prompted, trust the model. So I see inside of the agents.md system prompt that you just prompt it to behave the way that you want, and you just expect the model to behave it.

26:19Obviously, you have control of the model, so you can train it if it doesn't do well. But one thing that makes me question it is, how do you fit everything in context? Like, what if I just have a super long HSMD? You know, in your live stream, you had it demoing on, like, the OpenAI monorepo, which is just giant, right? Like, so how do you manage caching and context windows and all that? Yeah. I mean, would you believe me if I told you right now that it all fits in the context window? Not the OpenAI repo. No, but sorry, everything that the agent needs. Right, so you reify the agents ID, put it at the top, right?

27:03It's just like another system prompt. No, actually, it's a file that the agent knows how to, like, graph and set for, right? Okay. Because there might be multiple ones. And so you can actually see it in the work log, right? It's, like, going to look for, it very aggressively looks for an agent ID. It's been trained to do that. I'll say it's been really interesting joining OpenAI and seeing how when you're thinking about where models are going and what AI products will look like years from now, you design products in a different way. Before OpenAI, especially when you don't have access to a team of researchers and many, many GPUs, you're building these deterministic programs, a lot of scaffolding around how this operates.

27:49But you don't really let the model operate as full as capacity, right? A lot of it was interesting when I just joined actually I got a lot of pushback saying like, hey, why don't we just like hard code, like, listen you keep using this tool wrong let's just say in our prompt, don't do that and then the researchers will be like, no, no, no we don't do that, we're going to do it the right way, we're going to teach the model why this is the right way to do it and I think that's like related to this overall thought on like, where do you put the deterministic guardrails in and where do you really let the model think, right?

28:23Similar conversation around planning. Should we just have an explicit planning stage where it's like think out loud first, write down what you're going to do, and then go do it? Sure, but what if the task is really easy? All right, do you really want to think this whole time? What if it needs to like replan as it goes? Like do you have all these like if-else conditions, heuristics to do that? Or do you train a really good model that knows how to switch between those modes of thinking? And so it's tough. Like I definitely have advocated for like little guardrails here and there until like the next training run's done.

28:52But I think that's really like we're really building for this future where the model is able to make all these decisions. What's really important is that you give it the right tool, right? You give it ways to manage context, manage memory, manage ways to explore the code base. Those still are really important. Yeah, that's like, yeah, super well said. Like I think building here is like super fun and different. And like, you know, the model isn't all the product, but the model is the product, right? And you kind of, like, need to have this kind of, like, humility in terms of, like, thinking about, like, okay, well, what are the things that, like, we, there's, like, three parties, right?

29:32There's the user, the developer, and, like, the model maybe, right? What are the things that the user, like, just needs to decide up front? And then what are the things that, like, we, the developer, are going to be able to decide better than the model? And then what are the things that the model can just decide best, right? And, like, you kind of, every decision is, like, just has to be one of those three. And, you know, it's not like everything is the model. Like, for instance, we have two buttons in the UI right now, like ask and code. And, like, you know, those probably could get inlined into, like, the decisions the model makes.

30:05But, you know, right now, it was just really, like, it made sense to kind of just give the user choice up front because we spawn a different container for the model first based on what button you press. So, like, if you ask for code, we put all the dependencies in. I'm going to oversimplify here. But if you don't ask for code, if you're just asking a question, we do a much quicker container setup before the model gets any choice. And so that's maybe a user decision. There's some places where user and developer decisions kind of come together around the environment. But ultimately, a lot of agents that I see are really impressive.

30:41But it's basically part of what's impressive is it's a bunch of developers building this really bespoke state machine around a bunch of short model calls. and so then the upper bound of like complexity of problem that the model can tackle is kind of actually just what can fit in the developer's brain. And over time we want these models to capture like or to solve for much more complex problems you know just by themselves on like more and more complex individual tasks and then eventually you could really imagine that you get like a team of agents working together maybe with like one agent that's kind of managing those agents and you know the complexity just explodes and so we really want to like get as much of that complexity as much of that state machine as possible like pushed into the model.

31:19And so you end up with these kind of two modes of building. Like in one place, you're like building product UI and rules. And in the other case, you still have to do work to get the model to learn something. But rather what you have to do is you have to figure out like what are the right things that this model needs to see during its training to like learn something. And so it's still a lot of human work to like figure out how to get that change. But it's like a very different way of thinking of like, we're going to get the model to see this. But how do you build the product to get the signal?

31:47So if you think about the code in Ask, it's almost you're basically getting the user to label the prompt in a way, right? Because they say, Ask, this is an Ask prompt, code, this is a code prompt. Are there any other kind of like fun product designs, like as you built this, of like, okay, we think the model can learn this, but we don't have the data. This is how we architect codecs to kind of help us collect that data? I think file context and scoping is, we don't have great built-in things like that right now. but it's like one of the obvious things that I need to add is another example of this, right?

32:19Like you could have, we're often usually pleasantly surprised that it's, oh, it was able to find the exact file that I was thinking about, but it takes some time, right? And so a lot of times you'll shortcut a bunch of chain of thought by just saying, hey, I'm looking at this directory. Can you go through? So I think that'll probably be there for a bit until you have some better architectural indexing and search capabilities. Yeah, I'll add to this. I'm actually going to double down on my thing about how do we think about it. So one thing we might consider is context window management, right? And should we intervene here?

32:56And so we could do a product intervention, right? Like write some code to intervene. And then kind of the next level of thinking, maybe like a little bit more AGI-pilled, is like, okay, let's get the model to see context window management stuff in this training. I can't even come up with an example now at this point because I'm too AGI-filled, but like, I don't know, we could come up with something that it has to see to like learn how to manage its context. But it's like specifically tasks related to context window. But then the most AGI-filled thing to do is like to be like, we don't actually need to think about this problem.

33:24The model will just figure it out. All we have to do is give it harder and harder problems. And then like it will just have an emergent property of managing its own context because that's the only way it can like solve these problems, right? So I'm kind of slightly like oversimplifying here. But like, you know, basically the model learns to manage its context. And so when you were talking about it working in the monorepo, it learns how to be efficient with the way that it spends its tokens as it's browsing and setting. In your example of there's a giant agent's.md, I guess we would just need to show it some versions where there was that.

33:57And so it learns it shouldn't read the whole thing every time and should first figure out how many lines it has, et cetera. So anyways, summarizing, we just need to keep giving it harder and harder problems. And a lot of these things that we might be very tempted to build a sub-intervention for, it will just have to figure out. And if it doesn't figure it out, maybe it didn't matter. Sure. Yeah, I totally get that. I think, like, we don't really have online models yet, right? And that's kind of what you need for your vision to be real. And for what it's worth, I wasn't thinking about, like, a giant AgentsMD.

34:30I was just thinking about, like, hierarchical nested AgentsMD with, like, a lot of code. um and like i think one issue where you have this version where the model is the product is your your dev cycle as the codex team like the two of you like you have to kind of it's not as tight because you have to be like okay every time there's a bug all right now i need to go get data um and where do you get the data i don't know like maybe employees use it maybe you have like you buy it from vendors and like you hire some human raters or whatever and then you have to train it in and then you have to go test it again it's it's very slow isn't it and like expensive yeah i think it's definitely yeah i think it's definitely like from a from a building perspective you have to do this when you're really willing to play the like the long-term vision of like we're going to build a better model maybe even a better model bespoke for a certain like functional purpose like codex one and then we're going to generalize the learnings from that model into like an even bigger model that's like getting all these other like learnings from other functional purposes and like these together will like become a really powerful thing and that's kind of like the philosophy we have with training models so far and it has been working but it's definitely like a long term play you know like another example we do do this on occasion like for example recently we released GPT 4.1 like really good coding model and again that was like based on like working we were like hey we want to invest better in this area let's hang out with a bunch of developers understand their feedback how things work, you know, create some evals.

36:05And, you know, like you said, this is like, it's a lot of work to do that. But then we end up with a great model and even more exciting, we can then like take those learnings and like put them into our mainline models and then everything benefits. And you kind of, the sort of philosophical view, I don't know if I can like factually prove it or not, maybe someone here can, is that like, if you can like build, do something very specific for like a specific purpose, actually when you bring that and you bring into the generalized model, you might even get outsized returns on that because there's transfer from all these different domains.

36:37Okay, cool. I think we had a couple factual things to wrap up on just the codex itself, and then we wanted to double-click on the compute platform stuff, which I think, Josh, you wanted to cover more on. So I noticed in the details, it was between 1 to 30 minutes in length. Is that a hard cutoff? Have you had it go for longer? Any comment on the task time? Yeah, I mean, I just checked the code base before this. Someone else had a similar question. Our hard cutoff is an hour right now, although don't hold us to that. It may change over time. The longest is I've seen two hours when in development mode and the model went off the rails.

37:17So, you know, but I think 30 minutes is a great ballpark for the kind of tasks that we're trying to solve, right? These are hard tasks that require a lot of iteration. and testing and the model needs that time. Yeah, I mean, yeah, I think actually like our average is like pretty, it's significantly lower than 30. But if you give it a hard task, you'll end up at 30. Yeah, I mean, you know, I think there's a couple analogies here. One, I think the operator team released a benchmark where they had to cut off for two hours. And then the other one is the meter paper, which I don't know if has been circulating, where they estimated that the current average autonomous time is like an hour and it's maybe doubling every seven months.

38:00So like an hour sounds right, but also, I mean, that's the median. So there's going to be some that go longer than that. Yeah, totally. Is this part of the, you had cutoffs for a few, like 23 Sweetbench verified examples that were not runnable. Was that part of it in terms of length or was it just something else? Yeah. To be honest, I'm not exactly sure, but I feel like there's a bunch of Sweebench cases that actually are like, invalid might be too strong of a word, a little bit not sure, but I feel like there's issues with running them, so they just don't work. Okay. And then max concurrency, is there a concurrency limit if I have 5, 10, 100 simultaneous codex?

38:415 and 10 is totally fine. Do we actually have a, I feel like we did introduce a limit for fraud reasons. I don't know what it is. Yeah, I think right now it's 60 an hour. Wow. So one per minute, I'm just going to... Yeah, but look, this is literally the point, right? It is. So long-term, we actually don't want you to have to think about if you're delegating or pairing with AI. If you imagine an AGI super assistant, you just talk to it, and it just does stuff. It answers quickly if it needs to. It takes a long time. And you also don't have to only talk to it. It's also just present in your tools, right?

39:17So that's the long-term thing. but in the near term like yeah this is a tool you delegate to and um the way to use it that we see like you know going back to the i guess that maybe the title of this podcast of like best practices it's like you must have an abundance mindset and you must think of it as like like not using your time to explore things and so like you know often when something a model is going to work on your computer and it's going to work on your computer you'll like really craft the prompt because you know then it's it's going to use your computer for a while and maybe you can't but the way we see people who like love codex the most using it is they don't they think for like maybe 30 seconds max about their prompt it's just like oh i have this idea like boom oh like there's this thing i want to do like boom oh like i just saw this bug or like this customer feedback thing like and you just send it off and so yeah like the more you're running in parallel actually i think that i mean the happier we are and like the happier we think like users are when they see it like that's just the vibe of the product really yeah i would i would pass my own anecdote so i I was on the trusted testers team for this thing, as both of you well know.

40:20And I was using, I found out I was using it wrong. I was using it like cursor. Like I had my chat window open and I watched it code. And then I realized I wasn't supposed to. And I was like, oh, like you guys are just firing the things off and like, you know, going on about your day. And yeah, that was a change of mindset. Yeah, one. Yeah. Real quick, I'll keep it brief. Like one thing that's quite fun is like use it on your phone. because somehow just being on your phone just flips the way people think about things. So we made the website responsive. We'll pull it into the app eventually. So try it.

40:53It's actually super fun and satisfying. Okay, so yeah, it's not... There was a voice... There was one of the videos that was showing the mobile engineer coding with it on his phone, but it's not available in ChatGPT's app. Okay, yeah. Yeah, not yet. Just one question I got from the mobile. I got the notification that I get. When it starts the task, it says starting research the same way the deep research notification is. Is it using deep research as a tool or did you just reuse the same notification? We just used the same notification, yeah. So you mentioned the compute platform, you mentioned how you share some of the infrastructure with RL.

41:33Can you maybe just give people a high level of what the codex has access to, what it doesn't have access to? Like, it doesn't look like people can run commands themselves. They can only instruct the model to do it. Any other things people should keep in mind? Yeah, so, and I'll say it's an evolving discussion as we figure out what parts we can give folks access to and the agent and what we need to, like, hold back for now, right? And so we're learning and it's really, we would like to give humans and agents alike as much access as possible within safety and security constraints. What you can do today, right, is as a human, set up an environment, set up scripts that get run.

42:19These scripts typically will be installing dependencies. I expect that to be maybe 95 % of the use case there. And just really get all the right binaries in place for your agent to use. We actually do have, like, a bit of an environment editing experience where, as a human, you can drop into a REPL, try things out. So, you know, please don't abuse it. But there's definitely ways for you to interact with the environment there. We laugh about that because, like, earlier I mentioned scope creep. Like, we weren't planning on having a REPL, like, to, like, interactively update your environment. But, like, you know, anyway, we tried.

42:54Josh was like, oh, man, we need this. And so that was, like, an example. Scope creep. Thanks for doing it. We do have rate limits in place. and we do monitor that very carefully. But, you know, there's interactive bits of there to, like, to get that going. But once the agent starts running, right, what we actually do today, and we're hoping to, like, evolve on this, is we'll cut off internet access because we still don't fully understand what letting loose an agent in its own environment is going to do, right? You know, for now, the safety tests all have come back very sturdily, like, you know, it's not susceptible to sorts of certain exfiltration attempts on prompt injection.

43:28But there's still a lot of risk to this category, so we don't know. And that's why to start, we're being more conservative there. And when the agent's running, it doesn't have full network access. But, you know, I'd love to be able to change that, right, allow it to give limited access to certain domains or certain repositories. And so all this to say, it's like something we're evolving as we build out the right systems to support that. Not sure that quite touches on your original question. The last thing, though, that I do want to mention is, like, there's an interactivity element with like, you know, as the agent's running, sometimes you're just like, oh, I want to like correct it, tell it to like go somewhere else.

44:07Or, you know, let me maybe fill this part out and then you can take back over, right? We haven't quite solved those problems either. What we really wanted to start was to like shoot for the fully independent, just like deliver massive value in one shot kind of approach. But yeah, we're definitely thinking about how we can weave human and agents together better.

44:34I mean, for what it's worth, I think the one-shot thing is a good angle that the other people, you know, like this is me comparing you to alternatives like Devin and Factory and all the others. there are more focus on multi-shot human feedback all these but like you know i so i have a website i'm working on and i gave it a request and i compared it all the others that was my test for codex and it did one shot it i posted it the screenshot as a tweet just earlier today and it's uh i think it's really good especially if you're running 60 at a time so i think that that really makes sense but it's it is a very ambitious goal because like human feedback is a crutch that we like to use.

45:20It also, I think, makes us write more tests, which is annoying because I don't like to write tests, but now I have to write tests. Fortunately, I'm now getting Codex to write my own tests. And I really like on the livestream as well, you can just kind of ask it to just look at your code base and just suggest stuff to do because I don't even have the energy to figure out what I should be doing. Yeah, and delegated delegation. I thought that was a great line. Yeah. And we're not saying one form factor is better than the others, right? Like, you know, I love using Codex CLI. And it's really, we just want, like, as we talked about in our interview, when I was interviewing OpenAI, like, you really want both modes.

45:59But I think what we see as, like, the role of Codex here is to really push the frontier on that sort of single shot autonomous software engineering. Yeah, I kind of think of the research preview as like our thought experiment. It's like, you know, what is coding agent like in its purest, most AGI-pilled or scale-pilled form look like? And then maybe, I mean, for me personally, I don't know, part of what excites me about working at OpenAI, it's not just solving for developers, but it's just really thinking about how does AGI benefit all of humanity and what does that feel like to non-developers as well?

46:35And so for me, what's really interesting is like thinking of Codex as an experiment for like what it'll feel like to be in other functions, you know, doing work. And like the goal for me to build towards is a vision where it's like, you know, we do the work that's like ambiguous or creative or hard to automate in whatever way. But otherwise we just have like agents like that we're delegating most of the work to. But these agents, they're not like this like long horizon thing versus short horizon. They're just like kind of ubiquitously available with you. So, yeah, we decided to take the purest form to start, which we thought would be the smallest scope thing to ship and probably isn't.

47:14But, yeah, then we're going to bring these things together. Okay, I think we have time for a couple questions. I'm just going to double click on the research preview a bit. It is a research preview. What is left? What do you think would qualify it to be a full release? you know on the live stream Greg mentioned this seamless transition between cloud and CLI is it that or are there other things on your mind? I mean to be completely honest the part of why we believe so much in iterative deployment I can give you some of my thoughts now but also we're really curious to see because this is such a new form factor but you know some of the items that are top of mind for me are like multimodal inputs you know we've talked Yeah, I know you like that, right?

48:01Yeah, like another example would be like, you know, just giving it a little bit more access to the world. You know, a lot of folks have requested for like forms of network access. You know, I also think that right now kind of the UI that we shipped is actually one that we iterated around. It's like a fun story there, but like, and it's one that people find useful, but it's definitely not the final form of what it is. And like we would love for it to be like much closer with the tools that developers spend time in. So like, those are some of the themes we're thinking about. But, you know, to be clear, we'll iterate and figure that out.

48:38I wanted to ask, why did you put finding a typo as one of the onboarding things? Because I used it, and then I saw it, and it's literally just grepping for potential type. It's like searching, grepping for like selenium with an N, or like something with like some TGN instead of NG. but it went through like 50 of these and then finally found will misspelled this w-i-l without the two thing but it was really cool to see what it thought like default spelled as like d-e-f-u-a-l-t it's like it just grabbed all these different things and then eventually got there like why why did you pick that task honestly t when i were talking about it and he was like Like, listen, it would be funny if I had a typo as I type this prompt out and just make it a little bit meta.

49:31You know, nervous fingers on a live stream. So maybe optimize for that. I have noticed it. You know, it likes to do TEH to look for the. And so it's a work in progress. That was great. Any parting thoughts, call to action? Are you growing the team? Do you want specific feedback from the community? Yeah. I think for me, the one thing that is really on my mind for getting better at over the next few months is really helping you customize your environment in a more high-fidelity manner. It turns out the good news is the agent can do a lot of really good work with only the basics, right? It's much like if your dev machine is borked and you're sort of looking at your editor but none of the type checks are working, a lot of folks can still actually do a lot of good work.

50:24But how do you get close at last 30%, 40 %? It's really hard because there's such a wide variety of environments out there, but especially would love feedback from folks on how they would like to see their environment customized. Do they want to just ship us a darker image? Would they rather have us support dev containers? So the form factor of how you do, the DX of how you do environment customization, still very much an open question that we need to improve on. Yeah, big plus one to that. And I think for me, maybe the thing that I'm most interested in is, hey, this is like a new shape of tool to collaborate with.

51:03And I'm just really interested for people to try working with it in as many different ways as possible and kind of figure out where does it work well in your workflow? Like you mentioned earlier, you were trying to use it kind of like your IDE and then you realized it was different. So I would just love for people to take advantage, especially now, like we're very intentionally just providing like very generous rate limits so people can try it. Like we just want you to try it and figure out what sticks, what works, what doesn't, how do you prompt it? And then we want to learn from that and use that to lean in.

51:30So yeah, I guess my parting call to action is like, please go try it out in JatchBT, use it as much as you can, especially now, and then let us know like how you like to hold it basically. I'm worried about the pricing when it happens, but yeah, I'm going to abuse this. Why not? Why not? Yeah. Send us feedback on pricing too. Yeah. Okay. It's too early to talk about pricing, right? Yeah. It's too early now. Yeah. Okay. All right. But yeah, based on cloud code, that's the thing that people are worried about, right? And cloud has started to introduce some kind of fixed pricing and variable pricing.

52:08And I think it's a huge mess. Like there's no right answer. Everyone just wants the cheapest form of code attention they can get. So yeah. Good luck. Thanks. I mean, my take is, I don't know what's going to make it into it. But, like, we aim to deliver a lot of value, right? And it's on us to show that and really make people realize, like, wow, this is, like, doing very economically valuable work for me. And I think a lot of the pricing can fall from that. But I think that's where the conversation should start. Are we actually delivering that value? Yeah, awesome. All right. Well, thank you so much.

52:46Yeah, thanks for working on this. and thanks for sharing your time. It is, it's been, it's been a long time coming, but I think people can start seeing like, OpenAI in general is getting very serious about agents. It's not just coding, but coding obviously is the one loop that is self-accelerating that I think, obviously you guys are super passionate about. It's really inspiring to see. Yeah, super excited to, super excited to just like, shift everyone this coding agent and then yeah, like bring it together and to just like the, you know, the general AGI super assistant. yeah so thanks for having us on thank you guys thank you

From the publisher

ChatGPT Codex is here - the first cloud hosted Autonomous Software Engineer (A-SWE) from OpenAI. We sat down for a quick pod with two core devs on the ChatGPT Codex team: Josh Ma and Alexander Embiricos to get the inside scoop on the origin story of Codex, from WHAM to its future roadmap.

Follow them: https://github.com/joshma and https://x.com/embirico

Chapters

- 00:00 Introduction to the Latent Space Podcast
- 00:59 The Launch of ChatGPT Codex
- 03:08 Personal Journeys into AI Development
- 05:50 The Evolution of Codex and AI Agents
- 08:55 Understanding the Form Factor of Codex
- 11:48 Building a Software Engineering Agent
- 14:53 Best Practices for Using AI Agents
- 17:55 The Importance of Code Structure for AI
- 21:10 Navigating Human and AI Collaboration
- 23:58 Future of AI in Software Development
- 28:18 Planning and Decision-Making in AI Development
- 31:37 User, Developer, and Model Dynamics
- 35:28 Building for the Future: Long-Term Vision
- 39:31 Best Practices for Using AI Tools
- 42:32 Understanding the Compute Platform
- 48:01 Iterative Deployment and Future Improvements

More from Latent Space: The AI Engineer Podcast

All 247 episodes
ChatGPT Codex: The Missing ManualLatent Space: The AI Engineer Podcast
Listen in VO