In short
Latent Space Podcast Episode Summary
Podcast Details
- Title: Latent Space: The AI Engineer Podcast
- Description: A podcast by and for AI Engineers focusing on news, papers, and interviews related to cutting-edge AI technologies and frameworks.
- Episode Title: ⚡️GPT5-Codex-Max: Training Agents with Personality, Tools & Trust
- Guests: Brian Fioca and Bill Chen from OpenAI
Episode Overview In this episode, Bryan and Bill share insights from the frontlines of OpenAI's Codex and GPT-5 training teams. They discuss the development of Codex Max, a long-running AI coding agent capable of extensive background tasks, context management, and parallel task execution.
Key Discussions
- Introduction to Codex Max
- Capabilities:
- Functions for over 24 hours continuously.
- Manages its own context and spawns sub-agents for parallel processing.
- Purpose of Naming:
- "Max" implies maximalist qualities in speed and efficiency.
- Training for Personality and Trust
- Importance of developing a coding agent that communicates effectively and can plan its tasks.
- Behavioral attributes like checking work and strategic planning are integral to building trust with developers.
- Differences Between Codex and GPT-5
- Codex is optimized for coding tasks, while GPT-5 is more general-purpose and can adapt to various tools.
- Codex has strong preferences for specific tools (e.g., preference for `rg` over `grep`).
- Tool Utilization and Model Training
- Emphasis on training agents to efficiently utilize tools.
- Model habits can be developed; for example, naming conventions can enhance tool-call performance.
- The Abstraction Layer Shift
- Transitioning from model-centric architecture to agent-centric frameworks.
- Full-stack agents can be integrated into platforms like VS Code to streamline coding workflows.
- Rise of Sub-Agents
- Codex Max can create additional instances of itself to handle tasks more efficiently through delegation and context management.
- Real-World Evaluations
- Shift from academic performance benchmarks to applied evaluations focusing on real-world impacts.
- Approximately 50% of OpenAI employees use Codex regularly.
- Multi-Turn Evaluations
- Exploring the next frontier in evaluating AI performance through multi-turn tasks.
- Introduction of concepts like "job interview eval" to assess coding abilities under various scenarios.
- Personal Automation Beyond Coding
- Coding agents are expanding into personal automation tasks, such as organizing files and managing workflows.
Future Vision (2026 Predictions)
- Increased computer use and trust in coding agents.
- Democratization of advanced coding capabilities for all firms, not just top-tier companies.
- Agents capable of handling complex refactors and building integrations seamlessly.
Key Takeaways
- Codex Max represents a significant advancement in coding agents, focusing on long-term operation and task management.
- Trust in AI agents is built through clear communication and behavioral training.
- The evolution of coding agents is enabling more comprehensive personal and professional automation tasks.
- The aim is for coding agents to be accessible and beneficial across various industries and organizations, enhancing productivity and capability for all developers.
Links and Resources
- OpenAI Codex: [OpenAI Codex](https://openai.com/index/openai-codex/)
- Latent Space Podcast:
- [Latent Space on X (formerly Twitter)](https://x.com/latentspacepod)
- [Latent Space Substack](https://www.latent.space/)
Episode Chapters
- 00:00:00 Introduction: Latent Space Listeners at AI Engineer Code
- 00:01:27 Codex Max Launch: Training for Long-Running Coding Agents
- 00:03:01 Model Personality and Trust: Communication, Planning, and Self-Checking
- 00:05:20 Codex vs GPT-5: Opinionated Agents vs General Models
- 00:07:47 Tool Use and Model Habits: The Ripgrep Discovery
- 00:09:16 Personality Design: Verbosity vs Efficiency in Coding Agents
- 00:11:56 The Agent Abstraction Layer: Building on Top of Codex
- 00:14:08 Sub-Agents and Multi-Agent Patterns: The Future of Composition
- 00:16:11 Trust and Adoption: OpenAI Developers Using Codex Daily
- 00:17:21 Applied Evals: Real-World Testing vs Academic Benchmarks
- 00:19:15 Multi-Turn Evals and the Job Interview Pattern
- 00:21:35 Feature Request: Batch Multi-Turn Eval API
- 00:22:28 Beyond Code: Personal Automation and Computer Use
- 00:24:51 Vision-Native Agents and the UI Integration Challenge
- 00:25:02 2026 Predictions: Trust, Computer Use, and Democratized Excellence
---
These notes provide a comprehensive overview of the podcast episode, capturing critical insights and discussions around the advancements in AI coding agents and their implications for the future of technology and programming.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:03Okay, we're here at AIEcode and we have two of our speakers, Bill and Brian. Welcome. Hi. Latent Space. Thank you for having us. Bill O 'Brien, I know you've been a listener for a little bit. Oh, yeah. What's your take on Latent Space? What role does it perform in your function at OpenAI? Yeah. I mean, first of all, love the name. I'm a massive Latent Space context management person. Tell the story behind the name, maybe the chance. Yeah. So we never had Latent Space as a name at the start. It was called L-Space. Interesting. And one of my readers donated the domain name Layton.space. He's like, you want it.
0:42I'm like, yeah. Awesome. So Layton just like came accidentally. So you're in the ether, but like I didn't have the domain. Yeah. So I just like called it L-space. L-space is like a Viz school. Nice. So it's for my name. Yeah, no, it's amazing. I love it because you're like always on the cutting edge and it goes into a lot of detail about all the things that like I should be keeping up with as part of my job. And there's so much to keep up with, right? So there's only so many sources of really good, high-quality information for what's happening on a deep level. Well, you guys have your own podcast now.
1:13So I'm like, hey, prepotition. Yeah, well, I still listen to yours. And I still think yours is really good. So you guys, I guess, are representing startups team, Codex. Yeah. All the things. You just launched Codex. Yeah. The Clarence Max. Yep. Jardini yesterday. Yep. We're going to name him. I do. People do make friends. I think Thibault was like, yeah, you know, we're good at a lot of things, but not only me. I was like, well, why call it Max? Was there any like internal discussion? Yeah, I mean, it's complicated because it needs to be differentiated from the previous one. And the idea is like Max can run for a really long time.
1:52We can go 24 hours or more. I've actually like sort of had it gone for more than that. And the name is, you know, it's Inside Codex on the web. is that uh how do you when you say a really long time 24 hours oh i on my on my oh that's i think that was on the web inside of credits i'm not sure but i've actually done it on my local computer for for quite a bit longer than 24 hours over the course of a couple days but closing my laptop and reopening it but but the the name you know you could come up with something like pro but pro is is sort of like slower more thoughtful max is about sort of like speed and maximization like maximalist.
2:30For this Mono, it can run for a long time, but it can also actually, for the same types of problems, it can actually get to the right answer faster. It's simply better and faster. So I think part of what you guys are speaking about is the training that goes into something like, yeah, that's right. Bigly, people just kind of wave their hands and say, or, but what specifically have you learned about what's What's a good path you saw us on? So I got to, I mean, this sounds weird to say, but I was lucky enough to be really close to the training team while GP5 was training. And one of the big things that we focused on, Bill was there too, we focused on personality, right?
3:15So it's really important to build trust with developers for like how a model works. And if a model doesn't act the way that you expect it to do, or if it doesn't work alongside of you as well, you're not going to really trust it. You're not going to get as much out of it. So for coding, we thought, okay, well, what is the best personality for a coder, for a pair programmer, for somebody you trust? And how do we like eval against that? How do we come up with behavioral characteristics? And we came up with things like communication. It needs to keep you impressed of what's going on while it's working.
3:48Planning, like come up with a strategy, do some searching around, like figure out context gather, figure out what to do before you just dive in, if it makes sense to. and then check your work. And so these are just best software engineering practices that turn out to be behavioral characteristics and we can measure the model's performance on those behaviors and grade it that way. Yeah, I will say that another key aspect to how we trained the model is you work really, really closely with some of our coding partners and a lot of those folks that lead on the bleeding edge And so they have a lot of understanding of what particularities they need.
4:31And we really focused on sort of those areas and really dive deeply into those. Yeah, that's right. Especially tools, right? So like different harnesses have different tools. Some people have context, like semantic search. Some people have different ways of doing code edits. And initially, you know, our models are trained the way they were trained to use tools. And that kind of bakes in a habit. And so we've been getting the models better at using different types of tools. Yeah, there's a lot to follow up on. But I'll go tools first, and then I'll go back on the personality base. But the news-wise, I think the communication where the 5 Codex just came out was, well, this is the model trade for our Codex, not necessarily your choice, right?
5:15Has that message changed for other startups using the 5 Codex model? Right, no. So Codex is, just to be clear, Codex is the frontier coding model that we have that is optimized for its harness. The Codex team is very focused on creating a coding agent, and they want it to work perfectly inside of the shape of the harness and API that we have. So they're completely unbounded. It's open source. Yes, that's open source, and the model is available in the API. So that's what they focus on. And then the conflict is, well, you just said other startups have other tools, and obviously I know that. It is possible.
5:49One thing to mention here is I think we can probably Disimpang a little bit on sort of The codex Apart from the sort of the mainline models The codex models are sort of focused On the agents itself Right, like the codex agent itself The model has been trained with the agent specifically in mind It actually turned out to be Somewhat, even sometimes easier To integrate because we come into It with a firm opinion on what the sort of Best way Of using it looked like And so some folks that we work with I actually really appreciate that we come into it with that opinion. While for the other ones that has more of a general or specific tools that they definitely need, the mainline model is one that's more general in a sense.
6:35And that's sort of what Brian was referring to when he talked about GPT-5's tools. Yeah, so the 5.0 non-codex is more general across the board. It can respond to things that are, it's much broader than just codec. It has coding capabilities that are also mirrored in codecs and they work together to keep that trued up. But since it's more general, it does have more steerability to different types of tools. And when you're implementing tools, the model can get bogged down if it hasn't seen a tool that it's used to, and it might take more time thinking about how to use it or make more mistakes. So our recommendation is, if you're wanting to go bleeding-edge coding focused, pay attention to the Codex line and the Codex SDK and the Codex models, because that's the one that's really aimed at that.
7:23You'll have to do some work to look at how we're implementing our tools inside of Codex to maximize its capability without bogging it down. But people are having success bending it in ways that maybe we haven't thought of. If you think of the mind, I always want to pry. Sure. Do you have any examples? Is it bending in ways you haven't thought of? Yeah, so I think Codex is trained with terminal tools in mind. And so what we've thought would be the case is you will essentially only have to strip out all of the tools except for the terminal tools. But we found some partners of ours, like they discovered that what you can do is that you can actually still have a lot of the tools just named in the same way as the terminal tools, as well as having the same input and output.
8:16And all of a sudden, the tool call performance jumped up by a lot. Yeah, and Koditz loves RIP-GREP. So if you make a RIP-GREP tool and tell it to use it, it'll use it. So if you call it GREP, it actually does a little bit worse. But if you call it RG, it actually does really well. Right. Yeah, this is something that we ourselves only discovered. This is one of the coolest things about model training is literally they develop habits. Just like a person does. Like if you're like working on some podcasting tool, right? You're really good at editing. And then somebody makes you use a different one.
8:48It's going to slow you down. You're going to get kind of bogged down and make mistakes. Sure. But I don't know if like, yes, that's very humid. But I don't know if I'd call it cool because it's supposed to generalize. Well, right. That's the end goal. Yes, of course. And so that's what we're doing with the five series of models. They're way more general. And Codex is focused on maximizing coding. and those are the sort of two horizons that we're working on. Yeah, awesome. I want to go back on personality. I know you hate that word sometimes. It means different things to do. Yes. And when it comes to people who are very keen on model research, model personality is much more like what I think VB1-Eatropic would say is like it warms your friend.
9:36It is for your... I agree about just having people's emotional state whatever. And so it's really jarring when that is also applied to total agents, where like, well I want to talk to Vashii, right? Silica value HP Ultra is also saying Anton, but it could be they don't do it with the free cost. I think the other thing is also, but what doesn't matter because you said a lot of things about like commenting is that you're going to use your engagement or that. Doesn't matter if it's so coronavirus anyway, right? Like you're going for 24 hours, you're closing your laptops. You have like the extra high parameter now.
10:11It doesn't matter. Exactly. So here's, we're in this world right now where we're in between a situation where people don't quite have, like the models don't quite have the trust of senior engineers or engineers doing like very important work. And so we found, our customers have found that people really want to follow along with what it's doing so they can like interject or stop it, or at least understand what it's thinking so they don't waste all the kinds of time like doing a rollout that they have to throw away. So with the 5 series, because it's more general, and it's just about as good as coding as Codex for a lot of things, we've taught it to be more communicative.
10:48And so it has preambles before tool calls, it'll say things like, I'm about to go look for this. Yeah, and you can steer that really well. I actually really like it. I have I've created like a personality I tweeted about this, I created a personality for my coding agent, because I really like my tools to be kind of like fun to work with if I'm in there with them. And so I have it. It's got this like it gets really excited if we do something together. And like because I want to wake up in the morning and be like, oh, I'm going to go work on this project with my my buddy 5.1. Right. But some people don't like that.
11:18And also for, like you said, long running agentic tasks that can get in the way. You're burning tokens that don't really matter if it's running in the cloud. So 5.1, you can turn that off. You can prompt it not to do that. But the Codex model can't actually do that. And it relies on the reasoning summarizer to give you that update. But I guess more broadly, why should people know or think about in terms of what will be asked with voting models in general? More broadly than just like the media but experience release, just like what trends are you seeing? What discussions are active? Our talk today is focused on talking a little bit about sort of the trend that we're sort of seeing.
12:01is the abstraction layer really moving, starting to move upwards from the model layer towards the agent layer. As I said, we train our models starting to be a little bit more opinionated, especially with regard to Boeing model, like Codex. And the models are really good at doing certain things when inside of a certain furnace, a certain type in search shape. And so we're packaging that up more closely. So we're actually shipping this entire thing entire agent altogether, then you can actually build on top of that agent. That's one of the patterns that we're seeing here is rather than focusing on optimizing with every single model release, you're actually just being able to plug in an agent like Codex into your platform and be able to use an app box.
12:46Yeah, and you're seeing Zed use this. GitHub VS Code lets you just package a whole agent to work inside of it. That way, if you're building a coding tool, like Zed, and you don't feel like having a whole team keep up with all every single model release and every single API change and how to update the harness to do different kinds of sandboxing and all that kind of stuff, you can just build one layer above. And that is actually super powerful because coding is just like one agentic behavior. It turns out it's a really nice one to start with because you can measure the performance sometimes easier with a lot of other ones.
13:22But it also gives the model the capability, right? So we started out with like chatbots. Like you're having a conversation. Let's give the chatbot a tool to use. Okay, so now you have an agent that can run commands. Well, let's give the chatbot agent a codex to use. So now if it doesn't have a tool, it can make a tool that it needs to solve a problem, right? So that's like another layer of abstraction, and it's not just coding. You can write software that has an agent that can spin up a codex instance and write a custom plugin for your software for that customer's API, right? And so now your software is self-customerizable because it has its own team of people inside that can do integrations at launch.
14:04Yeah, solving integration engineering is a GI. Yeah, one thing, one theme I'm finding at this conference so far, even early, like the first heat rocks, I think people are starting to really explore sub-agents, agents that use more abstractly, agents that use agents. And we used to call it multi-agent. I don't know what we do now. I don't know if there's any thoughts on your end about this where like you get to call i guess like a very basic example is what you just said which is that the agents can uh create another instance of codex the case a tool and then drop it just use the tool okay is there a case for skating like some agents yeah i think so i mean codex max was designed for that right so it has its own compaction and context management.
14:53Codex Max manages its own context window. And so it can run basically forever without you having to worry about it while it's inside of the Codex harness. And that lets you do a lot of different things. You can essentially have it hand off its own context to other sub-agents, right? So letting it sort of like spawn different agents to do more of its work in parallel and all kinds of things like that. So it's built for that. I mean, we're just sort of like starting to see the indications of like what that means but that's i think the future and we're really excited about that yeah uh it's really i think like as i said uh the the trend that was sort of serving here really moving up the attraction layer to the agent to the agent layer really allows you to do a lot of cool things like brand new spaceship spending a few ages created new abstractions as things.
15:46As the long-running agent workflow continues, and right now we're building all the primitives as well as models, specifically with animates. Yeah, and it's really about moving the threshold up further, right? Like I was saying before, I now trust Codex to do some of my hardest work. I haven't written a single line of code by hand in months because I know what I can trust it to do. You're the Forbes person that said that in the last way for us. Yeah. No, it's real. I mean, I've actually launched something. There's an open source project that I did. There was a Codex upgrade pack for migrating from completions to responses that was totally written by Codex.
16:24And I didn't write a single line of that code. And now it's out there. It's open source. Actually, most of the folks that open AI, well, initially when Codex first launched, it was around 50 % of folks that open AI started using it. But now they go, but those folks that open AI. That's very true. We use it every day. the way that we do it is we're really good at evals right like in order to develop trust and like build a product that can do more than you design it for which is really what we're talking about here you're making an agent that can like solve its own problems um you have to get really good at figuring out how to build those guardrails and evals around you know what is it doing what is it allowed to do and check it in production so we have all of this platform tooling now around agent traces and rollout traces and and coming up with evals for that and building you know graders and all the things you need to sort of like maximize the pipeline.
17:12So you can let it go and then be like, okay, I don't really like the way it did that. Great, have it Metaprompt itself so that next time it actually does a better best practices. One of the biggest views in terms of what is the organizational capabilities that OPIC investigate is Azure Pyroxy. Can we say more about that? Like, why is that suddenly a big priority now? Obviously, I think there was, OVA always did internal emails, but now it's like a team that is more over-facing and then maybe go to this random era. The path to AGI really goes through evals and well, I'm sorry, that was a little...
17:50It's so true. It was repeated way too many times but I think there are a lot of academic evals, right? There's like sweep ends, there's other like you name it but I think there's a slightly lack of evals off the real world on sort of what people care about the most. And we want to make sure that whatever we're developing model-wise as well as product-wise are aligned and are actually making the most amount of useful impact on this world. And Applied Evals is really in that direction, capturing all of those sorts of real-world use cases and things for us to hill climb together. I like to think of it as like we have, I mean, people say it's a PhD in an API, right?
18:35But if you hire a PhD student, they don't know how to do the job. You have to give them a job description. Okay, that's a prompt, right? So now you have your policy, and then you have them do the job, and they're going to kind of flail around, right? So they need mentorship, they need guardrails, they need evals, performance reviews on how to do their job, the best practices. And so what we're doing is we're trying to put our models out there and see what they're good at, what they're not good at, talking to our customers. They're like, oh, we could really use your model for more things if it could do this one thing, here's our eval for it.
19:07Or help us build those evals with you so that we can see where we're deficient and go back and train the model to be able to do that job in the way that we wouldn't normally get to see it form. Yeah. How do you do multi-turn evals? Because I think that's the really hard thing that, I mean, sometimes you need multi-turn if it doesn't get around the first go, but it could just get around the first go, then it's no longer multi-turn, right? So then what? Do you want to take, I have some ideas. Oh, yeah, you go. I mean, I've built a few myself. This is sort of like my personal work. I think this is like an area that people are just now getting into, right?
19:45We have LLM as a judge. You can use LLM as a judge to look at an entire trajectory and see, okay, over the course of all of this, like how well did it perform, what did it do? And then you could maybe like walk it back a step to the part where you don't like, and then you could have the model run the next step with the instructions, grade it on that, and then have it improve itself. Oh, I don't like the way that you... We do this all the time inside of harnesses. It's like, that was a good answer, but I don't really like how long it took you to get there. So can you give yourself better instructions for doing that next time?
20:18And it'll write something and we'll add it in there and then suddenly it's better, right? So that's one way of doing it. Yeah, I think multi-turn evals, most of the companies or startups that we work with, like these days the agent runs in a multi-turn way right and then so therefore if you can build an agentic harness that works in a multiple turn way you can eval it and then there are like also academic benchmarks already does this in some ways like cow bench and now we have like tau square bench that does this like particularly well and would definitely certainly take inspirations from that i have this idea i call it like a good job interview eval i haven't finished it but really if you're evaluating a coding agent, what do you want it to be able to do?
21:04You want it to be able to take an underspecified... Imagine you were interviewing a developer. You give them a problem. Hey, go implement a string reverse or whatever. And then it's up to them to ask for, okay, well, I need more information. What are the constraints here? And then you judge them on that. And then they start implementing it. You give them some modifications. You grade them on that. You can imagine building with an LLM rollout that is promptable and the model responds and then you can kind of grade the whole thing. One thing I would love and this is like the feature request part of the podcast is batch multi-turn eval API.
21:44So batch API is single turn but you can't really batch multi-turn requests. Is that already doable? Batch multi-turn requests, I don't believe it. You can't do it yet but yeah I think that's like a really valid. because you need evals to be cheap as possible. Yes. They're not that time sensitive and you want to run it overnight when the things are cheapest. Yes. Well, feedback taken. Feedback taken then. But that's the thing. Like every day we're trying to make the platform better and right now evals is certainly part of it. Literally how we make product feature updates is we talk to people like you.
22:18They're like, hey, can you do this? I mean, it's super like, yeah, if I'm going to throw thousands of runs at this thing, you know, I should probably spend some time worrying about costs. Speaking of which, what are you trying to do though? I mean, Devin and Cascade. So I have a personal side project where I want to make Devin for non-coding. I love Devin so much. I think Slack, my kind of semi-hot take that I'm floating around because just to see how it feels is I think Slack is the ultimate user interface. Yes. For work, right? I don't want to read email. I just read Slack all day. I interact with my email agent through Slack.
22:55So basically, I'm building a dev info email. Yeah. Well, that's the thing is you can use devin to do that, right? Like a coding agent, like Codex, a CLI, it used to be back in the old days. Like I started out in the 90s working at IBM as a system administrator, and I had to write my own custom software and bash scripts and whatever to actually solve real-world problems every day. And so I had this toolkit of scripts that I made, right, that were like organizing file directories or doing like other random things that weren't necessarily writing code. Yeah, yeah. And so you can get... For not putting use cases.
23:30To just like sort through your email using like Elm or something, right? In the terminal. Or like have it generate like snippets of video clips from YouTube that you can watch later or things like that. You know, I never thought about that, but I do that all the time as part of Lanespace. Yeah. I should probably invest in that tooling. I had Codex go through my really messy directory of all of these experiments that I was running and completely organized them and put them into shape. And it was so wonderful. I used it for something that's more boring, organizing my desktop. Yeah. We have a lot of files on the desktop and Codex is really good.
24:08People think I'm all in. Codex, my MG0416.jpg. Yeah, well, just find all the images and put them in one folder. I think that even that's something Codex can do. I think that's one of the big things that we're also seeing. Coding tools are breaking out of coding and just like everything there, personal automation. Exactly. Because the way, if you can think about before graphic user interfaces and browsers, like what did we, how did we interact with a computer? We did so through a terminal and we did so by writing commands and writing code and stringing them together inside of the terminal. So what you think about it is are those coding agents are actually a computer use agent, but for the terminal.
24:48Yes. Yeah. They're actually incredibly general. I would say that coding ages today are still not vision native enough. Like you have to try to get it to use vision. And oftentimes it fails still. We should use vision a lot more. Yeah, I would say, you know, I was going to end the episode with asking for your 2026 predictions. Like we sit down this time next year. What do you want to see? You know, what do you hope to see? I'll just kick it off with the easy one. Yeah. More computer use. And I think like where you say things like, oh, we'll have a coding agent build its own integration to your application.
25:22A lot of applications don't have APIs, don't have NCPs. The only thing you have is a UI, right? Yeah. Because they're a legacy or because they don't want you to take the data. But while the data is yours, you just have to, like in a non-profession way, take it by the user. Yeah, and I can continue just by sort of like saying that that's definitely going to be something I think is going to be something that we'll be capable of in 2026. But also the other thing that I am sort of really like looking forward to are codecs being able to do more, right? We're already starting to talk about how codecs or like coding agents can sort of use computers in novel ways.
26:06we're going to be able to sort of see more general and general use cases like that coming along as well and more sensible ways for you to build with those sub-agents as well. I really want to see the trust level go up even further, right? Like, at OpenAI, I get to work with some of the most amazing developers I've ever worked with in my life. They're incredible, like some crazy tech leads. I wish every company, no matter whether it's like a small dev shop in Alaska where I worked for a while or OpenAI be able to have on their team capabilities that you would only be able to get at a top-tier firm, right?
Read the full transcript
26:40So all of my teammates at all of these places could turn to a coding model and be like, hey, how do we do this crazy, awful refactor that we have to do to support this new customer that we have? Or like, wow, there's so much of a mess here. Or like, what's the best way to actually implement this new technology? And have it be so trusted and so right and so smart that we can actually perform better than we could normally get access to. Yeah, see? I think that's going to be any if I know a call session. Oh, yeah. We're Brian and Bill at OpenAI. And yeah, feel free to find us on our Twitter, socials, whatever.
27:18And then let us know how you're building. Yeah, and we love working with startups. And anytime you have feedback about you really wish the model could do this or their product could do this and you could unlock some massive capability, just let us know. Yeah, amazing. Will do. that's it thank you guys thank you
From the publisher
From the frontlines of OpenAI's Codex and GPT-5 training teams, Bryan and Bill are building the future of AI-powered coding—where agents don't just autocomplete, they architect, refactor, and ship entire features while you sleep. We caught up with them at AI Engineer Conference right after the launch of Codex Max, OpenAI's newest long-running coding agent designed to work for 24+ hours straight, manage its own context, and spawn sub-agents to parallelize work across your entire codebase.
We sat down with Bryan and Bill to dig into what it actually takes to train a model that developers trust—why personality, communication, and planning matter as much as raw capability, how Codex is trained with strong opinions about tools (it loves rg over grep, seriously), why the abstraction layer is moving from models to full-stack agents you can plug into VS Code or Zed, how OpenAI partners co-develop tool integrations and discover unexpected model habits (like renaming tools to match Codex's internal training), the rise of applied evals that measure real-world impact instead of academic benchmarks, why multi-turn evals are the next frontier (and Bryan's "job interview eval" idea), how coding agents are breaking out of code into personal automation, terminal workflows, and computer use, and their 2026 vision: coding agents trusted enough to handle the hardest refactors at any company, not just top-tier firms, and general enough to build integrations, organize your desktop, and unlock capabilities you'd never get access to otherwise.
We discuss:
What Codex Max is: a long-running coding agent that can work 24+ hours, manage its own context window, and spawn sub-agents for parallel work
Why the name "Max": maximalist, maximization, speed and endurance—it's simply better and faster for the same problems
Training for personality: communication, planning, context gathering, and checking your work as behavioral characteristics, not just capabilities
How Codex develops habits like preferring rg over grep, and why renaming tools to match its training (e.g., terminal-style naming) dramatically improves tool-call performance
The split between Codex (opinionated, agent-focused, optimized for the Codex harness) and GPT-5 (general, more durable across different tools and modalities)
Why the abstraction layer is moving up: from prompting models to plugging in full agents (Codex, GitHub Copilot, Zed) that package the entire stack
The rise of sub-agents and agents-using-agents: Codex Max spawning its own instances, handing off context, and parallelizing work across a codebase
How OpenAI works with coding partners on the bleeding edge to co-develop tool integrations and discover what the model is actually good at
The shift to applied evals: capturing real-world use cases instead of academic benchmarks, and why ~50% of OpenAI employees now use Codex daily
Why multi-turn evals are the next frontier: LM-as-a-judge for entire trajectories, Bryan's "job interview eval" concept, and the need for a batch multi-turn eval API
How coding agents are breaking out of code: personal automation, organizing desktops, terminal workflows, and "Devin for non-coding" use cases
Why Slack is the ultimate UI for work, and how coding agents can become your personal automation layer for email, files, and everything in between
The 2026 vision: more computer use, more trust, and coding agents capable enough that any company can access top-tier developer capabilities, not just elite firms
—
Bryan & Bill (OpenAI Codex Team)
http://x.com/bfioca
https://x.com/realchillben
OpenAI Codex: https://openai.com/index/openai-codex/
Where to find Latent Space
X: https://x.com/latentspacepod
Substack: https://www.latent.space/
Chapters
00:00:00 Introduction: Latent Space Listeners at AI Engineer Code
00:01:27 Codex Max Launch: Training for Long-Running Coding Agents
00:03:01 Model Personality and Trust: Communication, Planning, and Self-Checking
00:05:20 Codex vs GPT-5: Opinionated Agents vs General Models
00:07:47 Tool Use and Model Habits: The Ripgrep Discovery
00:09:16 Personality Design: Verbosity vs Efficiency in Coding Agents
00:11:56 The Agent Abstraction Layer: Building on Top of Codex
00:14:08 Sub-Agents and Multi-Agent Patterns: The Future of Composition
00:16:11 Trust and Adoption: OpenAI Developers Using Codex Daily
00:17:21 Applied Evals: Real-World Testing vs Academic Benchmarks
00:19:15 Multi-Turn Evals and the Job Interview Pattern
00:21:35 Feature Request: Batch Multi-Turn Eval API
00:22:28 Beyond Code: Personal Automation and Computer Use
00:24:51 Vision-Native Agents and the UI Integration Challenge
00:25:02 2026 Predictions: Trust, Computer Use, and Democratized Excellence




