In short
TWIML AI Podcast Episode #730: How OpenAI Builds AI Agents That Think and Act with Josh Tobin
Episode Overview In this episode of the TWIML AI Podcast, host Sam Charrington speaks with Josh Tobin, a member of the technical staff at OpenAI. They delve into OpenAI's innovative approach to building AI agents, discussing three key agentic products: Deep Research, Operator, and Codex CLI. The conversation explores the development of reasoning models, the transition from traditional LLM workflows, and the implications for human-AI collaboration in software development.
Key Concepts and Topics
OpenAI's Agentic Products
- Deep Research: A tool designed for comprehensive web research.
- Operator: A system for navigating websites and automating tasks.
- Codex CLI: Enables local code execution and interaction with codebases.
Transitioning to Reasoning Models
- OpenAI's shift from simple LLM workflows to models specifically trained for multi-step tasks using reinforcement learning.
- Reasoning models allow agents to recover from errors, enhancing their ability to execute complex processes.
Practical Applications of AI Agents
- Unexpected use cases for AI agents beyond initial expectations.
- Potential for improving human-AI collaboration with concepts like "vibe coding" and enhanced context management in AI-enabled IDEs.
Challenges in Trust and Safety
- As AI agents become more autonomous, trust and safety considerations are crucial.
- The need for guidelines and mechanisms to specify how agents can use sensitive tools (e.g., credit cards).
Detailed Discussion Points
The Current Landscape of AI Agents
- Misconceptions: Many believe AI agents can fully automate tasks, but they often struggle with compound errors across multiple steps.
- End-to-End Training: Training agents from the ground up to manage workflows effectively and recover from failures.
Enhancements in AI Model Capabilities
- The use of reinforcement learning to train models for specific tasks is a game-changer, allowing models to learn through experience.
- The ongoing development of more advanced models that can effectively manage complex tasks and reasoning.
Future Directions for AI and Software Development
- The emergence of "vibe coding," where users provide high-level functionality goals while AI manages the technical details.
- An anticipated shift in the roles of developers, with increased emphasis on overseeing AI outputs rather than writing code manually.
Integration with Tools and Context Management
- Efficient context management is critical for enhancing the performance of AI agents.
- OpenAI's exploration of integrating tools and frameworks that allow agents to operate with greater autonomy and efficiency.
Codex CLI and Its Open-Source Nature
- Codex CLI offers a unique opportunity for developers to leverage AI for coding tasks, emphasizing collaborative learning through community contributions.
- The potential for Codex CLI to become an integral part of CI workflows and automate coding processes.
Conclusion The episode highlights the transformative potential of AI agents in various domains, particularly in software development. As these agents become more capable, the need for trust, context management, and collaboration between humans and AI will play a pivotal role in shaping their future applications.
For complete show notes, visit [TWIML AI Podcast](https://twimlai.com/go/730).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00In, you know, let's say 2023, 2024, the way most people are trying to build agents. is a human designs a system and that system kind of breaks the problem down into multiple steps. It assigns each of those steps to an LLM. Maybe there's some rules built in, but oftentimes like the real world is messy and that workflow that you built might be an oversimplification of the actual process. You know, a real expert at this task would follow. And when models are able to learn how to do the process by being rewarded for succeeding at the process, they're able to figure out in many cases something that's better than you could easily sit down and design yourself.
0:46All right, everyone, welcome to another episode of the TwiML AI podcast. I am your host, Sam Charrington. Today, I'm joined by Josh Tobin. Josh is a member of technical staff at OpenAI. Josh, it is great to have you on the show. We last spoke, can't believe it's been over five years now. That's amazing. Yeah. It is amazing. In fact, yeah, I think it was, I think we spoke in Vancouver, like right on the tail end of NeurIPS or during the NeurIPS there. I think that's right. Yeah, I had a paper there that we were talking through. I actually have it noted here, geometry aware neural rendering. Yep.
1:27Yep. Yeah, that was actually the last thing that I worked on and my last stint at OpenAI before rejoining again in September. Awesome. Awesome. Well, you've done a few things since then. I think full stack deep learning was one of the things that folks may have heard your name pop up on. And then you were working on a startup. Tell us about what you've been up to over the past few years. Yeah, so I left OpenAI back in 2019 and co-founded Gantry. We were a machine learning infrastructure startup. And when we were kind of thinking about sort of where to take that business at the sort of middle of last year, ended up kind of through a bunch of different circumstances winding up, coming back to OpenAI.
2:19And here I've been leading the agents research team. So our team builds the models that power our agentic products like Operator and Deep Research and the Codex CLI, which we launched a few weeks ago. And we're going to spend a bit of time talking about agents and those agents' agentic products. I'm curious in thinking about Gantry and what you were doing there and I think Gantry was in kind of a crop of ML infrastructure, ML ops companies that I think ran into this like all of the air getting sucked out of the room by Gen.A.I. I'm curious if you have any takes on that space and how it relates to what we're all trying to do with Gen.A.I.
3:10now. Yeah, absolutely. I mean, I think it was, you know, in the pre-ChatGPT era, I think a lot of companies' mental models was that every company is going to need to be training models. um you know before we had uh gpt3 and later models that did this more effectively the thinking was like um the way to create actual value with ai is to train models that are fit for purpose for the thing that your business needs to do um and uh so there's a whole kind of like category of infrastructure um that was built under the assumption of like every company is going to be you know every business out there more or less is going to be training its own models in the future.
3:54And so what tools will they need? I think what turned out to be true is that general purpose models, GPT-3, GPT-4, other large language models are pretty, are like quite good at most tasks that you can specify. For any given domain, you might be able to build a better model than GPT-4, but it's going to be pretty hard and pretty expensive. And if there's an off-the-shelf model that can just do those things well for you, it's so much faster, cheaper, more efficient to build like to what every business needs to build on top of that like commercially provided model instead of, you know, building up all the infrastructure expertise and data and sort of domain specific knowledge that you need to in order to build the model yourself.
4:41And so I think that kind of, you know, didn't fully invalidate, but it sort of, it made the business model that a lot of that crop of ML infrastructure startups was thinking about a lot less feasible because now I think my operating assumption is like when I advise companies on how to think about ML, I generally tell them, don't even think about training your own models until you've exhausted what you can do with the models that OpenAI or any other foundation model provider will sell you. Yeah, that's absolutely a smart approach. I recently wrote an article that reflected on a paper that I wrote back in that 2019 timeframe and I called out this idea that I thought that I got both right and wrong.
5:24The core idea was that in order to be competitive, businesses would need to become model-driven. This core idea of pulling patterns out of data and putting them into operational workflows in order to make decisions faster and more accurately than humans could was like a key element of what ML was offering folks. And I think, you know, businesses being model driven is even more true. But the idea that like everyone was going to build those models themselves, as you just noted, turns out to be less true because of foundation models, right? Yeah. And I think you're starting to see kind of a crop of CEOs now you know, publishing their thinking and the way that they're kind of urging their companies to use AI internally.
6:22And so I think that is like absolutely true and is actually coming true now that there are important companies that are sort of walking the talk when it comes to adopting AI and making it a pretty core part of a lot of what they do, which is kind of like what we all hoped or thought was going to happen back in the day. It just happened a little bit differently, I think, than maybe some of us anticipated. Yeah, I think you're referring to a tweet, an ex that went viral by Shopify's CEO, I forget his name. And he talked about kind of trying to drive urgency among their teams to adopt AI internally.
7:09And I think, you know, that maybe sets the stage for one of the things that we want to cover in this conversation, which is like what it's going to take to make agents real more generally. I think the presupposition is that, you know, we're not fully there. Like they do a lot of interesting things, at least, you know, when when bounded by, you know, whatever appropriate scope. But I think the general vision of, you know, my sidekick agent that, you know, sits there waiting to or even anticipating my needs and then going out in the world and operating on my behalf without constraint, you know, we're still a bit a ways from there.
7:53How do you think about what agents are really good at today, you know, relative to what we might want them to do? Yeah, you know, I think we're closer to that than it seems. I think that the misconception that a lot of folks have about agents is, before I came back to OpenAI, had spent a little bit of time tinkering, building agents on top of LLM APIs. And the problem that we ran into, which I think is the problem pretty much everyone runs into when they try to build their own agents, you know, using like workflows on top of LLM API calls, is that it's very tempting to think, oh, you know, LLMs are good at kind of making point predictions, making decisions.
8:39And so we build a workflow around this LLM. Then we can sort of automate a process and take a bunch of decisions and turn it into something that just happens on its own. But I think what a lot of us found trying to do this is that the temptation is there because you can create a demo of this very, very quickly. That looks great. but then when you try to actually deploy this and get people to stop doing this work themselves and instead delegate it to an agent, you start to run into all kinds of edge cases and failure modes and getting things to work reliably is really hard and sort of the root cause of the problem is that LLMs, most LLMs historically have not been trained to do agentic work and what that means is that at any given step of the process, maybe they're relatively accurate because they're pretty smart, general purpose AI systems.
9:34But as you run a process that requires many steps, the small errors at one step compound as you take multiple steps. So even if you're 90 % accurate on one step, if you have to take 10 steps, then your accuracy will fall off. And so that's, I think, kind of the core challenge of building agents on top of traditional LLMs. The missing ingredient has been that we need to directly train these agents end-to-end to do these workflow-like tasks. By doing that, you can train the agents in such a way that they see failures during their training and they learn to recover from those failures. So even if each step is only 90 % accurate or 95 % accurate, it.
10:24Now the model has seen what it looks like to fail at that step and it's able to reroute itself. It's able to think, oh, this doesn't look right. Like, let me go back and try that again. Do you have a canonical example that kind of illustrates this problem? Yeah. So a good example of this would be like, if you are, you know, if you're trying to kind of like build an agent that does research for you. So if you do a single web search and maybe you get the search term wrong, like the user doesn't know exactly what to search for. So you try to search for terms that seem relevant and you start to pull back a bunch of docs that don't have information that really gets to the heart of the problem.
11:08Then a naive agent might just get confused by that and it might think, well, okay, I searched for the term that I thought that they meant and all these docs are irrelevant. So maybe this is, you know, maybe what the user asked for doesn't make sense. Or maybe I need to go down this other rabbit hole. Whereas an agent that's trained to do web research and has been trained using reinforcement learning to sort of be good at this multi-step process, will in its training have seen many instances where it searched for the wrong term. and the training has incentivized it to learn to recover from those instances and instead go back and think, oh, I searched for this term, but I got results that weren't relevant.
11:56Maybe that means that I had the wrong search term. So let me go try again and pick a different one. And so the key is really like reasoning models, power that are trained end-to-end to solve the kinds of tasks that users need them to solve using reinforcement learning so that they're able to kind of like see these multi-step processes, see the kinds of failures that happen in training and learn to recover from them. It strikes me that the creation of models that are specifically designed to do this is one powerful difference that we're seeing in this new generation of models, but also it's like a collapsing of capability, a bunch of stuff around the model that was trying to like keep it in the right path to just like making the model better.
12:46You know, the deep research, I forget the name of the paper, but there was a paper that kind of proposed this deep research idea really early on. And it was like, you know, a multi-agent system with one agent that, you know, figured out search terms and the other agent that like reviewed the results and kind of rated them. And, you know, there were several agents in this loop and they would ultimately write, you know, whatever the report is. The impression that I get is that as we're training these models to more directly take on these tasks, we need less, it becomes less of a multi-agent type of solution.
13:29The way I see it is like in, you know, let's say 2023, 2024, the way most people are trying to build agents is a human designs a system. and that system kind of breaks the problem down into multiple steps um it assigns each of those steps to an llm maybe there's some rules built in and um uh there's two problems with that the first is the compounding error problem that we talked about before um but the second is just that like you know i think i think i'm uh borrowing this idea from uh an old idea from andre carpathy but like uh good bottles are just much better at us than designing these types of systems and so you know you can you can kind of imagine how you want to break the problem down and think through the steps that you might take to do it but oftentimes like the the real world is messy and you're that model that you built the model that you built that sort of that workflow that you built might be an oversimplification of the actual process you know a real expert at this task would follow and models are able to learn when models are able to learn the process learn how to do the process by being rewarded for succeeding at the process, they're able to figure out, in many cases, something that's better than you could easily sit down and design yourself.
14:47When you think about applying that idea broadly among developer teams writing software and startups and enterprises, like the ability, as you alluded to earlier, like the ability for open AI to create a model that's designed to do these things that improves these things is very different from, you know, a startup or, you know, enterprise academic institution. Like, I guess part of the question is like, do we just have to wait? Or like, what can we do now to build more robust agents? No, you can go use like real working, actual useful agents now. I mean, deep research, I think, has been incredibly useful for a lot of people.
15:42It's, you know, people are using it for kind of like business and research workflows, but also scientific research, travel and shopping, programming. Sure. And not to take away at all from deep research or codex, but like research and writing code are two of like many, many problems that I have or might want to tackle with agents. Like, how do we get to a more generalized agentic model that is easier for, you know, your users to, you know, build their own models with? And, you know, maybe so that's the general question, like a sub question is like, is that, you know, GPTX or OX or whatever? Is it, are all models going to, you know, get better at agentic capabilities?
16:47is this, you know, error correction, instruction following kind of melee that you just referenced? Or should we expect to see like agent specific types of models? Well, I do think it's very much like as the models get better, building these types of workflows will be a lot easier. Already, I think 03 is quite good compared to older models at being able to kind of understand the level of complexity of instructions that you often have to have to automate a process or build a complicated workflow. And it's better at adhering to those instructions and using tools over multi-step trajectories than anything else that can be for it.
17:37I still think there's a ways to go before that is a generally useful agent. But I think that, yeah, as the models get better, building custom agents is going to get dramatically easier. Are there characteristics of the models? Like, is it always going to be the case that you need a certain size or complexity of model to do agentic things well? Or, you know, can these types of behaviors that we're talking about, like be distilled into smaller models just as easy as kind of the core language fluency that, you know, we've been able to distill into smaller models? I don't think we've really exhausted our limit to of how good we can make small models.
18:27But I think the advantages that large models have are that large models tend to be better at generalization. um so if you're and and you know most of the time if you're building something custom on top of models you know that that thing might not be something that the developers of the model anticipated so um larger models will tend to work a little bit better um and uh i think the other the other feature of models for um agentic use cases that's really valuable is is reasoning because a lot of agentic tasks have this quality where there's a range in difficulty levels of solving the task and doing the right thing at each step of the task is very important to make sure that the overall task succeeds.
19:16And so letting the model choose how much reasoning effort it wants to apply to figure out the current step of the workflow So to me, it feels like a pretty important component of what's going to make these systems work well. And is that particular aspect new in 03 relative to 01? I've noticed with 01, like, you know, it was a solid reasoner, but it reasoned for the same amount of time or on the order of the same amount of time all the time. whereas O3 can give you a snappy-ish response, but also take a bunch of time to think through things. It's not new, but we're continuously getting better at making models that are smarter and knowing how much to think.
20:06So I cut you off when you were talking about deep research and some of the ways that folks were using it. I think probably everyone listening gets the core value proposition thesis behind deep research? Are there things that you find that surprise folks that it can do that, you know, folks might not know? Or are there, you know, examples of how folks have kind of pushed the edges with it? Yeah. I think when people think about deep research, they think about using it for market research and scientific literature review and kind of other tasks where the model just has to go out and think and read a lot of documents to come back with a synthesized answer.
21:01But some of the other things that people are using it for, which were initially surprising to us are coding. So it turns out, even though we didn't really design the model for this use case, there's a lot of, you know, a lot of code available, like publicly available on the internet. And because of the capabilities in the base model, the model is able to understand it very well. And it's able to search and kind of like find code on GitHub and understand how different pieces of the code base work and come back to you with descriptions of the code or plans for how to implement a feature or things like that.
21:36Um, and then I think the other kind of underrated use of deep research is for kind of finding very rare facts on the internet. Um, so I think people like probably the bulk of what I think most people use deep research for is going broad and then synthesizing, like, you know, find, find a, uh, go read about this topic and come back to me with a summary, like covering these points where the model has to like go touch a bunch of different topics. But the model is also quite good at finding information that is just kind of buried in the corner of the internet somewhere. Like if you have a question about, you know, if you really want to know, to recall like the details of some television episode that you saw a long time ago where, you know, maybe it's something specific enough that it wouldn't just be on the Wikipedia page.
22:25Often the model will be able to find kind of that like one fan page or something like that from 10 years ago and point to that information too. Yeah, even 4O is really good at that. I'm trying to imagine how obscure you have to get in order to go to deep research for that kind of problem. And remind me, deep research is not accessible via an API currently. It's only via the, you know, one of the surfaces. Okay. Is that anticipated? We found that mostly the kind of use case that we were imagining for deep research was kind of you're in ChatGPT for something that you normally use ChatGPT for. And then there's something where you just really want a much more thorough answer than you get out of the box with ChatGPT.
23:19That's kind of the use case that we designed it for. But you're right. I mean, there's a lot of other things that you could use it for where having an API would be useful. We talked just a second ago about kind of the spectrum of compute that you might want to throw at a response and how, you know, O3, for example, can, you know, kind of start to manage that. Like, at some point, I might want the model to just know if my question calls for deep research and just do it. Right. Yeah, that would that would be the dream. Right. I think the thing that I would love ChatGPT to become is a place where you can just go and it's like talking to your friend and your coworker and your personal assistant and your coach all in one.
24:15to where you can just, you can ask a question like you would to a person. And just like a great coworker would, it knows, you know, when to come back to you with a really quick answer off the top of its head versus like from the context and from knowing you when it should actually go and do a bunch of research and come back to you with something more thorough. Or both, like, here's my initial impression, but I can go research that for you. Yeah, exactly, right. It's like, yeah, here's the answer off the top of my head, but like, I'll go dig into this a little bit more. Or it could, you know, come back and like in deep research, there's this phase in the beginning where it'll ask you a few follow-up questions.
24:52But I think eventually it should be able to just come back to you in the middle of its research and say like, hey, here's what I'm finding so far. I'm going to keep going, but is there any other, is there any feedback? Do you have any feedback on this? Is this what you're imagining? Or are there any, you know, now that you've seen this, are there any other avenues that you want me to explore? So I think like evolving ChachiBT into an entity that feels a lot more natural to co-work with is one of the main things that I hope we're able to achieve. You mentioned the follow-up questions that happened at that initial phase of deep research.
25:31Can you talk a little bit about how explicit was like building that or tuning the model to do that? Yeah, what we found is that deep research is, since it's designed to go deep and collect a lot of information and come back to you with a very detailed report that is really excellent at covering all the finer points or the nuances of the sort of initial framing of the question that you wrote. And the quality of the results that you get are very sensitive to, they're not very sensitive, but you get better results if you put time up front into thinking about what you really want to see at the end.
26:16Because the model is very good at adhering to those things. And it's very good at knowing how to incorporate the different parts of the question that you asked into its search. And so the model will also come back with reasonable answers if you kind of just fire off a quick question to it. But the upfront sort of back and forth is meant to help flesh out a more detailed thing to run by the model because we think it produces more compelling results. Okay. And was that an emergent, you know, phenomenon? Like you train the model with an angle of like producing good results and it figured out that it needed to ask these follow up questions or was there a degree of direction in there that?
27:04I don't know how you might do that. I mean, yeah, any number of ways you might try to do that. Yeah, the upfront questions were something that we added kind of after the fact, as like from observing the way that users interacted with the model. And, you know, I know from experience, you can suggest that you're not sure how to frame the question and that you might want some, you know, you can prompt it to, you know, ask you questions. And that works just fine as well. Yeah. And one thing that we've seen a lot of people do is to use 01 or now 03 or 04 mini to help craft the question. So if there's, you know, if there's a topic where you don't even really know enough to frame a good question, a lot of times folks will have like a conversation with a smaller model.
27:50And then they'll use that to sort of flesh out all the details of the question that they want to research. And they'll put that into deep research and have a go, you know, spend 10 or 20 minutes coming up with a much more research response. So deep research was the first agentic offering. We also have operator. You consider that an agentic system? I do. Yeah, I think of an agentic system as any AI system that is able to go work on tasks for you that take longer than a few seconds. And that has to interact with the real world in order to solve. your problem. So yeah, Operator, Deep Research, and Codex CLI, I think, are all good examples of agentic systems.
28:39Got it. So let's talk a little bit about Operator. I tend to think of it primarily, I've not used Operator, actually. Like, I've thought of it as like a computer use analog, but it sounds like it's a higher level system that you interact with and have it do. I mean, I've seen demos and that's the way it was presented, but I thought that those were demos of an underlying API as opposed to a service that I might want to use. Yeah, no, Operator is also a service inside of ChatGPT. And so you can go to Operator. There's a link to it from ChatGPT, but it's a separate kind of web app that works a lot like ChatGPT.
29:20And so you can type in a request of something that you want to do, book a restaurant reservation, let's say. And then you can kind of watch in a virtual browser. The agent kind of navigate to the web page, kind of click around the web page and do the task that you asked it to do. So it's really, really cool. It's like super fun to watch because it's just, you can kind of feel the intelligence and the thinking as you watch the model, like click through the web much like you would. Are you finding that there are folks that are using it productively versus education and entertainment? Yeah, absolutely.
29:59So Operator, I would say, is like the technology to make Operator work really well is an incredibly difficult thing to build. And so Operator, you know, is not intended to be the thing that, you know, every single person in the world uses every day on day one. It was meant to be kind of like an early launch of this technology so that people could get value from it and can also just get a feel for where this is going. But despite that, there are a bunch of people who do use this a lot for, you know, who like power users. I think of it a lot like kind of like early GPT-3 API, if you remember that, right?
30:40Where it's like people use it and it was, you know, I don't think OpenAI would have framed it this way at the time, but in a lot of ways it was kind of a technology preview. Like it was something where most people tried it and they were like, oh, this is really amazing. It's amazing that humanity has created this thing, but, you know, they didn't immediately find it useful. But even on day one with GPT-3, there were a bunch of people where, you know, they were just really drawn to the technology and they tried it, you know, they played around with it a lot and they got to the point where they figured out how to make it really useful for them.
31:15I think operator is kind of more at that stage of model development where, like, there are power users who love it. But it's not for everyone yet. Yeah, that is my impression as well. It's not surprising to hear you put it like that. I've seen some examples of like, I think the theme that I've seen the most is like arbitrage. Like I'm going to, you know, find a bunch of, have this operator find a bunch of stuff on eBay that I can sell for more on Craigslist. Or like find Airbnbs and like message all the hosts to try to get a discount or something like that. like I looked at them and I was like, they were novelties, but, you know, maybe someone is actually doing that for or something along those lines for, you know, something that they think is useful.
Read the full transcript
32:02Yeah, you get a lot of value from it if you if you sort of put in the work to figure out how it's useful for you. Yeah. And then beyond figuring out how it might be useful for you, are there kind of non-obvious ways to make it useful? Meaning I've got my use case, like, Like, how do I need to think about the world or my problem or operator in order to, you know, make it work? Yeah. So one thing that I found to be really helpful for getting the most value from it is there's a way to add sort of site specific instructions, add or customize site specific instructions. and so you know if you've if you're imagining operator visiting a website for the first time it's kind of like you're you know you're seeing this website for the first time without any context or the very limited context of the world or who the user is and so there's a lot there's like a lot to figure out but you can provide instructions that help the model understand how to use this site to solve their problem.
33:03And that tends to make it like a lot more repeatable and a lot faster. So like, don't just say, click upload in the menu, say, you know, the menu is the hamburger thing on the right and upload. Like how granular do you need to, first of all, are we talking about, you know, location, localizing capabilities on, you know, the page? Or are we talking about other types of context that is useful for the agent? And then how granular does that tend to need to be? Yeah, I think it depends a little bit on the site. And, you know, sometimes it takes a little bit of tinkering to get exactly right. But I find mostly just like helping to clarify the intent.
33:43Like, oh, if I want to, you know, if I want to book a flight, here's how you do that, that kind of thing. I guess there's Codex is the third agent in the portfolio. and with Codex, the Codex CLI. Codex has been around for quite a while. I think I, when did I talk to Greg about the launch of that? That was August 2021 when we caught up about that. This is a different Codex. We're using the name, but it's not intended to be. There's no underlying model. I think the original Codex was like the first kind of code completion model. Yeah. this is like a throwback to that name but a very different model and yeah Codex is our local code execution agent and so it's just a package that's like fully open source that you can install on your laptop and then you can talk to it and you can ask it questions about your code you can or you can ask it to implement things in your code base.
34:54And this agent is able to use your computer to navigate the file system, find the relevant files, understand how they work, write code on its own, apply those patches to your files, run tests, try things out, and come back to you when it has an answer to your question or it's written some unit tests or it's written some code or whatever you really want it to do. So I think it's a really powerful demonstration of where I think AI systems for coding are going. And people are loving it so far. People at OpenAI use it a ton. There's like more than 20 ,000 stars on GitHub, which is amazing to me. we've got like around 100 people have contributed to it and so I think like a lot of people are kind of feeling the magic of it In a world where you've got Cursor and GitHub Copilot and WindSurf and you have maybe hundreds more of coding agents slash IDEs with coding plugins like what differentiates Codex CLI, Cloud Code type of experience from those others?
36:18Yeah, I would say most people, most software engineers here use Codex CLI and potentially like some AI powered functionality in their IDE. They're kind of used for slightly different things, I think. The IDE functionality is great because when you're inside of the IDE, well, when you're writing code, you're most of the time inside of the IDE anyway. And so those tools are amazing for just having like very quick access to the AI system, like in your flow state. So if you're kind of like, if you're in the middle of writing some code, you can get a suggestion or you can, you know, delegate something to a model and then just keep going with what you're doing.
37:02these uh these agentic tools um are better for like um or at least i find myself and a lot of folks find themselves using them more for like de novo type work um where you you know maybe you're just at the start of the project um or you're working in a code base for the first time or you have an idea for a feature that you want to build that is um not like literally the thing that you're working on at the moment. But it's something that, you know, you might want to delegate to someone else. That's when like this agentic paradigm becomes really powerful. Because, you know, today, the way the Codex CLI works is, you can, you know, you give it a task to do.
37:49And it just does it on your laptop. One of the really cool things about it is that it's able to do it in a network sandboxed way, meaning that it can just go and kind of work for you in a way that is safe. Because it's not able, you know, it's not hitting the network. And so it's not able to, you know, accidentally, you know, run some unsafe command or something like that. But I think over time, but it's still kind of like running on your laptop. And so you still need to have your laptop open and sort of watch it as it goes. I think over time, like where this is going is that interacting with these coding agents should feel more and more like delegating to someone where they're able to they're able to take larger and larger chunks of work and execute them more and more autonomously to where you can just, you know, if you have like 10 things on your to do list in the morning, maybe you go work on one and then you come back at the end of the day and the other nine have been finished for you.
38:47So that's kind of like, I think, the paradigm of agent decoding. And, you know, we're not there yet, but where we're trying to get to. When I think about the idea of it being best for de novo types of projects, I think of, OK, you know, we've got these three classes now. We've got the cursors and the like. We've got your V0 bolts, which I think of as like de novo web. And now the codex CLIs of the world are maybe like de novo backend or something like that. I guess when I say de novo, I don't necessarily mean the whole project is de novo. Actually, one of the amazing things. Or it could be a feature.
39:22Yeah, a feature. I guess is what you're saying. Yeah, just something, you know, where it's distinct enough from what you're, like, literally the code that you're typing right now that it's worth kind of like starting a new thread for yourself to, you know, to start working on this independently. With that in mind, I was going to raise a question around context. One of the really interesting things that I've seen recently that helped me understand things that I've experienced is this idea that when you're interacting with AI-assisted code generation tools and IDEs, It's like one of the things that they're, you know, limited by and doing is like managing the context of, you know, that is ultimately passed to the LLM.
40:17And that's why you get, you know, different results between like just typing in a prompt and cursor, for example, and, you know, versus copying your whole file and, you know, putting it into chat GPT or whatever, you know, chat model you use. And, yeah, it's been suggested by some that, like, one of the things that these kind of CLI-based agents, you know, do better somehow is context management. Like, is that true? How do you think about context in this context and the role that it plays in development tools? Yeah. Yeah. I think that like the Codex CLI is, you can almost think of it as like a contextless tool where my mental model for giving a task to Codex CLI is, it's kind of like you're giving a task to a, like a superhuman intern who has never seen your code base before.
41:24That's only a little oxymoronic. Yeah. An intern because, you know, you can't really trust it yet with like really, really complicated stuff that, you know, where you need a lot of state or a lot of, you know, experience working on this particular type of project where there's, you know, a lot of different pieces that need to be broken down. I think you kind of want to delegate like intern-sized chunks of work to it. Superhuman because it's able to read code, write code, understand code, understand coding patterns much faster than any human can. And seeing your code base for the first time because it is contextless.
42:10And so when you start a new task in the Codex CLI, the model has to explore the code base on its own before it can start working on the task. And the way it does it is just, it's pretty amazing. It's not using any kind of bespoke context management tools or ideas. It's just using the same command line tools that, you know, people have been using for decades. Yeah, like you'll see it said files. Exactly. Yeah. To explore the file system. And it turns out, which is, I think, really surprising in some ways that the model is just able to build an understanding of how to navigate the code base and how to find the information that it needs and where to write code and what the patterns in the code base are extremely quickly just by doing that.
42:58And coming back to what we were talking about before of like the magic of training models using reinforcement learning to achieve an outcome, I think this is kind of like, the type of magic that emerges when you do that type of training is models, um, are able to learn how to use tools really effectively to achieve their goals. Um, and so I, you know, I, my impression is that these models are much more efficient than humans are at like, um, how much code they need to read before they're able to actually build something. Um, so that's, uh, that, that's, that's kind of how I think about context in, in, uh, for, for these models.
43:38But, um, So that being said, I do think that we could improve on this by giving the models more context or having richer customization. So right now there's a way to customize the behavior of the models by sort of placing a file in the repository. But you can imagine all kinds of other types of customization too. Like you can imagine giving APIs or MCPs to the model. You can imagine giving the model memory so it can remember things between rollouts. These are lots of ideas that we're kind of exploring for how to make the models kind of even better at understanding the specifics of your code base and get smarter as they interact with it more.
44:23And how do you think of Codex, you know, particularly with it being open source, like in the spectrum of, you know, research preview here to get you thinking, you know, to, you know, first steps to something that reflects what, you know, we ultimately want to deliver? Yeah, I think the open source piece is like really critical to how we're thinking about building Codex. um so we're we're very much like uh trying to do this the right way as an open source project um we're kind of building with the community um you know we're accepting tons of contributions from folks i think there's like yeah around 100 contributors now um and so we are like uh we we have sort of opinions and a vision for where we want to go with it but we're also um letting the community kind of help guide the direction that we take uh take with it um but in terms of like readiness, like Codex CLI is really useful now.
45:21Yeah, people at OpenAI use it all the time. I think a lot of folks, kind of a lot of the early adopters have found a lot of success with it. I think especially with O3, it's able to just kind of like solve a lot of sort of, yeah, smaller, medium pieces of coding work that might have taken you a lot of time to do otherwise, or might have just been really hard because you don't understand that part of the code base or, you know, it's some front end work and you're a backend engineer or vice versa. So yeah, people, people who are using it are getting a lot of value from it. One of the complaints, if you will, is about like cost and cost opacity with using these tools.
46:01You know, everyone's miles going to vary in terms of value, but like, how do you think about that? How do you see, is that an area or is that something to you anticipate evolving over time? Yeah, I think if you look at the history of models, the cost has been falling pretty dramatically. And also, I think we're in the early days of the Codex CLI is brand new because the models have only recently gotten good enough for this to be a really good experience. we're also in the early days of model capability um as the models get more as the models get better and more consistent at producing high quality results the cost is not going to matter as much because i think already in many cases um it's uh it's saving people a lot of time um and so you know uh like the you know paying a few dollars to save you hours is you know depending on what you're doing is often a very easy trade to make.
47:12But yeah, I think also like if history is any indication, you know, I think it would be reasonable to predict that the cost will also come down. Yeah. Yeah. I was trying to get a sense for whether the predominant like trend or answer to that question is, you know, the cost is just going to go down naturally like we've seen or yeah, we hear that and, you know, we're going to be creating more visibility like, oh, that request is going to cost you 42 cents. Oh, yeah, that's a cool idea, too. Yeah. Have you seen integrations with Codex CLI and other systems like, you know, Slack-like or Devon-like Slack interface or something like that?
47:58Like, does the open source aspect lend itself to folks doing kind of interesting things and connecting it to other systems? Yeah, I mean, that's one of the things that we really hope is going to happen. And that's part of why we decided to open source it is, you know, I think that there's a lot of places where you can imagine just programmatically triggering a coding task to start happening. So, you know, one of the obvious ones would be in CI. Like if you have a build that's failing or, you know, some tests that aren't passing for whatever reason. We've seen people starting to build Codex CLI into their CI workflows to try to, you know, take a first pass automatically fixing those issues.
48:48But there's all kinds of other things where it's like, oh, you have a... Or even an issue posted, like, give it the issue test and, you know, what might be causing this. Yeah, where you have like some kind of like feed of potential coding problems that are coming through. And, you know, and kind of like with deep research as well, like the, I think that primarily what people use these tools for initially is to do work that they just otherwise would not have done. like if you're you know if you're kind of a busy software engineer and you have 10 things on your to-do list maybe you're only going to do two or three of them and that's not that the other seven aren't valuable it's just that they're kind of like below the cut line of what you can really realistically prioritize and so I think that's that's one of that's the thing that I'm most optimistic about using these agentic tools for in the near term is it's just kind of like doing all the stuff that you want to do and that would be good to do, but you just can't, you can't do as a, as a person with limited time.
49:52Yeah. The intern analogy I think is the most popular, but I always like the sous, the sous chef analogy. Like, you know, there's this thing that I'm really good at. I have my, my recipes, whatever. Um, but like, don't make me cut the carrots. Like just give me the chopped carrots and, you know, let me do my thing. That is another place where we've seen people use codex cli a lot is like yeah doing the kinds of um engineering work that you just don't like doing um again yeah oftentimes i see this with like you know back-end engineers who don't like writing javascript or you know very like product oriented engineers who don't like doing a bunch of data plumbing stuff um are very happy to just you know sometimes you have to do that stuff for your work and so uh makes people happy to be able to have the model take a first pass at Yeah.
50:40What do you think are the biggest things on the horizon that are going to change the capability of these coding models? I think folks who have used them, maybe let me ask this question, kind of the same question, but let me ask in a very different way. What's your take on the whole vibe coding thing? I think that we are in the early phases of a dramatic shift to the way that software is built.
51:17Where, you know, look, I don't think writing code is ever going to go away. You know, but I think that the process of manually writing code, or you have to write down kind of every single line of code yourself and manually test it and think through the architecture and kind of do each of these pieces by hand, it's going to become rarer and rarer. I think that most code, I think the vast majority of code is going to be written by AI systems much sooner than people think. I think the job of developing software is going to become a lot more like thinking about the functionality that you want the system to have, thinking about the trade-offs, thinking about the edge cases, guiding the AI system, providing feedback to it, finding ways to validate its work.
52:19So I think that just the nature of the work is going to change and people who write code are going to become much more productive. Some of those tasks that you outlined sound like engineering tasks, but some of them sound like product tasks or product manager tasks. Do you think that this new world precipitates some kind of shift in the relationship between those roles? I think that it is part of a shift that's already happening. where like I think the as
52:58as programming has moved up in abstraction over time the ability of one person to kind of bridge the entire like span of programming has gotten bigger and so we're already seeing the proliferation especially in the startup world of like design engineers like folks who are good great designers and can also build the things they design or like technical product managers or you know full stack engineers engineers who can like write front end code and back end code so I think this kind of accelerates that where it's it's going to be less of your engineering skill set, I think, is still really valuable, but less of your mental energy will be devoted to kind of like learning the nuances or the ins and outs of a particular framework or language.
53:58And more of your mental energy will be devoted to thinking about how things work, what they do, why we do things a certain way, how do we know it's working well. So that gives people space to think about a bigger chunk of the problem than they did before. Do you see that the way people acquire those new era skills is different than how they've acquired them in the past? School and on-the-job experience? That's a good question. I think that by far the best way to learn any subject now is using ChatGPT. and O3 in deep research. It's amazing because it's like you have access to effectively a world expert in many domains that can spend infinite time with you and can get to know you really well.
54:54And there's no question that's too basic to ask it. So I see how people learn these skills changing dramatically regardless of this phenomenon. I think it'd be much easier to learn programming now with ChatGPT than it was to learn programming when you didn't have access to a tool like this. Though we're pointing to learning programming isn't necessarily learning programming, it's learning architecture and patterns and all these other things. I don't think that changes your point that JetTPT can be a great resource for learning these things. But there's also this element of like what it is to learn.
55:37This is like a controversial opinion, but I still think it's important for people to learn programming, even if it's a smaller percentage of what they're going to spend their time doing if you want to build software. Because it's kind of like, I don't know, like in grad school for machine learning, there's like kind of, you know, some of these rites of passage type activities. like building a computer from scratch or building a back propagation library from scratch, from first principles that like, you know, okay, is this the single thing that you can do that will make you, you know, that will like maximize your ability to perform your machine learning research or your machine learning job the best next week?
56:20No, like you don't really, 99 % of the time when you're doing machine learning, you don't need to understand the way that backpropagation works at a very fundamental level. However, like when things go wrong, being able to kind of spelunk all the way down the stack and think through like, okay, well, this isn't working. Why isn't it working? Having like fundamental understanding of pieces lower in the stack is really important to being able to do that efficiently. And so, you know, I think like a lot of people who have gone through that exercise will tell you like, oh yeah, every once in a while, actually, I do have to kind of like think about this.
56:59And I do realize like this is where the bug came from is something about the way that backpropagation is working. And I think that, you know, the same will be true for code for a long time, where like if you aspire to be really good at this, there will be a lot of cases where, you know, being able to look at the code that the model wrote, even though you'll do it less and less, is still very much worth their time. And the people who are able to do that well will be better at this like more vibe coding style of coding than the people who can't. Makes me think of me spending years in grad school learning how to model MMKQs and never really using that.
57:37But I am great at picking lines in grocery stores and airports. Yeah, these weird skills that we pick up along the way, right? We've talked a little bit about, or at least you mentioned it, MCP and tools and the role of tools and, you know, working with and building agents. Like, what's your take on that landscape? MCP is, you know, maybe six months old and like really no one paid attention to it for, you know, three or four of those six months. and now like everyone's building an MCP server and like it's become a really hot topic. Yeah, it seems to be really taking off. Yeah, I mean, I think exposing tools to model, like, you know, if you think about the formula for useful agents, in my mind, it's like you have to have a really smart sort of general purpose or reasoning model.
58:33And then you have to expose it to the tools it needs to do the thing it needs to do. And then hopefully you should provide It's a relatively small amount of high quality task specific training to teach the model how to do the thing that you want it to do.
58:55And so being able to expose the tools to the model in a way that it's able to use them both for training and also for doing the task in the real world is like a really critical piece of how the ecosystem fits together. And so I think, yeah, I mean, everyone is thinking about this now for good reason. And are there obvious gaps to you that, you know, MCP or some other protocol or even just best practice needs to fill in order to kind of bridge this gap between tools and agents? One thing that feels not fully developed yet is the ability to specify a level of trust for the agent using certain tools.
59:53um so for example like the you know i think um the future that everyone is picturing is they'll just ask chat gbt to go book a vacation and you know it'll know you really well and it'll go do a bunch of research and it'll come back and it'll just charge your credit card and it'll like send you the itinerary you know hopefully asking you a bunch of questions and stuff along the way um if it needs to but um the the act of like doing something like using your credit card is a very high trust act. And, you know, just kind of exposing your credit card to an agent that you don't trust or exposing your credit card to an agent without specifying how you wanted to use that and when you wanted to use it is a very scary idea and potentially super risky.
1:00:42So one thing I feel like the field doesn't have a great handle on yet is how do we how do we kind of like build in that kind of like ways of specifying and enforcing certain levels of trust between the human, the agent, the tool and the task so that you can kind of like go just ask the agent to do something and it knows kind of when it needs to come asking for permission to use certain tools based on your interactions with it so far? Do you have a sense for what that looks like? And I'm asking because, like, you know, we want to, I think, build these models, you know, based on, you know, data and training and not, you know, heuristic and, you know, systems and, like, you know, the Carpathian analogy you used earlier.
1:01:48but it seems like you know we've seen the research that says that you know the models will kind of misrepresent the you know their thinking and thought traces and like disregard instructions like it seems like there might need to be some extra verbal way of enforcing these kinds of constraints yeah I think at a fundamental level what it comes down to is having a set of guidelines that you can provide to the agent about when it needs to come get your permission to use certain tools.
1:02:29And so, you know, the most strict version of that would be like, so it's providing the guidelines and also having some mechanism to enforce the guidelines or like some level of trust that the guidelines will be followed. Um, so like the, the, the simplest example of this would be like, if you just, if you have a guideline that says, um, you have access to a credit card tool, but every time you call the credit card tool, you need to come ask me to approve.
1:03:04then you don't need to have a lot of trust in the model to use that tool because you always will have the opportunity to review what it did. Now, there's still problems of deception because the model might try to represent what it did or what its goals of using the tool are. But at least you have the opportunity to say no. So you still have to have a high degree of trust in the agent, even if, you know, you specify, you know, some set of parameters. And I think it's an interesting question whether, you know, the current, you know, model of, you know, the current crop of models or model architecture or the way we train models or the state of alignment, like, warrants that trust.
1:03:52And it depends on the task. It depends on the model. And so I think building that trust to do high-risk actions will need to be something that model developers think really carefully about, product developers think really carefully about.
1:04:18and users will need to build that trust. Even after all that, users will need to build that trust iteratively through interacting with the model. And so, you know, I think there's, I think there's still a lot of thought that we need to do as an industry to figure out how to make that work well. It's funny that you use that example because like maybe an hour ago, Harrison Chase posted about, CEO of LineChain posted, about some news with Visa. They are partnering with Visa on Visa Intelligent Commerce, which is enabling AI agents to shop for you. And, you know, in their little, this I think is a mock-up, but like, there's a slider and you can say what your max spend is.
1:05:08But my response to that tweet was, is Visa going to allow me to like do a chargeback when I tell the agent that I want X and Y's Y. Yeah, I think, you know, that's both kind of pointing out the timeliness of this conversation, but also the fact that there are like technical and political solutions to the problem. Like, actually, a visa is going to be like, you know, take the role of providing some degree of trust by allowing me to, you know, beat them up when the agent does stupid things. That could be interesting and useful as well. I don't think they will, but. Yeah. Yeah, I mean, there will be residual risk there.
1:05:55And so I do think there is an interesting question of like, to what degree do these different, to different parties, should different parties bear that residual risk? Like, is it a risk that I just take on as a user? Is it something where like the model provider is sort of saying that, you know, it's okay? maybe the payment provider. It'll be certainly interesting to see how it plays out. Well, Josh, it has been wonderful catching up and chatting about all things Genetic in your world. Yeah, great to catch up. Appreciate you taking the time. Thanks. Thanks so much.
1:06:39Thank you.
From the publisher
Today, we're joined by Josh Tobin, member of technical staff at OpenAI, to discuss the company’s approach to building AI agents. We cover OpenAI's three agentic offerings—Deep Research for comprehensive web research, Operator for website navigation, and Codex CLI for local code execution. We explore OpenAI’s shift from simple LLM workflows to reasoning models specifically trained for multi-step tasks through reinforcement learning, and how that enables agents to more easily recover from failures while executing complex processes. Josh shares insights on the practical applications of these agents, including some unexpected use cases. We also discuss the future of human-AI collaboration in software development, such as with "vibe coding," the integration of tools through the Model Control Protocol (MCP), and the significance of context management in AI-enabled IDEs. Additionally, we highlight the challenges of ensuring trust and safety as AI agents become more powerful and autonomous.
The complete show notes for this episode can be found at https://twimlai.com/go/730.




