OpenAI Just Released ChatGPT Agent, Its Most Powerful Agent Yet

22 Jul 2025 · 38 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Episode Notes: OpenAI Just Released ChatGPT Agent, Its Most Powerful Agent Yet

Overview Podcast Title: Training Data Episode Title: OpenAI Just Released ChatGPT Agent, Its Most Powerful Agent Yet Hosts: Sonya Huang and Lauren Reeder, Sequoia Capital Guests: Isa Fulford, Casey Chu, Edward Sun (OpenAI’s ChatGPT Agent team)

Episode Description The episode focuses on the new ChatGPT Agent developed by OpenAI, which can perform complex, multi-step tasks for extended periods (up to an hour). The team discusses the integration of various AI capabilities and its implications for the future of AI interaction.

---

Key Concepts ChatGPT Agent Features

  • Unified Architecture: Combines functionalities of Deep Research and Operator, allowing seamless integration of tools.
  • Multi-Step Tasks: Capable of performing tasks that typically require human effort, thanks to access to a virtual computer.
  • Shared State: Tools can share state information, improving efficiency and flexibility during task execution.
  • Versatile Tool Access:
  • Text Browsing: Efficiently accesses and searches information.
  • Visual Browsing: Handles interactive web elements like forms and graphical content.
  • Terminal Access: Runs code and interacts with APIs for data manipulation and artifact creation.

Training and Development

  • Reinforcement Learning Approach: Utilizes reinforcement learning to allow the agent to discover optimal task strategies independently.
  • Safety Mitigations: Implementing safety protocols to monitor and regulate agent actions to prevent harmful behaviors.
  • Team Collaboration: Close collaboration between research and applied teams has accelerated the development of the agent.

---

Discussions and Highlights Evolution of AI Agents

  • The episode discusses the transition from traditional AI models to more agent-like systems that can interact and perform tasks over longer durations.
  • Emphasis on the need for personalization, memory, and proactive task management in future AI agents.

Use Cases and Applications

  • Practical Applications: Participants share various applications they've tried with the ChatGPT Agent, including:
  • Data analysis and report generation in fields like ancient DNA research.
  • Online shopping assistance, event planning, and financial modeling.

Challenges and Limitations

  • Safety Concerns: Discusses the importance of safety measures, especially when agents can perform real-world actions that may lead to harmful outcomes.
  • Complexity of Task Execution: Some tasks, such as date picking, are still challenging for AI systems, illustrating the current limitations of technology.
  • Operational Stability: Training the agent involves managing numerous virtual machines and ensuring stability in task performance.

Future Directions

  • The team envisions ongoing enhancements to the agent's capabilities, including:
  • Improved accuracy and performance in diverse tasks.
  • Exploration of new interaction paradigms with users.
  • Further development of personalization features to enhance user experience.

---

Key Takeaways

  • Multi-Turn Conversations: The new agent excels in maintaining context over extended interactions, enhancing its usability.
  • Emerging Use Cases: Open-ended design allows users to explore diverse applications not initially anticipated by the developers.
  • Transformative Potential: The ChatGPT Agent represents a significant step forward in creating embodied AI assistants capable of performing complex tasks collaboratively with users.

---

Conclusion The episode provides insights into the advanced capabilities of the ChatGPT Agent, highlighting how it merges various AI functionalities to create a more interactive and productive user experience. The discussions also emphasize the team's commitment to safety and continuous improvement as AI technology evolves.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00I think this model is actually very good at multi -term conversations and it's very nice to continue working on a task with. I think that's one of the deficiencies of deep research. A lot of people will do multiple deep research requests in a single conversation but it doesn't always work so well. So I think we're really happy with this model's multi -term ability and we just want to improve even further. And then I also think personalization and memory for agents will also be very important. And right now every agent task is initiated by the user, but in future it should also be doing things for you without you having to even ask in the first place.

0:53Today we're exploring the evolution of AI agents with ESA Fulford, KC2, and Edward Sun, the open AI team behind the new chat GPT agent. You'll learn how they got to a huge link forward in capability by unifying the architecture across deep research and operator, allowing for multiple tools to share state, giving users fluid transitions between visual browsing, text analysis, and code execution all within a single environment. We'll discuss their training approach. Rather than programming specific tool usage patterns, they let the models discover the optimal strategies through reinforcement learning across thousands of virtual machines.

1:29They've created an agent that can work alongside you for hours, asking clarifying questions and accepting mid -task corrections, expanding the ways that we can interact with AI agents. The team shares fascinating challenges around safety, guard tolerance around agent activities, and why things like date picking still remain mysteriously difficult for AI systems. They've revealed how a small -focused teams are achieving breakthrough capabilities through careful data curation, suggesting that we're now renting a new phase of AI development where product insights matter just as much as compute power.

2:01Enjoy the show. ESA, Casey, Edward, thank you for joining us today. Thank you so much for having us. So you're the team behind the ChatGibberish Agent or Agent Mode. What is it? Yeah, so this has been a collaboration between the former deep research and operator teams. We've created a new agent in ChatGibberish that's able to carry out tasks that would take humans a long time. and we gave the agent access to a virtual computer. And through that, it has a few different, two different ways to access the internet, or actually more ways, but we'll get to that. It has a text browser, which is similar to deep research tools.

2:41So it's able to efficiently access information online and search through things with this very fast text browsing tool. And then it also has a virtual browser, which is similar to the operated tool. So it actually has full access to the graphical user interface. And it's able to click and type things into forms and scroll and drag and all these kinds of things. So together it's much more powerful than either of those two tools because one's more efficient and one's like much more flexible. And then we also gave access to terminal. So it's able to run code and analyze files and create artifacts where you like spreadsheets or slides.

3:24We also, through the terminal, it's able to call APIs, so either public APIs or private APIs. If you sign in, it could access your GitHub or Google Drive, SharePoint, many other things. And the cool thing about this tool is all of the tools have Shad State, so it's similar to if you're using a computer, like all of your different applications have access to the same file system and things like that. that it's the same for the tool. So the model can do quite flexible things. And yeah, we'll talk more about this later, but I think it's just a very flexible way for the model to do very complex tasks on behalf of users.

4:03Tell us a little about the origin story. How did this get started? Well, our team worked on operator. And our team worked on deep research. And so back in January, we released our first agent operator. This is a product that can do internet tasks for you, like buy things on the internet shop for you, this kind of thing, and then two weeks later. We really steep research, which is a different model that's, or a different product that's able to extensively browse the internet and synthesize information, and it creates a long research report with citations for you. And we were kind of thinking through our roadmap, and we were kind of like, hey, this is kind of a matchmaking made in heaven here.

4:43So, you know, operator is really good at visual, you know, interacting with a web page, but it's less good at kind of the text browser, like reading long articles. Whereas deep research is really good at reading long articles, but it has a tougher time with like interactive elements or like highly visual things. Because the tools are different, so deep research has a text browser, so it's able to really efficiently read information and search and, since size information, but it's not able to like scroll and click in the same way or fill out forms in the same way that operator is because it has actually full access to the GUI browser.

5:19And as Casey was mentioning, like deep research has some things that operator doesn't have. And then similarly, one of the biggest requests for deep researchers further model to be able to access like pay world sources or things that you have to pay a subscription for and operator is able to do that. And also one of our members of our team Eric, he was running an analysis on the types of prompts that people were trying on operator. And we realized that it was a lot of deep research type tasks, like research this trip for me, then book it. So it really is a natural combination. In what ways, 1 plus 1 equals 3?

5:54So in deep research, we always wanted to figure out how to let deep research have access to a real browser that can loading all the real content that previous deep research cannot have access to. That's funny that you bring up the 1 plus 1 equals 3. because not only did we combine deep research and operator, but we also threw in a bunch of other tools that basically everything we can come up. So the terminal tool is there, so it can run commands to do calculations. The image -gen tool is a fun one. If it wants to spruce up its slides by making an image, it can do that. Can cool APIs. Can produce PowerPoints?

6:36Yes, I can do a lot of different things. Tell us a little bit how people are using it. Knowing it still early days. So I think the cool thing about it is we have some ideas of how we think people are going to use it. But I think we intentionally kept it quite open -ended. I mean, it's called agent, that's so vague, partially because we are excited to see how people end up using it. So I think some of the things that we specifically trained it for were, of course, deep research type tasks and things where you want a long report on a topic, operator type task where you want it to do something for you, like book something or book a flight, buy something for you, and then also task to make slide decks.

7:15We also, you know, spent a lot of effort on making spreadsheets and doing data analysis, but I think there are also just so many other things the model can do, so we're just excited to see how people use it. Kind of similarly to how when we launched deep research, we saw a lot of people using it for code search, which was really surprising to us. We're hoping to see a lot of new use cases that we didn't even think of ourselves. Would you guess it'd be more consumer or kind of be the beta type use cases? Or is that a false security? Hopefully both. Okay. I think we're kind of aiming for like the pro -sumer.

7:46Like someone who's willing to wait like 30 minutes for like a detailed report, but that can be in the like yeah, in the consumer case or the you know at your job. Yeah. I think it could be both good for both. Do any of you favorite things you've used it for? For me, it's more like you know pulling data from you know our like spreadsheet or Google Docs like a document in our like a expand -how log and then make some slides to present the data or like organize the data. It's pretty useful. I've been doing a deep dive into ancient DNA. What am I, you know? What am I interested? And there's actually a lot of exciting work going on like these past like five years.

8:23They're like sequencing all this DNA and like discovering all these facts about like, oh where did like this group of people come from and like, you know, historical stuff. The problem is that everything is so new that there isn't a reference source material for to summarize a survey of these materials. But agent can go out and pull together all these sources and synthesize it into a report that I can read or slides that I can read. And I think it's kind of made for this topic. Yeah, I like it for consuming use cases. I've used it for online shopping. I think especially because a lot of websites require using visual browser because it will have a search filter or something that it needs to go through or the model.

9:07Like, actually needs to be able to see what the item looks like. And then also for planning events, it's been pretty useful. What's your favorite shopping query? I think I was using it for clothes shopping.

9:24Love it. And you guys also showed us a really cool use case right before we filmed this episode. didn't you show that one? Yeah, sure. So that was actually something that one of our coworkers Tageel shared with us. She asked the agent to estimate, open -eyes, valuation, and create based on things that it found online, creates a financial model with projections. So create a spreadsheet, also create a summary analysis, and then also create a slide deck presenting the results. And so hopefully the model is correct, because it had quite an ambitious projection for us. It was an impressive slide deck.

10:01One thing I want to point out about this trajectory was that it reasoned for I think 28 minutes. And yeah, I think this is kind of opening up a new paradigm where you ask the agent for a task and then you step away and it comes back with a report. And yeah, I think as agents become more agentic, it will be longer and longer tasks. And this is a good example of one. Are these the longest running tasks you as have launched so far? I would say so. Like I just did one that was an hour long and I don't think I've ever seen that. I didn't know how long codex can run for. That's true. Yeah. Is there anything special that goes into making an agent run for so long without kind of flying off the rails?

10:43We have some tools to enable the model to be able to further extend these context lines. beyond what's original, like the harder limit. So that's the model is able to perform task by documenting what it's doing and step by step, like increase the time like it can do, the task of the task it can do without the human's interruption. Yeah. It's also the flow to go back and forth between the modeling human also is very nice. so I can correct it as it's going, right? Yeah, so this model is very flexible and collaborative and that was very important to us. So it's modeled after how you would interact with someone if you asked them to do a task for you.

11:31So imagine you're asking someone on Slack to do something for you, you'd probably give them instructions and then they'd ask you some questions and then maybe start doing the task and then maybe in the middle of the task, they'll say, oh, actually, can you clarify this to me or can you sign into this thing for me or what am I allowed to do this for you? And similarly, you might remember something that you forgot to say when you first gave them the task and you might want to interrupt them and just say, oh, hey, please also do this. Or you might want a status update if they're taking a long time to do it.

12:04Or you might want to redirect them if they're going on the wrong path. So that's what we modeled it after. And I think it's very important that the user and agent are both able to initiate communication with each other. So I think what we have now is probably the most basic version of what this could be, but it's better than anything We've released before in this area because at first the model can or the agent can ask You clarifying questions similar to deep research, but it's more flexible so it doesn't always ask you clarifying questions and then You can interrupt the model so you can say oh can you summarize what you've done so far or oh?

12:43I forgot to say I actually only want blue sneakers. And then if the model is going to take some kind of destructive action or if it needs you to log into something, it will also ask the user if it's allowed to do that before doing anything. On this topic, we kind of built this kind of computer interface you guys saw it where you can kind of watch along with what the agent is doing. And that actually persists for beyond the conversation. So, once it's done with the task, you can actually go back and ask it follow up questions and ask it to fix something or do another task. And you can also take over that computer so you can click in and then now you have access to its environment and you can click for it or log in for it or insert your credit card information or things like that.

13:36And so, yeah, I like to think of it as like looking over your core or shoulders and like being able to take over if necessary. Thank you for enabling the micromanager in me. Just kidding. So we'd love to talk a little about how this works, the extent that you can share. Yeah, so this agent is trained with the same technique as the one that we're supposed to learn. So we give this agent model the other tools we have like implemented in the same virtual machine like a text browser, like a GUI browser terminal and the image into and then the model will try to solve the task like we created like a which are pretty hard task that the model has to complete with using these tools and then kind of like we give we wrote the model if the model completed the task efficiently and correctly and for example like after this training, the model is able to, it should learn to switch between these tools like fluently, for tasks of, you know, you ask the model to research some restaurants and maybe booking a spot for you, it will first do a deep research style text -based browsing and then we'll probably also use the GUI browser to view the image of the food and also view the availability, like which is usually written in JavaScript that you have to use with a real GUI browser.

15:06And then, or, for example, if you ask it to create an artifacts, it's usually, you know, can pull sources from a website and then use them in the terminal. Yeah, I think the cool thing about this tool, compared to tool use implementation in the past, is that all of the tools have shared state. So it's all, it's like you're using, when you're using your computer and you have many different applications, you know, like if you download something, it's gonna be accessible to other applications. It's very similar, so the model can open a page in the text browser, which is more efficient, but then maybe it realizes it needs the visual browser, so it can just seamlessly switch, or it could download something using the browser and then in terminal it manipulates it or something like that.

15:47It can run something in terminal and then open it in the browser. It's very flexible. And so it's just giving the model a more powerful way of interacting with the intonate and files in its file system and code and things like that. Yeah, and one interesting thing to emphasize is that, like, we essentially give the model all these tools, and then lock it in the room and then get experiments. We don't really tell it when to use what tool. It figures that out by itself. It's almost magic. Is the technique, it sounds very similar to deep research we had you on the podcast before. Should we think about this as the standard technique of how open I think that agents will be trained going forward?

16:26I think we can take this really far. You know, this was, we haven't, our teams haven't been collaborating for that long. We even framed this model run as kind of minimum shipable de -resk. That was mostly for PR reasons internally, but this is really the most basic version we could make together. And I think we have so much further we could push this with these methods. For example, the slides capability is a new capability. It's a very, you know, already impressive. It's a great work from Adon, Paloma, Martin, a bunch of other people. But, you know, there's so much further we can push that and improve using the same techniques.

17:15But I think we can take it further, but we probably need other things too. Yeah, I feel so far it's pretty magical, like the same I agree with them. Just works on like O1 reasoning, like deep research with tool core. And then now like a more advanced computer use, browser use agents. Where does it run into the limits with this strategy and with this model specifically as well? I think the interesting thing with this model is that because it's taking, it's able to take actions with external side effects. There's a lot more risk. So for deep research, it was read -only. So there's kind of a limit to what the model could do in terms of like data exfiltration and other things.

17:59But with this, in theory, the model could successfully complete a task, but take a lot of harmful actions along the way. Like you could ask it to buy you something and it decides to buy just like a hundred different options To make sure that you're satisfied exactly or you know, you can think of many examples like that So I think that safety and safety training and Masegations was kind of one of the really cool parts of our process with this model and yeah Yeah, I was gonna mention that kind of along the same lines It's like this contact with the real world that makes things difficult We have to train this on a bunch of VMs.

18:40It's like thousands of VMs maybe. And things break and as soon as you're hitting a real website, the website's down or you're hitting all these capacity limits and load testing and this kind of thing. Yeah, it's really the very beginning. We're gonna iron out all these details and continue, but that's a major limitation. How do you think about from the safety perspective building in the right guard rails and like how do I make sure the model's not you know logging into my bank account and sending them all off to a Nigerian prince. Yeah that's a that's a very good question. Yeah this is definitely definitely an emerging risk where you know the internet's a scary place you know there a lot of like attackers and like scammers and this kind of thing fishing attacks like this goes on and on and yeah our model is a bit like it can it can reason about these things, like if you tell it to be careful, we've done some safety training to make this more robust, but sometimes it can get fooled, and sometimes it is a bit too over -eager to complete your task.

19:49We have a long list of mitigations, and the team has worked really hard to stack together a bunch of techniques to really try to make the model as safe as possible. So one example that I'll call out is that we have a monitor that looks over its shoulder and sees if anything looks funny, whether it's going on a weird website or anything like this, kind of like anti -virus for your computer. It's just kind of persistently watching. And then if it looks like there's anything suspicious, then it'll stop the trajectory and stop there. You know, of course we can't catch everything and this is a major area that we'll continue on, continue to iterate on.

20:33We do have like a protocol for if there are new attacks in the wild that we discover or we encounter, then we can rapidly respond and update these monitors, kind of like you would update your anti -virus software, like it would like kind of pick up on these new attacks and hopefully keep you safe. Yeah, I think the cool thing about the safety training is that it's been a really cross -borg effort from the safety team, governance team, legal team, research team, engineering team, like so many others. We have so many mitigations at every single level. We did a lot of external red teaming and internal red teaming, but yeah, as Casey mentioned, there's more, surely when we release the model, there will be new things we uncover, so we we just need to make sure we also have robust weights of detecting those and then mitigating those.

21:25For some of these models, there's a risk of what you can do with the models, whether it's creating biohazards or otherwise, how do you guys manage some of that? Yeah, it's actually bio has been heavily on our mind. Yeah, the team has been really thoughtful about, yeah, this agent, we think it's very powerful, it can do research, it can really speed up your work, But that also means that it could speed up harm. And kind of one of the top things that our team has been looking into is the risk of bio -risk, so like creating bio -weapons, this kind of thing. And yeah, the team has been really thoughtful about how to mitigate against this and generally being very cautious, we did like many weeks of red teaming to make sure that this model cannot be used for those harms.

22:17And a bunch of other mitigations in place, shout out Karen, who spearheaded this effort. And yeah, in general, I think we're very aware of this and just trying to be very cautious. Yeah, makes sense. Tell us a little about the team that came together to build this. So as Casey mentioned earlier, we had deep research, research team, and then deep research applied team, and operator, research team, well, So, computers using agent research team and operator applied team and we effectively merged everybody. We all work really closely together both the research team and the applied team. And the vibes have been great.

22:57It's been so far. Yes, and I have been transfer -long time. Yeah. So, like, it was a natural, like, it goes great. Yeah, it was a great one. How many of you are there? On deep research for the majority of the time, three or four, now we have some new people which is very exciting. And then on Kua. On Kua, I think around six to eight somewhere around there. On the research side. And then we have an amazing applied team, so like engineering product design, led by Yosh Kumar, and then he has just a really cracked engineering team. So it's been very fun to work really closely. I think that's one thing that's made this collaboration really special is that the research and applied teams work so closely.

23:40and even from the beginning when we're defining what the product should be able to do, it's very much a collaboration between research and products and design. So we go backwards from the use cases we want to be able to solve, to training the model and building the product. And obviously it's able to do, it's not able to do all of those things fully yet, and it can do some things that we didn't plan. But I think it's a good framework for us when we're starting a project. It's very grounded in how we want people to use it in the real world. It's a way smaller team than I was expecting. Small teams can do amazing things.

24:15So you felt a lot. Yeah, we haven't been working together for very long. It's been a few months. And actually the boundary between the research team and the applied team are not, you know, like a very, like a thesis because like, you know, during the model training, like lots of applied engineers, they are helping us training the model. And also after we changed the model, we are like some research team members are also working on the, you know, like a new set of models in the deploy the model to the real users. What was the hardest part about training this agent? Yeah, I think one of the biggest challenges we have is how to make training stable, especially given, you know, like when we change the research, it's only using browsing and the Python, it's like a, it's already, It's pretty mature tools there.

25:01We've been using it for a while. But when training the agent model, it has some new tools like computer. And also, the terminal like a bundle in the same container, in the same virtual machine as the computer. So it's actually quite hard to train because we are literally set up hundreds of thousands of virtual machines at the same time. And then they all like, you know, bidet to the internet. And we, so it's one of the big challenges. We see that actually the Chinese sometimes we feel but like, finally, we are very happy that we get this model. China. The VMs. Yes, yes. I remember. All back to the end.

25:44I'm not saying. Tell us about what's next. More sources, more tools, better model. How do you think about it? Well, I think one thing I like about our agent framing is that you can ask it to do whatever you want and you know you can ask it to do like every possible task you can imagine it just might not do it well and I said you tell it like go make me money on the internet you can tell it that you can try should we try that right but yeah I think it's really a matter of like improving the accuracy like the performance of task of of the whole distribution of tasks that anyone does on a computer.

26:27Right. Which is a lot of tasks. And then, yeah, and as soon as it's like an iterative deployment, we are very excited to see what's the new capabilities that our user will find in our agent, like the coding ability in deep research or deep research ability in operator. You were using the agent mode for coding. Yeah, I use it for coding a lot because I feel it's actually not, you know, it's I could very always try to rewrite my whole code base. It just actually have some small editing. And also, it actually reads the original docs of different functions pretty well. So I feel it loosened it less on the function quality.

27:11Oh, interesting. How do you choose when to go to codex versus when to go to agent for that? For the agent, we have to multiply it to how I use for O3. So it's more like an interactive experience. For the codex, it's more like, you know, you have some, you know, where you design the problem that you want a co -worker to solve. And then it will make a PR for you. But for the agent, it's more like just give you a function or give you a suggestion. Cool. And it can do code search because it can access GitHub through the API connector. So code search kind of things. Yeah. It almost feels like, you know, the agent roadmap up until now, you've built the different, almost like appendages of what it would take to have an agent, and by combining them all, this really is like the first fully embodied agent on a computer.

27:58I think it's very exciting. Yeah, I think another area that we're excited to push on is the experience of collaborating with the agent. I think this model is actually very good at multi -turn conversations, and it's very nice to continue walking on a task with. I think that's one of the deficiencies of deep research. A lot of people will do multiple deep research requests in a single conversation, but it doesn't always work so well. So I think we're really happy with this model's multi -turnability, and we just want to improve even further. And then I also think personalization and memory for agents will also be very important.

Read the full transcript

28:37And also right now, every agent task is initiated by the user, but in future, it should also would be doing things for you without you having to even ask in the first place. Yeah, I'm also pretty excited about the UI and UX surrounding the agent. Because right now, I think obviously we're working in a chat TVT world. It's like you start a conversation and it goes. But you can imagine a lot of different modes of interacting with an agent. And I'm really excited to explore different ways of interacting with the agent. Do you see this as always being a single mission super agent or will there be the financial analyst subagent and the personal party planner subagents?

29:23What's your vision for how that kind of plays out? I think people have different opinions on this. I think in the limit if you could just ask one thing and it can figure out what it needs to do to finish the thing that you want it to do for you, that seems like it would be easiest. If you just had a really amazing sheep of stuff who knows how to route things correctly and basically Can do anything you need that seems like it would be Pretty easy. I think I agree with that take and like even in some archer trajectories where like I don't know you're asking about I don't know maybe like a shopping task like sometimes it'll go into terminal and like do some like like calculations, budget.

30:06And I think the model should be free to use all the tools that wants, it doesn't need to be a financial analyst to have the financial analyst like toolset. Yeah, I feel like when you launch the product, it sometimes makes sense to have some GPS, like a customized model or customized instruction to put the model into a specific role. But in general, when you're training the model, there are lots of positive transfer between deep research core operations, also slice generation like all of these scales are transferable. So it makes much more sense to just have a single agent as an underlying base model.

30:47Totally. I guess even though people do different types of work, we're all fundamentally sending emails, we're making slide decks, we're doing a lot of the same work in front of a computer. I'd love to understand some of the learnings from the reinforcement learning perspective. It seems like that's the method that that seems to really be working for you guys with agents. Was it like very data -intensive to kind of get to this point of having an agent that's so good at, you know, such a wide variety of tasks? Or like what were some of the learnings from an RL perspective? Yeah, so, you know, actually we create a bunch of very diverse set of tasks, like some task to find some, you know, like a very niche topic, or a very niche answer in the internet, or some tasks, you know, like just very similar to deep research, like you need to write a whole four -day -length article and also lots of, you know, tasks, like just all the tasks that we want the model to be good at.

31:44And yeah, so far we think that as long as, you know, like you came, you know, great this task, we are like, after the model give you a result, you can judge whether the, you know, good or not, you came kind of like reliable, a trained model to be like even better on this task. Was there anything especially needed to do to make sure it had good turned by turn interaction with users when doing that training, or was it just about the type of trademarks you collected? Yeah, so like I think most of the time we focus on end -to -end performance, like you know, like a front, front, a rear -space, best value prompt how to accomplish a task.

32:26And somehow it's very good at working with users. To your question, the reinforcement lining is very data efficient. So that means that we're able to curate a much smaller set of very high quality data. The scale of the data just so miniscule compared to the scale of pre -training data. So we're able to teach the model new capabilities by just curating these much smaller high quality data sets. I will say to get the operator piece to work well. You know, before we do RL, the model has to be good enough to have a basic completion of tasks. And our team has spent a lot of time in the past, like over the past two, maybe three years, getting the model to that point where it's able to actually reason about a page and like kind of understand visual elements really well.

33:19So this model is built on all that as well. Actually, could you say a little bit more about that? Because I remember early days of opening I, this was always part of the world of bit stuff. And you're trying to rl the mouse paths. And it was just way too unbounded of a problem. What's changed now for that to be kind of solvable? Yeah, that's great that you point out the world of bits. That's, this project does have a very long lineage dating back to 2017 or so. So actually our codename is World of Bits 2, for the computer use part. That's awesome. And yeah, what's changed? I think essentially the scale of the training has changed.

34:04We have, I don't know the multiplier, but it must be 100 ,000X or something like in terms of compute. The amount of training data we've done, both in pre -training and RL. So yeah, I really think it's just scale and the scale catching up to our ambition, I guess. Wow, scale is all you need. I believe it. We have some good data. Are there particular capabilities or functionality that you're especially excited about in agent mode? Yeah, so this model is actually a pretty good at doing some real research like data science and also, you know, summarize the reports or like the findings in a spreadsheet.

34:46So we have some evaluation like, you know, data science bench, we evaluate the model, and it's actually outperform the human baseline. So in some sense, it's actually superhuman in some research tasks and we can rely on the model to, you know, perform some basic analysis for us. And this is an area that John Blackwin and our team was really pushing on, like spreadsheets and data science, so shout out John. spreadsheets and data science you are elevating us out of a job over here elevating us out of a job and hausing Another thing I'm excited about is um uh you know we released operator in January and um I you know it was decent at clicking around but I think we substantially improved that capability where it's like much more accurate and just kind of getting the basic things right um is what I'm actually excited about where it can like like reliably fill out a form and, you know, do those kind of things.

35:41Date picking? Date picking still needs a bit of work, but. Because the reason date picking is just the hottest task. It's hard for humans. They look like picking a date in the calendar drum dance. Okay, last question. It seems like you guys have the overall framework and structure in place for something really interesting here. What's ahead? Where do you go from here? I think the thing that we're really excited about is that this tool that we've given the model access to is very general. It's basically most of what you could do on a computer. And if you think about all of the tasks that a human can do on a computer, it's very extensive.

36:22And so now we kind of feel like it's a matter of us making the model good at all of those tasks too and figuring out a way of training on as diverse of tasks as possible with this very general tool. So I think there's a lot of hard work ahead of us, but we're very excited about it. I think what was so excited about pushing different forms, ways of interacting with the agent. I think there'll be a lot of new interaction paradigms between users and use virtual assistance or agents. So a lot of exciting times ahead. I can't wait to see it. Thank you. Thanks for joining us. Congratulations on the lunch.

37:03Thank you so much. Thank you. Thank you for having us.

From the publisher

Isa Fulford, Casey Chu, and Edward Sun from OpenAI's ChatGPT agent team reveal how they combined Deep Research and Operator into a single, powerful AI agent that can perform complex, multi-step tasks lasting up to an hour. By giving the model access to a virtual computer with text browsing, visual browsing, terminal access, and API integrations—all with shared state—they've created what may be the first truly embodied AI assistant. The team discusses their reinforcement learning approach, safety mitigations for real-world actions, and how small teams can build transformative AI products through close research-applied collaboration.

Hosted by Sonya Huang and Lauren Reeder, Sequoia Capital

More from Training Data

All 110 episodes
OpenAI Just Released ChatGPT Agent, Its Most Powerful Agent YetTraining Data · 38 min
Listen in VO