In short
Podcast Episode Notes: OpenAI’s Deep Research Team on Why Reinforcement Learning is the Future for AI Agents
Podcast Overview Title: Training Data Description: The podcast hosts discussions on AI technologies and their implications with industry leaders and researchers. Hosts: Sonya Huang and Lauren Reeder, Sequoia Capital Episode Description: OpenAI's Isa Fulford and Josh Tobin discuss the innovations behind the Deep Research agent, including its end-to-end training, high-quality data usage, and its potential impact on knowledge work.
---
Key Themes and Concepts
- Introduction to Deep Research
- Definition: Deep Research is an AI agent that can perform extensive online research and generate detailed reports.
- Capabilities:
- Processes complex tasks that typically require hours of human effort.
- Generates responses in significantly less time (5 to 30 minutes).
- Provides in-depth answers with specific sources.
- Development and Inspiration
- Origin Story:
- Development stemmed from previous successes in training models using new reasoning paradigms.
- Focus shifted to tasks requiring extensive online research.
- Team Involvement:
- Developed by a team including Isa Fulford and Josh Tobin, with contributions from skilled engineers.
- Use Cases and Applications
- Target Audience:
- Knowledge workers, including researchers, medical professionals, and individuals engaged in detailed tasks such as market analysis or even personal planning (e.g., birthday parties).
- Examples of Use Cases:
- Medical research and analysis.
- Market understanding and competitive analysis.
- Consumer research for shopping and travel planning.
- Unexpected applications in coding and technical documentation.
- Technical Aspects
- Training Methodology:
- Utilizes end-to-end reinforcement learning on complex reasoning tasks.
- Employs a browsing tool to access real-time web information.
- Model Capabilities:
- Synthesizes information and maintains clarity in reports.
- Offers citations for transparency and trustworthiness.
- Future Prospects
- Expansion Plans:
- Focus on improving browsing capabilities and expanding data sources.
- Potential integration of private data in future iterations.
- Vision for AI Agents:
- Aim to create versatile agents capable of handling a wide array of tasks with human-like flexibility.
- Reinforcement Learning (RL)
- Importance in AI Development:
- Seen as critical for building powerful AI agents, especially for optimizing performance in complex tasks.
- Perspective shift towards RL is fueled by advancements in foundational models that provide a robust base for further training.
---
Key Takeaways
- End-to-End Learning: The shift towards training models end-to-end rather than using predefined operational graphs enhances flexibility and effectiveness.
- Quality of Data: High-quality datasets are pivotal for training successful models, underscoring the age-old adage in machine learning.
- Trust and Transparency: Features like citations integrated into AI responses enhance user trust and clarity.
- Impact on Knowledge Work: Deep Research is expected to significantly reduce the time spent on various tasks, offering substantial time savings and efficiency improvements across multiple domains.
- Future of AI Agents: There is a strong belief that AI agents will evolve to take on more complex tasks, fundamentally changing the landscape of knowledge work and personal productivity.
---
Conclusion The discussion emphasizes OpenAI's commitment to advancing AI through innovative methodologies like reinforcement learning. The transformative potential of Deep Research and similar agents signals a future where time-saving AI tools play an integral role in both professional and personal spheres.
---
References
- Yann LeCun’s Cake analogy (referenced during discussions).
- Mention of potential future developments and complementary products related to AI agents.
Note
The content of this podcast episode does not constitute investment advice or offers related to investment services.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00A lesson that I've seen people learn over and over again in this field is like, we think that we can do things that are smarter than what the models do by writing it ourselves. But as the field progresses, the model come up with better solutions to things than humans do. The number one lesson on machine learning is like, you get what you optimize for. And so if you're able to set up the system such that you can optimize directly for the outcome that you're looking for, the results are going to be much, much better than if you sort of try to glue together models that are not optimized end to end for the task they are trying to have them do.
0:32So my like long -term guidance is that, you know, I think reinforcement learning tuning on top of models is probably going to be a critical part of how the most powerful agents get built.
1:01We're excited to welcome Issa Fulford and Josh Tobin, who lead the Deep Research product at OpenAI. Deep research launched three weeks ago and has quickly become a hit product used by many tech luminaries like the call -sins for everything from industry analysis to medical research to birthday party planning. Deep research was trained using end -to -end reinforcement learning on hard browsing and reasoning tasks and is the second product in a series of agent lunches from OpenAI with the first being operator. We talk to Esa and Josh about everything from deep researchers use cases to how the technology works under the hood to what we should expect in future agent lunches from OpenAI.
1:38Esa and Josh, welcome to the show. Thank you. Thank you so much for joining us. This is a great pleasure. I'm excited to be here. Thank you for having us. So maybe let's start with like what is Deep Research? Tell us about the origin stories about this product is doing. So deep research is an agent that is able to search many online websites and it can create very comprehensive reports. It can do tasks that would take humans many hours to complete and it's in chat GBT and it takes like five to 30 minutes to answer you and so it's able to do much more in -depth research and answer your questions with much more detail and specific sources than regular chat to be to your response would be able to do.
2:19It's one of the first aidness that we've released, so we released operator pretty recently as well, and so deep research is the second agent, you know, will release any more in future. What's the origin story behind deep research? Like when did you choose to do this? What was the inspiration and how many people work on it? Like what did it take to bring this to fruition? Good question. This is before my time. Oh yeah. So I think maybe around a year ago we were seeing a lot of success internally with this new reasoning paradigm and training models to think before responding. And we were focusing a lot on math and science domains, but I think that the other thing that this kind of new reasoning model regime unlocks is the ability to do longer horizon tasks that involve like like agentic kind of abilities.
3:11And so we thought a lot of people do tasks that require a lot of online research or a lot of external context. And that involves a lot of reasoning and discriminating between sources and you have to be quite creative to do those kinds of things. And I think we finally had models or a way of training models that would allow us to be able to tackle some of those tasks. So we decided to try and start training models to do first browsing tasks. So using the same methods that we used to train reasoning models, but on more real world tasks. Was it your idea and Josh, how'd you get involved? At first it was me and Yash Patil, who as our opening eye is working on a similar project that will be at release at some point, which we're very excited about.
3:57And we built an original demo and then also with Thomas Dipson, and who's one of those people who just isn't an amazing engineer, like, well, dive into anything and just get loads of things on. So it was very fun. Yeah, and I joined more recently. I rejoined opening I about six months ago from my startup. I was an eye opening I in the early days and I was looking around the projects when I rejoined and got very interested in some of our agentic efforts including this one and got involved with that. Amazing. Well, tell us a little about who you built it for. Yeah, I mean, it's really for anyone who does knowledge work as part of their as part of their day -to -day job or really as part of their life So we're seeing A lot of the usage comes from people using it for work doing things like you know research as part of their jobs for you know understanding markets companies these real estate.
4:58A lot of scientific research, medical. I think we've seen a lot of medical examples as well. And one of the things we're really excited about as well is this style of like, I just need to go out and spend many hours doing something that where I have to do a bunch of web searches and collated bunch of information is not just a work thing, but it's also useful for shopping and travel as well. So we're excited for the plus launch so that more people will be able to try deep research and maybe we'll see some new use cases as well. It's definitely one of the products I've used the most over the last couple weeks.
5:30It's been amazing. Using it for work? For work definitely. Also for fun. What are you using it for? Oh, for me, oh my goodness. I'm starting thinking about buying a new car and I was trying to figure out when the next model was going to be released for the car. And there's all these speculative blog posts like there's patterns in the manufacturer and so I asked deep research, can you break down all the gossip about this car and then all of the facts about what they've done what this automaker said before and it put together an amazing report that told me maybe wait a couple months, but this year, like in the next few months, it should come out.
6:01Yeah, like one of the things that's really cool about it is it's, it's, like, it's not just for going broad and gathering all of the information about a source, but it's also really good at finding like very obscure, like weird facts on the internet. Like if you have something very specific you want to know that you might not just turn up in the first page of search results, So it's good at that kind of thing too. So that's cool. What are some of the surprising use cases that you've seen? Ooh. I think the thing I've been most surprised by is how many people are using it for coding. Yeah. Which wasn't really a use case I'd considered, but I've seen a lot of people on Twitter and in various places where we get feedback using it for coding and code search.
6:41And also for finding the latest documentation or a certain package or something and helping them write a script or something. So yeah, I'm like, I'm kind of embarrassed that we didn't think of that as a use case because it's like, you know, for chat, you be teasers, it seems so obvious, but it's, I know, it's impressive how well it works. How do you think the balance of business versus individual use case will evolve over time? Like you mentioned the plus launch that's happening, you know, in a year's time or two years time, would you guess this is mostly a business tool or mostly a consumer tool?
7:09I would say hopefully both. I think it's a pretty general capability, which, and I think it's something that we do both and work and personal life. So, yeah, I'm excited about both. I think the magic of it is like, it just saves people a lot of time. You know, if there's something that might have taken you hours or in some cases, we've heard like days, people can just put it in here and get, you know, 90 % of what they would have come up with on their own. And so, yeah, I tend to think there's more tasks like that in business than there are personal, but I mean, I think for sure it's going to be part of people's lives in both.
7:51It's really become the majority of my usage for chat. You've just always picked deep research rather than normal. So what are you seeing in terms of consumer use cases? And what are you excited about? I think a lot of shopping, travel recommendations, I've first used the model a lot. I've been using it for months to do these kinds of things. We were in Japan for the launch of Deep Research, so it was very helpful in finding restaurants with very specific requirements and finding things that I wouldn't have necessarily found. Yeah, and I found it like when you have something, it's like the kind of thing where, you know, if you're shopping maybe for something expensive or you're planning a trip that is special or you want to spend a lot of, that you want to spend a lot of time thinking about, it's like for me, you know, I might go and spend hours and hours like trying to read everything on the internet about this one, this product that I'm interested in buying, like scouring all the reviews and the forums and stuff like that and deep research can put together kind of like something like that very quickly.
8:58And so it's really useful for that kind of thing. The models will so very good at instruction following. So if you have a query with many different parts or many different questions. So if you don't you want the information about the product, but you also want comparisons to all other products, then you also want information about reviews from a, you know, Reddit or something like that. You can give loads of different requirements, and it will do all of them for you. Yeah, another tip is like just ask it to format it in a table. It will usually do that anyway, but it's like if you it's really helpful to have like a table with a bunch of citations and things like that for all the categories of things that you want to research.
9:38Yeah, there are also some features that hopefully we'll get into the product at some point, but the model is able to, the underlying model is able to embed images so it can find images of the products. And it's also, this is not a consumer use case, but it's able to create graphs as well and then embed those in its response. So hopefully that will come to charge you to you soon as well. Nurti consumer use case. Yeah, yeah. Well, speaking of nerdy consumer use cases, also personalized education is a really interesting use case. If there's a topic that you've been meaning to learn about, if you need to brush up on your biology or you want to learn about some world event, it's really good at putting all the information about what you feel like you don't understand, what aspects of it you wanted to go do research on it and it'll put together a nice report for you.
10:31One of my friends is considering starting a CPG company and he's been using it so much to find similar products to see if specific names are already, you know, what the domain has already taken, market sizing, like all of these different things. So that's when fun to heal, share the reports with me and I'll read them. So that's really fun to see. Another fun use case is it's really good at finding like a single, like obscure fact on the internet, like if there's like a, you know, like an obscure TV show or something that you want to, you know, to like find like one particular episode of or something like that, it'll go and it'll go deep and find, but like one reference to it on the web.
11:14Oh yeah, my brother's friend's dad had this very specific fact. It was about some Austrian general who was empowered during a certain death of someone during a battle, like a very niche question. And apparently, Chargit B .T. had previously answered it wrong, and he was very sure that it was wrong. So you went to the public library and found a record and found that it was wrong. And so then deep research was able to get it right. So we sent it to him, and he was excited. What is the rough mental model for what deep research is excellent at today? And, you know, where should people be using the O series of models?
11:53Where should they be using deep research? What deep research really excels at is if you have a sort of detailed description of what you want, and in order to get the best possible answer, it requires reading a lot of the internet. If you have kind of like more of a big question, it'll help you kind of clarify what you want, But it's really at its best when there's a specific set of information that you're looking for. I think it's very good at synthesizing information that encounters. It's very good at finding specific, like hard to find information. But it's maybe less, and it can make some new insights, I guess, from what it encounters, but I don't think it's necessary.
12:37It's not making new scientific to discover is yet. And then I think using the O -Series model for me, if I'm asking for something to do with coding, usually it doesn't require knowledge outside of what the model already knows from it pre -training. So, how you would usually use O1 Pro or O1 for coding? Or O3 Mini, hi. And so, Deep Research was a great example of where some of the new product directions for opening our gulling. I'm curious how can today's set you can share how does it work? The model that powers deep research is a fine tune version of 03 which is our most advanced reasoning model and we specifically trained it on hard browsing tasks that we collected as well as Reason other reasoning tasks and so it also has access to a browsing tool and Python tool So through training and to end on those tasks, it learned like strategies to solve them.
13:39And the resulting models is good at online search analysis. Yeah, and like intuitively the way you can think about it is you make this sort of this request, ideally a detailed request about what you want. The model thinks hard about that. It searches for information. It pulls that information and it reads it. And it understands how it relates to that request and then decides what to search for next in order to get kind of closer to the final answer that you want. And it's trained to do a good job of pulling together all of those, all that information to a nice, tidy report with citations that point back to the original information that I found.
14:16Yeah, I think what's new about deep research as an agentic capability is that because we have the ability to train and to end. There are a lot of things that you have to do in the process of doing research that you couldn't really predict beforehand. So I don't think it's possible to write some kind of language model program or script that would be as flexible as what the model's able to learn through training where it's actually reacted to live web information. And based on something it sees, it has to make a change strategy and things like that. So we actually see it doing pretty creative searches.
14:53You can read the chain of thought summary, and I'm sure you can see sometimes it. It's very, very smart about how it comes up with the next thing to look for. So John Carlson had a tweet that went somewhat viral, how much of the magic of deep research is real time access to web content, and how much of the magic is in chain of thought? Can you maybe shed some light on that? I think it's definitely a combination. I think you can see that because there are other search products that don't necessarily, that won't train to end to end. So, it won't be as flexible in responding to your responding to information in Acanters won't be as creative about how to solve specific problems because they won't specifically train for that purpose.
15:43So it's definitely a combination. I mean, it's a fine -tune version of O3, O3 is a very smart and powerful model. A lot of the NASA's capability is also from the underlying O3 model training. But so I think it's definitely a combination. Before OpenAI was working at a startup and we were dabbling in building agents kind of the way that I see most people describe building agents on the internet, which is essentially, you construct this graph of operations and some of the nodes in that graph are language models. And so the language model can decide what to do next, but the overarching logic of the sequence of steps that happen is defined by a human.
16:28And what we found is that it's really, it's like a powerful way of building things to get quickly to a prototype, but it falls down pretty quickly in the real world because it's very hard to anticipate all the scenarios that the model might face and think about all the different branches of the path that you might want to take. In addition to that, the models often are not the best decision makers at nodes in that graph because they weren't trained to do to make those decisions. They were trained to do things that look similar to that. And so I think the thing that's really powerful about this model is that it's trained directly end to end to solve the kinds of tasks that users are using it to solve.
17:11So you don't have to set up a graph or make those node -like decisions on the architecture on the back end. It's all driven by the model itself. Yeah. Can you say more about this? Because it seems like that's one of the very opinionated decisions that you've made. And clearly it's worked. There's so many companies that are building on your API, kind of prompting to, you know, to, you know, solve specific tasks for specific users. Do you think a lot of those applications would be better served by kind of having, you know, trained models end to end for their specific workflows? I think if you have a very specific workflow that is quite predictable, it makes a lot of sense to do something, like Josh described, but if you have something that has a lot of edge cases or it needs to be quite flexible, then I think something similar to deep research is probably the better approach.
18:04Yeah, I think like the guidance I give people is, the one thing that you don't want to bake into the model is like kind of hard and fast rules. Like if you have a database that you don't want the model to touch or something like that, it's better to encode that in human written logic, but I think it's kind of like like a lesson that I've seen people learn over and over again in this field is like, we think that we can do things that are smarter than what the models do by writing it ourselves. But in reality, usually the model, like as the field progresses, the model come up with better solutions to things than humans do.
18:41And also, like, you know, the, like, probably number one lesson on machine learning is like you get what you optimize for. And so if you're able to set up the system such that you can optimize directly for the outcome that you're looking for, the results are going to be much, much better than if you sort of try to glue together models that are not optimized end to end for the task they are trying to have them do. So my like long -term guidance is that, you know, I think like reinforcement learning tuning on top of models is probably going to be a critical part of how the most powerful agents get built.
19:16What were the biggest technical challenges along the way to making this work? Well, I mean, maybe I can say as like an observer rather than someone who was involved in this from the beginning, but it seems like kind of one of the things that ESA and the rest of the team worked really, really hard on and was kind of like one of the hidden keys to success was like making really high quality data sets. It's another one of those like age -old lessons in machine learning that people keep relearning, but the quality of the data that you put into the model is probably the biggest to her many factor in the quality of the model that you get on the other side.
19:47And then have someone like Edward Sun, his other person who works on the project, who just any data set, he will optimize. So that's a secret to success. Find your Edward. Yeah, great machine learning, the model training. How do you make sure that it's bright? Yeah, so that's obviously a core part of this model and product is that we want it to be, users to be able to trust the outputs. So part of that is we have citations and so users are able to see where the model is citing its information from and we during training that's something that we actually like try and make sure is correct but it's still possible for the model to make mistakes or to eliminate or trust a source that maybe isn't the most trustworthy source of information.
20:45So that's definitely an active area where we can want to continue improving the model. How should we think about this together with, you know, O3 and operator and other, other different leases? Like, does this use operators? Do these all build on top of each other? Or are they all kind of a series of different applications of O3? Today, these are pretty disconnected. But you can kind of, you can imagine kind of where we're going with this, right, which is like the ultimate agent that people have access to at some point in the future should be able to do, you know, not just web search or using a computer or any of the other types of actions that you'd want, like kind of a human assistant to do, but should be able to fuse all of these things in a more natural way.
21:31Any other design decisions that, you know, you've taken that are maybe not obvious at first glance? I think one of them is the clarification flow. So if you've used deep research, the model will ask you questions before signing its research. And usually, it's actually, maybe you'll ask you a question at the end of its response, but it usually doesn't have such that kind of behavior up front. And that was intentional because you will get the best response from the deep research model if the prompt is very well specified and detailed. And I think that it's not the natural user behavior to give all of the information in the first prompt.
22:11So we wanted to make sure that if you're going to wait for five minutes, 30 minutes, that your response is as detailed and you're satisfactory. So we added this additional step to make sure that the user provides all the detail that we would need. And I've actually seen a bunch of people on Twitter saying that they have this flow or that they will talk to O1 or O1 Pro or to help make their prompt more detailed and then once they're happy with the prompt, then they'll send it to deep research, which is interesting. So you were finding their own work list, so how do you use this? So there's been three different deep research products in the last few months.
22:50Tell us a little bit about what makes you guys special and how we should think about it. And they're all called deep research, right? What's up with that? Yeah, not a lot of naming creativity in this field. I think people should try all of them for themselves and get a feel. I think the difference in quality, I think they all have pros and cons, but I think the difference will be clear. But what that comes down to is just the way that this model was built. And the effort that went into constructing the data sets and then the engine that we have with the O series models, which allows us to just optimize models to make things that are really smart and really heck -quality.
23:36We had the O -N team on the podcast last year, and we were joking that OpenAI has not that good at naming things. I will say this is your best -named product. Deep researches. At least it describes what it does, I guess. Yeah. So I'm curious to hear a little bit about where you want to go from here. You have deep research today. What do you think it looks like a year from now and what maybe are complementary things you want to build along the way. We're excited to expand the data sources that the model has access to. We've trained the model that's generally very good at browsing public information, but it should also be able to search private data as well.
24:14And then I think just pushing the capabilities further, so it could be better at browsing, it could be better at analysis. And then thinking about how this fits into our agent roadmap more broadly. I think the recipe here is something that's going to scale to a pretty wide range of use cases. Things that are kind of surprise people how well they work. But this idea of you take a state of the art reasoning model, you give it access to the same tools that humans can use to do their jobs or to go about their daily lives. and then you optimize directly for the kinds of outcomes that you're looking that you want the agent to be able to do.
24:58That recipe, there's like really nothing stopping that recipe from scaling to more and more complex tasks. So I feel like a GIs like an operational problem now. And I think, yeah, a lot of things to come in that general formula. I'm so Sam, how the pretty striking quote of deep research will kind of take over a single a little bit of a percentage of all economically viable tasks, or all the valuable tasks in the world. How should we think about that? I think about it as like, it's, deep research is not capable of doing all of what you do, but it is capable of saying you like hours or sometimes in some cases, days at a time.
25:40And so I think like, what we're hopefully relatively close to is deep research and the agents that we build next and the agents that we build on top of it, giving you you know, one, five, ten, 25 % of your time back depending on the type of work that you do. I mean, I think you've already automated 80 % of what I do. It's definitely on the higher end for me. We just need to start writing checks, I guess, yeah. Are there entire job categories that you think are kind of more at risk because they're but more in the strike zone for what deep research is exceptional. So, for example, I'm thinking consulting, but are there specific categories that you think are more in the strike zone?
26:22Yeah, I used to be consulting. I don't think any jobs are at risk. Like, I don't really think of this as a labor replacement kind of thing. At all, like it's... But for these types of knowledge work jobs where you are spending a lot of your time kind of looking through information making conclusions, I think it's going to give people superpowers. I'm very excited about a lot of the medical use cases. Just the ability to find all of the literature or all of the recent cases for a certain condition. I think I've already seen a lot of doctors posting about this or they've reached out to us and said, oh, we used it for this thing.
27:01We used it to help find a clinical trial for this patient or something like that. So just people who are already so busy just saving some time, or it's maybe something that they wouldn't have had time to do. So now they are able to have that information for them. Yeah, and I think the impact of that is maybe a little bit more profound than it sounds on the surface. It's not just like getting 5 % of your time back, but the type of thing that might have taken you four hours or eight hours to do, So now you can do for a chat, chattach, if you see subscription in five minutes. And so what types of things would you do if you had infinite time that now maybe you can do like many, many copies of?
27:44So like, should you do research on every single possible startup that you could invest in instead of just the ones that you have time to meet with, things like that? Or on the consumer side, one thing that I'm thinking of is, the working mom that's too busy, the plan of birthday party for her toddler. Like now it's, now it's doable. So it's, I agree with you, it's way more important than 5 % of your time. It's all the things you couldn't do before. Exactly. What does this change about education and the way we should learn? And you know, what will you be teaching your kids now that we're in the world of agents and deep research?
Read the full transcript
28:16Education's been not like one of the top few things that people use it for. I think it's, I mean, this is true for a chat with you generally. It's like learning things by talking to an AI system that is able to personalize the information that it gives you based on what you tell it or maybe in the future what it knows about you. It feels like a much more efficient way to learn, and a much more engaging way to learn than reading textbooks. We have some lightning round questions. All right. Okay. Your favorite deep research use case. I'll say yeah, like personalized education, just like learning about anything I won't learn about.
28:55I've already mentioned this, but I think a lot of the personal stories that people have shared about finding information about a diagnosis that they've received or someone in their family received have been really great to see. Okay, we saw a few application categories break out last year. So for example, coding, being an obvious one, what application categories do you think will break out this year. I mean, clearly agents. Agents are going to say to. Okay. 2025's the year of the agent. I think so. Yeah. And then how do you think about what piece of content that you should be recommending people reading to read to learn more about agents or where the state of AI is going?
29:37It could be an author too. Training data. This podcast. Not biased. I think it's like, it's so hard to keep up with the state of the art in AI. I think that the general advice I have for people is pick one or two subtopics that you're really interested in and go curate a list of people who think are saying interesting things about it and how to find those one or two things they're interested in. Maybe actually that's a good deep research use case. Go use it to go deep on things that you want to learn more about. This is a bit old now, but I think a few years ago I watched there. I think it's called like foundations of RL or something like this from Peter Abiel.
30:23And it's a few years old, but I think that it was a good introduction to Reinforcement Learning. Yeah, it would definitely second any content by Peter Abiel, my grad school advisor. Oh yeah. Okay. Reinforcement learning. It kind of went through a peak and then felt like it was a little bit of a dull dream again and is speaking again. Is that the right read on what's happening with RL? It's so back. Yeah. So back. Yeah. Why? Why now? Because everything else is working. Like I think if you, maybe people have been following the field for a while, we'll remember the John Lekun cake analogy. If you're building a cake, then most of the cake is the cake.
31:07And then there's a little bit of frosting and then there's a few cherries on top. And the analogy is that like unsupervised learning is the cake. Supervised learning is the frosting and reinforcement learning is the cherries on top. When we in the field were working on reinforcement learning back in 2015 -2016, it's kind of like, I think, Yann LeCoon's analogy, which I think in retrospect is probably correct, is that we were like trying to add the cherries before we had the cake. But now we have language models that are pre -trained on massive amounts of data and are incredibly capable. We know how to do supervised fine tuning on those language models to make them good at instruction following and like generally doing the things that people want them to do.
31:47And so now that that works really well, it's like very ripe to tune those models for any kind of use case that you can define or reward function for. Great. Okay. So from this lightning round, we got agents will be, you know, the breakout category in 2025 and reinforcement learning is so back. I love it. Thank you guys so much for joining us. We love this conversation. Congratulations. Thanks for having us. see what comes of it. Thank you. Thank you.
From the publisher
OpenAI’s Isa Fulford and Josh Tobin discuss how the company’s newest agent, Deep Research, represents a breakthrough in AI research capabilities by training models end-to-end rather than using hand-coded operational graphs. The product leads explain how high-quality training data and the o3 model’s reasoning abilities enable adaptable research strategies, and why OpenAI thinks Deep Research will capture a meaningful percentage of knowledge work. Key product decisions that build transparency and trust include citations and clarification flows. By compressing hours of work into minutes, Deep Research transforms what’s possible for many business and consumer use cases.
Hosted by: Sonya Huang and Lauren Reeder, Sequoia Capital
Mentioned in this episode:
Yann Lecun’s Cake: An analogy Meta AI’s leader shared in his 2016 NIPS keynote




