[State of RL/Reasoning] IMO/IOI Gold, OpenAI o3/GPT-5, and Cursor Composer — Ashvin Nair, Cursor

30 Dec 2025 · 23 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Summary: Latent Space: The AI Engineer Podcast

Episode Title

[State of RL/Reasoning] IMO/IOI Gold, OpenAI o3/GPT-5, and Cursor Composer — Ashvin Nair, Cursor

Overview In this episode, Ashvin Nair, a recent addition to Cursor and former member of OpenAI, discusses his journey from robotics to working with language models, the current state of reinforcement learning (RL), and the evolving landscape of AI models. The conversation touches upon the dynamics of robotics and AI, the implications of recent developments in RL, and the shifts in model architecture and capabilities at OpenAI.

Key Themes and Discussions

  1. Transition from Robotics to Language Models
  2. Ashvin Nair reflects on his background in robotics and how it aligns with current trends in AI, particularly with language models.
  3. He notes that many people from robotics are transitioning into AI, emphasizing the data-driven nature of both fields.
  1. Current State of Robotics
  2. Nair expresses skepticism about the current utility of robotics compared to language models, suggesting that LLMs (Large Language Models) are likely to generate more economic value in the near future.
  3. He believes that robotics companies are currently undervalued compared to their AI counterparts, and that robotics is at a nascent stage, similar to the early development of language models.
  1. Reinforcement Learning and Model Performance
  2. The discussion shifts to RL, where Nair shares insights from his time at OpenAI, specifically on the Codex team and the development of models capable of achieving IOI Gold in programming competitions.
  3. He highlights the pitfalls of benchmark overfitting and the importance of understanding why certain models perform well on specific tasks.
  1. The Challenge of Generalization
  2. Nair outlines the difficulties in RL, particularly regarding its generalizability beyond training distributions and benchmarking tasks.
  3. He mentions the need for models to be designed with context in mind and to integrate various aspects of user interaction.
  1. OpenAI's Evolving Model Strategy
  2. The conversation includes a critique of OpenAI's approach to model development, especially their shift away from the "one model fits all" philosophy.
  3. Nair discusses the implications of this shift and how it reflects broader trends in AI development.
  1. Cursor's Unique Approach
  2. Nair explains his motivation for joining Cursor, emphasizing the potential for closer collaboration between product and ML teams.
  3. He discusses the innovative work being done at Cursor, particularly with the Composer model, and how it aims to automate software engineering processes.

Key Takeaways

  • Robotics vs. LLMs: The current landscape favors language models economically, but there’s potential for robotics to catch up as technology progresses.
  • Benchmarking and Overfitting: There is a need for the AI community to rethink benchmark practices to avoid overfitting and ensure that models generalize well.
  • AI Model Strategy: OpenAI's shift from a singular model approach reflects a growing recognition of the complexity and variability of AI applications.
  • Cursor's Vision: Cursor aims to create a more integrated approach to AI product development, focusing on user interaction and real-world applicability.

Conclusion The episode encapsulates a dynamic conversation about the intersection of robotics, AI, and RL, with Ashvin Nair providing valuable insights drawn from his experiences in leading organizations. His observations on the evolution of AI models and the importance of context in learning highlight the ongoing challenges and opportunities in the AI landscape.

For further insights, visit [Latent Space](https://latent.space) for full show notes and additional episodes.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Transition from Robotics to Language Models

0:45 to 3:30

Discussion on Ashvin's background in robotics and its relevance to language models.

“He was famously opening as his first intern.”

Highlights from NeurIPS and Robotics Insights

3:30 to 6:00

Ashvin shares insights from NeurIPS and thoughts on robotics professionals.

“I think there's a ton of excitement around robotics right now.”

The Current State of Robotics

6:00 to 9:30

Exploration of the current state and future potential of robotics and LLMs.

“So yeah, actually, I was pretty burnt out for my PhD and I was like, okay, I'm going to go to this chill research lab.”

Challenges in Robotics and AI

9:30 to 11:00

Discussing the challenges faced by robotics compared to language models.

“And a lot of the methods that people were really excited about is like, you know, off policy learning, like value functions, like these kinds of things.”

Insights on AI Achievements and Expectations

11:00 to 12:50

Reflections on AI achievements and the gap between expected and actual progress.

“Those mathier ideas also give you these like kind of implicit knobs to tune that allow you to like overfit.”

The Future of Reinforcement Learning

12:50 to 14:01

Ashvin discusses the future of reinforcement learning and its impact.

“So I think what we had to do is bring the world of economically useful tasks in distribution for RL if we commit to using RL as a tool.”

Generalizing AI Beyond Coding Competitions

14:01 to 16:48

Explore the challenges of adapting AI models for various economic tasks beyond coding competitions.

“But I think in a sense of generalizing beyond coding competitions to economically useful tags, that is it.”

OpenAI's Shift from One Model Fits All

16:48 to 19:42

Discuss the transition within OpenAI from a single model approach to a more specialized model strategy.

“Another conversation that I think has really come to a phase this year is kind of the depth of one model fits all.”

Navigating AI Governance Challenges

19:42 to 22:25

Examine the complexities and concerns surrounding AI governance and decision-making structures.

“Maybe you'd rather just have a thing like the Microsoft board, which is probably all the pensions of the world.”

The Evolution of Reinforcement Learning at OpenAI

22:25 to 24:39

Analyze the development and significance of reinforcement learning techniques within OpenAI's models.

“Yeah, I think in general, OpenAI is really good about having conviction in something and just really, from first principles, going after it.”
Show all 23 chapters

Internal Dynamics of AI Research Progress

24:39 to 27:36

Delve into the internal processes at OpenAI that foster smooth and steady AI research advancements.

“and OpenAI is really good about once you decide that something is good, then you just scale it up all the way.”

Unexpected AI Achievements and Predictions

27:36 to 28:01

Reflect on surprising AI advancements and the discrepancies in predictions regarding AI capabilities.

Conference Insights on AI Predictions

28:01 to 29:06

Learn about the discrepancies in AI predictions and the mindset of experts.

“Yeah, one funny thing that kind of happened is while this was happening, I went to this conference called The Curve, which is about like kind of AI progress.”

Skepticism and Optimism in AI Development

29:06 to 30:02

Explore the balance of skepticism and optimism in AI advancements and future predictions.

“one interesting aspect is that, yeah, I think people still seem pretty miscalibrated in different ways.”

Reactions to Deep Seek and NVIDIA's Role

30:02 to 30:59

Discuss the implications of Deep Seek's revelations on NVIDIA and AI modeling.

Cursor's Unique Position in AI Development

30:59 to 32:34

Understand why Cursor is attracting top talent and its approach to AI models.

“I think it's more like, okay, I'll do the steel man, that side, which is, well, you don't need the top-of-the-line NVIDIA's.”

Challenges of Reinforcement Learning in Organizations

32:34 to 34:08

Examine the challenges of implementing reinforcement learning in large organizations.

“Yeah, I've actually already talked about this so far, I guess.”

Cursor's Vision for Software Engineering

34:08 to 36:43

Learn about Cursor's goals for automating software engineering processes.

“Recently, Jacob Jackson had this blog post about online tab where we're doing policy updates.”

The Role of Internal Tools in Machine Learning

36:43 to 38:02

Discover the importance of internal tooling for effective machine learning practices.

“It's like, you know, it's just like 20, 25 people.”

Continual Learning and Its Implications

38:02 to 41:15

Investigate the concept of continual learning and its potential impact on AI.

“Any example test that maybe Composer doesn't solve yet, but you're really motivated to solve?”

Information Theory and AI's Memory Capacity

41:15 to 42:00

Explore the intersection of information theory and the memory capabilities of AI.

“but also just kind of like in context learning but with infinite memory or something so that you don't...”

Exploring the Capacity of Neural Networks

42:00 to 43:49

The discussion focuses on the theoretical and practical capacity of neural networks and their information storage capabilities.

“It feels like if you could learn enough about those million tokens that you're actually in deployment on, I don't think you should need...”

Interview Strategies in Reinforcement Learning

43:50 to 44:36

Insights into effective interview techniques for gauging RL candidates at Cursor.

“I'm kind of springing this on you, so you can take some time.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Okay, we're here at NeurIps. We're recording a special land space coverage of the folks at NeurIps. and we're here with Asheran from Cursor. Welcome. Hi. Yeah, thanks for having me. So I guess the, like, Asheran from Cursor is like a new identity. I didn't even know if I should say that because you only joined Cursor for three months. Before that, your OpenAI worked in 0103. Before that, Berkeley, PhD in RL, just, but focus on robotics. Robotics, yeah. Is it weird searching for robotics to language models? I mean, this is kind of interesting because a lot of people have been kind of doing this.

0:31I mean, OpenAI is for robotics. I actually was at OpenAI in 2017 also. Where's he on robotics? Yeah, I was interning right before my PhD, where I worked on robotics there. 2017, is that like Japan and that's Japan? He was famously opening as his first intern. Oh, really? Okay, then he might have been before. But yeah, there was like 15 interns. It was a very different company. It was just like Robotics, Dota, and like 15 interns that summer all having pretty exciting individual products. like yeah that set of interns if you look over there now it's kind of cool yeah um but yeah anyone from that class that like you would shout out um like uh there's just like a lot of cool papers that came out like lo pinto now is at uh nyu um uh yeah the uh the person who leads um reasoning at xai i forgot his name uh well he left no eric but yeah i forgot his name but um He worked on KFAC and stuff.

1:29Vision dude? Greg? Not Greg. It was an exciting time to be there. I think robotics is a pretty good fit for LLMs because the switch ends up being pretty... You kind of do similar things. You want to look at a lot of data. It's kind of hard to get stuff working in robotics world. I think it kind of builds very gritty people who look at data a lot, that kind of thing. So, yeah, for whatever reason, I think that transfers. It's happening a lot, and I think it makes a lot of sense. One of my NeurIps highlights so far, I had dinner with, it's like a small group dinner with Lex Freeman yesterday. And Lex used to be in robotics.

2:09And he was like, my assessment of robotics people, robotics people are the best to talk to at a NeurIps because they're most rounded, he says. Because they don't have a choice. They work with the real world. So they look at data. And then the most unhinged, the most detached from reality are the simulation people. I see. Yeah, yeah, yeah. Yeah, I think I agree, yeah. Yeah, and I kind of, I actually did a little bit of both during my PhD. Like, I work in kind of, like, you know, like, prototype ideas in sim and then get them working on real-world robotics. And, yeah, I mean, probably, like, robotics is where you kind of feel AGI'd the least, right?

2:43Because it's just so far away from working. Now, I think, like, I think over the last year, maybe, there's been demos that have been super interesting from, like, physical intelligence and, like, Sunday and stuff that, yeah, I'm starting to be like, okay, this kind of feels sick. Have you ever seen Sunday robots themselves? I haven't seen them. Apparently, they've been doing demos. I'm pretty keen to seeing them. Yeah, I've seen the physical intelligence ones live. And yeah, it's pretty impressive. Just on in someone's living room, folding laundry and stuff. You can just toss it in there. Everybody must be materialized.

3:14Yeah, yeah, yeah, yeah. Okay, and last thing on robotics, and we can kind of pivot to O103. just OmniI is restarting a robotics team. Is that serious? I actually know very little about it because I was in a pre-DiffRZ org. I think it's serious. I think there's a ton of excitement around robotics right now. I'm actually kind of curious what drives it because I don't think I fully understand. There's been crazy raises and stuff recently for robotics companies. I guess my own view on it, and so when I left robotics in 2022, I thought I would actually come back to robotics. But I think my view on it now is that it feels like LLM agents are going to be like a trillion dollar market before robotics is maybe even like a$10 billion market.

4:05And this is just because, so I mean, LLM agents already create value out in the world. Robotics, it's kind of hard to make the case that AI robotics does anything that useful yet? And then once it does something useful, then you have to make the unit economics work out. And I think that's also quite hard. Reliability, these robots have to be fixed and this kind of thing. So I think it's kind of hard. I would say the market is kind of efficient in that the software LM companies are raising tens of billions. And then the robotics companies are raising hundreds of millions. I think this, very recently, it's been like single-digit billions.

4:45Oh, really? Okay. Yeah, so I think that's the maybe surprising thing to me, is that it feels like... It's ahead of where it's actually at. Yeah, I would say that robotics is in kind of like the GPT-1 to GPT-2 era right now. Okay. And I haven't worked on robotics closely. What task would qualify as like, oh, that's the inflection? It's a little bit like you know it when you see it. I thought the Sunday demos were kind of cool. like maybe it's like starting to get there where and the details matter a lot where it it's kind of like it it can't be it has to be in like a new scenario like in one that you haven't seen before and maybe on like generalize yeah exactly and i think that was kind of what gpt2 was too right is kind of like you start to see hints of like cool generalization but um but like uh and i think that's fine like you know it doesn't have to like work out of the box but um yeah i think at this point especially it still feels like in robotics you're not exactly investing in a technology probably you're just investing in a team yeah yeah i'm not in the space whatsoever but like that's kind of my impression it's actually nice when you're not in there because you're like you know as much as like most basically everyone else yeah yeah so we just kind of speculate exactly remind people there's a robotics team at open ai yeah uh so coming back to language models uh did you join 401 or were you like um so i joined right before tatchewt in like um i think september of 22.

6:07So yeah, actually, I was pretty burnt out for my PhD and I was like, okay, I'm going to go to this chill research lab. And then ChatGPT happens and everything kind of blew up and a lot of stuff got kind of refocused. But what, I guess, obviously ChatGPT surprised OpenAI. What did they tell you they were looking for you to do? And then obviously it changed. Yeah, I mean, so I joined on the CodeGen team The Codex. Yeah. The Codex. Exactly, yeah. It was like the team that shipped Codex. But by the time we were working on, like by the time I joined, we were kind of more so working on the model doing tool use and these kind of things.

6:46Yeah. And so like very related to the chat. Like we're kind of like a sister team to the team that made like a chat. Yeah, exactly. So yeah, so we're just kind of working on making the models like smarter, like kind of programming competitions, like yeah, how to do like SIT for that, that kind of stuff. the word and IOI gold has felt reachable in that title. Oh, yeah. Crazy. Like, I think, and this is something I've, like, repeat to people again and again. And these days, like, if you told me that we could have gotten IOI gold then, I would have just assumed that we could all just go on vacation.

7:21Like, you know, it's all over. Like, AI is solved. Like, no point in working anymore. We got it. Yeah. It feels like nothing's... Nothing that much has changed, right? Like, life is still the same. Yeah. Yeah, so I think that's, like, super interesting. Yeah. I don't have a great way to explain it, but I think that's actually what I spend a lot of time thinking about, is why is that the case? You kind of see this again and again in AI, right, with solving chess, and then it doesn't really matter, and solving Go, and yeah, so you keep seeing it, but yeah, I think it surprises you every single time.

7:52Yeah, I think maybe, I think one is we keep moving to the GoPos, we're very good at that, and then two is I think actually our definitions of what constitutes AGI is bad. And we don't actually mean what we say when we say, oh, when we have achieved this, then we have AGI. So like clearly when we have achieved our goal with a language model, we have AGI. It's wrong. Yeah, and I think shifting the goalpost to some extent is correct. Like we keep good-hearting whatever goalpost we have. And I think it's kind of hard to like... To be good-heart is like too negative. It's like I will cheat to do what you asked me to do.

8:28But I don't think it was cheating. just scaling test time compute. At a meta level, I think the community, not cheating, but makes a lot of implicit decisions to go after the evals and benchmarks that matter the most. So Subir is verified, for sure. Yeah, exactly. But, yeah, IOI, hopefully not that good-hearted. Well, but it kind of clearly is to some extent, right? Because most programmers in the world cannot do IOI at any decent level. but we're still struggling to automate most programming jobs or there's a lot left to do. So it's like where language models are here, like junior, senior, dev, and then suddenly for IOI you're like spike.

9:09Exactly. And there's something suspicious about that. I kind of saw this at a meta level also with RL research. So yeah, I did my PhD with Sergey Levin at Berkeley from like 2017 to 22. And that era of RL research was like super interesting because Ayo was like super hyped, right? Like starting from about DQN in like 2015. And a lot of the methods that people were really excited about is like, you know, off policy learning, like value functions, like these kinds of things. And somehow that stuff hasn't really panned out, I would say. And it's not exactly clear why, but in the academic literature, we thought we were making a ton of progress.

9:50And I think in retrospect, I had to say that we probably kind of overfit to the benchmarks pretty heavily. And, you know, how I see this in retrospect is that we gave ourselves a lot of, like, new knobs to tune and then implicitly kind of tuned those to fit the benchmarks. Everyone knew that we were doing that at some level, but I think it's hard to appreciate, like, that it's not just happening for single paper at kind of, like, a meta level for the whole community that's happening too. And I think the result is that, like, I don't know, like, a lot of the RL research that came out of that era I don't think is, like, that used, you know?

10:24And I think it's kind of for a similar reason that basically you were kind of like benchmark maxing. I will full out say there was RL Winter. Like entire startups that were founded based on premise at the time basically gave up. Some of them died, some of them pivoted, whatever. Yeah, yeah. Yeah, I think because I was in academia, there was still quite a lot of excitement over it. But yeah, it still felt quite academic. And yeah, I in that era was a little bit frustrated because I felt like, you know, one of the pitfalls of academia is that it doesn't really reward like simple ideas that work.

10:59And instead kind of tends to reward like kind of mathier ideas. Those mathier ideas also give you these like kind of implicit knobs to tune that allow you to like overfit. while the things that actually work tend to be kind of simple ones that have less knobs and just generalize to many things. There's just less secret sauce to it, apart from just throw a lot of compute in. Exactly. But those are the things that tend to... It's not intellectually interesting. Yeah, exactly. And from an academic point of view, it's like, oh, why am I sitting in school? I think for a lot of people who do PhDs, they're kind of wired in a way they want to think about interesting new stuff.

11:36And yeah, the scaling era kind of like, you know, probably stuck to that. Scaling era. Is scaling era over? I think I've just been paged into that from the ASUSGiver's interview. I don't think it's over, but there's definitely something interesting happening. The thing I was saying about IOI and IMO, I think we'll still continue more or less on the same track. Clearly, these labs are releasing their new pre-trained models and they're still doing much better than before. So I think scaling is still happening, but I think it's happening in a different way. It's worth seriously interrogating why is it that we're not just automating all jobs right now.

12:26I think my view is something like RL, the way it's applied to LLMs right now, is kind of a weird, funny tool where it doesn't really generalize beyond the training distribution that much. It generalizes to some extent and generalizes in interesting ways, but it's very peaky. It can kill the training distribution completely. It can be best in the world at it with not that much effort, really. But yeah, it doesn't really generalize. So I think what we had to do is bring the world of economically useful tasks in distribution for RL if we commit to using RL as a tool. And it might be the case that maybe there's some cool continual learning thing or something that shifts the paradigm next year or something like that.

13:06but it really feels like if RL is a tool then yeah a big thing that needs to happen is like it's not it doesn't feel like intelligence of the models is the bottleneck it's more like you just have products that bring the entire context of what someone wants to do into the product so that the LLM can like see it and then you use RL on top of that yeah have you seen GDPVal? yeah I've seen it Is that basically what you're envisioning? Yeah, I haven't looked at GDP eval closely. I actually haven't seen exactly. Roughly, yeah. To recap, it's 128 tasks across any white-collar job that takes more than 5 % of GDP.

13:50And they basically created all the context, the eval on it, and evaluated every model. Famously, OpenAI's evals to whoever runs that one always finds the anthopics where they're the best. Yeah, props to them for doing that. It's actual science. I think it's good. But I think in a sense of generalizing beyond coding competitions to economically useful tags, that is it. I think that is it. What is more important for GG6? Yeah, what I'd like to do is kind of like, I just haven't read the GDP of all traces closely. It's not clear to me that, you know, what does the job of an accountant entail and what kind of context needs to be in the product so you can actually do it.

14:35They have PDFs. I see, I see. So they try to go as close to source documents as possible. I see, I see. Yeah, I think roughly operating in this kind of thing is what I envision. Yeah, because it can't be an artificial, like, oh, let me clean up this data for you to make it easy for the LLM to process. No. PDF in an agent, go. Yeah, I think that's roughly the right shape of the thing. And I guess how I imagine this being operationalized is that you'd want to co-design the product and the model so that the product, for whatever it is, coding is maybe the easiest first step because most of the context that you care about is just your code base and being able to run stuff in the terminal and that kind of stuff.

15:13And still, we're not that close, really, to automating it necessarily. But for all the other jobs, the context is insane. It's all the conversations you've had with your coworkers, your Slack messages. So at OpenAI, I was working on kind of like hyperparameter scaling research. And I actually wrote not that much code. Like grid search or neuroarchitecture search? No, more like understanding how different, like science of deep learning in like 2020, where it's like, oh, you have to like initialize the layers in a particular way to get good scaling a lot of, kind of the analog for that for RL. The thing is like, I didn't write a ton of code.

15:51So the LM, you know, like writing code is not the bottleneck, But it's more like, over the course of a year, I run sweeps, look at the interaction between different hyperparameters, and kind of build up that knowledge for a year of just different graphs. And to do my job, the model would also need all those things in context to successfully automate my job. And you would kind of want a product that allows you to bring all that context in. Did you have to build it for yourself, or was there an existing one? No, I mean, those graphs are just sitting in my head. Yeah. Right. So I think it would be pretty hard to go automate that job.

16:34But I think what you need to do is build a product that kind of brings that context in. And then you want to RL on top of that to understand, to teach the models to use that context. Yeah. Another conversation that I think has really come to a phase this year is kind of the depth of one model fits all. I feel like the point of the G in the AGI is like one model fits all. I think OpenAI has clearly abandoned that this year. Oh, what did you say? Fiji Simo writing a blog post of the title that we are no longer doing one model fits all. Okay, interesting. And I think Mark Chen or one of the other senior people that are not Sam also saying this in the podcast.

17:16So basically the idea was you started with Codex. Someone else was doing Instruft GPT. then we launched 4, 4, 0 I guess 0-1 and 0-1 was kind of a supposed to be like a reasoning one model fits all and there we merged the 4-0 and 0-1-0-3 line into 5 and now we're splitting it out into 5-5 codecs again. It's like it's just a weird Well, Ophi is very guilty. I mean, you know I don't think you should interpret those as like scientific facts about the universe it's just more like OpenAI has a tendency to ship the org chart, basically. Yeah. Right? The world has the tendency. Yeah, exactly. So I think a lot of it is related to that.

18:00But yeah, I see what you mean by like, yeah, actually, I do wonder if, yeah, like the current reasoning paradigm, the current reasoning paradigm is just kind of fitting itself to this kind of peaky in certain areas thing, right? I don't think it's so much a matter of like model capacity, though. It's just more of another kind of organizational thing that like, if you care really, a lot about coding, you probably don't have the data to do all the other stuff. I don't think it's so much a matter of, if you had all the data, probably you would benefit from just training on all of it, and you'd get some generalization between these.

18:32But it's hard to find one organization that cares about all these at once. So before I double-click on the O-series in OpenAI, I do like to ask OpenAI people who are there, do you have a favorite blip story? The blip was crazy for me. Like, yeah, I was, it was like Thanksgiving. Like everyone remembers where they were, what they were wearing. Yeah, yeah, exactly. I was at Thanksgiving with two opening eye friends, actually. And then one of them on like Friday afternoon is like, oh, like Sam Altman just got fired. We were just like co-working together. It's like, I'm like, what? Oh, ha ha. Good joke.

19:12And then, yeah, it was crazy. And then, yeah, it was just like a crazy weekend of just like ups and downs. like you know we thought yeah like uh like you signed a letter um yeah i did it's like 95 percent people saying yeah yeah yeah um yeah i thought you know like i went to microsoft or well i think maybe i had a slightly more complex like i actually do think that governance feels really important to me yeah uh because it does feel like no matter if we hit agi in like two years or 10 or whatever it's not clear that we have a good structure for the governance of it okay and so it is a question that i think we like probably should spend more time on and i was like during that period just pretty willing to be like you know what like let's forget about the like equity and stuff like you know i think it's like good and healthy to have a conversation about like how exactly the government should work okay you care about this yeah right so now the open i non-profit has this like secret shadow board of members that determine when we've reached hgi yeah yeah better um yeah i don't have a like like maybe i would say i don't have an answer like you know like it's just it's not above my pay grade but like yeah and even even even back then i was kind of like well i don't care i i do care quite a lot um when the blip happened one of my reactions was like well you know this non-profit board stuff like actually if it takes such somewhat surprising, maybe erratic actions.

20:39Maybe you'd rather just have a thing like the Microsoft board, which is probably all the pensions of the world. But it's true, serious people, but also the stakeholders are kind of like the whole world because everyone's through their pensions or something invested in it. Maybe that is a bit more of a democratic way to run things than having seven people run it. But yeah, I don't really know. It feels like we haven't solved governance at all, though, right? Forget AI. Even stuff like unhealthy food or social media, it kind of feels like whatever the capitalistic incentive is doesn't actually capture good outcomes for society, maybe.

21:22Yeah. So about the transition into reasoning, right? You shocked me by mentioning that the reasoning team is 300 people. it's it's kind of like you know now that like you know when O3 was kind of structured as a product like I think it just like kind of gets like larger and larger how many people worked on it so yeah I think I've like lost track of the numbers but yeah like a lot of people contribute to the different aspects of like safety and whatever eval original O1 like I saw the videos like a dozen people yeah even then like if you look at all the contributors it was probably more like 5200 people okay yeah so I mean Let's tell that story from your point of view, figuring out what does RL mean there.

22:07And I guess, was this a branch of any other prior work that you wanted to credit? Yeah, so I think, yeah, like setting the scene, I guess, you know, in like 2023, people were kind of talking about, oh, like is scaling laws dead, this kind of stuff. Every year, every year. Yeah, yeah. But especially, I think especially that year, it felt pretty like serious, you know? Yeah, I think in general, OpenAI is really good about having conviction in something and just really, from first principles, going after it. And I think the people who are most responsible for that is probably Ilya Suskiver and Jakob Pachocki.

22:44I think even Dota was more or less the same template in some ways. And that was 2017. And so a lot of the people there have this like AGI in their bones kind of point of view. And they've basically been convinced that like RL would be the way to get there. So I think for a long, long time, people have been convinced that something like that should work. And it's just that it started to work once the training got good enough. Okay. Yeah. I think human feedback is kind of like a bit of like a side branch because you can't really pour that much compute into it, right? You take the model and you elicit it to be a little bit better in terms of personality.

23:28But the people there were really convinced that at some point it's not about copying the internet. You can go do RL and that's the path to getting much better intelligence. So I think there was a long line of returning to RL in different ways. And then it's just that around 2023 is when it started really clicking. And it's kind of interesting because even, you know, it's not like those initial models performed like way better than the existing models because they're like smaller scale. But people were very good at being like, oh, like this is kind of interesting. Like, you know, the reasoning trace that you see here is kind of not something that you've really seen be so accurate in other models like this one.

24:15And kind of similar to how I think a lot of people didn't really think of GPT or GPT-2 as something that was super compelling, probably. I know that I personally didn't like GPT-2 that much of GPT-2. I was like, okay, whatever. And then GPT-3 happened. I'm like, oh, well, I feel a lot of fun listening to my PhD. It's kind of that where I think it takes a bit of first principles conviction to decide that, oh, this thing, there's something here and we should really go scale it up. and OpenAI is really good about once you decide that something is good, then you just scale it up all the way. Yeah, was there an internal prototype pre-01 that was like, okay, this is the thing, we'll fund it to scale it up, right?

24:55Like there usually is. Yeah, yeah, exactly. What was the thing? What was the demo that really sort of sold it? Just this running RL on even a pretty small model, producing very interesting reasoning traces and getting surprisingly good scores on math. In a way that we couldn't have done without a bunch more pre-training. And then once that looks good, then more and more resources want to just scaling up that new law. And things like adding tool use and this kind of stuff. I think a lot of people make a lot of headlines on the large models, but I think it's very underappreciated, the minis, how well this solution works.

25:41Any comments on just discovery? Yeah, nothing much to say there. I was also not super involved in the mini stuff. I think maybe one thing, not exactly related to that, but it seems like externally people are kind of very like, oh, research seems to come in these big leaps. But I think internally at OpenAI, it feels very smooth. Like you have a bunch of experiments. Some of them have inconclusive results, but maybe you stack them. Yeah, exactly. You stack them and you keep scaling. You keep having different runs that get a little better each time. Okay. So I think that's maybe one other aspect that's like a little underappreciated is that like, I don't know, like in the media, there's just these wild swings between like, oh, it's so like, yeah, exactly.

26:25And I think like internally at Big Labs, it's just kind of like, oh, we're just like chugging along. Like maybe this month is a little better than last month or something, but it's like not as crazy up and down. I think the question is, there used to be more of this, and now I know this less, which is, well, the stuff we've released, we're like, you know, internally, we're like six months ahead. But you say, part of the reason why people, OpenAI wasn't that excited about ChatGPT's launch was because they already had GPT-4. They're like, oh, we'll just put this out. We're already way ahead. I think now people are just releasing things as they have them.

26:59I think, yeah, especially because there's some competitive pressure, right? Yeah, yeah, yeah. I think people are probably pretty worried that if you let a lead linger for too long, that will grab a lot of market share. I don't know. Nano Banana Pro right now is probably pretty good. It's pretty good. A month. So I would say now the internal to external lead time is about one to two months. Yeah, yeah. Which is tiny. Pretty sure, yeah. Tiny. Anything else on reasoning side? I guess you can talk about the work on coding. anything surprised you or like is an external misconception on a 103 side um before we go to cursor well um not really like uh yeah i mean it's yeah pretty cool like uh i think you know it felt already by like maybe early 2024 like oh wow like this recipe like really works and we can see how far we take it and um so i think um you know it was like very steady progress and you know by that point it was probably pretty predictable that we could like you know really like uh smash like things like IMO or IOI.

28:01Yeah, one funny thing that kind of happened is while this was happening, I went to this conference called The Curve, which is about like kind of AI progress. And Joseph Gordon-Levitt. Yes, I went last year. This was before the O1 stuff was released. And I like went to this thing where people were kind of making bets on where we would be on Epoch AI's, like the Epoch AI math exam and like Humanity's last exam and stuff like that. and their estimates were like oh we'll be at like 10-20 % in like 2027 and I think at the time there was like models internally that were like already better than their estimates so it's like off by like two years or something and the interesting thing is like those are also people who are kind of like you know predicting that there would be like Dyson spheres by like 2035 or something okay so like their current estimate is way under yeah they're too pessimistic in the short term too optimistic take the long term?

Read the full transcript

28:58Yeah, well, I don't know. There might be distance for 2035. I don't really know. But I think that is one interesting aspect is that, yeah, I think people still seem pretty miscalibrated in different ways. I do really appreciate how that community makes predictions, though. Because I think most of the rest of the world just kind of cynically says, oh, I saw this the whole time. Yeah. I do appreciate that. Is this EA adjacent? yeah I think yeah exactly it's like that group yeah yeah I like that they like to sort of register their opinions ahead of time and I think like broadly the people who've been you know the capabilities predictions in that group have been broadly correct if you look you know from like 2015 to 2020 or something like where I think a lot of people kind of thought that AI was like a sham or like you know not really going to be that useful for a long time and actually you know it is it's somewhere in the like 2030 ish thing that like it will probably reach like human level intelligence yeah it's weird so like i i i feel like a skeptic when i keep saying like everyone always predicts that aj happens in their lifetime and then it's very convenient for whoever and like we have a consistent view of history where you make see like people in the 1800s and 1900s making predictions it somehow always lands in their lifetime whatever the thing is but like this time it might happen almost surely right like i don't know i'm pretty sure here yeah yeah uh yeah so so it's it's an interesting observation like how different are we from our predecessors in terms of developing our technology yeah um did the deep seek moment this year also this year crazy uh change anything internally uh not really yeah i think that was i think more so just like surprised that um it created such a moment like it was kind of confusing right it was like deep seek shows that nvidia chips are actually more useful than previously thought, and like NVIDIA's stock goes down a bunch.

30:59It was kind of like... I think it's more like, okay, I'll do the steel man, that side, which is, well, you don't need the top-of-the-line NVIDIA's. You can just use the sort of previous generation or the shackled ones they sell to China to do an equivalent amount of work for a model. I see. Yeah, but then it was also, I guess the feeling in OpenAid that like, well, I think we had a better model already at the time, right? So, and it was quite valuable. Like smarter models were clearly quite valuable. So you kind of wanted to be at the frontier. Okay, so I wasn't quite framing this as like a race dynamics thing between labs.

31:42It was just also more like, well, what they write? What their approaches write? They had R1.0, which is kind of a really cool branch. So more like commentary on what we learned about RL this year in 2020. Yeah, well, it does seem like basically a lot of the labs have kind of like converged onto some similar-ish way of doing RL, and they're all kind of back at the same level of like Frontier again, like even the anthropic models, like the Opus 2 4.5, it has this kind of like, there's this like RKGI 2 plot that looks exactly like the OPI ones, right? Like, so I think everyone seems to be converging on a pretty similar form of RL.

32:21Yeah, it's kind of interesting. I think people basically figured out in one way or another to achieve more or less the same thing. Let's talk about the move to Cursor. Why is Cursor accumulating and drawing so many cool RL people? Yeah, I've actually already talked about this so far, I guess. I think from the perspective of Cursor, it's nice not to be dependent on external labs or everything. And I think there's also unique opportunities to co-design the product with the model in ways that we couldn't do unless we actually built the model ourselves and had access to making it good. So, yeah, that's kind of like broadly why Chris is so excited.

33:06I'll push back a little bit. OpenAI has infinity resources, infinity data, has codecs. You could have just stayed. Yeah, yeah. Well, actually, right around when I was leaving is when, like, I think people started actually, like, using Codex a lot. So that was kind of like, it happened right after a lot. So that was kind of funny. So mostly people are using Cursor internally, maybe a bit of Winsurf because it's left over from the previous thing. Sure, yeah, yeah, exactly. So it wasn't that obvious. But actually, I think more to the point, this thing I was saying about, like, RL is kind of a tool that doesn't really generalize that well.

33:44So what you want to do is bring the entire test distribution inside your training distribution. I saw the opportunity to do that at Cursor directly. I think the Cursor folks are also really excited about that kind of vision. And it's just a small place where the product people sit right next to the ML people. And I think there's a lot of potential there. You can kind of see that. Recently, Jacob Jackson had this blog post about online tab where we're doing policy updates. It's every two hours. Exactly, like a policy update every two hours or something. And I think that's the type of thing that, I think it's a little hard to do.

34:22It's very hard to imagine that at OpenAI, for example, just because the product is this kind of complicated thing and also the product people and RLs people are pretty on different sides of the org. I think if you put your mind to it, you would. It's like, tab is an autoconplete, it's a smaller model, it's not as complex, I guess. Um, blah, blah, blah. Yeah, but I don't think that's really this, like, you know, I don't think that's why Cursew was able to do it. It's actually more about just the org itself being kind of smaller and a bit more focused. Yeah. Well, I mean, since you're indulging this, I think the question about continual learning, which obviously is a big theme, it's always been a big theme, it's bigger this year, is, well, don't you need to carry your data?

35:04You can't just, like, chuck whatever your users are doing in straight in because that tends to get you towards the middle of the distribution and actually you want to spike it. I guess it depends how you're thinking about contouring. I mean, I don't know. Humans are quite good about dealing with bad data too, right? You can see someone doing something dumb and decide you're not going to do it. Filter it out, yeah. Yeah, but it's not even actually filtered out. You have presumably some kind of value function that says that if you see someone touch a hot stove, you're not going to go. It's not just filtering it out.

35:36you're actually not going to do it, right? You could rediscover hot stoves on first business. Yeah, but you don't need to. So I think there's something pretty deep there. Yeah, it seems like we're kind of like a few orders of magnitude of data efficiency, basically, away from that kind of like, you do something once or you make a mistake, you introduce a bug in your code, you're not going to do it again. But the models will happily just keep doing it. even within the same context, but definitely, you know, of course, across context. So I think there's something like interesting and deep there is like maybe, yeah, I suspect that it will be kind of like paradigm shifting in the next year or something, but I have no idea like, you know, what it might be.

36:19Yeah. So is primarily you worked on Composer, Tab, and maybe Search? So I've actually just worked on Composer. Yeah. And that's kind of like the main focus of the company, basically, or like the ML group is shipping a better composer. Can you describe, I guess, the impressive, brag a bit about the ML group? Yeah, yeah. I mean, I think the ML group is great. It's like, you know, it's just like 20, 25 people. And, you know, I was like, honestly, pleasantly like very, very surprised at like how good composer is, like given the size of the group. And, you know, it's not like a big research lab yet.

37:00And yeah, I think it's like a really good model. You can kind of see that in the reception. And I think it's kind of the start of hints of like co-design with the product in some ways. Because I think one of the reasons that people really like it is it's smart enough that you actually want to use it. And it's also fast. So you kind of like stay in the loop with the model while you use it. Because I think all the other smart models had this kind of, they're slow that you want to go kind of context switch away and come back. and that sucks as a programmer it just sucks to context switch it gives you ADHD it's really terrible and I think it's one step in the direction of being able to be more sync and I think that's the whole company is just really full of people who want to code, even the co-founders actually the co-founders are often some of the best high taste testers, which also kind of gives you a lot of reassurance that you're going to ship good stuff.

38:02So yeah. Any example test that maybe Composer doesn't solve yet, but you're really motivated to solve? Ironically, I feel like I'm actually a low-taste tester in some ways, because I don't know, I just write slow machine learning code and just think about algorithms and stuff all day. I think more broadly, I'm super excited about co-designing the product so that you can actually not just... Right now we're getting better and better at answering user prompts, and I think that's why Composer 1 is quite good. But what we're really aiming for is more like automate software engineering as a process where you write code, you go look at Datadog, look at what's happening, then come back and maybe have some hypotheses about what's better, rerun stuff.

38:50I think that's the type of thing that we actually want to make them all do. And I do think that Cursor is uniquely positioned to do that in the sense of if a lot of what a software engineer does ends up in the product, I think we can use that to get better and better at not just writing code, but the whole job. Yeah, I think that's very inspiring. Just to double-click on any RL insights, Sasha and me have talked a lot about the internal tooling that you've had for all the cluster visualizations. Is that helpful? Is that what EveryLab has? Yeah, I think the tooling at Cursor is actually really good.

39:28I think because it's just kind of like people are just down to vibe code stuff. They do test their own stuff. So we just have a lot of good tooling where you can have a SSH session into our own user environment or something and see if code runs the way that users got it to run, this kind of thing. I think that's actually quite nice. I think basically one of the big lessons in ML in general is that you want to be really close to your data and understand your data well. And yeah, I think because there's kind of uniquely positioned to do that well. Especially because it's all internal tooling, you're not buying anything.

40:06Yeah, it's just internal. And part of it is just that we're also working on a product where you can understand it really well because it's a code product. Well, if, I don't know, in OpenAI, if I was to look at a biology question, I have no idea what this is about. yeah yeah interesting okay so i think that's a good overview of everything i guess other than the we covered open ai and cursor just interesting rl work that other people are doing that you that you're like still mulling over it's influential to your thinking good papers anything like that yeah you know um unfortunately i've like kind of gotten the habit especially open ai of like not reading that much external work and just like reading people's like slack posts internally as the main way to learn new stuff.

40:54No super inspiring recent things have popped up to me. I do think that this kind of vibe of continual learning does feel like, and there's something super interesting there, and it feels like maybe even in academia, people could make a big crack at it. And continual learning specifically meaning kind of what Tab is doing? Yeah, maybe what Tab is doing, but also just kind of like in context learning but with infinite memory or something so that you don't... Once you experience something in context, it should just be in your weights and you shouldn't have to make that same mistake again. That kind of thing.

41:29Why do you think there's... So it should be in your weights but there's a finite capacity for the weights to remember things. You will forget things if you do that too much, right? I mean, you start out by memorizing or learning from trillions of tokens. Now you're going to experience thousands or maybe millions of tokens and somehow we can... The million tokens are kind of... You only need one epoch. Yeah, exactly. Crazy. It feels like if you could learn enough about those million tokens that you're actually in deployment on, I don't think you should need... I don't think there's a risk of overloading the capacity of your model, right?

42:10Because you can train on a trillion tokens and it's fine. Right, right. So proportionally, it's a drop in the water. Yeah, exactly. The water bucket. Yeah. Unless you run it for years. And, you know, at some point, it's hard. Maybe, yeah. Yeah. So basically, I find it very curious. I've only had one podcast on information theory of language blindness. Like, what is the theoretical capacity? How much are we using? And you should probably track that. Yeah. Yeah. That's a good idea. Yeah. Like, treat the weights. If you want to store things in weights, okay. Treat it as a hard drive. What's the capacity of the hard drive?

42:40How much can be stored in there? We know the capacity. it is the number of bits that I, you know, but the parameters physically cannot store one in that. Yeah, yeah, yeah. And it's, yeah, I've heard that there's this kind of like someone recently at Cursor Jacob kind of brought up this view. I don't know if it's like a more public view. It's like, oh, there's kind of like a hard drive view of, you know, neural networks and kind of like a CPU view of neural networks where, you know, is what's happening the way it's like, yeah, memorizing stuff or is it like you're like having some like few circuits that do a lot of work?

43:13this kind of thing. And yeah, I don't know. Yeah. You know, I would love to, yeah, there's like actually so many of these kind of more scientific questions that I would like love to explore sometime, but then it really kind of conflicts with like empirical stuff, you know? Like, unfortunately, at any given moment in time, it doesn't seem like the most root for like, you know, improving something especially in the short run, but even in the next couple of years is like understanding some of these questions. Yeah, I mean, I guess this is technically supposed to be the role of academia, but it's also hard to explore those ideas there without enough compute.

43:46But yeah, actually, I would love to go at some point and return to exploring these fundamental science ideas. Okay. I'm kind of springing this on you, so you can take some time. What is a good RL interview question that if somebody can answer, they should join Cursor immediately? Ooh, it's a hard question. I assume you do interviews. Yeah, yeah. Well, actually, at Cursor, we do work trials, and it's two-day work trials. that I actually think that that's more representative. Because you plug in and see how they behave. Exactly. So I actually think it's more valuable. This is honestly less of a thing about how you understand RL and a bit more like were you around in the 2017 to 22 era.

44:25But it's like why is off-policy RL unstable is kind of, I think, a good question to dive into. I don't actually know, so I'm going to dig into it. Yeah. Cool. Thank you. That was a great conversation. Do you have any call section? Yeah, I mean, we are definitely hiring a cursor. So if you're interested in working on, especially data and rewards for code, I think that's a huge need. Yeah, please get in touch. Yeah. That's it? Yeah. Thank you. Sweet.

45:08Thank you.

From the publisher

sdw

More from Latent Space: The AI Engineer Podcast

All 247 episodes
[State of RL/Reasoning] IMO/IOI Gold, OpenAI o3/GPT-5, and Cursor Composer — Ashvin Nair, CursorLatent Space: The AI Engineer Podcast
Listen in VO