OpenAI’s IMO Team on Why Models Are Finally Solving Elite-Level Math

30 Jul 2025 · 30 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Episode Summary: OpenAI’s IMO Team on Why Models Are Finally Solving Elite-Level Math

Podcast Overview Podcast Title: Training Data Description: A podcast focused on AI, hosted by Sonya Huang and Sequoia Capital partners, featuring conversations with leading AI builders and researchers.

Episode Overview Episode Title: OpenAI’s IMO Team on Why Models Are Finally Solving Elite-Level Math Description: In this episode, Alex Wei, Sheryl Hsu, and Noam Brown from OpenAI discuss their remarkable achievement of achieving gold-level performance at the International Mathematical Olympiad (IMO) within just two months.

Key Themes and Discussions

Achievement Overview

  • OpenAI’s model reached a gold medal standard in the IMO, an unprecedented benchmark in AI's mathematical problem-solving capabilities.
  • The achievement highlights the rapid advancements in AI, particularly in mathematical reasoning.

Team and Methodology

  • The team behind the success was small, consisting of only three core members.
  • General-Purpose Techniques:
  • The model utilized reinforcement learning techniques instead of traditional formal verification tools, allowing it to tackle hard-to-verify tasks.
  • The progression from struggling with basic math to solving elite-level problems showcases significant improvements in AI capabilities.

Insights on AI Progress

  • The team discussed the pace of progress in AI, comparing past performances to current abilities, with models evolving from handling basic math to engaging with complex competition problems.
  • The model demonstrated surprising self-awareness, recognizing its limitations (e.g., not attempting to solve problem six of the IMO).

Problem-Solving Capability

  • The team emphasized the distinction between solving competition problems and achieving breakthroughs in genuine mathematical research.
  • Problem Six: The hardest problem in the IMO was acknowledged as a significant challenge, illustrating the gap that remains in AI's mathematical abilities.
  • The model's ability to admit when it could not solve a problem marked a notable increase in its self-awareness, diverging from previous models that would attempt to generate plausible but incorrect answers.

Future Prospects

  • There are ambitious goals for future AI capabilities, including solving long-standing mathematical challenges and potentially contributing to broader scientific reasoning.
  • The team remains cautiously optimistic about the future potential of AI in mathematical research, despite acknowledging substantial challenges ahead.

Technical Challenges

  • The podcast touched upon the challenges of scaling computation time for complex reasoning tasks, noting the difficulties in verification as task complexity and required time increase.
  • Multi-agent systems were discussed as important for scaling up computation and allowing for more thorough problem-solving processes.

Closing Thoughts and Future Directions

  • The conversation concluded with discussions on the next steps for the model and the potential for wider application in mathematical research.
  • The excitement surrounding their methods and results indicates a promising future for AI in both competition settings and formal mathematical research.

Key Takeaways

  • The achievement of gold-level performance in the IMO exemplifies a significant leap in AI's mathematical abilities.
  • OpenAI's focus on general-purpose techniques and scaling computation is critical for future advancements.
  • The acknowledgment of limitations and self-awareness in AI models marks a vital turning point in their development.
  • Ongoing collaboration with mathematicians and researchers is essential for furthering AI's capabilities in solving complex problems.

---

This summary captures the essence of the podcast episode, highlighting the discussions, key milestones, and future implications of OpenAI's advancements in mathematical problem-solving through AI.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00The pace of progress is really, I think you see it so clearly in math. And I think Alex tweeted about this where even a few years ago, these models were struggling with like grade school math. And then we, I remember even in 2024 that like GSMAK was used as like the standard EVAL when everybody would release a model. And then it was like math for a short period of time and then it became Amy and then it became USAMO. and the case that it's just gone, blown through all of these math benchmarks, this is really astonishing.

0:47Today we're joined by Alex Wei, Cheryl Sue, and Noam Brown. the trio behind the OpenAI model that just achieves gold medal performance at the International Math Olympiad. The IMO gold is one of the most important milestones in the race to artificial superintelligence, and what makes this breakthrough particularly fascinating isn't just the mathematical chops but the underlying architecture. General purpose techniques for scaling test -time compute and handling hard -to -barify tests that extend far beyond competition math. We've now gone from models that can reason about math for a tenth of a minute just a year ago, to systems that can reason and concentrate on the order of 100 minutes.

1:21The hope for superintelligence is that as we scale reasoning to thousands or hundreds of thousands of hours, we can begin to solve humanity's greatest unsolved problems in math, the sciences, and more. Alex, Cheryl, and Noam joined us on training data to talk about their approach and share some of the behind -the -scenes fun and learnings behind this historic result. Enjoy the show. Alex, Cheryl, Noam, thank you so much for joining us today. We have with us the team behind OpenAI's first gold medal at the IMO. Congratulations to you all. It's a momentous achievement. Thanks, Steve. Yeah, thank you.

1:56I'd love to get into a little bit of the origin story behind this. I know that the IMO gold has just been this elusive thing that everyone in AI has been chasing for a long time. I remember back when Sam pitched us in 2021. It was on the slides, and I remember thinking, oh, that seems really far away. I'd love to understand, you know, the more immediate origin story for this specific effort. When did you guys start thinking about this and how did it come about? Yeah, I think it's like one sort of like something that we've been thinking about for a long time. I remember in my, you know, first week at OpenAI, Noah asked me like, you know, when do you think the model will get I am a gold?

2:35I thought, you know, like it was really unlikely in 2025. But I feel like it's something that's always been on our minds, as you said, like, you know, STEM, like many years ago as well. But this specific effort, I think, you know, we thought it, I think it was really only, like, you know, maybe like a couple months since like, just a couple months. Like the sort of last sprint to like get everything ready for this year's IMO. And of course, we've been working on like improving our algorithms, the ideas for this started coming together maybe like six months ago, but like really like the last push, like you know, we're gonna try to do something for this year's IMO was only a couple months long.

3:18It's amazing. And how big is the team involved? I mean, so we're like you know, definitely building on like a lot of folks work at OpenAI, like this is not possible without like, you know, a lot of help from people from, you know, the people working on like inference in the scaling org, the people who trained the pre -training and the RL training. But in terms of the core team, I would say it's just three of us. So it was a super small, scrappy effort here. That's crazy. Just the three of you. Also, it was mostly Alex. Alex had been working on this technique for a while. And Sheryl and I were happy to help out as we were getting closer to the IMO to make it a reality.

4:00That's so cool. And how does this even come about? Do you self -direct and self -choose? I want to work on IMO gold and I'm going to get us there. How do you even raise your hand to work on something like this? I think it was something that it just felt like maybe it's possible. Maybe if we push a bit for a couple of months, we can just get there. One of the nice things about OpenAI is that I think the researchers are really empowered to do the kinds of research that they think is impactful. And Alex had this pitch that like, hey, there's just a new technique that I think could help out a lot.

4:33And honestly, there's a decent amount of skepticism. I think some people were supportive, but everybody felt like we should give them the freedom to be able to explore this and pursue it. And then it started showing some strong evidence, and I think people still were a little skeptical, but more people were getting excited about it. And eventually it turns into something more substantial. And I think now people are obviously very excited about it. Can you say a little bit more about the strong evidence? What were some of the early signs that you all were seeing that made you lean in? I think it's just like progress on hard to verify tasks, where I think previously we know a lot of art was more focused around just like, if you have these verifiable rewards, what can you do?

5:22You know, we were just seeing more improvement on these harder to verify tasks, is I think what made us excited. Maybe on that front, huh? Did you even verify that the results you had were right? And I saw that you published the proofs on GitHub. But can you just say a little bit more about how you even know that you've discovered the answers? Because my understanding is that they've done a bit differently from how a human might answer them. Yeah, I do think the style of the model outputs is a little atrocious. just... A Trojus isn't the word I was gonna use. Is it creative? Like an alien language?

5:59Yeah, yeah, it's a little, yeah, I think it was, I think you know, we could have, I think you know, it was a very, like, small scrappy effort. And so, we didn't optimize as hard for like, you know, human readability, but that's something that, you know, we know how to do, like, you know, we can, like, we can do the same stuff to like, in the same way that like, you know, chat, GVT, like, is very readable. We can do the same things here. Do you even need to optimize for human readability? Is that even important? I think if you're showing this to humans, they prefer readability. We were actually discussing, you know, we got the proofs, like, okay, because you could actually just, like, run them through a chat chat chat chat chat chat chat to, like, rewrite them in a more readable way.

6:39And it's like, the proofs are still correct, they're just, like, a little bit more readable. And we were like, oh, should we, when he post these online, should we, like, post the more readable version that's, like, run through chat chat chat chat chat or should we just post the raw version? And we decided, you know, I think for full transparency, which people just like post the originals and people will figure it out. You guys have a bunch of IMO medalists and participants in the staff at OpenAI, right? Do you guys like moonlight in your spare time grading the answers that the model produces? Like during, I mean during like you know, the testing, like you know, yeah, we like we read a lot of samples.

7:13But look for grading these specifically, like we hired external former IMO medalists. So each proof was graded by three medalists and for each one they reached unanimous consensus on the correctness. I should also say that, for me, I don't know about Cheryl, but for me, the proofs are beyond my ability to comprehend. I was a math major and I never really did competition math. I already, the stuff that this model is writing about is beyond my ability to grade. Yes, I think that's what makes it even more amazing, just how smart the model is. Totally. What about problem six? How come none of the models at this year's IMO had a solution and your model didn't even attempt problem six?

7:57Can you say more about what makes that problem? And traditionally problem six is always the hardest at the IMO, is that right? Yeah, I think problem three or problem six usually. Okay. Why just say a bit more about what made problem six different and what you learned from, you know, I think you tweeted that the fact that your model knew that it couldn't solve problem six was one of the things that gave you hope. So just say a bit more about that as well. For problem six, I mean it's It's just a really tough problem, I think. If you gave me months to think about it, if you even gave me a big hint about the main idea to solve problem six, I don't think I'd be able to get there.

8:31It's just crazy, tough problem. There's so many things you can do, and there's very narrow path to finding the proof. And I think it's one of those things that I think math is just hard. Yeah. And we threw a lot of compute at problem 6. But I think it was good to see the model doesn't try to lose Senei or try to just make up some solution, but instead we'll say no answer. I mean, it is kind of disappointing when you're like, it's done so much work. Just to say no answer. But I think it's good that actually acknowledges that. Yeah, that's an amazing level of self -awareness of your own kind of ceiling.

9:12Because I remember at least a couple of years ago with these models, they would always try to be helpful and make up an answer, right? And so to see this is just like, I think an amazing level of software and this from these models. When we released the reasoning models, I talked to some professors at mathematicians, computer scientists, and I was asking them, like, you know, are you finding value in these models? And the answer was, you know, frequently yes, but the one thing they would complain about is like, if they would ever ask the model a question that it didn't know the answer to it would just like output a very convincing but wrong answer.

9:45And they would have to like go through it very carefully to figure out was it exactly correct or was there like you know some flip of an inequality or something that the model snuck in there. And it's nice to see that this model like if it doesn't know it will just like acknowledge that it doesn't know at least more frequently. I guess internally did you guys have like a betting like a poly market or something going on whether you guys were going to win I am a gold this year and like what was the internal vibe. I think we felt like we had a strong shot. But I think we also thought that it wasn't like a lock where there's definitely a distribution of questions where the models probably struggle more than a few minutes.

10:26But then you know there's another distribution of questions where the models would be really really strong. And I think this year was somewhere in the middle where you know, like problem six, like I think it's just out of reach of state of the art models today. And I think maybe in general, like, you know, like these hard, like, combinatorics problems, which problem six was, I think, more challenging. And that's still something that the models struggle with. What is it about combinatorics that makes it challenging versus, you know, the, like, geometry, for example, which seems like you guys do well at?

11:02I think with commentarics it's probably because it's a little more abstract, a little more high -dimensional. And I think oftentimes, like, commentarics problems sort of require like leaps of faith, or leaps of insight that, you know, the models are ex -good at, I think the models are more good at, like, you know, problems that require like a bunch of smaller steps, for example. was that from your guys perspective? Was the internal vibe optimistic or not that you all were in the end goal? I feel like it wasn't super optimistic. Like I think they definitely knew that like it could happen. But I think like even like a month for like two months back, it definitely felt like it would have to like improve quite a bit, which I guess we did.

11:46I remember I was talking to another researcher at OpenAI, like maybe two months before the competition and we were like, you know, saying like, okay, if we were to bet, you know, I'm a betting man. I have to have happened to bet. Yes you are. And I was like, what odds would you take? Because I was willing to bet. I'm like, we were going to get gold here. And he was like, there's really no chance. And you know, and you know, he said that he would gladly take like two to one odds against like the model winning. So like, you know, less than one third chance. But he didn't want to bet against us. So, you know, he thought it would be bad vibes to bet against the team winning.

12:24So he didn't go for the bet. So I do make some pocket change, don't I wish I wish I wish I mean you need it so because I mean you guys were I think you tweeted 12 % on Amy like 15 months ago right so it's even though you want never want to bet against scale and opening eye it's just it's just astounding slope of what you all have accomplished here. The pace of progress is really I think you see it so clearly in math and I think Alex tweeted about this where you know even a few years ago these these models were struggling with like grade school math. And you know, and then we, you know, I remember even in 2024 that like, GSMAK was used as like the standard evil when everybody would release a model.

13:06And then it was like math first, we're up here at a time and then it became Amy and then it became USAMO. And the pace that it's just gone, blown through all of these math benchmarks is really astonishing. Yeah, I remember training a model on GSMAK two years ago. Yeah, we'll pass those days, huh? that's where the e -vails, what's next? Do you think, I mean, at this point next year, you think we'll be solving Millennium Prizes? I think those are still very far away. I think on one hand, you think about how much math progress has been made since like GSM 8K, which is like, you know, just like two years ago, was sort of a standard that people were trying to push on.

13:46You know, that's like an astounding level of progress, but also you think about like how much time it takes for people like, you know, GSM -8K problems, they're like, great school math, you know, it takes someone to get a math like a couple seconds. And now we've gone from like a couple seconds to something that takes like, you know, these brilliant students an hour and a half per problem on average, you know, the IMO's three problems for an a half hours. And then, you know, research math is gonna be, like, you know, these same, you know, brilliant students, They've grown up their researchers and it's going to take them like 1500 hours So there's like you know a thousand acts of like more thinking time and then Millennium prize problems have taken entire fields like you know people's lifetimes Of thinking and you know, we still don't have much progress on most of those and so it's On one hand like you know Super exciting that we've made so much progress one other hand It's sort of also like humbling to see like how much further, you know progress has to go from like an hour and a half to like no tens of thousands, hundreds of thousands of hours of human thinking.

14:55Totally. No, I think you deserve a lot of credit for seeing the future on this. I remember you visited us before you even joined OpenAI talking about the results from gameplay and you know what happens if you let a model think for hours and tens of hours and credit view you've really seen the future on this. Thank you. Yeah, I mean it's exciting to see it actually happen. What are the hard things that happen as you scale compute time, inference time from the order of 0 .1 minutes to the order of 100 minutes? I guess at a high level, because most of our listeners are not AI researchers. But what are the hard things that happen to keep the model on the rails, so to speak?

15:39I don't think I think we can point to you. It's like pretty clearly a challenge is that if you have the model thinking for like 1500 hours, then in order to e -valid, you have to have it think for 1500 hours. And so eventually, the evaluation of the models becomes a significant speed bump on progress. So we're not really at that point. If we have the model think for an hour and a half, it's no big deal. We can run those tests. But run a test for the model is thinking for a month. It takes a month to finish that test. And so progress can only advance so fast if you want to wait for those kinds of results.

16:16I think both of you are on the multi -agent team. Help me understand what the role that multi -agent systems play in this is. Yeah, so in addition to having the model think for a very long time and make a lot of progress on hard to verify tasks, this also involved scaling up parallel compute. And so there's a multi -agent component to that. We're probably not going to be able to go into too much detail about the exact techniques. But that was certainly one way that we were able to scale up test on compute for the IMO. By the way, one thing I'll add for the multi -agents scaling parallel compute thing is that the way that we did it, we really tried to prioritize generality in our techniques.

17:02I, for example, like, you know, I worked on AI for Poker. Alex and I actually both worked on AI for diplomacy. So Alex was on the team that worked on Cicero. Cicero, yeah. And, you know, those were projects that I'm really proud of. But they were also projects that we spent years working on to, like, achieve that result. And with the pace of AI progress being so fast, it felt like that wasn't the best use of time to, like, develop a very bespoke system that could only do that one task. And so we all like really prioritized general purpose techniques and all this and you know the techniques that we used for Everything for scaling up the thinking time for Working on hard to verify tasks and for the parallel compute are all general purpose techniques that we're either planning or have used for you know other systems as well And is that the reason you all chose not to do this in lean like my understanding is the official kind of IMO AI track was was a lean interpretation this year.

18:03Is that why you guys chose not to go with lean? Yeah, that's right. I mean, there is certainly, I think there is a lot of value in lean as a tool. You know, mathematicians find it useful, for example. But the priority for us is really general purpose reasoning capabilities. And lean has its limitations. And so that's why we wanted to prioritize natural language. My layman's understanding is lean is a formal verification tool. Does your result here basically say that like informal verification with scale can, you know, can perform at the same level or even surpass formal verification? Is that the right takeaway?

18:37I wouldn't say you. I would not say that's the right takeaway. I don't know Alex, you have thoughts. I see that the user just like, you know, sort of two like orthogonal sort of components here. We're like, I think, you know, I think we found the informal math sort of an interesting problem because it represents a kernel of difficulty around scaling up test time compute, hard to verify tasks. That represented something like difficulties from a very broad set of tasks that we were interested in from a general purpose standpoint. I think lean is a little bit more narrow where a lot more of the world can be approached with informal reasoning than is formalizable.

19:22I don't think there's anything wrong with narrow AI. Nero AI can be very effective and obviously far surpass general purpose AI in certain domains. And I think the right way to think about it is in the same way that humans, human mathematicians find a lot of value and lean. General AI can be compatible with a more narrow system that's focused on formal mathematics. and the combination I think can be better because of it. I think I saw on Twitter from multiple folks at OpenAI and I think you guys have mentioned this as well, that this system was built with a very similar approach and infrastructure to many of the recent launches from OpenAI, like we had Issa from the ChatGPT agent lunch on the podcast last week.

20:14Can you say a little bit more about what that similar kind of foundation approaches? I think infrastructure -wise, we all kind of just use the same infrastructure. But I think as far as the core of this question, no, Malik said there's nothing that's very bespoke to IMO here. And the hope is really that we can use the techniques that Alex worked on as far as nonverifiable tasks and as far as just scaling up test time compute and being able to apply this to other areas of reasoning. or other areas of model capabilities in general and just build stronger models, keep improving agent, keep improving chat GPT and everything else.

20:55Tell me about the actual experience of IMO Day. What was it like? Yeah, I mean, we were waiting for the problems to come through because once the participants finished the exam, then they get posted. And so we plugged the problems into our model. and that was around like, I guess, pretty late at night, maybe like 1 a .m. or something. And honestly, I went to sleep because it's like, you know, it's 1 a .m. I'm not gonna stay up for four and a half hours to like see the output. I'll just like go in the morning and see. But I think these two actually stayed up and got to watch the model and see it come in in real time.

21:32It was a lot of fun. Did anyone call them? I'm like, wake up, wake up, we got this. There were a couple of moments where Alex was so exhausted that he decided to take an app, but we told him, OK, just make sure your phone is on silence so that we need to wake you up and call you. And at one point, we did actually have to call him, but you don't think you work on it. That's awesome. It must have been such a thrill and such a high, especially for that to come up through it. So you started at 1 a .m., so you must have known 9 a .m. then? Oh, it's 1 .5 hours. For however. So it's like, yeah, I don't know.

22:08I mean, we can kind of see the problems come in. So I'd just be making sure the systems are staying stable. And Alex is over there reading and seeing what they're not, how the model's doing. So you were doing the live human proof checking to see if it was actually. I was naturally very anxious about the results. So I was just looking at the, you know, like the partial progress the model was making you can sort of like we can sort of observe that. And then like, you know, I also like, you know, hand check things. But like, you know, we were, we were going to send these out to the graders, but I was also just like, can't check them because I was so curious.

22:49Well, call me next time. I want to come hang out there for now. I'm not gonna sleep. That sounds awesome. What are the cool things about these models is like, you know, I can't understand the proofs. But when you see the model like thinking about it, it will express its uncertainty or its confidence in natural language throughout the process. And it will just kind of say words that will like hint at its like, you know, if it's like really confident that it figured out, I'll say good a lot, you know? And if it's like unsure, it'll like throw in a lot of question marks. And so it's like cool that I can kind of follow along and see how the model is like, you know, feeling about about its progress, even though I can't really tell if it's like, got a correct or not.

23:29It's like the dreaded seems hard.

23:34You got that in problem six. Scott got a lot. No progress hard. Same thing. Keep going too bad. Wonderful. I guess looking ahead, you've gotten the pinnacle result in competition math. I guess you can go to Putnam next year, but you're basically at the top, right? And so what's next? Yeah, so actually for Putnam, the problems, I think, since the exam is, like, less time per problem than the IMO, and it's a little more knowledge -heavy, We actually found in our e -vows that the model was really, really good at putting them problems, like better than it was at IMO problems. And so I think the frontiers here are really not about these very time boxed competition problems anymore, but it's about problems that really take longer periods of time and more deep thinking to solve.

24:32It's really cool. Okay, so you're going to start proving novel theorems now? I think there's this intimidating gap between these time box competition problems and a real research breakthrough which takes a year's worth of work and a year's that's on the order of like 1 ,500 hours instead of 1 .5. Yeah, totally. I guess, relatedly, I was listening to the Demis podcast last night and he mentions that You know, the hardest thing is actually coming up with the interesting problems to solve. And I'm curious if you all agree with that. I think there's some truth to that that, you know, these models are really good now at solving these problems.

25:19Coming up with them is, you know, still a challenge. But I think it's also worth noting the incredible piece of progress that we're seeing. And there's always a next hurdle. And originally when LMS came out, it was like, well, how do we get them to reason? And then we got them to reason, but then how do we get them to reason on hard to verify tasks? And now they can reason on hard to verify tasks. And I think the next hurdle is going to be like, okay, well, how do we get them to come up with these novel questions? Like even creating an IMO question is a challenge. And it takes a lot of extra mathematicians, a lot of work to do that.

25:58But I don't see any fundamental barriers that block us from getting there. I love that. Do your results in math? Do they just fully generalize to, you know, you're just going to be better at scientific reasoning, you're going to be better at general reasoning? You know, does being great at composition math make you, you know, be great at everything else? I think how we approach this was not like, you know, we should be like, you know, great at competition math, but really, I think it's like we were focused on like developing like general purpose techniques to make a reinforcement learning better.

26:39And I think those we are very excited to like improve our models in other domains beyond math. And so, and hopefully like make models more useful for like, you know, us in like everyday usage. This is like a, you know, it's a pretty late breaking results. It's honestly, it was surprised even to people internally at OpenAI. And so, the next step is to incorporate this more broadly into our models and, you know, improve the recent capabilities across the board. But, you know, it's going to take some time to go through that process and deploy it to the world. So, I think it's going to come. But, yeah, I'll just take a little bit more time.

Read the full transcript

27:21Is it harder for these models to do the IMO or the physics Olympiad? I think definitely the physics Olympiad because the physics Olympiad has I think like an experimental section or you have to. We need silverbond expi. I didn't realize that. Okay. I thought it was just done with piece of paper. Yeah, so I think it's the one will probably be good at the on the paper part, but yeah, I think we'll be a little bit of time before it can do the experiments. Not with a world model. Okay, cool. Are you going to release this model for customers to play with? Reloaf's son is a math Olympiad kid and he's like, I want access to the Math Olympiad model.

28:06Will people be able to play with this? So we want to make this accessible to mathematicians to use. We're still trying to figure out the exact details of how we make that happen. But I think it's really cool that we've developed this system that is incredibly good at math and it makes sense that we want to see what mathematicians can do with it. I've actually already been emailing with the Stanford professor, mathematics professor. He actually emailed me about a year ago before we announced a one and he was like, hey, do you want to do a collaboration on solving hard math problems? And basically what I told him is like, I think we just got to advance general reasoning capabilities and eventually they're going to be able to help you with your hard math problems.

28:42And I think that's actually the most promising route to getting there. He was a little skeptical. But every model release, every reasoning model release, he's emailed me with a follow -up and is like, can it solve this problem now? And I've been plugging them in. And I don't know what the output is, but I email it back to him. And he says, yeah, that's wrong. And he emailed me a follow -up this time with the same problem like asking, hey, can it solve it now? It still can't solve it, but at least this time, it recognizes that it can't solve it. So I think that's a big step. But we're curious to see if there's like a lot of other problems out there that mathematicians Want to challenge this model with and see if it can take them on Amazing Congratulations to all I think this is a momentous result that the entire field has been waiting for for a very long time and The fact that it was accomplished by a team of three people in a span of two months.

29:32It's it's extraordinary Congratulations, and thanks for joining us on training data Thank you. Thanks for having us

From the publisher

In just two months, a scrappy three-person team at OpenAI sprinted to fulfill what the entire AI field has been chasing for years—gold-level performance on the International Mathematical Olympiad problems. Alex Wei, Sheryl Hsu and Noam Brown discuss their unique approach using general-purpose reinforcement learning techniques on hard-to-verify tasks rather than formal verification tools. The model showed surprising self-awareness by admitting it couldn’t solve problem six, and revealed the humbling gap between solving competition problems and genuine mathematical research breakthroughs.

Hosted by Sonya Huang, Sequoia Capital

More from Training Data

All 110 episodes
OpenAI’s IMO Team on Why Models Are Finally Solving Elite-Level MathTraining Data · 30 min
Listen in VO