How OpenAI Beat Every Human Team at the World's Hardest Coding Competition

1 Oct 2025 · 53 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The Neuron: AI Explained - Episode Summary

Episode Title

How OpenAI Beat Every Human Team at the World's Hardest Coding Competition

Episode Description

In this episode, hosts Grant Harvey and Corey Noles interview Ahmed El-Kishky, a research lead at OpenAI. They discuss OpenAI's historic win at the International Collegiate Programming Contest (ICPC), where their AI system solved all 12 problems, outperforming every human team. The conversation reveals insights into the AI's capabilities, the blend of GPT-5 with experimental reasoning models, and the implications for the future of programming and AI-assisted science.

---

Key Highlights

OpenAI's Historic Victory

  • OpenAI's AI system triumphed at ICPC, solving all 12 problems and marking a first in AI defeating human teams in such a high-stakes competition.
  • The AI's performance surprised the team, especially its ability to solve difficult problems that human competitors could not.

Behind the Scenes

  • The team did minimal preparation, akin to a hackathon, with three core members participating on-site in Azerbaijan while others supported remotely.
  • Initial submissions were correct for the first 11 problems, but the team faced intense pressure to solve the final, more complex problem, which took additional attempts.

Technical Insights

  • The AI combined GPT-5 with an experimental reasoning model, showcasing how machine learning can enhance reasoning capabilities.
  • The experimental reasoning model uniquely tested its solutions against simpler brute-force approaches before submitting final answers, demonstrating advanced problem-solving strategies.

Competition Structure and Experience

  • El-Kishky described the atmosphere as akin to competitive sports, with a community of researchers and developers engaging in the event.
  • The competition format mirrored traditional programming contests, but also allowed for heuristic problem-solving, demonstrating the AI’s adaptability.

Implications for Future Programming and Science

  • The conversation highlighted the potential for AI to assist in automating scientific discovery and solving complex problems beyond programming contests.
  • El-Kishky emphasized the importance of AI models developing reasoning and tool-use capabilities to contribute meaningfully to research and problem-solving in various fields.

Reflections on Education and AI

  • The hosts discussed the evolving role of education in light of AI advancements, advocating for continued learning of fundamentals even as new tools like Codex emerge.
  • El-Kishky noted that while AI can aid in problem-solving, the foundational skills of reasoning, coding, and persistence remain crucial for programmers.

Future Directions

  • OpenAI aims to leverage insights from competitive programming successes to enhance AI’s capabilities in broader domains, particularly in scientific research.
  • The potential for AI to collaborate with humans in competitive environments or in solving real-world problems was a central theme of the discussion.

---

Key Takeaways

  • OpenAI's victory at ICPC marks a significant milestone in AI capabilities, demonstrating advanced reasoning and problem-solving skills.
  • The collaborative approach within the OpenAI team and the combination of different models played a critical role in their success.
  • There is a strong belief that AI will continue to evolve, contributing to fields such as science and mathematics, and further changing the landscape of education and competitive programming.
  • Future competitions may blend human and AI participation, evolving into collaborative problem-solving events.

---

Conclusion This episode of The Neuron offers a comprehensive look at how AI is reshaping the landscape of programming and scientific discovery, highlighting both the achievements of OpenAI and the ongoing challenges and opportunities that lie ahead in the field of artificial intelligence.

For more insights and updates, don't forget to subscribe to The Neuron and check out their newsletter at [The Neuron Daily](https://www.theneurondaily.com/subscribe).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00We didn't tell it to do that. It had never done that before. Can our models actually discover some new knowledge? One of the problems that no human contestants solved, our models solved. GPT-5 tried seven times to solve that last problem. It was just like something else. It was mind-blowing. They'd write solutions that were so bad that the little virtual computer, the sandbox would give it, would just crash.

0:27Ahmed El-Kishky:The world's most elite programmers trained for years to compete at ICPC, the Olympics of coding. This year, OpenAI's system solved all 12 problems, outscoring every human team in the world finals. Today, we're joined by one of the researchers that made that happen. Welcome, humans, to the latest episode of The Neuron. I'm Corey Knowles, joined as always by Grant Harvey, writer of the Neuron Daily Newsletter. This episode is brought to you by Whisperflow. More on them to come later on. But today, we are chatting with Ahmed El-Kishki. Ahmed is a research lead at OpenAI, where he focuses on advancing reasoning, tool use, and programming capabilities in AI models.

1:04He holds a PhD in computer science and has worked at Twitter and Meta on large-scale data mining, multilingual NLP, and cross-lingual representation.

1:14Ahmed El-Kishky:Most recently, and what's brought us together today, though, Ahmed was a part of the team behind OpenAI's historic win at the International Collegiate Programming Contest, where their system solved every problem in the World Finals, marking the first time an AI has defeated all human teams. Ahmed, welcome to the Neuron. Thank you for having me, Corey and Grant. First off, congratulations. How did it feel when the moment that the system cleared that 12th problem and you realized, you know, wow, we've just made history? Is that what you expected? Actually, we weren't expecting to get all of them right.

1:47So let me give you a quick like summary of like how these competitions go. So we generally don't really like prep for them too much. Usually it's a, hey, like if we have enough time, a group of like two to three people can like form a ragtag team and just try to get something like this going. So in this case, we had like three or so people from my team. They flew out to Azerbaijan, which is, you know, way on the other side of the world. We had somebody like, you know, back home here, making sure everything was up and ready systems wise. and yeah like they were almost like hackathoning it they were just like up you know sitting there with everybody in the competition um they had the same five hour limits uh they were just like eagerly looking at the problems like as you know the model came back with a solution they'd like you know eagerly submit it and uh the cool part is like the first 11 that were submitted like Everyone was correct on the first try.

2:48Amazing. So it was mind-blown. We didn't even think that would happen. Yeah. And then we got to the last problem, and the last problem was so difficult. We kept submitting it every 10 or so minutes, and it would be wrong. So it was a little bit nerve-wracking. It already, like four hours, it passed. It had already gone to the last hour of the competition. And we're like, okay, we're probably like 11 out of 12 is pretty good. That's already really great. We were looking at the scoreboards. And then like in the last like little bit of the competition, we saw that answer accepted. And it was just like something else.

3:23It was just it was mind blowing. Just seeing all 12. It was like really a pivotal moment for like just AI, but also like all the hard work we've been doing so far.

3:33Ahmed El-Kishky:That's that seems so exciting and fun. Like like just like playing high school basketball or something. I know. I was going to say I could see the sports movie like of this and playing in my head, you know. definitely yeah biopic actually it's kind of interesting um we've been doing like you know this year we've done several of these programming competitions and it is almost like sports i've never actually like you know closely followed competitive programming like i tried it out in college and uh in high school um but like we competed in japan at atcoder uh earlier this year and uh it was obviously on like japan time uh and we were trying to see how well our models could do on like these new types of programming competitions where there isn't really a correct answer.

4:17There's only like how good you can do, how well you could do. So the competition was like up until like, you know, three, four in the morning. It was a 10 hour competition, which is crazy. Oh man. But everybody just like stayed up all night. Like we had this like channel on Slack and like hundreds of researchers were just following along. It was in Japanese, so we couldn't understand it. but we could hear like OpenAI's name being like yelled out but people were just like following along, looking at the solutions. It really felt like a sports game or like a high school basketball game.

4:50Ahmed El-Kishky:How did OpenAI get into doing these competitions in the first place? I know this is the second big one we've seen fall this year, right? This was the third actually. Third? Yeah. So I'll tell you a little bit about it. I joined OpenAI a couple of years ago and like programming competitions were like my first big project I was working on. When I first started, like I was like, okay, like maybe this is like a way to sort of see how good our models are. And there's some like hint of truth to that. When I joined GPT-4, I'd just been released. It was kind of abysmal at these competitions. It was, I think like 300, 400 ELO, if you know chess terminology, which like places it solidly at like sub-novice.

5:36somebody that just like you know just started um so it wasn't really good at all and it was kind of in fact quick funny story when we try to actually try these programming competitions problems they were so difficult and the model was so bad at them they'd actually kind of like crash the computer and so they'd write solutions that were so bad that the little virtual computer the sandbox would give it would just crash it ran out of memory it would just hang and we'd have to like, you know, delete it. It was that bad. But a lot of people at OpenAI really come from the competitive programming community.

6:13From like high school, they would do Olympiads, international like informatics Olympiad. In college, they would do ICPC. They would participate in these competitions like Google Code Jam. And so there was this like deep seated community of competitive programmers. And I think fundamentally they had a belief that competitive coding was a really good, I don't want to say benchmark, but a good way to sort of measure how well you were improving at reasoning. So when we saw that it was really bad at them, many people were just like, hey, wouldn't it be great if our models could compete at the same level as like some of the brightest people in the world, some of the brightest students, competitive programmers, that would be a good measure of like progress towards reasoning.

6:59So yeah, like we have Yaka Pachoki, who is our chief research scientist. We have Mark Chen, who is our chief research officer. And they really like heavily pushed for, you know, starting this work stream. That's awesome. Yeah, as you mentioned three. So is that does that include the International Math Olympiad or is it three programming competitions specifically? Oh, yeah. So three programming competitions and then one math competition um got it got it yeah so the first one we tried this year uh was at coder which is like this premier competitive programming competition that's held in japan every year the world finals there um we decided we wanted to mix it up there we didn't compete in the traditional like algorithmic competitive programming uh where there's usually a correct solution a solution that's you know is able to answer all the questions within a certain time complexity under certain memory constraints.

7:50That's like the traditional programming competition, which is like ICPC or IOI. This one was a, it's called the heuristic competition. Heuristics are necessary to solve like these problems that they're impossible to get a correct answer to. Like finding the best answer is just fundamentally possible. Like theoretically, it's difficult to do. It's super difficult. And so the best that people can do is writing these heuristics. They're programs that sort of find better and better solutions, but there's no idea of like a best solution. And so in these competitions, we were, yeah, we're trying to find like, you know, really good solutions.

8:33The model has to be creative because it isn't like the traditional, this is correct, this is incorrect. The model just have to, you know, submit solutions and then looking at like how well it did, it would have to then one up itself. So it was really in a competition against itself to sort of get better and better solutions. And we'd never tried these out before. And so we wanted to sort of see how we did. One example of these competitions or these problems are like games. Sometimes some games, there's no like, you know, best, you know, game or anything or best program, but you can sort of just get better and better at it.

9:05Ahmed El-Kishky:Okay. So let's talk about a tool that's completely transformed how I personally work. That's Whisperflow. Imagine being able to write full articles, emails, even take complex notes just by talking. That's what Whisperflow lets me do. It's hands-free writing that's smart, accurate, and ridiculously fast. For me, it solved a decade-long problem. I used to cover baseball as a beat reporter and had to file my story deep in the middle of the night. I always wanted a quality dictation tool that would let me get started on the way home. I wanted to be able to just talk my ideas into a dot. I tried several tools, they all came up short.

9:43Ahmed El-Kishky:But with Whisperflow, I can dictate an entire piece, have it cleaned up, and file all while I'm on the go. No more waiting to get back home, fire up a laptop in the middle of the night to rush something out. I was able to take time that I already had and use it. But it's not just about accessibility, it's about productivity. You'll save hours, get your thoughts down instantly, and stay in flow without breaking a note. So whether you're a writer, a founder, or just someone who needs to capture ideas quickly and accurately on the go, Whisperflow makes it effortless. Seriously, you're going to want to check this one out.

10:16Ahmed El-Kishky:You'll wonder how you ever worked without it. It's available on Mac, Windows, and iPhone. Visit whisperflow.ai slash neuron today and get started for free. That's whisperflow.ai slash N-E-U-R-O-N. Tell them the Neuron sent you. Well, then exactly like how do you prepare for a contest like that? You know, like what was happening behind the scenes, you know, as you're leading up to this? And, you know, if you could tell us a little bit about who else is on the team. I know a couple of them like Mustafa and Boris, but yeah. Yeah, yeah. I'll talk a little bit about ICPC. It's honestly just a team sport here.

10:53I don't even want to say it's just like a handful of people. um it's the byproduct of like you know pre-training our models you know and then doing like rl on them in general to make them really great reasoners really great tool use um uh and then afterwards you know there's so many people that contributed different ml aspects to the models one of the experimental reasoning model we used was a byproduct of the imo efforts by Alex Way, Cheryl, Noam. So there's so many people sort of playing a part here. But the core people that actually went to Azerbaijan, it's almost like a volunteer experience.

11:32We're just like, who wants to try this out? And Mustafa was one of them. We had Robin, we had Andrew, and Boris was helping out from London. So they formed almost like the core team of people that were, you know, driving this. A lot of it was sort of making sure the models were ready to go because there's no room for error here. It's a five-hour competition. You basically, it's live. So it's not like, oh, we're having technical issues. Let's push us back by an hour. We were like in the same situation as a student. People were nervous. They get the problems in the same PDF format. And so everything had to be just ready to go.

12:12And so they spent like, you know, a week before the competition, getting things up, trying it out, making sure that like things were going, doing a dry run, you know, spending some of the like spending long nights trying to make sure everything's working and then just flying out. And just, you know, after that, it's, you know, hope that you prepared well enough and the models can take it from there.

12:33Ahmed El-Kishky:You know, I understand you guys combined GPT-5 with an experiential reasoning machine system. What makes that a powerful combination for ICPC? We wanted to see how well models that are available to the public can do. So GPT-5, we just released it not too long before. And so we wanted to really see how well could these models, given enough compute and attention, actually solve some of these problems. And so it was kind of important to us to try out GPT-5. The experimental reasoning model is stuff that we worked on. We're always, you know, trying out, uh, to trying to improve our models. Like we're always running, uh, new experiments, training larger models, uh, more RL and seeing how far we can push reasoning.

13:24Some of these models never end up making it, uh, into like chat GPT. Um, but we're always learning from them. Uh, so in this case, we knew that, you know, GP5 can go a long way, but we wanted to, you know, not just, you know, limit it there. We wanted to also test how well our reasoning models perform. And so we tried out both. There's no reason to do one. We were looking to see a system that could push the frontiers of these competitive programming competitions. And so it was a very simple approach here. It was just like both GPT-5 and the IMO model, the experimental reasoning model, would take each problem and they would try to solve it.

14:02They would try to solve it multiple times. They would give like a collection of solutions. And then the experimental reasoning model would be like, I like the solution the most. Let's submit this one. Wow. That's awesome. It was a good model or it was a good approach just because we had two models. So the two models, you know, they have some diversity there. They think a little bit differently. So that helps. But the experimental reasoning model is just a more powerful model. GBD5 is, you know, a bit faster. So I was able to get the answers quickly. But ultimately, it was just a testament to how well the experimental reasoning model could sort of be like, okay, this is a good solution.

14:44Let's just submit this one. And then also, the last problem, GBD5 couldn't handle it. So we needed to fall back to the powerhouse, the reasoning model to get that one. Yeah, well, let's talk about that. So 11 problems were solved on the first attempt, like you said. What does this tell us, in your opinion, about the reasoning capabilities? And maybe you could expand a little bit on the 12th problem, too, if that has any insights there. It's basically a testament to how well we've managed to get just tier one quality reasoning models into the hand of the public. Like just a couple of years ago, it would have been unfathomable to sort of have a model as powerful as GPT-5 in just the hands of everybody.

15:28And now like people all over the world can just go on there and just try it out. It's honestly amazing. Like the progress, I look at it and I'm just like, wow, how far like we've come. So it's, yeah, just a testament to how well just the standard models that everyone has at their fingertips can reason, can solve these difficult problems. And it isn't just competitive programming. It's just all over. They do well in mathematics. Academics use it all the time to help them in their research. People use it in their day-to-day jobs in engineering and science. um so uh yeah it was just a testament to like what has now been commoditized um and uh the reasoning problem or the the reasoning model itself uh obviously like it's one of our you know newer models uh internally stuff we were experimenting on um it's it's not quite you know as fast as gpd5 because we're just you know training it for research um yeah but it's demonstrably better.

16:26Like you could tell, uh, GPT-5 tried seven times to solve that last problem. And each time it's failed. Uh, so you can imagine, you know, like Mustafa and Boris and Robin and Andrew sitting there, uh, submitting every, you know, like 10, 20 minutes and me and me, okay, this one's wrong too. Um, and all the while the experimental reasoning model is just thinking, it's thinking for hours and hours. It hasn't even like, I'll put it a single answer yet. so it's three hours later and it's still thinking it's just grinding on this one problem uh and then eventually it gets to uh solving it we start seeing it like you know get some answers uh we submit the first one it's wrong as well and then the second one that um it's we submit from that model is correct so gbd5 tried seven times and got it wrong and then you know on the second try uh the experimental model the imo model uh just you know got it amazing how close wait I think you said this before, but how close was the countdown clock?

17:23Like, how many minutes did you have left? I don't really know, but it must have been within, like, 30 minutes of, like, the competition. Wow.

17:29Ahmed El-Kishky:You know, a year ago, I mean, when we sit down and think about this, like, JGBT isn't three years old. I mean, like, I mean, just saying that sounds crazy. And, you know, and I think about AI struggling with even fairly easy contest problems, like, as recently as a year ago. Were there any, you know, insights or breakthroughs along the way specifically that really enabled this big leap we've seen over the last 12 months or so? I think we announced it about a year ago, but the biggest leap was just the introduction of reasoning models. If you remember when ChatGPT was released, I guess almost three years now ago, the models or ChatGPT, you'd ask it a question and immediately it would just give you an answer.

18:14It would just always give an answer. Sometimes the answer would be correct. Sometimes it wouldn't be. It didn't matter how hard the problem was. It would just always immediately give you an answer. So back then, people had these approaches. They called it chain of thought prompting. They'd be like, show your steps. Think through step by step. They would just ask the model to. And people noticed that when you just asked it to show your work, the benchmarks would just improve a little bit. It would be a tiny bit smarter. Yeah. So we worked on that. We were like, okay, what's happening here? And the idea is that like, when you have a difficult problem, no human just immediately like gives an answer.

18:56If I asked you to multiply like a four digit number by a five digit number, you wouldn't like give me an answer instantly. You would like, you know, get your piece of paper and pencil. You'd work it out. You'd think through, maybe make a mistake and fix it. And you give an answer. If you asked a scientist to work on a problem or a mathematician, they do the same thing. Hard problems require, you know more time more thinking um so we decided to lean into reinforcement learning as a way to get our models to actually you know think longer uh think better um and that's actually where the breakthrough came in um i thought what a breakthrough it was too it was a it was a crazy one like um we wanted to see like we wanted to a little bit mimic how a human you know thinks through these problems.

19:44They try things out and maybe they go down wrong directions to dead end. They course correct. And OpenAI has always been into like reinforcement learning. From the days of like, you know, Dota playing video games, they really leaned into reinforcement learning as a tool that would bring about sort of next level intelligence. And there'd been some attempts to apply reinforcement learning to LLMs, but nothing at this scale. So the, I guess, the strawberry efforts, the O-series models were the first, like, very serious attempt at getting reinforcement learning working on these large language models.

20:24And it was honestly, like, amazing to sort of see it from the beginning, like, you know, when it started to now being at a performance level where it's competitive with some of the best competitive programmers. I guess like as just what a couple more questions about the competition. So which problem in the contest surprised you the most? Like either because it was harder than you expected, like the 12th one, or maybe was it one of the first 11 that you were like shocked that it was so easy?

20:54Ahmed El-Kishky:Or the approach it took even. Yeah. Yeah. So, I mean, to be fair, I don't think I'm even at the level where I can judge these problems. That's fair. Yeah. Like these college students are like, there's something else. I mean, they really should be proud. These are some of the most brilliant college students in the world. And if I took these competitions, maybe if I'm lucky, I'd get one. But that's being generous. But I mean, after the competition, I looked at the problems. And it was really interesting to me that one of the problems that no human contestant solved, our model solved pretty straightforward.

21:32the problem that we struggled with the most. It was also a difficult problem. The people that solved it among the human contestants took nearly all the time to solve it. So it shows me that maybe what's difficult for an AI isn't necessarily what's difficult for a human, and maybe what's difficult for a human isn't necessarily what's difficult for an AI. But yeah, so that was sort of the interesting aspect to me. What did those iterations reveal, if anything, that was new to you, like about the systems reasoning process? Did you feel like you'd learned anything new about how it thinks, like from reading that back and trying to understand, like kind of like going through after the fact and seeing what it did?

22:14I guess I do have the privilege of being able to peek into the chain of thoughts from these models. And it's honestly, sometimes it's a little bit crazy to read through because they're so, yeah, like they do like such clever things. Yeah. So I'll tell you a little bit of a story, actually. Let's make this a story over like how it started versus where it's at now. When it started, it would basically barely even reason. It would just try to output the answer when we first started like impetus programming two years ago. um when you started doing rl like just getting started there it would do a little bit of planning it would be like i'm gonna try this uh and we saw an improvement there over time the reasoning that the models would do would actually become way more elaborate uh and it was really interesting because we have a lot of like ioy gold medalists icpc competitive programmers and they'd be like oh yeah that's that's a strategy that i would use um so one of the coolest um the coolest things in my opinion is when the model decided it wanted to test itself before it submitted.

23:19And so what the model would do is, so the solution would have to be a very complicated algorithm. It would have to have great complexity. It would be very memory aware, very complicated. The algorithm would be like so complex that it's easy to make a mistake. So what the model would do first is write a brute force solution, which is a lot easier to do. It can be like a few lines. And then it would generate some inputs, like pretending to be a, like, you know, the judge would generate inputs in the same way. And then it would compare the outputs of the brute force solution to the really complicated algorithm that it wrote.

23:56And it knows that for it to be correct, the outputs have to match. So it couldn't submit the brute force solution because, you know, there's strict time limits and memory limits. It wouldn't like be correct, but it would spend a lot of time thinking and then trying to match the outputs of the brute force solution. And I remember when I noticed that and just showed it to people on the team, the competitive programmers at OpenAI would be like, yeah, that's a valid strategy. And what was really cool is we didn't tell it that. We didn't tell it to do that. It had never done that before. What had happened was just over time, as we trained our models, it just decided, hey, this is a valid strategy.

24:33I'm getting really good results when I do this. So let's keep doing it. And that's the beauty of RL. Yeah. And when you look at that. That's so awesome. Yeah. When you look at the reasoning, I mean, it's doing tricks like this all over the place. You know, obviously, IMO happened earlier this year. That was the International Mac Olympiad. And, you know, both you and Google were competing in that. And they did eventually release, I believe, a version of that model that they ran IMO with called DeepThink. And I think there's a version of that that certain users of Google can use now. And I don't know if you mentioned that you're using the same IMO model.

25:09or a version of that that you used for this competition. Do you think OpenAI will eventually release like a similar reasoning model from the ICPC system available in production in chat GPT, perhaps as a future model? Or what are your thoughts there? One of the big powerhouses of the ICPC was GPT-5. So that's already kind of available to the public. Right. So the exact model that was used for like IMO and ICPC, the extreme reasoning model, that one probably is not going to be specifically released, but the insights that were used to train that model will be eventually incorporated into later models.

Read the full transcript

25:50Generally, when we do research, our goal is to bring it into the main models that we host on ChatGPT eventually. So that's a big process. We have to make sure our models are safe, aligned. They're, you know, pleasant models to use and, you know, more things than just like competitive programming or math. Right. But our hope is to take the insights that we use to train those models and incorporate it into the next models. So maybe it'll happen in a later version of GPT-5. Maybe it'll happen in GPT-6. But ultimately, our goal is whatever sort of we're experimenting on should eventually make it into not just, you know, a experimental model just available to a small group of people, but to everybody.

26:37Ahmed El-Kishky:So here's a question, and it may even seem a little silly, but I want to ask it anyway. What do you do on a day at OpenAI when you're not making these big breakthroughs or smashing competitive coding records? Like what's a normal day look like at OpenAI? It's honestly like being a graduate student sometimes. So we all have like, you know, projects that we think are incredibly useful and projects that are going to make our models a lot smarter. Usually we talk about them. We discuss them with our colleagues. We get some good constructive feedback. Sometimes people are like, actually, no, that's not going to work.

27:18Or maybe you have to actually, you know, push through. Nobody bets a thousand. Exactly. And then we basically, like, we're just scientists. Like, it's a bunch of nerds just really liking the work we do, really believing in AI. We, you know, are working on it. We run these models. We train them. We sort of do an experiment, see if our hypotheses are actually correct or not. So we're measuring, like, oh, did this intervention make the model smarter this way? Or did it, like, cause the model to do this? And it's honestly very data-driven, very scientific. And sometimes it takes a bit of coding. Like we're not all just like running experiments.

28:00We code a bunch. We sort of see how well these models then, you know, solve something. And we just reiterate. Once we're sure that what we propose is actually correct, makes our model smarter, makes our model better in some aspects, eventually it gets incorporated into what becomes tri-GPT. But a lot of it is just, you know, going about your day, just banging your head sometimes being like, why won't this work? It's honestly, in that way, it's very similar to a lot of other work, but it's really satisfying when you get to those situations where you see the fruits of your labor. That could be like, you know, weeks, months, sometimes years away from when you start a project.

28:44So it's often worth it in the end. I am sure like being able to like take a historic victory is a good fruit of your labor. I guess like switching gears just a tiny bit here. There's been buzz about the new version of Codex and, you know, a lot of people have been talking about how it can run for seven hours straight in some instances. how do you think about like scaling you know from a five-hour icpc problem to these more longer horizon agent tasks like in terms of what you're going to be doing next you know so shout out first to andrei mischenko hansen katie shi who and the rest of the codex team they've done an amazing job with the new codex model and i see it all online people just raving about it yeah so one of the things I think they're really excited about is how much it thinks.

29:37So if you ask it an easy problem, it responds really quickly. If you ask it a very difficult problem, it can take hours and hours to sort of think about it, work it, try things out in a very agentic manner. And I think that's really like cool. And they're incredibly excited by it. And they want to see how far they could push this um and i think it really it's uh it's along the like the our goals for as a company um we want to sort of see how far we can push the limits of like autonomous just agentic work and that involves both uh thinking reasoning uh but also engaging with like you know for example a computer or uh an environment um we think that's sort of the next frontier uh competitive programming like we our models are now like you know among the best uh for imo uh we've shown that hey we can you know play at the same level as some of the best mathematicians in like high school and these are some really tough competitions yeah um the the next frontier we're thinking about is like what if these models took longer than hours what if they took days weeks months to even like you know solve a problem some of the hardest problems in the world i don't think they're gonna be solved in hours um i agree yeah uh if you want to uh further um sort of human knowledge um almost like a phd student would spend four or five years working on this very narrow subject with the goal of just like you know making a tiny increment into you know the sphere of human knowledge um we're hoping our models sort of do the same thing uh we want to give them difficult problems and we want them to maybe like you ask the question and then you leave you uh come back like in a day or two or maybe like you get an email that's sort of what i personally feel the world is going to um you're going to have these be a amazingly powerful tool um that can you know solve way more difficult problems um one thing that i'm really excited about is like ai for science right now science has sort of been you know bottlenecks a bit but imagine giving a tool to scientists and maybe mathematicians where it can you know work with them you basically have some sort of idea or you want to investigate an area and you have the ability to spend a lot of compute sort of thinking about it trying things out and coming up with something novel, like something that humanity has never seen before.

32:15And that's sort of where I see things moving, I guess, in the near future. There's going to be a lot of focus on making sure our models can solve these very difficult problems that will take longer than a few hours. And I think the Codex achievement is just a stepping stone towards this. We've shown that it is possible. Each day, we're sort of like, you know, solving tasks that take longer and longer. Open AI's chief scientist, Jacob Pachocki, Pachocki, forgive me if I butchered his name.

32:44Ahmed El-Kishky:I'm so sorry. Mentioned moving from this win to automating scientific discovery over months and years. And as part of the team that pulled off this win, what do you think are some key technical challenges along the way in getting to where you can do that? We are very bullish on, you know, scale still at Open AI. I know some people mention, oh, maybe scaling isn't the answer, but I think scaling does sort of help. Having models that maybe are trained with more RL. We posted in our blog when we first announced O-Series. We had this really amazing scaling law that we sort of like showed in our charts.

33:28Besides pre-training, as we like put more and more trained compute into RL, we see better performance. So in this case, as we just train longer for reinforcement learning, our models get smarter. But then also, as we let the models think longer, in this case, sort of expend tokens, they're also able to get smarter. So we're really heavily betting on that. So we want to make sure that our reinforcement learning algorithms are working well, that we have the compute necessary to continue scaling these. to get better and better performance. But another one is just good tool use. Having a model that can actually engage with the world in meaningful ways is a key component here.

34:19If we constrain the model to just like thinking in text and never getting that experience, maybe trying stuff on a computer or maybe eventually trying stuff in the physical world, we will sort of be limiting what the model can actually tackle and solve. So my personal opinion is that the things that are necessary is to make sure that as you scale, things continue working. When you throw a lot of compute, a lot of computers at problems, things always start breaking down. You know, just how it goes. As things get more complex, everything gets way more difficult. So making sure that the ML is still stable, that the software is still great.

35:04Making sure, you know, we have a really good way to interact with what we're trying to solve. Because the models themselves, you know, we would like them to, you know, if they're working on something with chemistry, like it should, you know, get some feedback from like running a chemistry experiment. Yeah. So these kinds of things are, I think, the next sort of frontier. Wow.

35:27Ahmed El-Kishky:Yeah, I keep thinking of that as, oh, sorry, Grant, go ahead. No, we both have, It clearly sparked something. I was like, I keep thinking of that as a thing involving sensors and other things. Like, how do you teach it what a spring day feels like, for example? Like, how do you do that? And I think there's so much frontier still to grow in different types of data and sources of data that these things can be fed. that it makes me so optimistic about the future and the work that you guys and other labs are doing right now, especially in science. Science is the one that Grant and I kind of always watch.

36:06Yeah, on the outside, you know, I'm just like, that's cool, but I don't understand it. Actually, my question is kind of related to what Corey is saying. So, like, when it comes to tool use, I've always wondered, like, how exactly, and, you know, you can just kind of, like, you can talk about this at whatever level of abstraction, surface level you want. But how is it that the AI is actually able to use tools? Is it just like getting the responses of the feedback like streamed back to it? Like that part still kind of blows my mind how that works. Overall, with some of these tool uses, yeah, it tries out the tool.

36:37It just, you know, makes a call to it with some text. And yeah, the outputs from the tool just gets put back into the context. And so now it has in its context, oh, this is what happened when I used this tool. And so it can just like use that to continue making decisions. Like if you need to make a decision in a situation, you need to, you know, have in your context what's the, yeah, what the state of the world is. The world can be your personal computer if it's like, you know, it's trying to look at your directories and see what's in the specific folder. That would be sort of put back into the context and now it knows what's there and it knows, hey, you need to make like an edit to this file, for example.

37:20Right. And so if it was doing that in a science on like a science model, like you mentioned, like if it's doing like a physics simulation or something, it would still be getting the data back and then it would still be kind of working the same way.

37:31Ahmed El-Kishky:Yeah, yeah. Wow. That's so cool. Here's a question that's probably a little less, it's definitely less technical, but it's a, I think it's a fair question. And what message would you want university level programmers like you were not that long ago, those who trained to compete for years, getting ready for these type of coding events to to take away from the result? And how how does seeing an AI beat human teams affect your personal journey? Like looking at chess, you know, AI has surpassed many chess players. I mean, all of them actually at this point. Yeah. But chess is still incredibly popular.

38:14There is a tremendous amount of value in being a, you know, competitive programmer, someone that tackles these algorithmic problems. And just because AI can do really well, that doesn't take from it. there's a lot of valuable learnings that people get from going through that journey. You get the ability to think really deeply to solve a hard problem. You need to work logically. You need to develop grits as you sort of like hit a dead end and need to sort of like power through to sort of get something that just isn't coming easy to you. All that stuff makes for, you know, makes for great researchers, great thinkers, and, you know, the next generation of people that will discover the next big thing.

39:04And so what I would tell, you know, college students is keep doing what you're doing. If you look at, you know, our peers here at OpenAI, so many people were competitive programmers, competitive mathematicians. And what I really see is they bring a unique way of thoughts. And these unique, you know, ways of thinking are what lead to breakthroughs. Whenever they see a problem, they're not like, oh, this is a problem. It's impossible. Their, you know, experiences tell them, hey, this is a problem. Let's, you know, try different approaches, get some feedback, and then try to push through it. And in fact, a lot of these people are incredibly excited by the ability to sort of piggyback on a powerful AI to solve even more difficult problems.

39:52I'll give you an interesting, a funny story. When we were training our models to be really good competitive programmers, we had, you know, our two, you know, C-level researchers there, Mark Chen and Yaka Pachalki, the chief research officer and chief research scientist. Mark Chen, like, coaches the U.S. Olympiad team. He takes them there, makes sure that they're ready to compete on the world stage. And Jakob was a very accomplished, you know, competitive programmer, an ICPC, a Google Code Jam, winning so many. So they were always kind of our goal. And so we sort of made a chart. The chart itself was sort of like the ELO score of our model.

40:33And it was like 300 or 400 ELO when we started out. And we had this line that was just a Mark Chen line, but it was like above 2000 ELO. we had a nice little picture of Mark Shen and then we had a line a little bit above that of Jakob Ochaki at like 2600 ELO and it seemed so far away and basically both Mark and Jakob would like you know encourage us because they were excited by you know these models getting better because they knew the implication there of having a powerful model and like over the year or so that we were like you know trying to see how good our models are it would just like inch up a little bit on that chart so like oh this model is you know 600 elo this one is a thousand fifteen hundred until one day i remember uh messaging mark be like i have some bad news and it was we saw that like that model i just like slightly edged uh and just craft above mark's elo score on code forces you sir have fallen that's amazing afterwards like not even like maybe a month month and a half later um yakub you know was also slightly beaten and um Yeah, they basically were incredibly excited because they knew the implication.

41:41And as our models got better, we were able to do more. I know because AI has sort of like come at us almost like opening the floodgates, it becomes sometimes difficult to realize like how amazing this all is. Like I look back five years ago and I know that like I would have been astounded by the progress made. And right now, like having something this powerful at your fingertips, it lets you do so much more. We're all much productive coders at OpenAI, and I assume all over the world because we have these models at our fingertips. And as they get better and better, we can use them to just, you know, make a bigger impact on the world.

42:20And that's sort of how I see it and why I believe that students should keep doing what they're doing, because that is the path to growing and improving yourself. And when you grow and improve yourself, you will be able to better use these tools to, you know, address whatever, you know, issues are in the world or the state of the world at the time. It just becomes a really powerful tool. But you need to have, you know, the same skill sets, the fundamental aspects of being able to reason, being able to solve problems, being able to have the grit to sort of through whenever you hit blockers. All of that is stuff that students can benefit from.

42:56So what I didn't hear there was whether or not you're over or under on a college degree. And I say that, I say that jokingly, but I do actually want to know, like, do we like, like, based on where we're at now, like, do we need to rethink education to a certain degree for human programmers now that AI can excel at these types of contests or all the things where AI is like continually improving? Like, is it still worth it to go to school to learn programming? I mean, this could be your personal opinion. Like, well, I do think there is value in, you know, school and education, particularly in that, like, it exposes you to many different people from many different walks of life.

43:37And there's sort of a collaborative human effort. I do think, however, that whenever a new tool comes out, education does need to change. Like when the calculator came out, you know, people had to reassess things. Suddenly, you know, doing just arithmetic was not necessary. You still need to learn the basics of arithmetic, but now you could use the calculator to go one step beyond, one step more abstract. Same thing with computers. Whenever you have a new technological innovation, it needs to be incorporated into the educational curriculum. And I think it's the same situation for AI. AI has made things a lot more accessible, but I don't think that devalues the low-level understanding of how things work, fundamentals of coding, fundamentals of algorithm.

44:29Like right now, for example, we still have human judges that say, hey, this code is what I want. This is what I need. This is why I think it's useful. And that comes from a good education, good experience on your own. And I think it's hard to predict the future, but I find it difficult to believe it'll be a situation where you just don't need any education. I think humanity will always need education. It will always need the ability to think and grow and reason. They also need to adapt to an ever-changing world. And I think AI is now a part of this world.

45:08Ahmed El-Kishky:And the need to think row and reason is arguably more vital than ever. And I really think that we're going to see those become an increasingly important skill set in the world as we go. And, you know, I love that you mentioned that, you know, maybe education needs to shift with the times because I think that's a really vital thing right now. Do you think that these types of competitions, you know, a few years out will remain primarily human showcases or evolve into joint contests, maybe parallel paths or something where everyone's AI is competing over here and humans are over here? It's hard to predict.

45:53I think there will always be competitions that are human focused. No one wants to watch, you know, an AI play chess against other AI these days. They want to watch humans play chess. But I do think there is a world where competitions change as well. Like maybe in addition to the algorithmic stuff, the algorithmic growing competitions, maybe there's like a couple new problems where to solve the problem, it's not only like, you know, very fundamental algorithms and data structures, but maybe it incorporates an AI somewhere in there to solve a more complicated problem. I think the ACKoder competition was a really good example of that.

46:34you can sort of build out you know strategies and the AI can maybe come up with strategies as well so maybe there's some sort of harmonious meshing of like the two maybe also there's a new class of like problems that we don't even know about maybe you know people can you know modify their program with AI to sort of like compete against each other so that way you incorporate AI somewhere I think it's hard to know where the world's heading, but I really believe there will continue to be these competitions where, you know, students and individuals are, you know, pushing the limits of what humans can do and sort of improving.

47:16Ahmed El-Kishky:What's next on your list? Are there upcoming competitions, milestones, things you're, you know, zeroing in on and thinking, I want to knock this over? The thing that actually interests me the most is pushing human knowledge. I think it'd be really great if our models were contributing to humanity. I think we've kind of pushed to the limits of competition. I think the competitions are not really an end goal in themselves. There is sort of a way to benchmark progress. if we're able to sort of improve on these that's um that's that's great but uh we don't want to train ai just to play uh and compete in competitive programming competitive math we want to make sure our models are smart enough to do things like be a valuable for example coding assistant so software engineers be valuable uh you know engineering partners to other forms of engineering you want them to be an aid to scientist academics something that's you know really interesting to me is yeah can our models actually discover some new knowledge it can maybe be a maybe a new algorithm that no human has ever discovered or you know a math proof that's stumped mathematicians for a long time or yeah even biology something in chemistry all of these i think are things that I think would be great if the models could sort of contribute to.

48:45And that's a little bit where I see things going. Just all over the world, people are starting to think about using these models to, you know, sort of push beyond what humanity has currently discovered and known. Is there a, like, a competitive science competition? And I was thinking, well, the Nobel Prize. Like next, next goal unlocked is AI wins a Nobel Prize, right? Maybe like in the future, we'll have like, you know, papers where the first author is like an AI or something, and it wins a Nobel Prize one day. I could see it. I don't know what we need between now and then to get to that point, but I could see it happening.

49:27Are you familiar with the X Prize, either of you? Like it could be like the X Prize, right? Where the X Prize is like groups of scientists, researchers try to compete to solve these really hard problems and they get money from it. So maybe that's your next challenge. Win the XPRIZE.

49:41Ahmed El-Kishky:It'd be great for all of us, wouldn't it? Yeah. All right. We're going to ask our favorite question here, which is what's in your personal AI stack right now? I mean, obviously, we know you're using open AI models, but are you personally using codecs? And if so, how? Like, are you having agents just automate everything now? Are you still coding by hand? What are you doing? I mean, it's pretty basic. But yeah, Codex is really great internally because let's just say you want to make some change to our code base. We can just write something in there and just leave it and then come back and you have a pull request just ready for you or ask it to explain some part of the code that you don't understand.

50:22Honestly, it's kind of shocking how useful it is and how much you grow to rely on it. like we didn't have codecs you know like not too long ago when i joined everything was sort of done manually but now it's sort of such an accelerant and then obviously like we use gpd5 internally and just every time i have a question i always ask it to you know explain things for me that i don't understand or you know work on some math for me so my personal stack is just gpd5 and then codex for anything related to code. Do you think it's like good for a person who's interested in learning how to code to use codex or should they wait, learn the fundamentals and then get introduced to codex because of how much it can do for you?

51:06I don't know if they should be first or like wait, but they should definitely learn the fundamentals. One shouldn't rely on something as a crutch to not learn your fundamentals. The fundamentals are always useful even beyond the immediate, you know, outcome. It's very, you need to think long-term and long-term, you know, improving oneself involves getting the fundamentals. But then once the fundamentals are sort of achieved and you've mastered them, it's an accelerant. You should use it as a tool and then try to pick up something that's more abstract, more complex, or you can rely on this tool to sort of accelerate that.

51:46Yeah, so I'm always pro-learning the fundamentals and I'll never say otherwise.

51:50Ahmed El-Kishky:Ahmed, thank you so much for joining us. It's been a real pleasure. I've loved the chance to chat with you and pick your brain on all this. You guys have so many exciting things going on right now. Thank you for having me. It's been a pleasure, Corey and Grant. I loved answering these questions. And yeah, hopefully we can show you some exciting stuff in the coming months. Awesome. Yeah. I can't wait to see it. Can't wait to see it. I'd like to give another quick shout out to Whisperflow for making today's episode possible. Remember to go try AI voice transcription tool for free. at whisperflow.ai slash neuron today.

52:24Ahmed El-Kishky:Yeah, and if you enjoyed today's episode, please take a moment to like, subscribe, leave a comment. Also, make sure to check out the Neuron's Daily AI newsletter at theneuron.ai. We really appreciate every one of you for joining. And that's all we have for today. Farewell for now, humans.

52:50I'll see you next time.

From the publisher

In this episode, we're joined by Ahmed El-Kishky, research lead at OpenAI, to discuss their historic victory at the International Collegiate Programming Contest (ICPC) where their AI system solved all 12 problems, beating every human team in the world finals.


We dive into how they combined GPT-5 with experimental reasoning models, the dramatic last-minute solve, and what this means for the future of programming and AI-assisted science.


Ahmed shares behind-the-scenes stories from Azerbaijan, explains how AI learns to test its own code, and discusses OpenAI's path from this win to automating scientific discovery over months and years.


Subscribe to The Neuron: https://theneuron.ai

WisprFlow: https://wisprflow.ai/neuron

OpenAI: https://openai.com

More from The Neuron: AI Explained

All 106 episodes
How OpenAI Beat Every Human Team at the World's Hardest Coding CompetitionThe Neuron: AI Explained · 53 min
Listen in VO