In short
AI capabilities research and how fast AI is progressing, using math as a test bed. The episode argues that AI’s benchmark performance has continued improving steadily (often appearing linear after statistical “stitching”), driven largely by scaling compute and training data, while “discontinuities” from algorithmic self-improvement are not yet evidenced. It also discusses how AI is moving from solving known-style exams to solving previously unsolved problems (e.g., Navier-Stokes), and what that implies for understanding AI’s true limits.
Guest backgrounds
Greg Burnham leads capabilities research at Epoch AI, developing methods to track AI strengths/limitations (including an “Epoch Capabilities Index”) and benchmarks like Frontier Math.
Key claims
AI progress shows no measured slowdown; benchmarks are correlated and can be unified into a composite trend; discontinuous jumps are mostly not observed; AI math solutions rely on prior literature + persistence, not clear “new idea” generation yet; Epoch’s Frontier Math is designed to verify solutions automatically and detect creative breakthroughs.
Notable examples
Frontier Math tiers; O1’s math spike; solving the Erdos unit distance problem; Navier-Stokes as a Millennium Prize problem; Go “move 37” as an analogy for a genuinely novel move.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOAI's Rapid Progress in Mathematics
0:07 to 2:45
Explore how AI has advanced in solving complex mathematical problems.
“Today's codebases have grown beyond human comprehension.”
The Importance of AI Capabilities Research
2:45 to 4:35
Understand why measuring AI capabilities is crucial for future developments.
“And these, like, AI was perfectly good at these already, end of 2024, beginning of 2025.”
AI Performance and Benchmarking Insights
4:35 to 7:31
Dive into the nuances of AI benchmark performance and its implications.
“at, you know, certainly we're seeing this broad performance.”
Generalization in AI: Deep vs. Shallow
7:31 to 14:00
Learn about the concepts of generalization in AI and their impact.
“It seems like that, like normalizing across different benchmarks to produce a result that is coherent, that sounds like a very challenging problem.”
Statistical Trends in AI Models
14:00 to 15:00
Explore how the progression of AI models can be viewed through statistical trends.
“statistics to be clear but this sure enough looks very regular in its own statistical analysis terms over across generations of models.”
Discontinuities in AI Progress
15:00 to 17:42
Discuss possible discontinuities in AI progress and their implications.
“function, you know, improvement over GPT-5.6, you know, linear implies to me that it's kind of incremental and there's really been no step functions.”
Tracking AI Innovations and Research
17:42 to 20:45
Learn about methods to track AI innovations and research breakthroughs.
“But it's the sort of thing that has spooked the people inside the labs, the AI companies, I mean, saying like, gosh, this could be right around the corner.”
Mathematical Challenges and AI Progress
20:45 to 24:45
Examine the impact of high-level math challenges on AI development.
“You know, certainly Navier Stokes has been in the news quite a bit recently.”
AI's Remarkable Math Competence
24:45 to 27:30
Discover how AI has rapidly improved its math capabilities over time.
“And at that level, we started, this was all just this year, like 2026.”
Understanding AI's Jagged Frontier
27:30 to 28:00
Analyze the complexities of AI's performance in various mathematical domains.
“And, you know, we're all living through this trajectory, but to step back and see, how much has happened in the past couple of years is pretty nuts.”
Show all 24 chapters
Superhuman Capabilities in AI and Mathematics
28:00 to 30:00
Exploration of AI's capability to solve complex mathematical problems and the nuances of human and AI approaches.
“in the sense that they'll be superhuman on some axis and not superhuman, still inferior in capability to human on some other axis.”
AI Problem-Solving Techniques
30:00 to 33:00
Discussion on how AI tackles mathematical problems, including the role of brute force and the use of existing knowledge.
“When humans solve these math problems, they write up the solutions in like papers, proofs that, you know, this is true or this is not true.”
Impact of AI on Mathematical Progress
33:00 to 36:10
Insights on how AI builds on human mathematicians' work and the potential future of AI in creating new mathematical concepts.
“there's an element of brute force there.”
Defining New Ideas in AI and Mathematics
36:10 to 41:40
Exploration of how to assess AI's creativity and ability to generate new mathematical ideas compared to humans.
“a new thing, a new mathematical concept, and that unlocks a lot for us.”
Evaluating AI Solutions and Human Efforts
41:40 to 42:00
Discussion on the challenges of evaluating the results generated by AI compared to human efforts in solving mathematical problems.
“And you could imagine doing a more rigorous version of this experiment.”
The Impact of AI on Predictability
42:00 to 44:40
Explore how AI can lead to both incremental and discontinuous improvements in fields like chemistry.
“Not always, but its impact is often hard to predict.”
AI Innovations: Lessons from AlphaGo
44:40 to 46:32
Discuss the surprising moves made by AI in games like Go and their implications for mathematics.
“Yeah, we've got a couple, two new sets of frontier math problems.”
New Frontier Math Problems for AI
46:32 to 48:26
Learn about the new set of frontier math problems aimed at evaluating AI's creative problem-solving abilities.
“this is maybe the big piece of work on our side, was to make it easy to verify automatically whether an AI system has solved this problem.”
Measuring AI's Progress: Challenges and Methodologies
48:26 to 50:23
Delve into the complexities of measuring AI's progress on unsolved problems and the methodologies used.
“It's intuitive that, OK, the best humans were able to do was X.”
AI's Capabilities in Creative Tasks
50:23 to 54:04
Examine the limitations of AI in creative tasks like data insights and infographic generation.
“Say, which problems did it make partial progress on?”
Learning on the Fly: AI vs. Humans
54:04 to 56:00
Compare how AI systems learn in dynamic environments versus human learning capabilities.
“We have a benchmark that's taking a complicated board game.”
Exploring AI's Learning Capabilities
56:00 to 1:02:32
Learn about AI's ability to learn strategically and its limitations.
“versus harness thing in and of itself is, you know, practically brand new.”
The Current State and Future of AI Capabilities
1:02:32 to 1:06:06
Discuss the evolving capabilities of AI and monitoring its growth.
“They kind of group into, like, how do I assess where AI is now?”
Practical Impacts of AI and Upcoming Benchmarks
1:06:06 to 1:07:35
Examine the real-world effects of AI and the importance of benchmarks.
“look, cyber is maybe the latest example, where it's like, look, we see their ability in controlled settings to do hacking going up smoothly.”
Transcript
Automatic transcript. May contain errors.0:00This episode is brought to you by Blitzy, the autonomous software development platform built for enterprise scale. Today's codebases have grown beyond human comprehension. Millions of lines, decades of tech debt, and complexity existing tools just can't fathom. With Blitzy, thousands of specialized agents reverse engineer the codebase, mapping the architecture, dependencies, and business logic. With that context, the platform then autonomously executes entire epics. writing, validating, and testing the code for every project. The result? Fortune 500 enterprises are able to modernize legacy systems and ship new features five times faster.
0:40Want to try Blitzy on your code today? Unlock 1 million lines of reverse engineering and 25K lines of code generation by visiting blitzy.com slash sandbox.
0:53AI's progress in mathematics has been remarkably fast. Not too long ago, frontier models struggled with grade school math. Now, they're contributing solutions to problems that have resisted mathematicians for decades, including, most recently, the Navier-Stokes equations. But as AI moves from tests we already know the answers to to more open-ended research challenges, it gets harder to understand what these systems are truly capable of and what their progress tells us about where AI is headed. My guest Greg Burnham leads capabilities research at Epoch AI, where he and his team are developing new ways to track AI's rapidly evolving strengths and limitations.
1:32Here's Greg on why measuring and understanding AI's capabilities matter so much right now.
1:38Greg Burnham:AI capabilities writ large have gotten to such a point where what they're good at and bad at is starting to have real impact on the world, like on things we care about for other reasons. We're past the academic exam phase of understanding AI capabilities. And for that, we need high quality benchmarking, high quality evaluations. And one of my big questions is just can AI come up with new ideas? If AI systems gets sort of superhuman at this, that's a big deal for the world, both in terms of economic and scientific progress, but also in terms of becomes harder to predict what AI systems are capable of.
2:18Greg Burnham:We use math as a test bed for this. I'm Sam Charrington, and this is the TwiML AI Podcast. For over a decade, I've been exploring the ideas and innovations shaping the future of AI through conversations like this one that help you understand what's real, what's next, and what matters. Let's jump in.
2:45let's talk a little bit about ai capabilities research broadly how you pursue that research
2:50Greg Burnham:and why you think it's important we're getting to a point in the in the trajectory of ai overall where ai capabilities can have a real impact on the world so maybe the transition could be characterized as school to work in maybe even up through 2025, most ways that we tested what AI could do looked more like human exams, like that you might give to, you know, to kids or to maybe advanced graduate students. I was going to say. Right, right. The kid exams, It is remarkable. The bar exams, the medical exams. And these, like, AI was perfectly good at these already, end of 2024, beginning of 2025. So we really saw this transition happening over time, to be clear.
3:46But this is certainly the year where AI capabilities are impacting work activities that humans were engaged in anyway for their own purposes.
4:01Greg Burnham:And so as these capabilities grow, it's just important for us to keep tabs on them. They grow very rapidly in some qualitative sense. We see no measurement we have shows any slowdown in how AI is getting better and better at any tasks that we're able to measure. And so we think that just understanding the impact AI will have on the world, it's important to know what it can do, what it can't do, and keep tabs on how that trajectory is going. I want to push back on one thing you said in terms of we're not seeing any slowdowns that can be looked at broadly, like in terms of maybe the number of, you know, benchmark data points that we're looking at, you know, certainly we're seeing this broad performance.
4:52but is it also true that like on a particular, I feel like on a particular benchmark, we're getting like the, you know, for expected reasons, when we went from, you know, GPT-2 to GPT-3, like we were taking these huge steps up these benchmarks and now like there's just a lot less room left and we're making more incremental progress. Do you see that as a slowdown in AI performance or would you characterize it differently?
5:22Greg Burnham:I would say, well, I'd say there's two issues here. One is on the benchmarks we have, what does performance look like? Two is, are there things our benchmarks don't measure that maybe there has been less or could even be negative improvement on? And is there some reason why it's the things that are hard to measure where there's been the least progress? But on the first point, because this is a very, this is a much more cut and dried statistical point, we absolutely see continued progress on benchmarks, like a benchmark that was challenging for GPT-3, you know, many years, whatever, six years ago, is now completely, you know, solved, aced by GPT-4.
6:11Greg Burnham:Benchmark's challenging for GPT-4, same for GPT-5, now we've got GPT-6. So we do some statistical work to try to aggregate benchmarks over time because the benchmarks that people used to track progress years ago, now every AI model off the shelf that you've heard of that you might use is too good at them. So people make new benchmarks. And then you have to sort of do this like stitching of like, okay, well, a model that was pretty darn good at this benchmark was not so good at this next one. And you can use like, you can stitch those together and get this sort of composite score. We call that we have a methodology for doing this, we call it epoch capabilities index.
6:52Greg Burnham:And that's a unified way of tracking benchmarks over time. And we really do see this thing continuing to go up, you know, pretty smoothly, linearly. And that march of progress is steady. That's an important point. That's truly if we have, we've made more benchmarks and AI systems that are harder, that models start off at zero, and then within a year or two years, they're at 100 % and we have to make a new benchmark to keep up. And that's really been the story of the last, certainly the last three years, I'll say with confidence there. It seems like that, like normalizing across different benchmarks to produce a result that is coherent, that sounds like a very challenging problem.
7:45Like you, you, you describe the result as like, you know, kind of continued linear improvement. And I imagine if that's like a constraint, you can kind of map backwards to, you know, to factors that will show you that continuous linear improvement. But, you know, starting from, you know, a set of disparate benchmarks that have their own kind of, you know, ranges and challenges and, you know, where AI and, you know, where AI sits in them and, and kind of mapping this across different generations of AI and coming up with a, you know, something that makes sense, you know, objectively at the end and then having that thing show continued linear progress.
8:30That sounds like, um,
8:35you know, almost like too good to be true and like suspicious that it is true.
8:39Greg Burnham:I, I, yeah, I, I agree. So, so let, let me go into that a little, cause I think it's, uh, It's a killer point to understand. And it is surprising. So I think you're very right to be surprised. In fact, if you look at AI benchmarks, they are all correlated with each other. Even if the benchmark, the scores across models are correlated with each other, even if the benchmarks are in nominally different domains. So maybe you don't expect, you don't find this so surprising if it's like I've got one benchmark that's, you know, graduate chemistry questions and I've got another one that is coding puzzles.
9:18Greg Burnham:And the models that are good on chemistry are also good on software engineering. Maybe that's not super surprising. They should be correlated largely by the models themselves. Like as the models get better, they're going to get better across these, you know, disparate benchmarks. Exactly. But that's the surprising point. So all that you need to be true to see this linear trend over time is basically every time, you know, Opus 4.5, Opus 4.6, 4.7, GPT 5.2, 5.4, every time these come out, if they get better on all benchmarks at once, then you're going to see this big over time linear trend. Like that's two sides of the same coin, statistically.
10:00Greg Burnham:Why does this happen? Like, this is a huge question. Like, you're right to find that suspicious. And I would say there are two explanations that we see. And it's an important question, which of these explanations is more true. And of course, the truth is some mix, but let me give the two explanations. One is the AI companies are making darn sure that their model gets better at all of these benchmarks with each new generation of model. And the main mechanism you might imagine them doing that if they're like looking around, okay, people care about chemistry, people care about software, people care about operating, but you know, answering my email, like they, they would collect training data, one form or another, that would help the model learn how to do these, how to do these tasks.
10:51Greg Burnham:and that's something of a very manual human in the loop process of just saying, hey, we're doing product development. Like they want this feature, they want that feature, our users want all these features. And the way you do that in machine learning is you collect training data and throw it in. I'm oversimplifying, but that's that. It's sort of like a shallow way to make it happen. So I'll call that shallow. The other one is deep. The deep way to make it happen is if the model general, is if the AI system can generalize very, very well. So yeah, you never trained it on chemistry, but you trained it a lot on math and software engineering, and it just got super smart.
11:31Greg Burnham:And with a modicum of chemistry basics, it was able to derive the rest from first principles sort of thing. This is like a deep form of generalization. Both of these happen to some extent, like kind of the revolution with GPT, even GPT-3 around that era, this is like quite a while ago was, oh, if you train it to like predict the next word, then it actually gets better at a wide, wide range of language oriented tasks that people had been trying to handle individually. Suddenly it just got good at everything. And that wasn't because anyone was like, oh, let's have data specifically around sentiment analysis or syntactic structural analysis.
12:14Greg Burnham:Like it wasn't anything like that. It was just give it this very general thing. but now we're onto this world where you have to like train it specific right now there's a huge industry of collecting all this rich training data and it's really unclear how much generalization you're getting from you know in this deep way uh and and if and anyway that's that's sort of a big question because that deep generalization means you could really see if that happened in a in a rich way where you just didn't have to even be trained on something at all to gain capability in that area, then you might have a radically superhuman AI that would be pretty hard to understand, to bound its capabilities.
12:55It's still a little counterintuitive that this would be linear. And maybe part of the question is, are we talking about linear across generations of model releases or linear within some regime, like segment, you know, segmentally linear or? The thing I can say is we have like a statistical methodology.
13:23Greg Burnham:So benchmark score, like what are we talking about here? Our unit here is a score from zero to 100%. And typically the way these scores go just historically, empirically is like they follow like an S curve, like they start out sort of bad, then they get better, they get better pretty rapidly and then you can only get up to 100 % so then they sort of level off takes them a while to get that last you know little chunk and then they're at 100 % so you've got these like sigmoids is what that s-curve is called we've just got a statistical methodology for stitching sigmoids together and what comes out is a linear function and this looks like this is an artifact of the statistics to be clear but this sure enough looks very regular in its own statistical analysis terms over across generations of models.
14:12Greg Burnham:Like, absolutely. Like it's been, if anything, it accelerated a little with the development of so-called reasoning models around late 2024, where the models started to get good at math and software engineering and things that require lots of logical reasoning. Maybe like it's a higher linear slope, but this is just like, that's a statistical artifact. You now have to ask, okay, well, if your score is like, one model, the score is 150 and the next model score is 152. Like what are those two points mean on your benchmark aggregate index? Now, I think part of what I'm asking myself as I hear this described as linear is, does it say that in spite of the model, the frontier lab vendor claims that, you know, Fable is like dramatically better than Opus or that GPT-6 is like the step function, you know, improvement over GPT-5.6, you know, linear implies to me that it's kind of incremental and there's really been no step functions.
15:19Greg Burnham:This is a really great question. And yeah, we believe this is one of the big uses of this Statsy tool we've got is you can try to find, you know, trend breaks and you can try to find accelerations. And we mostly don't see that. Yeah. Like one way to put it, that is everything just about is on trend. Fable, Astra, whatever, is on trend. But the trend is crazy. Like you got to you got to hold both of these together. Right, right, right, right, right. This is like this is sort of the zoomed out view EPOC takes with everything is a lot of this, we think the engine of AI progress is heavily mediated by raw inputs like GPUs, the training compute, like training data.
16:12Greg Burnham:Maybe those are the big ones. And then there's this like factor of the labs come up with algorithmic, like innovations, like research breakthroughs. But we really think like a lot of it is coming from the scale up in the compute and the scale up in the training data. So until you see the data and the GPUs, data and compute supply chain ramp significantly or discontinuously, to be more precise, you're not likely to see discontinuities on the model side. I think that's right. A big thing we're keeping our eye on, to be clear, is are there AI abilities emerging that could cause a discontinuity on the model side?
16:54Greg Burnham:So this is why people talk about recursive self-improvement, or sometimes they call this a software-only intelligence explosion. And the point of these concepts is, well, sure, right now there's a comparatively slow process for building more GPUs and collecting more human training data and whatnot. But if an AI system got really good on the algorithmic innovation side, could it do more with less or do more with a fixed compute budget? And that could change the dynamics here. So the feedback loop might no longer run through the physical world. It's not clear we see evidence of this yet. I'd say on balance, we don't really see evidence of this yet.
17:42Greg Burnham:But it's the sort of thing that has spooked the people inside the labs, the AI companies, I mean, saying like, gosh, this could be right around the corner. That's like a bit of the fear. And if that happens, then the pace we're used to, we might get that discontinuity. Have you developed a particular benchmark to identify and track this, or is it more the way you look at the benchmarks that we've been talking about, the statistical model that we've been talking about for an indication that there's some kind of exponential in the curve? And that's probably one of the likely causes for that kind of exponential or just continuity.
18:25Greg Burnham:All of the above. We want as many tools in our arsenal as we can for this, so that this Statsy-like acceleration detector is definitely a good tool. But we also sort of get down in the weeds with some of our specific evaluations and benchmarks saying, I'll give you an example. This is some forthcoming work we have that I'm happy to preview. We take a recent AI research breakthrough or innovation from humans and we take AI systems, we don't connect them to the internet and we choose AI systems that were developed and released just before this recent, whatever the recent human innovation is. And we basically say, look, here's the metric that that innovation, the number that it makes go up, it improves this sort of efficiency, whatever, some metric.
19:23Greg Burnham:And we say, AI system, your goal is to improve this metric as much as you can. But we don't tell it the innovation. And so we see, is it capable of replicating that human innovation that was highly relevant for AI, you know, research and development? And so far, we don't see them able to do this. There's this nebulous idea called research taste, which is like, where's the good idea? Like, what experiment should I do next? Where should I go hunting for big improvements more than the mundane improvements of just more compute, more data? Like, where should I look for a breakthrough? And AI systems don't seem great at doing this, at least in an AI R &D context.
20:09Greg Burnham:But they're getting a lot better at a lot of things that are adjacent to this area. So we really do think this is an important, sometimes I almost call it a tripwire, that we're, you know, the future is foggy, some foggy landscape. We want to put a tripwire out there so that if AI ever does cross some threshold where it's able to do rapid improvements in AI algorithms themselves, then we want that tripwire to sound. So we go, okay, gosh, watch out, like maybe something, maybe we're about to see a big improvement, a big, you know, uptick in AI capabilities. And how does this relate to AI's performance on some of the math challenges?
20:52You know, certainly Navier Stokes has been in the news quite a bit recently. Prior to that, there were some ERDOS results and others. How do you think about the role that those math challenges play in this broader view of trying to understand AI's trajectory?
21:11Greg Burnham:Yeah, very good question. If I may, I would rewind history just a bit here to say the trajectory we've been on, because once you step back, it's wild. Uh, so up until, yeah, up until fall of 2024, so we're talking two years ago, uh, grade school math was still a challenge for these LLM-based AI systems, the kind of AI we're talking about here. and all and the benchmark there was literally a benchmark of 8 000 you know word problems so and so you know has so many apples that sell for this gsm 8k gsm 8k that's the one and gpt4 which was 2023 had like done pretty well on this so people started looking at so like adding into the mix a benchmark just called math like four capital letters uh that um started pulling from some high school math competitions.
22:18Greg Burnham:And there was like slow progress on this, but not a lot. And then at the end of fall of 2024, OpenAI comes out with the O1 preview model, full O1 by the end of 2024. And this showed a huge spike in math competition capabilities. But we're still just talking like medium hard high school math competitions, like something, you know, thousands of math nerd, nerdy math 16-year-olds were into across the U.S. And this became the new territory. My company, EPOC, around that time released a benchmark called Frontier Math. At the time, that's all we called it. We now call it tiers one through four to distinguish from later versions I'll get to in a bit.
23:03Greg Burnham:But the original Frontier Math consisted of problems that I would say go from advanced undergraduate to advanced grad student. So sort of early career research warm-ups almost. These are problems that are no longer thousands of high schoolers would tackle them. These require deep background knowledge in various areas, niche areas. Like we're talking a problem from, you know, elliptic curves that you'd give to a grad student, like, shortly before they're going to embark on their novel thesis research, just to get them familiar with the details of the niche that they're going to be trying to work in and make an original contribution to.
23:49Greg Burnham:Over the same period, we also saw high, like, so over 2025, we also saw high school math contests get completely aced. We got the gold medal on the International Math Olympiad by the, even by the beginning of 2026, we saw most of these frontier math problems solved though not all of them had in fact been solved by an AI system at that point. And so just like, even just here, this is a very rapid trajectory from over the course of just a little over a year going from, you know, like hobbyist high schooler to graduate student. Like very competent kind of you know, up their graduate student in terms of the problems they could solve.
24:36Greg Burnham:Really rapid AI progress. At the same time, we started to see people test AI systems on unsolved math problems. Like the frontier was no longer something, this is what I mean, we'd gone from exam to on the job, where it wasn't like, here's your qualifying exam as a grad student anymore, it became more interesting to say, well, here's a small problem that maybe a mathematician, a professional thought about for an hour, didn't solve AI systems. Can you do it? And at that level, we started, this was all just this year, like 2026. At the beginning of the year, there were like starting to be claims of, oh, maybe here's a problem that was posed by a famous mathematician.
25:21Greg Burnham:This is maybe the Erdos problems. There's this, just for background, is this very prolific mathematician, Paul Erdos, who proposed over the course of his life, maybe like well over a thousand, maybe a couple thousand problems that people still like haven't fully cataloged all the questions he asked. Some of those became central to whole fields of math. Some of them were, you know, no one really paid much attention to, including Erdos himself. And AI started to notch a few solutions to some of the easier ones of those. Fast forward to May, of this year and an internal model, probably something like the Astra model that we now have publicly available, was able to solve a big one of those.
Read the full transcript
26:05Greg Burnham:This was sort of a historic moment, the so-called unit distance problem, which was a problem about how many points you can fit in the plane that are just one unit, distance of one apart from each other. It was sort of like very cool that you could frame it. It's such a simple problem, but there hadn't been much progress on this problem despite a lot of attempts since Erdos had posed it decades and decades before. And it was a serious problem. It wasn't one of these, no one had thought about it. It definitely had received a lot of attention and an AI system solved it. And it was really this, this was like a aha moment of, oh my goodness, like AI can solve problems that humans have cared about, just independently.
26:48Greg Burnham:And we've seen a lot more of that since. And the, you know, with a sort of the recent wow moment being the solution of the Millennium Prize problem. This was, you know, the Navier-Stokes problem. This was a collection of these Millennium Prize problems, a collection of seven problems that mathematicians sort of, or at least one math institute set as like good goals for the new century. Like they were, they were, they're much older than the year 2000, but, but they were sort of canonized in the year 2000 as, hey, like these are worthy targets of entire subfields of mathematics and it would be a big deal.
27:24Greg Burnham:And so one of them was solved. So it's been a hell of a, hell of a ride. The historical context is super interesting. And, you know, we're all living through this trajectory, but to step back and see, how much has happened in the past couple of years is pretty nuts. Yeah. You have to, as one of my colleagues said, you have to not get frog-boiled by it. The temperature is rising and it's easy to, I don't know. Anyway. So there is this interesting gap so far. You might hear the phrase jagged frontier, where the AI capabilities are jagged in the sense that they'll be superhuman on some axis and not superhuman, still inferior in capability to human on some other axis.
28:16Greg Burnham:And that remains the case so far in math. In particular, there's a couple of things. In hindsight, the analysis by mathematicians of the solutions AI has found to these big, important, previously unsolved math problems, it's always, in the first place, it's been nothing too super duper surprising. Maybe there was a direction that humans had overlooked or could have invested more in and AI sort of had the persistence, maybe this is a good word, or the bravery or the temerity to like wade into these very in-depth intricate computations and calculations and sure enough came out the other side with a solution and humans look back and are like ah that's not a crazy direction to try at all and maybe if we'd like gone in that direction we would have come up with this this isn't something that like would be surprising what i'm curious about is are we seeing the AI systems come up with a solution to the problem or an algorithm to create the solution to the problem?
29:35And I think I'm trying to like ask a lot of different things here. You know, one is like creativity. One is maybe levels of abstraction or the way it's like, you know, thinking, abstracting, you know, around these problems. Yeah, I'll kind of pause and let you react to that.
29:55Greg Burnham:Yeah, let me throw out a couple things there. Let me know if you want me to expand on any of them. When humans solve these math problems, they write up the solutions in like papers, proofs that, you know, this is true or this is not true. And those proofs are not usually, like not exactly algorithms. They're written in natural language. They have sort of, they leave the reader to fill in some of the gaps. And AI systems are coming up with just the same kind of thing. Like, it's the same products. There's a lot of complaints about their writing style, but like the arguments, the substance of the arguments really are the thing we expect to see from humans, such that if a human came up with this, and yeah, maybe cleaned it up a little, I don't know, like would be lauded for having made a very insightful connection between two different areas or having gotten an intricate argument just right and balanced different considerations to make everything come together to land the proof.
31:00Greg Burnham:A hundred percent, this is something we're past the point where we can say, well, AI systems are sort of not really doing something that humans would be impressed by on human terms. AI systems are doing things. And is there a way to categorize or qualitatively are we finding that, and I think you spoke to this a little bit, like, you know, is there a qualitative difference between, you know, the best human approach and the AI approach in a sense of like, you know, AI brute forced it. You kind of spoke to this versus like pulled in things from some other domain and like created some elegant solution or are we past that point as well?
31:44Greg Burnham:Here I'd give more nuance. There's an emerging sense of, at least for the moment, what an AI-shaped problem looks like. Maybe I could say AI has two big advantages right now. One is it knows everything. Like it knows all the prior work. It's got the literature memorized. I'm speaking loosely, but this is basically valid. The other is it's extremely patient and persistent. So if there is some intricate calculation you have to do, I don't mean like numerically or algorithmically. When I say calculation, I just mean lots of equations and you have to make all the pieces work out just as a very high level conceptual logical kind of computation, but still something very intricate.
32:38Greg Burnham:They just have the patience to go for that. And if you take the Navier Stokes example, I think it's not a coincidence that OpenAI's reported system that solved it included a swarm of thousands of different instances of AI systems trying lots of different things and sharing ideas where clearly it was able to, like, there's an element of brute force there. So there's degrees of brute force. This is where It wasn't like somehow, you know, you asked it to, like, it wasn't solving this in, by any means, in the dumbest, most low-level way. Like, we're past that point as well. They used to solve some problems in, like, kind of surprisingly low-level, like, grind it out kind of ways.
33:25Greg Burnham:Like, that's not really what these look like anymore. but like the high level, a higher level version of that still is maybe what you might say here. So, and I like the frame of precise and intricate computations here. It's important, I'll say one more thing about the Navier-Stokes solution. It's important to recognize that humans sort of put in place a lot of the structure through previous work that AI then went and as far as we can tell, I mean, humans are still figuring out exactly what goes into the AI solution. But as far as we can tell, it builds on a lot of prior human work. The AI didn't like build it up from the ground.
34:03Greg Burnham:That kind of theory building, as mathematicians call it, we still don't see that very much from AI systems. Last point I'll make before pausing on this is just remember its point in time. Like, remember that trajectory. We've seen very rapid. so mathematicians who are saying we have to adapt our profession to the fact that now you can push a button and get what used to have been a career-making result like you know are also I think cautioning and we can expect the capabilities will freeze right now we we see you know that that linear whether it's linear or exponential as a matter of perspective but it's going up and it's probably like we see no reason to expect it to stop there so this is a big thing Epoch is looking for.
34:47Greg Burnham:Can AI move from intricate computations plus lots of background knowledge to more of a developing new ideas, more from whole cloth? And if so, and that's a big thing we're watching for in our own math benchmarking activities. Yeah, it's interesting to think that at least the way I characterize this as brute force versus pulling in things from other domains those are not mutually exclusive. Like for an AI that has access to all of the research from every domain, an approach that could be very fruitful is to brute force the exploration of adjacent domains and pull that into the solution to a given problem.
35:38Greg Burnham:That's it, I think, a lot of what we see. And maybe just to say the other thing that would really supercharge them and bring them closer to full human parity or even dominating human capabilities would be something like coming up with a new idea. And maybe like in calculus, the idea of the derivative or the integral of a function. That was something that maybe mathematicians in the 1500s had been grasping for this idea, and then Newton or Leibniz or whomever comes up with, like, hey, here's like, I can describe a new thing, a new mathematical concept, and that unlocks a lot for us. That lets us solve problems we previously couldn't solve.
36:24Greg Burnham:And so that's the sort of thing we don't yet see AI doing. It's also the sort of thing that's rarer for humans to do. A lot of human math progress has been, well, I sat down and I put in the work and I did the computation and I worked out which was the dead end and which was the good path, and here's the answer. Now, XYZ is true. I've proven it. Or I studied one area and I noticed something and I thought it might be applicable to this nominally, superficially different area. Humans do that sort of stuff all the time. Maybe that's 90%, vaguely speaking, of human mathematical work from what I hear from mathematicians.
36:58Greg Burnham:The new idea stuff is rarer, so maybe we should expect it to be not the first thing AI gets good at. But so far, we don't see AI really doing that yet. How easy or difficult is it to define a new idea in such a way that we can easily distinguish them? Like, you know, if we think about kind of the evolution of this discourse around like, is it AI creative? You know, initially it was like, yeah, it's very creative. It created this like poem and, you know, and pirate like that's creativity. That's, you know, that was a new idea. But that's certainly not what we're talking about when we're talking about like advancing science and advancing mathematics, et cetera.
37:47you know so a how do we like how do we define that but also you know from your perspective as someone who's studying this is it like you know we've got this long line of folks that are coming that says here's this ai with a new idea and you're like yeah well let me evaluate it against my new idea criteria no it's not really that or are we actually like is it pretty clear to everyone that the AI is not having new ideas and like we're waiting for the first person to come and say that AI has the new idea? Like, how do you think about that whole space?
38:22Greg Burnham:Very, very challenging nut to crack. So we, you know, try to approach it from different angles. One thing you can try to say is, I've got a problem. I don't know if a new idea will be required to crack it. And this could be a problem in math or in science or any quantitative area. But humans have tried to crack it. Humans have tried to make this number go up, whether that number is the efficacy of some drug at fighting some disease or something like that, or it's a math result of some conjecture that has stumped mathematicians for a long time. And you can say, okay, well, I don't know if, I'm not sure how I'm going to measure a new idea, but I can at least tell if this problem has been solved or this number has been made to go up or whatever the concrete metric is.
39:17Greg Burnham:And you can then, then you're, you can at least check if that has happened. Now, I was probably a year ago more optimistic that there would be a tighter relationship between math problem that humans have failed to solve and post hoc qualitative assessment of creativity. But, you know, this is sort of the combination, the like, you know, double punch we're trying here. First, you set a target that's hard on human terms that humans care about. Humans have tried to be creative to solve. And then you do post hoc review with experts and say, what do you make of the solution? Why hadn't you found it?
39:58Greg Burnham:No offense. You know, what is in the AI solution that humans had missed and so on. And this is risky, like risky in terms of humans being honest about it, because, you know, your ego might be wrapped up in it, or you might, once you see the solution, it feels more obvious in hindsight. So would you have described it? Like, very risky. But you still, you try your best. And, you know, you might say things like, well, how, like, here's an interesting approach that we haven't operationalized in detail, but it's a nice heuristic, I think. How big a hint would a human have needed to solve this in hindsight?
40:37Greg Burnham:So, for example, that unit distance, a very prominent mathematician, Fields Medalist Timothy Gowers, did this sort of exercise on this unit distance problem once AI had solved it saying, how big a hint and how specific of a hint do I think a mathematician would have needed now in hindsight to solve this problem? He came away with an interesting exploration of, well, actually AI sort of disproved this conjecture as opposed to proving it true. It said, actually, it's false. Even that is a huge hint. Humans had mostly thought it was true. They were wrong, but they'd been trying to prove that it was true instead of looking for a counterexample.
41:17Greg Burnham:So even just the hint of try to find a counterexample is pretty big. And then there was like one more piece. And he was like, I think, though, very hard to say, we can't really do the counterfactual experiment. I think with just that, with like this sort of two part hint, it would have been possible for like a human would have been like, oh, okay, I know what to do now, blah, blah, blah. And like probably would have at least hastened a human solution. So you can do things like that, maybe still very qualitative, but you try to be honest with yourself and be rigorous. And you could imagine doing a more rigorous version of this experiment.
41:48Greg Burnham:And anyway, things like that. But then the fact is math is still, for the most part, a sort of discipline removed from physical real world impact. Not always, but its impact is often hard to predict. There are domains where impact is not at all hard to predict. Like I'm in the chemistry lab and I'm trying to improve the yield of my synthesis process. And if AI makes incremental improvements. I expect incremental yield. And if AI suddenly comes and says, no, no, no, you're doing it all wrong. Let me like redesign your whole thing from scratch. I've got like a, you know, maybe creative idea. Well, at some point you don't care if it's creative or not in the qualitative sense.
42:26Greg Burnham:You just care that you saw discontinuity in the yield of your chemical substance or whatever. And that's, you know, that's a bigger impact for the world anyway. What I took from that is that we filter on the problem and how hard humans have kind of banged their heads against the problem. Then we look at the results that AI has come up with and try to kind of subjectively say, you know, is this based on a new idea or is it more brute force or something else like. But it's kind of imperfect and messy and like we're not really sure, but we're we don't think that we've come up with, you know, great examples of AIs coming up with you know, quote unquote new ideas.
43:17Greg Burnham:That's exactly right. One reference point I'd give is way back in, gosh, well, the teens, I forget which year, the famous Go playing system, the game of Go, AlphaGo. I had that thought earlier as we were discussing this. Like, if I remember correctly,
43:37part of the solution was like, Like, oh, that was an innovative move that no human ever would have done.
43:43Greg Burnham:There was a specific one. Yeah, it was a very specific move in there. Yeah. Yeah. They call it move 37 because it was the 37th move in this game where. Yeah, apparently this was a very surprising move. Like you said, no human would have done it in this particular position in this in this game of Go. And I think it's encouraging for our admittedly quite subjective program here that Go experts looked at it and were like, oh, my goodness, that's new. And then it also became clear that that was pivotal, that that move gave the AI system playing Go a decisive advantage. And so that's sort of what we're looking for in math.
44:27Greg Burnham:And at least maybe we have some hope that humans will be able to spot it when it happens. uh but but yeah that's that's what we haven't seen quite yet and again the trajectory is fast the things are always a little more qualitative than you know that you know than they are uh or continuous than discrete uh but you know so anyway this is this is where we are now which is kind of wild in its own right earlier when we were talking about frontier math you kind of characterize it as you know, one to four, I think, and then advanced and we never circled back to your more recent work? Yeah, we've got a couple, two new sets of frontier math problems.
45:08Greg Burnham:We ditched the tier system, whatever. So now these are, I'll lump them together because it really is the, this is our main tool. So we have, you know, our latest iteration on frontier math consists of a bunch of problems that humans have tried and failed to solve. They're curated to be of a special interest to mathematicians, where, and some of them we've characterized even on a bit of a scale from, like, this would be moderately interesting, or this would be, like, maybe what that means is two teams of mathematicians have tried this problem, it's at least 10 years old, but it's not central to any research program.
45:52Greg Burnham:It's just something that people would be happy to get an answer to. All the way up to like a major breakthrough. Four tiers here. Breakthrough is the highest. None of the breakthroughs have been solved yet, but don't get me wrong. If Navier Stokes had been in this problem set, it would surely be a breakthrough there. Regardless of the fact that humans had like done a lot of work and gotten maybe within spitting distance of the solution before AI finished it off. But yeah, we really view this as just a way of finding AI, a way of finding cases where AI might have had a creative idea, because we can tell easily whether it solved these problems.
46:32Greg Burnham:this is maybe the big piece of work on our side, was to make it easy to verify automatically whether an AI system has solved this problem. It's not obvious when you think about it that no human knows the answer. So how do you tell without a human looking if AI has found the answer? But there's a couple techniques you can use for this, and we employ different techniques for this. It can go into that, but it's something of a technical detail. And so when a new model comes out, Astra came out, we hit go. And it's like, oh, it solved three new problems from our list. Okay, let's go look at the solutions.
47:05Greg Burnham:Let's share them with mathematicians. Let's see what's going on here. And then we analyzed that with the help of mathematicians, analyzed that for like, what kind of qualitatively, what kind of solution was this? What does it mean when the models are so good that we have to shift from, you know, these benchmarks that we understand to like, these collections of problems that we just don't have the answers to? It's definitely a phase transition in the science of AI capabilities measurement, which is what I work on. But it doesn't change the game all that much. It just means we have to look for harder problems.
47:44I think I joke sometimes that even certain forms of optimization problems, like allocating scarce resources efficiently, that's a benchmark of sorts that will survive the singularity.
47:56Greg Burnham:Like even if AI is like off going crazy, running rampant in the galaxy, it'll still have to, you know, it'll still have to decide how to, you know, optimize. So optimization is always going to be a benchmark, even if we're superhuman. And basically, it used to just be we knew the answers ahead of time. Now we don't. So it takes this extra trick of, OK, how do we tell when AI has gotten something right? But this isn't insurmountable, especially like I was saying with numbers in a scientific context. It's intuitive that, OK, the best humans were able to do was X. But we just say to AI, hey, try to do Y greater than X.
48:34Greg Burnham:And if it does, those numbers go on and on. So it's no problem. But yeah, you're right to sort of note the phase shift and to be somewhat surprised by it. And is there a methodology, a concrete methodology for assessing and awarding partial credit? Is that part of the way you think about benchmarks in this post-phase shift? Yeah, it's a good question. It depends on the benchmark. So sometimes it's something like, well, you just want to make this number go up and the higher the better. So partial credit is very continuous and very, very natural. in those cases. For math problems, sometimes it's just like, well, what we care about is, is, you know, if X does Y always hold, like, is this, is this always true?
49:21Greg Burnham:And then it's kind of, sometimes humans will find and publish partial results for the AI systems. We usually just say, you know, all or nothing, but we have a large enough collection of such conjectures that, you know, we hope to have some, you know, smooth gradient of, well, okay, this one got two more out of 100 or something. And so it looks like a pretty smooth signal. That's typically how we approach these things. Yeah, across benchmarks. Yeah, maybe what I was thinking was, once you're in the regime of very challenging, unsolved problems, is it possible that AI systems are doing surprising or novel or creative things, but still not quite getting them to the full, the full solution?
50:08And do we care about that? Do we want to track that? And is that part of the way you've constructed the benchmark?
50:15Greg Burnham:It's a great question. It's not central to how we construct the benchmark. It's certainly worth tracking. The problem is it's just labor intensive. You can have AI review the AI. We do this occasionally. Say, which problems did it make partial progress on? What if we let it think for longer on those problems? And nothing yet. It doesn't seem good at understanding how close it is to a solution. Humans aren't necessarily very good at this either. I did want to mention that while we focused on maybe the area where AI is the very best, like has come the farthest, the fastest math, we do see lots of areas where we also think it's important to measure AI capabilities, where the progress is somewhat fuzzier, less clear, and certainly less superhuman, even in some, well, maybe.
51:08Greg Burnham:be. Still superhuman in subways, but the ways in which it's not yet are maybe more obvious and feel less rarefied. What are some examples? A couple examples. One area that forthcoming work we're excited about is just, it's a question everyone must be wondering, can AI take my job? So our job at EPOC often involves many things and all sorts of research, not just AI benchmarking, but other topics about trends in AI and writing research reports and doing data analysis, making infographics that make a point in an easy to digest way. We've been collecting examples of these from within Epoch, where the judgment at the end of did the AI system manage to do this Epoch task is too messy for traditional benchmarking.
52:04Greg Burnham:So it wouldn't be obvious, like, how to say, did it get the right answer or not? It's not about making a number go up. It's not about satisfying some logical chain of deductions. It's like, here's the infographic to go along with this report. Like, does it meet our style guide? And does it look good? And does it, you know, is it better or worse than the human reference that we, you know, one of my colleagues did? um and here we see like an interesting mix of uh some things i think we we sort of all our experience is like yeah if you ask it for a summary of the literature it's going to do a pretty good job it's like very good at that but but there's other tasks where it's um you know hard like creating compelling visuals that follow our style guide or this is a good one coming up with a new short-form data-driven insight.
52:59Greg Burnham:We call these data insights. You can find them on our website. And these, often we try to boil down a single important observation into a single chart and a single sentence. And this is hard, but we don't see them doing a great job at this. Come up with a new project for EPOC to undertake and do a prototype, very open-ended, not so good at this. Like these more open-ended tasks, we do see, again, my guess is if we will do this, go back in time some and see, have models gotten better at this? My guess is we'll see progress. But the level compared to like humans is still a little, still a little meh.
53:41Greg Burnham:I want to give one other example, which is learning on the fly. This is another case relevant for can AI take my job? Because you weren't great at your job on day one, but as you practiced, you got better, and humans have this sort of, they learn on the fly. We take a, we like to measure this for AI. We have a benchmark that's taking a complicated board game. We did a human study for saying, how many times do humans have to play this board game in a row before they sort of figure out how it works and can get a very high score on it? We do the same with AI. We said, play it once, take notes, same AI instance, like play it again, take notes, play it again, take notes, and see how you can do.
54:26Greg Burnham:AI systems have gotten good at this board game, but not by playing it repeatedly. It's more of an intergenerational thing, like GPT 5.6 scores X out of the box and then stays flat on its playthroughs. GPT 6 scores much higher. It's like much better at it, but it's still out of the box. And then it stays flat in its playthroughs and is still less than the top human. So this kind of on-the-fly learning is something we also don't see AI systems as being good enough at, even to just pick up a board game, let alone a messier job, a real-world job like that. So these are sorts of things where we don't see these capabilities and we think it's important to try to catch when they emerge.
55:05My gut on the ladder, the AI learning on the fly is that it's maybe a harness solvable challenge versus a model solvable challenge. And maybe, you know, no one's focused on that. like with the right set of memory structures and like some kind of sidecar heuristic thing like you know we can probably make a big step improvement in an AI's ability to like you know improve game to game but it's not uh it might not require dramatic model changes Do you think about that? And is, do you have, you know, in this domain or in other domains, like formal or structured ways of thinking about, you know, what, you know, model versus harness, you know, granted, this whole model versus harness thing in and of itself is, you know, practically brand new.
56:07but it's proving to be important in distinguishing where these improvements come from.
56:15Greg Burnham:Yeah, definitely. We've done some experiments on this for the board game case in particular, where we've tried a very simple harness, the first party harnesses like Claude Code or Codex. We've tried multi-agent setups. We give them note-taking tools and make it clear, like, okay, your goal is to learn the game so you can do really well in your final plays of the game. So explore, take notes, figure out. We try to prompt them well. I mean, we want to find good results if we can because that's a big deal for the world. The game is a test bed, but real jobs would also benefit significantly. Risks would also go up if this capability, you know, you test in the lab, is it a great hacker?
57:05Greg Burnham:You find maybe it is, maybe it isn't, but you find that it isn't, and yet it can learn on the fly, then okay, once you deploy it, like, look out. So anyway, we've tried pretty hard and have mostly not found anything that cracks this, that cracks this on the fly learning. There's been some general research about this as well, not just us, about, you know, people taught, like there was, people still use these so-called skills where you like just write some notes to an AI system of, look, when you do work I ask you to do, you're going to need to use this tool or produce this kind of output. And like, here's my notes on how you could do that.
57:41Greg Burnham:Well, so you don't have to like learn it on the, like pick it up fresh every time. So use these skills and these show some improvement, but then there's like a plateau, like pretty quickly, at least in the research I'm familiar with of like, you can get some improvements. My senses, this is very loosely speaking, it's like low-hanging fruit like if you're telling it oh you're gonna have to navigate this like finicky website that my company makes me use and like just you're gonna get stuck and like you have to look in this menu and that's where the button you're looking for like straightforward advice that's like obviously useful but when it comes to higher level strategic uh thinking kind of hard to bake that into the harness uh where or or like where where the the the board game plays like like has this, you know, sort of brings this out where it's like, there's not any perfect advice that, like, you have to find the advice for yourself, in a sense, like, you have to learn the game, like, so I will say, if we give it extremely detailed strategy guides for, here's how you beat this game, like, here's everything you could possibly want, like, it's a cheat by humans, like, no one would play the game that way.
58:50Greg Burnham:They do very, they do very well, they do much better than if they basically have to find that for themselves. Right, creating the strategy guide. They can't create the strategy guide for themselves, which is like pretty interesting. So like we do expect eventually they'll get better at this, but it's again, it's a big deal if that happens because now it means, okay, go back to that first benchmark suite. Okay, they're not great at making charts for us or coming up with data insights for us, you know, without guidance. but now put them in on the team for a little while and let them learn on the fly and maybe they do become good.
59:29Greg Burnham:So you sort of see these as a combination. Like if on our board game where it's easy to measure, they're not getting better over time, then it's okay for us over in the real world, like human in the loop evaluation studies for us to just do one try and say, okay, well, they weren't great and we don't think they learn on the, so that, you know, but really, you know, those go together. And is the board game a synthetic one that you created for this challenge? Or is it a real board game that people play? The real board game, it's called, or the first one we did this with, we'll do this with more, is called Earthborn Rangers.
1:00:04Greg Burnham:It's fairly obscure. If you go to like boardgamegeek.com, it's like 500th most popular or something. So it's not like up there. It's a small devoted community to it. And AI systems don't seem to know that much about it. Like they've heard of it. They can give you some facts about it, but it doesn't seem like something they've been trained on specifically. It's impossible for us to tell this with high certainty, with high confidence, but that's sort of why we chose it. It's just hard to make these things from scratch and know whether they're hard or whether they're broken. So that's why we, but the risk, which you might have in mind is, what if AI systems like, you know, just know about it from their training data.
1:00:45Greg Burnham:It's a risk on our mind. It's just a risk. One thing we will do in the future is we're moving on to video games. Incidentally, it's another area where AI is like very much not like human reaction time and visual and spatial reasoning. Again, coming along fast, but not really there yet. So video games, you can always test on a new game, like a game that just came out, not in the training data probably, probably, or at least, you know, not as robustly in the training data. And so we'll be testing, can AI systems do well on those? Do they do better on like, you know, the older version of a video game than like, you know, Grand Theft Auto 4 versus 5 or whatever, like when GTA 6, whatever it is, comes out, like, I don't think we'll be doing this one, but, you know, are they worse at that?
1:01:32It's interesting when you bring up video games, it immediately calls to mind kind of the classical reinforcement learning approaches, which is kind of learning on the fly, but kind of different.
1:01:44Greg Burnham:Yeah, we call the second sometimes in context learning, where the only like lever the AI system really has is managing its context. And like, look, when we give it the strategy guides, and they do much better, clearly, this is a powerful lever, but it's not as powerful as updating the weights. And like, so yeah, I mean, like, what's the big picture here? It's, it's, if you have this fast loop of in-context on-the-fly learning, then AI capabilities might sort of improve much more rapidly than the slower loop of collecting more training data, doing more reinforcement learning, which is done on language models as well these days, and then deploying a new model.
1:02:25Greg Burnham:Now, AI might automate that whole loop, but it'll still be slower than if it can just, you know, work out its own context on the fly, yeah. Yeah, it's interesting when you zoom out on some of the questions that you focus on, and we haven't talked explicitly about it, but you have a blog post where you kind of write out these nine big hairy questions that you're tracking. They kind of group into, like, how do I assess where AI is now? That's like, can I do my job? Like, how is it, you know, competitively? and, you know, where's the puck going? Like, can it learn on the fly? You know, can it do research?
1:03:07Can it come up with new ideas? These are all, it is kind of where it is now, but also like what, you know, is it positioned to like dramatically self-improve or like, you know, what's the trajectory?
1:03:22Greg Burnham:That's exactly right. The, I mean, I think for people thinking more broadly about what they need to know about AI. This is, the second topic is, you know, just as important as the first. But it's really, you know, even if you take some of the recent incidents, like the hugging face incident, like the sort of think of two axes, alignment of is the AI doing what I want it to do holistically? Obviously hacking into some other companies' computers is not well aligned. And then capabilities, just like, what can it get done? Can it hack into, you know, reasonably well-secured other companies' computers?
1:04:09Greg Burnham:What that incident was, was like an early example, very, like, I think pretty robust example of, well, we don't have the alignment thing solved, and capabilities are getting so big, so strong, that this is becoming a problem. And so paying attention to capabilities like what we should expect is becoming more important sort of day to day. And we're just trying to produce benchmarks that both say, where are we now? Exactly like you said, like, what can it do now? This is maybe relevant for more mundane economic utility or whatever, as well as, you know, do we see the dynamics of the system changing in a way that might be alarming, basically.
1:04:53I'm kind of also thinking about it as like first and second derivative of capability.
1:04:58Greg Burnham:Yeah, I think that's right. Some of these things make it the slope goes, you know, much faster, hyperbolic growth or whatever in some of these cases versus merely just like, yep, it's growing and it continues to grow. Both are important. I mean, I think we've seen the disruption we've seen to date has come from these kind of unpredictable thresholds. Like when, you know, GPT-2, GPT-3, suddenly you had the chat GPT moment and you're like, oh, this thing, I want to talk to this thing. I might use this thing instead of a search engine. This thing is useful to me for finding information. And then like a little while years later, whatever, two years later, you had the, it's getting better slowly and steadily at helping me with coding or with like using my computer, but suddenly - And thinking models.
1:05:47Greg Burnham:With the thinking models. And then you had the clawed code moment where it was like, oh, suddenly this thing can actually do software projects for me, even if I'm not an engineer myself. And it's hard to predict when these thresholds will be crossed. So it's useful to still track, but it's useful to track the mundane capabilities today and say, look, cyber is maybe the latest example, where it's like, look, we see their ability in controlled settings to do hacking going up smoothly. And at some point, it's going to cross some threshold where, oh, shoot, it can hack Hugging face or whatever. Hard to predict where those thresholds are, but still very useful to track those capabilities.
1:06:26Greg Burnham:So even the first derivative, as you said, version of this, we find very useful. We've got a fun benchmark we hope to release, maybe by the time this is published, we'll have it out on a furniture assembly where we build some IKEA furniture, make some mistakes and see if the AI system can like just given photos, like say, oh, like, wait, you made a mistake and step back in step four, like stop, go back. And they've, you know, we see the same smooth capabilities like Astra, GPT-6 Astra, pretty good at this, actually. And, you know, but it wasn't out of nowhere. It was like from a smooth and this is, you know, important because we haven't seen AI have much of an impact in physical industry yet.
1:07:09Greg Burnham:Like it's been mostly digital. and even before robotics hits the scene, we might expect that, you know, Claude in your glasses or something is saying like, hey, you know, I'm going to guide you through repairing your car or whatever. I'm going to help you fix this machine in the factory that broke. So, you know, we want to track those capabilities too, just for the mundane impact. Anyway, impact all over the place coming. Benchmarks help us say so. Awesome, awesome. Well, Greg, thanks so much for jumping on and sharing a bit about what you're working on. It's super interesting stuff and a very thought-provoking way of thinking about, you know, AI and where it is, where it's going.
1:07:48Greg Burnham:Thanks, Sam. Really appreciate it. Awesome. Thank you.
1:08:09you
From the publisher
AI systems have gone from struggling with grade-school math to helping solve research problems that have resisted mathematicians for decades, including Navier-Stokes.
In this episode, Greg Burnham, who leads AI capabilities research at Epoch AI, joins us to examine what that progress says about where AI is going. We look at how these systems are solving hard math problems, how much they rely on persistence and prior human work, and whether they are starting to produce genuinely new ideas.
We also discuss how to measure progress as traditional benchmarks become less useful, why capability gains appear surprisingly steady across model generations, and where models still struggle with open-ended work, learning from experience, and identifying promising new research directions.
🗒️ Full show notes: https://twimlai.com/go/778.




