In short
AI’s rapid jump in solving advanced, proof-based mathematics (despite still failing at basic arithmetic) is triggering an existential crisis about what mathematicians do, how math knowledge is produced, and whether AI will “mow down” open problems without expanding the field. The episode centers on OpenAI’s internal model “Astra” and its “10 advances in mathematics and theoretical computer science,” plus broader concerns about funding, training, verification, and hype/attribution.
Guests
Robert Hart, London-based AI reporter at The Verge. He interviews multiple leading mathematicians (including Fields Medal winner James Maynard) and cites others such as Gary Marcus, Johannes Schmidt, and Oxford’s Andras Juhas; also references mathematicians who signed the Leiden Declaration.
Key claims
AI can generate verifiable proofs (e.g., via Lean) but still can’t reliably do elementary counting/time; repeatability and provenance of model methods are unclear; crediting/attribution was mixed; the field may become sterile if AI solves questions without creating new ones.
Notable examples
OpenAI’s Astra “10 advances”; the unit distance conjecture; “strawberry” counting failures; Lean proof-checking; Watson’s failed cancer attempt as a cautionary analogy.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOAI's Effect on Mathematics
2:58 to 4:06
Exploring the existential crisis mathematicians face due to advancements in AI.
“Robert Hart, you are a London-based AI reporter here at The Verge.”
AI Advancements in Math
4:06 to 5:00
Discussing AI's improvement in solving advanced mathematical problems.
“And it's gone from being very terrible to seemingly genuinely quite good at a professional level in a very short space of time.”
The Range of Mathematical Skills
5:00 to 8:00
Delving into the varied capabilities of AI in different mathematical disciplines.
“My conspiracy theory is that the strawberry thing is hard-coded into all the models.”
The Astra Model Breakthrough
8:00 to 8:54
Robert Hart explains Astra's breakthrough in solving significant math problems.
“I think that tracks broadly with sort of the rise of AI in every field.”
The Astra Model Breakthrough
10:28 to 11:21
Robert Hart explains Astra's breakthrough in solving significant math problems.
“We need to pause here for a quick break.”
The Astra Model Breakthrough
11:31 to 12:52
Robert Hart explains Astra's breakthrough in solving significant math problems.
“Your innovation and how you protect it depends on the technology you choose to trust.”
Response to Astra's Achievements
13:18 to 14:00
Discussion on the mixed reactions to OpenAI's Astra model results.
“We're back with The Verge's Robert Hart, talking about AI's math capabilities.”
Attribution in AI and Academic Standards
14:00 to 14:40
Discusses issues of credit and attribution in AI research compared to human researchers.
“There was one that drew particular attention for how they credited it and also how they'd kind of announced all of this in their blog post.”
Evaluating AI's Problem Solving
14:40 to 16:00
Explores the legitimacy of AI's breakthroughs in mathematics and the repeatability of results.
“So yeah, it was impressed with what it did, if not the credit for some of it.”
Understanding Mathematical Proofs
16:00 to 17:40
Examines the nature of mathematical proofs and the implications of AI's involvement in generating them.
“I think it was just a poor press release, to be honest.”
Show all 21 chapters
Concerns Over AI's Impact on Math
17:40 to 20:00
Discusses fears among mathematicians regarding AI potentially outpacing human contributions.
“I keep saying prove a lot here, but it will test the rigor and the assumptions of everything going on there.”
The Value of Mathematical Inquiry
20:00 to 22:45
Investigates the philosophical implications of AI in mathematics and the nature of mathematical exploration.
“What do we know about how the models did this?”
The Shock of Rapid AI Advancements
22:45 to 24:45
Describes the shock within the mathematical community due to rapid advancements in AI capabilities.
“I think it was James Maynard that said that the most interesting kind of discoveries in this isn't that you've solved something, it's what evolves from that.”
The Shock of Rapid AI Advancements
26:15 to 26:55
Describes the shock within the mathematical community due to rapid advancements in AI capabilities.
“back in just a minute excuses are easy an epic movie night we don't have enough snacks dinner party with the girls?”
Introduction to the AI Crisis in Math
28:00 to 28:20
Discussion of concerns mathematicians have about AI's impact on their field.
“From athletic to athletic-ish, Sierra's got it.”
Skepticism and Generalization in AI
28:20 to 29:39
Exploration of skepticism around AI's capabilities and its generalizations across disciplines.
“Gary Marcus is a reliable skeptic of AI.”
The Impact of AI on Mathematics
29:39 to 31:07
Discussion on how AI might democratize access to mathematics and the reactions from mathematicians.
“And I mean, that said, there has been an undeniable trajectory, I think, in the last few years of a broadening capability increase.”
Concerns Over AI-Generated Research
31:07 to 34:02
Mathematicians express worries about the quality and legitimacy of AI-generated papers.
“One of the bigger arguments about AI in general is that it democratizes access.”
Resistance Against AI Hype in Mathematics
34:02 to 36:30
Mathematicians' organized efforts to resist AI hype and its implications for the field.
“your piece about the nature of the AI labs and how they are talking about math.”
The Future of Mathematics in the Age of AI
36:30 to 39:26
Concerns about the future of the mathematics field and the potential economic implications.
“And that turns into, I don't know, yet another way to launch rockets.”
Optimism and Uncertainty Among Mathematicians
39:26 to 41:20
Balancing optimism about AI's benefits with uncertainty about its effects on the profession.
“One great thing about The Verge is our commenters are vast.”
Transcript
Automatic transcript. May contain errors.0:00Nilay Patel:Support for the show comes from ServiceNow. AI is moving fast across the enterprise, but without visibility, it's just chaos. Different tools, different models, different teams using AI in completely different ways. ServiceNow turns that chaos into control. With the AI control tower, you see all your AI across the business in one place. What it's doing, what it's done, and what it's about to do. So you stay in control. To put AI to work for people, visit servicenow.com.
0:57Nilay Patel:Fancy a chat? Come try all the voices free through September 12th at deepgram.com slash keep talking. Terms apply. We secure the world's AI future by securing the world's AI leaders. AI is being adopted faster than companies can secure it. Sensitive data is moving across AI models and attackers are using AI to strike with unprecedented speed and precision. That's why CrowdStrike built the agentic cybersecurity platform for the AI era. To secure AI and to stop breaches. So you can embrace AI with confidence, not compromise. CrowdStrike. We secure AI. We stop breaches. Hello and welcome to Decoder.
1:38Nilay Patel:I'm Neil Iyipotel, editor-in-chief of The Verge, and Decoder is my show about big ideas and other problems. Today I'm talking to Robert Hart, The Verge's London-based AI reporter, about what AI is doing to the field of mathematics and the existential crisis that many leading mathematicians are having about it. The news here is that OpenAI just published a set of solutions to longstanding problems in math that went off like a bombshell in the field. It's caused a huge debate in the math community, and Rob talked to some of the most accomplished mathematicians of our time about it. It's funny that AI systems are still pretty bad at elementary school arithmetic, but they're getting increasingly good at very high-end abstract math.
2:14Nilay Patel:All of this raises some big questions for the field of advanced mathematics. If AI can do math of this caliber, does that mean AI labs can transfer those skills to other domains? What good are the academic grants and university programs designed to train new generations of human mathematicians to identify new problems as they try to solve existing ones, if the frontier models simply solve all of the outstanding questions? And perhaps most importantly, what if all this attention to math is just another marketing exercise for the Frontier AI labs, who don't seem to care less about what happens to one of the oldest and most fundamental academic disciplines there is.
2:49Nilay Patel:There's a lot here, and Rob talked to a lot of people with a lot of views on all of it. Okay, Verge AI reporter Robert Hart on what AI is doing to math. Here we go.
3:11Nilay Patel:Robert Hart, you are a London-based AI reporter here at The Verge. Welcome to Decoder. Thank you for having me. I am very excited to talk to you. There's a lot going on in particular with AI and math that you recently dove into. You spoke to a lot of leading mathematicians about the crisis in mathematics due to AI. It feels like a lot and also like there's a lot yet to know and discover about the interaction of these two things. A full existential crisis, which is pure Decoder bait. Broadly, tell us what's going on. I think a lot sums it up quite well. I mean, basically, a bit of an existential crisis within what is mathematics, what are mathematicians doing, what is the role of mathematicians going forward.
3:55And a lot of that has been spurred by kind of a phrase transition in what AI is capable that has kind of, I think, exploded, would be a reasonable way of saying, in the last six months to a year. And it's gone from being very terrible to seemingly genuinely quite good at a professional level in a very short space of time. And so I think it's a lot of what these other fields have been struggling to deal with for the last five years in a very compressed period of time.
4:24Nilay Patel:I would put that next to software engineering. We've been living through the AI crisis in software engineering for some amount of time. But as recently as 2024, even last year, the conventional wisdom was that AI models were particularly bad at math. The famous example is they could not count the number of R's in the word strawberry. Even just counting sort of eluded them. What has happened to make them better at math? Are they still bad at general arithmetic and they're good at advanced math or is it something in between? Yeah, I mean they are still truly, truly terrible at some areas of math.
4:55I did check. They can do strawberries now. I think someone's tweaked that.
4:59Nilay Patel:I think strawberry is hard-coded. I want to be very clear. My conspiracy theory is that the strawberry thing is hard-coded into all the models. I think so, too. That is a conspiracy I'll buy into. But yeah, I mean, it's still terrible at those kind of things. I mean, it's math. It's arithmetic. I mean, even the days of the week, my boyfriend was saying the other day, he's like, it keeps thinking it's Wednesday. It's not Wednesday. Or time. I mean, Alyssa Welly for us a few months ago, I think, wrote the chat GPT, can't tell time still can't that's not all of maths like so there's this sort of disconnect i think we always kind of equate or you to be good at math you've got to be good at counting or adding or multiplying and a lot of it is reasoning like i mean if you look at academic maths papers a lot of the time you won't see numbers which kind of sums that one up i think so they're still terrible but they're now also very good at this other part and as to why i mean i think at some point you reach a critical mass of what these systems can do and we've seen it as we said with writing we've seen it with programming and they're very good at kind of forging connections between different areas or applying old methods in new ways or those kind of things and it appears that the newer models they're training have apparently reached that level where it clicks and now it can do maths i mean it's important to say as well like it's we speak of it and especially from the outside as a sort of unitary discipline but i mean imagine say biology you've got something that would range from like literally watching animals and describing behavior all the way through to like cellular mechanisms and biochemistry like math is not a unitary discipline either some bits is really good at some bits like counting still really bad um and even on the more kind of abstract levels of that i mean some experts told me that they floated topology is one area that it's apparently still quite bad at i can't verify that to be honest it's beyond my own expertise but um it's still, yeah, it's a bit of a mixed bag.
6:54Nilay Patel:So you're describing mathematics as a huge field, obviously, many, many academic areas of interest. And there are some parts where the models have gotten quite good. Some parts, maybe the basic parts that people think of as math, which is simply counting where they're still struggling. And there's a wide range in the middle. Is it the wide range in the middle where the existential crisis is? People don't know what's going to happen. Or is it at the parts where it's really good? Bit of both, which I feel is going to be a running theme through this. I mean, no one's really afraid of it being a mediocre mathematician, but obviously there's a huge element of what this field does.
7:26And like in terms of the research elements, like the cutting edge kind of, as we see with a lot of the results that generate hype, like what can it do? There are areas now where it seems to be producing work that is on par with good mathematicians, alongside other parts where, yeah, it can't count. So I think it's, and all caught up in that is whether it's going to kind of rewrite Right. Employment structures, funding structures. I mean, you also raise the murky question of like, what is mathematical knowledge? And the roles that these workers will be doing as well. So it's kind of all of that wrapped into one.
8:00Nilay Patel:I think that tracks broadly with sort of the rise of AI in every field. Where you can just add horsepower or compute to a problem and there's some kind of verifiability, it seems like the models continue to get better. Everything in the middle where you might need some world knowledge or the models might need some actual intelligence about the world itself, they seem to struggle. The parts of math, at least reported out in your piece and what the labs are talking about, they seem to be almost entirely self-contained theoretical problems where the models can generate a proof or solve a problem that no one's been able to solve and then try to verify that that has existed and it can just run it again and again and again.
8:36Nilay Patel:That brings us, I think, to May of this year, where an internal OpenAI model, which we have not really seen, disproved the unit distance conjecture, which is an 80-year-old problem. And then just recently we heard about Astra from OpenAI. Astra is the one where it seemed like the switch flipped and everyone decided it was an existential crisis. What did Astra achieve and why is it a big deal? I'm also pretty sure that Astra was probably behind the earlier one as well. OpenAI just listed it as an unnamed internal model. It's probably Astra. they didn't answer me when I asked but there we go open ai a few weeks ago dropped a blog along with a lot of paperwork um proving it I think several hundred pages that they described they called it 10 advances in mathematics and theoretical computer science it was basically an array of disciplines that they claimed the newest model astra had solved in some capacity that ranged from I think one was in quantum game theory uh which I don't know how to begin to explain and even less how to explain is kind of there was seer packing in higher dimensions and warm three dimensions and there was sort of a lot of other different disciplines as well and it was i mean yeah it caused a lot of stir in the community it was i think a bit of a bombshell i mean as we'd said there'd been these individual breakthroughs that happened but to drop 10 in one go and they were quite big ones i mean researchers had sort of told me that yeah these were if a human had solved these, we would be impressed.
10:05If a human had solved all 10, we probably wouldn't believe it. Their problems they actually care about as well, I think, is an important one. A lot of previous ones have been accused of areas that mathematicians didn't really bother with. And these are ones that mathematicians, good mathematicians, have spent a lot of time trying to solve and had them.
10:28Nilay Patel:We need to pause here for a quick break. We'll be right back.
10:39Nilay Patel:Support for the show comes from ServiceNow. AI was supposed to handle the parts of the job you hate. Instead, it just describes them, suggests what to do about them, and then leaves you to do it. That's not help. That's homework. ServiceNow's AI specialists are different. They're not a tool. Think of them as digital teammates who actually do the work from start to finish. Cases get resolved, requests get processed, loops get closed, and most importantly, no extra work for you. Because when you can truly delegate to AI, you can get back to the work only you can do. The work that requires a person with ideas and judgment.
11:19Nilay Patel:And you know, a pulse. To learn how to put AI to work for people, visit servicenow.com. Time is your most important asset and AI is transforming the speed of business. Your innovation and how you protect it depends on the technology you choose to trust. 70 % of the Fortune 100 trust CrowdStrike's leading AI-native cyber security platform built to secure their business and the AI fueling their innovation. AI is changing the world. We're securing it. CrowdStrike. We stop breaches. Support for the show comes from Rippling. Imagine you just found out your sales team is at risk of missing quota. No need to panic.
12:07Nilay Patel:Just ask Rippling AI. Since it's built on your real-time people and business data, Rippling AI can pull metrics from Rippling and Salesforce into a meeting-ready dashboard showing quota attainment, headcount trajectory, and monthly revenue to quota by region. In seconds, you can see exactly what's behind your quota risk and fix it before it's missed. Question answered. Action taken. Crisis averted. So when you have critical business questions that need answers, don't just file a ticket and wait weeks for an outdated report. Describe what you need and have Rippling AI build it instantly from your live people and business data.
12:43Nilay Patel:Whether it's a dashboard with detailed charts or automated workflows with the right triggers, conditions, and approvals. Ready to rule your business? Head to rippling.ai slash decoder to get the only AI built to give you full visibility and take complex actions across your entire organization. That's R-I-P-P-L-I-N-G dot AI slash decoder. Sign up for exclusive access today.
13:18Nilay Patel:We're back with The Verge's Robert Hart, talking about AI's math capabilities. and what happened earlier this month with an all-new AI model called Astra. Let's talk about these 10. You reported them out. OpenAI did produce some documentation. But they're not all entirely horsepowered out of nothing, right? They're based on previous work. There's some question of attribution. What was the response? Is it, oh, the models did this? Or was it the response we see to so much AI work, which basically boils down to, well, you stole this and didn't attribute anyone. and you've built on the shoulders of giants without mentioning it.
13:52Nilay Patel:How did the response land? By and large, it was generally one of being quite impressed from the people I spoke to. It was, as I said, these are problems mathematicians cared about. There was one that drew particular attention for how they credited it and also how they'd kind of announced all of this in their blog post. OpenAIR initially, I think, had said that these are 10 problems. There have been no progress in the last 10 years. and then if you actually read the papers one of them it quite clearly says oh we build on progress from these two researchers and so i mean that was later changed quite quietly as well but a few of the researchers i spoke to were quite unimpressed and they did feel it was an element of well yeah you've not credited something that you've used heavily here and by your own acknowledgement that said one of the researchers i did speak to who was one of the ones named was a bit ambivalent on the whole thing as well so it was a real mixed bag but I mean the general vibe I'd say other than this perhaps was oversold in terms of what came before which was corrected to their credit I don't think it was anything other than I think the word sloppy was what was described to me by one person the general impression was quite impressed like these were actual breakthroughs that bothered people.
15:08And it did move the field forward in a way that, yeah, as I said, if it was a human mathematician that had done these, I think, I mean, several researchers actually said that, well, if a researcher had done any one of these problems, they'd probably be set for an academic career. So yeah, it was impressed with what it did, if not the credit for some of it. Yeah.
15:29Nilay Patel:Well, it's funny, you know, credit and attribution in academia is like the whole game. And it seems like the AI companies get away with being sloppy in a way that no human would be able to get away with being sloppy. Did the scale of the discovery or the work overcome the sloppiness? If a human had accomplished the same goals and had been as sloppy, would the reaction be the same? I mean, part of me always wants to lean on the whole like, oh, it looks like plagiarism kind of element. But like, if you actually read the papers, they kind of produced. And I mean, One of the researchers I spoke to said there's probably about 50 people in the world who are going to bother reading through this in depth.
16:05It is very clear. It doesn't attempt to plagiarise. I think it was just a poor press release, to be honest. As much as I love to go in on it sometimes. I mean, I find that having, I suppose, having covered science for a decade plus, the press releases are often overselling. what discovery has actually been made and the import of it and the novelty of it. And I think that's just another case of what happened here.
16:33Nilay Patel:Does this seem repeatable? There's some proof that they provided that they solved 10 problems that were unsolvable. Do they provide any proof that they can solve another 10? That's the question. I mean, it's also so the big unknown from, oh God, the near dozen people I spoke to for this was well how many did they try to get these 10 who knows um i mean they know but they won't say but that is the big question here is like it's unclear how quite how many attempts it took to get these 10 impressive as it is it's not like that was i mean i would be very impressed if it was the kind of first thing they sort of go and then out come these 10 impressive results it's unclear kind of what areas they would focus on next and why i mean there are i imagine business reasons as to putting together which problems they are choosing to publicize so their models can can do so it's yeah and they've all the ai labs have been hiring a sort of cohort of senior mathematicians uh behind the scenes so it's anyone's guess i think as to whether they do it again but on the question of proof i mean maths is i suppose a bit it's an odd discipline in science in that a repeatability is kind of not the same thing like with experimental sciences a proof is a proof and if it works it works it's the problem here is well can people follow through what they've done each field is quite highly specialized so there'll be individual mathematicians who are in those fields that go through those i spoke to that worked in some of the fields that were covered here say it all looks very legit there's also in maths there's a programming language slash computational proving type thing called lean where you can basically codify the mathematical proofs and run them through.
18:15And it kind of, well, proves it. I keep saying prove a lot here, but it will test the rigor and the assumptions of everything going on there. And they've published that as well. So it does appear to hold. No one I've spoken to once, they may say that the press release has a lot of hype or there's a lot of hype around it. No one seems to be doubting the kind of essential breakthroughs that they're claiming here.
18:36Nilay Patel:I want to stay on this subject for one more second. There's the mathematical proof. We've generated a proof. And that is, as you say, just repeatable in a way that math is just logic. You can just go through the steps and say this proof worked. And anybody listening to this who had to suffer through writing a proof in calculus in high school probably remembers that process. There's something there that's pure logic. Then there's a part of it that is software code, as you're describing in Lean, where you can take the pure logic, you can express it in code, and you can run it to see if it works.
19:05I understand how AI is theoretically good at all of that.
19:08Nilay Patel:You're just going to run the reasoning and the reasoning is going to generate some code. You're going to run the code. You're going to get some verifiability. We've seen this play out in software engineering where the code runs or not, it's verifiable or not, and the models can just reason out about it. Then there's, to me, the big question that you alluded to. How many times do you have to run this? Can we verify that the models did this and they weren't directed by human mathematicians who've been hired at high rates by the labs in a way that suggests the field is going topsy-turvy? You have a quote here from James Maynard, who has won the Fields Medal, the highest prize in mathematics, who said he's been soul-searching.
19:44Nilay Patel:And I keep looking at the quote. You've got similar quotes from all these other mathematicians in the piece. And it seems like they're soul-searching against a thing that kind of hilariously they cannot verify, which is how did the models do this? And is that thing scalable in a way that threatens mathematics? What do we know about how the models did this? I'd say as much as we normally do and do not know about this, it's, I mean, there are two, I think there are a few issues kind of there. One is the nature of the models. Well, this is an unreleased model, so good luck to anyone wanting to independently test it.
20:17I mean, the same if that goes with anything proprietary, really. That said, I am inclined to kind of almost give the benefit of the doubt that they're not lying in some capacity about the models they're using. um as for the other part i think is perhaps more noteworthy on the kind of how do we prompt this or how is it being guided um and whether that's by a mathematician that knows what they're doing and i think that's probably a key factor here is a lot of mathematicians i spoke to when they've tried using these and this is often the consumer models but still it kind of speaks to a broader landscape and they say that if you know what you're doing and you can kind of point things out it's good or you can use it as a tool in a way that you kind of want and in a way that you wouldn't be able to if you didn't really know how to fact check it i mean the same time as if i mean i've had it where i've had chat gpt or saying doing a basic sum and i'm like that number is not right it's like i'm so sorry uh you're right it's this and it's still wrong but um that's kind of still needed at this level as well i think it does allude to a kind of broader problem as you said kind of this almost soul searching of well what if we can automate that our way?
21:25And what if it gets to a point where we don't understand it? And that really cuts to a deeper question of like, well, what is mathematics? Why do we do it? Why do we value it as a field? I mean, everyone will have different answers to that, I think. But a fear of a lot of people I spoke to was that this might kind of move beyond a realm of human interest. And in which case, well, maybe they just won't engage with it. Or it'll be something that interested people will go through and then the rest will kind of continue as normal.
21:49Nilay Patel:You got a quote here from a researcher in Zurich named Johannes Schmidt, who says we might be headed toward a situation where the math problems get, quote, mowed down by AI, but we don't actually push the field forward because humans are taken out of the loop and they're not either checking or they don't understand it or they don't know what the future breakthroughs might be. How likely does that feel? Is that a big concern? There is an element of that, of the mowing down of the problems, and especially those are used as sort of a training field for younger mathematicians coming up and kind of cut their teeth so to speak but i also think that kind of this problem solving idea in maths that that's that's what maths is is a very much outsider's perspective of the mathematical endeavor i think so a lot of the mathematicians i spoke to sort of found that almost tick box part the least interesting and valuable part of the field like the the areas that are valuable for them aren't oh you've solved something or you've proven something it's what happens from that and i I think it was James Maynard that said that the most interesting kind of discoveries in this isn't that you've solved something, it's what evolves from that.
22:54So is it like sometimes they open entire new fields of research that no one ever thought was possible? Or, oh, this is a new tool that you can apply everywhere in fun and exciting ways. I think if we look as well about what almost the kind of popularization of maths, like even kind of what I'm thinking of is kind of those theorems that people have posed. It's the questions that endure, not the solutions. like it's always kind of Fermat's last theorem, not like, well, here's the solution to whatever the last guy proposed. I think the concern here is that, well, they're going to kind of tick off all of these questions, but normally in the process of doing so, one would hope would branch out into all of these new exciting areas or pose new questions, but it won't do that.
23:35That's the concern. And then that would kind of leave the field quite sterile and it will have all of these things that have been done and maybe nothing left to pursue. And the general consensus was, well, the jury's out. It's too early to tell. Even with human mathematicians, it takes a lot of time to realize the impact of these kind of things. And as I said, it's kind of exploded in the last six months to a year, which I mean, math is not a fast moving discipline at the best of times, but it's too early to tell really whether that will then be a kind of concern. But it is a concern and a big one.
24:07Nilay Patel:You have another quote from Maynard here saying, if the standard for a publishable paper in math is something that an AI cannot do, particularly when a PhD is typically four years, the challenge is you're not trying to come up with a problem that AI can't do now. It's an AI in four years' time. And so this is really related to the rate of improvement of the models, which, as you say, particularly math, it seems to be increasing, but not at an even rate across all of the domains of mathematics. That appears to be what is causing the soul-searching, right? If you're a student and you start today and you pick some obscure domain that maybe the AI isn't good at sometime halfway through your PhD thesis or your PhD research, the AI will just solve it and you'll be done.
24:48Nilay Patel:And that is a real problem for you. Is the field reacted to that yet or are they just in the shock of, oh, the models can start to do things that we didn't think they were capable of? Yeah, I think it's shock really. I mean, the little notebook I have whenever I do these interviews, I keep one to the side to just do broad feelings. And I've written like shell shock in it because it just feels that it's and it's far from universal but it's it feels like it's happened so quickly that it has just taken a lot of people by surprise i mean even if it's they kind of knew in in theory that well this is coming they've seen all these sort of start ai math startups kind of going they've seen colleagues being moving around to different labs or kind of areas of work but yeah it's just happened very very quick and so it's given them very little time to kind of figure out what and it's not necessarily even the fields that it might be good at it's more just like like what what can it do like it's just that's how quickly it's happened is that like if i think like what six months um i'm thinking of like an academic year in the uk from what it goes from like october through to october but i can't imagine how when i was studying if something like this had come out and literally in the space of half a year it just upended what was possible and over the summer as well so students are possibly coming back to a completely different discipline after a break we have to take on a short break we'll be back in just a minute excuses are easy an epic movie night we don't have enough snacks dinner party with the girls?
26:36Nilay Patel:We'd have to decorate. Surprise date night? Nothing to wear. But Amazon's Prime same-day delivery lets you say yes before the moment slips away. Try that new popcorn maker, order those cheeky drink glasses, get that new perfume, and turn that I wish we could into an I'm so glad we did. Visit amazon.com slash prime to find millions of items delivered fast. Same-day delivery. It's on Prime. Available in select areas. Terms apply.
27:30Nilay Patel:dot com slash podcast. Terms and conditions apply. Need a hiring hero? This is a job for Indeed Sponsored Jobs. Sierra has all the best active and outdoor brands you need. From athletic stuff, like a full court pickup game, swish, to athletic-ish stuff, like a half mile stroll. Get those steps in. And for morning hikes up the mountain trail, good pace. It's a nighttime ghost choice from the camping chair. What a twist. Whatever level of active, Sierra loves it all. Head to Sierra or Sierra.com for the brands you want at the prices that let you do it all. From athletic to athletic-ish, Sierra's got it.
Read the full transcript
28:11Nilay Patel:We're back with The Verge's Robert Hart, discussing why mathematicians are so concerned about what AI is doing to the discipline. There's some skepticism here in the world. Gary Marcus is a reliable skeptic of AI. And he pointed out that over and over again, what you see is AI accomplishes something in one domain, and then it's used to generalize AI's ability across every domain. Good quote here. As we learned a decade ago from AI's shambolic and ultimately failed attempt to turn Jeopardy winning Watson into a cancer fighting machine, success in one domain does not guarantee success at all. I can read this two ways.
28:46Nilay Patel:One, there's AI has solved math, which is not true, as you've pointed out, in several ways, all the way down to it's still bad at counting. And then there's AI has solved math, and that means necessarily it's going to come for everything else. It will come for physics. It will come for law. It will come for whatever you want in the world. You can see if there's any verifiability, AI can solve it because you can just run it in this way. I understand both sides of that argument, right? that obviously success in one domain does not guarantee success in every domain. And then the arc of AI is, well, it keeps collecting domains.
29:20And if there is any verifiability, it is more likely to connect those domains than not.
29:26Nilay Patel:How do you see it? I'm not entirely sure that the people that Marcus is criticizing here have actually said quite what he says they are saying as well. There's a lot to say about the hype. But yeah, success in one domain does not even equal success throughout that domain, let alone other domains. And I mean, that said, there has been an undeniable trajectory, I think, in the last few years of a broadening capability increase. And I don't think you need to be on the whole AGI train to acknowledge that and to acknowledge that that will have an impact. I mean, within mathematics, I mean, it is the case of like, I think it was Andras Juhas, I've probably butchered his name there, but one of the professors at Oxford that I quoted in the story.
30:09but something else he'd said to me was that he's been kind of toying around with chat gbt a bit and he's like i don't think it has any geometric intuition whatsoever and and he's like which might explain why there have been a very kind of limited amount of progress in fields like topology i'm in no position to um verify that claim in terms of the maths of it but i think it illustrates it quite well i mean it is it's almost like what is it they call it a jagged edge and it felt like an easy argument for me the whole let's let's criticize the whole the singularity is near and i mean i think that's what elon musk said in response to to the astra thing which it's a lot i also think it's perhaps the least generous interpretation of that argument you can take to argue against i think if you take a more nuanced element that does acknowledge that there has been clear progress here and quite quickly and as you said it is racking up domains i feel there's a trajectory there that is like a reasonable one to like consider rather than just dismiss out of hand.
31:07Nilay Patel:One of the bigger arguments about AI in general is that it democratizes access. I was not a great software developer in my days trying to write software code. And now I can vibe code apps at will to do all kinds of dumb stuff in my house. Is there a similar argument here where a bunch of people who had mathematical intuition, but did not have the formalized language or training of academic mathematics can now access a model and push the field forward? Because that is usually the thing that undercuts the criticism from the professionals is, well, many, many more people now have access to this thing that only you had access to because of your money and your training.
31:43Yeah. I mean, annoyingly, I am going to say it's a two-pronged thing again, but yes, I mean, on a broad sense, yes, it is. I mean, a lot of the mathematicians I spoke to were almost quite weary of this, actually. They love the idea and theory of democratizing access. They're also quite fed up of AI-generated slash assisted papers that are flooding every publication and manageable, as well as the preprint servers that they kind of use in these fields. I mean, at least two I think I spoke to were like, oh, I got three emails this week alone with people being like, hey, is this legit? Because they thought they'd solved something with ChatGPT or with Claude.
32:20And they also don't have the mathematical skills with which to check whether they've actually solved something. I mean, on the flip side, there are parts of where they said, well, we've got a talented undergrad who's done something that a talented undergrad would probably have never managed. And here they are doing grad level work and they've produced a paper that is legit. And on the kind of bigger scheme of things, a few I suppose said, well, yeah, a lot of these kind of are in the ivory tower. Having access to this kind of thing globally could really boost access to the kind of things here.
32:47I mean, on the flip side, the cost. I mean, these things cost to run. It's, I mean, it's always easy to forget. I think when you use, say, a free version of ChatGPT or Claude or something to the higher levels, these things cost money. And whilst they may not necessarily cost a lot of money, I mean, OpenAI claimed, I think it was 2k for these 10 results, which again, doesn't factor in literally anything else once they've got these. So it's a very generous number, but even taking that figure, maths is quite a poor discipline, even at very well-off institutions. I mean, Colva Rooney-Dougal, uhst andrews who i spoke to she said she's like well a lot of the time i don't bother getting a research grant i don't need one i just have a blackboard um and so like if you're not even getting say a research grant two grand is a lot to to put up and so it could lock out researchers that way even at quite well-funded institutions not to mention that the speed of which this is happening that virtually no one would have been able to bake any of this into a grand proposal yet was another theme that I came across a lot.
33:59Nilay Patel:Actually, Rony Dougal has another great quote in your piece about the nature of the AI labs and how they are talking about math. She said, they're treating our discipline as an advertising playground. A bunch of mathematicians have signed something called the Leiden Declaration, which is an open letter to pledge not to buy into hype around AI, these things are running right at each other. The AI labs are not going to stop using every discipline as an advertising playground. And a bunch of mathematicians saying we refuse to buy the hype certainly does not seem to be stopping the hype. There's just a piece of this that is organized professional resistance to a thing that is upending a field that has, as you say, been pretty cheap to operate and now might be getting cheaper or easier to access or easier to upend day by day.
34:45Nilay Patel:Do mathematicians feel that that is going to be effective? historically mathematicians are not like savvy political operators there's a part of me that says oh they're just going to get run over i don't know the history history of maths actually i think a lot of them were quite were quite savvy isaac newton is the one that always comes to mind for that although quite a petty political operator as well but uh yeah i mean that is the fit i mean it's a few that i spoke to um and one really comes to mind as they say that there's often this belief that maths is kind of the pinnacle of knowledge um we're not live so i can be free on this one but he was like well that's bullshit uh because um it is and he wasn't alone in kind of illustrating that sentiment but it is good for showcasing and it's a lot neater as a discipline and a lot cheaper as are than i mean you mentioned with with marcus saying as i ibm's watson is this cancer curing thing like well that involves lots of messy experiments including on people you don't need that in this so it's a really easy discipline to kind of come in throw your weight around and then move to somewhere more lucrative if that's what you want i'm not saying that that's what they are i mean a lot of the people at these companies have kind of been hired i don't i mean i don't doubt their credentials for sure and i I don't doubt their motivations as well, but it does raise a question long-term as to how viable is this?
36:12Because, I mean, let's be clear, as a field goes, I cannot imagine them being a very lucrative enterprise customer for these companies.
36:23Nilay Patel:The thing that might be lucrative is pushing a field forward to turn it into something economically viable. Right? We push mathematics forward as a field that turns into some engineering or physics breakthrough based on that mathematics. And that turns into, I don't know, yet another way to launch rockets. Some circle happens in there that I don't quite understand. But that is the history of innovation, right? From research to engineering to products or services that make money. Is that on the minds of any of these mathematicians that pushing the boundaries here is upstream of something radically economically lucrative?
36:58The immediate counter that would come to mind here is that a lot are scared that it's closing off the field. So by definition, those breakthroughs that lead to something surprising and new that you can say, oh, this works here, may not be happening anymore. And so like, if anything, that kind of lucrative endeavor of like applying maths to then this entire new field that may have a lot of money in it. I mean, it remains an open question as to whether that would, anything like that would be possible if, if we're closing off avenues rather than opening them up.
37:27Nilay Patel:Right. If the economic incentive of solving the unsolved problem is reduced because you personally won't get rich if a computer is just solving every unsolved problem, like something very fundamental breaks there. I mean, a lot of it comes down to, though, like, it's not just solving problems. I mean, a lot of these things, as we said, like, it's about what solving that problem tells you elsewhere. And it's those elsewears that, like, if these were very lucrative problems to be solving, I imagine more people would be trying to solve them than have sort of left them for decades. so like but it's possible that in and this is the nature of a lot of kind of pure science and it's i'd say a broader criticism of what is going perhaps with the trump administration's approach to science policy at the moment in that it's very applications focused there is something to doing pure research that can yield potentially very big dividends that is by definition utterly unpredictable as well but you cannot plan for it and the fear is i think with maths is that in that solving all of these problems and then also in doing so not opening up new areas of research in that sort of i suppose that's that open question well what are you left with as even if it's from more lucrative kind of like what are you going after point of view if you're not opening up new areas of research and you're just taking off old ones it kind of just leaves a big sort of question mark as to what might be left in its wake and i think the sentiment of almost, even from those I was excited that I spoke to that were very excited about what's happening, they said that even they don't really know what's happening.
39:00And they're excited from like a personal level because, oh, we might be able to do this, might be able to do that. But there was still this lingering uncertainty of like, well, what, where does this leave the field? Especially for more pure disciplines like research mathematics, it's tougher to say what comes next. Because in a lot of the other sciences, you can say, well, okay, well, they'd shift onto more engineering problems or applying that. But if you solve all the problems at the ground and there's nothing being built up from that, where do you go from there?
39:27Nilay Patel:One great thing about The Verge is our commenters are vast. They're very knowledgeable. And there was a comment from a mathematics researcher on your story that I just want to read to you and see if you think this is the right framework. Here it is. Quote, I have no doubt these models will bring massive change in the field, but in their current state, they won't yet drive us to obsolescence, just occupy a particularly a useful spot in our bag of tricks. My apprehension comes from not knowing where these things will peak, but overall I remain optimistic. I think AI will be a net boon for math when used properly.
39:56Nilay Patel:I feel like AI will be a net boon for X when used properly is just where you land in life in a lot of things, but that's the most optimistic response that I've heard. If we get it right, it's going to be great. Is that kind of the vibe or is it still more shell-shocked than that? Yeah, I'd say shell shock is still the still the overriding impression i mean i think perhaps the gut response to that is like well it will be a net positive for whom and what is properly all of those are quite legitimate questions here i think and i mean some of the bleak responses from graduate students i saw in essays posted online where's their place in this as future researchers do they have a place in this is it as glorified ai proof checkers um that will be quite an unsatisfying career i imagine or maybe not.
40:43I don't know, but that's, we will see. I think it, I think anything used properly will be in that boon, but yeah, I think it all comes down to what properly means and for whom we're talking about.
40:55Nilay Patel:I suspect over the next year or so, things will come into focus because at some point, OpenAI will have to show people how they did the things of the models. And perhaps more importantly, the other labs are going to want to either replicate these results or show that they can push farther, which will necessarily have to lead to a little bit more transparency and yet more mathematicians having a crisis with you. Rob, thank you so much for being on the show. We'll have you back for you soon. Thank you for having me. I'd like to thank Rob for taking the time to join me on Decoder and thank you for listening.
41:29Nilay Patel:I hope you enjoyed it. If you'd like to let us know what you thought about this episode or really anything else at all, drop us a line. You can email us at decoderattheverge.com. We really do read all the emails, or you can hit me up directly on Threads or Blue Sky. We're also on YouTube. You can watch full episodes at DecoderPod. We also have a TikTok and Instagram, same handle at DecoderPod, and they're a lot of fun. If you like Decoder, please share it with your friends and subscribe wherever you get your podcasts. Decoder is a production version, part of the Box Media Podcast Network. The show is produced by Kate Cox and Nick Stat.
41:54Nilay Patel:This episode was edited by Ursa Wright. Our supervising producer is Greg Ott. Our editorial director is Kevin McShane. The Decoder music is by Breakmaster Cylinder. We'll see you next time.
From the publisher
My guest today is Robert Hart, The Verge’s London-based AI reporter. Robert recently wrote a fantastic story for us about the debate raging inside the world of mathematics — and the existential crisis over what it means that new frontier AI models have become very good at math in a shockingly short period of time.
I wanted to dive into all of this with Robert, who actually spoke to some of the most accomplished mathematicians working today to figure out what’s hype and what’s real — and to get a sense of just how exciting, and how scary, all of this is.
Links:
The AI takeover of mathematics has begun | The Verge
Ten advances in mathematics and theoretical computer science | OpenAI
An unreleased Anthropic model made progress on the Riemann hypothesis| TechCrunch
OpenAI’s amazing — but vastly oversold — new model Astra | Gary Marcus
OpenAI’s math breakthrough played to AI’s strengths | Understanding AI
Why the legendary Erdős problems are falling to AI | Quanta
Subscribe to The Verge to access the ad-free version of Decoder!
Credits:
Decoder is a production of The Verge and part of the Vox Media Podcast Network.
Decoder is produced by Kate Cox and Nick Statt. This episode was edited by Ursa Wright. Our supervising producer is Greg Ott, and our editorial director is Kevin McShane.
The Decoder music is by Breakmaster Cylinder.
Learn more about your ad choices. Visit podcastchoices.com/adchoices




