Terence Tao – Kepler, Newton, and the true nature of mathematical discovery

20 Mar 2026 · 1 h 24 min · 31 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Dwarkesh Podcast - Episode Summary: Terence Tao on Kepler, Newton, and the Nature of Mathematical Discovery

Overview In this episode of the Dwarkesh Podcast, host Dwarkesh Patankar interviews Terence Tao, a renowned mathematician. The discussion revolves around historical mathematical discoveries, particularly focusing on Johannes Kepler's laws of planetary motion, and how these relate to modern-day scientific discovery processes, especially in the context of artificial intelligence (AI).

---

Key Themes and Discussions

  1. Kepler's Discovery of Planetary Motion
  2. Historical Context
  3. Kepler built upon Copernican heliocentrism, originally proposing that planets revolved around the sun.
  4. He created a theory involving geometric shapes (Platonic solids) to explain the spacing of planets.
  • Data Collection
  • Tycho Brahe collected precise astronomical data, which Kepler used despite initial difficulties in accessing the data.
  • Kepler's eventual realization that planetary orbits were elliptical (not circular) was a significant breakthrough.
  • Empirical Verification
  • Kepler's laws were derived from extensive data analysis rather than theoretical deduction, showcasing the importance of empirical evidence in scientific progress.
  1. Role of AI in Scientific Discovery
  2. Verification Loops
  3. AI is expected to expedite scientific discovery through tight verification loops; however, historical cases show verification can take decades or centuries.
  4. For instance, even with better theories, previous models may persist due to a mixture of judgment and heuristics not easily captured by AI.
  • AI's Limitations
  • Current AI systems exhibit "cleverness" but lack "intelligence." They can generate theories but may not comprehend or build upon them as humans do.
  • Future of Human-AI Collaboration
  • Tao suggests that humans paired with AI will be the dominant force in mathematics and science for the foreseeable future.
  • AI can handle data analysis at scale, but human intuition and understanding are essential for deeper insights.
  1. Mathematical Progress and Historical Perspective
  2. Shifts in Scientific Method
  3. The historical emphasis on theory and experimentation has evolved with the rise of big data analysis, changing how mathematical and scientific discovery occurs today.
  4. The discussion contrasts traditional methods with modern data-driven approaches where hypotheses are derived from data rather than established first.
  • Importance of Narrative in Science
  • Effective communication of scientific ideas is paramount; historical examples illustrate how narrative and presentation affect the acceptance and progression of theories.
  1. Future Predictions and Advice for Aspiring Mathematicians
  2. Adaptability in Mathematics
  3. The evolving landscape of mathematics suggests that new practitioners may find opportunities to contribute at earlier stages of their careers, aided by AI tools.
  4. Tao encourages a flexible mindset, emphasizing the importance of curiosity and exploration in learning.
  • Impacts of AI on Mathematical Discovery
  • Tao predicts that AI will soon be able to solve many mathematical problems, but emphasizes that the most significant breakthroughs will likely still involve human-AI collaboration rather than complete autonomy.

---

Timestamps of Key Discussions

  • 00:00:00 - Kepler's innovative approach to planetary motion
  • 00:11:44 - AI's role in scientific discovery and verification loops
  • 00:26:10 - The concept of "deductive overhang" in scientific theory
  • 00:30:31 - Selection bias in AI-generated discoveries
  • 00:46:43 - AI's contributions to enriching but not deepening scientific papers
  • 00:59:20 - Need for a semi-formal language in scientific communication
  • 01:09:48 - Tao's personal time management and productivity
  • 01:17:05 - Future of human-AI hybrids in mathematics

---

Conclusion The episode provides deep insights into both the historical context of mathematical discoveries and the potential future roles of AI in science. Terence Tao's reflections on Kepler's work serve as a lens through which to evaluate contemporary issues in mathematics and the evolving landscape shaped by AI, emphasizing the importance of human intuition and narrative in scientific progress.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Kepler's Discovery of Planetary Motion

0:45 to 4:14

A deep dive into Kepler's laws of planetary motion and his methods.

“planets were perfect circles and his theory kind of fit the observations that the Greeks and the Arabs and Indians had worked out over centuries.”

The Role of Data in Scientific Discovery

4:14 to 6:32

Discussion on how data and theory interplay in scientific advancements.

“Where Newton comes up with this explanation of why the three laws of planetary motion must be true.”

Evolution of Scientific Methods

6:32 to 8:48

Exploration of the shift in scientific methods from hypotheses to data-driven approaches.

“And he was using Euclidean geometry and the most advanced mathematics he could use at the time to match his models with the data.”

AI and the Future of Scientific Discovery

8:48 to 14:00

Insights into how AI transforms the landscape of scientific discovery and validation.

“And then a few years later, he gets Brahe's data.”

The Challenge of AI Paper Validation

14:00 to 15:00

Explores the difficulty of recognizing significant advancements among numerous AI-generated papers.

“You know, for each individual paper, we can discuss it with, you know, have a debate among scientists and get to a consensus in a few years.”

The Role of Historical Context in Science

15:00 to 16:00

Discusses how historical context affects the reception and evaluation of scientific ideas.

“And how would you identify it among millions of papers which might actually constitute progress, but which have much less general unifying ideas?”

Understanding Scientific Theory Evolution

16:00 to 17:40

Examines how scientific theories evolve and often appear less accurate at their inception.

“And it was the first deep learning architecture that really was sophisticated enough to capture language.”

Misconceptions in Scientific Theory Acceptance

17:40 to 19:10

Analyzes how initial misunderstandings can hinder the acceptance of correct theories.

“that just either make no sense because they're wrong and we realize later on why they're wrong or they're correct but seem wildly implausible at the time.”

Revolutionizing Understanding of Intelligence

19:10 to 20:50

Discusses the shift in understanding intelligence in relation to AI and historical perspectives.

“And Copernicus' theory was a lot simpler, but much less accurate.”

Communication's Role in Scientific Progress

20:50 to 23:10

Explores the importance of effective communication in the dissemination of scientific theories.

“And so, you know, trying to fit AI into sort of our theories of scientific progress and what is hard and what is easy, we're struggling quite a lot.”
Show all 31 chapters

Data Analysis and Scientific Insight

23:10 to 25:10

Investigates how data analysis helps scientists extract deeper insights from complex information.

“He wrote in English, in natural language.”

Exploring the Pairing and Ordering of Blocks

28:00 to 28:50

Learn about Sean's innovative methods for pairing and ordering blocks in AI models.

“First, pair the layers into 48 different blocks, and second, put those blocks in the right order.”

Measuring Scientific Engagement through Citations

29:00 to 30:26

Explore how clever studies measure scientists' engagement through citation practices.

“And they could infer whether an author was actually just copying it, cutting and pasting a reference without actually checking it.”

AI's Progress in Solving Mathematical Problems

30:26 to 35:36

Discuss the recent achievements and current plateau of AI in tackling ERDOs problems.

“But then I think, I don't know if it's still correct, but as of a month ago, you said that there had been a pause because the low-hanging fruit had been picked.”

Transforming Mathematics with AI

35:36 to 40:58

Understand how AI tools are reshaping mathematics and the balance of breadth vs. depth in research.

“So I see very much a future of very complementary science.”

The Future of Mathematics in an AI-Driven World

40:58 to 42:00

Explore the implications of AI on mathematical practices and the potential for large-scale data analysis.

“And it would be interesting to understand how much progress one can make simply from using existing techniques.”

AI's Role in Mathematical Problem-Solving

42:00 to 44:30

Explore how AI tools are impacting mathematical problem-solving techniques.

“and you have to add one more wrinkle to it.”

Evaluating AI's Effectiveness in Math

44:30 to 46:20

Discuss the varying success rates of AI in solving complex mathematical problems.

“All these problems that haven't been solved before for decades, now they're falling.”

AI and Mathematics: Productivity Changes

46:20 to 48:50

Understand how AI has altered the productivity and methodology of mathematicians.

“It would be like a colleague in mathematics or?”

Understanding Intelligence vs. Cleverness

48:50 to 51:40

Delve into the distinction between artificial cleverness and true intelligence.

“And eventually, you know, we sort of, we've systematically mapped out what doesn't work, what does work, and we can kind of see a path forward, but it's evolving with our discussion.”

The Future of AI in Mathematical Discovery

52:55 to 56:00

Discuss the potential future dynamics of AI and human collaboration in solving math problems.

“One big question I have is, how plausible is it that if we just keep training AI that get better and better at solving problems in Lean, that they will continue to solve more and more impressive problems.”

Understanding Mathematical Proofs with Lean

56:00 to 57:30

Learn how formalizing proofs in Lean allows mathematicians to isolate and study important lemmas.

“we would be able to apply it in all these different situations.”

The Future of Proof Writing in Mathematics

57:30 to 1:00:00

Discover how AI tools can transform the writing and refinement of mathematical papers.

“One thing that will change quite a bit in the near future is that until recently, writing papers was the most time-consuming and expensive part of the job.”

Statistical Patterns in Prime Numbers

1:00:00 to 1:03:49

Explore Gauss's conjecture and the emergence of statistical patterns in prime numbers.

“We have a few sort of mathematical ways to model this, Bayesian probability, for example, but you often have to set certain base assumptions and there's a lot of subjectivity still in these tasks.”

Theorems and Conjectures about Primes

1:03:49 to 1:08:15

Understand the implications of prime number conjectures on mathematics and cryptography.

“So he conjectured what we now call the prime number theorem.”

The Process of Learning Mathematics

1:09:49 to 1:10:00

Discover methods for mastering new subfields in mathematics as an autodidact.

“So in some sense, you're also one of the world's greatest autodidacts.”

Learning New Subfields in Mathematics

1:10:00 to 1:12:20

Discover Terence Tao's approach to delving into new areas of mathematics.

“What is your process of learning about a new subfield in math?”

The Role of Serendipity in Academia

1:12:20 to 1:14:30

Explore the importance of unexpected interactions and experiences in academic growth.

“How long does it take you to write a blog post?”

The Shift from Physical to Digital Research

1:14:30 to 1:16:40

Learn about the transition from traditional research methods to digital tools and its implications.

“And more often than not, I feel like I've gotten a positive experience, which is not something I would have planned for.”

AI's Impact on the Future of Mathematics

1:16:40 to 1:21:05

Understand how AI is changing mathematics and what it means for future mathematicians.

“I actually surf the internet a lot more.”

Advice for Aspiring Mathematicians in a Changing World

1:21:05 to 1:23:43

Get insights on navigating a mathematics career in the age of AI and rapid change.

“Anything is possible, really, at this point.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Okay, today I'm chatting with Terence Tao, who needs an introduction. Terence, I want to begin by having you retell the story of how Kepler discovered the laws of planetary motion, because I think this will be a great jumping off point to talk about AI for math. Okay, yeah. So I've always had an amateur interest in astronomy, and so I've loved stories of how the early astronomers worked out the nature of the universe. so Kepler was building on the work of Copernicus who was himself building on the work of Aristarchus so Copernicus very famously proposed the heliocentric model that instead of the planets and the sun going around the earth that the sun was at the center of the solar system and the other planets were going around the sun and Copernicus proposed that the orbits of the planets were perfect circles and his theory kind of fit the observations that the Greeks and the Arabs and Indians had worked out over centuries.

0:56I think Kepler got interested, like he learned about these theories in his studies and he made this observation that the ratios of the size of the orbits that Kuhlunka predicted seemed to have some geometric meaning. I think he started proposing that if you take say the orbit of say the earth and you enclose it in I I think maybe a cube. The outer sphere that encloses the cube almost matched perfectly the orbit of Mars and so forth. And there were six planets, none of the time, five gaps between them. And there were five perfect platonic solids, the cube, the tetrahedron, isochedron, octahedron, and dodecaedron.

1:35And so he had this theory, which he thought was absolutely beautiful, that he could inscribe these platonic solids between the spheres of the planets. And it seemed to fit. And it seemed to be to him like, you know, God's design of the planets was matching this mathematical perfection of the platonic solids. So he needed data to confirm this theory. And at the time, there was only one really high quality data set almost in existence, which was the, so Tycho Brahe, this Danish astronomer, very wealthy, eccentric astronomer, had managed to convince the Danish government to fund this extremely expensive observatory.

2:12In fact, an entire island where he had taken decades of observations of all the planets, Mars, Jupiter, every night, at least every night for which the weather was clear, with the naked eye, actually. He was the last of the naked eye astronomers. And so he had all this data which Kepler could use to confirm his theory. And so Kepler started working with Tycho, but Tycho was very jealous of the data. He only gave little bits of it at a time. And I think Kepler eventually just stole the data. He copied it and had to have a fight with Brahi's descendants. But he did work out, he did get the data, and then he worked out to kind of his disappointment that his beautiful theory didn't quite work.

2:52Like the data was sort of off from his platonic solid theory by about 10 % or something. And he had all kinds of fudges moving the circles around and things. It didn't quite work. But he worked on this problem for years and years. And eventually he figured out how to use the data to work out the actual orbits of the planets. And that was an incredibly clever, genius amount of data analysis. And then he eventually worked out that the ellipses were actually ellipses, not circles, which was shocking for him. And then he worked out the two laws of planetary motion, ellipses, also equal areas, sweep out equal times.

3:33And then 10 years later, after collecting a lot of data, the furthest planets like Saturn and Jupiter were the hardest for him to work out. But then he finally worked out this third law also, that the orbits, the time it takes for a planet to complete its orbit was proportional to some power of the distance to the sun. And these are the three famous capital laws of motion, and he had no explanation for them. It was just all driven by experiment. it. And it took Newton a century later to give a theory that explained all three laws at once. The take I want to try on you is that Kepler was a high temperature LLM.

4:17Where Newton comes up with this explanation of why the three laws of planetary motion must be true. And of course, the way that Kepler discovers the laws of planetary motion or figures out the relative orbits of the different planets is, as you say, a work of genius. but then you know he's through his career he's just trying random relationships and in fact the in the book in which he writes down the third law of planetary motion it's sort of an aside on the harmonics of the world which is this book about you know all these different planets have these different harmonies and the reason there's so much famine and misery on earth is because the earth is me for me that's the node of earth and so all this random astrology but in there is the cube square law which tells you what relationship the uh the period has to a planet's distance from the sun which is as you're detailing uh if you add that to newton's f equals ma and then the equation for centripetal acceleration you get the inverse square law and so newton works that out but the reason i um i think this is an interesting story is i feel like lms could do the kind of thing of like 20 years let's try random relationships some of which make no sense as long as there's a verifiable data bank like Brahe's data set where, okay, I'm going to try out random things about like musical notes.

5:28I'm going to try out random things about platonic objects. I'm going to, all these different geometries have this bias that there's some important thing about the geometry of these orbits. And then one thing works. And as long as you can verify it, it can then drive, these empirical regularities can then drive actual deep scientific progress. Traditionally, when we talk about the history of science, idea generation has It's always been kind of the prestige part of science. So, I mean, a scientific problem comes with, there's many steps. You know, you have to identify a problem, and then you have to identify a good problem to work on, a fruitful problem.

5:59And then you need to collect data. You need to figure out a strategy to analyze the data, to make a hypothesis. And at this point, you need to propose a good hypothesis. And then you need to validate, and then you need to write things up and explain. There's a dozen different components. but yeah the ones we celebrate are these of eureka genius moments of idea generation and yeah so so kepler certainly had to to as you say cycle through many ideas and and several which didn't work and and i bet many that he didn't even um publish at all um because yeah they just didn't fit and that's an important part of the process um trying all kinds of random things and seeing if they worked um but as you say the um you know the uh it had to be matched by an equal amount of verification otherwise it's it's slow you know i mean um we celebrate kepler but we should also celebrate brahi for for his his assiduous data collection with which was ten times more precise than any previous observation and it was um that extra decimal point of accuracy was actually essential for Kepler to get his results.

7:08And he was using Euclidean geometry and the most advanced mathematics he could use at the time to match his models with the data. So all aspects had to be in play, the data and the theory and the hypothesis generation. I'm not sure nowadays that hypothesis generation is the bottleneck anymore. So sciences has changed in the century since. So classically, sort of the two big paradigms for science were theory and experiment. Then in the 20th century, numerical simulation came along. And so you can also do computer simulations to test theories. But then finally, in the late 20th century, we had big data.

7:55We had the era of data analysis. and so a lot of new progress is actually driven now by analyzing massive data sets first collecting large data sets and then drawing the patterns from them to to deduce laws which is a little bit different from how science used to work where you make a few observations or you just have one out of the blue idea and then you collect data to test your idea that's the classic scientific method now it's almost reverse you collect big data first and then you try to get hypotheses from it. I mean, Kepler was maybe one of the first early data scientists, but even he didn't start with Tycho's data set and analyze it.

8:34He had some preconceived theories first. But it seems like this is less and less the way we make progress just because the data is just so much more massive. It's just so much more useful. Oh, interesting. I actually feel like the mold the 20th century science that you're describing is actually very well describes what happened Kepler, where he did have these ideas, 1595 and 96 is where he comes up with first polygons and then platonic objects theory, but they were wrong. And then a few years later, he gets Brahe's data. And it's only after 20 years of just trying random things that he gets this empirical regularity.

9:15And so it actually feels closer to Brahe's data is analogous to some massive data bank of simulations. And then now that you've got the data, you can keep trying random things. But if it wasn't, Kepler would be out there just writing books about harmonics and platonic objects, and there would be nothing to actually verify against. Yeah. So the data was extremely important. But the distinction I was trying to make was that traditionally, you make a hypothesis and then you test it against data. But now with machine learning and data analysis and statistics and something you can you can start with data and um through say statistics work out um um laws that um were not present before so and so kepler so kepler's third law is a little bit like this except that uh for the third law instead of having the thousand data points that brahi had kepler had like six data points um like every planet uh you knew the length of the orbit and the distance of the sun and there was like five or six data points and he did uh what we would now call regression you know he could fit a curve to these six data points and he got a square coup law which was amazing uh but actually he was quite lucky i mean that these six data points gave him the right conclusion um you know it's uh that's not enough data to be really reliable um there was a later astronomer uh johannes borde who took the same the same data actually um the the distances to to the planets uh and inspired by Kepler, I think, he had a prediction that the distances of the planets formed basically a shift to geometric progression.

10:48He also fit a curve, except there was one point missing. So there was a big gap between Mars and Jupiter. His law predicted that there was a missing planet. So it was a kind of a crank theory, except when Uranus was discovered by Herschel, the distance Uranus fit exactly this pattern. And then Ceres was discovered, this asteroid between, I think, in the asteroid belt. And it also fit the pattern. So people got really excited that Bill had discovered this amazing new law of nature. But then Neptune was discovered and it was completely like way off. And basically it was just a numerical fluke. There were six data points.

11:30So maybe one reason why Kepler didn't highlight his third law as much as the first two laws is that maybe instinctively, even though he didn't have modern statistics, he kind of knew that with six data points, he had to be somewhat tentative. with the conclusions. But maybe to ask the question about the analogy more explicitly, does this analogy make sense to if we have, you know, in the future, we'll have smarter and smarter AIs and we'll have millions of them. And then they can go out and hunt for all these empirical regularities. It sounds like you don't think the bottleneck in science is finding more things that are for each given field, they're equivalent of the third law of planetary motion.

12:09So that then later on, somebody can say, oh, we need a way to explain this. Let's work out the math. Here's the inverse square law of gravity. Right. So I think AI has basically driven the cost of idea generation down to almost zero. In a very similar way to how the internet drove the cost of communication down to almost zero, which is an amazing thing, but it doesn't create abundance by itself. Yeah, so now the bottleneck is different. So we're now in a situation where suddenly people can generate thousands of theories for a given scientific problem. And now we have to verify them, evaluate them.

12:45And this is something which we have to change our structures of science to actually sort this out. So, you know, in fact, traditionally we build walls, you know, so in the past, you know, before we had AI slop, you know, we had sort of amateur scientists, you know, have their own theories of the universe, many of which were basically of very little value. and so we've built these like you know peer review publication systems and things to kind of filter out and try to isolate the high signal ideas to to test but but now that we can generate these these these these possible explanations at massive scale and some of them are good and a lot are terrible I mean human reviewers we just it's just they're already being overwhelmed actually I mean many many journals are reporting ai general submissions are just are just flooding their submissions so it's great that we can generate all kinds of things now with ai but it means that we have to the rest of the rest of the aspects of science have to catch up yeah so verification validation um and and assessing uh what ideas actually move the subject forward and and what which ones are dead ends or or red herrings um and that's that's not something where we've we know how to do at scale.

14:02You know, for each individual paper, we can discuss it with, you know, have a debate among scientists and get to a consensus in a few years. But when we're generating, you know, a thousand of these every day, yeah, this doesn't work. Yeah. So I think there is this incredibly interesting question of, you have billions of AI scientists, not only how do you gauge which ones are real progress, but how do you, I mean, this is actually a question that human scientists had to face and we've solved somehow. And I actually am not sure how we solve this, but in any given field, let's say in the 1940s and there's, if you're at Bell Labs or if you're just generally trying to, there's these new technologies coming out, pulse code modulation, basically how do you transfer signals?

14:39How do you digitize signals? How do you transfer them over analog wires? And then, but there's like all these papers about the engineering constraints there and the details. And then there's one which is like, comes up with the idea of the bit, which has implications across many different fields. And And you need some system which can then look at that and say, OK, we need to apply this to probability. We need to apply this to computer science, et cetera. And in the future, the AIs are coming up with the next version of this kind of unifying concept. And how would you identify it among millions of papers which might actually constitute progress, but which have much less general unifying ideas?

15:14So a lot of it's the test of time. So many great ideas didn't actually get a great reception at the time that they were first proposed. It was only after some other scientists realized that they could take it further and apply them to their own. Deep learning itself was actually a niche area of AI for a long time. The idea of getting answers entirely through training on data and not through first principles reasoning was very controversial. And it took a long time before it actually started bearing fruit. You mentioned the bit. I mean, there are other proposals for computer architectures than the 01 that is universal today.

15:50I think there were tricks, you know, zero, one, three-valued logic. And, you know, in an alternate universe, maybe a different paradigm would have showed up. People have argued that, you know, the transformer, for example, is the foundation of all modern large language models. And it was the first deep learning architecture that really was sophisticated enough to capture language. But it didn't have to be that way. There could have been some other architecture that was the first to do it. And once that was adopted, it would become the standard. So I think one reason why it's hard to assess whether a given idea is going to be fruitful is that it depends on the future.

16:29It depends also on the culture and society, like which ones get adopted, which ones don't. The base 10 numeral system in mathematics is extremely useful, much better than the Roman numeral system, for instance. But again, there's nothing special about 10. it's a system that we it's useful for us because everyone else uses it and we've standardized it and we've brought all our computers and our number of representation systems around it and so we're stuck with it now actually you know people are some people occasionally push for other systems than decimal but it's this this is no this is no there's too much inertia so you can't look at any given scientific achievement purely in isolation and give it an objective grade without being aware of the context, both in the past and the future.

17:22And so it may never be something that you can just reinforce and learn the same way that you can for much sort of more localized problems. Yeah. It seems often in the history of science, when a new theory comes up that in retrospect we realize is correct, it seems to make implications that just either make no sense because they're wrong and we realize later on why they're wrong or they're correct but seem wildly implausible at the time. So as you've talked about, Aristarchus had heliocentrism in the third century BC and then the ancient Athenians were like, this can't be because if the earth is going around the sun, we should see the relative position of the stars change as we're going around the sun.

18:07And the only way that wouldn't be the case is if they're so far away that you don't notice any parallax, which is actually the correct implication. But there's times when actually the implication is incorrect and we just need to graduate to a better level of understanding. So Leibniz would, you know, chide Newton and disagree with Newton's theory of gravity on the basis that it implied action at a distance. And then there's we don't know the mechanism. And Newton himself was sort of stunned that inertial mass and gravitational mass were the same quantity. So all these things were resolved by Einstein.

18:36Yes, yes. But it was still progress. And so the question for a system of peer-reviewed for AI would be, even if you can falsify a theory, how would you notice that it still constitutes progress relative to the thing before? Yeah, so often actually the ultimately correct theory initially is worse in many ways. Yeah, so Copernicus's theory of the planets, it was less accurate than Tomli's theory. So geocentrism had been developed for a millennium by that point, and they had made many, many tweaks and increasingly complicated ad hoc fixes to make it more and more accurate. And Copernicus' theory was a lot simpler, but much less accurate.

19:16It was only Kepler that made it more accurate than Tomlin's theory. I mean, science is always a work in progress. So when you only get part of the solution, it looks worse than a theory which is incorrect, but somehow it has been completed to the point where it kind of answers all the questions. As you say, Newton's theory had big mysteries, the equivalence of mass and action at distance, which were only resolved with a very conceptually different approach centuries afterwards. often progress has to be made not by adding more theories, but by deleting some assumptions that you have in your mind.

20:02So one reason why geocentrism held on for so long is we had this idea that objects naturally want to stay at rest. This is the Aristotelian notion of physics. And so the idea that the Earth was moving, how come we weren't all falling over? Once you have neutrons in motion, object in motion remains in motion and so forth, then it makes sense. But you had to, so conceptually, it's a very big conceptual leap to realize that the Earth is in motion. It doesn't feel like it's in motion. And like the biggest advances, you know, Darwin's theory of evolution, you know, is the idea that species are not static.

20:42But, you know, it's not obvious because you don't see evolution in your lifetime. Well, now we actually can, but it seems permanent and static. Right now we're going through a cognitive version of the Copernican revolution where we used to think that human intelligence is the center of the universe and now we're actually seeing that there's very different types of intelligence that are out there with very different strengths and weaknesses and so our assessment of which tasks require intelligence and which ones don't has to be reordered quite a bit. And so, you know, trying to fit AI into sort of our theories of scientific progress and what is hard and what is easy, we're struggling quite a lot.

21:29We have to ask questions that we've never really had to ask before. Or maybe the philosophers had, but now we all have to deal with it. This actually brings up a topic I've been very curious about. So you mentioned Darwin's The Era of Evolution. There's this book, The Clockwork Universe, by Edward Dahlnack, which covers a lot of this era of history we're talking about. And he has this interesting observation in there that The Origin of Species is published in 1859. The Principia Mathematica is published in 1687. So The Origin of Species comes out basically two centuries after The Principia. And conceptually, it seems like Darwin's theory is simpler.

22:02There's a contemporaneous biologist to Darwin who reads The Origin of Species, Thomas Huxley, and he says, how stupid not to have thought of that. And nobody ever says that about Principia. They're chiding themselves for not having beaten Newton to gravity. And so there's a question of, well, why did it take longer? It seems like a big part of the reason is that the evidence for natural selection is cumulative and retrospective, whereas Newton can just like, here's my equations. Let me see the moon's orbital period and its distance. And if it lines up, then we've made progress. And so Lucretius actually had the idea, this idea that species adapted their environment in the first century BC, but nobody ever really talks about it until Darwin because Lucretius can't run some experiment and people are forced to pay attention.

22:47And so I wonder if we'll, in retrospect, end up seeing much more progress in domains which have this kind of tight data loop where you can verify them quite easily, even though they're conceptually much more difficult. I think one aspect of science is it's not just creating a new theory and validating it, but communicating it to others. So Darwin was actually an amazing science communicator. He wrote in English, in natural language. No lean. Okay. I have to sort of get out of my technical mindset. He spoke in plain English. Didn't use equations. and he synthesized a lot of disparate facts. So, you know, little pieces of evolution had been worked out in the past, but he had this very compelling vision and again, still missing things.

23:42Like he didn't know the mechanism for hereditary DNA. Yeah, but his writing style was persuasive and that helped a lot. Newton wrote in Latin. he had invented you know entire new areas of mathematics just to explain what he was doing he was also from an era which was where scientists were much more secretive and competitive so you know academia is still competitive it was even worse back in Newton's day so he he held back some of his best insights because he didn't want his rivals to get any advantage he was also actually somewhat unpleasant person from what I what I what I gather actually so it was actually only a couple decades after newton where other scientists explained his work in much simpler terms that they became widespread um so um yeah the the the art of exposition and making a case and creating a narrative um is uh is also a very important part of science um and um if you have the data and it helps but but people need to be convinced otherwise they will not push it further or they want to take initial investment to uh to to learn your theory and really and really explore it um and that's another thing which is really hard to reinforce and learn on uh yeah how can you score how persuasive you are okay well okay there's the entire marketing departments who are trying to do this so maybe it's good that ai are not yet optimized to be uh persuasive so yeah there's There's a social aspect to science.

25:18Even though we pride ourselves on having an objective side to it where there's data and there's experiment and validation, we still have to tell stories and convince our fellow scientists. And that's a soft, squishy thing. It's a combination of data and painting a narrative. And it's a narrative of gaps. I mean, so even Darwin, as I said, there are pieces of his theory he cannot explain. But he could still make a case that in the future, people would find transitional forms, that they would find the mechanism of inheritance. And they did. Yeah. I don't know how you can quantify that in such a precise way that you can start to reinforce some learning.

26:06Maybe that will be forever the human side of science. One takeaway I had from reading and watching your stuff on the Cosmic Distance Ladder, by the way, I highly, highly, highly recommend people watch your series with Thru the One Brown on the Cosmic Distance Ladder. But one takeaway was that the deductive overhang in many fields could be so much bigger than people realize where if you just had the right insight about how to study a problem, you might be surprised at how much more you could learn about the world. And I wonder if you think that's sort of a product of astronomy at the particular times in history that you're studying, or is this that based on the data that is incident on the Earth right now, we could actually divine a lot more than we happen to know?

26:51Right. So astronomy was one of the first sciences to really embrace data analysis and squeezing every last possible drop of information out of the information they had because data was the bottleneck. I mean, it still is the bottleneck. I mean, it's really hard to collect astronomical data. So astronomers are the best, almost, world-class in extracting, almost like Sherlock, extracting all kinds of conclusions from little traces of data. I hear that a lot of quant hedge funds, they're preferred hires in astronomy PhD. They also are very interested for other reasons in extracting signals from various random bits of data.

27:34Okay, speaking of clever ideas, one of my listeners, Sean, solved the puzzle that Jane Street made for my audience and posted a great walkthrough on X. For context, Jane Street trained AriseNet and then shuffled all 96 layers and then challenged people to put them back in the right order using only the model's outputs and training data. You can't brute force this. There's more possible orderings than atoms in the universe. So Sean broke the problem into two different parts. First, pair the layers into 48 different blocks, and second, put those blocks in the right order. For pairing, Sean realized that in a well-trained resonant, the product of two weight matrices in a residual block should have a distinctive negative diagonal pattern.

28:16And this arises as a way for the model to keep the residual stream from growing out of control. From this insight, he was able to recover the right pairings. For ordering, Sean noticed that the model seemed to improve if he sorted the blocks by the size of their residual contributions. Starting with that rough approximation, he combined a clever ranking heuristic with local swaps to recover the exact right order. His full walkthrough is linked in the description. Don't worry if you didn't get to this puzzle in time, though. There's still one up about backdoor LLMs that even Jane Street doesn't know how to solve.

28:44You can find it at jainestreet.com slash dworkash. All right, back to Terrence. We do underexplore sort of how to extract extra information from various signals. like just to pick one random study I remember reading once that people had discovered were trying to measure how often scientists actually read these citations the papers that they cite so how do you measure this you could try to survey different scientists but they had some clever tricks so many citations have little typos like a number is wrong or punctuation symbol is wrong And they measured how often a type of word got copied from one reference to the next.

29:32And they could infer whether an author was actually just copying it, cutting and pasting a reference without actually checking it. And so from that, they were able to infer some measure of sort of how much attention people were paying. So there are also clever tricks to extract. you know so these questions you posed earlier of you know how can we assess whether a scientific development is fruitful or or interesting or represents real progress you know maybe there are really useful metrics and or footprints of this of this of this of this phenomenon in in a data data service you know we can we can examine citations and and like how often something is mentioned in a conference or something and maybe that there's there's a lot of uh social sociology of science research to be to be done and and that could actually um detect these things um yeah maybe we usually get some astronomers on the case actually um okay so i think this brings us uh nicely to the progress that from the outside it seems like ai for math is making and i think you had a post recently where you pointed out that over the last few months AI programs have solved 50 out of the 1 ,100-odd ERDOs problems.

30:47But then I think, I don't know if it's still correct, but as of a month ago, you said that there had been a pause because the low-hanging fruit had been picked. First of all, I'm curious if actually that is still the case, that we have picked the low-hanging fruit and now we're at this plateau currently. It does seem so. I mean, there's still activity at the ERDOs. Yeah, so 50-odd problems have been solved with AI systems, which is great, but there's like 600 to go. and people are still chipping away at one or two of these right now. We're seeing a lot fewer sort of pure AI solutions now where the AI just one-shots the problem.

Read the full transcript

31:23So there was a month where that happened and that has stopped. Not for lack of trying. I know three separate attempts to get frontier model AIs to just attack every single one of the problems simultaneously. And they picked up some minor observations or maybe they found that some problems are already solved in the literature, but there hasn't been any further AI purely powered solution yet. People are using AI a lot currently. So someone might use AI to generate a possible proof strategy. And then another person will use a separate AI tool to critique it or rewrite it or generate some numerical data for it or do a literature survey.

32:04And some problems have been solved by an ongoing conversation between lots of humans and lots of AI tools. But it does seem like it was this one-off thing. So maybe one analogy for these problems is like, imagine like there's all these, you're in some sort of mountain range with all kinds of cliffs and walls. And maybe there's a little wall, which is maybe like three feet high and one that's six feet high and then there's 15 feet high and then there's some mile high cliffs. and you're trying to climb as many of these cliffs as possible but it's in the dark uh we don't know which ones are tall which ones are short and um so you know we try to light some candles and make some maps and and slowly we kind of figure out uh some of them are climbable some of them we can identify some some partial um track in the wall that you can reach first um and then these these ai tools they're kind of like these jumping machines that can kind of jump you know two meters in the air you know higher than any human and sometimes they jump in the wrong direction and sometimes they crash but sometimes they they can reach um um the tops of of the lowest um you know um walls that we couldn't reach before and so we basically set them loose in this mountain range hopping around and you know and then there's this exciting period where they could actually find all the um all the low ones um and they could reach them um but then uh there's been no I mean, maybe if the next time there's a big advance in the models, then they will try it again and maybe a few more will be breached.

33:42But it's a different style of doing mathematics than sort of the, you know, so normally we would hill climb and, you know, we would make little markers and try to identify partial things. And, you know, these tools, they either succeed or they fail. and they've been really bad at creating sort of partial progress or identifying intermediate stages that you should focus on first. Again, going back to this previous discussion, we don't have a way of evaluating partial progress. The same way you can evaluate a one-shot success or failure of solving a problem. So there's two different ways to think through what you've just said.

34:22And one of them is more bearish on AI progress and one of them is more bullish. and bearish one being, oh, they're only getting to a certain height of wall, which is not as high as humans are reaching. And the second is that, well, they have this powerful property that once they achieve a certain waterline, they can fill every single problem that is available at that waterline, which we simply can't do with humans, where we can't make a million copies of you and give each of them a million dollars of inference compute and have you do 100 years of subjective time research on a hundred different problems at the same time or a million different problems at the same time.

34:58But once AIs reach Terence Tower level, they could do that. And once they reach intermediate levels, they could do the intermediate version of that. So the same reason that we should be bearish now is the reason we should be especially bullish, not even when they achieve superhuman intelligence, but just when they achieve human level intelligence because their human level intelligence is qualitatively wider and more powerful than our human level intelligence. I agree. yeah so they excel at breadth and humans excel at depth um human experts at least yeah so um i think they're very complementary um but our current uh way of doing math and science is focused on depth because that's where the human uh expertise is because humans can't do breath um but uh yeah so we have to redesign uh the way we do science to take full advantage of of this breadth capability that we now have um so as i said we do we should have a lot more effort in creating very broad classes of problems to work on rather than one or two um really um deep important problems i mean we should still have the deep important problems um and humans should still be working on them um but but now now we have this other way of of of doing um of doing science you know i mean we can explore entire new fields of science by by first getting the these broad um moderately competent AI to sort of map it out and clear out all the easy, make all the easy observations and then identify certain islands of difficulty, which, you know, then human experts can come and work on.

36:28So I see very much a future of very complementary science. Eventually you would hope to get both breadth and depth, you know, and somehow get the best of both worlds. But I think we need practice with the breadth side. It's too new. We don't even have the paradigms really to make full advantage of it. But we will. And then science will be unrecognizable after that. To this point about complementarity, programmers have noticed that they're way more productive as a result of these AI tools. And I don't know if you as a mathematician feel the same way, but it does seem like one big difference between Vibe coding and Vibe researching is that with software, the whole point of the thing is to have some effect on the world through your work.

37:20And if it leads to you better understanding a problem or you coming up with some clean abstraction to embody in your code, that is instrumental to the end goal. Whereas maybe with research, the reason we care about solving the Millennium Pipers problems is presumably that in the process of solving them, we discover new mathematical objects or better new techniques and those who understand our civilization's understanding of mathematics. And so the proof is sort of instrumental to the intermediate work. I don't know if you agree with that dichotomy or if that in any way will explain the relative uplift we'll see in software versus research.

37:58Right. Yeah. So certainly in math, the process is often more important than the problem itself. The problem is kind of a proxy for measuring the progress. and I think even in software there's there's different types of software tasks I mean you know like if you just kind of create a web page that does the same thing that a thousand other web pages do um there's there's sort of no skill to be learned well um there's there's some skill maybe that the individual programmer could pick up um but you know for for kind of a boilerplate type code definitely um um you know it's it's it's something that you should definitely offload to AI.

38:34But, you know, sometimes once you make the code, you know, you still have to maintain it and there's issues with upgrading it and making it compatible with other things. And that, I think, I've heard that programmers are reporting, you know, that even if an AI can create the first prototype of a tool, making it mesh with everything else and making it interact with the real world in the way they want, I mean, that's an ongoing process. And if you didn't have the skills that you pick up from writing the code, that may impact your ability to maintain it down the road. So certainly mathematicians, we've used problems to build intuition and to train people to have a good idea as what's true, what to expect, what is provable, what is difficult.

39:25and so just getting the answers right away may actually inhibit that process I mean so I made a distinction between theory and experiment before so in most sciences there's an equal division between there's a theoretical side and experimental side but math has been almost unique it's almost entirely theoretical we pay the premium on sort of trying to have coherent clean theories of why things are true and false. And we haven't done much experiments as to, like, you know, maybe we have two different ways to solve a problem. Which one is more effective? We have some intuition, but we haven't done large-scale studies where we take a thousand problems and we just test them.

40:10But we can do that now. So I think AI-type tools, we really will actually revolutionize the experimental side of math, where you don't care so much about individual problems and the process of solving them, but you want to gather just large scale data about what things work, what things don't. Same way that if you're a software company and you want to roll out a thousand pieces of software, you don't really want to handcraft each one and learn lessons from each. You just want to find what are the workflows that you scale. So we don't yet... The idea of doing mathematics at scale is at its infancy, but that's where AI is really going to revolutionize the subject.

40:52Interesting. I feel like a big crux in these conversations about how much how good AI will be for science is, I think you said this, it's like, oh, they're using existing techniques and modifying them. And it would be interesting to understand how much progress one can make simply from using existing techniques. If I looked at the top math journals, how many of the papers are coming up with whatever coming up with the technique means, doing that versus using existing techniques in new problems? And what the overhang is, where if you just applied every known technique to every open problem, would that just constitute a humongous uplift in our civilization's knowledge, or would that not be that impressive and useful?

41:35This is a great question, and we don't have the data to fully answer it yet. Certainly a lot of work that human mathematicians do, when you take a new problem, one of the first things we do is we just find, we look at all the standard things that have worked on similar problems in the past and we try them one by one. And sometimes that works. And that's still worth publishing sometimes because the question was important. Sometimes they almost work and you have to add one more wrinkle to it. And that's also interesting. But then the papers that go into the top journals are usually ones where the existing methodians methods can kind of solve, you know, 80 % of the problem, but then this is 20%, which is resistant.

42:16And a new technique has to be invented to fill in the gaps. It's very, very rare now that a problem gets solved with sort of no reliance on past literature, where all the ideas come out of nowhere. You know, that was more common in the past, but math is so mature now that it's just so much of a handicap to not use the literature first. So AI tools are getting really good at the first part of that, just trying all the standard techniques on a problem, often now actually making fewer mistakes in implementing them than humans. They still make mistakes, but I've tested these tools on little tasks that I can do, and sometimes they pick up errors that I make, sometimes i pick up errors that they make it's about a tie right now uh for um but um yeah i i haven't yet seen them take the next step you know so so when there are holes in in in the argument where none of the things are working to to how then what do you do um and then they can kind of suggest random things and it but it it um often i find that trying to chase them down make them work and finding they don't work it wastes more time than it saves yeah so um now so i think some fraction of problems that we currently think are hard will will fall from this this method um i mean especially the ones that haven't received enough attention um so like with the urdish problems you know like almost all of the 50 problems that were solved by ai's were ones for which basically there was no literature i mean irdish proposed from once or twice um i think maybe some people tried it casually and they couldn't do it, but they never wrote up anything.

44:03But there was a solution and it was just, you know, maybe combining this one obscure technique that not many people know about with some other result in the literature. And that's the kind of median level of what AI can accomplish. And that's really great. It clears out 50 of these problems. So I think you will see some isolated successes. But what we found, so people have done large scale sweeps of these early problems. If you only focus on the success stories that get broadcast on social media, it looks amazing. All these problems that haven't been solved before for decades, now they're falling.

44:38But whenever we do a systematic study, any given problem, an AI tool has a success rate of maybe 1 % or 2%. It's just that they can buy a scale, and if you just pick the winners, it looks great. So I think there'll be a similar thing happening with, you know, there are hundreds of really prestigious difficult math problems out there a couple may make um you know some am may get lucky and actually solve them and there was there was some some backdoor to solve the problem that everyone else missed um and that will get a lot of publicity um but then people will try these fancy tools on their own favorite problem and they will again experience the one to two percent success rate right so um there will be a lot of noise and amongst the signal of sort of when they're working, when they're not.

45:26We have to do, yeah, it's increasingly important to collect these really standardized data sets. You know, there are efforts now to create a standard set of challenge problems for AI to solve and not just rely on the AI companies to only publish their wins and not disclose their negative results. So that will maybe give more clarity as to where we're actually at. Well, I think it's worth emphasizing how much progress in AI constitutes already to have models that are capable of applying some technique that nobody had written down is applicable to this particular problem. The progress is simultaneously amazing and disappointing.

46:02It is a very strange feeling to see these tools in action. But also acclimatized really quickly. I remember when Google's web search came out 20 years ago and it just blew all the searches out of the water. You're just getting relevant hits on the front page like perfectly almost exactly what you wanted and it was amazing and then after a few years you just took for granted that you could you could just google anything um and yeah so a lot of yeah i mean 2026 level ai would be stunning in 2021 and a lot of it you know face recognition natural speech uh yeah doing you know college level math problems we just take for granted now right yeah okay so speaking of 2026 yeah you made a prediction in 2023 that i think by 2026 What was it?

46:49It would be like a colleague in mathematics or? Yeah, a trustworthy co-author if used correctly. Which is looking pretty good in retrospect. Yeah, I'm pretty pleased. So, you know, let's see if we can continue the streak. You personally are 2x more productive as a result of AI. What year would you say that? Yeah, so productivity, I think, is not quite a one-dimensional quantity. like I'm definitely noticing that the style in which I do mathematics is changing quite a bit and the type of things I do so for example my papers now have a lot more code a lot more pictures because it's so easy to generate these things now so some plot which have taken me hours to do now I can I can do in minutes but in the past I just wouldn't have put the plot in my paper in the first place I would just talk about it in words so it's hard to measure what 2x means um so yeah on the one hand you know i think the type of papers that i would write today if i had to do them without ai assistance they would definitely take five times longer but interesting but i would not write my papers that way 5x so yeah that's but but it's it's because but these are sort of uh auxiliary i mean you know the um yeah so so things that yeah things like like um doing a much deeper literature search, supplying a lot more numerics.

48:13I mean, they enrich the paper. So, yeah, the core of what I do, like actually solving the most difficult part of a math problem, that hasn't changed too much. I still use pen and paper for that. But, you know, there's lots of silly things. I use an AI agent now to reformat. Like sometimes all my parentheses are not quite the right size. You know, I just manually change my hand and I can get an AI agent to sort of do all that quite nicely now in the background. So, yeah, they really sped up lots of secondary tasks. They haven't yet sort of sped up the core thing that I do, but it's allowed me to sort of add more things to my papers.

48:58yeah but by the same token like if i were to write a paper i wrote in 2020 again and not add all these extra features but just have something of the same sort of level functionality yeah then that doesn't have hasn't saved that that much uh to be honest uh yeah so it's made the papers sort of richer and broader but not necessarily deeper you made this distinction between artificial cleverness and artificial intelligence and i would like to better understand those concepts what is an example of um uh intelligence that is not just cleverness yeah so um it's intelligence is famously hard to define it's one of these things that you kind of know it when you see it um but when i when i talked to someone um and we're trying to collaboratively solve a math problem together um there's this conversation where you know we neither of us knows how to solve the problem um initially but um one of us has some idea and and it looks promising and and so then then we have some sort of prototype strategy and then we test it and then it doesn't work but then we we modify it and there's some adaptivity and and um and and uh continual improvement of of of the idea over time.

50:17And eventually, you know, we sort of, we've systematically mapped out what doesn't work, what does work, and we can kind of see a path forward, but it's evolving with our discussion. And this is not quite what the AI is. The AI can kind of mimic this a little bit. So to go back to this analogy of these jumping robots, you know, so, you know, they can jump and fail and jump and fail and jump and fail, But what they can't do is they kind of jump a little bit and they reach some handhold, but then they sort of stay there and then they pull other people up and then they try to just jump from there.

50:55There isn't this cumulative process which is sort of built up interactively. It seems to be a lot more trial and error and just repetition brute force, which it scales and it can work amazingly well in certain contexts. But yeah, this idea is sort of building up cumulatively from partial progress is kind of what's still not quite there yet. Interesting. You were saying if Gemini 3 or Claude 4.5, whatever, solves a problem, it is not the case that its own understanding of math has progressed. Or even if it works on a problem without solving it, it's not that its own understanding of math has progressed.

51:36Yeah. You run a new session and it's forgotten what it just did. Right. It has no new skills to attach to, to build on related problems. Maybe what you just did is part of 0.001 % of the training data for the next generation. So maybe eventually some of it gets absorbed. So Terence talks about the importance of decomposing particularly in Arlie problems into a series of easier chunks. Even if this doesn't result in the full solution, approaching problems in this way helps you build up the intuitions and practice the techniques that you'll need to keep making progress. But models today tend to struggle with these kinds of problem-solving techniques.

52:13That's where LabelBox comes in. LabelBox helps you train models not just to get the right answer, but to think the right way. They've operationalized these reasoning behaviors into rubrics, giving you the ability to evaluate every important dimension of a model's output. These rubrics go beyond simple correctness. Did the model reach for the right tools? Did it check its own work and explore alternative paths? How clear was its response? These skills are useful across domains. math, physics, finance, psychology, and more. And they're becoming increasingly important as models take on harder, open-ended problems, some of which have multiple solutions and some of which we don't even know the solutions to.

52:48LabelBox can get you rubrics tailored to your domain, helping you systematically measure and shape how your models think. Learn more at labelbox.com slash Dwarakash. One big question I have is, how plausible is it that if we just keep training AI that get better and better at solving problems in Lean, that they will continue to solve more and more impressive problems. And then we will, in retrospect, be surprised at how little insight we got from some Lean solution to proving the rebound hypothesis or something. Or do you think it is a necessary condition of solving the rebound hypothesis, even by an AI that is totally doing it in Lean, that the constructions which are made, the definitions which are created, even in the Lean program, have to advance our understanding of mathematics.

53:34or do you think it could just be assembly code gobbledygook? Oh, yeah, we don't know. I mean, some problems have been basically solved by pure brute force. A full-color theorem is a famous example. We have still not found a conceptually elegant proof of this theorem. It basically, and maybe we never will. I mean, some problems may only be solvable by just splitting into some enormous number of cases and doing a brute force unincitible computer analysis on each case. I mean part of the reason we prize problems like a hypothesis is that we're pretty sure that that something amazing has to a new type of mathematics has to be created or a new connection between two previously unconnected areas of mathematics has to be discovered to make this work we don't even know what the shape of the solution is but it doesn't feel like a problem that will be solved just by exhaustively checking cases or something I mean it could be false actually.

54:32So we could actually, there is an unlikely scenario that the hypothesis is false and there's this, you can just compute, oh, here's a zero off the line and a massive computer calculation verifies it. That would be very disappointing. I don't know. I do feel that, you know, fully autonomous one-shot approaches are not the right approach for these problems. I mean, I think you will get a lot more mileage out of the interplay between humans collaborating with these tools.

55:07And I can see one of these problems being solved by some smart humans assisted by some extremely powerful AI tools. But the exact dynamic may be very different from what we envision right now. I mean, it could be a collaboration of a type that just doesn't exist yet.

55:29yeah, I mean, there may be a way to generate, you know, a million variants of the human's data function and do some data analysis, AI-assisted data analysis, and we discover some pattern between connecting them, which we didn't know about before, and this lets you transform the problem into a different area of mathematics. I mean, there could be all kinds of scenarios. So suppose the AI figures it out, and latent in the lean, is some brand new construction, which if you realize the significance, we would be able to apply it in all these different situations. How do you even recognize it, right?

56:06Like if you just, again, a very naive question, but if you come up with the equivalent of like, Descartes comes up with this idea, oh, you can have this coordinate system where you can unify algebra and geometry, but in lean code, it would just look like R to R and it wouldn't look that significant or something. Or similarly, I'm sure there's other constructions which have this kind of property. Well, the beauty of formalizing a proof in something like Lean is that you can take any piece of it and study it atomically.

56:34So when I read a paper with my humans, which shows some difficult problem, there's often some big sequence of lemmas and theorems and things. And so ideally, the author will talk their way through what's important, what's not. But sometimes they don't reveal what steps were the important ones and which ones are just kind of boilerplate standard steps. But you can study each lemma in isolation and some of them I can say, oh, this looks fairly standard. This resembles something I'm familiar with. I'm pretty sure there's nothing interesting going on here. But this lemma, oh, that's something I haven't seen before.

57:08And I could see why if you had this result, that would really help prove the main result. You can assess whether some things are really sort of key to your argument or not. and lean really facilitates that you know you can you can you can you know the individual steps are identified really precisely um i think in the future there'll be um you know there'll be entire professions of of mathematicians who might take a giant um lean generally proof and maybe you know do some ablation on it or something i try to remove steps of parts of it and um and try to find it find more elegant ways you know you know maybe some other ais to sort of do some reinforcement learning how can you make the proof more elegant and and and uh um maybe other ais will grade whether this proof looks better or not.

57:55One thing that will change quite a bit in the near future is that until recently, writing papers was the most time-consuming and expensive part of the job. And so you did it very rarely. You only wrote up your results once everything was all the other parts of your argument were checked out and things, because just rewriting it again, refactoring was a total pain. But that's one thing that's become a lot easier now with modern AI tools. So you don't have to have just one version of your paper. Once you have one, people can generate hundreds more. So yeah, one giant messy lean proof may not be very meaningful or understandable on its own, but other people can refactor it and do all kinds of things with it.

58:39We have seen with the Erdős problem website, an AI will generate a proof and then here's 3 ,000 lines of code that verify the proof. but then people got other AIs to summarize the proof and people write their own proofs. There's actually post-processing, once you actually have one proof, we actually have a lot of tools now to deconstruct it and interpret it. It's a very nascent area of science or mathematics, but I'm not as worried about, so some people are concerned, what if the real hypothesis is proven with a completely incomprehensible proof? I think once you have the artifact of a proof, we can do a lot of analysis on it.

59:21You posted recently that it would be helpful to have a formal or semi-formal language for mathematical strategies as opposed to just mathematical proofs, which is what Lean specializes in. I would love to learn more about what that would involve or look like. We don't really know. I mean, we've been very lucky in mathematics that we have worked out the laws of logic and mathematics, but this is actually a fairly recent accomplishment. I mean, it was started by Euclid. millennia ago, but only in the early 20th century did we finally list our Q, the axioms of mathematics, or the standard axioms of what we call ZFC, and the axioms of first-order logic, and this is what a proof is, and this we've managed to automate and have a formal language for.

1:00:05But there could be some way to assess plausibility of certain, you know, so you have a conjecture that something is true, you test a few examples, and it works out, like how does this increase your confidence that the conjecture is true? We have a few sort of mathematical ways to model this, Bayesian probability, for example, but you often have to set certain base assumptions and there's a lot of subjectivity still in these tasks. So it's not clear. I mean, this is more of a wish than a plan to develop these languages. But just seeing how successful having a formal framework in place like Lean has made deductive proofs so much easier to automate and train AI on.

1:01:01If there was some similar framework, so the bottleneck for using AI to create strategies and make conjectures is we have to rely on human experts and the test of time to validate whether something's plausible or not. If there was some semi-formal framework where this could be done semi-automatically in a way that isn't sort of easily hackable. Of course, it's really important with these formal proof assistants that there's no backdoors or exploits that you can do to somehow get your certified proof without actually proving it because reinforcement learning is just so, so good at finding these backdoors.

1:01:51But yeah, if it's not a framework that sort of mimics how scientists talk to each other in a semi-formal way, using data and argument, but also constructing narratives and there's some subjective aspect of science that we don't know how to capture in a way that we can insert AI into them in any useful way. Interesting. So yeah, this is a future problem. I mean, there are research efforts to try to create automated conjectures and maybe there are ways to benchmark these and get some way to simulate this. But this is it's all very, very new science. Can you help me get some intuition for. I have two step questions.

1:02:43One, it would be very helpful to have a tangible sense of. It would be helpful to have a specific example of what something like this would look like that the way scientists communicate that we can't formalize yet. And two, it seems almost definitionally paradoxical to say, building up some narrative or building up some natural language explanation, and then also having something which you could have formalized. And I'm sure there's some intuition behind where that overlap is, and I'd love to understand that better. all right so so an example of of a conjecture so um gauss um was interested in the prime numbers and he computed he created one of the first mathematical data sets he just computed the first hundred thousand prime numbers or so um hoping to find patterns um and he did find a pattern but maybe not the pattern he was expecting he found a statistical pattern in the primes that if you count how many primes there are up to 100 1 000 um um one million and so forth they get sparser and sparser, but the drop-off in the density was inversely proportional to the natural logarithm of the range of numbers.

1:03:57So he conjectured what we now call the prime number theorem. The number of primes up to x is like x divided by the natural log of x. And he had no way to prove this. It was data-driven. So this was a conjecture. it was revolutionary for its time because it was maybe the first really important conjecture of math that was statistical in nature you know so normally you talk about patterns like maybe the spacing between the primes has a certain regularity or something but yeah but this was really something which it didn't tell you exactly how many primes there were in any given range it just gave you an approximate approximation that got better and better as you went further and further out but it it helped so it started the field of what we call analytic number theory but it was the first in many conjectures like this many of which got proved which sort of started consolidating the idea that the prime numbers actually didn't really have a pattern, that they behaved like random sets of numbers with a certain density.

1:05:02I mean they had some patterns, like they're almost all odd and they're not actually random they're what's called pseudo random i mean there's no random number generation involved in creating the prime numbers but um over time it became more and more productive to think of the primes as as if they were just generated by some some some god rolling dice all the time and just creating this this random set um and this allowed us to make all these other predictions um so there's a still open conjecture in in number three called the twin prime conjecture that there should be infinitely many pairs of primes that are twins distance two apart like 11 and 13.

1:05:37We can't prove that, and there's actually good reasons why we can't prove it, but because of this statistical random model of the primes, we are absolutely convinced it's true. We know that if the primes were sort of generated by flipping coins or something, that we would, just by random charts, just like infinite monkeys at a typewriter, we would see twin primes appear over and over again. And we have, over time, developed this very accurate conceptual model of what the primes should behave like based on statistics and probability, but it's all mostly heuristic and non-rigorous but extremely accurate.

1:06:10So the few times when we actually can prove things about the primes, it has matched up with the predictions of what we call the random model of the primes. So we have this conjectural concept framework for understanding the primes that everyone believes in. And it's the same reason why we believe the real hypothesis is true, why we believe that cryptography based on the primes is basically mathematically secure, things like that. It's all part of this belief. In fact, one reason why we care about the Riemann hypothesis is that if the Riemann hypothesis failed, we knew it was false. It means it would be a serious blow to this model that it would mean there's a secret pattern to the primes that we were not aware of.

1:06:53And I think we would very rapidly abandon any cryptography based on the primes because if there was one pattern that we didn't know about, there's probably more. and these patents can lead to exploits in crypto and it's going to be a big, big shock. So we really want to make sure that doesn't happen.

1:07:14So we've been convinced of things like agreement of offices and things over time, but some of it is experimental evidence, some is the few times we've been able to make theoretical results, they've always aligned. It is possible that the consensus is wrong and we've all just missed something very basic. You know, there have been paradigm shifts in the past in scientific history. Yeah, but we don't really have a way of measuring this. I think partly because we don't have enough data on how math or science develops. We have one timeline of history and we have like, you know, a hundred stories of turning points in history.

1:07:53If we had access to a million alien civilizations and each of the different development of history and of science in different orders, then maybe we actually have a decent shot at an understanding of how do we measure what is progress and what is a good strategy. And we could maybe start formalizing it and actually having a framework. Maybe if what we need to do is actually start creating lots of mini universes or simulations of AI solving very basic problems, in arithmetic or whatever, but coming over their own strategies for doing these things and having these little laboratories to test. I mean, there are people who investigate like what's the smallest neural network that can do 10-digit multiplication and things like that.

1:08:41I think we could actually learn a lot just from evolving small AIs on simple problems. We could learn a lot. I was super excited when Mercury reached out about sponsoring the podcast because I've been banking with them for years. I think I opened my first account with them in 2023. Something I've come to appreciate over the last few years is that Mercury is constantly updating things and adding new features. Take their newest feature, Insights. Insights summarizes your money in and out, showing you your biggest transactions and calling out anything that deserves extra attention. Like maybe your revenue from a particular partner has gone down, or you've got a big uncategorized purchase that needs to be investigated.

1:09:15It's a super low friction way for me to keep tabs on my business and make quick decisions. For example, I tried to invest any cash that I don't need on hand to keep running the business. With insights, with just a couple of clicks, I was able to see exactly how much money I spent in each month of 2025. And that lets me know exactly how much cash I'll need for the next year or so of operations. And then I can go invest the rest. Mercury just keeps adding new features like this. Go to mercury.com to check it out. Mercury is a fintech company, not an FDIC insured bank. Banking services provided through Choice Financial Group and Column N.A., members FDIC.

1:09:48You have to learn about new fields, not only very rapidly, but deeply enough to contribute to the frontier. So in some sense, you're also one of the world's greatest autodidacts. What is your process of learning about a new subfield in math? What does that look like? Yeah, so I certainly identify with kind of the, yeah, we talked about depth and breadth before. And it's not purely human AI distinction. I mean, humans also split, so I think it was Irving who split them into hedgehogs and foxes. And a hedgehog knows one thing very, very well, and a fox knows a little bit about everything. So I definitely, I think of myself as a fox.

1:10:31I work with hedgehogs a lot, and sometimes I can be a hedgehog if need be. But yeah, so I've always had a little bit of an obsessive streak. If there's something which I read about, which I feel like I should understand, I have the capability to understand this, but I don't understand why it works. There's some magic in it that, you know, so someone was able to use a type of mathematics I'm not familiar with and get a result which I would like to prove, and I can't do it by myself, but they could do it by their method. Then I wanted to find out what was their trick. It bugs me that someone else can do something which I think I can do, but I can't.

1:11:08So I've always had that kind of obsessive completionist type streak. I've had to wean myself off computer games because I start a game, I want to play it to completion, so both the levels.

1:11:23So that's one way in which I learn new fields. I collaborate with a lot of people who have taught me other types of mathematics. I just make friends with another mathematician who's working on another area of mathematics, and I find their problems interesting, but they have to teach me some of the basic tricks and what's known, what's not known. And I learned a lot from that. I found that writing about what I've learned, I have a blog where I sometimes record things that I've learned. Because in the past, when I was younger, I would learn something and do this cold trick. And I'd say, okay, I'm going to remember this.

1:12:02And then six months later, I'd forgotten. I remember remembering it, but I can't reconstruct my arguments. And the first few times it was so frustrating to have understood something and then lost it. I sort of resolved I should always write down anything cool that I've learned. And this is part of how this blog came about. How long does it take you to write a blog post? It's something I often do when I don't want to do other work. You know, like there's some referee report or something. There's something that feels slightly unpleasant for me to do at the time. And so writing a blog, it feels creative and fun.

1:12:37like it's something that i do for myself um so maybe depending on on the topic it could be a quick you know half an hour or several hours but i um it doesn't because it's something that i do sort of voluntarily it doesn't feel like it it it doesn't feel uh time flies when i write these things as opposed to sort of doing something which i have to do for administrative reasons but it's just that it's it's drudgery okay those are tasks that ai is really helping with nowadays actually Is it if like civilization could from first principles decide how to use Terry Tau's time? You know, it's like a limited resource.

1:13:17What is the biggest diff between if the veil of ignorance got to decide how to use Terry Tau's time versus what it does now? This podcast wouldn't be happening. Yeah. As much as I complain about certain tasks that I don't want to do, but I have to do. So as you get more senior in academia, you get more and more responsibilities. You get some more committees and whatever. But I have also found that a lot of events that I kind of reluctantly went to because I was obliged to for one reason or another. Because it's outside my comfort zone, I often find interactions with people who I wouldn't normally talk to, like you, for instance.

1:13:56And I would learn interesting things and have interesting experiences. and I would have opportunities to then network with other people that I would never have done before. So I do believe a lot in serendipity. I mean, I do optimize my time when I... So there's some portions of a day where I do schedule very carefully. But I have been willing to sort of leave some portions just, okay, I'm going to do something which is not my usual thing and maybe it'll be a waste of my time, but maybe I will learn something. And more often than not, I feel like I've gotten a positive experience, which is not something I would have planned for.

1:14:41And yeah, so I believe a lot in serendipity. And maybe there's a danger actually that, you know, in modern societies, it's not just AI, but we've become really good at optimizing everything. And maybe we are optimizing, we're not optimizing a lot of optimization. that you know with with with covert for example um we we switched like we switched a lot to remote meetings um and so everything was scheduled now and so uh we kept busy at least in academia you know we met almost the same number of people that we met in person but everything had to be planned um you had to schedule things in advance and what we lost out on was sort of the the casual like knocking on the hallway, just meeting someone while getting a coffee.

1:15:27And there's serendipitous interactions that you may think are not optimal, but actually are really important. When I was a grad student, I would go down to the library to look for a journal article. I had to physically go down to the library, check out the journal, and read your article. And sometimes the next article, you can just browse through, and the next article is also interesting. sometimes it wasn't but but you could accidentally find interesting things um which is something which has basically been lost now because you can just type in you know if you if you want to access an article now you just type it into to a search engine or even an ai and you can get instantly what you want but you don't get so the accidental things that you might have have gotten if you'd done it more inefficiently um so um yeah there have been times when i'm um i spent a year once at the Institute for Advanced Study, which is a great place to, you know, there's no distractions.

1:16:27You're there to just do research. And like the first few weeks you're there, like it's great. You're getting all these papers written up that you've been wanting to do for a long time. You've been thinking about problems for blocks of hours of a time. But I find if I stay there for more than several months, like I run out of inspiration somehow. Like I get bored. I actually surf the internet a lot more. You actually do need a certain level of distraction in your life. It somehow adds enough randomness and temperature, high temperature if you need. So yeah, I don't know the optimal way to schedule my life.

1:17:03It just seems to work. I'm very curious when you expect AIs that can actually do frontier math better than at least as good as the best human mathematicians. I mean, in some ways, they're already doing frontier math that is super intelligent. that humans can't do, but it's a different frontier from what we're used to. I mean, you could argue that calculus were doing frontier math that humans could not accomplish, but it wasn't, you know, number crunching. But replacing Terry Tao completely. I mean, what do you want me for? You'll just go on all the podcasts after.

1:17:47I'm not sure we've, It might not be the right question to ask.

1:17:56I think within a decade, a lot of things that mathematicians currently do, we spend a lot of the bulk of our time doing it, and a lot of stuff we put in our papers today can be done by AI. But we will find that that actually wasn't the most important part of what we do. you know a hundred years ago a lot of mathematicians were just solving differential equations like people needed physicists needed some exact solution to some system and they were just they hired a mathematician to go through the calculus and work out the solution to this fluid equation or whatever a lot of what a 19th century mathematician would do you could make a call to Mathematica or Wolfram Alpha or a computer algebra package, or now more recently an AI, and it would just solve the problem in a few minutes.

1:18:50But we moved on. We worked on different types of problems after that. Once computers came along, computers used to be human. People used to liberally create log tables and work out primes as Gauss did, and that has all been outsourced to computers. But we moved on. in genetics, you know, to sequence the genome of a single organism, that was an entire PhD of a geneticist, you know, so carefully, you know, separating all the chromosomes and whatever. And now you can just spend$1 ,000 and send it to a sequencer and get it done. But genetics is not dead as a subject. You move to a different scale.

1:19:30You know, maybe you study whole ecosystems rather than individuals. I take your point. But on the question of, well, when is most mathematical progress, almost all mathematical progress happening by AI. So that if you find out, oh, this year, a millennium price problem has been solved, you would put, you know, a 95 % odds that an AI did it autonomously. Surely there will be such a year. I guess, I mean, I do believe that hybrid human plus AIs will dominate mathematics for a lot longer. It will depend, it will require some additional breakthroughs beyond what we already have. So it's going to be sarcastic.

1:20:10I think AI is currently very good at certain things, but they're really terrible at others. And while you can sort of add more and more frameworks on top to kind of reduce the error rates and make them work with each other a bit more and so forth, it feels like we don't have all the ingredients to really have a truly satisfactory sort of replacement for all. intellectual tasks. It is complementary currently. It's not a placement. But maybe, I mean, because current-level AIs will accelerate science in so many ways, hopefully, I mean, new discoveries, new breakthroughs will happen more quickly. I mean, it's possible that also by somehow destroying serendipity, we actually inhibit certain types of progress.

1:21:07Anything is possible, really, at this point. I think the world is very, very unpredictable at this point in time. What is your advice to somebody who would consider a career in math or is early in a career in math, especially in light of AI progress? How should they be thinking about their career differently, if at all, as a result of AI progress? Yeah, so we live in a time of change. It is, as I said, we live in a particularly unpredictable era. And I think things that we've taken for granted for centuries may not hold anymore. So, yeah, the way we do everything, not just mathematics, will change.

1:21:53And, you know, so I think, which is, you know, I mean, in many ways, I would prefer the much more boring, quiet era where things are much the same as they were 10 years ago, 20 years ago. But so I think one just has to embrace that there's going to be a lot of change and that, you know, the things that you study, some of them may become obsolete or revolutionized, but some things will be retained. And so you somehow always have to keep an eye on, there'll be a lot of opportunities for things that you wouldn't be able to do before. So, I mean, in math, you previously had to basically go through years and years of education, be a math PhD, before you could contribute to the frontier of math research.

1:22:45But now it's quite possible at the high school level or whatever that you could get involved in a math project and actually make a real contribution because of all these AI tools and Lean and everything else. So there'll be a lot of non-traditional opportunities to learn. So you need a very adaptable mindset. You know, there'll be one for pursuing things just for curiosity, for playing around. And, I mean, you still need to get your credentials. I mean, I think for a while it'll still be important to sort of still go through traditional education and learn math and science the old-fashioned way for a while.

1:23:28But you should also be open to very, very different ways of doing science, some of which don't exist yet. Yeah, so it's a scary time, but also very exciting. Awesome. That's a great note to close on. Thanks so much. Yeah, pleasure.

From the publisher

We begin the episode with the absolutely ingenious and surprising way in which Kepler discovered the laws of planetary motion.

People sometimes say that AI will make especially fast progress at scientific discovery because of tight verification loops.

But the story of how we discovered the shape of our solar system shows how the verification loop for correct ideas can be decades (or even millennia) long.

During this time, what we know today as the better theory can actually make worse predictions.

And the reasons it survives this epistemic hell is some mixture of judgment and heuristics that we don’t even understand well enough to actually articulate, much less codify into an RL loop. Hope you enjoy!

Watch on YouTube; read the transcript.

Sponsors

- Jane Street loves challenging my audience with different creative puzzles. One of my listeners, Shawn, solved Jane Street’s ResNet challenge and posted a great walk-through on X. If you want to try one of these puzzles yourself, there’s one live now at janestreet.com/dwarkesh.

- Labelbox can get you rubric-based evals, no matter your domain. These rubrics allow you to give your model feedback on all the dimensions you care about, so you can train how it thinks, not just what it thinks. Whatever you’re focused on—math, physics, finance, psychology or something else—Labelbox can help. Learn more at labelbox.com/dwarkesh.

- Mercury just released a new feature called Insights. Insights summarizes your money in and out, showing you your biggest transactions and calling out anything worth paying attention to. It’s a super low-friction way to stay on top of your business. Learn more at mercury.com/insights.

Timestamps

(00:00:00) – Kepler was a high temperature LLM

(00:11:44) – How would we know if there’s a new unifying concept within heaps of AI slop?

(00:26:10) – The deductive overhang

(00:30:31) – Selection bias in reported AI discoveries

(00:46:43) – AI makes papers richer and broader, but not deeper

(00:53:00) – If AI solves a problem, can humans get understanding out of it?

(00:59:20) – We need a semi-formal language for the way that scientists actually talk to each other

(01:09:48) – How Terry uses his time

(01:17:05) – Human-AI hybrids will dominate math for a lot longer



Get full access to Dwarkesh Podcast at www.dwarkesh.com/subscribe

More from Dwarkesh Podcast

All 94 episodes
Terence Tao – Kepler, Newton, and the true nature of mathematical discoveryDwarkesh Podcast · 1 h 24 min
Listen in VO