AI That Can Prove It’s Right: Verification as the Missing Layer in AI — Carina Hong

26 Feb 2026 · 1 h 4 min · 34 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The MAD Podcast with Matt Turck - Episode Summary

Episode Title

AI That Can Prove It’s Right: Verification as the Missing Layer in AI — Carina Hong

Podcast Overview In this episode of The MAD Podcast, Matt Turck interviews Carina Hong, a young math prodigy and CEO of Axiom Math. They discuss the innovative approach of Axiom to formal verification in AI, aiming for a system that not only generates solutions but can also prove their correctness.

---

Key Themes and Topics Discussed

  1. The Need for an AI Mathematician
  2. Carina argues that the world needs AI mathematicians to tackle a vast number of unsolved mathematical problems efficiently.
  3. She believes that solving math can also help in addressing other complex issues like verification and optimization.
  1. AxiomProver's Achievements
  2. AxiomProver achieved a perfect score on the Putnam Exam (12 out of 12) within months of its establishment.
  3. Successfully solved four open research conjectures autonomously, marking a significant milestone in the use of AI for mathematical problem-solving.
  1. Understanding the Technology
  2. Lean: A programming language for proofs that Axiom uses. Carina explains how Lean allows for formal verification of mathematical statements and is essential for Axiom's approach.
  3. Axiom's strategy differs from other companies like DeepMind and OpenAI by focusing on formal reasoning rather than informal or brute force approaches.
  1. Formal vs. Informal Reasoning
  2. Carina emphasizes the distinction between informal reasoning (natural language) and formal reasoning through systems like Lean.
  3. The importance of auto-formalization (converting informal math problems into formal language) is highlighted as a key component of their system.
  1. Challenges and Opportunities in AI Verification
  2. The discussion includes the "reward hacking" problem and the difficulty of ensuring an AI's reasoning process is verifiable and reliable.
  3. The aim is to create AI that is 100% accurate, solving the AI hallucination problem in mathematical and logical reasoning.
  1. Future of AI in Math and Beyond
  2. Carina mentions a "math renaissance," where verified reasoning systems could revolutionize not just math but also verified code and hardware verification.
  3. The potential for AI to achieve significant breakthroughs in mathematics, possibly earning accolades like the Fields Medal, is explored.
  1. Cultural Insights and Leadership
  2. Carina shares her journey from a competitive math background in China to founding a startup focused on AI.
  3. She emphasizes the importance of building a strong, collaborative team culture and the challenges of transitioning from a researcher to a CEO.

---

Key Takeaways

  • Verification as a Core Component: Axiom's unique approach emphasizes the importance of formal verification to ensure the reliability of AI outputs.
  • Interdisciplinary Applications: The principles of mathematical reasoning can extend beyond pure mathematics to fields like software and hardware verification, enhancing their reliability.
  • Cultural Dynamics in Startup Success: Carina highlights the importance of fostering an inclusive and intellectually stimulating company culture to drive innovation.

---

Conclusion This episode of The MAD Podcast showcases Carina Hong's visionary approach in the intersection of AI and mathematics. Her insights into the necessity of verification in AI systems paint a promising picture for the future of computational reasoning and its potential applications across various fields.

For listeners interested in the evolving landscape of AI, mathematics, and their integration, this episode provides a compelling narrative of innovative thinking and groundbreaking advancements.

---

*For more insights, consider subscribing to The MAD Podcast for future episodes featuring leaders in the AI and data landscape.*

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Need for Verification in AI

0:00 to 0:39

Explore the importance of verification in AI and mathematical reasoning.

“The world is realizing that we need verification.”

Why We Need an AI Mathematician

1:18 to 2:05

Karina explains the potential impact of an AI mathematician on solving problems.

“Please enjoy this great high-signal conversation with Kuro Nahong.”

Axes of Mathematical Difficulty

2:05 to 2:50

Discussion on the different axes of difficulty in mathematics.

“Yeah, I think people generally start with competitive math because it's kind of, you know that you have like a non-solution.”

The Putnam Exam Success

2:50 to 3:41

Karina shares Axiom's experience tackling the Putnam math exam.

“You guys are ultimately a young startup, but you've already had incredible success.”

Solving Open Research Conjectures

3:41 to 5:10

Axiom's achievements in solving significant research conjectures autonomously.

“And we all kind of gathered at the Axiom office and decided to put Axiom Prover in the real-time test.”

Intuition and Breakthroughs in Mathematics

5:10 to 6:06

Exploring the blend of intuition and systematic approaches in mathematics.

“One is there are like some really important mathematical breakthroughs that happen like, you know, once in a decade or maybe perhaps more once in a decade for each domain probably.”

The Role of AI in Solving Problems

6:06 to 6:58

Discussing the implications of AI's unique problem-solving methods.

“So a little bit of intuition, a little bit of sort of just kind of like one step at a time.”

Understanding the Lean Programming Language

6:58 to 10:51

Karina explains the Lean programming language and its applications in math.

“And is the system solving problems in a predictable way?”

Comparing Approaches in AI for Math

10:51 to 14:00

Differences between Axiom's approach and that of other AI players.

“People may have heard of OpenAI and Google DeepMind sort of winning IMO, the International Math Olympiad, and other very hard to crack kind of like math problems.”

Exploring the Combinatorics Problem

14:00 to 15:00

Discussion on the unresolved combinatorics problem and the importance of different reasoning approaches.

“And in 2025, no one solved the one combinatorics problem either.”
Show all 34 chapters

Informal vs. Formal Reasoning in Math

15:00 to 17:10

Understanding the differences between informal reasoning in natural language and formal reasoning through code.

“I mean, it wouldn't be super readable to humans.”

Challenges of Translation in AI

17:10 to 18:50

Discussion on the difficulties of translating natural language reasoning into formal language and its implications for AI.

“I think that's, you know, we believe in doing things at big scale and internet scale data set of Lean.”

Importance of Verification in AI

18:50 to 20:30

Exploring the need for verification in AI and its relationship with logical reasoning in mathematics.

“you know he just basically guessed three questions correctly versus the rest of us need to like reason it through and like Jesus Christ he just put like a zero in there and then somehow that answer is indeed zero.”

Real-World Applications and Challenges

20:30 to 22:10

Discussing the practical implications of AI in mathematical reasoning and potential challenges faced.

“So completely solving the hallucination problem or the stochastic issue.”

Growing Up as a Competitive Math Kid

22:10 to 24:10

Sharing personal experiences of growing up in a competitive math environment and its impacts.

“But I think in a lot of the fields such as from math and code verification and sort of verification applied to many different domains, it's incredibly valuable.”

Experiences at the Ross Math Program

24:10 to 28:00

Insights from attending a prestigious math camp and its influence on mathematical understanding.

“Because you're also a year old company, not nine, ten months old?”

The Role of Action Prover in Mathematical Proofs

28:00 to 29:31

Learn how Action Prover learns and improves its capability to prove mathematical theorems.

“So we're given like about 25, 30 problem sets.”

Academic Journey and the Rhodes Scholarship

29:32 to 31:56

Explore Carina Hong's academic path, including her choices in neuroscience and law.

“Then you went to the UK on a Rhodes scholarship to study neuroscience.”

Interdisciplinary Studies: Math and Law at Stanford

31:57 to 33:51

Discover how Carina integrated her interests in math and law during her PhD and JD studies.

“I think there's a lot of things, actually.”

Unpacking the Components of Action Prover

33:52 to 36:51

Get insights into the architecture of Action Prover and its components for mathematical reasoning.

“So let's actually go into the product now.”

Challenges and Innovations in Lean Theorem Proving

36:52 to 40:27

Understand the challenges faced in Lean theorem proving and the innovations being developed.

“It's interesting because there are a lot of sort of grassroots effort from the open source community to try to provide infrastructure tooling for Lean theorem proving.”

Current Focus and Future Directions for Action Prover

40:28 to 42:00

Learn about the current research focus and future directions for the Action Prover team.

Current Research Challenges in Mathematics

42:00 to 43:12

Explore the ongoing math problems and research focus at Axiom.

“really great mathematicians telling us how we should think about certain research problem targets And we currently have really hard research math problems in-house that we are tackling.”

The Complexity of Mathematical Problems

43:12 to 44:02

Understand the depth of current research problems and their implications.

“Roughly by, I mean, journal submissions, a lot of other factors, obviously.”

Scaling Complexity in AI Models

44:02 to 45:50

Discover how AI models are evolving to tackle more complex math problems.

“So on the easy end of the Pundum problem, we have 40 nodes.”

Axiom Prover's Impact on Mathematics

45:50 to 47:36

Learn about Axiom Prover's ambition to solve longstanding mathematical problems.

“They are not solved, not because they are like or they're not auto formalized.”

Potential of AI in Scientific Discovery

47:36 to 49:33

Examine the potential for AI to lead groundbreaking scientific discoveries.

“I mean, obviously it's doing some, but like in terms of humanity altering kind of groundbreaking discovery.”

The Interconnection of Math and Code

49:33 to 52:15

Understand how mathematical reasoning can enhance programming and code verification.

“They conjecture this based on their observed or like physics like phenomenon that I know, frankly, not very much about.”

Emotional Responses to AI Developments

52:15 to 54:02

Discuss the emotional reactions of mathematicians to AI advancements in the field.

“okay so through all my good friends telling me about Lean telling me about Howard Correspondence I believe math is code he believes code is math what does that mean?”

The Future of Verification in Mathematics

54:02 to 56:01

Explore the challenges and acceptance of AI-generated proofs in the math community.

“The parity of differentials for services of genius zero and one by algebraic geometry paper.”

Verification Challenges in AI and Mathematics

56:01 to 57:26

Explore the difficulties of verifying AI-generated mathematical proofs.

“So I think a lot of the like adverse reaction about from the math community about AI is actually coming from the fact that they cannot verify an informal solution.”

Transitioning from Math to Leadership

57:27 to 1:01:49

Learn about the challenges and insights of transitioning from a math background to a CEO role.

“just one sort of final chapter in this conversation.”

Building an Exceptional Team Culture

1:01:50 to 1:03:08

Discover how to attract and retain top talent in a collaborative environment.

“but it was those conversations that basically made me realize I have to do this.”

The Frontier of AI: Discovery and Verification

1:03:09 to 1:03:26

Understand the future possibilities of AI in generating and verifying knowledge.

“We still feel like we cannot fully elaborate and emphasize the thing that we are seeing that is the next frontier of AI.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Carina Hong:I think that we are at a threshold of mathematical renaissance, which is to realize that there are so many unsolved problems that will currently take, say, researchers months to crack that we believe is not sort of out of reach for today's AI technology. The world is realizing that we need verification. Like, you know, we can have very cool, like, lovable websites. But, like, how do I have a web code nuclear reactor? Our world view is math reasoning is a true reasoning layer of AGI. From math, you kind of get to code. And from code, you can run a lot of real-world experiments in the software stack.

0:32Carina Hong:The beautiful vision is for all the theoretical problems to be resolved in a satisfactory way. Hi, I'm Matt Turk from FirstMark. Welcome to the Mad Podcast. Today, my guest is Karina Hong, founder and CEO of Axiom. Karina is a 24-year-old Rhodes Scholar who's had an incredible personal journey from competitive math Olympian in China to MIT, Oxford, and Stanford Law. She started Axiom less than a year ago to build an AI mathematician, a system designed to be 100 % correct, 100 % of the time, effectively solving the AI hallucination problem. We talked about how Axiom's AI aced the notoriously difficult Putnam math exam and then went on to autonomously solve four open research conjectures, why verification is becoming the missing layer in AI, how Axiom works under the hood, and what this could all mean for the future of AI.

1:18Please enjoy this great high-signal conversation with Kuro Nahong. Hey, Karina, welcome.

1:24Carina Hong:Hi, great to meet you. So you are building an AI mathematician. Why does the world need an AI mathematician? Yeah, so I think that the idea where you have like infinite number of mathematical reasoning agents going out to the industrial society to solve like all the theoretical problems, I think that's incredibly compelling. I think through solving math, we also realize that it can solve a lot of other problems, such as verification, such as optimization, etc. I think that math is great. And if you solve math, you can have physics, you can have a lot of logic and property, and you can extend to a lot of things.

2:01Carina Hong:The world needs more math. Great. And what kind of math are we talking about? Is that high school math? Is that competitive math? Or is that deep research math? Yeah, I think people generally start with competitive math because it's kind of, you know that you have like a non-solution. And then you kind of start hill climbing the infinite sort of like infinitely high mountain of math. There are actually two axes of difficulty. One is how creative the solution is. And the other one, roughly speaking, is how abstract the mathematical object is. So say a qualifying exam can be incredibly abstract, but the sort of creativity required to solve each problem might not be that high.

2:41Carina Hong:It might be very standard. On the other hand, IMO problem, while it's very sort of easy to understand even by high school students, not very abstract, but it's incredibly creative. You guys are ultimately a young startup, but you've already had incredible success. So let's talk about the Putnam last year first and maybe define what the Putnam is for people that may not be in the math world. Yeah, 100 percent. So we started out in July, mid-July. And so Putnam was December and we were like a four month startup. We were kind of like looking at Putnam as this like really hard math competition. and most of the people actually got zero.

3:21Carina Hong:So over 50 % of humans got zero. And I think over the 100-year history of Putnam, there's only five human-perfect scores. And you have six hours to do it and 12 questions? Is that how it works? That's right. You have 12 questions. You have three-hour sessions. You have morning and afternoon. So, yeah, that's the setup. I think it was a Saturday. It was December 6th. And we all kind of gathered at the Axiom office and decided to put Axiom Prover in the real-time test. So it's not a benchmark. It's, you know, we got the exam from the proctor of the PUNAM exam, and then we just basically throw it to the prover.

3:57Carina Hong:And we announced that we got a perfect score. So 12 out of 12. That's right. Eight within the time limit, and then 12 out of 12. And then the more recent challenge that you guys saw, I think you saw four challenges, maybe. Talk to that. That just happened, right? Yeah, that's right. So I think that because we have like a lot of mathematician friends and they all have a lot of really hard research conjectures. So, for example, Professor LG Fell, he has like four fell conjectures and this is the last one still standing. And he's a Israeli professor at the Technion University. And there are also Dawid Chen, a Boston College professor who's an algebraic geometer.

4:36Carina Hong:And he knows actually Professor Ken Ono, who's our founding mathematician for four years. and they recently met at this joint math meetings conference. So he also supplied a problem. So people start sending problems to us and then we just put the system into test. And recently, I think like a couple of weeks ago, we just announced that Axiom Prover solved these four research level open problems. And it's quite interesting because it's probably the first AI to solve a research conjecture completely end-to-end and self-verified. That means the output are fully verified, 100 % correct. And that was without human intervention?

5:10Carina Hong:Without human intervention, that's right. Okay. My very uneducated understanding of world-class level math is that a lot of the top mathematicians today in history ultimately operate through a combination of, you know, sheer IQ, deep knowledge and all the things, but a lot of intuition and a little bit of serendipity. Yeah. Where does that fit? Yeah. I think that kind of two parts. One is there are like some really important mathematical breakthroughs that happen like, you know, once in a decade or maybe perhaps more once in a decade for each domain probably. But those require like very, you know, interesting sort of eureka moments, deep intuitions and very sort of almost like lucky kind of moments.

5:55Carina Hong:And there are a lot of other questions that are sort of proficiently, routinely applying the standard bag of tricks. And I think that a lot of the research questions can be solved by a combination of both. So a little bit of intuition, a little bit of sort of just kind of like one step at a time. I think that we are at a threshold of mathematical renaissance, which is to realize that there are so many unsolved problems that will currently take, say, researchers months to crack or even technical lemmas in those really longstanding conjectures that we believe is not sort of out of reach for today's.

6:33Carina Hong:high technology, but only through, I think, very intricate design of system and hybrid kind of use of different methods. And that's what Axiom is trying to do. These are the first batch. We hope to have a lot more coming. We actually have a few more research conjectures that's being proven every week just by the supply of mathematicians from the world. And we try to put those problems into use. Fascinating. And is the system solving problems in a predictable way? Where I'm going with this is the whole Move 37 discussion where you find AI solving problems in an almost alien kind of way. Is that part of what you're doing?

7:16Is that what you're seeing? Or is that something that's coming up in the future?

7:19Carina Hong:It's interesting because there's, I think, like a hindsight problem. So it's like we didn't know how to do it. I definitely have no hope in solving those conjectures. and our founding mathematician professor, oh no, didn't know how to do it either. Okay, now we saw the Lean code, right? Like thousands of lines of Lean code and sort of read through it, understand it. And maybe we see some of the techniques as sort of standard, but I think the application of them and the combination of them is also not entirely, I think it's somewhat at the level of a junior math professor, say like, you know, a postdoc or a junior researcher.

7:58Carina Hong:obviously a lot of junior researchers do amazing work and they have their move 37 moments in some very long-standing open questions but there are a lot of sort of day-to-day of research that it feels like it's at that level there's also this question of like because it's solving the problem in limb which is a kind of machine language not a human natural language uh the proofs actually look quite different so we actually analyzed all 12 problem solutions of the putnam exam and we found that a lot of the solutions actually differ from the human solution. So because it is a, you know, lean based system, it is really good at sort of routine bookkeeping.

8:37Carina Hong:And it will actually choose a lot of the more mechanistic, you know, arguments over the ones that require like a clever, say one picture solution. And on the other hand, you know, there are a lot of these sort of case work that humans shy away from. It's just very, you know, easy for the machine. There might be like a slightly bit of what's difficult for humans versus AI are different. So you mentioned Lean a second ago, and I guess that's going to take us a little bit into how the product and the model work. Maybe for people, again, who are not in the math world, what is Lean? So Lean is a programming language for math proofs.

9:13Carina Hong:I think that's kind of the one line high level explanation. There's this kind of concept called Caryl Howard correspondence, which basically allows you to kind of code math up as computer programs and so lean is like similar to python but it also can serve as it's like self-verifying function so in the computer science analogy it's roughly both the c language and the gcc compiler so so two in one it's a formal language there are a lot of sort of others they're improving language before lean such as isabel such as um coq um now now called ROC, R-O-C-Q, and such as whole is actor. So there's this sort of like family of like formal languages and Lin is one of them.

9:56Carina Hong:And it's a very popular language. There are a lot of people, mathematicians around the world that use Lin that choose to code their proofs up in Lin and they can just run it and then they will see a check mark, which shows that it's a correctly, logically correct proof. And if there's an error message, maybe there's like a bug somewhere or there is some sort of syntax or, you know, type mismatch, just like any other programming language. The fun fact is you can actually use Lean as a functional programming language. You can write an autograd in Lean, for example. And that's very interesting because basically it allows you to both math and code at the same time.

10:36Carina Hong:So if you think about a security protocol, you can try to implement the code in Lean, but also prove its soundness in Ling. So it's a very flexible, adaptive language. Great. People may have heard of OpenAI and Google DeepMind sort of winning IMO, the International Math Olympiad, and other very hard to crack kind of like math problems. How does their approach differ from what it is that you guys are doing? Yeah, I think the sort of concept form of their improving actually existed, like automated and improving as a field existed before deep learning. So I think there are a lot of researchers, a lot of them in Europe in 20, say 2018.

11:21Carina Hong:And even before then we're doing sort of like automated reasoning without the LM component. And we actually have some of these people on our team, like they were the authors of ATP Boost. And it's a very interesting time. And in 2019, François Charton and Guilherme Lampeau, the co-founder of Mistral, they have a paper which tries to put transformer on sort of symbolic integration and realize that it can beat like computer algebra systems such as MATLAB or Mathematica. I think Ilya was actually the reviewer of the paper and he actually tweeted about it. There is this other field medalist, Tim Gower said it was either amusing or game-changing because it was an open review.

12:04Carina Hong:People don't know if it's correct or not and that was not amusing. That was the beginning of AI for math and Francois is now also at Axiom. This is like a long history of what people are trying to do with it. And I think Google, you know, started the alpha geometry effort in 2021. And that was very exciting effort. They realized that if you convert the figures and lines, triangles, circles, intersection points in the symbolic expression in the vector language, a specific domain specific language for geometry, Euclidean geometry, you can actually try to do those geometric problems like a lot easier than using machines.

12:43And that's very interesting

12:44Carina Hong:because it's just like drawback to my childhood time where I was doing Mass Olympia and I could never solve one Euclidean geometry problem. I don't know what's wrong with my brain. Like it's usually the easiest problem of every competition. So it's like, if you go to a mass competition, you don't solve the geometry one, then like obviously you can't solve the inequality one and obviously you cannot solve the number theory or the commentaries like you know holy grail like it's like that's like the one problem you must know how to solve and i don't know how to solve it and i remember my teacher the coach taught me how to do the complex coordinate one which is a very tedious way of converting everything that shows up on the figure into a complex coordinate and just basically manipulate those algebraic expression and through that i can solve it i will solve it a lot slower than other people but at least I will solve it.

13:30Carina Hong:But it's a very interesting philosophical point, which is you can convert like, you know, geometrical figures into algebraic expressions. And I think that's what they did. I mean, not exactly the human version, but alpha geometry. And then that led to alpha proof. In 2024, I think Google sort of silver medal in alpha and IMO missing the gold medal only by one point. That was my moment. That was my moment of IMO at least. And then they couldn't solve the two combinatorics problem. And in 2025, no one solved the one combinatorics problem either. So 2025, there's only one combinatorics problem. And I think that was kind of the history of things.

14:08Carina Hong:There were also other players in the field and also a lot of really great academia labs doing it. We kind of take the approach of, you know, we think that it's important for the system to be able to reason both informally and formally. and in a way sort of bridge across these different abstractions from high-level intuitions to low-level, more lean sort of formal checking. What does that mean, formally versus informally? Yeah, so informal is like, say, reason in natural language, in English. And mathematicians, quite fascinatingly, they have been doing reasoning in English for thousands of years.

14:41Carina Hong:I mean, they write formulas, exactly. But mostly they write arguments, proofs in English. And I think it's to us, math is code in a way. Mathematicians have been coding in English for centuries and thousands of years. The formal language means, you know, lean and, you know, the sort of output will be lean machine code. I mean, it wouldn't be super readable to humans. But I think there are two beautiful things going on here. One is first time this sort of formal proving kind of comes in to assist mathematicians, which is a traditionally informal reasoning subject. The second thing that's interesting is you can play the strength of both informal reasoning and formal reasoning.

15:23Carina Hong:So you're kind of like, you can bridge across these different sort of like level of abstractions. And auto-formalization, which is the sort of capability of converting the natural language reasoning to, say, the formal language. And that's harder than translation because it's different than, say, translating between two programming languages. You're translating something that cannot be verified, natural language, into something that can be. And that direction is obviously very challenging, but also very promising. There's also auto-informalization, which is kind of translated back from Lean to English.

16:01Carina Hong:That's easier than auto-formalization because most of the machines' AI have seen a lot more English than Lean. Just to make sure I understand, are we saying that the OpenAI and Google DeepMind approach would fall into the informal? Correct. Whereas you'd be falling into the formal. Google was doing, I think, formal until I think as of the previous year's IMO 2024, the alpha proof system was a formal system. So it's English and my words, not yours, but like more of a brute force kind of approach versus what you do is more neuro symbolic. Is that accurate? This is how I would describe it. I think first of all is that we are, you know, we're supportive of scaling.

16:40Carina Hong:We think scaling works in a lot of the scenarios. there's also this question of sample efficiency which is kind of you know how effective uh scaling is potentially um and i think for the sort of informal like way to solve mathematics require like vast amount of like training data uh you basically throw everything you can possibly find uh on the internet to it now my question of that is what if you also throw these vast amount of math text data to your AI, but you throw the Lean version of them as well. I think that's, you know, we believe in doing things at big scale and internet scale data set of Lean.

17:19Carina Hong:I think that in addition to the internet, you know, scale data set of math is going to be quite interesting. And I think that we shouldn't do pre-training. We shouldn't try to just only train from scratch. I think we're kind of focusing on post-training, reinforcement learning can potentially get us better performance gain. And to that exact point, like help us contrast and compare. So RLVR versus, which is reinforcement learning with verifiable rewards against what it is that you do. Is that, are those just completely different approaches? Because they both seem at the same thing, which is to basically get to perfection.

18:00Carina Hong:Yeah, I think the world is realizing that we need verification. Verification means very different things in math. I think in early 2025 or late 2024, it means the numerical answer associated with each problem. Now, the thing is, one is reward hacking. We have seen from, say, Frontier Math and other benchmark, which only compels a numerical answer, that it doesn't actually necessarily reflect the model's capability in the logical reasoning. so it's able to get to the answer without reasoning through it which is quite fascinating I mean there's always these I did math Olympia before and there's always this like classmate who's really good at guessing the answer I don't know like AIME which is this exam that all the answers are between 0000 to 999 like I remember there's one year where my friend told me that like you know he just basically guessed three questions correctly versus the rest of us need to like reason it through and like Jesus Christ he just put like a zero in there and then somehow that answer is indeed zero.

19:00Carina Hong:It sounds very unfair. And, you know, in a way in high school, the teacher will like ask you to show your work. So for a while, I think like verifiable reward means that final numerical like output. I think that like people are now realizing it doesn't scale to the sort of math AGI and however you define it. Most of the sort of adult mathematics, mathematical research are proof-based, require sort of step-by-step rigorous deduction based on logical reasoning. And a lot of them don't even have a numerical answer. Like a lot of the problems are proof that something exists, proof that something cannot exceed a certain value.

19:37Carina Hong:You know, very seldomly, I mean, beyond the math Olympiad kind of high schooler's context, you would have a math question where getting to the answer is end of it. It's a lot more difficult to get verification reward for the intermediate steps, right? And so if you want to have a reasoning engine that really truly masters at logic and mathematical reasoning, then you need to somehow get verifiable reward for the proof steps. Coding is great. Like, I mean, people have seen like RL and coding have incredible gain. And can we turn math into code? And Lean, which we just talked about, the Carl Howard correspondence, exactly turned proof into a computer program.

20:13Carina Hong:So that makes RL VR, like, you know, possible in our setup as well. And just to drive it home for people, in case that's not obvious by now, Now, what we're talking about is building an AI that is 100 % correct, 100 % of the time. So completely solving the hallucination problem or the stochastic issue. Make AI perfect, the perfect prover. How generalizable do you think the approach that you guys are using is? I think the one thing, if you talk about perfect AI, I think people's first reaction is, wow, that's really valuable. I think a lot of the, you know, different labs are trying to reduce a hallucination or increase sort of the accuracy through many, many different ways.

21:01Carina Hong:If you have a lot of the sort of industries where mistakes are extremely costly, that's a block to AI deployment if you don't have that sort of approval guarantee. And now that's the value of like, say, catching the edge cases. And there is this additional value of trusting that your edge case can be covered. So two additional layers of value to sort of, you know, reliable, consistently correct AI. In terms of like how general this is, I think we start with math. Our worldview is math reasoning is a true reasoning layer of AGI. And I think a lot of the labs share that view, labs across US, China, Europe.

21:34Carina Hong:And, you know, from math, you kind of get to code. Math gives you proof of property and code gives you output. output and property affiliated to it are two quite important parts of the digital world. And so from math, you go to code. And from code, you can run a lot of real world experiment in the software stack. Then you can have a lot of other things. We don't claim to be doing things that are in the physical world at all. We are obviously not doing things that are non-verifiable, say, just like, you know, sometimes mathematicians are stereotypically not the best sort of writer. we are not kind of building an AI that's really good at literature.

22:11Carina Hong:But I think in a lot of the fields such as from math and code verification and sort of verification applied to many different domains, it's incredibly valuable. And do you need a lean equivalent for each one of those domains as you expand? It's a very interesting question. I think so, as you can see even within math, right, sometimes creation of domain-specific language, like the vector language for Euclidean geometry, have its gains. There could be the case where in other domains, something that is not exactly the abstraction of Lean are the right sort of medium. But in that case, you know, you can sort of do code translation and you can kind of build out, you know, the sort of stack that's required to use your like Lean-based theorem proving engine.

22:56Carina Hong:I think the sort of gap between, say, for example, Lean and another strongly typed language like Rust is a lot closer than the gap between Lean and English. And I think that's a lot of the commercial value. I mean, if you can sort of reason in between informal and formal space, that I think is going to unlock a lot of the things beyond just the power of a formal theorem prover. Yeah. And do you have a sense for where that threshold is? So if you have math on one side and English on the other side, effectively with your approach, you're going to be able to cover kind of like all of science. And the second you start getting into non-scientific fields, then the approach doesn't work anymore or you don't know yet and you're about to explore.

23:45Carina Hong:I think we want to try to figure out like what are things that can be done in software. One is I think math and code, they really complement each other very well. I mean, there are a lot of great code generation company. We can provide approval guarantee and code verification. There are a lot of other like domains where just that sort of verified generation capability is incredibly valuable, like hardware. And then I think if you can have a lot of theory and you can have partners who are really good at real world testing, then that is you know ai for science and i think that's also incredibly promising i feel like this is a generational effort it's going to take like a long time we're going to see the dna of the company remains math we're going to see best first market maybe verification best second market i don't know what that is could be optimization a lot of the things i think are waiting to be explored but just the generation verification loop i think itself is going to have a large term and then And I think, you know, there are a lot of things that we're also learning together with the potential customers.

24:48Yes. Because you're also a year old company, not nine, ten months old?

Read the full transcript

24:52Carina Hong:Seven months old. Seven months old. Yeah. Amazing. Before we go further, you mentioned your background a couple of times in passing. And as I was prepping for this, it's just so fascinating. I want to spend a few minutes talking about it. You covered it in some other podcasts, but I think the story is just amazing. So taking it from the top. So you grew up in China and you were a competitive math kid. Yeah. Just tell us that story. Competitive math kid is such a great term. Because you can break in different ways. Competitive math kid. Competitive math kid. So which one was it? Both. Well, I think I like to win.

25:35So walk us through what is that experience and how formative was this?

25:41Carina Hong:It was extremely informative. I think years of math Olympiad training, you have one goal that is to score as high as possible on whatever that is next, math competition. You have people that are in the same sort of community circle that are also doing the same thing. You are friends with them. You are competitors with them. There are a lot of sort of like, you know, background reading, learning exercise you need to do to over prepare for every competition. I remember I did 75 exam papers to prepare for a competition that I didn't know if I will be selected for. And I didn't end up being selected for it.

26:18How old were you? I think it was like 14, 15.

26:21Carina Hong:I think I learned a few things. One is resilience. it's just like I think you get addicted to pain and suffering so that like the word resilient is like almost a paradox because it's like you like it it's I feel it's like a given I think um throughout that time I mean just the exam get keeps getting harder and the number of people competing keeps shrinking like elementary school I have like 1000 friends competing for you know the spots for the middle school then middle school is like 90 okay and then like high school 25 that means vast vast majority of your friends lost the opportunity to compete and that's a very interesting thing i think it does to a child but then i also like learned some like you know other side you know not just the math olympia when i was i think 14 15 i got into the ross math program which is one of i think the best like high school uh math camp in the states uh ohio state university um i think my first trip to the united states um was that summer it was like you know summer of like eternal joy like every day i will be learning cool research math like they taught us undergrad math and they asked us to like deduce everything from the grounds up like we were asked to prove zero times everything's zero um by like a limited number of axioms and that was very defining it felt very different from math competition it's not like how many people will win that award it's like there is a vast amount of math that you just have no idea about and you get to build it yourself, almost like one brick by another, right?

27:56Carina Hong:From the limited number of theorems you are provided, you prove new things. So we're given like about 25, 30 problem sets. And each of the problems that have problems that like are probably just bookwork theorems, and you would just learn it in college. But instead of presenting it as something that is like a given fact, it asks us to prove it. So our world of mathematical knowledge is constrained to how much we can prove. And that's actually what's going on right now with Action Prover. Like Action Prover has access, obviously, to a lot of the world's information. But because the lean data is so scarce, it's like a lot less than, say, the amount of code data out there.

28:35Carina Hong:It's only two-digit million number of tokens out there in the open world. Action Prover learns to prove things. And it kind of self-improved in a way where all the things that it proved got fed back into it. into a kind of a skill library can be like, apply for the next challenge. There's also this sort of like self-challenging conjecturing component that keeps giving it harder problems, just like my camp counselor gave me the problem set. And this is a very beautiful process. I think like without this sort of right order, sequential order of introduction, I wouldn't like kind of go this far in math or love math as much as I do.

29:14Carina Hong:I think the sort of curiosity and discovery is a basic human need and that definitely exists in like back then the tenured child and that being used to motivate and inspire mathematical learning I think that was a very beautiful process. Amazing and then on some other incredible things that you've done so you did MIT in three years I believe Then you went to the UK on a Rhodes scholarship to study neuroscience. Why neuroscience? Was it all part of a grand plan towards AI or was it just your interest to naturally carry you? Yeah, I think the Rhodes program like did a really good job and probably too good a job to encourage us to just like shift direction.

30:02Carina Hong:Like it has this sort of, I mean, broad sort of belief that you need a lot of disciplines and studies to help you become a global leader. That's what the Rhodes Scholarship is trying to sort of nurture. People who have a background in STEM, they will encourage you to go into liberal arts. I wasn't fully encouraged to go into liberal arts. So I picked something that's kind of in the middle, like neuroscience. Obviously, I think at Oxford, I have a lot of math friends. And so kind of still math was part of the equation. I was also trying to apply math in my neuroscience study, specifically maybe because I just am afraid of animal experiments.

30:42Carina Hong:Like I'm not gonna, I'm probably just gonna stay in, you know, data analysis, computational neuroscience. I was quite interested in topological data analysis and persistent homology. But later I realized, I think it was like two, three months after the school year started, I realized there was something called UCL Gatsby. UCL Gatsby is this like premium, like, you know, AI hub in London and Oxford to London is a short train. And there's so many world-class faculties there, like doing really cool research in say like theoretical machine learning, like, you know, analyzing the neurodynamics of stuff.

31:21Carina Hong:They're also doing like various other like applied AI research and some from like a cognitive science motivation, but really the core is AI. Then you topped all of this with a joint PhD in math and JD in law at Stanford. All of this has been fascinating, not just in terms of achievement, but in terms of range. And I'm just curious how you were able to do all of this and whether there's any lesson for anybody else. I mean, clearly there's an element of the chart IQ, but there must be something else. I think there's a lot of things, actually. My junior year, I mean, my last year at MIT, I'm like, I kind of grew up with a very sole focus and goal to do Mass Olympiad and then do mass research.

32:12Carina Hong:And the very next step is to probably go to grad school directly and probably not even do the Rhodes Scholarship. I'm not sure because a lot of the Rhodes Scholars are like, you know, politicians or aspiring lawyers, judges. I'm like, I just want to have some fun intellectually. At the end of the neuroscience, I was like, OK, I want to go to Stanford to start my math PhD. But also there's this incredible opportunity of Stanford Law School, one of the two first ranking law schools in the country. And they have really good IP professors who marry like AI and copyright law. They have really good Professor Mitchell Polinsky in law and economics, where you basically are doing differential equations, but you are like analyzing like deference and like retribution, the ratio of each sort of criminal law measure.

32:58Carina Hong:And there are a lot of like other things, cool methods to apply textualism to constitutional law that's very similar to looking up definitions in math textbooks. And so I was like, OK, that's very cool. So I did my I mean, the J.D. PhD is that you have to spend one full resident year in the law school. So that was my first year. I spent like one year being a diligent law student. I was even trying to apply for clerkships. It was a fascinating year and I learned so much. And then the second year, which is kind of a very interesting year where I was browsing all the AI for math research paper and realized that, wait a second, there's so many ideas.

33:33Carina Hong:I mean, from DraftSketch Proof to, I think, STP, Southlake's Air Improving. There's so many exciting papers. And I wish I just have the resources in industry to execute. And that was when, I think, very, very soon after, I just basically decided to do Axiom, fully focused on the company. Fascinating. Thank you for that. So let's actually go into the product now. So we alluded to some of this. Let's unpack how it actually works. So you mentioned there were three components. What is the architecture? What do those components do? So I think like our kind of very broad vision is that we are going to have a conjecture.

34:11Carina Hong:We're going to have a prover. And then there is knowledge base. In a way, your knowledge base is like, let me take a metaphor. Because this is like quite, I think, like in the niche area of like very like subfield of AI. But suppose you are like selling on an ocean, right? And like, where do you know where to go? and sort of your ship that basically decides where to navigate. That's your conjecture. And then you like, you know, sail and tour one direction and then you land at this island. Okay, well, like, do you know if you have been on this island before? You don't necessarily know. Basically, you need to look up your knowledge base.

34:46Carina Hong:You want to make sure that this is indeed uncharted territory. And then once you realize that it is uncharted territory, like, how do you know if it's like, say, India or West Indy, right? So is it, I don't know, is it going to have some rare metal? That's kind of where your prover starts coming in to basically prove this new conjecture that is not in the knowledge phase, that is mathematically correct and has merit. And so that's kind of the, and then there's all the formalization, which is the ability to reason across informal and formal space, kind of weaving all these. Great. And so the conjecture part, is it LLM-based or are you in a completely non-LLM world?

35:28Carina Hong:So we do post-training on, say, like, you know, open source, like, LLM. Okay, so there is an LLM. So how does that work then? What creates the conjecture? Is that front-based? Like, what goes into it? Yeah, I think I would say here, like, the conjecturing part is still the underdevelopment part. We have been, in the last, I think, seven months, very focused on Prover and also made a lot of progress on the knowledge base. So in the Pundam exam, you don't need to conjecture. You have 12 problems. They're incredibly hard. And they are basically tests for your Prover. So Action Prover tried on Pundam exam, got perfect score.

36:06Carina Hong:The underlying system is a system of ensemble of models. And there's also a set of deterministic tools. And also there's like a proprietary data set that's very large. So kind of a combination of these three things that led to that success specifically. I think that for the deterministic tooling, it's quite interesting because these are actually written for Lean in the language of Lean. So a bit like metaprogramming. And that's very interesting. We are actually going to release them on a public API on all these dozens of tools. very, very soon, beginning of March. Great. It's a big announcement today on the podcast.

36:51Carina Hong:I mean, it's releasing the infrastructure for mathematical reasoning. It's called Axiom Lean Engine Excel. It's interesting because there are a lot of sort of grassroots effort from the open source community to try to provide infrastructure tooling for Lean theorem proving. Because Lean is a really sort of, you know, it's a relatively new language and there are a lot of reasons why it could be a bit slow sometimes. There could also be, you know, things where if you assume like an axiom that's mathematically incorrect, like if you assume n plus n equals n, then you will be able to prove 2 plus 2 equals 2.

37:29Carina Hong:You don't want that, right? 2 plus 2 equals 4. So a lot of the sort of like verify, verify proof is actually, you know, one of our prover tools that's about to be released. And that's actually 100 times faster than the other counterparts that are the open source effort called Comparator. So a lot of them are hopefully going to make everyone prove more theorems than Lean. And in this architecture that you just described, compared again to what seems to be becoming the norm in other parts of AI, fundamentally this pre-training LLM plus post-training system, Is there a trade-off to your system in terms of, is it more or less compute intensive?

38:11Is it more or less fast or slow?

38:15Carina Hong:We had a little bit of a sort of cold start problem, right? I think the data is quite scarce. So while there are more than one trillion tokens of code, it's probably a lot less for Lean. So we had to basically take our bold database to generate a lot of Lean proprietary data. So that's one difficulty. And let's double click on that again. So how did you do that? So you created synthetic lean data? That's right. So it's interesting because, you know, when people talk about synthetic data generation in the unverified domain, you really don't know the quality, right? How do you know the synthetically generated financial advisor data is actually good?

38:56Carina Hong:Then they have human experts to try to label it and grade it. Here you have lean. So you know that your thing is correct, at least. And if you do good quality control on the statements, then you will have things that are off mathematical merit. And when we kind of take data bets, we use things like auto-formalization to convert existing math from informal language to formal language. We also do things kind of that are more formal system-inspired, such as repair, fuzzing, exactly, to make there a lot more synthetic variants of the existing formal data that we currently have. So the other difficulty, I think, is like Lean runs on CPU and then the sort of LM part runs on GPU.

39:38Carina Hong:So you have a little bit of CPU, GPU means engineering. Just like a very interesting effort where ideas are out there. You need a very strong industry strength engineering team to execute the many good ideas. Maybe some academia researchers have produced. Some of our researchers have produced this sector. in terms of kind of how compute it's not like horribly compute intensive like definitely not compared to pre-training I think that like data is a large part of it I think that like good sort of infra engineering is another part of it so as I was researching this there were some numbers on a Putnam question basis where there were millions of tokens give us a sense for like the order of value really really varies I think there are like you know there have been cases where a hard problem and penance takes like 1 million tokens and stuff but there are also like a lot of other things we could do we in-house have something that can shorten proof so for example you can shorten the proof significantly 20 times shorter it's like different levels of how you would like to count like how kind of bulky a proof is and then uh still bearing in mind that you're a very company what is the current state of the product are you mostly focused on mvp kind of product that can solve this amazing problem but that's not industrialized yet what what part is research versus what part is engineering and product so far we focus a lot more on i say so there's this sort of like you know team of really strong machine learning researchers and engineers and they're all like both researcher and engineer in one they're like really amazing um we have a lot of really good people from meta from google brain from anthropic etc and we just um we keep hiring you know more and more uh sort of frontier lab researchers um this part i think is like you know focused on developing the core capability of the system so we want to basically push the the goal post like forward right so from putting them perfect score that was four months in then two months later was the four research conjectures and then you know during this middle we also have tested something that is transfer learning from math to code verification so another evaluation only a community recognized benchmark we want to kind of try where we can get because we have really great mathematicians telling us how we should think about certain research problem targets And we currently have really hard research math problems in-house that we are tackling.

42:16Carina Hong:It's showing some problems. It's also obviously getting stuck. So there's this part. And this part is like, you know, the current focus of the company. Once we know where the frontier is, then we can try to say, okay, let's make it robust. Let's make it sort of production grade. So when 1 ,000, 1 million people hit it, you know, it doesn't break. But this part is kind of sort of an effort that's surrounding that. And sometimes people like jump between different tasks as well. Now we learn something about the applied use cases. We also have a lot of subject matter experts and we are hiring, we are hiring subject matter experts to join us to work on trip verification, to work on code verification.

42:58To just double click on something you just said. So is there a long list of just pure math challenges ahead, you know, for the uninitiated? Is there an Everest in math?

43:10Carina Hong:There is. There is. So what is it? Roughly by, I mean, journal submissions, a lot of other factors, obviously. But, I mean, you can think about currently the batch of papers Axiom Prover has autonomously proven and mathematicians have written. You can probably get into Journal of Number Theory, Journal of Algebra, like that level. Well, that's a very different question from Endless of Mathematics or JMS, invention needs. That's one big jump. I think, you know, to get to that sort of result requires a lot of pushing. And we are not pushing it to just chase the amazing feeling, which is quite amazing, of proving something that is like grand and open for a long time.

43:55Carina Hong:But also at the same time, we are basically teaching the model things that it could not do before, such as a more complex reasoning tree. So on the easy end of the Pundum problem, we have 40 nodes. On the hard end of research questions in-house, we currently have a research problem with thousands of nodes. So it's a much wider and much deeper tree. And we want to see, are we going to hit a limit or not? We currently are not seeing one. So we really want to basically scale the complexity of the reasoning of the problem. We want to make the AI be able to do library learning. That is, we have seen it actually quite promisingly, auto-formalized definitions, which is really, really hard.

44:38Carina Hong:So in math, you have theorems, proofs, lemmas, propositions. Basically, you have definitions and that's very hard to ground. So you want to be able to auto-formalize definitions. You want to be able to have the model system explore definition to still be relevant for the proof. To progress through this series of problems. So if this is not super GPU intensive and this is not super... Fairly. Yeah. And if you've built a way to create synthetic data that works, what is the fundamental bottleneck? Is that doing more of the same thing across more domains or is there an architecture evolution? They're scaled up and they're scaled out.

45:21Carina Hong:So we're currently scaling up in difficulty. We believe that is a defensible move. I believe we are currently doing certain things, interestingly, and are ahead of a curve in terms of how we get rid of running out of context, this kind of problems, how we scale learning from experiences, how we scale inference. I think we are doing that. And we are also scaling out in a way of both there are some math problems. They are not solved, not because they are like or they're not auto formalized. Actually, there are interesting things you can do. You can choose an existing math result and try to auto-formalize it, or you can choose an unsolved math problem and try to prove it.

46:02Carina Hong:Both are incredibly valuable. I think people talk about unsolved problems all the time, and there's a lot of value in actually picking good targets to try to auto-formalize it. A lot of these are unsolved or, you know, not completed because they are very complicated in terms of the sheer volume of that result. So that's kind of scaling out within the domain of mathematics. and then there's also scaling out from math to other domains such as code verification and hardware verification. Do you think that AI can win a Fields Medal? There is this friend who taught me a lot of things about math and he said that you don't celebrate when you win the Fields Medal, you celebrate when you get into the shortlist of Fields Medal.

46:48Carina Hong:So obviously there is only a finite number of awards And I think that we really want Accent Prover to be able to solve one longstanding problem in mathematics that you can objectively say, even though if it's an AI, you know, or double blind, whatever, that will be in the shortlist. And just to unpack that, why shortlist? Why is the... Because then there are reasons, you know, whether... Oh, because then it gets political? Well, not quite political. I mean, sometimes fairness, for example, if a certain domain just got, you know, it's, yeah. Yeah, but to the short list, that is sort of the objective standard.

47:25And to the broader creation that I guess we alluded to a little bit earlier in the conversation of just training brand new sort of groundbreaking science. Like you think AI is well on its way. I mean, obviously it's doing some, but like in terms of humanity altering kind of groundbreaking discovery.

47:47Carina Hong:The first couple, well, not the first couple of months, a couple of months before we actually start executing. It was incredibly exciting for me on an intellectual level. Like every day I have this sort of excitement. Like it's like I drink like six cups of coffee kind of excitement for months. And the main source of that excitement, which I would tell you actually about, you know, my colleague Shubo, a good friend, he's excitement of this in a bit. But my excitement is like, just like we're now realizing that we are at the threshold of mathematical renaissance, we could also be at the threshold of theoretical discoveries in science.

48:27Carina Hong:Massive, massive scientific discovery at the theory level. And I think what I mean by that is we have been in a very math poor world. The supply of outlier mathematical reasoning skill is so lacking that people are like in the scarcity mindset. Like you will like hear discussions of, oh, like this problem is so interesting. Unfortunately, I'm solving that problem. They should all be solved. everything that human mind can conjecture find interesting find tasteful to be solved by hopefully majority of them by action prover and then you have the question of high physicists when they talk to mathematicians generally they will have interesting opportunities for collaboration like i actually have this paper with professor kenono and and others xin tong jiang and Michael Mertens, which addresses like the elliptic expansion moonshine conjecture.

49:29Carina Hong:And that kind of stands from like, you know, three, I think three theoretical physicists. They conjecture this based on their observed or like physics like phenomenon that I know, frankly, not very much about. But I can solve the math part. They come to the conclusion that what they believe are a beautiful phenomenon that they find it worthy to formulate as a conjecture and publishes a paper have a proof because they know some mathematics. That doesn't seem right. Like, I think that the really, the beautiful vision is for all the theoretical problems, all the curiosity, all the lack of understanding to be resolved in a satisfactory way across all scientific subjects.

50:12Carina Hong:And beyond this, right, there are things that we still cannot solve that we will get a close form sort of, not a close form solution, We'll get a very like, you know, approximation as precise as possible. I mean, there's a lot of value, for example, to know, say, what is after the 1000 or 10 ,000 digit compared to what is after the third digit. The word is actually a lot of the times not diminishing return. The last mile carry a huge amount of value. Like in search, for example, if you cover some edge case, you likely win. You will be a market winner in, for example, writing, right? If you write just that extra bit or any sort of creative art, getting that extra mile correct or done has a lot of value, like optimization, you know, precision.

50:54Carina Hong:We can try a lot of these things as well. And then as math kind of helps with both first principle understanding and try and error, it just kind of this cycle. Like you have some first principle understanding, you try testing it, you try some try and error, and you maybe give some sort of, you know, risk bond, uncertainty, principle, robustness estimation. And then you go back to your first principles and then you go to your trial and error again. You have this sort of circle of discovery. And this is really not the end of it. So that's why I'm already very excited. I think this is going to be amazing.

51:28Carina Hong:Ideas can diffuse between different fields. A bit like you have like Abacus and now you have like trade and commerce. You have calculus integration. You have like thermodynamics, mechanics, industrial revolution. You have the Babbage engine, which is to calculate log tables faster. And, okay, well, you have the prototype of computer science. The rest is history. You have number theory. You have RSA. Like, you have all these kind of mathematical tools, you know, kind of open up new discoveries and new use cases and in turn demanding more mathematical tools. Beyond this cycle and this cycle marrying science, here comes code.

52:03Carina Hong:We haven't even talked about that. And that's why actually Shubo is very excited. so Shubhou, CTO of Axiom before that was like you know long term like meta veteran he was an IC director he believes in code is math okay so through all my good friends telling me about Lean telling me about Howard Correspondence I believe math is code he believes code is math what does that mean? okay so it means that you can try to fulfill the dream of Donald News literary programming have computer scientist programmers enjoy the luxury of mathematicians where they can reason in natural language and this is kind of starting to happen right web coding like front end right like you know we can have very cool like lovable websites but like how do i have web code nuclear reactor how do i web code control flow how do i web code complex systems that require like quite honestly superhuman hierarchical reasoning skill it's interesting because you're not in code alone You have code and you have math.

53:05Carina Hong:So you have, in addition to the flywheel we already seeing in the coding companies, an additional layer of flywheel of verified code. And sort of math starts to come in. And this kind of flywheel of data keeps compounding. You have actually two, even counting the science parts, three orders of flywheel. How far do you think we are from that world where we have armies of it? We have to execute like, you know, something every couple of months. Like we have to move extremely fast. Like there's so much to do. I think Axiom is a very, very young company and we are at like, you know, the very, very beginning tip.

53:42Carina Hong:And we are already, I personally feel some sort of shock and emotional response. And I know some of my mathematician friends, Scott Commoners is a Harvard microeconomics professor, also Morgan Price winner. We're good friends. We all have this sort of emotional response when Axiom Prover proved Vals conjecture, proves that almost all primes are partially regular, partial Vandiver conjecture, which is one part of the original Vandiver conjecture that has been open for 90 years. The parity of differentials for services of genius zero and one by algebraic geometry paper. We are really just leaping across a point.

54:20Carina Hong:I mean, Pundit marked, I think, the end of AI trying on Math Olympia. We are very glad that we got a perfect score. It's a really good period point. Pundum 2025 is by a lot of sort of experts grading harder than IMO 2025. So it's the hardest reward math Olympia test. And now we are leaping. We are leaping to research. And I think I'm going to have another similar emotional response if it really does solve one of those breakthrough mathematics problems. Interestingly, I think there are a lot of experts in domains that are currently overlooked. by AI development. So if you're a software engineer, you feel like, oh, web coding really changed and improved your quality of life in a meaningful way.

55:02Carina Hong:There are people who are in industries where because of lack of provable guarantees, they couldn't use AI. And there are, for example, AeroAstro, for example, like I think Defense. For example, there is no partial credit for a mostly verified GPU. It's all or nothing. Or mostly flying plane. Yeah, yeah, yeah. a mostly verified, formalized hypervisor. I think for these experts, they are currently kind of hand-holding a lot of the traditional goals. Their life has not been changed. And I want to see the AI for math movement that Axiom is hopefully leading and to transfer to these domains and to try to solve some of those problems as well.

55:47Speaking of emotional response, Is the entire math field super excited about math AI or do they feel like the rest of us possibly disinhibited and replaced by AI?

56:01Carina Hong:So I think a lot of the like adverse reaction about from the math community about AI is actually coming from the fact that they cannot verify an informal solution. So suppose like GPT generate a math proof a million lines. No one's going to check that. Like I'm not going to check that. Like, you know, like I know there are a lot of really great sort of data labeling, you know, services. They couldn't have that sort of level of expert to verify a proof of that link. It's just hard to do. On the other hand, if you have a formal certificate, like a stamp certifying that this is a correct link proof, I think people are a lot more receptive to it.

56:39Carina Hong:That's why actually a lot of mathematicians, especially almost all of the new school of mathematicians are accepting lean. what they want to do is to have human to formalize lean. Now, it's really fun for humans to do lean. We have a lot of the initial Maslip people here at Axiom, and it's a very fun team to see them formalize some statements. But when it comes to, say, like hundreds of thousands of lines of lean code, which is already not a hypothesis, like automated reasoning team at some big tech currently have 260 ,000 lines of theorem proving code, not lean to verify one component of a hypervisor for CPU utilization.

57:19Carina Hong:So that cannot, I mean, they actually did write it by human, but it's just not quite sustainable. So maybe just to close, just one sort of final chapter in this conversation. So while I was listening, I was thinking about how someone with such a deep background in math becomes a CEO and what that transition was like and whether there's any lesson for anyone considering making that jump or any founder out there. I remember that reading the anecdote of Hamilton that he writes down all his flaws and shortcomings every night and forcing himself to correct them. I know what my habits and flaws and shortcomings are.

58:05Carina Hong:I'm a very spontaneous person. I don't have the best ideas when it's planned, which I think sometimes that's like you know in the course of research you have like math research you have these kind of very interesting eureka moments I think so I try to overcome that in my day-to-day I try to do scheduling very intensely I try to surround myself with people who inspire me to execute faster and that's basically the entire team and I think it's a great honor to be working with them every single day I think like I go into the office and I look around and I hear the discussions and I'm like wow I'm like so I'm so lucky to be here uh and I think that's something that in this kind of talent market I think people like to work with people who value their intellectual judgment respect their voice and opinion and I think Axiom has this sort of very flat like non-hierarchical everyone's a member of technical staff or a mathematician, right?

59:06Carina Hong:If you're earlier founding plus that. This culture is really great for sort of, I think, the sort of open communication of ideas and debate of ideas. And that helps with the pace. It helps with the pace of iteration, helps with being more correct. A lot of things I'm learning very rapidly. I mean, it's been very interesting roller coaster, like seven months. When I was in math, I mean, there is this sort of idea where you have like taste and And sometimes these tastes can be interpreted as arrogance. And I learned that ego is a really, really bad thing. And we should just basically get rid of it.

59:42Do you hire people for taste? And if so, how do you evaluate it?

59:46Carina Hong:Yeah, we hire people for taste, but not for ego. And I think that's a very interesting kind of... How do you select people? How do you evaluate someone's taste? Is that their prior work? Right. So one is, basically, that means I have to learn every day because I need to have a certain ground, you know, a basic amount of taste and also the other people who are senior at the company need to have, research scientists need to have that sort of amount of taste and what makes us excited. And when we are excited, we just go after that person. We have had a lot of, I think, extraordinary hires from amazing and interesting backgrounds.

1:00:26Carina Hong:And we like, we really like this team. And I think that we raise the bar for hiring actually continuously. So we recently have an uptick in the number of people who would like to join us and our candidates that we're excited about. We take recruiting very, very seriously. And how did you secure that founding team initially? Because that's kind of ridiculous in terms of like caliber of just world-class mathematicians, world-class. How did that all come about? One of them was your mentor. So now technically you mentioned that you're very flat, but like technically working for you as CEO. How did that come about?

1:01:07Carina Hong:So quite interestingly, I think at the start of this, I realized that there is a movement. And this movement of AI for math is very much by and large in academia. Or actually people are hiding in labs, like secretly doing AI for math, where their day job is something else. So when I talked to each of these people, like it was a very, I think, mutually exciting feeling from both parties. And basically it's like two whales and just kind of realized that only they can communicate in that like frequency range. And that happened like multiple times. I got very inspired by this. Like, I think there were a lot of times at the beginning where fundraising was quite hard.

1:01:45Carina Hong:I think, I mean, I'm a nobody. No one should trust me with their large amount of money. And I think that was very difficult and challenging, but it was those conversations that basically made me realize I have to do this. I just have to. And like, you know, and I have to be the best deal for this team because my team deserves the best of the world. And those kind of like, you know, intellectual alignment were like a main theme. And the other thing I realized was that they found, say, maybe the other AI for math opportunities is not particularly attractive for one reason or another. Maybe they don't want to move geographically to London or to China or maybe, you know, it's just a different kind of vision.

1:02:20Carina Hong:So kind of gathering them was a relatively natural process. And then after that, I think when you have a bunch of really smart and nice people, you just attract other smart and nice people. And especially people who are adventurous, like rebellious. People who come from co-gen want to disrupt co-gen. People who come from math want to disrupt math. Disrupt and I guess like elevate as well. Because they do have a lot of affection still with that feud. So I think it's a very interesting, it's almost like a tribe. Axiom is like a tribe. And we have more and more people joining us. And sometimes I feel like both in those initial conversations and still now, and even after all these sort of like, you know, talking to the world about what we are doing, we still feel like secret keeper.

1:03:09We still feel like we cannot fully elaborate and emphasize

1:03:13Carina Hong:the thing that we are seeing that is the next frontier of AI. That is a generation and verification loop. that is the discovery of verified knowledge. Feels like a wonderful place to live in. This is all incredibly fun, compelling and inspiring. Thank you so much for spending time with us. Yeah, thanks so much. Yeah. Hi, it's Matt Turk again. Thanks for listening to this episode of the Matt Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing if you haven't already or leaving a positive review or comment on whichever platform you're watching this or listening to this episode from.

1:03:46This really helps us build a podcast and get great guests. Thanks and see you on the next episode.

From the publisher

What if AI didn’t just sound right — but could prove it? In this episode of the MAD Podcast, Matt Turck sits down with Carina Hong, a 24-year-old former math olympiad competitor and Rhodes Scholar, and the founder/CEO of Axiom Math, to unpack how AxiomProver earned a perfect 12/12 on the Putnam 2025 and why formal verification (via Lean) may be the missing layer for reliable reasoning. Carina argues we’re entering a “math renaissance” where verified reasoning systems can tackle problems that currently take researchers months — and potentially push beyond math into verified code, hardware, and high-stakes software. They go inside the “generation + verification” loop, what it means to build AI that can be trusted, and what this approach could unlock on the road to superintelligent reasoning.


(00:00) Intro

(01:25) Why the World Needs an AI Mathematician

(02:57) Scoring 12/12 on the World's Hardest Math Test (Putnam)

(04:05) The First AI to Solve Open Research Conjectures

(06:59) Does AI Solve Math in "Alien" Ways? (The Move 37 Effect)

(08:59) "Lean": The Programming Language of Proofs Explained

(10:51) How Axiom's Approach Differs from DeepMind & OpenAI

(16:06) Formal vs. Informal Reasoning (And Auto-Formalization)

(17:37) The AI "Reward Hacking" Problem

(20:18) Building an AI That is 100% Correct, 100% of the Time

(23:23) Beyond Math: Verified Code & Hardware Verification

(25:12) The Brutal Reality of Competitive Math Olympiads

(29:30) From Neuroscience to Stanford Law to Dropout Founder

(33:57) How Axiom Actually Works Under the Hood (The Architecture)

(37:51) The Secret to Generating Perfect Synthetic Data

(40:14) Tokens, Proof Length, and Inference Cost

(42:58) The "Everest" of Mathematics: Scaling Reasoning Trees

(46:32) Can an AI Win a Fields Medal?

(47:25) "Math Renaissance": What Changes if This Works

(55:47) How Mathematicians React to AI (And Why Proof Certificates Matter)

(57:30) Becoming a CEO: Dropping Ego and Building Culture

(1:00:42) Recruiting World-Class Talent & Building the Axiom "Tribe"

More from The MAD Podcast with Matt Turck

All 44 episodes
AI That Can Prove It’s Right: Verification as the Missing Layer in AI — Carina HongThe MAD Podcast with Matt Turck · 1 h 4 min
Listen in VO