Axiom’s Carina Hong: Solving Math’s Hardest Problems With AI, And AI's Problems With Math

2 Apr 2026 · 39 min · 21 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Axiom Math CEO Karina Hong explains “verified AI” for math and code—using Lean-based formal proofs to solve hard problems reliably, avoiding “Schrodinger’s superintelligence” from unverifiable LLM outputs.

Guest background

Karina Hong is Axiom’s founder/CEO; she won top US undergraduate math prizes at MIT, was a Rhodes Scholar at Oxford, did law and a PhD (then left Stanford), and shifted from math toward AI-for-math after exploring theory/practice and meeting AI-for-math researchers.

Key claims

Math and code are “twins” (math is code; code is math). Verification is not just error-catching; it enables stronger “superintelligence” via formal properties. Formal systems provide verifiable signals that physical AI lacks.

Notable examples

Axiom’s Putnam (Math Arena) result: 120/120 perfect score, real-time competition; Lean metaprogramming tools (14) to formalize the exam quickly; code-verification benchmark transfer learning claim: Axiom Prover ~99% vs DeepSeek Prover ~11%.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Pursuit of Truth

0:00 to 0:17

Curiosity drives the quest for truth in understanding complex concepts.

“I don't want the perception or the Schrodinger superintelligence.”

Axiom Math's Rapid Rise

0:49 to 1:54

Discussion on Axiom's impressive growth and valuation in just eight months.

“And you're less than a year old, I think even younger than Upstarts.”

Karina Hong's Academic Journey

1:54 to 4:02

Exploration of Karina's illustrious academic background including her achievements.

“And Axiom Math is an AI company in a way.”

The Mission of Axiom Math

4:02 to 6:15

Karina explains how Axiom uses AI to revolutionize mathematics.

“So verification actually to me is about the super intelligence part.”

The Importance of Verification

6:15 to 8:16

A discussion on the significance of verified AI in critical applications.

“And then you kind of run the reward experiments starting from your software stack, right?”

Math as the Foundation for AGI

8:16 to 10:40

Karina argues that understanding mathematics is key to advancing AGI.

“There's a theoretical bound, but you can do a lot.”

The Beauty of Mathematics

10:40 to 13:14

Exploration of what makes mathematics beautiful and its impact on curiosity.

“And another final kind of touch on why math is AGI.”

Challenges and Recognition at MIT

13:14 to 14:00

Karina reflects on her experiences and achievements while studying at MIT.

“are like every single country's like gold medalist.”

Understanding the Morgan Prize

14:00 to 15:36

Learn about the significance of winning the Morgan Prize in mathematics.

“Explain how hard it is to win this prize.”

Transitioning from Math to Neuroscience

15:36 to 17:26

Discover the journey from mathematics to neuroscience and AI research.

“Issuing an invoice, sharing routing numbers, and processing the wire with just a couple clicks.”
Show all 21 chapters

Pursuing a PhD and Law Degree

17:35 to 19:58

Explore the challenge of balancing a PhD in mathematics with a law degree.

“This is a couple of years ago at this point?”

The Intersection of AI and Mathematics

19:58 to 21:45

Learn about the growing influence of AI in mathematical research.

“And you had already co-published a few papers, right?”

Exploring AI for Math Research

21:45 to 24:10

Discover the research landscape and community in AI for mathematics.

“So at the time, actually, it felt very interesting because I would write an email to someone And it felt like two travelers travel a very far distance to meet each other.”

Starting Axiom and Its Vision

24:10 to 25:58

Understand the motivations behind starting Axiom and its research objectives.

“I do think it's very much research lab, like kind of like motivated at the beginning.”

Challenges in the PNM Exam

25:58 to 28:00

Hear about the unique challenges faced during the PNM exam.

“Pacific, which is the time that the Stanford Pundam takers will walk out of the room.”

Challenges in Scaling Lean for Math

28:00 to 29:28

Explore the difficulties and tools developed for scaling Lean in mathematical operations.

“I mean, all the objects have types and there's a homotopy type theory, dependency type theory behind it.”

Experiences with the Putnam Exam

29:28 to 30:56

Hear about the challenges faced during the Putnam exam and the team's evolving confidence.

“If you're like a frontier lab, you can just use it for free.”

Growth and Excitement in Axiom's Journey

30:56 to 32:28

Learn about the excitement and growth Axiom has experienced from investors and customer interest.

“Like 80 wouldn't necessarily make Putnam Fellow in a year where Putnam makes them easy.”

The Future of Code Verification

32:28 to 33:38

Discuss the implications of achieving advanced code verification and its diverse applications.

“And when you when you achieve that, what does the code verification at that level unlock for the world?”

Transitioning from Mathematician to Entrepreneur

33:38 to 36:28

Understand the challenges of moving from academia to entrepreneurship and learning on the job.

“so there's obviously iterative refinement.”

The Impact of Verified AI

36:28 to 37:48

Explore the value of verified AI and its potential to change how coding and verification are approached.

“And when you think about the real world impact for our audience in the months to come, where Where can Axiom be helping them in the world from an impact standpoint?”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Carina Hong:Curiosity is like, I want truth. I mean, I don't want a spike of truth. I don't want the perception or the Schrodinger superintelligence. Like, I would actually argue, for example, in a lot of the handling of large language models solving mass open problems, it's Schrodinger's superintelligence. Welcome to the Upstarts podcast, our weekly show where we talk to emerging startup founders about their upstart moment. Upstarts are challengers who punch above their weight to take on the status quo and improve the world, all while building a big business, too. I'm your host, Alex Conrad, founder and editor of Upstarts Media.

0:35I'm delighted to be joined today by the founder and CEO of Axiom Math, Karina Hong. Karina, thanks so much for joining us.

0:42Carina Hong:Thank you for having me. Great to see you. This podcast is brought to you by Mercury, banking redesigned from the ground up. You just raised$200 million at a$1.6 billion valuation for a startup in the math world. And you're less than a year old, I think even younger than Upstarts. Is that right? Yeah, yeah. We are. We started July last year. Wow, that's crazy. So a trajectory from Inception all the way to Unicorn in just eight months. Yeah, seven, eight months. Yeah, we're a young company. What exactly does Axiom do? I mean, you've had this crazy academic career. You know, reading the highlights made me feel really stupid.

1:17because you went to MIT where you won math prizes for the top undergraduate in math in the country. Then you did a Rhodes Scholarship at Oxford. Then you did a law degree and PhD and dropped out at Stanford. Ridiculous. Were you intending to be a mathematician or what was the goal with all of this?

1:36Carina Hong:I was definitely intending to be a mathematician. I think this startup thing is kind of a complete accident. I did a little bit of a calculation of how much I compute I would need if I want to do an AI mathematician. And it very quickly became clear that I have to do it in a for-profit setting. That's so interesting. And Axiom Math is an AI company in a way. It's a math company as well. What's the way you explain it to normal people like me? Yeah. So we built an AI that can do math at a super intelligent level. That means, you know, it should be able to crash all math Olympiades. It should be able to crack a few unsolved research problems.

2:15Carina Hong:We should be able to know kind of exactly what are the blockers to cracking a really, really hard millennium prize problem. And then from there, we want the outputs to be verified. So everything that we produce is actually the output you can trust it to be correct. So that's really the promise of the sort of verified AI or, you know, perfect prover. What does the math verification unlock in a business context? If you think about like machine critical areas in various parts of the industrial society, then there are certain cases where you cannot deploy the code unless you know that it is always, always, you know, satisfying certain property you want.

2:53Carina Hong:Now, you know, people do testing, but testing covers like finally a number of cases. And there are potentially, you know, edge cases that are missed by testing will be extremely expensive, sometimes a loss of human lives. Now, I think even more than the mission-critical areas where there are large enterprises, they have certain regulatory and compliance requirements. They will want their code to satisfy certain properties as well. And then I think you can extend from there, which is that if you do knowledge generation in a verified way, you can do knowledge generation better. So that part is actually, I think, quite beautiful in that it's not about erasing hallucination or catching mistakes about how to scale and compound superintelligence.

3:34Carina Hong:So this is like the part that I think a lot of people often miss. Like they think of verification or verified AI as like, oh, well, something's lousy, you go fix it. Well, actually, it's not, right? If you think about, take a math example, Ramanujan became a really powerful mathematician after he went to Cambridge, studied with Hardy and Littlewood to know how to do proofreading. So that kind of, you know, moved from his brilliant intuitions to something that he can land as rigorous results, that he becomes much powerful mathematician after that. So verification actually to me is about the super intelligence part.

4:06And so in the short term, you can use the math to keep AI from screwing up as much basically, right? Yeah. But then that can unlock a lot of other cooler things.

4:16Carina Hong:We believe that if you do the sort of informal-formal hybrid approach, you can, for example, get a perfect score on PUNNAM, which by Math Arena. Explain to our normies what PUNNAM is. So PUNNAM exam is the hardest undergraduate math test. So it's actually, this year is 2025. It's harder than the International Math Olympiad. So no AI was able to get PNAM exam like ever. We know that last year, so 2024, I heard rumors like other industry players like Google DeepMind tried and they didn't go anywhere. PNAM is known to be much harder in terms of getting a perfect score than say the IMO gold medal.

4:51Carina Hong:And we got a perfect score. We were the only AI that we competed in real time. And we were the first and probably the only AI that competed in real time and got that result. Math Arena, which is this nonprofit academia lab, they evaluated various large language models on this year's Putnam many months after. And they found DeepSeek to get the top score. And that's 103 over 120 versus we got a 120 out of 120. But why is it important that these companies, whether they're a big one like Google or a Chinese company like DeepSea, their creators, or a new startup like you guys, I can understand why from an ego standpoint, from an excitement standpoint, it's cool to get really good at this math.

5:34But why is it so important that your software be able to ace these tests? What does that do for you?

5:40Carina Hong:Yeah, so slowly getting to the point that I don't think we will be able to get a perfect score if it's not with the power of a formal system. which is quite different from all the large language models approach. To answer the question, we kind of share the sort of cultural mission statement that I think DeepSeek also has, which is that math is AGI. So saying math is AGI feels like a provocative statement. I feel like maybe some of your peers in the AI world might find that a hot take or maybe not everyone would agree. What exactly do you mean by that? And where would maybe someone say, Karina, you're crazy?

6:14Carina Hong:Yeah, I mean, a few things. My work wheel is we start with math. You go to code, software. Math and code is twins. Math is code and code is math. We're going to get to that in a bit. And then you kind of run the reward experiments starting from your software stack, right? So you go to physical work. You go to physical AI. And then you go to everything else. So this is like the order. And then the two domains that you can get verifiable signals are math and code as a twin. Okay. And physical AI. Like I drop an egg like gravity. Like it gives me a reward. Yeah. But there's no other thing to have that property.

6:50Carina Hong:And software controls the physical one. So then, okay, I'm going to get to math and code, and code is math. So this is a line that is, I think, in Shubo's, one of his keynotes, the last slide. Mathematicians have been writing proofs, almost like coding for thousands of years. They're coding in English. What I mean is all the lines of the proof have meanings, and they are hooked together in a perfect way. There is a rigorous deduction process that's actually going on and there's objective truth and not truth. And in a way, proof is not computation, but it's close. So computation gives you output and proof gives you property.

7:33Carina Hong:And the output is meaningless if it's not associated with the property. So a lot of the, I think, people working in the formal verification, program verification space actually share the belief that, you know, instead of vibe coding, we need very coding, very coding, V-E-R-I and coding. The dream is that you can generate code and verify the formal verifier, like checker running the back the same time. And what that does is that creates a much more advanced, trustworthy system. When you say math is AGI is because that... No code review, okay? No extensive testing, right? There's a theoretical limit to that.

8:20Carina Hong:You cannot verify all code. There's a theoretical bound, but you can do a lot. You can do a majority of the code that matters. I have seen actually a demo of that. After I saw it, I cannot see it. This is what I mean by it's a code. There's also, I think, a question of whether we will still need mathematicians after, say, you know. What's your take? Do your math careers root for you? I'm using that to show a point, which is that curiosity is a fundamental human need. Even if you have action proof, we're basically generating all the math proof. We're not there yet, but if that day comes, mathematicians will still like to understand it, to auto-informalize from lean to natural language or do informal summarizer to understand it.

9:07Carina Hong:And then from that, they will be able to conjecture more. It's a never-ending cycle in flywheel. And because curiosity is a basic human need, I could also actually argue that reliability or the sort of requirement for consistently correct or trustworthy is actually the flip side of curiosity. Because curiosity is like, I want truth. I mean, I don't want a spike of truth. I don't want the perception or the Schrodinger superintelligence. Like, I would actually argue, for example, in a lot of the handling of large language models solving mass open problems, it's Schrodinger's superintelligence. 2024, I think this guy Dan, who's the head of the Center of AI Safety, Dan Hendricks, I think his name, put on Twitter the pandemic sample of 2024 and, like, put, I think, open AIs, either 03 or 01 Pro was the hot model that was being evaluated.

10:04Carina Hong:And, like, the rollouts are on Twitter. People disagree whether it got it right or not. Like just really like superior mathematicians who were put them takers themselves could not agree on how to grade that. And I think the final consensus got zero out of 12. The open-air people claim it's 10 out of 12 or 11 out of 12. I think no one pretty much believed that. Okay, now like this is Schrodinger's, like something cannot be both zero and 10. And curiosity being the fundamental human need and the pursuit of truth in the verified form being something that also speaks to that. I think that's quite beautiful.

10:40Carina Hong:And another final kind of touch on why math is AGI. Math is a sandbox, and especially lean, where you can try all your frontier techniques. We, I think, are ahead of the curve for a few of the things that people are doing in frontier labs because we have this perfect environment to try. So it's a better sandbox. It's a better sandbox. You can have faster iteration cycles. If you want to push for scientific discovery and scaling learning from experiences, your data trails, how to approach a really complicated mass problem can be that perfect environment. And I talk to a lot of people who are in AI for science, actually.

11:17Carina Hong:They have, like, physical lab. They have, like, they're working on, like, material science, biology sector. Lots to talk about because they are also seeing, because they have that real world of real words. They are seeing a little bit of that, but obviously with much larger spending because, I mean, we don't need robots doing experiments. When you were, you know, a student doing the Olympiad, thinking about going to MIT someday. Why was math exciting to you and what was the potential that you saw in math, you know, at a younger age? Math is like very interesting because in a way I think when you're a mathematician doing deductive proof, like, you know, step by step, like this follows from the last step and then kind of finally arriving at the goal you want to prove.

11:57Carina Hong:What you're doing are two things. One is you're doing knowledge generation. So even if that theorem has been proven before, when you are working out a proof, you're actually exploring uncharted territory in your intellectual boundary. And so there's this exercise of if you take a real analysis textbook, like Rudin, a lot of students will just cover the proof, use another piece of paper to cover the proof and re-derive all these very classical, beautiful results. So that's one thing you're doing. And do you think they're beautiful? I think they're beautiful, yeah. What made you... I'm one of the very few mathematicians who think analysis is beautiful.

12:32Was that always the case? I mean, you're obviously a brilliant mathematician, but were you always just really good at it and like, this looks cool, I love this, or what triggered that?

12:41Carina Hong:I always thought I was okay at it, but I'm always then like, the moment I thought I'm like kind of getting it, I got thrown into another more competitive environment. It's actually pretty bad in terms of like, you know, it's very depressing. You look around and everyone is like an international math olympiad medalist, and a lot of them have like three medals, and like I've never been. So that was like a very humbling moment at MIT. Like the first week I remember it was before the actual orientation was an international student orientation. And then the international student orientation are like every single country's like gold medalist.

13:17You ended up winning multiple big prizes at MIT and being awarded a Rhodes Scholarship. How did you reorient your focus and find the new challenges that you kind of kept moving once you achieved that goal?

13:28Carina Hong:Yeah, that's a great question. So when I got to MIT, I think when some other people probably feel like they're not the smartest kid for the first time, I have felt that like three times, like, you know, in the previous Olympia experience. I was actually coping OK. I was kind of like, well, OK, this is what I expected. I mean, I had the choice of whether I wanted to go to Stanford back then and then that would be less mass Olympia people. But I didn't want to do it. I still want to kind of learn from the most brilliant people and collaborate with them. And after I got the Morgan Prize, I do think actually that thing kicks in, which is like, well, if someone is like a nobody for the entire duration of their limited number of years of life, okay, suddenly you got a recognition.

14:12Explain how hard it is to win this prize. Yeah. For people who are not in math, it's a huge deal.

14:16Carina Hong:Yeah. So the Morgan Prize is awarded to one undergraduate student each year, I think around the world or North America, I think it's US, Canada, Mexico. It's just one person. Just one. Yeah. And it was you. Yeah. So after I got Morgan, I'm like, okay, well, I want to try something else. Like immediately that sort of the loss of focus started kicking in. Right. So I'm like, I was in a neuroscience program at Oxford. I was like really random. I was doing computational neuroscience. Really fun. Casual, you know, piece of cake compared to math, I'm sure. Well, no, I know nothing about biology. But you threw yourself into the deep end with this at Oxford anyway.

14:53Carina Hong:Right. And then there's obviously the UCL Gatsby Institute, which is a very vibrant sort of AI hub. I would take the train from Oxford to London, and I took many, many of those trips. So thanks to Professor Andrew Sachs, whose lab reimbursed those train trips. I was really poor as a grad student at Oxford. Like, a calculator is like, we cannot spend more than 23 pounds each day. Like, that's it. Like, that's your entire road step. And it's hard. It's hard to pay for three meals under 23 pounds, if you think about it. As a year-old startup, we experienced a lot of firsts at Upstarts Media. So when it was time to process our first international wire following a London event, I braced myself.

15:30This is going to be a painful lesson. Instead, Mercury made the process easy. Issuing an invoice, sharing routing numbers, and processing the wire with just a couple clicks. No phone calls, no paperwork, no learning curve. We're still figuring plenty of stuff out. When it comes to our banking, Mercury's already got it all figured out for us. Visit mercury.com to learn more and apply online in minutes. Mercury is a fintech company, not an FDIC-insured bank. Banking services provided through Choice Financial Group and Column NA, members FDIC. You're kicking, you know, whatever word we want to use in the academic world.

16:07And you say, OK, I'm going to go to Stanford now. I'm going to do a PhD and a law degree at the same time. Was that just a proof that you could?

16:15Carina Hong:I never had a liberal arts education. Like I was, MIT is a very big engineering school. and I was taking mostly only math classes, frankly. And it felt a little bit like, I mean, I went to Rhodes and then I saw some of my friends who went to Columbia and they start off their freshman year reading the big books. So from like Pluto to like Socrates. And like, I kind of wanted that. I wanted to read and write a little. I kind of felt like illiterate, you know what I mean? Like English is not my first language. My first semester at MIT, I was learning how to write an English email. like by the winter quarter or like the IAP, which is a winter recess period.

16:56Carina Hong:I called my mom. I'm like, mom, I know how to write an email in English. What was the end goal still? Was it, okay, I'm going to then go back and finish the PhD and then I'm going to become a great mathematician? Or what was the vision? The law school year, the end goal has nothing to do with math. Like I think it is in a way I did only one year of road. I didn't do the full two years. Okay. So I go to this coffee shop every weekend. and this is called Verve Coffee Roaster they have the best mantra advertisement for Verve I am a regular there and every weekend I will go and I will carry my big like books that's my like law school textbook because I have readings to do and it's like I cannot keep up with the reading speed that the law school demands because I read slower than my peers and so then I go and I read my intention is to read the books but then once I get there usually I will just open my laptop and check out the like papers there are a lot of like AI papers around the time.

17:51When would this have been?

17:53Carina Hong:This is throughout the law school year. Okay. This is a couple of years ago at this point? So it's 2023, 2024. 2024, okay. Yeah, because I think of the Gatsby kind of training, I was quite interested. There are naturally people I follow, their work, deep learning, theoretical deep learning. And then I think one day I met Shubo, and then he's like, theory of deep learning is useless. like you know get into practice which is what some of other my friends told me that around the same time a conversation i had i think that was quite influential was with greg young who was a um ex-co founder uh morgan price honorable mention 2018 is that how by the way people who win math prizes talk about each other they're like prize winner 2017 third place 2019 i completely co-emailed i mean that that's like i don't have other connection to him to him it might be just one conversation with this like you know random person for me it was very influential he said try to grab as much gpu as you can try to do like actual ai research and that was very interesting because he kind of had his like period of focused locked in on math he had this period of exploration where he went off being a dj and then he kind of came back to math briefly and then jumped into ai I mean like mu-p, like you know hyperparameter.

19:16Yeah, yeah.

19:17Carina Hong:And then obviously XAI. For me I was at a time where I'm like, okay, well I'm doing neuroscience, I'm doing law, you can definitely call that my exploration period. I'm not following the expectation of what someone might reasonably expect a math person to continue the path. And then I think that's why they say AGI, ASI is a cult. It's like once you see it, you cannot unsee it. the feeling that maybe AI can do math kind of starts kind of climbing into my head. And it is just that thought. My first reaction actually is who I can work with at Stanford for AI for math. And if I were able to find such a person, and if that person had a large amount of compute, I'm not sure I would jump that quickly.

20:00Carina Hong:I think I would still jump. And you had already co-published a few papers, right? I think you've published at least 10 at this point. Yeah, a lot of them actually under a supervision profile. Oh, no, it's quite interesting. I felt like I kind of wanted to explore a little bit. I do want to say that I think generally it's good to be a nobody for as long as you can. Why is that? Your learning curve is a lot steeper. So what was the jump moment to put aside the academics, start a company? So what I did was there was like hundreds of paper on the GitHub repo called AI for Math. And it's grouped into, actually I think they've done a really good job maintaining it.

20:41Carina Hong:It's grouped into different parts. There are like AI for mathematical discovery, so finding constructions and interesting examples. There's this one part about like, I think like hammers, like much to hammer, which is this sort of very niche concept within the formal language lean that we need to get into. there is this hammer that can help you handle the low-level computations or or or derivations I read all these papers abstract over like a very short amount of time like intensely so I was basically reading these people because I was like I don't again I don't have the prior for that so then I need to be able to see what the picture is like and I start to be able to like write on paper what I thought is the way to build such a thing.

21:28Carina Hong:And obviously in conversation with a lot of other amazing people. It's a little later on, but kind of gradually start to like talk to François Charton. Some initial conversations with other AI for math researchers were also very defining. Is this a tight-knit community? It is a tight-knit community now. It was not at the time. So at the time, actually, it felt very interesting because I would write an email to someone And it felt like two travelers travel a very far distance to meet each other. It sounds like if you're in this community, this is an obvious confluence of a bunch of people and ideas.

Read the full transcript

22:02Carina Hong:But they don't talk to each other. They don't talk to each other. Yeah, so that's interesting. And so for people who are not in this world, what made you excited when you started Axiom that you could do something new, meaningfully contribute with a startup? And obviously there are several venture-backed startups that say they're kind of doing a similar thing. I think in the startup space, there's Harmonic and Us. In the hyperscalar space, there's Google DeepMind, Large London Presence, ByteDance, SeedProver. Oh, that's an incredible team in China. DeepSeek has dissolved its formal team. Quen's in a mess.

22:40Carina Hong:Oh, I think there's like a lot of the very good talents are in Europe. So that's the landscape that we're seeing. So in the United States, probably there's Axiom, Harmonic, on the formal math side. Okay. And Harmonic, obviously, kind of co-founded by one of the Robin Hood creators. It's raised a lot of money itself, some very smart people. What was making you fire it up? Like, I can build something that will win here or do something really new and exciting? Those conversations. So we have the initial conversations. That's actually what we found was quite different. We have our research vision that is you cannot just work on proving.

23:26Carina Hong:I was telling everyone around the same time we were talking about we need auto-formalization and we need conjecturing and we need a knowledge graph. And this forms a self-improvement system. For the formal proving side, this is the self-improvement system. You also need discovery. So the constructing examples, counter-examples. So like DeepMind, they have Alpha Proof Team and Alpha Evolve Team, right? And we're the only sort of starter that have this proving team. And we also have the discovery team. We're actually releasing something next week. So quite exciting on the discovery side. So we think our vision is broader and we think we can do it right.

24:02Carina Hong:So it's hard to not sort of form a team ourself if you know exactly these are the people you want to invite on this incredible journey. It just felt like the right moment. I do think it's very much research lab, like kind of like motivated at the beginning. I mean, I was at a hedge fund for a while and I saw how math can make money. Okay. But that was, I never kind of thought that was a go-to market. I had this like feeling that it's going to be something that is a niche area that I'm not looking when I started it, not rather than quant trading. So, you know, the places that employ MassPhD are Wall Street hedge funds and like, you know, NSA cryptography, right?

24:42Carina Hong:Definitely see the cryptography one is not large enough. The pie is not large enough for it to be interesting venture bet by and large. You need something else. And I always thought there is something niche. And we're like always doing that sort of discovery. And we found that it's chip and code verification. Like it is very compelling. And then the sort of co-generation team that we have. We have really incredible researchers from Meta working on like compiler code gen. And we also have people who work on software testing. So Shubo, for example, his software testing work is used, deployed actually across Meta.

25:16Carina Hong:So all the automated generated unit tests stemming from the mutation-based unit test work shared by Shubo and a few other researchers there. There is this sort of dream of software verified generation that now feels like the people on our Axiom's team converge with the subject matter expert from AI for math. And it is a joint vision. And that feels very good. It feels quite serendipity is working. And it's a beautiful feeling. When we think about your upstart moment where you were most, you know, punching above your way or back to the wall, what has it been in the journey so far? PNM. The day of the Pundam exam, we were in the war room.

25:54Carina Hong:We got the exam from an official proctor who could only give it to us an hour and a half after the students got it. And we want to finish it by 4 p.m. Pacific, which is the time that the Stanford Pundam takers will walk out of the room. By 3.58, we have eight problems. So, for example, like in the previous gold medal announcements, like Harmonix gold medal is not in the time limit. So they announced the gold medal with extended time. Now, with extended time, we have a perfect score, but we want to log in exactly how we got within the time. Now, you can say that real-world clock time and compute clock time are very different things, but for us, we kind of just want to try.

26:36You had this goal in mind, yeah.

26:37Carina Hong:And the hard thing is lean is slow. So it's like, you know, even if you have your GPU and you have your CPU, right, and they're running the same time, and the CPU being, like, slow will drive down your sort of efficiency. Can you explain, for someone who hasn't seen this before, how does this literally work? So they hand you the test? They gave me, they sent us an email. Okay, so it's a digital test. Yes. And then you upload it to your system, basically? No, we need to formalize it in Lean. Okay. Right, so we need to formalize it in Lean to give it to the theorem program. Lean is a programming language.

27:10Yeah. It was used by mathematicians for years. Yeah. Okay.

27:13Carina Hong:Now we can take natural language. Got it. But at the time, when we were four months old, So July 15, we spent one month building the infrastructure. So we have these amazing Facebook engineers, and they're like, okay, suddenly we have nothing. Imagine how much tech stack Facebook has. We built that from the ground up. So really three months. It's computing the cloud that you're using, basically. Okay, and so you guys are on the clock. You're trying to get this done. Yeah, and Lean is slow. So we had sort of foreseen this issue, and we built a dozen, actually 14, metaprogramming tools. What that means is things that are written for Lean in Lean.

27:58Carina Hong:So instead of using Python to write a tool for Lean, we use Lean to do the tactic-level metaprogramming. It's extremely hard to do. Lean is a very finicky thing. I mean, all the objects have types and there's a homotopy type theory, dependency type theory behind it. So a lot of people are like scaling will never work for Lean because it's such a, you know, there are so many fundamental issues. We build those like, you know, 14 metaprogramming tools that will handle, for example, merging theorems, separating theorems, repair proof, all these stuff and verify proof, validate proof 100 times faster than what the FRO sort of like comparator tool, the gold standard.

28:39Carina Hong:is and all these tools like were on like we have a dashboard of how much they are used they're intense workout during the PNM exam day it was very exciting I mean it was incredibly exciting and so all these tools basically you know like made us not relying on ILM's call or like calling to the model there's no need for that to do those to do those things we just like handle it from pure engineering this is another thing I think Axiom is doing like, you know, significant lead compared to any other competitor. It's our engineering strength. And we actually released these 14 tools. So it's now free release last two weeks ago.

29:23Carina Hong:So anyone can use it. Anyone doing like large scale lean, you know, operations. If you're like a frontier lab, you can just use it for free. Okay. And so with Putnam, at what point were you like, we've got this? It's going to work. Yeah. So it didn't start off very well. We got the exam. And then we look at how many combinatorics problems there is, and we are like, we are screwed. And then we saw also there's a geometry problem. We don't have a geometry engine. And so we're like, it's definitely not that great. I remember there are a lot of quotes by the mathematicians who are kind of trying to eyeball whether the formalization is correct.

30:00Carina Hong:And they're kind of debating among themselves, well, is this an optimal formalization? And then I remember Professor Ono said something, this is a quote, I actually have a quote of the day from that very memorable day. He said, there's no room for purity. We are in a sports-like situation. It's really funny for a pure number series to say that. Yeah, so it didn't start off that well. And then very quickly, we got like four or five problems, which is our internal prediction market of how well we will do. So we're extremely happy. We're like, okay, this is great. Near the afternoon where everyone is already, like we're skipping lunch, obviously, like when everyone's a little bit more exhausted, yeah we got we got um we got the eighth problem that was that was hard the eighth problem came at 358 just under the wire yeah they didn't celebrate because then i have a problem like do i announce it or not because i thought we could get a ninth you know what i mean like i i just thought i don't i don't like it i mean eight a's would have placed a putnam fellow last year 2024 but only because 2024 is a hard putnam okay a score of 90 if we just get one more would have been the highest scoring individual of 2024 and would have made like Putnam Fellow in most of the years with a score of 90, right?

31:18Carina Hong:Like 80 wouldn't necessarily make Putnam Fellow in a year where Putnam makes them easy. I really want that 90. And like, so, but like, you know, time kind of just passed and we're like, okay, well, like, what do we do? So next morning I just did the right thing. I announced it on Twitter. Two hours later, we got the nine. Yeah. So that's interesting. That's exciting. Has the trajectory just been crazy since then? I mean, obviously, investors have given you a ton of money. Have customers been excited? Like, what have the last few months been like? Yeah, we're quite excited in talking to some companies or big companies or startups.

31:51Carina Hong:It's fascinating. I think it's like, this is a space where you need to look carefully enough. And once you do, you realize it's like five times larger than it is. Usually something is like, well, you look closer and then it's actually worse than what you thought it is. But this one is exactly the opposite. it. So every day I feel like I'm learning a lot and being educated by these, you know, industry practitioners who tell me why they will need this technology and how. I think there is like going to be a lot more, I think there's going to be a platinum moment for code verification. And that's what we are working toward.

32:27Okay, so that's the new goal is, okay. And when you when you achieve that, what does the code verification at that level unlock for the world? Like, why should people be rooting for you to be successful here?

32:39Carina Hong:I would imagine forward deployment things will be done. Basically, there are a lot of specific properties that people would like to verify. A bank needing it for compliance reasons will be very different from another sort of defense use case, for example. Yeah. So I'm really looking forward to get into the weeds of that. I think benchmark to reward is usually the sort of problem. The difference is in our case is the benchmark are quite hard to do, quite, quite hard to do. So we want to get there. We have seen actually Axiom Prover demonstrating excellent transfer learning capability. So same system, apply on the really community-recognized code verification benchmark.

33:21Carina Hong:Very now we've got 99%. 99%. That's pretty good. And DeepSeq Prover is at 11%. GoTo Prover, 11%, 12%. I think like the large language model they evaluated, the author evaluated, 3.6 % past one. and 22 % iterative. Now we have a system, so there's obviously iterative refinement. It's amazing. My last big question for you is, you've obviously been working with the members of the top of the math community for years, but now you're a first-time founder. Less than a year into your company, it's worth over a billion dollars. You have a team of at least 20 people. How are you leveling up and learning how to be the best entrepreneur and not just the best mathematician?

34:02That's a great question.

34:03Carina Hong:Everything is new, And so if your default is like, I don't really know what's the right playbook for this, then you are very curious and you're very low ego and you just want to learn. And I think that has worked quite well. I would say like the old contract negotiation, never done that before. Like I was at law school, but I was at litigation. I was not doing corporate work. But it's really fun because I'm like, OK, well, I'm on the other side and I'm learning what my law school classmates would have done as corporate lawyers. And that's interesting. I think being on podcast is like something that's completely new to me I didn't I'm not one of the podcast listeners when I you know before this and now I'm starting to listen to podcasts and learning this the writing skill I have picked up not fully from law school has helped a little bit in some blog writing creative content making video for the first time so so these these are interesting and these are fun and you like you like trying the new things yeah I try the new On the sort of internal end, realizing and being quite grateful that we have some of the best, best researchers who have also independently, you know, without the context of Axiom, been pursuing the same dream for years.

35:12Carina Hong:Like there are people who have been developing all the code verification benchmarks by themselves. Like there's formalized with tests, code with proof. There's one specific individual. So really leading visionary in the open source community, like meeting him and be able to have regular discussions with him is something I really am quite honored by. I guess that's a good point, too. It's like this is only a company that's less than a year old, but your team have been working on these problems for years. So from MathLib Initiative, Prime Number Theorem, and from my last theorem, Sphere Packing, so the leading formalization projects, Frontier Maths, Benchmark, CombiBench, Formal Conjectures, all these like ProofNet, ProofPiles, all these like sort of benchmark data, Lean Universe, all these projects on the sort of like hammers like Mug Schnuhammer or even on models like proof optimizer, compacting proof models, on the discovery and patent boost into it.

36:13Carina Hong:All these landmark AI for math projects, the core contributors are Axiom. So that's, when you ask like kind of how are we different from the competitor, we are an AI for math native team and we are a software generation and verification native team. And when you think about the real world impact for our audience in the months to come, where Where can Axiom be helping them in the world from an impact standpoint? Yeah, I think we want a world where vibe coding for complex systems and backend can be unlocked with the guarantee that you actually did it. So vibe coding without the uncertainty. In certain cases, you don't really need that guarantee.

36:58Carina Hong:But for more like enterprise, I think cases you do. Wouldn't it be nice if, for example, you know, Clockwork can continue to help you decompose the thing into finer and finer, fine-grained kind of components and at one point decide to call Axiom API, you know, and then we give you either a complaint saying that it's still too hard, you want to break that down further, or give you something that you know is verified. I think people are underestimating. So there are two values of a really fully verified and trustworthy AI output. One is the value of not missing or catching all the edge cases. The other layer of value is trusting and knowing that all the edge cases will be captured.

37:43Carina Hong:I think people are overlooking that second layer of value. Why do proof trackers like Lean have a place and gain its popularity? it's because the good feeling that you can instantly verify it and it can be sort of shipped to deployment in this case can be accepted into the it can be mathlib or your own GitHub repository and it also can, the tactics like grind can handle the low level computations for you and so it gains its popularity so I always want to use this as a thought exercise of like verified AI is really not about ensuring that's correct. It's also about having something that satisfies your curiosity and give you that reliability at the same time.

38:30Carina Hong:Verify knowledge generation. I think superintelligence is not that meaningful if it's not verified. Well, you're satisfying your own curiosity has led you to build this incredible fast-growing startup. So Karina, thank you so much for joining us on the show. Awesome.

38:54I'll see you again soon.

From the publisher

Just nine months ago, Morgan Prize-winning prodigy Carina Hong was still on an academic track, pursuing a joint law degree and math PhD at Stanford. Now, as the founder of startup Axiom Math, she runs one of the most promising challengers in a new, fast-paced category: AI for math.

Recently valued at $1.6 billion, Axiom has already solved some of math’s most challenging problems, and Hong hopes it can help researchers advance the field. Her bigger ambition? To power real-time math-based verification of AI-generated code, to do away with vibe-coded slop.

On this episode of The Upstarts Podcast, Hong shares her founder journey from immigrant at MIT to Oxford, and ultimately dropping out of Stanford; how she’s learning as a first-time founder to help Axiom compete in a red-hot new category; and her Upstart Moment when Axiom took the world’s hardest college-level math test.

Chapters
00:48 Intro to Carina Hong
2:02 What Axiom Math does
5:51 ‘Math is AGI’
8:48 Not replacing mathematicians
13:28 Winning the Morgan Prize
17:17 Origins of Axiom
22:22 The new math AI race
25:51 Carina’s Upstart Moment
31:47 Axiom’s business prospects
36:39 Why the future is verified coding

For more, visit https://www.upstartsmedia.com/

Season 1 of the Upstarts Podcast is presented by Mercury

Produced & edited by Eric Johnson from LightningPod

More from The Upstarts Podcast

All 23 episodes
Axiom’s Carina Hong: Solving Math’s Hardest Problems With AI, And AI's Problems With MathThe Upstarts Podcast · 39 min
Listen in VO