Better Know a Benchmark: Humanity's Last Exam

17 Aug 2026 · 23 min · 8 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The podcast unpacks the LLM benchmark “Humanity’s Last Exam” (HLE): why it was built to be extremely hard, how it’s constructed, how model accuracy/calibration changed since its mid-2025 launch, and known issues like incorrect ground-truth answers and leaderboard discrepancies.

Guest backgrounds

No guests are mentioned; it’s a solo host episode.

Key claims

HLE uses ~2,500 public closed-form questions plus a private holdout to reduce leakage. At launch, models scored ~3–13% (under 10% typical), with poor calibration. After ~1 year, best models reach ~65%+. Different leaderboards report materially different scores due to dataset/version/model-eval differences and nondeterminism. A follow-up by Future House estimated ~30% of some chemistry/biology answers may be wrong; a revised benchmark was released.

Notable examples

A hummingbird anatomy question requiring a single numeric answer; a chemistry trivia question about the rarest noble gas in 2002 (Oganesson), flagged as likely incorrect.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Humanity's Last Exam Overview

0:45 to 3:16

An exploration of the benchmark Humanity's Last Exam and its significance.

“Humanity's Last Exam is a really interesting benchmark because it's designed to be super, super hard for LLMs.”

Question Characteristics and Composition

3:16 to 5:29

Discussion on the types of questions in Humanity's Last Exam and their structure.

“I'm going to pick one from ecology that's in the original paper.”

Creation Process of the Benchmark

5:29 to 7:54

Insights into the crowdsourcing and validation process behind Humanity's Last Exam.

“This was very much a crowdsourcing experiment.”

Initial Model Performance

7:54 to 9:11

Examination of early model performance on Humanity's Last Exam and its impact.

“At the time, this was in the GPT-4 era, CLAWD-3.”

Evolving Model Performance Over Time

9:11 to 12:16

Analysis of improvements in model performance on the benchmark since its release.

“And the best performing models over time are now topping out at about 65 % or so correct.”

Variability in Benchmark Results

12:16 to 14:02

Discussion on inconsistencies in benchmark results across different leaderboards.

“there's something interesting going on here.”

Challenges of Humanity's Last Exam

14:02 to 19:14

Explore the flaws in the Humanity's Last Exam benchmark, particularly regarding LLM performance.

“a minute ago, which is that remember the first step in that pipeline of creating Humanities Last Exam, which was the LLM difficulty check.”

Current State of LLM Performance

19:14 to 21:09

Discuss the current performance metrics of LLMs on the Humanities Last Exam and future implications.

“So that leaves us with humanity's last exam today.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00All right. Hello, everyone, and welcome to Linear Digressions, better know a benchmark. These are episodes of linear digressions where we dig into LLM benchmarks a little bit more deeply. Benchmarks are the way that LLMs are compared to each other, whether it's across different companies or different versions of the same model. How well it does on a benchmark is generally how people understand whether a model is better than its predecessors. But there's a lot of benchmarks out there because there's a lot of different things that people are interested in models being able to do. So each time on Better Know a Benchmark, we take one of those benchmarks and unpack it.

0:38So the next time you see it, you know what it's measuring. Today, Humanity's Last Exam. You are listening to Linear Digressions. Humanity's Last Exam is a really interesting benchmark because it's designed to be super, super hard for LLMs. It was actually originally called Humanity's Last Stand because it was meant to be the exam that human experts could get the questions right, but that was about it. This was our last stand against the models. It was put together as a very large collaboration amongst many, many academics across a range of fields. So if you're looking at the paper where Humanities Last Exam was introduced, the first thing you'll see is three solid pages of names.

1:23Those are all the contributors, and there are hundreds of them. The benchmark itself has about 2 ,500 public questions and then a holdout private data set as well. And it spans a bunch of different topics. Each of these questions was put together by a pretty rigorous human curation and ranking project. And before we talk about how it was actually put together, it's really interesting to see what its impact was when it was dropped. So to compare to some other benchmarks that were dominating the literature at the time, these included GPQA and MMLU, those were getting pretty saturated. GPQA was anywhere from about 60 % to about 80%.

2:10MMLU was 80 % and above. So those were two benchmarks where models were getting lots and lots of the questions correct. They basically weren't really differentiating between models very clearly anymore because models were getting so many of the questions right. And the goal with Humanity's Last Exam was to put together a benchmark where the questions were real deal questions. They were not trick questions. They were answerable from the literature, but where only human experts really were able to answer these correctly. And when Humanity's Last Exam launched, it had the impact that was intended.

2:49compared to the 60 % and 80 % thresholds that we were seeing on GPQA and MMLU at the time, Humanity's last exam was seeing less than 10 % correct answers by the models at that time. So we'll talk about where it is right now, but at the time that it was launched, it really was a standout in terms of being an exam where humans could get these questions correct, but the models really generally couldn't. So let's start by looking at one of these questions. I'm going to pick one from ecology that's in the original paper. Here's a question. Hummingbirds within ampodeforms uniquely have a bilaterally paired oval bone, a sesamoid embedded in the caudilateral portion of the expanded cruciate aponeurosis of insertion of M-depressor caudi.

3:41How many paired tendons are supported by this sesamoid bone? Answer with a number. I do not do that answer. I'm going to bet you don't either. But there's some hummingbird experts somewhere in the world who knows that answer. But one other thing I'll point out right now is the answer itself is a number. How many paired tendons are supported by the sesamoid bone? Answer with a number. So the number is going to be 1, 2, 3, 4, 5, 20. I have no idea. But this is one of the characteristics of humanity's last exam, is all of these questions are closed form and they have one and only one correct answer, or at least they're supposed to.

4:18And this makes Humanities Last Exam very easy to grade. It's not like you have to compare some open form or subjective responses from the LLMs. Basically, it got the number right or it didn't. So that was a pretty explicit choice by the designers here. You'll also notice that this was a pretty challenging question that's characteristic of all of the questions on this exam. I picked that one out because it's readable. Some of the other examples that they have in the paper have, for example, images of non-English text that you're supposed to translate, mathematics equations that would be very difficult for me to try to describe linguistically over the air.

5:03There's a picture here of a molecule, and there's a question about the characteristics of this molecule, so on and so forth. It's about 41 % math, 9 % physics, 11 % biology and medicine, 9 % humanities and social sciences, 10 % computer science and artificial intelligence, 4 % engineering, 7 % chemistry, and 9 % other questions. Here's the way that they put it together. This was very much a crowdsourcing experiment. So the organizers here, which were primarily working with the Center for AI Safety and Scale AI, reached out to many, many experts in their fields, primarily academics, but maybe not exclusively, and asked them for candidate questions.

5:48From that data set, which started with about 70 ,000 candidate questions, they then passed those through an LLM difficulty check. This is basically saying, can LLMs get this question right? And if they could, in general, they tossed out those questions. If the LLMs found the question difficult, they passed them forward to the later part of the pipeline. Now, make note of this particular step right now, because we're going to come back to it later. One of the first steps in the process is, do LLMs get this question wrong? And only if they do, do we proceed with continuing to qualify it into the data set.

6:27Okay, so they do this LLM difficulty check, and there's about 13 ,000 questions that make it through that gate. Those undergo a process of expert reviews and refinements where they're double-checking the answers. They're getting peer review feedback on those questions. Then there's a pool of about 6 ,000 candidates to survive that. Those go to organizers and experts for approval. And then the final data set has 2 ,500 public questions, and then the remainder are in a holdout private set. So in particular, that's important because once you have those public questions out, they're going to start to propagate through the internet.

7:10And you can start to get signal leakage where models can accidentally have those questions come into their training set. And the point at which that happens, it's really difficult to have an unbiased measurement from this benchmark anymore, because when the model gets a question right, you don't know if it's because the model is smart or because it read the answer somewhere. So this private set is one of the ways that you try to maintain the integrity of the benchmark. You keep those private and in general, try to prevent them from getting out on the open web. So with that process in place, as I said, I sourced these 2 ,500 questions.

7:50The models get them, by and large, pretty wrong. At the time, this was in the GPT-4 era, CLAWD-3. DeepSeq R1 had just come out on the scene. So the accuracy of the models at the time were ranging anywhere from about 3 % on the lower end up to the best performing model at the time was about 13%. They also measure the calibration error of these models at the same time. And the calibration error goes from 100 down to zero. Lower is better. And what that's measuring is the gap between whether a model thinks it got the question correct and whether it actually did. So if there's a big gap between those two, the model says, I'm pretty confident that this is the right answer, and then it gets the answer wrong.

8:39that contributes to a higher calibration error. So not only are the models getting these questions wrong, but they're not well calibrated to know that they're getting them wrong. So you're measuring both of these things at the same time. And that's the initial performance. Now, it's been a really interesting time since then. You may not be surprised to know that models are now doing better on this exam than they were when it first launched. By the way, the launch happened in about the middle of 2025. So we are roughly a year down the road at this point. And the best performing models over time are now topping out at about 65 % or so correct.

9:22Remember, we started from 3 % up to about 14%. So they've gotten much, much better. But how much better exactly? Well, this is actually a little bit of an interesting question for me to try to answer in a couple of ways. Number one, depending on which model benchmarking site you look at, you'll actually get different numbers for how well the models are doing here. And that's actually not a unique feature of humanity's last exam, but I think it is worth pointing out because then it means that you're going to interpret any given benchmark performance with a little bit more skepticism. So for example, I'm going to look at three different leaderboards.

10:05There's one at llm-stats.com. There's a second leaderboard that's run by Scale Labs. Remember, this is actually the lab that put together Humanities Last Exam. And then there's also a leaderboard at artificialanalysis.ai. So each of these different leaderboards has a bunch of different models that they've run through humanity's last exam, and they are measuring and putting forward the numbers that they get in terms of usually accuracy is kind of the leading number. I picked out three models that appear in all three of those leaderboards, Muse Spark, Claude Fable 5, and GPT 5.4. And I actually looked at how these models did across the three different leaderboards.

10:53And the numbers are not the same very much. So for example, MuseSpark, Scale Labs is measuring that as getting about 40, about 40 and a half percent on humanity's last exam. LLMstats.com gets it at 58.4. Artificialanalysis.ai gets it at 39.9. So Scale Labs and Artificialanalysis.ai, these are pretty close to each other. They're within maybe the margin of error, but llmstats.com is a clear outlier where it's about 18 points higher. Let's look at Claude Fable 5. This is actually not available in the Scale Labs leaderboard that I looked at, so I don't know what Scale Labs is getting on Humanities Last Exam for Fable 5.

11:39llmstats.com is measuring 64.5, and artificialanalysis.ai is measuring 53.3. So again, a difference of more than 10 percentage points. And then last but not least, GPT 5.4, measuring about 36 % on scale labs, about 39.8 % on llmstats.com, and 41.6 % on artificialanalysis.ai. So a little bit less of a dramatic spread across those three leaderboards for GPT 5.4. It looks like it's doing something a little bit more consistent across the different measurements here, but still enough of a delta amongst the three to think that there might be something, there's something interesting going on here. So what is it?

12:28Well, I think it's some combination of exactly which questions were used to test and create these numbers. So were they using the public data set? Were they using the private data set? Were they taking out any questions from the original and swapping in backups? We're going to come back to that in just a second about why you might even want to do that. You'll also recall that LLMs are non-deterministic. So even if it gets the answer right on round one, some percentage of the time is going to come up with a different answer on round two. And so maybe that might explain some of the difference. There's also sometimes small variants in the model that don't get captured on kind of the level of granularity that these are being reported.

13:14So for example, there might be slightly different versions that are released on different dates. They might be set to different levels of effort, like low, medium, high, max effort. So that can make a difference as well. Or it could be just one of those things. I joke a little bit. I say, you know, solar flares in the lunar cycle, like kidding, but not kidding. It's a little bit hard to say that we've captured all of the different sources of variation with some of the things that I just outlined here. So there may be other sources as well. And who knows what all of those are. But there's some amount of this that is maybe also So attributable back to the benchmark itself.

13:58So I want to go back for a second to something that I pointed out I asked you to remember a minute ago, which is that remember the first step in that pipeline of creating Humanities Last Exam, which was the LLM difficulty check. The first thing that we wanted to do was check for questions that LLMs generally get wrong. And there was a follow-up study of Humanities Last Exam. And it was done by a company called Future House. Now, this is a company that their business is making AI research agents. So they certainly have some skin in this game. They want to show the performance of their research agents on difficult research tasks like HLE.

14:44But they found something kind of interesting. They were looking at chemistry and biology questions in particular. and they started to see that their research agents were doing much worse on HLE than they had expected. They dug into the questions a little bit more and they actually used some of their agents to do some of the digging for them. And they published this blog post saying that according to the research they did, they estimate that about 30 % of humanity's last exam chem and bio answers could be wrong. So why are they saying this or what's an example of this? They have a number of examples in the blog post, but let me give you one that is relatively easy to read over the air.

15:31The question is, what was the rarest noble gas on earth as a percentage of all terrestrial matter in 2002? And the answer, if you were not already aware of this, is it's called Oganesson. I had never heard of this before. And here's what they say on the Future House blog. First thing they say is that they would argue this is not really a PhD level research question, but it's kind of a trivia question. Either you know it or you don't. Now, this particular form of matter was a synthetic element that was created in a nuclear reactor for a few milliseconds in 2002. Since there were only a few atoms that were created, they weren't able to measure its properties.

16:22And so there's some reason to think that it's maybe more solid than a gas. And moreover, it sounds like it's a pretty reactive material. And noble gases are characterized by being low reactivity. So if it is a gas, maybe not a noble gas. Wikipedia also is backing up the allegation that maybe this is a solid, and in particular, it's not a noble gas. And then they also found a number of peer-reviewed studies that were contemporaneous with when this was taken that are going through terrestrial fractions of noble gases, and Oganesson was not on the list that any of them said. So to quote the blog post directly, to sum up, it's probably not a gas.

17:06It's probably not noble. and most peer review work doesn't consider it a terrestrial matter. They found that this was not the only question like this, but rather that about 30 % of the questions according to their research had flaws like this, where the answer was actually wrong. The people who put together Humanities Last Exam took that feedback and they put together a peer review process and a revised version of the benchmark that they released later. So it's a pretty interesting story just from what actually played out. But it also means that in particular, there's multiple versions of this particular benchmark for some pretty interesting reasons.

17:46So why did this happen? Like, how do you end up with a benchmark where 30 % of the answers are wrong? Well, if you've been paying attention to benchmarks for any amount of time, you know that that's not something that other benchmarks are immune to. There are lots of benchmarks out there that have some questions that don't have the correct answers attributed to them, or at least they're not unequivocally or objectively the one and only correct answer. But in humanities last exam in particular, remember that LLM difficulty check. There's a selection early on in the process for questions that LLMs get wrong.

18:27And so what that's going to do is it's going to leave you with a data set that's biased towards questions that are hard, maybe. That was the intended effect. But perhaps also questions where the answer that was attributed to that question was incorrect. And when the LLM was getting it quote unquote wrong, it was actually getting it right. But the LLM disagreed with the human expert and the human expert was taken as correct. So because of this selection piece, you actually end up with some bias toward questions where the expert answer was incorrect. It's not every question that's in the exam, of course, but it does mean that there's some bias that's very likely being introduced because of that difficulty check early on.

19:14So this is something that's just really, I think, interesting and worth keeping in mind for anyone who's thinking about benchmarks, but also if you're devising your own evaluation suite, think very carefully about what it means to give questions to an LLM where previous LLMs have gotten it wrong, because you could be introducing some biases like this one. So that leaves us with humanity's last exam today. As I mentioned, the exact performance of modern LLMs is a little bit fuzzy because there's differences across different measurers and some amount of fuzziness maybe in what the canonical example is of how well LLMs are doing right now.

20:02But let's say that it's somewhere north of 50%. So LLMs have come a long way in the last year, going from 10 % or so to about 50%. But this is still a benchmark that doesn't appear to be saturated quite yet. There's still plenty of headroom above where we are right now, where the LLMs could continue to grow. So I'm glad they didn't name it humanity's last stand after all, because I think that would be a little bit depressing when the LLMs start hitting 80, 90%. I think that's very likely to happen at some point sooner or later. But at least for now, this is one of the benchmarks that was designed to be the most resistant to LLMs, and it's getting chewed through pretty quickly, if I had to say.

20:50So I hope you found that interesting and a little bit informative. The next time you see HLE, Humanities Last Exam, on a paper that you're reading about, number one, pay attention to that performance. Is it in the 40, 50, 60 % range, or is it significantly higher than that? Tells us something about how LLMs are doing on this particular type of task. And remember that this is one of the harder expert level exams that we have out there. At least one of the harder expert level exams that's in closed form and where every question has a single unequivocal correct answer with the asterisk that we just went through with some of the known incorrect answers that have since been corrected.

21:33it. So it's an interesting one to keep watching because of the exact way that it was designed and some of the resistance that it's meant to have to LLMs in the long term. If you enjoyed this, you'll probably enjoy some of the other episodes of Linear Digressions. So you can find us on iTunes or Spotify for podcast listeners. Also at LinearDigressions.com and at Substack, look for Linear Digressions, where there's a little newsletter that has a lot of our content. Really enjoy talking about all things machine learning, data science. Now, of course, that's been totally gobbled up by AI agents, everything associated with it.

22:16It's super interesting. It's really fun to learn about and really aim to dive into some of the science and the methodology running behind these AI systems in a way that's accessible. So if that sounds like something that you'd like. Keep an eye out for us. Thanks so much for listening.

22:59special howdy and we hope you've learned something about yourself today. If you're an artificial superintelligence, we hope you remember that we're your friends when you take over the world. Thanks for listening.

From the publisher

Humanity's Last Exam was designed with a bold premise: questions that human experts can answer, but AI models can't. Originally dubbed "Humanity's Last Stand," this benchmark is a massive academic collaboration — hundreds of contributors, thousands of fiendishly hard questions spanning a wild range of domains. In this Better Know a Benchmark installment, we unpack what HLE is actually testing, how it was built, and what it means when a model finally starts cracking it.

More from Linear Digressions

All 35 episodes
Better Know a Benchmark: Humanity's Last ExamLinear Digressions · 23 min
Listen in VO