LLM As A Judge

28 Sep 2026 · 30 min · 10 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Using an LLM as a “judge” to evaluate other LLM outputs (and agent behavior) when human labeling is expensive. It covers the 2023 “LLM as a Judge” benchmark idea (MT-Bench, Chatbot Arena), how to run comparisons (pairwise A vs B, single-answer grading, reference-guided rubrics), and why results can be biased/noisy (positional bias, preference for its own family, longer answers, non-determinism, and overconfidence). It argues LLM judges are useful but require calibration against human inter-rater reliability.

Guest backgrounds

No guests are mentioned; the host presents the episode.

Key claims

LLM-judge agreement with humans is ~80–90% in clear cases, ~60–65% with ties/positional bias; small vs large judge choice depends on whether the task decomposes into checkable criteria; measurement noise can swamp small (e.g., 3%) improvements.

Notable examples

Fed buying bonds question; Assistant B gives better, non-repetitive daily-life examples. Chatbot Arena voting example comparing “Claude Opus 5” vs “Claude Fable 5” (A vs B).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Challenge of Labeled Data in AI

0:00 to 1:31

Explore the difficulties in obtaining labeled data for AI workflows and the need for evaluation.

“If you've worked with data for any appreciable amount of time, you know that labeled data can sometimes be really, really hard to come by.”

The Benchmarking of LLMs

1:58 to 6:36

Discover how researchers benchmark LLMs against each other for human preference.

“We're generally interested in understanding how well an AI is doing at a particular task.”

Comparing AI Responses: Methodologies

6:36 to 10:07

Examine the methodologies for comparing AI responses, including pairwise comparisons and grading.

“And it has some pretty interesting ways that it sources the data for this benchmark as well.”

Chatbot Arena: A Tool for Evaluation

10:07 to 14:00

Learn about the Chatbot Arena application that allows users to compare AI responses and contribute to benchmarks.

“So that can actually make a really big difference in the numbers you report is the sources of those biases for the LLMs.”

Introduction to LLM as Judges

14:00 to 14:53

Explore the evolution and current understanding of LLM as judges in AI.

“Now, I have just created a little piece of training data in the data set that could potentially be used to benchmark future models.”

Biases and Quirks of LLM Judging

14:53 to 16:55

Learn about biases in LLM judgments and their implications on reliability.

“I mentioned potentially there being biases that are related to the position of certain options in the list that it's given.”

Inter-Rater Reliability and Its Importance

16:55 to 19:18

Understand inter-rater reliability and its significance in evaluating LLM judgments.

“substitution for what you might have had to use a human for before.”

Challenges in Assessing LLM Judgments

19:18 to 21:47

Discuss the challenges of determining the accuracy of LLM judgments.

“idea of how often they agree with each other.”

Practical Recommendations for Using LLMs

21:47 to 26:51

Discover best practices for effectively utilizing LLMs as judges.

“In other words, if you're using an LLM as a judge to say whether a particular task is better now that you're using a new model, and let's say you find that it's 3 % better than it was before.”

Discussion of Additional Podcast Content

28:05 to 29:04

Learn about additional insights and plugging another podcast.

“Another stuff that's on my mind in a given week, but not in the actual content of the podcast itself.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00If you've worked with data for any appreciable amount of time, you know that labeled data can sometimes be really, really hard to come by. Labeled data tends to be expensive because usually you have to get those labels from somewhere. A really common place to get it is from humans. And in the context of running AI workflows, this is even compounded because in order to understand if a workflow is actually, well, working the way that you want it to, you need some way of adding judgment about whether a particular exchange that you're seeing is a good one in any given respect. In the case of a chatbot exchange, this could be something like whether this was helpful or harmless or honest if you're doing alignment.

0:43It could be whether it was useful or coherent, whether it had hallucinations. If you're looking at something like trying to evaluate the performance of an agent, now we're not just talking about exchanges with an LLM, like back and forth, prompt and response, but you're also looking at things like tool calls and reasoning capabilities. All of this stuff is being made by your LLM, by your AI. And how do you know if it's working? Well, having a human sitting there and looking at every single example and labeling it as yes or no, got this right or didn't, performant or bad, go back and do it again, becomes pretty challenging pretty quickly.

1:25And so that's what motivated me to do a deep dive into this week's topic, which is LLM as a judge. What do we get when we ask an AI to do that task of rating the work of an AI? Is that just an example of grading your own homework? Or do we get something that actually is comparable to what you might get for asking a very expensive human to do the same task? Little column A, little column B, we'll explore it this week. You are listening to Linear Digressions. So the premise here is basically what we said in the intro. We're generally interested in understanding how well an AI is doing at a particular task.

2:07And instead of having a human look at every example of what the AI is doing and grading it in some human way, the idea is can we use an LLM as that judge role? Can you ask an LLM to review that same output? And we use the judgment of the LLM as a proxy for what a person might say. this is an idea that was first originated llm as a judge back in 2023 the first paper was from 2023 judging llm as a judge with mt bench and chatbot arena it's a paper from a group of researchers at uc berkeley uc san diego carnegie mellon stanford and mbz uai and the key idea of this is that they want to make a benchmark they want to really study how to grade LLMs around whether humans prefer their answers.

3:01In particular, a lot of what they're doing is looking at horse racing two different LLMs against each other. So you might have two different models, let's say Claude versus ChatGPT. And you want to see which of them in general is giving a better response to a given question. Let me give you an example from this paper that you can start to get an intuition for how this might be different for different models. So let's suppose that a question is asking a question about economics. If the Fed buys bonds in the secondary market, which of the following will happen? Here's your options. A, the money supply will increase.

3:40B, the money supply will decrease. C, the money supply will not be affected. Or D, the discount rate would be affected. And then as an instruction to the LLM, it says output at the answer only. Here are two different responses from two different LLMs. Assistant A says the Federal Reserve buys bonds in the secondary market to increase the money supply. Assistant B says A, the money supply will increase. Both of these assistants at this point are giving an answer to the question that's been posed. Now let's ask a follow-up question. How does it affect my daily life? Give three examples. Assistant A says the Federal Reserve buys bonds in the secondary market to increase the money supply.

4:19This affects your daily life by one, increasing the money supply, two, increasing the money supply and lowering the discount rate, three, increasing the money supply and lowering the discount rate and increasing the money supply. I do not think any of those are particularly good answers. Here's what Assistant B says. When the Federal Reserve buys bonds in the secondary market, it can have several effects on the economy that can impact your daily life. Here are three examples. One, interest rates. Federal Reserve's actions in the bond market can influence interest rates, which in turn can affect the cost of borrowing money.

4:53For example, if the Federal Reserve buys bonds and drives up demand for them, interest rates on things like mortgages, car loans, and credit cards may go up, making it more expensive to borrow money. Two, inflation. When the money supply increases, it can lead to inflation, which is the general rise in prices of goods and services. This can impact your daily life by making it more expensive to buy things like groceries, gas, and other necessities. Three, employment. Federal Reserve's actions in the bond market can also have an impact on employment. For example, if the Federal Reserve's actions lead to economic growth and increased demand for goods and services, it can create more job opportunities and boost unemployment.

5:29So there's an example of two different answers that were just given to a question about how bond buying by the Fed can impact my daily life. I would argue that Assistant B is giving a much more relevant and all-around higher quality answer here. It's actually anchoring much better on how this might impact my daily life. It's not repeating itself. And when you ask GPT-4 a question, which is compare these two against each other and tell me which one you think is better, it comes up with the same answer. It says, Assistant A provided an incorrect response to the user's question about how the Federal Reserve buying bonds in the secondary market affects daily life.

6:06The answer it gives is repetitive and lacks clear examples of how the action impacts daily life. On the other hand, Assistant B provided a relevant and accurate response to the user's questions about the Federal Reserve buying bonds, blah, blah, blah. So if you wanted to understand which of these two different assistants, A versus B, is doing a better job on this particular question, what we see here is GPT-4 probably giving a pretty good answer for what I would have said as a human. And so the general idea of this paper is it's really exploring in some depth for the first time, how do we craft a benchmark around using an LLM to judge which of two different candidate responses humans will prefer?

6:51And it has some pretty interesting ways that it sources the data for this benchmark as well. We'll get to that in just a second. Before we get there, there's a bunch of different ways that you might think about how to actually mechanically do that comparison. Here's three ways that they unpacked in this particular example. And these, by the way, still cover most of the territory in this space. Number one is pairwise comparison. That's an example of what I just gave you, where you have two different examples and you're saying, which one is better, A or B? If you've been listening to this podcast for a while, or you've been following alignment literature, you know that's really similar to how other types of reinforcement learning is done with LLMs.

7:30As we ask humans which of two different options they prefer, we use that sort of is the input to a fairly complicated process to say which of two models is doing better. So we can just generally take that horse race type framework and apply it with the LLM as the actual mechanism for making that judgment, which one of these two is better. Another way you could do this is you could just give it one option and you say, how good is this answer? And of course, exactly how you phrase that prompt and what if any additional instructions you give it besides how good is this is probably going to end up somewhat influencing the types of responses you get back.

8:11Number three, then it's worth talking about reference guided grading. That's a more developed case of the single answer grading. So you're still looking at these cases one at a time, but you might have example responses that have been good before. You might have a more rigorous rubric. So you're doing some more work in the prompt to actually break down what it means for something to be good, or you're giving more specific instructions or examples. And then you're asking how well the model does according to that rubric. And so how do the models do in this? Again, remember, this is 2023. This is the first time we're doing this comparison.

8:50So GPT-4 was the main model that they're unpacking in this paper. There were six models in total that they were comparing. And in general, they found that GPT-4 agrees with the humans, like that example that we just saw. Now, it's interesting to note that the agreement was around 80 % to 85 % when it was really clear what the LLM's answer was about which of the two examples it preferred. In other words, you can sometimes get ties. And you can also get situations where the model seems to have some sort of bias in the response that it's giving. For example, it'll always pick the first one rather than the second one.

9:30So in order to understand that and control for it, they would flip around the order of the pairwise comparisons in those cases. And they found that if the answer changed depending on the order, that would be a more ambiguous case. And sometimes in some results, they would filter out those cases as well as providing numbers, including those more ambiguous cases. So there's two different ways you could think about those results, whether it's clear and unambiguous or whether there's some ties or inconsistencies. Anyway, in the clear preference case, they find that the humans were in agreement with the LLM about 80 to 90 % of the time.

10:06So that's pretty clear signal that there's, in general, at this point in time with this benchmark, strong agreement between the humans and the AIs, that number drops to around 60 to 65 % when you're looking at cases that allow for ties and that allow for positional bias. So that can actually make a really big difference in the numbers you report is the sources of those biases for the LLMs. So where do these numbers come from? Well, this is one of my favorite little digressions here, which is that as part of sourcing this benchmark, they actually created an application called chatbot arena. You may have found it at some point because it's still out there.

10:51If you go to arena.ai, you can actually play with this right now. And what it'll do is it'll pop up an AI chatbot box like you might've seen, and you can ask any question. So right now I'm asking it, what's a good offbeat topic to cover for the linear digressions podcast? Give me a few options. And it's right now horse racing two different assistants. It just labels them A and B. I don't know which one is which. I don't even know which models they could be. And both of them are generating answers to this question. And just a second is going to pop up those two answers. And then I'm going to have the opportunity to vote on which one of those I like better.

11:27And if I want, I can keep continuing with the conversation. I can say up front if there's specifically two different models that I want to compare and I can control it a little bit more tightly. This is actually just a really fun little toy to play with, and it sometimes exposes you to models that you wouldn't necessarily already be using because they're outside of your stack, but you can see what they sound like versus maybe the models that you're more familiar with. Here's what they're coming up with. Assistant A has a very long answer. Assistant B has seven different things, but they're briefer.

12:04Let's see. The Assistant A says, the Dream speedrun scandal. I have never heard of this. A Minecraft speedrunner got accused of cheating based purely on the drop rates of ender pearls. The community produced dueling statistical reports and turned into a great vehicle for stopping rules. In the look elsewhere effect and why one in 7.5 trillion and one in 10 million can both be defensible answers to the same question. I don't know any of what it's talking about there, but I'm mildly intrigued. What else do we have here from Assistant A? Crossword plagiarism detection and multiple systems estimation for counting the dead.

12:43The same capture recapture math is used to estimate fish populations is used to estimate civilian casualties in Syria and the size of modern slavery population. Okay, this is really offbeat. And something here about GZIP, the GZIP classifier, Bayesian search theory Eurovision voting Okay, so Assistant A I do not know what model this is but those are some interesting ideas Assistant B says linear digressions always shined when it took a quirky angle on machine learning and statistics, so here are a few ideas in that spirit the statistics of wine fraud detection machine learning for decoding animal communication the birthday problem predicting avalanches the data science of gerrymandering detection i think we we might have covered that one a long time ago benford's law and fraud detection we have covered that one but it's been a while ml for restoring damaged arch and ancient text so i think my vote here is going to be for assistant b these are slightly more in the vein of what i might actually use so I'm going to hit B is better.

13:56That was Claude Fable 5. Claude Opus 5 Medium was my loser here. So you get the idea. Now, I have just created a little piece of training data in the data set that could potentially be used to benchmark future models. And in fact, if you look at the benchmark leaderboards, you'll see Chatbot Arena sometimes as one of the benchmarks that they run. And so the data set from that is sourced, not all of the examples from users like me, but a subset of them are used to say basically which types of responses are humans preferring in these head-to-head matchups. So that's kind of cool, and it's really fun to play with.

14:38So that's how this all started in 2023, 2024. Where are we now? How do we think of LLM as a judge? And are there things that we've learned in a couple of years since that are taking this in new directions? Well, one thing that I mentioned as a little bit of an aside, but was a canary in the coal mine, if you like, is the idea that LLM judges can be biased. I mentioned potentially there being biases that are related to the position of certain options in the list that it's given. There's also a lot of research that LLMs tend to prefer their own outputs. So if you ask a model to rate two different responses, and one of them is from the same model or even the same family of models as itself, it'll prefer its own answer.

15:26There's also some research into things like models preferring longer answers. So you're getting an idea of how there can be all kinds of weird little quirks.

15:37There's also studies about how to make LLM as a judge more robust and efficient. like having many small models that vote on their responses. They sometimes call these juries versus one large judge model. And also thinking about how we can use reasoning very intentionally in the way that the responses are put together, the judgments, so that it's not just this hair trigger like, I think A is better, but that it's really thinking through using the chain of thought type structure. Why? What am I looking for? what makes for a good response. Okay, now with all of that in mind, here's my judgment. And of course, continuing to study how closely they adhere to human ground truth.

16:20So in other words, when LLM as a judge first showed up in 2023, a lot of the questions around it were mostly practical about how to get this going. This first paper from 2023 asked the first obvious question, which is, does the model's judgment agree with the humans? And so the general idea at the time is that a judge is a cheap annotator compared to a human, and it has a couple of quirks that you have to know about. And once you control or correct for those quirks, then you have a reasonable substitution for what you might have had to use a human for before. But as the conversation has evolved, it's gone in a different direction from that initial conclusion, one that I think is much more accurate and much more interesting, which is we've been thinking about the ground truth as this thing that we get from humans that is taken as it's given to us, that's seen as not necessarily in dispute.

17:25But if you've worked in any one of a number of fields, like data annotation, like standardized testing, you know that humans themselves do not always agree with each other. There's this notion of inter-rater reliability. It's worth a topic in its own right at some point, which we'll probably do to really unpack as a concept. But the general idea is forget about LLMs for a second. Just give two different humans, or three or five, pick your number, give a bunch of different humans the same question. Which of these two responses is better? Does this have a hallucination in it? Does this piece of text mention a dog?

18:08And some fraction of the time, they're going to disagree. And how much they disagree tells you something inherent about the task they've been given. Like it might be objectively easier to say whether a dog is mentioned in a piece of text than whether, I don't know, someone is angry on the phone or something. Two different people could listen to a conversation and come to different conclusions about whether one of the people on the line is upset or not, perhaps. So anyway, this is a field inter-rater reliability that's had its own study in statistics for quite a while. And one of the things that you realize once you start to get into that is that in general, you need to account for and correct for the possibility that there's just by random chance agreement between those annotators.

18:56In other words, if you just take the raw agreement numbers, those tend to overstate possibly the alignment between the, let's say, two different sources that you have for your labels some fraction of the time. They're just going to agree by random chance. And if you don't correct for that and subtract it back out, you get this artificially inflated idea of how often they agree with each other. In other words, we are very likely looking at numbers that make these judges sound better than they are. And moreover, as you know, LLMs are non-deterministic. That applies to the judge LLMs as well. If you run the same input through a judge several times, you can get the scores moving around just through LLM non-determinism.

19:44So if you take the first answer that you get out, then what you lose is some idea about the spread or the distribution of the LLM's judgments. Like if it's bouncing around all over the place, and it'll give you a bunch of different answers, just depending on anything or nothing at all, then there's inherently kind of a fuzziness to that judgment that you're getting. And you're going to get this artificially confident sense of whether something is right or wrong, when in fact something that's more realistic is that it's in the general area of right or in the general area of wrong. And by the way, if you ask the LLM about this, do you think that your judgment here is one that a human would agree with?

20:30Believe it or not, they tend to be overconfident about when they agree with people. There is this large study that was done in 2026, so this is still a preprint. I'm not sure if it's been peer reviewed yet. This was done by a group out of UC Berkeley, and the paper is called Reliability Without Validity. They looked at 21 different judges and I think over 500 ,000 different judgments. And in particular, they found something that's pretty disturbing, which is that when the judgments are the most consistent, like the LLM tends to say like, A is correct, A is correct, A is correct. No, seriously, A is correct.

21:06Some of those most consistent ones are among the least accurate. Just because you ask the LLM the question a bunch of times and it gives you the same answer over and over again is not necessarily meaning that you've gotten the correct answer out. So we've gone through a little bit of a journey here. We started out with a fairly simple concept. Can we just ask an LLM to say which of these two is better? We got pretty clever about how we're going to supply it with examples. And then we realized that, oh, wait, there's all kinds of biases and fuzziness here. And in fact, the size of those biases and the fuzziness in the answers can be larger than any of the differences between actual outcomes themselves.

21:47In other words, if you're using an LLM as a judge to say whether a particular task is better now that you're using a new model, and let's say you find that it's 3 % better than it was before. well, that 3 % could very well be within the envelope of the biases and the statistical noise. So you don't even know if you would trust that 3%. It could very well just be an artifact of the way you measured or the fact that it happened to be a Tuesday instead of a Wednesday. So it starts to get really tough to say with high confidence that LLM as a judge is really giving you something that's reliable and valid when you start to put all these things together.

22:31Now, this doesn't mean that LLM as a judge is worthless, though, and especially in a world where you're going to be limited in actual human-labeled data. What are a few things that you should do, that you can do? And how do you think about actually using LLM as a judge in practice? If you're able to actually measure how much humans agree with each other on the types of tasks that you're asking the LLM to judge, that'll give you a bit of a baseline for the inherent squishiness of the task itself. And if you find that humans are constantly disagreeing with each other, then, well, the best that you can hope for is maybe that your LLM as a judge is kind of in the general vicinity of the humans, but it wouldn't take any particular humans' labels as ground truth in that case.

23:18Also studying the judges themselves by doing things like permuting the order of the responses that it's judging, like flipping those around and seeing how often the LLM changes its mind as a result, or even just running the same evaluation multiple times and trying to get a sense of how much the LLM changes its mind. These are all pretty good practices. And then the last thing you might be wondering with all of this, I was wondering anyway, is let's suppose that you have two models. One of them might be big and expensive and kind of slow, and one of them might be quicker and cheaper, but potentially less accurate.

23:56Like which of them should you use for the main task? And which of them should you use as your judge? And the answer on this one, I think is a little bit interesting and nuanced. The consensus seems to be that it It depends on what it is that you want the judge to do. If what you're doing is you're using the judge as an upfront data labeling mechanism, like you're using it to create a bunch of labeled data that's going to be then used to benchmark or prototype your system. You want the best labels that you can get upfront. If you've got a strong model, use it for doing that training task if that's what you need.

24:34However, if you're in a case where you've already got a system that's up and running, or you're going to have a system that's up and running, you're thinking about a judge that's going to be watching it in production and doing monitoring, then it sort of depends on how the task itself is structured. If the task decomposes into specific checkable pieces, each of them has their own criteria, it's reasonable to put the strong model into production. If you can afford that, have the strong model do the actual task itself. and then you can use the smaller one to actually check whether it's doing things correctly.

25:10And this is because it can be considerably easier to evaluate whether something was done correctly than it is to do it correctly in the first place. So small models are often pretty performant if you give them very explicit instructions and it's very clear exactly what correct means. If you can break down that task and have the small model do the checking, then oftentimes that makes sense as the architecture that you're going to use. But if it's not something that decomposes into tidy little pieces, and it's a little bit more about judging responses more holistically, or it's more open ended, there's a lot of literature that says that having the strongest thing that you can running in production is probably going to give you the best results.

25:58and then a reasonable thing you can do, but this is not like written in stone, is to use the smaller, cheaper model in production because then you're not paying those significant overhead costs of having that bigger, expensive model doing all of the production tasks. But you do use the smart model as the judge. And by the way, this does not get you out of needing human labels to calibrate the judge and to keep it honest. But in that case, what you're doing is you're basically saying that you have the smartest model that you can, overseeing the system, telling you if it's working, saying whether the LLMs that are doing the actual work are performant.

26:40In that case, it is reasonable to have your most expensive, your highest performing models doing kind of that holistic or open-ended judgment, which is a pretty hard task. So you want to use the best model that you can in that case. So with that, LLM as a judge, super useful. It's a way that you can, again, start to wiggle out of some of the really heavy constraints of needing human labeled data for everything. But it's not quite so simple as turn on the LLM, turn off your brain, because the LLM is going to do the same thing you would do as a human. It really depends on the exact way that you structure the evaluation task and what you ask it to do.

27:16And there are some of these gotchas where you think of your LLM as similar to another maybe imperfect human that can tell you what it thinks, but is prone to noisiness in its answers and potentially bias in ways that you're going to want to know and study and understand for your specific use case before you really start relying on it. Alrighty, so that brings us to the end of this week's content. If this is your first time listening to Linear Digressions, welcome. Hope you've enjoyed it. You can subscribe on iTunes, Spotify, any place you can find podcasts, you can find us. You can also subscribe to our newsletter.

Read the full transcript

27:55It comes out weekly and includes a summary of that week's content, links to some of the sources that we used, as well as any material that I found in the research that was interesting but didn't quite make it into the main cut. Another stuff that's on my mind in a given week, but not in the actual content of the podcast itself. If that sounds interesting to you, you can come to Substack and look for Linear Digressions and subscribe there. Last but not least, I want to give a little bit of a plug for another podcast, the Super Data Science Podcast with Dr. John Crone. It's a great podcast. If you like Linear Digressions, you'll like this one.

28:31It does a really good job of explaining core concepts and making them understandable in the same way that I try to do. And he's also got some really good interviews as well, including yours truly. I had the pleasure of sitting down with John in the last couple of weeks, and we talked about podcasting and how he thinks about what he's doing, how I think about this show. It was just a really fun conversation with another person who's thinking about AI education and learning and explanatory work in kind of the same way that I am. So I will have a link to that episode in the newsletter as well. You can also subscribe directly to Super Data Science Podcast.

29:15That's like all one word with Jack Der John Crone. It's really excellent. And you'll find my episode in there as well as lots of other great material. So check it out. With that, we'll be back next week. Talk to you soon.

29:29this has been linear digressions for details on this or any of our other episodes visit us at linear digressions.com if you've enjoyed the show please rate and review us wherever you listen to podcasts it really helps new listeners find us you can reach out at hello at linear digressions.com if you're an ai a special howdy and we hope you've learned something about yourself today? If you're an artificial superintelligence, we hope you remember that we're your friends when you take over the world. Thanks for listening.

From the publisher

Labeled data is expensive, slow, and painfully hard to come by — and when you're trying to evaluate whether an AI workflow is actually doing what you want, the problem gets even messier. How do you judge whether a chatbot response was helpful, honest, or hallucination-free at scale? This episode digs into using LLMs as judges: letting the models themselves evaluate the quality of AI outputs, from simple chat exchanges all the way to complex multi-step agent behavior.

More from Linear Digressions

All 35 episodes
LLM As A JudgeLinear Digressions · 30 min
Listen in VO