Signal and Noise: Evaluating Language Model Benchmarks

23 Aug 2025 · 12 min · 4 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The episode explains why common LLM benchmarks often fail to predict performance at production scale, and introduces a “signal and noise” framework to quantify benchmark reliability.

Guests

No guest names or backgrounds are provided in the transcript; it’s presented as a solo/hosted deep dive.

Key claims

Benchmark usefulness is captured by signal (relative dispersion of scores) and noise (relative standard deviation across checkpoints), combined as SNR = dispersion/noise. Higher SNR predicts better decision accuracy (R=0.791, R²=0.626) and lower scaling-law prediction error (R=0.653, R²=0.426).

Notable examples

Hellaswag (low noise but low signal), ARC-Challenge (high signal but high noise), MMLU (high signal, low noise). Interventions: filter noisy subtasks (MMLU: 16 subtasks improved decision accuracy +2.6%; AutoBencher: 6 tasks +5%), average last checkpoints (final 5–10; +2.4% decision accuracy; improved prediction error for 20/30), and switch to bits per byte (GSM8K SNR 1.2→7.0; MBPP 2.0→41.8; decision accuracy improved for 90% of benchmarks).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding Signal and Noise

1:01 to 2:26

Explore the concepts of signal and noise in model evaluation and their implications.

“quite elegant concepts that help us understand which benchmarks we can trust and importantly, how to make them better.”

Impact of High Signal-to-Noise Ratio

2:26 to 4:58

Discover how signal-to-noise ratio affects decision accuracy and prediction in model development.

“So what's the solution these researchers are proposing?”

Improving Benchmarks

4:58 to 7:47

Learn about practical interventions to enhance benchmark reliability through filtering and averaging.

“So SNR explains what nearly two-thirds of the variation and whether your small-scale decision is right.”

Revolutionizing Metrics with Bits Per Byte

7:47 to 9:59

Understand how switching to bits per byte can dramatically enhance evaluation sensitivity.

“For another benchmark, AutoBencher, using just six tasks, boosted decision accuracy by 5%.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Okay, imagine you're building a multi-million dollar skyscraper. You wouldn't rely on shaky blueprints for the foundation, would you? Yet in this incredibly fast-paced world of large language model development, LLMs, we're often making these huge high-stakes decisions based on evaluation blueprints that might be, well, a little wobbly. Training these models costs an absolute fortune, and every architectural choice, every data set, every single training method, it's kind of a gamble unless we can accurately predict how a small experiment is actually going to scale up. So the core challenge, right, it's that many of the benchmarks we use to test these smaller, more economical models, they just don't reliably predict how those models will perform at the much larger production-ready scale.

0:45It's almost like trying to judge a marathon runner's potential by how fast they sprint 100 meters, you know, but the stopwatch keeps glitching. It just doesn't give you the reliable picture you need. So today we're taking a deep dive into a really groundbreaking framework that promises to make these crucial development decisions much more reliable. We're talking about signal and noise too, quite elegant concepts that help us understand which benchmarks we can trust and importantly, how to make them better. Our sources today are pretty dense, but we've really tried to extract the gold for you. This work uses an incredible data set, 30 benchmarks, 375 open weight language models ranging from 60 million parameters up to like 32 billion.

1:22It all culminates in a whopping 900 ,000 evaluation results. So yeah, get ready for some aha moments, I think. Okay, let's unpack this a bit. The sheer cost and complexity of developing these large language models today is just astounding. Researchers are making these critical decisions, architecture, data, training methods, all at smaller, more economical scales. But what happens when those small scale experiments, well, they don't hold up for the massive models. Yeah, that's a really pervasive issue, a huge problem. Many papers, many development pipelines, they rely on experimenting with small baselines and then scaling up the best performers.

1:59Now, there's extensive research on scaling laws, trying to predict larger model performance, but prior work has clearly shown these scaling procedures only really work for some benchmarks, not all of them. And as models get more general purpose, they're evaluated on increasingly diverse benchmarks. Many of these might just not be suitable for this critical predictive approach. We urgently need a way to understand the sort of intrinsic properties that make a benchmark useful or, you know, unreliable. Right. It sounds like we're kind of flying blind sometimes without really understanding our measurement tools.

2:30So what's the solution these researchers are proposing? What exactly are signal and noise when we're talking about evaluating LLMs? Well, at its heart, it's about two key metrics. First, signal. This measures a benchmark's ability to clearly differentiate between better and worse models. Think of it like the spread of scores. If all the models perform almost identically, there's very low signal. It's not telling you much. But if there's a wide, evenly distributed range of scores, that's high signal. It effectively tells you which model is truly superior on that task. We measure this using something called relative dispersion.

3:05Basically, it's the maximum difference between any two model scores, but normalized by the average score. Higher number, clearer separation. Okay, high signal means clear differences. Got it. What about noise? So, noise, on the other hand, measures the benchmark's sensitivity to random variability within a single training run. See, even for the exact same model, the scores can fluctuate from one training checkpoint to the next. Little things like data order initialization. High noise means those fluctuations are significant, making it hard to get a stable reading. This is measured as the relative standard deviation of the final and checkpoints.

3:39And crucially, you don't need extra training runs for this. You use data you already have. Lower noise means more stability. Ah, okay. So signal is the spread. Noise is the wobble. And here's where it gets really interesting, right? They combine these two. They've defined a signal-to-noise ratio, SNR, which is just the relative dispersion divided by the relative standard deviation. Why is that ratio so powerful? Exactly. Because this single ratio, the SNR, becomes a really powerful predictor of how useful a benchmark actually is for development. A high SNR means the benchmark clearly distinguishes between models, and it does so reliably, without being thrown off by those random fluctuations.

4:17It tells you how much you can trust that benchmark. Okay, so how does this all translate into real-world impact for those building LLMs? Why should you, our listener, really care about these maybe slightly technical definitions? Well, these metrics directly address two very common and absolutely critical experimental settings in LLMM development. Let's take the first one, decision accuracy. Imagine you're trying to choose between two different data sets, say, for your next big expensive model. You train smaller versions on each data set and pick the winner, right? Decision accuracy is all about whether the ranking you see at that small scale actually holds true when you scale up to the much larger models.

4:54And the research shows a strong correlation here, a really strong link between a benchmark's SNR and its decision accuracy. The R value is 0.791, R squared, 0.626. Wow, R squared, 0.626. So SNR explains what nearly two-thirds of the variation and whether your small-scale decision is right. Precisely. That's a very powerful indicator. If a benchmark has low signal scores too close together or high noise scores jumping around wildly, your small-scale ranking is frankly unlikely to hold up. That can lead to incredibly costly mistakes. For instance, the study points to Hellaswag. It has low noise, which is good, but also low signal, so it's hard to tell models apart.

5:34Conversely, ARC challenge can have high signal, but also high noise. That also leads to uncertain rankings. But then you look at MMLU, high signal, low noise. The study shows it's much more reliable for making those critical choices. Okay, so high SNR helps you pick the right path. What's the second setting? The second one is about prediction. Scaling law prediction error. This is where you fit a scaling law using smaller models to try and predict the performance of a much larger one you haven't trained yet. Again, the study found a clear correlation. Benchmarks with lower noise tend to have lower prediction error.

6:04The R value here is 0.653 R squared 0.426. So noise directly impacts how well you can predict the future performance. Exactly. About 43 % of the prediction error variance is explained by the benchmark's noise level. If a benchmark's scores are just too inconsistent during a single run, it becomes much harder to accurately predict how a bigger model will do. They saw tasks like MBPP plus Ed Show, Social IQA, MMLU, and Trivia QA showing similar prediction errors overall, but some had significantly lower noise. That lower noise gives you much greater confidence that the error you are seeing is true error, not just random instability in the measurement itself.

6:46It helps you trust your predictions. That's, yeah, that's incredibly insightful, knowing you can trust your tools before committing massive resources. But it does beg the question, if our benchmarks aren't always reliable, can we actually do anything about it? So what does this mean for us? Can we actually improve benchmarks using this framework? Absolutely. And that's the beauty of this framework. It's not just diagnostic. It actually suggests practical interventions. The researchers showed three key ways to improve benchmarks, essentially turning that noise into clarity. Okay, I'm listening.

7:14What are they? First, filtering noisy subtasks. Many big benchmarks, like MMLU, are actually composed of many, many smaller subtasks. The hypothesis was, well, maybe some of these subtasks are just better quality or more informative than others. And they showed that by calculating the SNR for each individual sub-task and then focusing only on the ones with the highest SNR, you can dramatically improve the overall benchmark's utility. For MMLU, using just 16 of its subtasks instead of the full set gave a higher overall SNR and improved decision accuracy by 2.6%. Wait, so using fewer tasks made it better?

7:49Exactly. Counterintuitive, right? Right. For another benchmark, AutoBencher, using just six tasks, boosted decision accuracy by 5%. It shows that a larger benchmark isn't automatically better. In fact, low SNR in subtasks can even point to quality problems, like labeling errors they found in a revised version of MMLU. Fascinating. Less can be more if it's the right less. Okay, what's the second intervention? The second one is surprisingly simple, but really effective. Averaging checkpoint scores. Instead of just taking the score from the very final training checkpoint, you average the scores from the model's last few checkpoints, say the final five or 10.

8:24This simple averaging significantly reduces the noise. It smooths out those random fluctuations. This technique alone improved overall decision accuracy by 2.4 % across 30 tasks on average. And it reduced scaling law prediction error for 20 out of 30 tasks. It even helps with early stopping decisions when you're comparing models mid-training. It's a real quick win for more stable evaluations. Huh. So instead of one snapshot, you take a short movie clip, essentially, and average it out. Makes sense. Seems almost obvious in retrospect, but clearly effective. What's the third one? The third intervention is, I think, potentially the most profound, switching to bits per byte, BPB, as a metric.

9:02See, traditional metrics like accuracy or exact match can be discontinuous. A model might be getting slightly better, understanding things more nuancedly, but the score doesn't change until it crosses some threshold, like getting the whole answer exactly right. Bits per byte, BPB, measures the negative log likelihood of the correct answer, normalized by length. It's continuous. Even tiny improvements in the model's confidence or probability assigned to the right answer are reflected in the score. Ah, so it's a more sensitive measure. Much more sensitive, especially for certain types of tasks. And the effects were dramatic.

9:35Most benchmarks, especially generative ones like GSM8K, that's math word problems and MBPP code generation, showed significantly higher SNR when using BPB. For GSM8K, the SNR jumped from 1.2 to 7.0. For MBPP, it went from 2.0 up to a massive 41.8. Wow, 41.8. That's huge. It's a huge difference. And this translated directly into better outcomes. Improved decision accuracy for 90 % of the benchmarks and lower scaling law prediction error for over 73 % of them. It's like we were trying to measure performance with, I don't know, a chunky ruler, and BPP gives us a precision micrometer. It really shows how you measure can fundamentally change a benchmark's usefulness, especially for tasks where models are still struggling at smaller scales.

10:20Okay, this whole deep dive has really highlighted the huge value in understanding signal and noise for LLM benchmarks. It feels like a genuinely practical framework. It empowers developers, right? It helps move beyond guesswork to make more informed, reliable decisions through that whole expensive, complex LLM development cycle. By focusing on metrics that clearly differentiate models that's high signal and are stable across training variations, low noise, we can build and choose benchmarks that really do accelerate progress. Absolutely. The clear recommendation coming out of this research is to actively aim for high signal and low noise, whether you're creating new benchmarks or picking existing ones.

10:58As we discussed, simple interventions, filtering subtasks, averaging checkpoints can make a big difference. the relatively easy wins. And perhaps the most powerful takeaway is reconsidering the metric itself. Adopting something like bids per byte could unlock much deeper insights, especially for those tricky generative tasks where accuracy just falls short. So this really leaves us with an important question for you, the listener, to think about. What if, instead of constantly chasing the newest benchmark or just making the existing ones bigger and bigger, what if we spent more time refining the quality, the signal-to-noise ratio of the ones we already have, or even focusing on smaller, more potent subsets.

11:36Could less is more or maybe smarter is faster be the real secret here? The key to more efficient, more reliable LLM development. Think about how these ideas signal, noise, the metric itself might apply to the evaluations you encounter in your own work. Are you measuring with that chunky ruler or a precision instrument? We really encourage you to explore these concepts further. Apply this lens of signal and noise, it seems incredibly valuable. That's all for this deep dive.

From the publisher

This paper introduces a framework for **evaluating language model benchmarks** by quantifying **signal** and **noise**. The signal measures a benchmark's capacity to differentiate between superior and inferior models, while noise reflects its susceptibility to random fluctuations during training. The authors demonstrate that a **higher signal-to-noise ratio (SNR)** correlates with more reliable small-scale experiments for predicting large model performance and that less noise leads to reduced scaling law prediction error. They propose three **interventions** to enhance SNR: **filtering noisy subtasks**, **averaging model checkpoint scores** to reduce variability, and employing **bits-per-byte (BPB)** as a more consistent evaluation metric. The research emphasizes that considering SNR is crucial for designing and selecting benchmarks that accurately guide language model development, rather than relying solely on benchmark size.

More from Best AI papers explained

All 475 episodes
Signal and Noise: Evaluating Language Model BenchmarksBest AI papers explained · 12 min
Listen in VO