Textual Bayes: Quantifying Uncertainty in LLM-Based Systems

12 Jul 2025 · 9 min · 6 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The episode discusses the paper “Textual Bayes: Quantifying Uncertainty in LLM-Based Systems,” which aims to make LLM outputs more reliable by producing calibrated uncertainty estimates. It argues that LLMs struggle in high-stakes domains because they often can’t quantify confidence, are hard to interpret when closed-source, and are sensitive to prompt wording.

Guest backgrounds

No guest names or biographies are provided in the transcript; only the paper authors are mentioned (Brendan Lee Ross and 10 other researchers).

Key claims

Treat prompts as “textual parameters” in a Bayesian statistical model; run Bayesian inference over prompt text to quantify uncertainty; incorporate prior beliefs via free-form text; use MHLP (“Metropolis Hastings through LLM proposals”), an MCMC-based algorithm that can work even with closed-source models by sampling prompts and observing outputs.

Notable examples

The episode highlights uncertainty calibration for tasks/benchmarks measuring uncertainty quantification (UQ), claiming improved predictive accuracy and better-aligned confidence scores (confidence reflects likelihood of correctness).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Challenges of LLMs

0:46 to 1:46

Discussing the challenges of uncertainty in large language models.

“okay, I'm 90 % confident about this part, but maybe only 60 % on that bit.”

Introduction to Textual Bayes

1:47 to 3:39

Explaining the concept of Textual Bayes and its relevance in LLMs.

“It's called Textual Bayes, Quantifying Uncertainty in LLM-Base Systems.”

Bayesian Inference in Practice

3:40 to 5:30

Examining how Bayesian inference can be applied to prompts in LLMs.

“Okay, meaning the LLM could potentially tell you, okay, based on the way you asked, I'm 95 % sure about this conclusion, but maybe only, say, 60 % sure about that recommendation.”

The MHLP Algorithm

5:31 to 6:33

Introducing the MHLP algorithm and its significance for LLM workflows.

“And they say this MHLP is a turnkey modification.”

Testing and Results

6:34 to 8:02

Reviewing the testing and results of the MHLP algorithm across benchmarks.

“The algorithm sounds practical, especially for closed models.”

Implications for AI Trust

8:03 to 8:53

Discussing the broader implications of quantifying uncertainty in AI.

“This textual Bayes paper introduces a really fresh perspective.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Welcome to the Deep Dive. Today we're looking at something, well, pretty fundamental in AI right now. Large language models, LLMs. Yeah, they're everywhere. Exactly. You've probably used them, seen what they can do. I mean, these AIs are just transforming how we tackle problems, right? From writing emails to, I don't know, complex coding tasks. Their ability to work with language is, it's really quite something. It really is. And, you know, it feels like they get better almost daily. but there's still this core challenge, especially when you think about using them for really important stuff. High stakes things.

0:38Exactly, like in medicine or finance. They have trouble telling us how sure they are. They might give an answer, but they can't reliably say, okay, I'm 90 % confident about this part, but maybe only 60 % on that bit. And that lack of certainty, that's a huge barrier, isn't it? It's massive, yeah. And it seems like it's made even trickier by a couple of things. First, a lot of the really powerful ones are, well, they're closed source, right? Right. Black boxes. We can't really see inside, figure out their process, which makes understanding why they give a certain answer incredibly hard. It's like, you know, trying to get inside the head of a genius who won't explain their thinking.

1:13Precisely. And then there's the other factor. The prompts. Yeah. How sensitive they are to the exact wording you use. This whole field of prompt engineering right now, it feels, well, almost like an art form sometimes. A very time consuming one. Lots of trial and error. Exactly. You're tweaking words, trying different phrases. Yeah. It makes the whole system feel a bit brittle, you know. That guesswork, the lack of transparency. Yeah, big hurdles for using them reliably. But that's what makes this paper we're looking at today so interesting. It offers a potentially transformative way to think about this.

1:46Ah, the source material. Right. It's called Textual Bayes, Quantifying Uncertainty in LLM-Base Systems. It's by Brendan Lee Ross and 10 other researchers. Quite a team. Okay, textual Bayes. And our mission here is to really unpack what they're proposing. They talk about a viable path to tackle this uncertainty problem. By looking at LLMs through a Bayesian lens, the goal being to get systems that aren't just clever, but actually more reliable and calibrated. Reliable and calibrated. Okay, let's get into what that really means for building and frankly trusting these AI systems. Right. So the absolute core idea here, the real innovation, is how they suggest we think about prompts.

2:29Okay. Instead of just seeing them as instructions, like write me a poem, they say we should interpret them as textual parameters in a statistical model. Textual parameters. Okay, that shifts things. So you're saying the actual words we use become like variables, things the model can analyze statistically. That's exactly it. The prompt isn't just a static command. it becomes a dynamic part of the model that can have probabilities assigned to it. They can be updated based on data. Wow. Okay. And if you treat prompts that way, what does that unlock? Well, it lets them apply Bayesian inference directly over these textual prompts.

3:03Now, Bayesian inference, you know, it's the standard way to update your beliefs as you get more evidence. Usually you do it with numbers, model weights. Right. The really novel thing here is doing it with the natural language of the prompts themselves. And they show you can do this even with a relatively small training data set to build up this probabilistic understanding. So it's not just about finding the perfect prompt anymore. It's about understanding the probabilities associated with different ways of prompting. What's the real world benefit there? The benefit is huge. It enables what the paper calls principled uncertainty quantification.

3:38Principled uncertainty. Okay, meaning the LLM could potentially tell you, okay, based on the way you asked, I'm 95 % sure about this conclusion, but maybe only, say, 60 % sure about that recommendation. That's a game changer compared to just getting a confident-sounding answer that might be wrong. Yeah, that level of nuance is exactly what's missing. And what's really interesting, too, is that this whole framework lets you incorporate prior beliefs expressed in free-form text. Wait, prior beliefs? You mean, like, human knowledge? Essentially, yes. Imagine you're setting up an AI for a task. You could tell it, not just the command, but also some of your assumptions or knowledge about how certain props should influence the outcome, just using plain English.

4:20So you can kind of guide its reasoning more directly with text. Kind of, yeah. It's about injecting that human insight right into the model's probabilistic framework. It really makes you think, doesn't it? How can we best blend our knowledge with these powerful algorithms? That's a deep question. Okay, so injecting knowledge quantifying uncertainty sounds powerful, but turning this Bayesian idea into something that actually works for LLMs, especially given how computationally heavy Bayesian stuff can be, that sounds like a challenge. It is, and that's where their algorithm comes in. They call it MHLP.

4:53MHLP. It stands for Metropolis Hastings through LLM proposals. It's a novel MCMC algorithm. Okay, MCMC, Markov Chain Monte Carlo. That's about sampling, right? Exploring possibilities without checking every single one. Exactly. Think of it like smartly exploring a huge landscape of possible prompts and their associated uncertainties. MHLP cleverly combines techniques we already use for prompt optimization, that tweaking we talked about, with these established MCMC methods from Bayesian statistics. Ah, okay. So it's blending the practical prompt tuning with the theoretical rigor of MCMC sampling.

5:30That makes sense. And they say this MHLP is a turnkey modification. Is it really that straightforward to plug into existing systems? Well, turnkey is maybe a strong word in practice, but the paper argues it's designed to be integrated relatively easily into current LLM workflows. It doesn't require you to fundamentally rebuild everything. And here's the really critical bit. They claim it works even for pipelines that use closed source models. Ah, the black boxes again. Sure. So even if you can't see inside the LLM, this MHLP method applied sort of from the outside, can still give you uncertainty estimates.

6:03That's the claim. because it operates on the prompts and the outputs. It doesn't necessarily need access to the model's internal weights or architecture. You're inferring the uncertainty based on how the model responds to different probabilistically chosen prompts. Wow. Okay. That is significant because it means all those teams using proprietary models could potentially adopt this. It opens the door much wider. It really could. Yeah. It potentially brings Bayesian uncertainty quantification to the mainstream of LLM development, not just the open source research world. So the theory is neat. The algorithm sounds practical, especially for closed models.

6:39Did it actually deliver in tests? What were the results? They tested it quite thoroughly, it seems, across a range of standard LLM benchmarks and tasks specifically designed to measure uncertainty quantification, or UQ. Yeah. And the results were positive. They reported improvements in both predictive accuracy, getting the right answer more often, and in the uncertainty metrics. So the models weren't just more accurate, they were also better calibrated, meaning their confidence scores actually reflected their likelihood of being correct. That's great. Concrete improvements. It really does sound like that viable path they mentioned.

7:13It does. It feels like a solid demonstration, as they put it, of how you can bring methods from the, you know, very rich history of Bayesian statistics into this new era of LLMs. Like building a bridge between these two powerful fields. Exactly. And the end goal, the big prize, is getting those more reliable and calibrated LLM-based systems. Right. Which could unlock their use in much more critical areas. Think about it. Legal tech, maybe more advanced medical diagnostics, managing critical infrastructure. Areas where right now you hesitate because an LLM might give you a very confident wrong answer.

7:51If you can actually trust the uncertainty estimate, know when the model knows it's unsure, that changes the risk calculation completely. It could open up applications we currently deem just too sensitive. Yeah, that makes perfect sense. Okay, so let's quickly recap then. This textual Bayes paper introduces a really fresh perspective. Treat prompts not just as commands, but as textual parameters. Parameters you can do Bayesian inference on. And they developed a practical algorithm, MHLP, using MCMC to actually do that inference. An algorithm that seems promising even for those tricky closed source models.

8:25Offering a real path towards quantifying and reducing uncertainty in LLMs, making them more reliable, more calibrated. Yeah, that's the core takeaway. Which leaves us with a pretty big thought to chew on, doesn't it? If we can genuinely start to quantify and trust an LLM's uncertainty, how does that change your fundamental relationship with AI, our trust in it? That's the question, isn't it? Once we move past the, is it right or wrong? And towards how confident is it? And can I trust that confidence? Huh? What does that unlock? What new possibilities, maybe even responsibilities, emerge when AI can reliably tell us not just what it thinks, but how sure it is?

From the publisher

This paper titled "Textual Bayes: Quantifying Uncertainty in LLM-Based Systems," available on arXiv. This paper addresses the critical challenge of quantifying uncertainty in large language model (LLM)-based systems, which is crucial for their application in high-stakes environments. The authors propose a novel Bayesian approach where prompts are treated as textual parameters within a statistical model, allowing for principled uncertainty quantification through Bayesian inference. To achieve this, they introduce Metropolis-Hastings through LLM Proposals (MHLP), a new Markov chain Monte Carlo algorithm designed to integrate Bayesian methods into existing LLM pipelines, even with closed-source models. The research demonstrates improvements in predictive accuracy and uncertainty quantification, highlighting a viable path for incorporating robust Bayesian techniques into the evolving field of LLMs.


More from Best AI papers explained

All 475 episodes
Textual Bayes: Quantifying Uncertainty in LLM-Based SystemsBest AI papers explained · 9 min
Listen in VO