When Does Trajectory-Level Supervision Permit Efficient Offline Reinforcement Learning?

27 Jun 2026 · 19 min · 11 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Episode topic: When offline reinforcement learning can learn efficiently using only trajectory-level feedback (final outcomes or preferences) instead of step-by-step rewards, and when it becomes impossible.

Guest backgrounds

No guests are named; the episode is presented as a host-led discussion.

Key claims

Replacing process-level rewards with a single final label increases sample complexity by a factor of the horizon H (even in deterministic environments) because information about individual steps is compressed away. OPAC (Outcome-Based Pessimistic Actor-Critic) learns a hidden stepwise reward model pessimistically from final outcomes. The same leading H penalty holds for preference feedback using the Bradley-Terry-Luce model (RLHF). But for generalized trajectory criteria like “all-success” (multiplicative/binary success), learning becomes effectively unlearnable offline, requiring exponentially many trajectories (omega(2^H)).

Notable examples

Hospital supply-chain management with quarterly “within budget” success; “thumbs up/down” preferences; pitch-black maze/all-success where one mistake yields zero.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Importance of Step-by-Step Feedback

1:00 to 2:20

Learn why step-by-step feedback is crucial in AI systems.

“And to understand why this matters to you, whether you're just casually tracking the AI industry or you're fascinated by how complex systems actually learn, we really should frame this around a real-world scenario.”

Understanding OPAC: A New Learning Algorithm

2:20 to 4:00

Discover how the OPAC algorithm approaches learning in AI.

“We are going to unpack the exact statistical cost of losing that step-by-step feedback.”

The Impact of Missing Feedback

4:00 to 5:00

Examine the mathematical implications of losing feedback in learning.

“Well, the way OPAC approaches that soup recipe is through the pessimistic part of its name.”

Deterministic Environments and Learning Challenges

5:00 to 7:00

Discuss the challenges of learning in predictable systems.

“Because I'm imagining a totally predictable hospital system, like no random dice rolls, no sudden pandemics, just strict cause and effect.”

Learning from Preference-Based Feedback

7:00 to 8:40

Understand how preference-based feedback affects AI learning.

“if adding up points costs an extra factor of each, what happens in the real world where we often don't even have points to begin with?”

The Exponential Wall in Learning Objectives

8:40 to 10:00

Investigate the exponential wall in complex AI learning tasks.

“Finding the gradient of improvement is all that matters mathematically.”

The All-Success Objective Problem

10:00 to 12:20

Learn about the difficulties of achieving all-success objectives in AI.

“Up until now we've been talking about cumulative rewards, right?”

Crucial Coefficients for Learning Success

12:20 to 14:03

Explore the two coefficients that impact AI learning in complex scenarios.

“You learn absolutely nothing about the boundary of success.”

Understanding Signal Propagation in AI

14:03 to 16:14

Learn how reinforcement learning algorithms propagate signals and the challenges involved.

“Kappa is basically measuring whether the evidence still exists at the crime scene after the fact.”

The Importance of Data in AI Training

16:14 to 17:08

Explore the significance of data collection methods in AI training to avoid failures.

“It really breaks down exactly where the failure points are when we try to train these systems.”
Show all 11 chapters

Philosophical Implications of AI Feedback

17:08 to 18:52

Discuss the parallels between AI feedback challenges and human psychological experiences.

“It puts the responsibilities squarely back on the humans designing the environments.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00You know the feeling. I mean, it is this uniquely frustrating, totally universal human experience. You get a test back, you look at the top of the page, and there's just a giant red C -. Oh, yeah. The dreaded C -. Right. And there are absolutely no other markings on the paper. Nothing. You're just staring at it, scanning the pages, thinking, like, where did I actually go wrong? It's maddening. It really is. Did you mess up the long division on question three? Did you completely misunderstand the essay prompt on page two? When you only get that final grade, figuring out the exact step where things went off the rails feels, well, nearly impossible.

0:39You are essentially left guessing in the dark. You really are. And it is a terribly frustrating experience for a student. But, you know, that exact frustration, that specific lack of step-by-step feedback is actually one of the biggest multibillion-dollar bottlenecks currently facing artificial intelligence development. Yeah, and that's why we are jumping straight into the deep end today. we are focusing on the world of offline reinforcement learning. And to understand why this matters to you, whether you're just casually tracking the AI industry or you're fascinated by how complex systems actually learn, we really should frame this around a real-world scenario.

1:16That's a great idea. So let's say you want an AI to manage a hospital's entire supply chain. Okay, high stakes. Very high stakes. Everything from ordering surgical gloves to managing life-saving medication inventory. Usually when AI learns in a highly controlled lab environment, like, say, a video game, it gets what the research calls process-level rewards. Right. And process-level feedback just means the AI gets a distinct grade after every single action it takes. So order the right amount of bandages. Plus one point. Forget to order saline. Minus one point. It's instant. Exactly. It is continuous, granular, immediate feedback.

1:52But a real hospital does not work like a video game. In the real world, the AI makes, I don't know, a thousand little inventory decisions over the course of a month, and it only gets one single final metric at the end of the quarter. Right. Did the hospital stay within budget while meeting patient needs? Yes or no? Exactly. And that is what the study calls a trajectory-level outcome. The system made a thousand moves, but it only gets that single C-minus at the top of the page. Which brings us to our mission today. We are going to unpack the exact statistical cost of losing that step-by-step feedback.

2:27Because it's not a small cost. Not at all. Replacing those granular rewards with one final outcome changes the fundamental math of how a system learns. And the research we are looking at proves that there are specific scenarios where losing that feedback doesn't just, you know, slow the AI down. It actually creates an insurmountable mathematical barrier. An impossible wall. So the first thing the study does is isolate that hidden cost. If we are forcing an AI to run a hospital supply chain using only quarterly performance reviews, we really need to know exactly how much harder its job just became.

3:00Right. And to figure this out, the researchers built a standard AI training setup aiming for a high cumulative score, but they completely hid the step-by-step points from the model. So it's flying blind during the process. Exactly. All the algorithm gets to see is the final label at the very end of the run. And to navigate this, they introduce an algorithm called OPAC. OPAC. Outcome-Based Pessimistic Actor-Critic. That's the one. I want to break down what OPAC is actually doing mechanically, because reading through this, it feels like trying to figure out the exact ingredients of a complex soup just by tasting a single spoonful of the final product.

3:38I like that, the soup analogy. Right. Because if the chef hands you a bowl of minestrone, you have to taste a lot of different variations of that soup to mathematically isolate, like, the flavor of the thyme versus the oregano. And OPAC is essentially trying to learn a hidden step-by-step reward model, a recipe using nothing but those final soup tastings. Exactly. But how does it actually do that? Well, the way OPAC approaches that soup recipe is through the pessimistic part of its name. Meaning it assumes the worst. Basically, yeah. When an AI doesn't know the value of a specific step, say, it has absolutely no idea if adding a certain spice was good or bad because it only saw the final soup, it assumes the worst.

4:20It artificially assigns a low value to unknown steps. Okay, so it's actively avoiding overconfidence. If it isn't absolutely sure a decision was good, it treats it as a bad decision until proven otherwise with more data. Exactly. It forces the AI to systematically stick to what it already knows works. It only ventures into unknown decisions when it has enough overlapping data to confidently update its internal recipe. That makes sense. And by doing this, the researchers were able to calculate a mathematically sharp finding. Yeah. They proved, with matching upper and lower bounds, that replacing step-by-step rewards with one final label costs exactly one extra factor of H.

4:57Wait, H being the horizon, right. Like, the total number of steps in the task. Yes, the length of the sequence. So if the hospital supply chain requires 100 decisions to reach the end of the quarter, it is mathematically a factor of 100 times harder to learn from the final quarterly review than it is to learn from daily step-by-step feedback. OK, wait, hold on. Let me push back on this for a second. Sure. Because I'm imagining a totally predictable hospital system, like no random dice rolls, no sudden pandemics, just strict cause and effect. If the rules of the environment are deterministic, meaning no random chance at all, shouldn't the AI easily figure out exactly which step went wrong?

5:38You would think so. Right. Because if I take the exact same actions, I get the exact same result. Why are we still paying this massive penalty factor of 100? It feels super counterintuitive, I know. But the researchers actually used a perfectly deterministic environment to prove their lower bound. Wait, really? Yeah. They showed this penalty exists even in the most predictable clockwork worlds imaginable. Because the extra difficulty isn't about random environmental noise confusing the AI. It is the pure mathematical price of information compression. Information compression. So cramming all those individual decisions into one single sum.

6:14Exactly. When you take a hundred separate reward signals and compress them into one scalar outcome, the information about the individual steps is literally erased through the active addition. Oh, I see. Think about it like this. If I tell you two numbers add up to 10, those numbers could be 5 and 5, right? Sure. Or 9 and 1. Or under negative 90. Even if the environment's perfectly deterministic, figuring out the exact contribution of step 42 versus step 89 requires massively more data. You have to observe that final sum across a huge number of different overlapping projectories to statistically isolate the variables.

6:53Wow. The act of adding it all up literally destroys the evidence of how you got there. Exactly. Okay, that makes perfect sense. But, and this is a big but, if adding up points costs an extra factor of each, what happens in the real world where we often don't even have points to begin with? Because I'm thinking about, you know, large language models. Humans don't give an AI an exact score of 87.4 for writing an email. No, we definitely don't. We just give a thumbs up or a thumbs down. We say, I preferred draft A over draft B. The thumbs up economy. This brings us right into reinforcement learning from human feedback or RLHF.

7:28Right. So the researchers extended their OPAC algorithm to see what happens with preference based feedback. They relied on the Bradley Terry Luce model, the BTL model, which is the foundational math for evaluating pairwise preferences. OK. And you would think losing an absolute score would make things drastically worse. Yeah, I would. But surprisingly, the leading penalty remains exactly the same. Wait, what? The mathematical guarantee preserves that same leading horizon penalty. It doesn't just blow up. It doesn't. The dependence on the horizon, H, doesn't fundamentally change just because you switched from a hard score to a preference.

8:04Right. You don't actually need calibrated exact scalar outcomes to learn efficiently. I am struggling with that a bit, I'm not going to lie, because a preference tells you even less than a score. If I just say meal A was better than meal B, how does the AI know if meal A was actually a delicious, perfectly cooked filet mignon or if both meals were totally burnt toast and meal A was just slightly less burnt? A preference gives you zero baseline. It gives no baseline, you're right. But uniform shifts in absolute value, whether both meals are great or both are awful, simply do not matter for optimizing behavior.

8:39Meaning, the AI doesn't need to know if the meal is Michelin star quality, it just needs to know the trajectory to make it less burnt. Exactly. Finding the gradient of improvement is all that matters mathematically. The algorithm's only job is to select the action that leads to the better outcome. As long as the AI can learn the relative differences between the two trajectories, it can find the optimal policy. Oh, wow. The mathematical engine of pessimistic offline control works beautifully with sufficiently informative comparisons. It focuses entirely on the gap between the options. So as long as the AI knows which direction the mountaintop is, it can keep climbing.

9:16It doesn't need to know its exact altitude above sea level. That is the perfect way to visualize it. And it just pays that same extra factor of age to figure out the delta. Exactly. Okay, up to this point, this research sounds incredibly optimistic. It sounds like AI can always get away with just having final outcomes or vague human preferences, provided we give it a bit more data to handle that factor of H. Right. But. We have to look at the sharp turn the paper takes next because the math completely breaks down when the final outcome isn't just an accumulation of points but a totally different nonlinear rule.

9:52Yes. This is where we hit the exponential wall. The exponential wall. The paper examines what they call generalized trajectory level criteria. Up until now we've been talking about cumulative rewards, right? Adding up the points from step one, step two, step three. Right. But what if the rule for success isn't additive? The deadliest example they focus on is the all-success objective. Okay, let's bring this back to the hospital supply chain. An all-success objective would be a scenario where the AI only gets a one at the very end of the quarter if every single surgical department gets exactly what they need.

10:27Right. If it messes up even one delivery to one department, the entire quarter is deemed a failure and it gets a zero. Exactly. Because the rewards at each step are binary, a 1 or a 0. And instead of being added together, they are multiplied. Oh, wow. Yeah. 1 times 1 times 0 is 0. And the research proves that this specific problem is, generally speaking, not learnable offline. Any offline learner might require omega of 2 to the power of 8 trajectories. The amount of data needed scales exponentially with the length of the task. Exponentially. So if the hospital task requires 100 steps, you don't need 100 times more data.

11:06You need 2 to the power of 100 times more data. Which is a number so large it's practically impossible to compute. Even with perfectly deterministic transitions and constant coverage of the data, the mathematical barrier is absolute. Let me use a really dark analogy here just to make sure the mechanics of this exponential wall are totally clear to you, the listener. Go for it. Imagine you are navigating a pitch black maze. If you step on a trap on step 2, you die. You get a zero. Right, you get a zero. Now, if you navigate perfectly for 98 steps, but step on a trap on step 99, you die. Still a zero.

11:40In both cases, the universe just gives you a game over screen. You have absolutely no clue which step killed you. And because one wrong move invalidates the entire run, there is no partial credit. It's brutal. And in an all success scenario, the final outcome label is informative on only an exponentially small set of trajectories. Almost every single trajectory the algorithm attempts is going to result in a zero. Right. And a zero tells you almost nothing about where you failed, just that you failed somewhere. The data is incredibly sparse. It is a desert of information. I mean, if you need 100 correct steps and one mistake is a zero, then 99.99 % of your data points are just zeros.

12:23You learn absolutely nothing about the boundary of success. It creates this insurmountable statistical barrier because the compression of information is just too extreme. The final grade obscures the entire path. Okay, so if all success objectives are practically impossible to learn, how do we ever train AI for complex, high-stakes tasks? I mean, a self-driving car doesn't get partial credit for almost stopping at a red light. Are we just doomed? How do we find a tractable path forward without hitting this exponential wall? We aren't entirely doomed. The paper actually identifies a specific learnable regime.

12:58They mathematically prove that you can solve these complex nonlinear goals, but only if the task respects two specific structural coefficients. Okay. These coefficients basically measure the information loss happening inside the system. Right. First one is the reward process coefficient, denoted by the Greek letter kappa. And the second one is the Bellman inverse coefficient, denoted by the Greek letter chi. Okay, whoa. We need to slow down and really untangle these concepts because this feels like the crux of the whole paper. Why do we need two different coefficients? Doesn't a bad final grade just obscure everything all at once?

13:30It seems like it, but it actually reflects the two very different jobs the AI has to do behind the scenes. Okay, let's break down kappa first. Sure. Kappa, the reward process coefficient, is entirely about the raw data. It measures how much information is lost when the raw step-by-step events are squashed into that final outcome. It essentially asks, does the raw data even contain the necessary signal after it's been aggregated? So in the pitch black maze, Kappa would be terrible because a game over at step two looks mathematically identical to a game over at step 99. The signal is completely lost in the data itself.

14:05Kappa is basically measuring whether the evidence still exists at the crime scene after the fact. That is a brilliant way to put it. Yes. Yeah. That covers the data side. But then we have Che, the Billman inverse coefficient. And this is about the dynamic programming side. It captures how well our mathematical models can actually propagate that signal once we have it. Stop right there. Dynamic programming. Propagating the signal. Explain this to me without the jargon. Like, what is the algorithm actually doing? Okay, fair enough. In reinforcement learning, the algorithm doesn't just guess randomly.

14:37It works backward from the goal using something called a Bellman update. Working backward. Right. Imagine building a bridge backwards from the destination to the starting point. The algorithm looks at the final reward and then calculates the estimated value of the step right before it. Then it uses that estimate to calculate the value of the step before that and so on, cascading backward all the way to step one. Okay, I am picturing the AI retracing its steps from the finish line back to the start. Exactly. Yeah. But every time it takes a step backward, it performs a mathematical operation. Chi asks a terrifying question.

15:11When we run generalized Bellman targets backward, does the mathematical operation itself wash out the reward differences? Oh, wow. It's like a mathematical game of telephone. If you make a tiny estimation error at step 100 and you multiply that error at step 99 and then again at 98, by the time you get back to step one, the original signal is completely unrecognizable. It's just noise. Precisely. The algorithm's own internal math accidentally blurs the signal as it calculates backwards. So there are really two distinct inverse problems the algorithm must solve. Right. First, can you recover the hidden reward from the final label?

15:48That is kappa. Second, can you successfully propagate that recovered reward backward through the chain of decisions without the signal deteriorating into noise? That is Che. And if the answer to both is yes, if the data survives the compression and the math survives the backward propagation, then the AI can learn efficiently without needing 2 to the power of 100 data points. A generalized version of OPAC actually achieves polynomial sample complexity again. Math holds. That is incredibly elegant. It really breaks down exactly where the failure points are when we try to train these systems. Because if we zoom out and look at the bigger picture, as AI takes on managing cower grids or writing complex software or running hospital supply chains, we are not going to be able to hold its hand step by step.

16:33No, we simply can't. We will have to rely on trajectory level outcomes. We can't give a red checkmark on every single line of code. We can only tell it if the software compiled and ran successfully. And that's why understanding the exact statistical price of these final grades is so crucial. It tells us what kind of data we actually need to collect in the real world. It tells engineers when they can afford to just use a thumbs up or a thumbs down, and when they absolutely must invest millions of dollars into gathering granular, step-by-step process data. Because if they face an all-success wall without that granular data?

17:08The model will simply fail to learn. It puts the responsibilities squarely back on the humans designing the environments. We have to structure the feedback intelligently. We really do. You know, as we wrap up this deep dive, I want to leave you, the listener, with a bit of a philosophical pivot. Oh, I like where this is going. We have spent all this time exploring the cold, hard math of how artificial intelligence struggles with sparse feedback. But think about what this means for human psychology. The parallels to human experience are definitely there. Right. If sophisticated mathematical models struggle exponentially with all or nothing outcomes, unless the data is structured perfectly, what happens when we experience a catastrophic all or nothing failure in our own lives?

17:52Yeah. Like a 10-year relationship ends abruptly or a business you built completely collapses. You just get a giant zero from the universe. Yeah. A pure trajectory level failure. And you don't get an itemized list of which specific conversations or minor decisions over the last decade led to that failure. Exactly. You get nothing but the final grade. So do we inherently have to generate our own internal step-by-step pseudo rewards? Do we obsessively replay the past, analyzing every single interaction, trying to assign a tiny plus one or minus one to our actions just to ensure we don't need a million lifetimes to learn from our mistakes?

18:29We essentially act as our own pessimistic algorithm. We try to infer the hidden reward model from the final devastating outcome, assuming the worst about our unknown missteps until we can prove otherwise. We have to, because if we just accepted the final C - at the top of the page without trying to reconstruct the red check marks ourselves, we would be stuck in that pitch black maze forever. We really would. Something to chew on the next time you are replaying a five-year-old conversation in the shower.

From the publisher

This paper discusses a statistical framework for offline reinforcement learning using trajectory-level supervision, where only final outcomes or preferences are observed rather than step-by-step rewards. The authors introduce OPAC, a pessimistic actor-critic algorithm designed to learn from these aggregated signals by estimating latent rewards and applying pessimism to account for distribution shifts. Their analysis establishes that moving from process-level to outcome-level feedback incurs a quantifiable statistical cost, specifically an additional horizon factor in sample complexity. The research also explores generalized RL objectives, proving that non-linear outcomes like "all-success" criteria can lead to exponentially difficult learning problems. To address this, they identify specific structural coefficients, $\kappa_\mu(\sigma)$ and $\chi_\mu(\sigma)$, which determine when efficient learning remains possible. Ultimately, the paper provides a theoretical boundary for when sparse, trajectory-based data can successfully guide sequential decision-making.

More from Best AI papers explained

All 475 episodes
When Does Trajectory-Level Supervision Permit Efficient Offline Reinforcement Learning?Best AI papers explained · 19 min
Listen in VO