Dual Goal Representations

14 Oct 2025 · 17 min · 9 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Dual Goal Representations (DGR) for goal-conditioned reinforcement learning (GCRL). The episode explains why feeding raw goal observations (e.g., goal images/sensor readings) adds exogenous noise (lighting changes, background motion, camera angle), hurting training efficiency and generalization. DGR instead defines a goal by temporal distance: for each state, the optimal shortest-path travel time to the goal, yielding a “noise-invariant” relational goal map.

Key claims

DGR is sufficient (provably retains information needed for optimal policy) and noise-invariant (ignores irrelevant observation changes like shadows).

Notable examples

discrete “Lights Out” puzzle and OG Bench tasks (13 state-based, 7 pixel-based). Results: on state-based tasks DGR beat prior representation-learning methods on average; some cube manipulation success rates improved ~3x; test-time Gaussian noise to goal images reduced performance gaps. Limitation: pixel-based visual puzzles sometimes reached 0% success due to late-fusion architecture versus early-fusion baselines.

Guests

No named guests are mentioned; it’s just the host(s) discussing the research.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding GCRL

0:45 to 1:56

Exploration of goal-conditioned reinforcement learning and its challenges.

“But there's this sticky point, a real bottleneck that keeps coming up.”

Introducing Dual Goal Representations

1:56 to 4:00

Discussion on the framework of dual goal representations and its key points.

“And that brings us to what we're really digging into today, this framework called dual goal representations.”

Shifting Perspectives on Goal Representation

4:00 to 6:07

How DGR focuses on temporal distances instead of raw features to define goals.

“And the philosophy behind this is actually pretty deep.”

The Challenges of Practical Implementation

6:07 to 7:48

Exploring the difficulties in computing dual goal representations in real-world scenarios.

“In a continuous space, like a robot arm moving freely, there are infinitely many states.”

Training the Distance Function

7:48 to 10:52

Details on the two-step training process for learning the DGR.

“much better, in fact, than what you might intuitively think of first.”

Evaluating DGR Performance

10:52 to 12:46

Results of DGR in various tasks, emphasizing its improved performance.

“Give the policy better goal description, get a better policy.”

Limitations and Architectural Trade-offs

12:46 to 14:03

Discussion on the limitations of DGR in visual tasks and the impact of architectural design.

“It really highlights how much that noise must have been hurting performance before.”

Exploring Architectural Trade-offs in RL

14:03 to 16:15

Learn about the trade-offs in representation learning for reinforcement learning.

“It seems counterintuitive if the representation is supposed to be better.”

Philosophical Insights on Goal Representation

16:15 to 17:11

Discussion on relational perspectives in defining goals for AI systems.

“And you mentioned it was inspired by that philosophical idea.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Welcome to the Deep Dive. We take complex research, pull it apart, and try to give you those core insights you need, like an instant knowledge boost. Yeah, hopefully. Today we're getting into something pretty cool in reinforcement learning. It's called goal-conditioned reinforcement learning, GCRL for short. Right. And the whole point of GCRL, it's, well, it's ambitious, right? You want to train one policy like a single AI brain that can handle lots of different tasks. Basically, get from any starting point to any target point and do it efficiently. Quickest time, lowest cost, that kind of thing.

0:38Yeah. Think about like robot navigation in a huge warehouse maybe. Or even manipulating objects, solving little puzzles with blocks. Exactly. But there's this sticky point, a real bottleneck that keeps coming up. It's all about how you actually tell the system what the goal is. Yeah. Most of the time right now, people just feed it the raw goal state. Maybe it's a picture of what the finished puzzle should look like. Or just raw sensor readings from the target location. Yeah, right. And that seems intuitive, but... But that's where the trouble starts. That raw data, that image, it's full of stuff the agent can't actually control.

1:10Stuff that doesn't matter for the task. Like what? Well, maybe the lighting changes in the room. Or there's, I don't know, some other robot moving in the background that's totally irrelevant. Or just the camera angle is slightly different from the training examples. Exactly. All that extra stuff, we call it exogenous noise. And it really messes with the learning, makes it way less efficient. Because the poor policy is trying to figure out if that flickering light in the corner is somehow important for solving the puzzle. Pretty much, yeah. It wastes a ton of effort trying to link these random background things to whether it succeeds or fails.

1:47so it takes longer to train and then it doesn't generalize well if the environment changes even a little bit. Okay, so that sets the stage. This is the problem we need to solve. Right. And that brings us to what we're really digging into today, this framework called dual goal representations. EGR. DGR. The goal is conceptually simple, but the idea behind it is clever. It wants to create the ideal goal representation, one that hits two key points. What are they? First, sufficiency. It has to hold on to all the information you actually need to figure out the best way to reach the goal. Okay, don't throw the baby out with the bathwater.

2:22Exactly. And second, noise invariance. It has to completely ignore, just filter out everything that's not relevant, all that exogenous noise we talked about. Sufficient but noise-free. That sounds like the dream goal representation. That's the idea. So how do you achieve that? If you're saying the raw image, the direct observation is bad, what do you use instead? This feels like a big shift. It is a fundamental shift. It's about changing perspective. Instead of describing what the goal looks like or is in terms of its raw features, we describe its relationship to everything else in the environment.

2:59Its relationship. How? We basically define a goal state by thinking about the set of temporal distances from all other states to that goal. Temporal distances. You mean like travel time? Exactly. How long does it take optimally to get from any state S to the goal state G? In a simple, discrete world, like a grid maze, the dual goal representation for a specific goal square would be a big list or a vector. And each number in that vector tells you the shortest path length from every other square in the maze to that goal square. Whoa. Okay. So instead of a picture of the exit, you're giving it like a map showing the travel time from everywhere to that exit.

3:40That's a great analogy. You're giving it the subway map focused on that one destination station. It tells you immediately how long it takes to get there from any other station. The optimal path info. Okay. I think I'm getting it. The dual goal representation isn't the goal itself. It's this intrinsic map of how hard it is to reach the goal from anywhere else. It's based on the environment's dynamic. Precisely. And the philosophy behind this is actually pretty deep. It connects to ideas, even in abstract math, like category theory, this notion that an object is kind of defined by how it relates to everything else.

4:12You know, dilemma, that kind of thing. Yeah, that sort of relational thinking. We're looking at the goal through its dual space of these temporal distance functions. OK, that's the core idea. Yeah. Now, you mentioned this approach has some big theoretical wins. You said three key benefits earlier. What were they again? Right. So benefit number one is invariance because this representation is built purely from shortest path times, from the actual physics or rules of the environment. Dynamics. The dynamics. Exactly. It doesn't matter how you look at the environment. Change the camera, change the lighting.

4:45The optimal travel times don't change. So the representation stays the same. It's invariant to the observation method. That makes sense. Benefit two. Benefit two, provable sufficiency. This is huge. They actually prove mathematically that this DGR definition keeps enough information to figure out the optimal policy. So you're not losing anything essential for solving the task. Correct. You're just repackaging the essential information in a better way. And the third one, the one we started with. Noise invariance. Because it's based on those intrinsic path lengths, it's mathematically guaranteed to be stable even if there's noise in the raw observation.

5:22Oh, so? Well, imagine you have two goal observations, G1 and G2. Maybe G1 is the cube in the right spot and G2 is the cube in the same right spot. But now there's, I don't know, a weird shadow over it. Okay. From the robot's perspective, the task, the required actions, the optimal path are identical. And because the DGR only cares about those path lengths, G1 and G2 will have the exact same dual goal representation. Ah, so the shadow, the noise just gets ignored by definition. It's like an automatic filter. Exactly. It filters distractors automatically because they don't affect the underlying dynamics.

5:56Okay. That theory sounds incredibly powerful. Perfect filtering, guaranteed sufficiency. But then there's the real world. Right. Always the catch. How on earth do you actually compute this? You said the representation is the distance from all other states. In a continuous space, like a robot arm moving freely, there are infinitely many states. You can't make that giant vector you talked about. You absolutely can't. Calculating it perfectly is impossible in most interesting cases. The ideal DGR is technically a function mapping that infinite state space to distances. You can't just store that. So what's the workaround?

6:34How do they make this practical? They use function approximation, machine learning to the rescue. Of course. Instead of calculating the exact distance function, DSG, they train a model to estimate it. This model usually takes two inputs. Okay. It takes features representing the current state S. Let's call that CSES. And critically, it takes a separate set of features representing the goal G. Let's call that phi-jaw. Ah, so phi-G is the practical, learnable version of the dual goal representation, a finite vector we can actually compute. Exactly. Phi-G is the thing we learn. It's a compressed, finite dimensional embedding that captures that essential relational information about the goal.

7:14The model then combines phi-jaws and phi to predict the temporal distance. So you learn two encoders, basically. One for the state, one for the goal. Pretty much. You learn feature extractors, yeah. And then you need a way to combine those features, say, she s, and phi to get that distance estimate. And how they combine them matters, I guess. Hugely. This was a key finding in the research. They tried different ways, and what worked really well was using the inner product. The dot product, so just psi s transpose times phi. Yep. that simple calculation, the inner product, turned out to be surprisingly effective, much better, in fact, than what you might intuitively think of first.

7:51Which would be like just calculating the distance between the future vectors, like Euclidean distance, FGAS2, that seems common. It is common, especially in representation learning. People often learn embeddings where distance means similarity. But think about temporal distance, about pathfinding. Okay. Is the cost or time to get from A to B always the same as getting from B to A? Definitely not. Think about driving up a steep hill versus driving down it, or like your cube example. Exactly. Rolling a cube off a table is easy. Rolling it back up may be impossible. Path costs are often asymmetric.

8:24Right. And a simple distance metric, like Euclidean distance, is always symmetric. The distance from A to B is the distance from B to A in that feature space. Precisely. But the inner product, the dot product, M-A-S dot 5-N-E-G, is not inherently symmetric. It can capture directional relationships. It respects the asymmetry of the underlying dynamics much better. Ah, that's a fantastic insight. The inner product allows the learned distance function to understand that going one way might be different from going the other. It aligns better with the physics or rules of the world. That's the key intuition.

8:59It makes the function approximator much more expressive for modeling these temporal distances. Okay, that makes a ton of sense. So we have the theory, the practical representation, phi, and this inner product idea. How do they actually learn this phi-g in practice? What's the training process? It's basically a two-step recipe. Step one is representation training. Okay. Here you take a big data set of trajectories, just lots of examples of the robot moving around, trying to reach various goals collected beforehand, maybe offline. Right, offline RL data. Exactly. And you use an offline RL algorithm.

9:32They specifically use a goal-conditioned version of IQRL, implicit cue learning, to train that distance predictor function, the one using the inner product,.psis,.phi g. The main goal here isn't to get the perfect distance value, but to force the network to learn good representations, especially that goal representation phi, that make predicting the distances possible. So learning the distance prediction task forces the network to embed the goals in a way that captures those temporal relationships. That's the idea. You learn phi implicitly by training it to predict optimal temporal distances. Okay, that's step one.

10:10Learn the goal representation. What's step two? Step two is policy training. Now you take your learned, hopefully noise invariant goal representation. If you freeze it, you don't train it anymore. Got it. And you plug it into basically any standard GCRL algorithm that trains the actual policy, the thing that decides the actions. They tried it with several, like GCVL, CRL, GCFBC. So the policy now takes the current state, S, or its features, PSY-S, and this nice, clean, dynamics-aware goal representation as input. Exactly. Instead of getting a raw, noisy image as the goal, the policy gets this pre-processed, information-rich feature.

10:48And the hope is that this cleaner input makes the policy learn much faster and better. That's the hypothesis. Give the policy better goal description, get a better policy. Right. Theory sounds solid. Practical recipe seems clever. Did it actually work? Let's talk results. Did this DGR approach deliver in practice? Yes, it really did. They started with some simpler discrete tasks first, like the lights out puzzle, places where they could actually compute the ideal DGR perfectly just to check the core concept. A sanity check. Right. And even there, using that perfect DGR massively sped up training and improved how well the agent reached goals compared to just giving it the raw state.

11:26So the fundamental idea held water. OK, good start. But the real challenge is the complex stuff, right? Yeah. The OG Bench benchmark. Exactly. OG Bench is designed to be tough. It's got navigation tasks like point masses or simulated ants in mazes. Ant mazes, yeah. Those are tricky. And also complex manipulation like moving cubes around, stacking them, rearranging objects in a simulated tabletop scene. Big variety. They tested on, I think, 13 environments where the state was given as coordinates or sensor values and seven pixel-based environments. Okay, quite comprehensive. What happened on the state-based tasks?

12:03Yeah. The ones without raw images. The results were pretty impressive. across those 13 state-based tasks when they paired their practical DGR approach with three different policy learning algorithms. The GCIVL, GCFBC you mentioned. Right. DGR consistently got the best average performance. It outperformed five other previous methods designed for representation learning in GCRL. So just by changing how the goal is represented, they got state-of-the-art results on average. Yes. And in some specific tasks, the improvement was dramatic. You mentioned the cube manipulation tasks earlier. On some of those, using DGR, boosted the success rate by like three times compared to the baseline that just used the raw state as the goal.

12:43Wow. A 3x improvement just from cleaning up the goal signal. That's huge. It really highlights how much that noise must have been hurting performance before. Absolutely. It shows the power of filtering out irrelevant information and focusing the policy on the actual dynamics needed to succeed. What about the noise invariance claim? Did they test that directly? The idea that DGR ignores distractions? They did. They specifically added Gaussian noise to the goal observations only during the testing phase, not during training. To simulate, like, unexpected sensor noise or visual changes when the robot is actually deployed.

13:19Exactly. Test time robustness. And the results confirmed the theory. DGR showed significantly better performance under this added noise compared to baseline methods, especially in challenging environments like the large amp maze and the complex scene manipulation task. So the noise invariance wasn't just theoretical, it showed up in the experiments too. It did. It held up really well. Okay, that sounds overwhelmingly positive. But you mentioned pixel-based tasks earlier. Was there a catch? Did it work as well when the input was raw images? Ah, yeah. That's where they hit a significant limitation.

13:52On the visual puzzle tasks within OG Bench ones requiring intricate visual understanding DGR, and actually several other representation learning methods too, didn't do well. Sometimes they got literally zero success. Zero. Why such a dramatic failure there? It seems counterintuitive if the representation is supposed to be better. It boils down to an architectural tradeoff, how the information flows through the neural network. Okay. See, in vision-based RL, often the best results come from something called early fusion. Early fusion. Yeah, you take the raw image of the current state and the raw image of the goal state, you stack them together, maybe channel-wise, and feed that combined input into one big convolutional network right from the start.

14:33Letting the network figure out the relationship between the pixels of the state and goal images together. Exactly. It gives the network the maximum raw context immediately. But methods like DGR, because they pre-compute the goal representation for you separately. In that first step. Right. They have to use what's called late fusion. The policy network processes the current state image through its own convolutional layers, and it takes the already processed goal vector feed, and it only combines or fuses this information later on deeper in the network. So it can't do that initial joint processing of the raw visual data.

15:08Correct. And it seems that for certain visually complex puzzle tasks, that early fusion approach used by the simple baseline, just feeding raw state and goal images, is crucial. and DGR's necessary late fusion architecture becomes a disadvantage. That's a really interesting trade-off. You gain the theoretical benefits of noise invariance and sufficiency from DGR, but potentially lose out on some visual processing power by being forced into late fusion. Precisely. It suggests that while DGR is powerful, there might still be work needed to integrate it perfectly with the demands of complex visual reasoning or perhaps develop hybrid approaches.

15:44It's the edge of current research. Okay, that makes sense. So let's try to wrap this up. If we zoom out, the big picture here is really about redefining the goal. Yeah, fundamentally. It's moving away from just showing the agent a picture of the destination. And instead giving it that map, that profile of travel times based on the world's own rules. A noise invariant map distilled down to the essential dynamics required for planning and acting. You get sufficiency for optimality, but you throw away the distracting noise. It's, well, it's elegant. It really is. And you mentioned it was inspired by that philosophical idea.

16:18An object is uniquely determined by its relations with every other object. Taking that relational perspective seems incredibly powerful here. Which leaves us, and you the listener, with a final thought to chew on. If thinking purely relationally, focusing on dynamics and travel times, works so well for defining goals for robots and filtering noise, where else could this apply? Right. What other really tough problems in AI or machine learning may be modeling really complex long-range dependencies or understanding abstract concepts where the raw inputs are super noisy and confusing? Could those problems also be tackled more effectively?

16:53By shifting our focus. By moving away from describing what things are in isolation and instead concentrating on defining them by how they relate to everything else in their world. It's a compelling question. Definitely something to think about. Indeed. And that's all the time we have for this deep dive. Thanks for joining us.

From the publisher

This paper discusses dual goal representations for goal-conditioned reinforcement learning (GCRL), a novel method for encoding a state based on its temporal distance relation to all other states within an environment. The authors theoretically establish that this representation is sufficient for recovering an optimal goal-reaching policy and is invariant to extraneous noise within the state observations. Building on this theory, they propose a practical implementation using an inner product parameterization and offline value learning, demonstrating that this approach consistently improves goal-reaching performance across a suite of robotic navigation and manipulation tasks, outperforming existing representation learning methods. The overall aim is to enhance the efficiency and generalization capability of GCRL agents by providing a robust and structured goal representation.

More from Best AI papers explained

All 475 episodes
Dual Goal RepresentationsBest AI papers explained · 17 min
Listen in VO