In short
Latent Space: The AI Engineer Podcast - Episode Summary
Episode Title
[NeurIPS Best Paper] 1000 Layer Networks for Self-Supervised RL — Kevin Wang et al, Princeton
Episode Description This episode features Kevin Wang, Ishaan Javali, Michał Bortkiewicz, Tomasz Trzcinski, and Benjamin Eysenbach, the authors of a groundbreaking paper that won the Best Paper award at NeurIPS 2025. Their research defies conventional wisdom by successfully scaling reinforcement learning (RL) networks to 1,000 layers, achieving performance gains that were previously thought to be impossible. The discussion dives into their approach to self-supervised RL, architectural innovations, and the implications for robotics and AI deployment.
---
Key Topics Discussed
Introduction
- Overview of the paper and the excitement around winning the Best Paper award at NeurIPS.
- Team introductions and the origins of their research at Princeton.
Challenges in Reinforcement Learning
- Deep networks have historically performed poorly in RL, remaining shallow (2-4 layers) for over a decade.
- Kevin Wang's skepticism about the efficacy of deeper networks, reflecting the community's long-standing beliefs.
Self-Supervised Reinforcement Learning (RL)
- Shift from traditional value-based RL to self-supervised learning.
- Objective: Learn representations of states and actions, pushing similar representations closer together, and dissimilar ones apart.
- This contrasts with traditional RL that focuses on maximizing rewards, which can be noisy and biased.
Key Discoveries and Breakthroughs
- Initial attempts at scaling depth resulted in degraded performance.
- Introduction of residual connections and layer normalization drastically improved results, unlocking the "critical depth" phenomenon.
- Scaling depth is more parameter-efficient and sample-efficient compared to scaling width.
Architectural Innovations
- The use of deep networks (1,000 layers) yields linear growth in parameters, while width scaling leads to quadratic growth.
- The architecture borrows from established concepts (e.g., ResNets) but applies them in a novel context.
Data Efficiency and Collection
- The use of Jax and GPU-accelerated environments allowed for rapid data collection (hundreds of millions of transitions in hours).
- Emphasis on the importance of data abundance to achieve significant performance improvements.
Implications for Robotics
- Potential for goal-conditioned RL without human supervision, making it scalable compared to traditional imitation learning methods.
- The architecture can facilitate more efficient data collection, avoiding the need for extensive manual data curation.
Future Directions
- Exploration of distillation: training large networks and then distilling them into smaller, efficient models for deployment.
- Potential applications in various industries, particularly in robotics and autonomous systems.
- The blurring of lines between RL and self-supervised learning, suggesting new pathways for developing intelligent systems.
---
Key Takeaways
- Scaling Depth: Achieving significant performance improvements in RL requires not just bigger networks but also innovative architectural techniques.
- Self-Supervised Learning: Transitioning from reward-maximization to representation learning can enhance scalability and efficiency in RL tasks.
- Data Collection Innovations: Leveraging modern computational power allows for unprecedented data collection, vital for training deep networks.
- Future of RL: The methods presented in the episode suggest a fundamental shift in how RL can be approached, akin to advancements seen in NLP and computer vision.
---
Closing Thoughts The episode concludes with reflections on the paradigm shift in reinforcement learning brought forth by the team’s research, emphasizing the potential to revolutionize AI applications, particularly in robotics. The speakers express excitement for future research directions and the ongoing evolution of AI and RL methodologies.
---
For more information and detailed show notes, visit [Latent Space](https://latent.space).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOCelebrating the Best Paper Award
0:45 to 2:24
Kevin shares his excitement about receiving the best paper award at NeurIPS.
“And yeah, I guess I led the project, started the project.”
Introducing the Research Team
2:24 to 5:04
Kevin introduces himself and his collaborators, detailing their backgrounds.
“When Kevin and Sean actually wanted to try really deep networks, I was kind of skeptical it was going to work.”
The Challenge of Deep Networks
5:04 to 7:50
Discussion on skepticism about deep networks and their application in reinforcement learning.
“is scalable in these different areas in deep learning.”
Scaling Reinforcement Learning
7:50 to 10:49
Exploring the scaling potential of self-supervised reinforcement learning algorithms.
“I think the thing I might clarify about the paper is, I think a lot of people reading the title are like, wow, big networks, they're great.”
Key Insights on Architectural Choices
10:49 to 13:20
Importance of architectural components like residual connections in improving performance.
“that we've seen is that it seems like by approaching RL in this different approach.”
The Intersection of Learning Methods
13:20 to 14:00
Discussion on how the research blurs the lines between reinforcement learning and self-supervised learning.
“I'm not familiar with the pre-existing literature.”
Model Depth vs Width: Performance Trade-offs
14:00 to 14:44
Learn about the differences in performance and parameter scaling when using depth versus width in neural networks.
“the number of parameters that your model has is going to grow roughly linearly.”
Bottlenecks in Reinforcement Learning
14:44 to 15:50
Discover how the structure of neural networks affects training time and performance in reinforcement learning (RL).
“And in general, of course, like more parameters is also going to be more expensive.”
Using JAX for Efficient Data Collection
15:50 to 17:48
Understand the advantages of using JAX for collecting large datasets in reinforcement learning environments.
“And making four passes through our network may not be the bottleneck.”
Scaling Data for Reinforcement Learning Success
17:48 to 19:29
Explore how scaling data collection can improve performance in reinforcement learning and its parallels with language models.
“And so I think that this serves as a really good test bed for us to be able to also find ways to scale up network capacity and get similar kind of gains.”
Show all 13 chapters
Insights from Language Models to RL
19:29 to 22:00
Learn about the potential parallels between language model training and reinforcement learning objectives.
“Did you get my meaning about the role model stuff?”
Future Directions in Reinforcement Learning Research
22:00 to 24:04
Find out about future research directions in reinforcement learning, including efficiency and scaling models.
“Let's talk about other future directions.”
Exploring Action Models and Robotics
24:04 to 28:00
Discuss the exploration of action models in robotics and the current trends in representation learning.
“what kind of compute budget did you have?”
Transcript
Automatic transcript. May contain errors.0:12Welcome to Lanespace. We are basically trying to provide the best optimal sort of podcast experience of Europe's for people who are not here And congrats on your paper. How does it feel? Yeah, it was very exciting Yeah, we had a poster yesterday and then today we'll have an oral talk Were you just like mobbed? Oh yeah, there was a lot of people. It's like three hours straight of like, you know, like waves of people to like That we were trying to do So I've never received the best paper Did you just find out on the website? Like what? I just like woke up one day and like checked my email and then they just they just like they was like oh like I saw an email you've been awarded best paper but maybe you know from the reviews as well right yeah we know from the reviews that we did well but there's a difference between like doing well in the reviews and getting best paper so that part we didn't actually know yeah okay so I skipped a little bit maybe we can go sort of one by one and sort of introduce you know who you are and what you did on the team.
1:11I'm Kevin. I was an undergrad from Princeton. I just graduated. And yeah, I guess I led the project, started the project. And then I was very happy to collaborate with Ishan and Nicole and Ben also. Right. And were you in the same research group? How do you... What's your social context? So yeah, so we're all from Princeton. Yeah. Thanks to Alan for booking you guys. So this project actually started from an IW seminar. so like an independent work research seminar that Ben was teaching and this was like actually like like one of my first experiences in like ML research so it was really valuable to like get that experience and then Ishan was also in that seminar and working on adjacent things so we collaborated a lot during that seminar and then yeah the project turned out to have some pretty cool results and then later on also like the Halt working on sort of similar things also joined it on the project and became like a good collaboration Yeah, and I don't know if any of you guys want to chime in on other elements of coming into deciding on this problem.
2:14So it's like probably my lab works on deep reinforcement learning, but historically deep meant like two or three or four layers. Not 1 ,000. When Kevin and Sean actually wanted to try really deep networks, I was kind of skeptical it was going to work. I've tried this before, it doesn't work. Other papers have tried this before and it doesn't going to work. so I was very very skeptical starting out I don't know if I conveyed this at the time but that was my prior going in because but do you view your job as like screening or like hey guys this is probably isn't going to work you should try a different idea you know like or should you be encouraging even if it's dumb it's selecting bets yeah and this was a bet I was willing to make what what made you willing to make a bet it seemed relatively low cost uh in that we Michal I had spent the past year developing infrastructure that made it a lot easier to run some of these experiments.
3:08The precedent was deeper networks should do a whole lot better. That's what the deep learning revolution has been over the last day. Yeah, I know. Why do we stop making them deeper? And reinforcement learning was like this one anomaly where we continue to use these really shallow networks. That's particularly true in the settings that we were looking at, where you're starting from scratch, you're starting from nothing. Any other perspectives you guys want to chime in with? I guess maybe I should just go over an overview of our project? Yes. Okay, sorry. Yes. So the way that I kind of view our project is that if you look at the landscape of deep learning, you have NLP, language, vision, and then RL.
3:44And as Ben kind of alluded to, in language, in vision, we've sort of converged to these paradigms of scaling to massive networks, right? Like hundreds of billions of parameters, trillions of parameters. And there's been a lot gained in deep learning from that. But then it seems like in the third sort of branch of deep learning in deep RL. That has not yet been the case. Like I was very surprised like coming into some, like, you know, Ben's class and seminar when I was looking at the networks. Oh, why were you just using like a simple two layer MLP for like these frontier sort of, you know, state of the RL algorithms?
4:16And so I was very curious, like, can we design RL algorithms? Can we sort of put together a recipe for RL that can allow it to scale in potentially, you know, analogous ways that language envisioned my scale? And so what we did is that we know that traditional RL, like let's say value-based RL, doesn't really scale. This is pretty clear from the literature. So we tried a different approach with RL called self-supervised RL, where instead of learning a value function, we're learning representations of states, actions, and future states, such that the representations along the same trajectory are pushed together, the representations along different trajectories are pushed apart.
4:51And this is just a different approach to RL that allows us to learn in a self-supervised manner. So we can solve task-reach goals without any human crafted reward signal. And so we know that self-supervised learning is scalable in these different areas in deep learning. So can self-supervised RL scale in similar ways? When we first tried it, it actually didn't work. Like we've made the networks deeper. The performance like totally degraded. But then we also, but then I separately was like, there's also some other work like in our literature, like we tried like residual connections and it is a few other architectural components that we had to put into the recipe.
5:29And then all of a sudden, one day, I ran this experiment and there was this one environment in which there was going from doubling the depth didn't really do anything, but doubling the depth again with these different components suddenly skyrocketed performance in this one environment. Getting this to work was very non-trivial in the sense that usually when you think about doing hyperparameter optimization, we try changing A, see if it makes it better, try changing B, see whether it makes it better. And if we just made the depth bigger, makes it worse. We guess that residual connections didn't make it better.
6:01And it was really this combination of factors that Kevin and Yishan figured out that really made this work. And as a precursor to that, we also tried scaling along different dimensions. So scaling the batch size, scaling the width of the network, so the hidden layers. Yeah, pretty much kind of similar to just scaling depth naively. And then once we started introducing residual connections, layer norm, these specific architectural choices, that's when we saw these significant jumps in performance, like these critical depths at which performance multiplies by a pretty huge factor. And that's where we really noticed unlocking some significant performance gains as opposed to scaling just along, which did yield some performance improvements.
6:41But when you look at the number of parameters that your network has as you grow width, it's roughly a quadratic as opposed to something like growing depth. So it's more, in some senses, more parameter efficient, also more sample efficient from the experiments that we conducted. Nice. In some ways, you're sort of replicating stuff that is seen in the wild, but on a very small model that you can study. Would you say that? Yeah, so I can add to what Kevin said earlier. We saw these huge performance improvements in language models, image generation models by making them larger, making them deeper, which seems very intuitive.
7:13Yeah. And so that's why our work, we draw from foundational research, right? Like residual networks, which employ residual connections to avoid vanishing gradients. And that's something that we show in some of our ablations in our paper, further down, probably in the appendices, where we did experiments without these residual connections. And so it's sort of borrowing these concepts that have existed in other fields and applying them to this setting with RL and showing that it works. Before Ben has to go, I'll leave the sort of last word to him. What additional work does this inspire that you want to push on next?
7:49I think there's one thing I'd clarify about the paper and then I'll directly answer the question. I think the thing I might clarify about the paper is, I think a lot of people reading the title are like, wow, big networks, they're great. I'll take big networks. You solved it now. We can just go. Yeah, we'll take big networks, add them to PPO, add them to SAC, add them to your favorite reinforcement learning algorithm. But I think that's actually not the main conclusion. I think the main conclusion is that using big networks not only requires these architectural tricks, but also, as Kevin mentioned before, it requires using a different objective.
8:17This objective doesn't actually use rewards in it. And so there's another word in the title, reinforcement learning, that also might be a little bit of a misnomer. Because we aren't directly trying to maximize rewards. Our code doesn't have a line of code saying maximize rewards here. And so is, at the end of the day, this a reinforcement learning method? I don't know. It looks much more similar to the self-supervised methods in other areas of machine learning. And so I think that the method and the work really stands in some sort of interesting intersection of reinforcement learning and self-supervised learning research.
8:51And we had this little figure on the bottom left of the poster, which was a screenshot of a slide from Yen Le Koon talking about how to build intelligent systems and whether that's going to be done by unsupervised learning or supervised learning and reinforcement learning. And I think what our paper really suggests is that the boundary between these things is really blurry. And maybe the keys to building intelligent systems are going to be leveraging insights from all of them. Yeah, the layer kick. Exactly. Well, thank you for your time. I know you have to go soon. Yeah, thank you so much for coming.
9:24I think that the insight of blurring things is interesting. I don't know if you were talking about the abstraction layer of representation learning. I don't know if that triggers anything in terms of the mix between self-supervised and reinforcement learning. Is that something fundamental that you've discovered or that people don't understand when they read the paper? Yeah, I think the best way that I would explain it is that we know that standard RL is not super scalable. And so why can this different approach or different objective RL be scalable? I think it's because we're fundamentally shifting the burden of learning from something like Q-learning or regressing to TD errors, which we know is quite spurious and noisy and biased, to fundamentally a classification problem.
10:09We're trying to classify whether a future state is along the same trajectory or along a different trajectory. And we do this with representation learning. And we know that classification, cross-entry loss, and representation learning is scalable in the deep learning literature. if we think about language and some of the objectives there. So in some sense, we're kind of blurring the lines. We're doing reinforcement learning. It's still an actor-critic reinforcement learning algorithm. It's like a goal-conditioned reinforcement algorithm. But the objective, the burden of learning, of solving that RL task shifts to something that's more similar to objectives that you might see in language and vision that we know have scaled so much.
10:46And so I think that's one of the fundamental insights that we've seen is that it seems like by approaching RL in this different approach. We were able to get so much more out of, we were able to scale our networks significantly beyond what was standard used in RL. Can I jump in? I will just give a bit more of context about the architecture because we use another objective, the influence, so the contrastive loss. However, the architecture is quite similar to the previous works of previous papers like Braw or Simba, Simba V1, Simba V2, Simba V1, Simba V2. So we also tweaked a bit of this architecture.
11:31However, it's not that we invented the wheel for the first time. It's the merging between the architecture and the objective that makes the scale really go up and performance follow the scale. I think that's something that we should probably mine deeper Do you think, I guess, like, what domains, what industry, like, you've applied it on multiple different types of networks or data sets. Is there a particular affinity that you think, like, is, like, kind of low-hanging food? Yeah. So, actually, if you look at a lot of our tests, they're particularly sort of, like, robotics tests. So, this is, personally, I'd be very curious about how a work like this could impact, like, the robotics field.
12:15Like, my understanding of robotics is that a lot of robotics right now, there's kind of a few different approaches. One approach is we want to train robots using imitation learning. So we try to collect an insane amount of data. We have a ton of human supervision and we try to scale up this data and we're learning with imitation learning. But on the other hand, perhaps there's another approach, which is, for example, goal-conditioned reinforcement learning, where we can actually train robotic agents and train RL agents to solve meaningful tasks with absolutely no human supervision, no demonstration.
12:45It's much more scalable, yeah. Right. So yeah, so this could serve as an alternate approach. And perhaps instead of scaling data, scaling manual human supervision, which is not super scalable, if there are ways to make goal-conditioned reinforcement learning scalable, and we can just scale the architecture, or we can scale... Because you're focused on all your objectives. Right, with certain different objectives. I think that could be very exciting to see how that can affect a field like robotics, for example. Yeah. Double-click on just one thing on the efficiency, which Igor was talking about.
13:14I would expect the deeper it is, it should be quadratically worse. I'm not familiar with the pre-existing literature. I'm just sort of working on intuitions. But basically, what are the trade-offs that you've found that I think you might want to warn people about? Because you are the guy who mentioned efficiency. Sure, sure. So I was referring to one of the figures on our poster, also in our paper, where we compare the number of parameters that models have as we see along the axis of depth. as we scale along the axis of width. From our baseline architecture, the most baseline one would be a width of 256.
13:51The hidden layers have 256 neurons, and then the depth is for four layers, for hidden layers. And so the point I was making there is that when you scale along depth, the number of parameters that your model has is going to grow roughly linearly. Whereas with width, you're making your network outputs wider, and then the input to the next network is also growing as well. And so the number of parameters your network's then going to have grows approximately quadratically. And so one of the experiments we did was sort of examining as we grow the number of parameters in our model by scaling along these two different choices, which one for the same approximate number of parameters yields a better performance.
14:28And the depth curve kind of goes like this. It jumps up pretty fast. That's present throughout our paper. For width, it grows a little bit more slowly. And so the kind of takeaway from that is that if you are a bit more resource constrained, scaling along depth might be better because there's fewer parameters with a smaller model. a smaller number tool learnable parameters. With is expensive. With is expensive, exactly. And in general, of course, like more parameters is also going to be more expensive. So that's just like another consideration to think about when using these networks, I suppose.
14:57Yeah. Any other sort of rules of thumbs like that that I can extract that this is just the most basic one that I could think of? Yeah. I don't know if there's any others. Yeah, I guess like your original question of like the trade-offs, like one of the trade-offs, one of the limitations that we say is like obviously if you make the networks bigger, it will take longer to run, right? So if you double the depth, at some level of depth, it might take twice as much to make a forward pass through the network, right? However, this is not... So within our paper, for most environments, we are able to saturate, get to almost perfect performance within just...
15:33We don't even need to get to 1 ,000 layers. Maybe just 64 layers, for example, is sufficient. And in this regime, the latency of the network is not necessarily a significant bottleneck. You can imagine there's a lot of tasks in which, especially in RL, that collecting data might be the bottleneck. And making four passes through our network may not be the bottleneck. And so in our environment, in our research, we specifically used the JAX GCRL environment, which is a JAX-based GPU Accelerator environment. So we can collect thousands of environment trajectories in parallel at the same time so that we're able to make...
16:10Oh, this is built in. Right, this is built in so that we can collect a thousand trajectories at the same time along all these environments and so make sure that we have enough data to exaggerate the learning from data. Wow. That's like work they have in columns. Okay. I don't know if you want to expound upon that on the Darkly Zero. And you know, most people are familiar with PyTorch, maybe less familiar with JAX. With JAX. I think JAX is getting the traction, especially in RL field because for online reinforcement learning, getting as much data as you can is the most important. There's got to be a PyTorch equivalent, but anyway.
16:50Any tips for other people also exploring this kind of rollout? Yeah, so I think I can also recommend, for GoConditioned RL, I'm recommending JAX-CRL, but there are also multi-agent JAX implementations and others. So going back to our paper, if you look at the plots, we only see this huge performance increase when we cross 50 millions of transitions gap. So I think the data is crucial here. I guess even to build on that, I like drawing analogies to successes in other areas of deep learning. For example, in large language models, the reason why we're able to scale to such large networks is that we found a paradigm in which we can leverage the entire internet scale of data to learn.
17:38And so data in RL traditionally has been hard to come by. But now with these GPU accelerator environments, we can collect hundreds of millions of time steps of data within just a few hours. And so I think that this serves as a really good test bed for us to be able to also find ways to scale up network capacity and get similar kind of gains. Are you saying that you have a difference, you would do pre-training differently in LLMs? Like what's the difference objective now very simply the paradigm that you're referencing is next word or next token, right? It's very robust How do you change that? Oh, I'm not saying that we're changing it I want to leverage insights from that to apply to morale I feel like you should go the other way You think you should go the other way?
18:28Maybe, I mean that would be a very interesting research direction too But actually, even on that point one of the things I was thinking about is that the way that our RL objective works is in some sense, it's not exactly next word prediction, but it's kind of like next state prediction, right? You imagine you're at some current state and you're at some current action. And we want to predict whether or not this future state, this certain state is a future state along the same trajectory or a different trajectory. And so in some sense, we are actually doing some sort of like... Implicit world model.
18:57Implicit like, you know, like... I don't know if that's a bad word. Or like in language, you do a cross-stattory loss to classify the next token, right? And here we're just doing a binary classification of whether or not some next state is some... Yeah, yeah, yeah. It's a classification. Yeah, yeah, yeah. And so I do see that there are some sort of parallels here that perhaps we should dig into deeper and see what is the core of what enables deep learning to scale and then how can we leverage that? How can we distill those insights and then apply those across all different fields, whether it's language or reinforcement learning?
19:28Yeah. Did you get my meaning about the role model stuff? Yeah, yeah, actually. And I heard, I think I might have heard Professor Eisenbach yesterday talking about this at a poster and he's explaining to a couple of people that because this is like doing representation learning and trying to learn these meaningful representations for a given state of action but for a given goal in some sense you can think of it almost like learning a model of environment learning a model of the world but without having to do any sort of like next frame prediction or stuff like that that's a little bit more high dimensional and complex yeah I would think like the angle that I'm trying to think about and push is instead of learn the next world, they basically generate a number of candidates, possible worlds, and classify them, to your point, which is exactly how I do things.
20:14Let's say I'm playing poker and I'm trying to classify what hands you have. Well, there's a range of hands based on what you're doing. And the more information I get, the more I resolve to, oh, I know exactly what hand you have based on what you're showing, or you're buffing, but that's a different thing. But you know what I mean? I feel like that is the ultimate sort of angle of representation, which is a world. But I don't know if that is too vague compared to the more concrete types of world models that, let's say, the video gen people are doing. And then, I guess, one other thing I'm also exploring, you mentioned the deep models being slower or more expensive.
20:52Yeah, that is a trend in the inference world of making models shallower, right? And I wonder if this short catchphrase I was thinking about, like, deep teacher shallow student would be a good deployment paradigm. Like you push the frontier capabilities with that and then you distill it back. Actually, this is a good point. If you go out to our website, this is one of the future directions that we list at the very bottom. We would love to see if we could get similar performance. We do achieve state-of-the-art performance on Goal Condition RL and Jack's GCRL by a significant amount. And so it was very exciting to see the sort of frontier of the ability to train RL agents sort of pushed.
21:38And if we can do that in a way that also sort of is just as efficient as a standard, you know, networks, that would be very cool. So, you know, like... Yeah, because training doesn't have to be the same thing that you deploy at inference. Right, you know what I mean? Yeah, so if there's ways to distill down to a smaller model or prune the model and maybe not still retain performance, that's a very interesting research structure in that we're using. Let's talk about other future directions. What else is your personal passions? Yeah, so currently I'm pursuing the direction of stitching in reinforcement learning.
22:11So we are trying to generalize reinforcement learning from shorter sub-behaviors so that they are stitched, merged during the test time. And yeah, I think this is one of my last papers that I will tackle during the PhD. Personally, I'm very curious of like, can we, like, what's the real, like, can we push I'm curious about like advancing the frontier as much as possible. So if you actually look at our paper, we focus on the scaling depth, but we notice that we see that scaling width actually also improves performance, and we also find that actually by scaling depth, we actually unlock the ability to scale along batch size as well.
22:50So this is one of, yeah. So I guess Right, so like, okay, I guess for context, like, in traditional RL, like a value-based RL. Scaling batch size is not super effective. But we also can see, there's also other work in other areas of deep learning that show that scaling batch size is only most effective when there's like a large enough network capacity to take advantage of the scaled batch size. And we actually find that, you know, perhaps, so one hypothesis might be like, perhaps the reason why scaling batch size isn't that effective in traditional RLs because like we've been using these tiny networks that haven't been able to capture that.
23:22And one of our experiments is that like because we are enabled successful training of deep network, we actually were able to, this is a great testbed for, you know, like testing this hypothesis. And we find that indeed, as we scale to network capacity, we also unlock this different dimension of scaling batch size. And so all I have to say is that I'm very curious for someone like with enough compute to like take some of these environments, scale up depth to the maximum capability, also scale along width, also scale along batch size. And let's like basically like in the same way that in language, we're scaling along so many different axes.
23:56can we unlock different dimensions of scaling as well and what capabilities and how far can we push the frontier of training these RO agents from doing them. Before we pass it, Sean, when you say enough compute, what kind of compute budget did you have? I just want to see what you guys got. Good question. So we wanted to make it such that it's quite accessible. So the nice thing is that all of our experiments, even the 1 ,000-layer networks, can be run on one single 80-gigabyte H100 GPU. so that's dollars right, right, right so everything can be run on one GPU but in theory if we had a distributed training setup and can just blast compute through this and really wanted to push the frontier it'd be very interesting to see how things go and I've actively been trying to learn as much as I can about vision language action models role models at Nureps I'm going to a lot of machine language action models?
24:49vision language and yeah, curious about applications of representation for these. Yeah, exactly. For robotics. I'm actively trying to explore more in that area. So just reading a lot of literature, talking to as many people. Yeah. We just released our episode with Genuine Tuition. Oh, okay. Awesome. Where if you know a bit about their history, they started as a gaming clipping company and they basically have a vision language action model, which I saw a preview. It was very impressive. I'm not sure exactly how transferable it is to embodied use cases, but it doesn't have to. Like, screen is fine.
25:27You know? Like, yeah, I don't know if you have any takes on... Yeah, that's an exciting research direction. Definitely. Yeah, I think the concepts of actions as something that you are outputting is actually not that popular in industry, right? Only because text has completely dominated the last three years. And tool calling which is just another form of structured text. And I feel like the action research is kind of like, I don't know what needs to happen in order to unlock the next phase in that. I don't know if you've seen anything interesting out here. Shut it up. Yeah, there's a lot of cool work on leveraging pre-trained BLMs.
Read the full transcript
26:13You freeze it and then you apply it. And then you're creating on top of that sort of experts to output actions. also like systems for doing like hierarchical planning maybe outputting some higher level plan and this is like a larger network that takes a long time, a little longer to do inference and so it outputs its plans with less frequency, like some sort of chunk and then from there there's like some sort of second system that operates a bit more fast. I think there's quite a bit of interesting research in that direction so that's what I'm looking forward to. Cool. Final question. Hardest question you were asked at the post-to-session or just favorite encounter?
26:48Anyone famous that you met? So I actually haven't gotten a chance to go to the conference that much. I'm actually working full-time now. Oh, damn. Yeah. So far, I actually literally just got my badge like a few moments before my interception. So I guess I wouldn't be the best to answer that question. No, no, no. Like people ask you stuff, right? Oh, oh, oh. At my post. People asking you or meeting you and like, you know, just give a vibe of like what people are saying. Yeah, people were very, I think it's sort of like a very eye-opening. I think that the general question is that people thought it was a very eye-opening paper because like the objective is quite simple.
27:25It's quite elegant. And for us to be able to like, you know, like I don't want to say like overturn, but like sort of challenge the conventional wisdom that like RL is not super scalable and push it to such limits like a thousand layers deep and see continuing to improve performance. I think the general impression that I've gotten is that, you know, this could be like a really cool, like if we can sort of build along this direction and that like we can really scale along all these different dimensions and push the frontier of the ability for RL. I'm very curious to see how that goes. All right.
27:58Well, thank you so much for dropping by. Congrats on the paper again. And good luck in your future work. Thank you. Thanks for having us. Yeah. Thanks.
28:14Thank you.
From the publisher
From undergraduate research seminars at Princeton to winning Best Paper award at NeurIPS 2025, Kevin Wang, Ishaan Javali, Michał Bortkiewicz, Tomasz Trzcinski, Benjamin Eysenbach defied conventional wisdom by scaling reinforcement learning networks to 1,000 layers deep—unlocking performance gains that the RL community thought impossible. We caught up with the team live at NeurIPS to dig into the story behind RL1000: why deep networks have worked in language and vision but failed in RL for over a decade (spoiler: it's not just about depth, it's about the objective), how they discovered that self-supervised RL (learning representations of states, actions, and future states via contrastive learning) scales where value-based methods collapse, the critical architectural tricks that made it work (residual connections, layer normalization, and a shift from regression to classification), why scaling depth is more parameter-efficient than scaling width (linear vs. quadratic growth), how Jax and GPU-accelerated environments let them collect hundreds of millions of transitions in hours (the data abundance that unlocked scaling in the first place), the "critical depth" phenomenon where performance doesn't just improve—it multiplies once you cross 15M+ transitions and add the right architectural components, why this isn't just "make networks bigger" but a fundamental shift in RL objectives (their code doesn't have a line saying "maximize rewards"—it's pure self-supervised representation learning), how deep teacher, shallow student distillation could unlock deployment at scale (train frontier capabilities with 1000 layers, distill down to efficient inference models), the robotics implications (goal-conditioned RL without human supervision or demonstrations, scaling architecture instead of scaling manual data collection), and their thesis that RL is finally ready to scale like language and vision—not by throwing compute at value functions, but by borrowing the self-supervised, representation-learning paradigms that made the rest of deep learning work.
We discuss:
The self-supervised RL objective: instead of learning value functions (noisy, biased, spurious), they learn representations where states along the same trajectory are pushed together, states along different trajectories are pushed apart—turning RL into a classification problem
Why naive scaling failed: doubling depth degraded performance, doubling again with residual connections and layer norm suddenly skyrocketed performance in one environment—unlocking the "critical depth" phenomenon
Scaling depth vs. width: depth grows parameters linearly, width grows quadratically—depth is more parameter-efficient and sample-efficient for the same performance
The Jax + GPU-accelerated environments unlock: collecting thousands of trajectories in parallel meant data wasn't the bottleneck, and crossing 15M+ transitions was when deep networks really paid off
The blurring of RL and self-supervised learning: their code doesn't maximize rewards directly, it's an actor-critic goal-conditioned RL algorithm, but the learning burden shifts to classification (cross-entropy loss, representation learning) instead of TD error regression
Why scaling batch size unlocks at depth: traditional RL doesn't benefit from larger batches because networks are too small to exploit the signal, but once you scale depth, batch size becomes another effective scaling dimension
—
RL1000 Team (Princeton)
1000 Layer Networks for Self-Supervised RL: Scaling Depth Can Enable New Goal-Reaching Capabilities: https://openreview.net/forum?id=s0JVsx3bx1
Chapters
00:00:00 Introduction: Best Paper Award and NeurIPS Poster Experience
00:01:11 Team Introductions and Princeton Research Origins
00:03:35 The Deep Learning Anomaly: Why RL Stayed Shallow
00:04:35 Self-Supervised RL: A Different Approach to Scaling
00:05:13 The Breakthrough Moment: Residual Connections and Critical Depth
00:07:15 Architectural Choices: Borrowing from ResNets and Avoiding Vanishing Gradients
00:07:50 Clarifying the Paper: Not Just Big Networks, But Different Objectives
00:08:46 Blurring the Lines: RL Meets Self-Supervised Learning
00:09:44 From TD Errors to Classification: Why This Objective Scales
00:11:06 Architecture Details: Building on Braw and SymbaFowl
00:12:05 Robotics Applications: Goal-Conditioned RL Without Human Supervision
00:13:15 Efficiency Trade-offs: Depth vs Width and Parameter Scaling
00:15:48 JAX and GPU-Accelerated Environments: The Data Infrastructure
00:18:05 World Models and Next State Classification
00:22:37 Unlocking Batch Size Scaling Through Network Capacity
00:24:10 Compute Requirements: State-of-the-Art on a Single GPU
00:21:02 Future Directions: Distillation, VLMs, and Hierarchical Planning
00:27:15 Closing Thoughts: Challenging Conventional Wisdom in RL Scaling




