In short
Podcast Episode Notes: Building the GitHub for RL Environments
Episode Overview
- Podcast Title: Training Data
- Episode Title: Building the GitHub for RL Environments: Prime Intellect's Will Brown & Johannes Hagemann
- Hosts: Sonya Huang, Pat Grady, and partners from Sequoia Capital
- Guests: Will Brown & Johannes Hagemann from Prime Intellect
- Description: Discussion on the transition from static prompting to environment-based AI development, focusing on Prime Intellect's Environments Hub aimed at democratizing AI training.
Key Themes and Concepts Institutional Knowledge as Training Data
- Preference for long-term experience: The value of having personnel with extensive institutional knowledge over just having the smartest individuals.
- Importance of compounding expertise over time, enabling organizations to build and improve upon past successes rather than starting anew each day.
Shift in AI Development Paradigms
- Transition from static prompting to environment-based AI development.
- Recursive Language Models that can manage their own context.
- Agentic Reinforcement Learning (RL) that scales through trial and error.
Vision for the Future
- Anticipation that every company will evolve into an AI research lab capable of pre-training, post-training, and fine-tuning their models.
- Emphasis on optimizing models for specific applications over generic off-the-shelf solutions.
Environments Hub Purpose
- A platform designed to provide access to advanced training infrastructure for startups, enterprises, and research labs, previously restricted to large labs.
- Focus on making frontier-level AI capabilities accessible to all.
Components
- Full stack research platform called Lab, offering tools for large-scale reinforcement learning.
- Community-driven approach, similar to GitHub, allowing users to share and collaborate on AI training environments.
- Integration of evaluation frameworks and model customization capabilities.
Environment Definition
- Environments are crucial for RL, encompassing tasks, evaluation metrics, and interaction protocols.
- Comparison with evals; environments serve both training and evaluation purposes, facilitating the improvement of model performance.
Reinforcement Learning (RL) and Post-Training Post-Training
- Concept of post-training as a significant phase in the AI development cycle, enabling models to improve continuously.
- Importance of using environments for effective post-training and performance optimization.
Role of Agent Harnesses
- Agent harnesses are integral to the environment, defining how models interact with tasks and systems.
- Evolution of harnesses to adapt to diverse application needs, including coding environments.
Customer Examples and Use Cases Notable Collaborations
- RCI: Focused on Frontier Open Models and collaboration on large pre-training and post-training efforts.
- Medical AI Labs: Developing benchmarks for medical capabilities and ensuring trust from medical professionals.
Popular Environment Examples
- Wordle: A simple yet effective environment for demonstration purposes.
- WikiSearch: An adaptable template for searching over various document types, encouraging user adaptations and iterations.
Research and Future Directions Recursive Language Models
- Interest in models that improve their context management.
- Potential to enhance long-term reasoning capabilities in AI.
Synthetic Data Research
- Exploration of how models can curate their own training data.
- Connection to continuous learning and enhancing model capabilities over time.
Open Source and Model Accessibility
- Discussion on the importance of open-weight models.
- Recognition that closed models can still benefit from environment-based optimization and testing.
Closing Thoughts
- The future vision includes empowering companies to harness AI capabilities and not solely rely on large labs for advancements.
- The overarching aim is to create an accessible environment for AI research and development, democratizing AI innovation and improving product optimization.
Key Takeaways
- Institutional knowledge is invaluable in AI training.
- The shift to environment-based training is pivotal for the future of AI development.
- Prime Intellect’s Environments Hub facilitates collaboration and innovation in AI research.
- The evolution of models and training practices will enable more efficient AI development across various industries.
Conclusion This episode provides a comprehensive exploration of the ongoing transition in AI development practices, emphasizing the importance of institutional knowledge, environment-based training, and the future potential of AI research labs within every company. Will Brown and Johannes Hagemann articulate a vision for democratizing access to advanced AI capabilities, setting a promising trajectory for the industry.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe Importance of Institutional Knowledge
0:45 to 1:50
Discussing the value of long-term expertise in development.
“We see the same happening for AI research.”
Prime Intellect's Mission
1:50 to 3:17
Overview of Prime Intellect's mission to democratize AI research.
“Maybe to get started, you are one of the leading research labs enabling customers to post-train their agents.”
Understanding Post-Training in AI
3:17 to 4:37
Explaining the concept of post-training and its implications.
“where we are like the winning applications are using AI for a specific thing, for some agent, for some workflow, where you also want to be able to deliver this at scale and with cost effective performance.”
The Role of Environments in AI Training
4:37 to 6:12
Exploring how environments are essential for reinforcement learning.
“So we have a full stack research platform called Lab.”
Optimizing AI Models with Environments
6:12 to 7:39
Discussing strategies for optimizing AI models using environments.
“of have this way of sharing the ability to improve performance across different tasks.”
Agent Harnesses and Their Importance
7:39 to 9:01
Examining the relationship between agent harnesses and environments.
“And can we say a word on, you said something interesting.”
Evaluating AI Performance
9:01 to 10:48
Differentiating between evals and environments for AI performance assessment.
“And that's kind of what I mean by it's uneval in that there is a system that it starts at some state, it interacts with the system, the environment, the harness, the agent, whatever you want to call it.”
Understanding Evaluation in Reinforcement Learning
14:01 to 15:14
Explore how evaluation metrics relate to environments in reinforcement learning.
“I think that's a less like controversial take.”
Post-Training Techniques and the Role of RL
15:14 to 16:40
Discover various techniques alongside reinforcement learning in model development.
“Meaning like if you're post-training models, are you doing reinforcement learning or are you doing other things?”
Customer Insights and Target Audience
16:40 to 18:35
Learn about the customers using the platform and their capabilities.
“Actually, on that note, do you guys have any favorite customer stories that you want to share?”
Show all 24 chapters
Customer Stories in Medical AI
18:35 to 21:18
Hear about collaborations in the medical field using AI for better outcomes.
“seamless and kind of quick as possible and efficient as possible.”
Building Cybersecurity Learning Environments
21:18 to 22:40
Understand the construction of environments for cybersecurity training.
“And yeah, those folks have built amazing environments in all kinds of verticals, in a sense, everything from verifiable software engineering and lean to medical physics environment to some cybersecurity environments.”
Evaluating Reinforcement Learning Environments
22:40 to 24:16
Explore how reinforcement learning can identify issues in environments.
“So there's someone in our residency program who actually does a lot of this stuff.”
Designing Realistic Environments for AI
24:16 to 28:00
Examine challenges in creating environments that reflect real-world complexities.
“And if we take the cybersecurity environment example one step further, by the way, I have no particular affection for cybersecurity, but it's something that like I love having a specific example in my head.”
Enhancing Developer Experience in RL Environments
28:00 to 29:00
Learn how to improve the developer experience when building coding agents.
“There's some that I think we're kind of ready to build for when we need to.”
Creating a Hub for RL Environments
29:00 to 30:40
Discover the benefits of a centralized hub for sharing and optimizing RL environments.
“Yeah, I mean, it very much seems like it kind of already is, where it does seem like a lot of the focus from the major labs has shifted to they're still using a lot of human data.”
Community Engagement with the Environment Hub
30:40 to 32:40
Explore how the community interacts with the environment hub, including sharing and modifying environments.
“So having proper evaluations, integrated with your environment, so you can like immediately test them across all frontier models as one of the features people are heavily using the environment hub then for, right?”
Future Research Directions in Reinforcement Learning
32:40 to 35:30
Examine the future of RL and its efficiency in research and model training.
“But it's designed to be this template you could use for agentic search more broadly.”
Open Weight Models and Their Impact
35:30 to 37:30
Understand the role of open weight models in RL and how they integrate with various systems.
“a sense on like the specific piece of how useful reinforcement learning can be in the coding domain.”
Recursive Language Models and Context Management
37:30 to 40:00
Learn about recursive language models and their potential in managing context for agents.
“And do you need, need, need the weights?”
Synthetic Data and Lifelong Learning
40:00 to 42:01
Explore the potential of synthetic data in improving lifelong learning for AI models.
“I've been exploring it as part of our research as well.”
Exploring Model-Driven Training Data
42:01 to 42:31
Learn about the potential of models curating their own training data for lifelong learning.
“Okay, we're going to close on an optimistic note.”
Vision for a Collaborative AI Future
42:31 to 43:58
Discover how Prime Intellect aims to empower entrepreneurs in the AI landscape.
“If everything goes right, what does the world look like?”
The Importance of Institutional Knowledge
43:58 to 44:15
Understand why institutional knowledge and expertise are vital for company growth.
“And I think that's how we've thought about approaching it, especially as software becomes easier for people to manipulate, as the barrier to entry for coding becomes easier.”
Transcript
Automatic transcript. May contain errors.0:00If data is the bottleneck, if having the real expertise is the bottleneck, would you rather have the smartest person in history work at your company or someone who's been there for 30 years? Sometimes you really want the person who's been there for 30 years. There's a lot of expertise that comes from really understanding a problem deeply and interact with it over a long time. And this is really what happens in training that is almost impossible to replicate in a short prompt. You really want the ability for institutional knowledge to compound over time, for best practices to compound over time.
0:27And this is how institutions and companies grow to be really powerful and successful is they stand on the shoulders of what they've done before rather than kind of resetting every day. And we want to have this be accessible to any company that wants to do this. And I think that's how we've thought about approaching it, especially as software becomes easier for people to manipulate, as the barrier to entry for coding becomes easier. We see the same happening for AI research.
1:07Will and Johannes work at Prime Intellect, which is one of the coolest neolabs in AI right now. Your mission is to make frontier lab training accessible to everyone, which I think is a very noble mission. You have really, really strong taste and just developer feel and just understanding how to, you know, that intuition for what developers care about. And then what you all launched with the Reinforcement Learning Environments Hub was like really, really differentiated and people were very excited about. And so I'm excited to chat about many topics with you all today. Post-training, reinforcement learning, agent harnesses, your platform, the RL Hub, and then big picture questions on what's coming next and post-training in RL.
1:47Does that work? Sounds great. Absolutely. Maybe to get started, you are one of the leading research labs enabling customers to post-train their agents. Can you tell me about what is that? What does that mean? Yeah, for sure. Happy to take that one and give a bit of a higher level overview of what our platform does as well, our research at Prime Intellect. As you already mentioned at the beginning, we try to make frontier infrastructure available to any startup, enterprise and Neolab as well. And basically the infrastructure that is currently locked behind the walls of the big labs where nobody really has access to them.
2:19And, yeah, we really start from, like, the compute layer and the compute orchestration layer and go all the way up then to the entire full post-training stack. So everything from, like, the training frameworks that are needed to do large-scale reinforcement learning to, yeah, the environments with, yeah, a bit of a more community approach of our environment hub to, yeah, other pieces that actually needed to do this, like, yeah, sandboxes for secure code execution and, yeah, evaluations as part of our environment hub as well. and, yeah, to offer this, like, as an end-to-end product in a sense. And what's the intuition for, why even, why pursue that mission statement of making all that infrastructure available to everyone?
3:00Yeah, there's a lot of reasons. I think one is, and I think something that we are very passionate about is just, like, open science as a way that humanity moves forward, where, like, a lot of the big scientific discoveries historically have been things that we talk about and as a world we can kind of build on top of. But kind of more practically speaking as well, there's a lot of value in model customization, where we are like the winning applications are using AI for a specific thing, for some agent, for some workflow, where you also want to be able to deliver this at scale and with cost effective performance.
3:29And so really the way to kind of really optimize these systems end to end is to be able to have access to the model weights directly, where you can then craft the model to be the best model for your problem rather than some model off the shelf where the crafting happens just in a prompt. And so it's really just allowing a deeper layer of customization than what you can do at the prompt level. And then is your vision for the future then that every company will be pre-training their own models, post-training their own models, fine-tuning their own? What do you think the future holds? We definitely think that every company will be an AI company.
4:02And we think most AI companies will want to have an AI research lab. And research can look like many different things. It can look like pre-training, especially if you're in a domain where you maybe don't just want text in, text out, if you want something more bespoke. In a lot of cases, it will be like post-training agentic or otherwise kind of focused models for specific tasks and workflows. And I think that's getting to the point where it's productionizable and cost-effective, where you can actually make this very practical for people to do at scale for the right shape problems that people want to solve.
4:35Awesome. And then can you say at a high level what your platforms us? Yeah. So we have a full stack research platform called Lab. And Lab is about giving everybody the ability to do the things that a frontier research lab can do internally, but for anyone in the world who wants to do this kind of research. And a big focus of Lab is the notion of an environment. And so I think a lot of people have heard the term environment in the context of reinforcement learning. And that's a big focus of it for us, which is that an RL environment is encapsulating the things you need to do to do reinforcement learning, where you can have a model improve via trial and error.
5:08But it's also more general beyond just reinforcement learning. And I think if you haven't heard of an environment before, it's essentially the same thing as the evals that get reported when people talk about new model releases. So Sweebench and Amy and Terminalbench, these are all examples of environments where there's a data set of tasks, there's a harness for the model to be in, and there's something called a rubric or reward function, which is responsible for grading the quality of the outputs. And so the same thing you'd use as an eval offline as your kind of your test set. You can use this in reinforcement learning as your train set.
5:39And this is a way to improve model performance interactively. And so our platform is really enabling people to use environments in their workflows for post-training, for evaluation, for synthetic data, for reinforcement learning. And it's also very much focused on as a community platform in the same way as things like GitHub are, where this stuff is new, it's complicated, and we want people to have lots of building blocks they can draw from and lots of examples and for it to be collaborative and for people to have reasons to kind of show off ways of using models or different tools in workflows as environments to allow other people to post train models in those environments to kind of have this way of sharing the ability to improve performance across different tasks.
6:18I think the general idea is to give more companies the actual ability and advantage that's currently only the big labs have in a sense of like this product model optimization loop where they can optimize their models for their specific product in a sense. And we see it as a thing where, yeah, that's the kind of reason why like a chat GPT was created by OPMI or like a cloud code was created by Anthropic. They actually have the capabilities to optimize models for their specific scaffolds in a sense and yeah, have their models work way better in their products. And some more popular as those kind of products also become in a sense like a cloud code becoming extremely popular right now um there is a big levels have yeah naturally less enough an incentive to actually make it work better for like other coding startups in a sense right and um the idea there is to to give them the tools to have like their own like model product optimization loop and um yeah i think there are early adapters on that front um i think yeah one great example um uh always in this case that i always give is scurr for example um that yeah realized that in my opinion quite early on they built their own like composer one model where they did like large-scale post training in a sense and um yeah really optimized a model where the environment was actually cursor itself so uh yeah it gives them all the tools that like you have in cursor in a sense as well and um yeah optimize a model inside of cursor and yeah we believe there's a lot of more startups that yeah will go this direction to um yeah on the one side optimize their current products in a sense, but also, yeah, build completely new products that are really not possible right now without having this product model optimization loop.
7:58Awesome. And can we say a word on, you said something interesting. You said environments are just evals. Can we dissect that statement? In my head, an environment is a state. It's a description of world state through which, you know, you observe what actions you take, you observe how world state changes, and therefore you update your world model. That to me is distinct from an eval, which is like, you should have gotten this answer on this set of questions. And so can you help me merge those two realities? Sure. Yeah. So I think there's a version of eval that's kind of where we were maybe a year or two ago where a lot of evals were like question and answer.
8:34And it's like this big bank of questions. And then maybe there's other notions of environment that people think about when they talk about like an old school RL with like Atari that is much more about like this kind of long running state interaction loop. And I think where we are now is it's both in the same thing where especially the environments that you want to do large scale training on, they do have this complex state. They maybe are simulating a web app. It's a full fledged kind of coding platform where you have an agent doing these things. But in the original RL games, there's always a reward.
9:02There's a goal. And so the notion of there being a goal of this problem, where it's not just running through some system and a human is going to kind of vibe check it, there actually is something that can measure progress and performance. And that's kind of what I mean by it's uneval in that there is a system that it starts at some state, it interacts with the system, the environment, the harness, the agent, whatever you want to call it. But there is some goal and there's a way to measure whether it's doing well or not. Okay. Got it. And then is there a difference between, you mentioned kind of cursor as a great example of somebody doing really great frontier work and reinforcement learning.
9:39Is there a good way to think about when should you be kind of constructing RL environments versus when should your actual application and your application states be the environment, so to speak? Yeah, I think there's definitely reasons where you want both, especially like let's say you're training a model to be a really good Rust coding model. Here you might want it to be good for lots of different applications where you'd have different environments that are focused on some like domain task. Maybe you're a company that wants a model that's going to be really good at calling your tools or using your specific domain language that then you can provide as a service to people who are building around that.
10:15And there's also ones where the product very much is like a user interface for an end user where the user is interacting with an agent, in which case it might make the most sense for the product to very directly become the environment where I think the companies who might want to be doing this are the same who care about whether they're using Cloud or GBT5 or Gemini or the ability to choose models and have internal systems to evaluate. whether a certain system prompt is good or whether changing out a model endpoint is good or whether using the mini version of a model is a better cost performance trade-off.
10:44The infrastructure to do that is that a lot of people have already been building at these kind of advanced agent companies is the same infrastructure you use to do reinforcement learning. And so I think that's kind of this kind of convenient world we ended up in where the training paradigm that makes the most sense for improving model capabilities is the same sort of thing that a lot of people have been building up the muscle for just without doing the training piece. And so that's kind of where we see RL being a very useful tool for people to then have this as an option in their toolkit for system optimization.
11:13Got it. Okay. I want to talk about agent harnesses as they apply RL. I feel like harnesses is like the theme of the moment, especially with you mentioned Cloud Code getting so much love. I think one of the things that they do exceptional engineering around is the harness. Harness and RL, are those things orthogonal, mutually exclusive? Like how do they relate? they definitely relate um i think the way i think of a harness is like a piece of the environment and so um there's some for any like eval or environment task there's some input a bunch of stuff is going to happen then there's some output state which is then going to be graded and so this whole intermediate piece whether it's uh interacting with some simulator or interacting with uh some another agent or physical world sim this is the environment and the harness is very much a piece of that, where it couples how the model interacts with any other pieces of the system.
12:05And I think depending on the application area, like I think for coding agents, we have a pretty clear like definition of what's the, you could say the harness is the CLI coding agent and the terminal is the environment. But this isn't necessarily going to be universal across like all different types of agents. In some cases, it's a system prompt and some tools is the harness. In some cases, it's something that is going to be spawning sub agents and those sub agents also have their own harnesses. And so there's a lot of complexity. And I think the way we've thought about it is like harnesses are going to keep evolving.
12:34There's going to be this Camperian explosion of ways people want to use models. And we want to take a pretty general approach in defining what you could do with a harness. And so what we are really thinking about and why we use the term environment as the abstraction is like agent is too narrow, harness is like kind of too narrow. And you can do all of these things within an environment, but the environment as an abstraction on the whole allows this, any sort of system model interaction is in scope. And then do you think then, do you think all companies should be post-training their models with environments?
13:09Are there specific kind of, you know, where the bullseye, like you absolutely need to be using environments versus like you could be post-training with a different method versus you should just be prompting your thing? I mean, I think environments are tools you can use to do all of these things. And so I think That's when I talk about kind of environments beyond RL. I think part of the reason why it's a useful abstraction is because it doesn't tie you to RL. Let's say I want to have a small model and I want it to be distilled from a big model. The way you can do this is you take the big model, you plug it in your environment, you let it run a bunch of times.
13:42Now you have all this data that comes from the same interaction protocol. You can use the same greater at the end to filter for the best examples and then do SFT fine tuning on that. You could do prompt optimization with an environment. You could A-B test different models with an environment just as you would with an eval. And so we really, I think the idea is that every AI company should be optimizing their AI systems. I think that's a less like controversial take. Okay. And maybe there's another way to frame it. Like the eval is almost how your agent performs in the set of environments that you expect your customers to face.
14:13Yeah. I think it depends on what you want to call, like, I mean, in kind of traditional machine learning terms, you don't want to overfit to the test set. And so in some ways, we kind of are already accepting that we're going to be like using the test set or the eval to measure, to have that kind of filter back into the model. And so I think it's a little tricky to distinguish even like what is the eval? What is the environment? We kind of think of them as one and the same. And we use the environment term very generically, where an eval is a type of environment that's used for measuring performance but not training on.
14:45and that's kind of how we see a lot of people thinking about when they're doing evals what they really mean they're doing is they have some way of measuring current performance and they're iterating on it and in some ways rl is this iteration applied at scale where you're automatizing the process of changing the model a little bit changing the prompt a little bit and having this be the way that you can kind of hill climb on some goal maybe another question Do you see reinforcement learning as synonymous functionally with post-training? Meaning like if you're post-training models, are you doing reinforcement learning or are you doing other things?
15:18There's definitely a lot of things. I think reinforcement learning is like the big thing now where it's like in many cases, practically speaking, if you're doing a large scale model, RL will be where you spend the most of your time and focus and compute. But it's not the only thing you want to do. There's a lot of things involved in the pull process of going from some initial model to the system you want to deploy. This can be prompt tuning. This can be SFT. This can be online distillation. There's a number of algorithms that all kind of are under this umbrella where RL is like the big one in the middle.
15:46But there's a lot of stuff around it. And really, I think exposing this toolkit to people and letting them have all these knobs they can play with is the way to unlock. And are you finding that it's the, you know, really smart researchers that actually know how to make this stuff happen in practice? Is it, you know, towards your goal of democratizing AI development for all, you know, does your average Fortune 500 know how to use the platform and get value out of it? I think most companies have people who can. I think any Fortune 500 company will likely have a team of AI engineers who are capable at following the latest tools, who are good at using cloud code, who have a lot of opinions about models and prompting.
16:28And those people certainly can do this. That's kind of the audience where we see as the target customer for this. Especially if you give them the right tools to actually do it. Like maybe some of those large companies don't really have anybody in there who can like debug your GPU cluster in a sense to actually kick off such a run or other components that are needed in there to like actually just make it easier in a sense to do large scale like agentic reinforcement learning with tool use, with code execution and pieces like this to just abstract away the entire infrastructure for them. Yeah, got it.
17:02Awesome. Actually, on that note, do you guys have any favorite customer stories that you want to share? Yeah, one of our favorite customers that I would like to point out in a sense is RCI and Neolab working on like Frontier Open Models have been working with us on the entire stack in a sense, right? We've been talking a lot about reinforcement learning, right? But yeah, we've been also doing a large pre-training in the past. that's basically where our history is coming from in a sense as well right so yeah I've been training with them some of the largest mixture of expert models and actually able to yeah open source them as well and yes well on the on the post training side with them as well maybe well do you want to share some more sure yeah so they're a very close collaborate of ours that we've I think I'll be all been friends for a long time but also like we've been I think they've been a way where we've had the they've had a lot of things that they want infra for.
17:55And that's been a way that we've been able to kind of force us to build out a lot of the pieces, both from compute orchestration to post-training to pre-training to inference around kind of just making everything that is needed to kind of be a frontier lab. And I think they're very aligned with us in kind of the openness mission. But I think their kind of focus, they are more targeted at like enterprises and the end user where they are kind of going to work more directly with customers in terms of like end to end kind of delivery of a certain artifact where I think where we come in is we are we are really focused on the developer experience at the infralayer and the ways that we can make put these tools into more people's hands where the process of going from idea to deployable model can become as seamless and kind of quick as possible and efficient as possible.
18:45Awesome. And then any other favorite customer stories, maybe on people that are using the environment's product specifically? Yeah. So we are definitely really focused on like the research community. And like some of this is like a lot of grad students use it. A lot of like students and people who are getting early in like their career learning how to do this stuff or using it. But also a lot of labs who are focused on a very specific domain where let's say you're starting a medical AI lab. We work closely with a number of groups in the medical space where they want to create more both benchmarks to understand how good our models at medical capabilities, both in terms of diagnosis or patient interactions or question answer about medical literature or agentic search over certain medical tasks.
19:27And so like SoFont, OpenMed being two that we work closely with, where the focus there is to they really care about like earning trust from the medical professionals. And so for them, they really want to focus on the domain and not as much on kind of the generic medical, the LM infra that is kind of like a headache for a lot of people. And so for them, like being able to kind of have this platform for creating evaluations, for showing them off, for being able to use them to then improve model capabilities that could then be deployed locally in a hospital or deployed locally for some end user where the ability to have this customization and have this kind of end to end trust and understanding of the data providence for tasks at hand or the ability to kind of customize models very directly is very key.
20:12Do you have any customers that are using you for more of like what you call the old school kind of Atari style, you know, learning from your environment type? Do you mean like, so when I think of Atari, I think of like non-LLMs. And so we definitely are focused on LLMs and foundation models that look like LLMs. There's definitely a lot of researchers who use our platform for these things that are more like, there's some examples we have on the platform that are much more like games. So, I mean, one fun one that I use is like a demo a lot is the game Wordle from the New York Times, where this kind of ended up just being the like a great hello world environment for people because it's the infras really simple, but it's very expressive where you can kind of get a feel for it.
20:52And you get this aha moment of seeing a model learn to think about the game as you give it rewards for doing better. and you can do it with a really tiny model too. Like it's in this sweet spot of difficulty where you can do these runs on like a couple GPUs in an hour and see a model actually learn how to get better at the game. Yeah, I would say like the more toyish game examples. Yeah, usually the ones that people are going for for actually learning how to build those reinforcement learning environments. And yeah, that's also what people are heavily using the environment hub for in a sense, just because we have all this infrastructure built around it to be able to actually test and your environments and yeah that's usually how they start out and then yeah go to actually building more complex environments later on another group I would love to give a shout out to in a sense that I've been building some of the more complex environments on the hub are people part of our reinforcement learning residency where yeah we have like a group we initially started out with like eight to ten people I think 14 to 16 are now in the group you have people grad students as well as yeah people working full-time that yeah part-time like building reinforcement learning environments, as well as doing novel research on top of the Environment Hub.
22:01And yeah, those folks have built amazing environments in all kinds of verticals, in a sense, everything from verifiable software engineering and lean to medical physics environment to some cybersecurity environments. And then also we give them the tools, obviously, now to actually do the training in those as well. Yeah. Awesome. Can you help demystify for me, I've heard that some of the foundation model companies are spending millions of dollars each on some of these environments. And so you mentioned cybersecurity at the end. What goes into constructing a cybersecurity environment? Yeah. So there's someone in our residency program who actually does a lot of this stuff.
22:43And so there actually is a lot of tooling in the cybersecurity world around these capture-the-flag games where there's some system that has some hidden bug in it. And this is like a challenge that originally these are built for programmers. or programmers will have like a little hackathon where they try to go find the bug in some system. But you can adopt these challenges to LLMs as well, where it then is a full software environment, where it's a terminal, where the agent lives in the terminal and has access to tools for doing bash commands. Maybe it's using Cloud Code or some other wrapper for an agent harness, but it's in a terminal full of files and can interact with these files.
23:14And then at some point, it marks that it's finished with the rollout. And then you can grade the state of the environment using other pieces of software, other code that executes. But we do like, we actually have a lot of people we work with who are in the data and environment space where we've found that there's a lot of interest in using reinforcement learning as a way to evaluate data quality, where the ability to measure what happens when you train a model on a set of tasks, set of rubrics, set of environments, allows you to understand bugs in the environments. Because there are issues that come up in reinforcement learning where maybe if your environment has a backdoor, a model can exploit this and kind of game the system.
23:55And so I think there's a lot of interest in people using RL to like in the pipeline from like idea to environment that ends up in some frontier labs, like foundational training run for the next GPT or cloud model. There's a lot of vetting that goes into this and doing these like smaller or like medium scale runs in like, let's say one environment allows you to really poke and see where the problems might be. Okay, super interesting. And if we take the cybersecurity environment example one step further, by the way, I have no particular affection for cybersecurity, but it's something that like I love having a specific example in my head.
24:27I could see how you could construct a toy example, right? These Capture the Fly examples, I would imagine they're toy examples. I would imagine they look nothing like the actual kind of corporate network environment of a, you know, real company with all of its cybersecurity products. And so I guess how do you construct an actual, you know, does what people can construct, does it actually scale up towards imitating and reflecting the complexities of the real world? Almost like the, you know, the robotics terms of real gap. Yeah. Is there a sims real gap in crafting these environments? And is it a solution just to like train on like real kind of corporate security environments?
24:59Yeah. So I would say there's less of a barrier in terms of the actual complexity. Like the, in principle, these can be as complex as you want. It's anything that could be on a computer. So just think of anything that's on a computer or network of computers as potentially an environment. What really kind of becomes the bottleneck in many ways is like cost of the simulator, where I think there's a lot of focus on identifying clever ways to kind of mock the right piece of the system. Where let's say you have an agent that you want to like, let's say the Internet, like let's say you want to have an agent that like does a lot of like web search.
Read the full transcript
25:30In some cases, you actually want to do the real web search. In some cases, you want to find ways where you can design tasks where you actually don't need the full thing. it's like this kind of I don't know I think of like the map games where there's this world you want to explore where there might be this whole map and like some of it's like dark but as you walk around like the light shines then you now see this part of it and so if you know which pieces of this map might actually be explored on a given task then this kind of allows you to decide which pieces are important or not and so like one example like there's some bench there's a benchmark called Tau2 bench that is very popular in the eval world where it's about customer service agents involving a database.
26:09And so the database is something where like in a real system, you might need the full database. But if you know that the agent is only going to ask certain types of questions for a task, you don't need a full database with millions of records or thousands or billions of like different examples. You can have kind of a mock database that is a cheaper to run, maybe in memory piece of software that only has the things that are expected by the agent or expected in the scope for a given task. And so identifying, like, the pieces of the system where the complexity is kind of overkill is part of this task design process to enable efficiency.
26:45And does your platform help with that? Like, it almost seems like, you know, if creating environments is the bottleneck to training these systems, then, like, being able to efficiently create these environments where, you know, you're lighting up the part of the map and it has to be lit up is, like, a core kind of almost platform competency. Do you guys assist with that? Yeah, I mean, it's definitely a way that we think about the design of everything. And we kind of go down a lot of these different like rabbit holes as we have to focus on different tasks and working with different people. And so like coding agents maybe is one example where like there's a lot of complexity that comes up when working with like sandboxes and terminal states and ensuring that you have like good snapshotting and all of these things and protocols to interact with different agent harnesses.
27:25And so we've built a lot around that. But it's also the sort of thing that we kind of know that the space of complexity is going to grow arbitrarily over time as people start getting more deep into these different domains. And I think we've tried to design everything in a way where we keep a lot of doors open, where like we have room to build features that kind of like there's there's base layers of like generic environments. Then you can go. I want a coding agent environment. I want a coding environment with a sandbox that's global across the run. or I want one per rollout or there's lots of these different branches you can go down.
27:57And so we kind of have anticipated like there will be a lot of these branches. There's some that we've built for. There's some that I think we're kind of ready to build for when we need to. And we also like think a lot about how do you make a good developer experience where let's say think about just documentation or skill files for agents. People are going to be using coding agents when they're building these. And so there's a lot of institutional knowledge that gets built up when you're doing this kind of research both for like a specific project or as a research team doing large-scale training runs.
28:25And being able to surface this information, some of it is directly in the product, some of it is in the way we design the library, some of it is in documentation or skill files that get shown to agents. And this is the sort of thing where we've designed it to be something that can evolve over time as the research literature evolves, as the best practices for different types of complex agents become more clear. I assume you guys read Sutton's Learning from, what is it, Age of Experience? Age of Experience, yeah. Scale AI era data labeling. Do you think that constructing environments is sort of the natural successor to that?
29:01Yeah, I mean, it very much seems like it kind of already is, where it does seem like a lot of the focus from the major labs has shifted to they're still using a lot of human data. And so human data doesn't seem to be going away because creating these environments kind of by definition is for things that models aren't good enough at yet, which means the humans are better than the models in some capacity. And so identifying which pieces of information the human can most uniquely assist the model in improving its skill on, I think, is really the key to target, which is how do you create this information flow from the human who really has the expert knowledge on how to do some sort of thing?
29:36what does a job well done look like and get it into the model. And it does seem to be that RL is the most effective way to do this right now, where having tasks with prompts to grade with an LM judge a rubric for what success looks like on a task is kind of the paradigm that's emerging for a lot of these cases where a human, sometimes it's the reward model, but kind of the most direct visceral version of it is there's a set of questions about yes, no, was this done in the LM's answer? and that turns into the reward score. And so that requires a lot of human data. Okay, awesome. Why create a hub for environments?
30:15Yeah, I think the hub idea started by seeing a lot of different, like open source reposers out there that had like overflowing implementations of all those environments in a sense. And yeah, and in general, like we already or we already created a nice verifiers framework before we even started the environment hub, right? Where you had different examples for environments in there. And it was very much an approach to standardize the whole process even more. And beyond that, just beyond just sharing those environments and having an open source platform for it, it's also about having a place where you can build a lot of infrastructure around them, which you can't really do if you just upload them to a GitHub repo or something like this.
30:57So having proper evaluations, integrated with your environment, so you can like immediately test them across all frontier models as one of the features people are heavily using the environment hub then for, right? You make it extremely easy to install one of those environments inside of a different trainer. So we obviously have our own like large scale trainer with Primer L, which we've been heavily optimizing for this. But yeah, we are very open there on the open source side to integrate with like a bunch of different trainers because people have different needs on the trainer side as well. And, yeah, that's how the whole idea of the environment up.
31:37Yeah. And what has the community behavior been? Like, do you see people forking, modifying these environments? Do you see them, you know, putting something 80-20 out into the public domain and then, like, you know, we're going to keep our secrets for ourself and, you know, not share that back? Like, what's the community behavior? Yeah, I mean, there's definitely a lot of people we work with who want to keep their environments private, as you understand. But the value for them of it being a hub is that they can do ablations on ones that are kind of known to be they can compare their private one versus some public one that might be on a similar type of task.
32:06Or for evals, there's a lot of value in having kind of mixed uniform implementations of popular benchmarks in a way that if you're doing a training run, you can plug in some known eval as a way to monitor the progress of your run. And so maybe you're doing well in your environment. You can see, does your environment also generalize to other tasks? And so having all of these other tasks available is a really helpful way for people to be able to understand not just their own tasks, but other things they might want the model to be good at as well. Yeah. Awesome. What are the most popular environments on the Hub today?
32:36Yeah, I think it tends to be the ones that are these kind of, one is the ones that we use as the examples in the documentation that naturally in software tends to be the most popular ones. So the Wordle stuff? The Wordles one's popular, but I think one that we see a lot of interest around and one where I think there has been the most degree of people kind of branching and forking and turning it into different versions is one that we call WikiSearch, which is doing search over Wikipedia pages. But it's designed to be this template you could use for agentic search more broadly. So there's a lot of applications where people want agents that know how to search over their internal documents or documents for a specific type of information.
33:10And having this kind of template for if you just have to swap out the documents. and now you have this environment ready to go where the rest of it is already kind of set up, that tends to be the sort of thing where we see a lot of value in people being able to bring other types of documents that they'd want to do a gender search over. Yeah. Got it. Okay. Awesome. I want to maybe shift towards future research, you know, big blue sky questions. Maybe the first one, Andre, I think is one of your angels as well. he kind of has that infamous quote of you know rl is you know it's it's amazing but it's it's quite inefficient and it's like sucking bits from a straw i guess do you agree with that and what do you think is going to happen in the research side to make rl more efficient yeah i mean i think um it's definitely true that rl is using a lot of compute to get a pretty kind of small signal in terms of pure information but i think in some ways that's part of the value of it as well.
34:09And I think one of the reasons that a lot of the labs have really focused on it is that one of the bottlenecks that's hard to scale is human data, especially high quality human data. And RL allows you to kind of trade off compute for data in a sense, where you can get a lot of value out of a smaller amount of data by using more compute. And so the supervision coming from this data is like small, but you can get a lot out of it more so than you can via pre-training or supervised fine tuning alone, as well as it's useful in cases where you don't necessarily have golden examples. So like if you have a bigger model to distill from, that's great.
34:44But if you're already at the biggest model size that you have access to, then you kind of need to go into untreaded territory. And exploration is really the heart of RL. It's how do you explore and try out different things. And maybe there are ways to do exploration that are more efficient than RL that people will kind of discover but this is kind of currently the the frontier of using compute to explore and improve capabilities and so that's what we got for now for sure and i think uh like i can't speak for andre obviously um but yeah i would be curious to hear his views like two months later after the latest like drugish podcast in a sense on um his views especially on the coding side for uh like uh logic reinforcement learning and how it actually helps there um if we look at something like like cloud code which definitely was already popular before but definitely popped up more over the last like one month, I would say, then yeah, his views might change in a sense on like the specific piece of how useful reinforcement learning can be in the coding domain.
35:42But yeah, generally, we don't think that's going to be the end in a sense, right? We generally think we want to be always at the frontier of like what comes next in terms of paradigms. Yeah, there's definitely lots of low-hanging fruit still on the like pushing agentical capabilities even further. but yeah we also like know the limitations in a sense right like some of the pieces we've been working on where we definitely see limitations is on the context side so yeah I think there yeah we just like have a hard limit right now of how much like tokens we can fit into a context and yeah I've been thinking there about ways on how to actually improve that in a sense yeah Yeah.
36:20Awesome. Switching gears a little bit. Open source, open weight models. What do you see the role open weight models play? Does your infrastructure kind of work only on or work optimally on open weight models? Could you kind of help people do post training around closed weight models? How does that all work? Yeah. So in many ways, the trainer itself is going to require having access to the weights. And so if anyone who has closed weight models wants to use it, we're happy to chat. But more broadly, the idea of the environment is general across different types of optimization. And so we can use the same infrastructure at the environment level for doing evals on closed models, for doing prompt tuning on closed models, for doing model selection, for evaluating agent harnesses.
37:10There's a lot of research that can be done around closed models just by having a way to do this experimentation. And so whether using open or closed models, you can create data that some platforms might let you upload some examples that maybe you're distilling from a particular model into another where you use the environment as this data engine. So there's a lot of different ways you can use the tools to optimize models. And do you need, need, need the weights? Like, for example, can you kind of LoRa? Yeah. So, I mean, the LoRa I would consider still part of the fine-tuning process. And that's what we recommend a lot of people do for RL anyways.
37:45And so, like, you don't need this necessarily with the full weights. But, like, you can't, I can't upload my own LoRa for GBT5. But I could bring an environment to the platform potentially. And so I think there's a lot of different ways you can do partial customization. And I would imagine that in many cases, the degree of like how many do you need or adaptor do you need full-fant tuning is going to depend on the training recipe as well as the goal of your optimization. What about can you do reinforcement learning on like the agent harness around a closed model? Yeah. So I would definitely consider this in the domain of like, like there's a world of prompt optimization that some people have been exploring in the research world that is in some ways kind of an analog of RL, but in prompt space where you have.
38:30Like the SPY? Yes, exactly. And so the GEPA algorithm got a lot of attention like late last year as kind of a what seemed to be a better way of doing this. And so we support we have support for that as well with our environments where you can do this around different pieces of the harness where this might be. What's the prompt used for a certain tool? What's the agent skill? What's the system prompt? There's a lot of these things that you can apply different types of optimization to. Awesome. While we're on the topic of DSPY, I think the DSPY authors also have this new thing that is the current thing.
39:02Was it recursive language models? What do you guys think? Yeah, we are definitely very interested in that. As I already said earlier, in a sense, we are very interested in longer horizon agents and so on and actually solving things for those type of use cases. And yeah, I've been internally thinking for a long time about how can we have models learn how to manage their own context. So right now, people are building a lot of scaffolds for context management. And yeah, we believe like something that is a bit more better, better lesson built in a sense is to have the model learn how to manage its own context.
39:40And yeah, I've been searching for different research in this kind of domain for a pretty long time. And yeah, the recursive language model research direction is one of the most promising ones in our opinion. We've been, yeah, since Alex Tsang, who's the original author of the RLM work, published it. We've been very interested in this kind of work. I've been exploring it as part of our research as well. And we had a blog post out a couple of weeks ago that basically showed using this RLM harness where you pretty much give a language model access to a variable in a persistent Python repl. So it can not have this whole context or the whole data as input in a sense, but it has it in this variable that it can then, yeah, retrieve, it can transform it and, yeah, manage the context through that and then also call other, like, sub-LLMs where the recursive part comes from to actually, like, yeah, manage it.
40:42And, yeah, that's the whole idea behind the recursive language model. We've been doing, yeah, some exploration there on this front to actually just give, yeah, current like frontier language model access to this yeah specific rm harness so not necessarily training in this harness but just giving it access to the specific um yeah um way of like dealing with its own context um and yeah it's been already shown to like improve benchmarks on um yeah very long horizon reasoning quite a bit and yeah we are very excited as like a new frontier in a sense to actually train in this as well to let the model train uh well train the model to actually use this harness.
41:23And yeah, that's what we're going to work on over the next couple of months. Super exciting. What else in the research domain? You guys have great research taste. What do you think is on the horizon? I'm really excited about synthetic data research. And it feels like there's a lot of stuff that feels like we should be able to do it. But you haven't really seen it emerge in the open in terms of creative ways of doing kind of self-reflection. And I think people talk a lot about continual learning as this kind of idea that we're going to have to get better at models kind of learning things on their own.
41:56And I think the idea of using other tricks that we already know in conjunction in different ways, things like prompt optimization, things like distillation, in conjunction with synthetic data, it feels like there's a lot of, I don't want to kind of go too deep into different directions, but it seems like there's a lot of room for exploration around having models curate their own training data, maybe curate their own environments, and understanding which versions of this are most effective for lifelong learning. Love it. Okay, we're going to close on an optimistic note. If everything goes right, what does the world look like?
42:33And what is the role that Prime Intellect serves in that world? Yeah, good question. How would you answer that on a high-level overview, in a sense? I would say we don't want to have a world where like all the like future value of AI and all kinds of verticals is just owned by the big labs. We have something where we like empower like entrepreneurs and enterprises and so on to actually not get steamrolled in a sense and like optimize their products. Yeah, better than they have the tools for doing so right now. And yeah, just enabling this and yeah, a lot more like cloud code moments, a lot more cursor for X type moments.
43:12that would be enabled to this. Every company is in the lab. In some ways, I mean, if data is the bottleneck, if having the real expertise is the bottleneck, like would you rather have the smartest person in history work at your company or someone who's been there for 30 years? And in some ways, sometimes you really want the person who's been there for 30 years. There's a lot of expertise that comes from really understanding a problem deeply and interact with it over a long time. And this is really what happens in training that is almost impossible to replicate in a short prompt where you really want the ability for institutional knowledge to compound over time, for best practices to compound over time.
43:49And this is how institutions and companies grow to be really powerful and successful is they stand on the shoulders of what they've done before rather than kind of resetting every day. And we want to have this be accessible to any company that wants to do this. And I think that's how we've thought about approaching it, especially as software becomes easier for people to manipulate, as the barrier to entry for coding becomes easier. we see the same happening for AI research. It's a really inspiring vision for the world. Thank you guys so much for joining today. You've really paved the way on environments and your Environment Hub.
44:21And thank you for taking the time to demystify what an environment is and share your vision for the future. Thanks. Thank you.
44:40Thank you.
From the publisher
Will Brown and Johannes Hagemann of Prime Intellect discuss the shift from static prompting to "environment-based" AI development, and their Environments Hub, a platform designed to democratize frontier-level training.
The conversation highlights a major shift: AI progress is moving toward Recursive Language Models that manage their own context and agentic RL that scales through trial and error. Will and Johannes describe their vision for the future in which every company will become an AI research lab. By leveraging institutional knowledge as training data, businesses can build models with decades of experience that far outperform generic, off-the-shelf systems.Hosted by Sonya Huang, Sequoia Capital




