In short
The episode argues that “robot-use agents” will emerge when general-purpose LLMs (e.g., coding-capable models) control robots via tool/code use, leveraging broad pretraining rather than robotics-only training. It reviews early work like RT-2/VLAs, “code as policies,” and the limits of in-context learning; then contrasts direct model control with harness-based systems that consolidate skills and reduce latency.
Guests
Hanmei (co-founder, Waddle Labs) and Vincent (Waddle Labs). They build LLMs + a “harness” that collects robot data and trains better models for robot control. Jay (co-founder, RoboCurve). They build physical-AI evaluation/evals that benchmark models across robot embodiments (hands, grippers, arms, humanoids, quadrupeds) and approaches (LLMs, VLAs, world/action models).
Key claims
Architecture matters less than training/data; harnesses distill in-context experience into reusable skills; computer-use/CAD data improves spatial intelligence; ICL saturates quickly and needs consolidation; general-purpose models may outperform robotics-specific ones if representations align.
Notable examples
RT-2 outputs end-effector pose (coordinates → joint commands); demos include unscrewing caps, uncapping pens, multi-robot communication, and Astra controlling arms to pick blocks and place them into a bowl using camera inputs and tool calls.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOFounders Introduction
0:23 to 0:35
Get to know the founders of Waddle Labs and RoboCurve and their focus areas.
“two groups of startups that are working at the very frontier of making robots more capable with LLMs.”
Research and Viral Videos
0:35 to 1:40
Discussion on the viral demonstrations of LLMs controlling robots.
“Maybe you guys want to briefly introduce yourselves and just say a little bit about what each of your companies focuses on.”
Early Research in Robot Control
1:40 to 4:00
Explore the early research connecting language models to robot control.
“And then maybe we can also show some demonstrations of what this can actually do.”
Advancements in Robotics and Coding
4:00 to 6:00
Understanding the advancements in robots using coding agents and policies.
“or lower Kolmogorov complexity code length to just use code and really compact code to just like, here's the code, just go.”
Data Bottlenecks and Learning
6:00 to 7:30
Discuss the challenges in data bottlenecks affecting VLA models.
“You're referring to a difference in data there.”
The Evolution of Code as Policies
7:30 to 9:30
Insight into how code is utilized as policies for robotics.
“So, you know, how does code as policies work and how did that inspire some of what you guys are now seeing as possible?”
Capabilities of Early Coding Agents
9:30 to 10:10
Analysis of the capabilities of early coding agents in robotics.
“What did some of those early ones, maybe even before Voyager, like code as policies, which I think the paper came out in 2022 at the end, right around when ChatGPT came out.”
In-Context Learning and Its Implications
10:10 to 14:00
Delve into in-context learning and its role in robotics and AI.
“And what was very surprising or what just excelled was that coding agents can do this very one shot.”
In-Context Learning in Robotics
14:00 to 16:01
Discussion on in-context learning and its implications for robot design.
“It's like, clearly we wouldn't do that, right?”
Astra's Role in Robot Control
16:12 to 23:27
Exploration of Astra's functionalities in controlling robots and learning tasks.
“We're watching Astra controlling the arms to pick a block off the table and put it into the bone.”
Show all 12 chapters
Future of General-Purpose Robots
23:27 to 28:04
Insights into the development and potential of general-purpose robots.
“Like, you know, that still we're benefiting from their lesson and we're still extremely good at code and we're still solving math problems, but haven't quite cracked the spatial intelligence needed here.”
Memory Compression and Robotic Learning
28:04 to 29:18
Explore the parallels between human memory processes and robotic learning techniques.
“And then you have distillation from those experiences back into an updated wait file that maybe isn't Astro, but maybe it's your own models.”
Transcript
Automatic transcript. May contain errors.0:00Waddle Labs Founder:One of the big surprises the last few years has been the generalizability of coding agents across different domains. And now frontier researchers are showing that this includes controlling robots. This led MIT professor Philip Isola to suggest in a recent viral essay that we may be entering the era of robot use agents, where general purpose models could make different robots more capable. So today, Francois and I invited the founders of Waddle Labs and RoboCurve, two groups of startups that are working at the very frontier of making robots more capable with LLMs.
0:35Waddle Labs Founder:Maybe you guys want to briefly introduce yourselves and just say a little bit about what each of your companies focuses on. I'm Hanmei. I'm from Water Labs together with Vincent.
0:43Philip Isola:We work on building LLMs that control robots. And we do this by doing two things. Building a harness that allows the LLMs to do this very effectively, collecting data and using that data to train better LLMs. Hi, I'm Jay, co-founder of RoboCurve. We are an evals company for physical AI. We measure everything, any robot, any model, including LLMs and also classical approaches such as VLAs, Vision Language Action Models, and World Action Models. And we evaluate all kinds of embodiments like hands, grippers, arms, humanoids, quadruplets, all kinds of stuff.
1:20Waddle Labs Founder:So in the last few weeks, videos from both of you guys went quite viral on Twitter and I think inspired. I think both of them are actually quoted in that Philip Isola essay talking about robot use agents. The videos showed things like LOMs being able to unscrew caps, being able to communicate between multiple robots, being able to do tasks like uncapping a pen, for example. And so I thought maybe what would be fun is for Francois and I to dig in with you guys on some of the research that led to this moment to begin. And then maybe we can also show some demonstrations of what this can actually do.
1:49Waddle Labs Founder:To start, why don't we talk about some of the early research around transfer between language models and robot control. So I want to talk about some of the early work on pre-training language models for decision making, and then also the RT2 paper, which established the VLA. So maybe, Jay, do you want to tell us a little bit about some of this early work involving being able to use language models to do any kind of robot control?
2:12Philip Isola:One of the earliest successful approaches of using AI on robots is the RT2 paper, where they use a pre-trained language model on web text and images and use that to control robots. And the interesting thing about that is that it's a fine-tuned version of a language model. And instead of outputting English, for example, they just output what we call an end-effector pose, which is the coordinates that you can then translate into join commands that can control the robots. In some cases, in some sense, it's very similar to what we see now with large language models, just that instead of fine-tuning, the models are good enough to just do it out of the box.
2:59RoboCurve Founder:And this is actually quite similar to the mapping that I have in my head is kind of like the COT moment. And so when we did GSM8K and you said, Susie had$8, she spent$5, how much does she have now? It had to output four hash marks and then the answer and then EOS. there was no ability for it to chain of thought. And so then we allowed it to actually have like, okay, let me go eight minus five is three. And so like, hash, hash, hash, hash, three. And so you allowed it to do this kind of thinking before it actually gave an action or an answer. And similarly now, like VLA is where basically like, in RT2 is basically like, it has to output an action.
3:39RoboCurve Founder:There's no, like, I can't allocate more compute for a more complex task. And then now I have this code chain of thought thing that I can do. And I can say, okay, even if you were using the example where you're not actually outputting the code, you're just giving it to Astra and let it think, think, think, think, think and then output an action. Similarly now, and then with code, it's a more Kolmogorov complexity or lower Kolmogorov complexity code length to just use code and really compact code to just like, here's the code, just go.
4:07Philip Isola:I totally agree with you on that. I feel like at the end of the day, it's a lot about a bit of less, right? If you give the agents So you give the AI model more autonomy and you, if you unshackle it a bit more and give it more resources, it can actually do a lot of the things that we fine tune it to do. I'm just going to hop in here as well. I think the really interesting thing about the bitter lesson here is that, um, like VLA is by design architecturally. They're built on top of language models as well, right? They can potentially reason they can write code. so perhaps the bitter lesson here isn't necessarily what architecture you build necessarily but it's what data is most useful right like people have been sort of nagging at this uh vla data bottleneck for years now and we're seeing very very slow progress and so the lesson here is maybe we we have a modality of data that we know works very well and this idea of transferring across different modalities which ham and i are super excited about like how do you take something that's traditionally out of distribution like robotics and make it something that's in distribution like maybe the way of harnessing a bit of lesson is saying, let's pick a data for which we know this modality is pretty blessed.
5:13Philip Isola:There's a lot of data, these models understand it well, and then use this as a way of unlocking a lot of other domains as well. Yeah. And to add to that, the reason why RT2 is such a good model is because it is using a language model that taps into the modality of all the data that a language model is trained on. So all those web images and web text actually improves the VLA compared to just training a specific robotics foundation model without any pre-training.
5:45Waddle Labs Founder:So you're benefiting in the case of RT2, these early VLA approaches from pre-training. I guess, what's the limitation in the, I guess, action-taking, fine-tuning approach that you can now get around if you can directly write code? You're referring to a difference in data there. What exactly is that difference in data you're referring to?
6:05Philip Isola:There's a few things that models have got much better at since RD2. Like one, I mean, they're much better tool use, for example. So, you know, one kind of data is now these models can, well, I mean, they write code much better. So now they can write complex policies as code. They're also much better using tools to explore the kinds of environments they're in, like what arms they have access to. And a lot of this comes from in-context learning. like the key difference between like these VLAs and like GPT-6 or LLM isn't necessarily the architecture, but kind of the approach you take towards training.
6:37Philip Isola:We want to be bidder less than pill, right? We want to benefit from all kinds of data. We want to pour in computer use data into our robot models. We want to pour encoding data into our robot models. But when you do that, you just end up getting what we think of as like these general purpose LLM agents. And then so to me, it's like, okay, why not just build a really good foundational LLM and then use that to control robots rather than training like a model that's specifically for robotics and that's more dependent on just robot data.
7:04Waddle Labs Founder:On that note, maybe one of the things that I think if we could rewind the clock a couple of years may have been a good sign that we were trending in a good direction to this is research on code as policies, right? Do you guys want to talk a little bit about when the research community refers to code as policies, what exactly that means? And I think it's worth putting this in the context of when this paper came out, which was coding agents just starting to work. I think this was still in the era in which people were like putting comments and auto-filling, you know, Python code blocks, which, you know, now we think of as ancient history, but this was like two years ago.
7:35Waddle Labs Founder:So, you know, how does code as policies work and how did that inspire some of what you guys are now seeing as possible?
7:40RoboCurve Founder:I probably rewind back to Voyager, where like Voyager was like the one of the first, what is, what does a coding agent mean? Coding agent requires good tool use and on the fly tool creation that's called code okay you have python built-ins let's just say i have only python built-ins those are tools and i have sort and i have like if and i have four and i have while and i have all these things and i need to like use those tools to assemble a new tool called a new francois.py right and like now that's a new tool and i get to use that that was voyager and like yeah i don't i'm not gonna say voyager was the first one to do this voyager was the most popular one to do this for minecraft and then they created tools that they could invoke on the fly uh to help them play the game better and compressed thinking and experience into a new tool that i can later invoke and then i think because everyone in 2024 had this insight was like okay if we just pour if you know daria and sam both all realized if i pour all of my heart and soul and resources and compute and intelligence into getting better at coding, we will have liftoff.
8:45RoboCurve Founder:Then I can automate the ML engineer. Then I can have liftoff and I'll have all the things. I'll solve all the things. And so then that's what everyone did. And I don't think that they had the insight. I don't think they were so insightful to know that, oh, if I did this, then we will have LLMs that would be good enough to write policies for robotics. And I can displace all of robotics with actual coding agents. And I could displace artists with like, you know, coding agents that are writing JavaScript to make beautiful, like, art. Like, I'm not sure that they had that insight, to be honest.
9:18Waddle Labs Founder:Can we talk a little bit about what some of these early code as policy methods were even able to do? So even, you know, this is well ahead of us having coding agents that are widely available and people intuitively understanding how they work. It's well ahead of people using RL methods to actually train those coding agents to be really good at tool calling. What did some of those early ones, maybe even before Voyager, like code as policies, which I think the paper came out in 2022 at the end, right around when ChatGPT came out. What was the capability of those compared to what you can see now?
9:43Philip Isola:Yeah, those CODIS policy papers, especially ones for Google DeepMind team, were incredible because what they did is they created these kind of functions, like pick up an object, lift up, or like move to a certain pose. Like they provided this list of functions to a coding agent in the form of like literary Python functions. And then the coding agent will be able to write code involving these functions to then control the robot to do very complex hacks. And what was very surprising or what just excelled was that coding agents can do this very one shot. They did not need additional robot data to, in order to work with this code, because they're already trained on so much coding data.
10:22Philip Isola:They already have a sense of like how to, what to do first, what to do second in order to like move a block into a bowl, for example. And I think it's this one shot ability, this in context exploration ability that really motivated a lot of later work to continue exploring, including us to continue and exploring how we can apply LLMs to robotics.
10:41Waddle Labs Founder:I know, Francois, you have this framework we've talked about of the various ways that learning can happen. There's going to be learning that's embedded into the weights versus being in context learning and so on. Do you want to quickly talk about that? And then maybe we can think about where all the research over the last few years has fit into that framework, especially the direction it seems to be going.
10:57RoboCurve Founder:I mean, I've been working on this experiment, and actually I haven't really crystallized it into a paper yet. But basically it's like, what is the most efficient way from an intelligence sample perspective to input a learning, let's say in this context, a state action result or state action reward tuple back into the policy. There's ICL. There is... So that's in-context learning. Yeah, it's in-context learning where I just append it. This is how most people are using LLMs. They're just like, oh, no, don't do it like this. Do it like this. And I was like, okay. And then it's still in the context and then I can remember, I can iterate.
11:35RoboCurve Founder:And that really only works if you train the LLM or at least post train the LLM to be able to learn and improve. And that self-refines a reflection like literature from way back that like actually allowed multi-turn into the training set. If you don't have that, it doesn't learn, it doesn't learn. But even then there's a limit to how good ICL can go. So I've done this experiment where like you take an LLM that we've trained on and I've held out a task. let's just say gsmak for for uh simplicity and and then i icl it and i want to measure on the val set how much it improves on a per sample basis and it improves greatly very cheaply it doesn't cost much i don't have there's no sgd so i don't have any flops and so very quickly i can adapt and improve on my val set example example example one is non-monotonic improvement which is wild so it gets worse it gets better it gets worse it gets better like pretty grat aggressively number two is that it caps out very quickly and so after like 20 30 maybe 40 examples it is basically saturated and more examples back into the context does not improve you're like you're constrained by the
12:43Waddle Labs Founder:model's ability to intelligently use all of its context exactly from the post training how many
12:47RoboCurve Founder:multi-turns did it actually get to be able to improve and do self-reflection and then the most important one it for surely can't do is after it hits the context window context length of the model that it was trained on if it was trained on 100 000 really that means you have a context window of like 50 ,000. Once you exceed 50 ,000, you don't improve anymore. You actually just get worse because the model can't attend over everything. There's a really good example you said, which I didn't really think about is like, well, if I can rag or I can load in my act into active memory, and this is very prime agent, continual harness kind of like thinking where I can pull stuff in on the fly, similar examples, then it's like me, like, hmm, I have an exam or I have an exam with a textbook.
13:29RoboCurve Founder:I'm going to do better if I have the textbook to look up as reference. And so that is definitely going to work. And it also scales, I would imagine, much better. And then the last two paradigms would be LoRa with rank one or rank two or whatever, LoRa with rank 10 or rank 100, and then all the way to full SFT RL. And if you're Tesla and you have all the data, infinite data, and you're learning, doing self-driving car by ICL, what are you doing? Like, are you kidding? It's like, clearly we wouldn't do that, right? And if you're figure or something like that. But it's amazing how good in low data regimes you can do with ICL.
14:11RoboCurve Founder:And then even cooler, we can compress the ICL into tool use, into tools, or like compress it into learnings. And that's the hierarchy that we haven't really figured out. that you guys are probably on the forefront of.
14:22Waddle Labs Founder:Yeah, do you guys want to elaborate on that a little bit? Because I imagine that's a central part of how you guys think about building Waddle.
14:27Philip Isola:There's a few things that are quite interesting here, I think, to unpack. Yeah, the first thing is we've been sort of thinking about this idea of a harness almost as a form of domain specificity. Like when you, for example, deploy a robot in a new environment, maybe it's in a wet lab and it just needs to do a lot of tests you're picking up, for example. You can learn a lot of this in context, but then one way of consolidating this context for future agents is sort of packaging each of these skills that you've learned into specific programs for example like writing out these skills and write out these memories is a form of consolidating it's it's like a form of distillation from past experience for your future agents to use so this is quite interesting the other thing that's quite interesting is it almost seems like there's some relationship to the broader meta-learning literature throughout in machine learning history.
15:18Philip Isola:Like you have this sort of bigger model that then programs a smaller model to do things. And one of the very interesting things is that, yes, these smaller models are going to learn in context. They're sort of wrapped inside the harness and they're doing a task. You can make them maybe smaller. You can make them run faster as long as the bigger model can sort of do this domain specialization of your harness quite well. So this is something that we're pretty interested in right now. And last point as well, I think the really interesting theoretical question is how far in-context learning can get you.
15:47Philip Isola:And I think it's really anyone's guess as to that. There's papers that show that these models sort of can approximate gradient descent during in-context learning as well. So this line between weight space and symbolic space, I think, is a little bit blurry.
16:00RoboCurve Founder:YC's next batch is now taking applications. Got a startup in you? Apply at ycombinator.com slash apply. It's never too early, and filling out the app will level up your idea.
16:12Waddle Labs Founder:Okay, back to the video. So what exactly are we watching here?
16:16Philip Isola:We're watching Astra controlling the arms to pick a block off the table and put it into the bone.
16:22Waddle Labs Founder:So here Astra is using just these camera inputs and presumably it's aware of the robot that it's controlling. And it's going to write code that controls the robot at its individual joints.
16:31Philip Isola:Yes, so it does look at all the cameras. It has all the camera feeds. Instead of code, it's more like a tool call. So it sends commands to the robot to control and put the block into the ball.
16:46Waddle Labs Founder:So in this situation, we're using Astra directly. But if we were using, say, Waddle's harness, what's the kind of difference between directly using Astra to control this versus doing this through Waddle's API?
16:56Philip Isola:Sometimes having Astra directly command the post a robot should go to is not the optimal tool to use. If it's a repetitive task, you don't want Astra in the loop. maybe you want to write code that can run repeatedly very, very fast. Or if it's a task, you've done a similar task before, you should be able to call a skill that has compiled and use that skill to do the task faster and handle edge cases better. So here we can see it seems to be approaching in on this guy. Oh, very dramatic. Yeah, you can notice that the latency is a bit slow because we are bottlenecked by the latency of Azure. But if you look at the trends, the latency of these models are improving very rapidly.
17:34Philip Isola:Right. So one thing that we saw is that for Babel class LLMs, their latency is improving by around 2x per month, which is very, very fast. Very fast. Yeah. If the trends continue, we could get real-time control by end of the year.
17:50Waddle Labs Founder:So now that we saw this task go from end to end, why don't we walk through what that actually entailed? So there's a coding agent in the loop here running this. Yes. What are the steps that the coding agent would have taken? And let's kind of contrast the direct coding agent version versus the waddle harness in the loop version. Yes.
Read the full transcript
18:06Philip Isola:So what just happened was that in a series of turns, Astra received the images from the cameras and it outputted the end effect of pose for the robot to go to. And it does it over multiple turns to complete the task. This is less of code as policies, but more of tool calls. the unscreen part is pretty repetitive so you could actually ultimately a lot of that using code as well as actually approaching the bottle picking it up like a lot of these things are pretty deterministic once you've done a task quite a few times the interesting part is where you build in the variation into the code as policy graph and so we have a few points of variation this can come from for example when you detect an object maybe you use a vlm as far as your tool call or when something fails for example how do you check that it's failed?
18:57Philip Isola:How do you then potentially do something else depending on the failure? Like there's kind of more flexible responses. We tend to put a VLM inside a loop to ensure that, you know, while the coding graph is sort of deterministic, there's points of variation that allow it to generalize.
19:13RoboCurve Founder:The biggest insight I think I've had a change in worldview on how to perform machine learning and how to get the rest of the distance on AGI was like a lot of conversations I've had with Francis Chalet. And in 2020, maybe even 2018, when he did On the Measure of Intelligence, he talked a lot about transduction is just wrong versus program induction. And transduction just means I'm learning a function theta that maps from X's to Y's. And why is that wrong? It's just slow and it's information inefficient. And so to go from X's to Y's, I need a lot of pairs of X's to Y's. And if I have a very small amount, then you need a lot of inductive bias and you what you need to act is a good a generator to map from x's to y's so now my theta takes in three and n pairs of x's and y's and emits the function that maps from x to y's yeah and that's what code is and so that's like if you give me a coding interview even whiteboard you say okay here's a coding problem here's some examples okay cool write the function and then we are emitting an f that will map from x to y
20:16Philip Isola:and and this is quite interesting right because i think i mean this was a lot of the traditional um program synthesis literature and i feel like a lot of reason why those methods didn't work as well was because that inductive bias as you said right it's you're shifting the difficulty of the problem from finding that mapping to finding the right inductive biases to map it to this like smaller code to then do the mapping but finding that set of inductive biases is super hard but maybe that's what astra is buying us like even in cocci there's a giant shift in the literature from you know, specific neurosymbolic methods to just using Astra to write code, to map.
20:49And so maybe that's a cool way to think about it.
20:51Philip Isola:Which is, it is neurosymbolic. That is neurosymbolic.
20:53RoboCurve Founder:We have neurons and then they're emitting symbols. Yeah.
20:55Waddle Labs Founder:I think this is actually a pretty good segue to a kind of final topic, which is around, inspired by this paper also from Philip Isola that he titles the Platonic Representation Hypothesis, right, where platonic, because as a reference to Plato's cave and, you know, the idea that, you know, they're seeing shadows of, various realities. And I think here the point he's making is that there's extensive evidence to show that language models and various representations of different data actually learn these distance mappings between similar things under different training policies that are overlapping.
21:33Waddle Labs Founder:And so like the idea there would be that as you train these systems on increasingly large amounts of data, they converge to a kind of consistent mapping of the world. And maybe this implies that we would expect language models to get better and better at things over time that would make them useful for new tasks like robot control. But maybe I imagine this representation hypothesis is actually pretty central to all of your guys' worldview or your view of your companies. And so maybe why don't you guys tell me a little about how you think about this. And also maybe we can use that to make some guesses as to why we think models like Astra seem to be so much better at robot control than previous models and where, you know, make some predictions for where we might be going here.
22:13Waddle Labs Founder:Yeah.
22:13Philip Isola:I think one way to make this concrete for the robotics models versus language models debate is that the very, very strong language models will have very similar representations of the world with very strong robotics models. And if that's the case, then if you have a really strong language model, you also have a really strong robotics model using, if we believe in the platonic representation hypothesis. And if that's true, then that is the best manifestation of the bitter lesson, right? Because you just need one really strong model regardless of architecture and they would be outperforming any specific models that is slightly weaker.
22:56Waddle Labs Founder:The interesting thing that's maybe not entirely intuitive and maybe I'd be curious to hear you guys all think about is why specifically it seems like these newest models, specifically Astra, seems to be so much better at spatial intelligence. You know, we've all seen the demos online of, you know, controlling Blender and making great 3D images. Now it seems like, you know, you showed Jay in a benchmark that is like a meaningful step function improvement on certain tasks. Where do you guys think the pre-training and post-training, I guess, of those models, what has likely changed about OpenAI's approach there that made them so much better compared to models even six months earlier?
23:29Waddle Labs Founder:Like, you know, that still we're benefiting from their lesson and we're still extremely good at code and we're still solving math problems, but haven't quite cracked the spatial intelligence needed here.
23:36Philip Isola:I think what Astra does incredibly well is its vision capabilities. It was probably pre-trained on way more computer use data than ever before. It's probably pre-trained on so much CAD data. And it's like all of these kinds of data probably teach the model similar understanding of the
23:55Waddle Labs Founder:physical world as a lot of robot data might. So the computer use data is an interesting point because it's not totally intuitive, I think, why computer use data is useful for understanding the physical world and that it tends to be like a computer where you're clicking around. Say more about why you think that is a big unlock because I totally buy they trained way more computer use data than ever before. Right.
24:13Philip Isola:I think immediately if you were just training a model for computer use, that would not be able to apply for robotics. But if you feed a computer use data into a big model like Astra and by computers data I mean like you drag a cursor around on a screen to orbit some cat object in order to design in blender This tells you Like how to reason about spaces at least tells you about like top down left right all these concepts that you need to control a robot Right. Um, and I think that's why this data helps so much for making these LLM so much better at robot use I think like from the other direction, like a lot of people in the robotics community were trying more and more different kinds of data as well.
24:50Philip Isola:Like egocentric videos kind of went from just teleoperation to like, you know, a broader range of this data because it's not just robot data that can teach a model how to use a robot. So if we take this to the extreme, it's like, why not feed every kind of data coding, computer use, egocentric into the same model? I think that's how we get to the most capable like robot use agent.
25:10RoboCurve Founder:I think it's so funny that like, you know, if you go to 1980s, your X park or whatever, like we literally made the graphical user interface to be more like the physical world so that we could interface with it. And like what we ended up doing is building an environment that was actually helpful for robotics to learn how to use the physical world. Right. We have file systems. We have like files, we have folders, we have, you know, the GUIs for like SolidWorks and like Autodesk to like spin around stuff to make it similar to the physical world. and then like we couldn't get robots to work in the physical world so we just trained on that and
25:45Philip Isola:now it works right exactly there were incredible papers uh i think from princeton and also cloud place robotics touches on this it's like where you just if you design a right harness where you make the tools that you face with the robot look like computer use tools like you have a agent drag a cursor around to control where the robot goes that improves how well the llm uh is able to perform on these physical tasks.
26:05Waddle Labs Founder:Where do you guys see, you know, given like a reasonable guesses as to where the base models are going to continue improving and your guys' own investments in either evaluating these models or building harnesses around them, what do you think is going to be capabilities that we maybe now see as challenging to do, but that are going to be increasingly possible or even trivial very few months from now? Like, you know, six months ago, even the demonstrations that I've seen you guys post on your Twitter, I think would have been kind of mind-blowing to imagine come from LLM.
26:32Philip Isola:There's some consensus within the Frontier Labs and also in the Robotics Foundation models companies that we will have general purpose robots within the next two years or even earlier. And this is something that society is probably unaware of or even unprepared for. And when we say general purpose robots, we mean something like if you give any natural language instruction, it can do what a competent teenager could do with their bare hands. It's robotics in terms of capabilities where it can generalize to unseen tasks and unseen environments.
27:06Waddle Labs Founder:For you guys, for the Waddle Labs guys, what does that mean for what you guys are building?
27:10Philip Isola:I think we're going to want to take steps to get there in about two years time. I think there's so many challenges that are very visible. For example, when I let people online talk about latency, this is a big problem. If you just have Astra being in the loop thinking at every step, this is really slow. It's not going to be economically useful. So how do you consolidate they kind of first pass by Astra, like this contextual learning into like a faster skill or policy that you can then run repeatedly at incredibly high throughput.
27:40RoboCurve Founder:I think it's the mappings to humans will get more and more like this, where like the optimal thing to do may be to put a new SAR, state action reward back into context because it's very quick. But then there needs to be some like go to sleep for a while, like almost everything that is intelligent intelligent sleeps like tell me an intelligent system that doesn't sleep right in some way and then during sleep compression happens the weird thing that happens from your hippocampus and it shortwave ripples to both lobes and like there's weird you know pass from memories that were compressed throughout the day to train the weights and similarly maybe what the right thing to do is similar to dagger a data segregation framework in our classic rl where you're going you're collecting a bunch of data and then you're somewhat reflecting on it and then you're using to update your wait file.
28:28RoboCurve Founder:And then you have distillation from those experiences back into an updated wait file that maybe isn't Astro, but maybe it's your own models.
28:36Philip Isola:It sounds very Dreamcoder-esque, but a lot of those specific tools I think that people used to build, to take Dreamcoder, for example, right? It had this library of skills. And then during the sleep phase, it basically refactored everything into a more compact representation and so on. I mean, we're seeing sort of similar things just happen not as rigidly as before in the space of programs. but now it's maybe like maybe you want to refactor traces maybe you want to refactor skills a lot of these robotic coders policy papers right they have this growing ladder of skills and it's anyone's guess really how you print them how you organize them and things like that so yeah i think as as waddle goes forward like how you manage our growing context of skills of deployment data all that is
29:17Waddle Labs Founder:going to be quite interesting i think with that i think this was an excellent discussion thank you so much Vincent, Tanming, and Jay for being here. I'm very excited for all the incredible robotics advancements I think we're gonna see over the next few years and I think you're totally right, Jay, that I don't think broader society is totally aware of how much is coming and it's gonna be a very incredible few years to come. So thank you so much.
From the publisher
One of the biggest surprises in AI over the last few years has been how well coding agents generalize beyond software. In a recent essay, MIT professor Philip Isola argued that we may be entering the era of robot-use agents: general-purpose models that can control different robots, write policies, and learn new physical tasks with little or no robot-specific training.In this episode of Decoded, we're joined by the founders of Waddle Labs and RoboCurve, two of the startups whose work helped drive this realization. They're working at the frontier of using general-purpose models to control robots, and their recent demos helped inspire the growing conversation around robot-use agents.
Together, we dig into the research behind that idea, from code-as-policies and vision-language-action models to the harnesses and evals needed to make these systems work in the real world.




