In short
Podcast Episode Summary: Building the "App Store" for Robots with Thomas Wolf
Podcast Title
Training Data Description: A series hosted by Sonya Huang, Pat Grady, and Sequoia Capital partners that delves into AI with leading builders and researchers to explore the evolving technologies and their impact on technology, business, and society.
Episode Title
Building the "App Store" for Robots: Hugging Face's Thomas Wolf on Physical AI Description: Thomas Wolf, co-founder and Chief Science Officer of Hugging Face, discusses democratizing robotics through open-source tools, datasets, and affordable hardware. He shares insights on Hugging Face’s project, LeRobot, and its implications for the future of physical AI.
---
Key Takeaways
Introduction to Thomas Wolf and Hugging Face
- Role: Chief Science Officer at Hugging Face, significant figure in advancing transformers and language models.
- Vision: Transitioning the community-driven success of transformers into the robotics space.
LeRobot Project
- Goal: Democratize robotics, making it accessible to a broader audience by providing open-source tools and affordable hardware.
- Components: Combination of software libraries, datasets, and hardware to empower software developers to become roboticists.
- Community Building: Aim to expand the robotics community similarly to how AI research has grown.
Current State of Robotics
- Inflection Point: Wolf argues that robotics is at a pivotal moment akin to the rise of language models a few years ago.
- Challenges: Data scarcity in robotics compared to language models. Robotics often lacks the vast datasets available for training language models.
Community Engagement and Growth
- Rapid Growth: The LeRobot community is expanding, with thousands of developers actively contributing and creating datasets.
- Types of Developers:
- Traditional roboticists frustrated with existing limitations.
- Software developers transitioning into robotics due to interest in AI.
- Hobbyists and educators exploring robotics through accessible platforms.
Hardware and Accessibility
- Initial Focus: Started with software, which led to hardware developments.
- Affordable Robotic Arm (S-100): Designed to be a low-cost entry point for budding roboticists ($100).
Data as a Bottleneck
- Main Challenge: Limited data diversity in robotics leads to difficulties in generalizing robot behaviors across different environments.
Future of Robotics and AI
- Market Maturity: Discussion of a potential "iPhone moment" in robotics where consumer robots become common.
- Diverse Use Cases: Anticipation of various form factors beyond humanoids, including more accessible and varied designs for different applications.
Open Models vs. Closed Models
- Current Dynamics: Open models are gaining traction, especially in China, due to competitive pressures and a culture of collaboration.
- Hugging Face's Role: Focus on enabling open-source communities while adapting to the evolving landscape of AI models.
Vision for the Future
- Ten-Year Outlook: Envision a world where everyone can engage with AI, much like the shift from passive media consumption to active content creation. The goal is to empower developers and hobbyists to innovate with AI tools.
---
Discussion Points
- Open Source in Robotics: The role of open source in fostering a collaborative environment for both software and robotics development.
- Safety Concerns: Importance of running models locally in robotics to avoid catastrophic failures due to connectivity issues.
- Comparative Analysis: The contrast between the growth trajectories of AI and robotics, with a call for more engagement and community-building in the latter.
Final Thoughts Thomas Wolf emphasizes the potential of robotics to reach diverse audiences and the importance of making these technologies accessible. By fostering an open-source community and addressing data challenges, Hugging Face aims to lead the way in democratizing robotics, ultimately reshaping how people interact with and utilize robotic technologies in their daily lives.
---
Conclusion This episode showcases the intersection of AI and robotics, highlighting the transformative potential of open-source initiatives and community engagement. Thomas Wolf articulates a vision for a future where robotics becomes as ubiquitous and accessible as software development, paving the way for innovation and creativity in the field.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00many, many startups just already being built on top of LoroBot. Just, you know, they want to build something. They have this idea of a manual test they can automate or they have an idea of something that could do in a physical world. And then they take LoroBot. They take already like the basic building blocks with SHIPT, which is just a very simple robotic arm. S -100 that we designed basically to be the cheapest LoroBot account to be $100. And they already like trying to start business around this. At the bottom, you know, it's the again face editors, which is you bring all this platform, all this basic belly blocks for people to be already crazy things on top.
0:34A robotrix is the same goal for us.
0:52In this episode we interviewed Thomas Wolf, co -founder and chief science officer of Hugging Face, the largest open source community for AI. Thomas has consistently predicted the future, and he's responsible for why Hugging Face invested heavily behind Transformers and Language models, enabling the LLM way a handful of years ago. Now he sees the same opportunity for robotics and physical AI, and is helping to shepherd hugging faces low robot project, which brings together policy models, data sets, and physical embodiments to help developers everywhere become roboticists across the diversity of use cases and form factors.
1:27We're excited to chat with Thomas about the robotics explosion, and the bottlenecks ahead for physical AI, the US versus China in the open model race, and a lot more. Enjoy the show. Tommas, in your role at Hugging Face, you help Hugging Face invest behind moon shots as the Chief Science Officer, one or two years ahead of the field. You mentioned to me last time we chatted that, you know, you have a spidey sense that we're at the same moment for robotics today that we were for Transformers and Language models a handful of years ago. Tell us what you're seeing. Yeah, I think honestly, I studied two years ago.
1:59I mean, we studied activities in robotics 18 months ago. And I think at this time, there was a couple of breakthroughs out of, I mean, these labs, you know, were a way to stand for them, basically, these things. We're starting to show robots that were able to tie notes, fall clothes, you know, cook, like through things in the air on the pan and grab them. Basically, all of this thing, we've in one way very little data, but also with also good perspective to be able to leverage some of the world models and things like that that we see really benefit from internet size data. So all of this kind of pointed to a short future where robotics were going to work in a way like hardware was already there and in my opinion it's been there for quite some time but the missing brick was really software that could adapt, that could be dynamic all of that.
2:48And we started to see that and that's why we started to work on the robot a little bit more than 18 months ago. And I think for us the success of the robot, like seeing like this huge community. For us, the big beds was, can you build a big community in robotics as well? There was small community of kind of hobbies or like, or people very seriously building robots for like, you know, factory lines, but it was kind of like a tiny vertical in my opinion. The question was, could you move this tiny vertical to like a full horizontal thing? Just like nowadays, every software developer is kind of an AI researcher almost.
3:22They all want to know how, you know, LLM works, how you train them, and there is this very smooth transition where the 100, 200 millions of software developers are becoming kind of AI aware. And I think there is a potential future transition where all of these people also become robotist in a way, if you give them the tools. And so that was kind of our goal. So we started with the software library and the success of that kind of brought us to also try to go in hardware, which is a big question as well. And that's why we acquired first hardware company, DC Apple and Robotics. And we shipped our, At least we opened the orders for our first robots.
3:56There was mid -Julys, that's one month ago. Really? And this went crazy. Can you tell us what LeBer robot is? Of course, yeah. So LeBer robots is our attempt of doing again the success of Transformer, reproducing the success of Transformer's library in the robotic field. So this idea is having kind of a thoughtful library that everyone would use, and that would bring in a very simple and accessible and easy way, all the latest technology, all the latest algorithms that people use to train robots, efficiently all the latest data set that they use to train this and also connect these to actuators, which is the hardware part of robotics.
4:36And that's this intersection and this mix of three aspects, the policy, the models, the data set, and the hardware which we try to combine in robots. But how does the role of hugging face change in robotics? For people building in the physical world, does hugging face play the same role or a different role than hugging face is played for people building in the digital world with LMS? I mean, our goal is to play the same role, which at a very, very high level is building communities and bringing people in society that AI can be open source and it's not something you only consume, but it's something you can tweak, you can train, you can control, you can host where you want.
5:14And actually hosting where you want is even more important in robotics because in the future where you have robots everywhere You kind of want a lot of these models to run locally because if your robot lose connection to the Wi -Fi or something and then run in the wall Or maybe I don't know run in your kids. It's gonna be much more dramatic than you know, just an LLM Elitinating so safety question in robotics. I think I have a good reason you may want to really be able to not you know depend on the distant API but have the models as possible to the hardware. So I would say our role is maybe even more important for safety and ensure for robotics than it is in 11.
5:52Could you say a word on the size of the community you have at Lerbaat? How many people are building, how many people are contributing data sets and things like that? Yeah, for sure, should have checked actually the latest number because it's exponentially growing. So it's several thousand people, I would say six to ten thousand. One event we did a couple of months ago, we did a hackathon, which was worldwide. We had a hundred locations in over six continents. So it's still not in the million of people, for sure, but it's way above several thousand people. And the main indicator for us is we can measure the number of data sets, for instance, on the hub.
6:26What we see in all of this thing, number of community members or data sets, we see kind of this exponential growth, which is, I think, very good indication that we are on the right track. And you have to keep in mind that the hardware that's available right now is still very much kind of a hobby hardware. So it's like 3D printed arms. It's like they're still wired everywhere. And that's why if starting this summer we wanted to bring much more mass market hardware I would say. So something that would interest not only the hacker and the people who are used to plug cables everywhere but also everyone in families, something that looks much more polished.
7:03What's the persona of developers in your lower bot community? And I'm curious how it's the same or difference from the people that have traditionally been building, you know, like controls, classical base systems. I would say there is three, three type of person now. I mean, one is the, the traditional robot is this, they think you want to use AI. So for a lot of them, they know how to build hardware, they know what they can use, but they've been frustrated by, you know, the limitation of the software stack, all the optimal control models. And all these are very limiting what you can do. So all of these are people who have really happily joined the band bag and we see many, we see the same effect we saw in transformers, which is many academic labs starting to use robots because it's it's it's a very nice entry point for all less students.
7:47And so and so this has been growing very strongly. The second community is is is much more interesting in my opinion. It's people who were not really into robotics, but because they they're into AI and robotics looks like a physical manifestation of AI, I kind of want to go into robotics. And so these people that covers, I would say, software developer, but even people who are just interested in robotics. So a good example of talking here is interesting, but a lot of investor have actually bought, you know, S -100 arm, just to try to understand physically, what is this robotic thing, what can it do?
8:19And because it seems so accessible, you get the arm and the software is just a Python curve. And now, you know, with a little bit of vibe coding, you can actually even tweak it or control it quite easily. We see people who maybe are not purely technical, but who want to understand what's happening in robotics and they use this entry point, which is low robots. So you can vibe code a robot? Yeah, let's read my goal. So you can already do that a little bit, but for the new robots, reach the meaning, I definitely want this to be one of the easiest way to use. I would love my kids to be able to vibe code a girl on the robots.
8:51What sort of phase of maturity do you think we're in for the robotics market writ large? You know, like, what do we have a chat GPT moment in the world of robotics? Yeah, that's the thing I'm looking for. I like sometimes called it also the iPhone moment. Maybe what would be the first use case? The first, you know, the first moment where everyone or like a large fraction of population will we think I want to robot in terms of consumer. I mean, I think the enterprise market is quite complex. There is in some place there is already a lot of robots in some kind of industry. The car manufacturer is a basic sum.
9:23Then there is this second part where there is like an entry of robots and here there is a lot of challenge around reliability. Will these robots be reliable enough to be deployed in retail store and basically be really useful? And the third part I'm much more interested in is actually entertainment, fun, demo, education where I think maybe this question around, I need a $3 ,000 robots because I need the our less pregnant and so you can take a robot that's really accessible. So which menu, for instance, is priced as $300? That's something that can definitely be like an impulsive buy. You buy it as a gift and you're not sure it is going to work on that.
10:06But for this price, what we want to find is, is there not a lot of potential in more entertainment, fun, education, learning AI through a physical interaction stuff? Instead of just coding on the chat but are coding on the screen. And I think that's something that has not been explored at all. I think there was a couple of tried, I mean, the mid media lab, since you had the G -Ball, for instance. But in the past, they were using their price to be high, I think, above 1000. And more importantly, I think the software that was there was very limited. So you would buy a robust, that would be firm, but you had maybe five or 10 behaviors.
10:48And once you've tried them all, that's it. And it's finished. And here the goal for each chimney is really to make kind of almost like a smartphone. So it comes, you have a couple of behavior, but just because you can tweak it and people can build new behavior and share them and plug all the new VLN speech models, chat model, activity to possibility, arcane, endless. It's kind of a open door on like rebuilding, you know, the up store of iPhone basically. So that's what I'm very excited. And to be honest, these last parties, they're very much a bad because nothing exists there. Like there is no, real proof.
11:21My major sense are all this exponential growth of community which make it quite plausible. So you see Richie Minnie as like the you know the re - incarnation of the robot dog dogs of the 90s was like how people can actually play and experiment and you know have like these robot companions and households. I mean to be fair this one is a big bet but something I was discussing just yesterday actually on robotics at Tech Babacu. Someone was telling me as an investor, you know, how just is so many, many startups just already being built on top of LoroBot, just, you know, people who are like, they want to build something, they have this idea of a manual test, they can automate or they have an idea of something that could do in a physical world.
12:05And then they take LoroBot, they take already like the basic building blocks with SHIPT, which is just a very simple robotic arm, S100, that we designed basically to be the cheapest robotic arm to be $100. And they already like trying to start business around this to start like to say all that something around this and Richie Minis also in a way designed for that it's a very wide simple robot kind of a white labeled thing and if you want to adapt it and if you think hey I have a business idea around this but I need a robot to interact with people I don't know in hospitals or something like that you can take this and you can actively start to build your A.
12:40And let's get at the bottom, you know, it's the again phase ethos, which is you bring all this platform, all this basic building blocks for people to build really crazy things on top. And robotics is the same for us. Really cool. I'm going to talk about data as a bottleneck. Like I think one of the big differences between language and robotics is you have trillions of tokens out in the public internet to train all amps that dynamic doesn't exist in robotics. And actually I think that's where hugging faces role and ecosystem could be much more interesting in in terms of decentralized data sets, curation, creation.
13:13Talk about what's happening on the data set side of LoroBot. Yeah, it's super interesting, I think. So, I mean, there's a couple of challenges in LoroBot X. And I mean, the general, the main challenge is in the data. There's just not enough data. There is some ways to use video on the internet as training data, but it's very limited. And in some way, we may be able to use models, but in some other if you want to automate a test, there is no way around just recording someone or the robot possibly just doing the task. I think here there is one possibility on one limitation. I mean, the main limitation is you can record a lot of tasks yourself, but usually what you will lack a lot is the diversity.
13:56So you will basically be able to train a robot to do something very well in your room when everything looks the same, but once you put it in the next door room where maybe the walls are green instead of red, the robots has a lot of troubles to general likes. So this is the main limitation. So our idea with the herb was that everyone could record data sets and if we managed to insensivize them to share the data, then we could maybe build a very multi -colored data set that would be extremely diverse. Any addition hopefully would be also very big. So I would say that's a long -term goal. We hope this can help.
14:35But another more direct thing we try to do is to work also directly with the actors of the community. So we release a couple of data sets to try to help them releasing some data set. We think in robotics, one of the nice aspects is a lot of people in the end want to sell the hardware. And so they can actually afford to even modern LLM. They can afford to share a bit of the software as open source if it brings all the field above because in the end, that's not really directly what they sell. So that's kind of what I'm trying to convince a lot of robotics company to do. And surprisingly, it's actually something a lot of them seems to be interested in.
15:13Super interesting. You tweeted out world models today. I think you and I met the same world model founder. What's happening in the world model open source and how does that help or not help? What's going to happen in robotics? And can I ask on what to do? Is there a why now in world models at the moment? because it feels like they've started to pop up recently. So what is interesting is it's, it feels like it's a couple of teams who have been actually working on that independently for, for Fumarth and they just happen to release this right now, right? Because when you talk about all of them, they are not really coping each other.
15:45I mean, I guess one thing was the, the advent of really cool, really good image generation and basically finally understanding how to fix this six finger things and basically just get a more reliable and more coherent world model for image, which naturally was transposed to video. And so we see now some really cool video models as well. And this one is just one next step. And a lot of the founders I've talked to in this field also say that they were helped by the advance of like open source video model generation and open source image generation. And basically they take this video generation model and then they fine tune them and then they train them to be able to react to some inputs, which is also what we do in robotics.
16:29There's a lot of common points between these two things. And it seems to work quite well. So you start to have this kind of very interesting in my opinion, totally new experience where you actually have a film that's controllable. that's like both photorealistic and also reacting in a very coherent way to the action you will input, which is either just moving around or just asking it to add something, you know, add a rider, a castle, a car driving, and you see this thing that just reacts very well. And you have here, I think, a lot of potential applications, you know, both obviously an entertainment, actually some form of entertainment that might be just totally new, something we've never seen, which is maybe the first time we create a new form of really unusual entertainment, but also a lot of application in business and how you can have interactive things.
17:19And one of these downstream applications is generating more data for robots. I mean, there is just two ways to generate data. One is to record it in the real world, which I think is still very interesting. And the other is to simulate it. And surprisingly, on the simulation, we have not seen a lot of, you know, I mean, There was some development, but it's not like they've been some really breakthrough recently on simulation. So maybe this is the first breakthrough I've seen on simulated generated data in quite some time. Yeah, I was very excited to see some of the, even like what DeepMind's doing with Genie to train their Mollied robots.
17:53Super exciting. Humanoid. Do you believe in humanoids as the kind of ultimate form factor? Yeah, big debates, big debates. What is sure is that I'm quite more excited about trying other form factor right now. I mean, the main problem with humanity, I think there is too many problems. The first one is it's always quite expensive, just because you need a lot of models and all the price in the robots is just the actuator. That's always like 70 % of the price tag. So when you have six actuators, that's just your bill is the year. And so it's really hard to drive humanity below the price of a car. I think the price of a car is still already quite a high requirement.
18:30If you buy something that's a price of a car, you do expect to have a lot of value out of it, right? And so that's why we're exploring smaller robots, like just one arm or just moving head and stuff. There is some possibility that we can get cheaper humanity at some point. And K scale was trying to do. You need trees they've been trying to cut the price and there's a lot of company trying to aim, but it's gonna be really hard, I think to get this under like 10k, 10k, $10 ,000. The nice thing about the humanity, of course, is once you've sold humanity, you solve like a lot of tasks at the same time.
19:04So if you solve the imaginary, you can do everything a human do, which is very exciting. And the main question is do you need to solve the human reads? So on my side, I'm more like I would like to see a galaxy of different form factors. I also think some of them are much more cute than a humanoid. I think for social adoption, I think the humanity is also asking a lot from people. It's like this, you're directing this kind of uncalee value with something that looks a lot like you move a lot like you. So I thought this would be a big limit for social adoption. Now to be honest, I've seen a lot of Unitry robots and then about you, but you kind of ignore them at some point.
19:41So I'm also much more confident on the people who just say, yeah, that's just maybe we too worried about it and can be very in robotics and maybe at some point once we start to, I've seen like a couple of robots, people we just accept them very, very easily. Okay, so we're going to see the low robot humanoid soon. I mean, the goal would be three chimini and our small robots work really well. At that point, we'll climb back to make it a humanoid form factor, which I kind of progressively as we've done, bringing the community along with us. As you imagine, the world in 10 years, like, how many robots do you think there are among us?
20:16Do you think it's like 80 % of the more humanoid than 20 % of this long tail of this diversity of hardware in use cases or like, how do you think the world plays out? Yeah, and I would love to see the second option because I think that's an option where we have much more robots in our life. What I would really not be super excited about is the future where kind of robots are kind of a elite thing because they cost 100 ,000 and so basically if you're rich, you have three robots at home and if you're not, you don't. I mean, hugging faces always always been also about the big communities that we care about that.
20:48And so for this reason, and much more excited to see a lot of form factors that are basically accessible to a lot of people, some of them are cheap, some of them are expensive, then just this single humanoid that cost a lot and if you can buy, that's nice and if you cannot do that for you. So I would say, I talking phase as the future we try to manage to first or I think also is much more fun because you in a way you're also restricting yourself, you're just like LLM if you just try to make them copy you man is one thing but if you try to think maybe they can do something that you and cannot do, it's also much more interesting in a way.
21:22Do you think we're heading towards a world of big foundation models that can kind of do everything and then be adapted quickly to any new domain with just like, you know, a few prompts or do you think that developers in your community are going to start from like a small base model and then do a lot of their own data collection, customization to adapt to their domains? Hmm. I think we'll see more and more both. I mean, I think as the field evolve, you know, we start to see really a long tail. So, for instance, if we take the downloads on Hiking Face, we see both very large state of the art models being downloaded, which are usually too large to run on a local laptop.
22:01But we see also some of the most downloaded models are actually just the right size that fits to run quickly on the laptop. So we see these two like really modality. And I think as the fill can match here, we'll start to see this more and more, which is it's not like you choose one or the other, It's just depending on what you need. You may choose locally or not. And I think GPT -5 with the router is a good example of this. Maybe the largest model are the most reasoning. The longest reasoning chain is not the answer to everything and you actually need to smartly select the one you want. So it can be behind the router.
22:36But it can also be just locally. You will run some models here. They might be extremely useful. And we know better and better how to train model that actually is extremely useful. But when you need something much more complex, when you need reflection for a long time, then you will turn to much larger models. One of the narratives that's been really popular over the last few years is this narrative of the battle between open and closed, you know, closed models versus open models who's going to win. And just in the last few weeks, OpenAI is now present on hugging face. And so I'm curious what to make of that.
23:11And sort of what it might imply about the future of open versus closed or maybe how they work together. I mean, we're super happy to welcome the back. They were there. I mean, first model, our work done. And the reason we switched from being a game company to an open source platform was GPT -1, which not a lot of people remember, but it was very funny because it was trained mostly on novels and romance novels. And so when you would put two characters in the continuation, they would always fall in love and some way on They shoot still miss a little bit this one and then and then Google took this idea and train it also on Wikipedia Which had a lot of world knowledge and then he expanded to GPT all of that But at that time there were very very pro open source and I think I think open source just like in software I think both both solution which is we just go exist and having company that open source both or that do both I mean, Google has been an example for quite some time, right?
24:09With the Gemaline and Gemini line. And some interesting moments were some time I heard that one Gemal model was actually so good that it was better than closed source models. So they had to not open source it. So the front here, what is sure is quite seen at the moment in a way. And the challenging new players mostly in China, but I think we start to see also some new foundation models team in the US. So I think we might see also some challenger in the US. The frontier we stay quite seen, I think, and both both think we'll see with a tiny difference of performance. The main reason right now, I think, is to be honest, at this exact point in time, I think we are not exactly in a kind of cost saving time of AI.
Read the full transcript
24:51So which means that for a lot of actors, I think moving to open source because it's safe costs is not the most important thing for them. So usually they move to open source right now because they're on data privacy, they want to be able to adapt the model, they have maybe a new idea or new, new, like this, this action model, for instance, they have a new idea of something that does not exist and they want to do that. So that's usually what we see right now. So we see a lot of a mullion of new, of new, of new exploration in open source way. What I do expect is as, as we go to a more mature market as well, then the cost and being able to run it maybe on faster hardware or this type of other hardware and then being able to out the model and to out also the full stack of where the models run with become actually more and more important.
25:36So I think just like in software, I think in the long term, open source is kind of a winning solution for many applications for many years' case. But we're still in the like the two -balance place. Yeah. How do you think hugging faces role in the LLM ecosystem has evolved as, you know, these models have pushed at the frontier and there's closed models. I remember back when it would be like, you could download the small brick model on Hugging Face and run it locally, right? And that was, there was a lot of the usage. How has your business evolved now that we're going towards, you know, as you mentioned, models that are too large to run on consumer hardware?
26:13And how do you see Hugging Face is really evolving? Mm. I mean, surprisingly, I was doing these stats at the end of last year. This burst model is still really used a lot. So a surprisingly interesting aspect of open source is also resiliency, which is once you have something that worked, that really work in production, you may not want to be forced to move to the new GPT, right? I mean, that was a little bit of the backslash around GPT 5, you know, just people actually wanted to keep using GPT for all from anything. Maybe they fell in love with this or were the main friends and some Reddit posts was around this, but also maybe they just had their like applications, probably well, and I don't want to redesign it.
26:51I think open source, I mean, the long term interest for SEOs to provide is very stable base, like you build something, you know, it will exist. And you know, you can keep this as a very stable base. And in general, I think in the community, our role has switched progressively from maybe pushing ourselves a lot of things, pushing our library, pushing our early product to more like enabling more and more the community in general. So we work now a lot with many many actors of the community. We work a lot with LMS CPP. We work a lot with VLLM We work a lot with all the big players to try to see you know how How this all ecosystem can be very efficient can work really well So like one model is released you want to be able to use it directly in VLLM You want to be able to use it directly in LMS CPP.
27:37So we try to have more and more this kind of role of Meta community builder where we try to align and to bring maybe all the players as at the same pace and to help them move in the same way. So in a way, we are much more focused on the community, on the hub than we were maybe a couple of years ago. Really cool. What do you think about what's happening? You mentioned China's had a lot of the open models recently. Like, why do you think that is happening? And what is the state of open model development in the West? Yeah, this is the most surprising thing that happened. and I think in the last two years, right?
28:16The fact that China would become a champion of open souls would have predicted that in 2020, right? And so I've been actually visiting them two weeks ago to try to understand a bit better on the ground, how it's happening. And the thing is just, it's a very, very competitive market internally. There's a lot of teams there that are extremely good, and it reminded me in some way of Silicon Valley, people are working extremely hard and they compete with each other. all of this model provider. One part on which they compete, which is surprising, is being the most open, is the open source aspect. So they are extremely proud of being very open and some of this company, and as top being open when was called Z -Poo, they decided to not open source it.
29:01And they saw an immediate black slash, I see mostly on hiring, like people, they don't want to come work there anymore. And so they went back to open source. So it's quite strong now, I would say strongly ingrained. So I would expect this to continue. I would expect also quite more team to come because I see a lot of, you know, I mean, we see that as well, right? When you, when you're the presentation of GPT -5, like a lot of people are, you know, actually, did doisters, some of them at Sincwai University, right? We know the, the team are here also, you know, partly partly with Chinese members.
29:33So they have extremely, extremely strong people and they already want to train the best model. What I think is interesting is to see the West coming back to open source very recently and to be honest, just over the summer, right? But this call for open source scene, OpenAI decided to come back. Now we're just waiting for on -tropic to maybe open source their first models. I think it's time to try to ask them to participate. Yeah, I would say right now, the situation for open source is pretty good. But yeah, it's never like it's like the Jedi and Star Wars. It's never one. We have to keep pushing this.
30:13We have to keep pushing a flag of openness. What's driving the resurgence of open source in the West? I think one thing is when you have in a way nothing to lose, open source is always a good solution when you're new team. So it can be, for instance, you create a new company and you want to quickly rise to the top, the new open social model, right? That's the mistrial recipe. How can you very quickly become a great player? But for the Chinese, for instance, it's also, for instance, in the West, almost nobody will use a Chinese API. So they don't sell API in the West anyway. So in a way, they have nothing to lose, you know, from the Western market by open sourcing their market there.
30:52So I think there is this thing as players. So and the consequence of that is also that when nobody's open source is like a market, there is an interest for someone to take the room, right? Say we're going to be the open source player. So meta was this open source player when everyone kind of stopped open sourcing. And I feel like they will always be this thing when some people stop open sourcing and then there is actually a gap to being the only, you know, the new top open source actor. or then someone we want to fill this volume. Thomas, you mentioned that Western companies won't use a Chinese model over Chinese API.
31:29What about, are you saying Western companies are actually willing or not willing to use Chinese open models when it's the weights and hosted on US servers? Like, is there still hesitance to do that? And is it well -founded or not? I don't see that a lot, to be honest. I mean, it's a good question. And I try to do regular polls. So I try to ask a lot of people regularly, you know, what do you think about that? Because it can be for sure a concern, right? There is when deep sea came out and there was very nice Anson's dog model from from from perplexity for instance um The thing is in many business cases.
32:05I don't think people read notice anything, you know So I think there is more than a general appetite for people to have a better way to understand um like the safety of a model. But every day it's quite general. People are a little bit worried about having a model that maybe will behave strangely in some cases. And so I think this is a general thing that a lot of company have been asking, which is can you guarantee this model will always behave well? Which we know is really hard, because even with GPD, sometimes you ask a number of R in strawberry, and it's just behave badly, and you're like, why?
32:40You're very smart, you should be able to know that. So I think this is a general thing that is needed soon. And there is a couple of teams working on that for sure. Can we talk about open science? Yeah, we build LLM like human, but what if an AI model could see infrared, could see some radiation we cannot. Does it think that human cannot do? So it's already superhuman. And for science, it's actually super interesting. So a lot of the AI model for science are already superhuman in a way because they can actually as a similarity or pretty thing that are just inaccessible to human. and I think it's a good ground to think outside of the human limitation of what we can do.
33:22You've been pretty passionate about open science for a while. So can you just say what about what is open science, what role does Hagenfeece have to play, and what is your passion for it come from? For me, it started a very long time ago. So before I was loyal, but before being a lawyer, I was a researcher in physics. And so I was working on this superconductive material. And surprisingly, in superconductive material, a lot of the great research had been done by the Soviets back in Soviet Union. And these people, I mean, those Soviet researchers have a very different way of inventing theory than the Western world has.
33:58And so they had some really great ideas and some really interesting thing. But I had to find this invention or this theory, I had to find them out, to track them down in the Soviets, GTP letters. and some of them were even still in Russian. And so from this time, I get this idea that, damn, accessing knowledge is hard. And if I can make this easier, that's gonna unlock a lot of really cool stuff. If I could just find where, you know, that this equation comes from and really be able to read this article, that would be crazy. And so when I joined computer science, I discovered archive, I discovered open source.
34:31I was like, this is really cool. Everything's just free. Basically, everyone just share things, reading in English, everyone can read it. It's even free. You don't even have to buy the publication. And I was very excited about that until I started to try to reproduce one deep -mind paper. And I discord that there was a limit because people publish what they want to publish. But they don't really give you all the tricks of the trade. And so when you try to reproduce that, they discover that it just doesn't work. And so open science for me was this extension, which is, it's nice to give open models to people so they can build things on top.
35:05but it's even better to explain how to train a model. It's a thing that's nice to give a fish to someone, to feed them, it's even better to teach them to fish. And that's basically what we want to do. We think in a very long term, AI is going to be just a fundamental technology that basically should be, just like physics, should be something everyone could learn by reading a book. If you want to learn today about general relativity, you can read a book and you can know about it, right? You don't have to pay to get access. I mean, you buy the book or you'll find it, but that's basically free access.
35:34I think AI, all the recipe to train an intelligent Object or artifact should also be something that everyone should know. So I mean, that's a very long term thing. But the very short term thing is if we teach people how to train great models Then they bring great models on the hub and then we have much great content to offer. So it's kind of all also just content Providing if you provide great model is nice. And so one example that we do For that is we we write very long blog posts We, that some of them even become books. We just published the books this summer on how to train on a thousand GPU and how to balance the load and how to do all of this parallelism thing.
36:11Another very long blog post we wrote was around how to make a very good quality data set. And so we made a data set to three trainer models called FineWeb. And it's using a lot of the recent model, the Quen models, the set of the used fine wave, for instance, you know. And then we also wrote how we build this data set, how we filter it, what is important and to understand when you want to be a great data to train models. So all of this, I think, just go together. And for us, it's a way to basically bring just better, better open source AI models in a honey phase. I want to go back to your physics and superconducting comments.
36:44Like it feels like a lot of the AGI labs believe that AI actually disrupting science is not that far out. There's been some exciting discoveries. I think, well, I think there's been exciting evidence so far in math. and then maybe extending into physics, material science. Do you think we're gonna see an inflection point in scientific discovery from these models? And what do you think open sources roll will be and driving that? Otherwise, there is some hype here because there is drives people, but I think sometimes we overestimate what's happening. I mean, math is a good example, right? There was this idea, A .I.
37:22is doing a new proof for some math theorem, and that's like inventing your science. I think as a scientist myself, I think that's really the wrong way to view science. The reason is I was a bad scientist, so I can tell about that. So I was a very good student, so when you give me a problem, I'm always pretty sure I can find the proof. I can find the thing. But I know this thing has a solution, so I just have to fill the gap and kind of grab a couple of things I know and then combine them together. And when I became a researcher, I discovered that I was a pretty bad researcher because was basically what I was not able to do was, I was not able to ask the right question.
38:01So if somebody asked me a question, say, can you demonstrate this theorem? I could do it. But if someone say, okay, what is interesting to explore now in math, I had no idea basically. And so in science, the main thing you need to do, if you want to do, I am talking about big, big breakthrough, right? Is you need to ask the right question? You need to find a way to ask a question. Nobody has asked before and a question that will open a whole new field of research. And that's basically a noble field, a noble price. A noble price is typically someone who just opened a new field of research because this person just asked the right question is maybe, you know, maybe the speed of light should be the constant.
38:39And let's explore what does it mean? It means actually we can create general relativity and then we can invent black hole out of it. And I think LLM right now asked here, extremely bad at this thing, at this kind of a tasteful way to ask the right question. which doesn't mean we cannot do really cool stuff with them. But the way I see them nowadays is really more as very useful helpers. So once you have human research, or we say this is something interesting to study, then you can use them to actually multiply by 10, 100 or 1000 the production you can do. You can use them to quickly do a full survey of what has been done in the past on this molecule, this protein.
39:18You can use them to say, okay, what would be the most logical way to test this hypothesis. But I still see this as kind of accelerator on the assistant of scientific research, then what I would love to see, which is an AI, that would say, hey, I have an idea on how to go faster than light. But for this, you cannot just write the answer on how to go faster than light. You have to ask the right question, what should we change to today's theory? What should we do today? What should we reconsider to invent something that has ground breaking? Do you point out asking the rare questions? What do you think are the interesting questions in the world of AI right now?
39:59Or maybe the questions that people are not asking that they should be asking? I mean, this is one question, I think. And it's related to something we talk a lot, which is this, um, CCO -Fancy, which is tendency of AI models to always agree with you. I think, I think a good researcher is actually a good example of a person who is a grieve for a lot of people. My former professor was a Nobel Prize, was very not friendly in how he could wear this. But I think that's part of it. You have to be extremely open and angry. So finding a way that this, you know, to push this model to have maybe in a way stronger opinion or maybe a taste in the opinion or like, I think for science will be a key.
40:40And it may, of course, this will be based on deep learning and LLM, but it may involve other way to train them or the way to think about them. I think that's one of the big questions. I think not a lot. There is a couple of people exploring that, but not a lot of people exploring. Okay. When you see the world in 10 years, like what is hugging faces rolling at? How much of your community you think is building with LLMs, with robotics? And it's hard to think in 10 -year time spans, but what do you think the world looks like in 10 years? Yeah, 10 years is very, very different. I mean, what I would love to see is in 10 years, the world where basically everyone feel like they can build with AI, and not just consuming AI, but they feel like they can be actor of this thing.
41:26A little bit like, you know, the difference between... We used to have a lot of media that were generating and created for us, and then we moved to the current era where everyone is actually able to create media, we saw that this created a whole new generation of people and YouTubers and influencers and people like to be making extremely interesting content. And I would love AI being the same, which is a very big community, like the software developer community, where everyone can create things with AI and they feel like it's just another tool in our box. They can code stuff, but they can also train a model and they can maybe adapt a model.
41:59So the nice thing about that is a big believer in basically the creativity and natural invention of just the community. I think it's something that's very beautiful to witness. So in 10 years I hope that people are not just consuming AI content and not doing anything, but they're actually exerting their creativity to build a really nice thing with a lot of AI tools around them. To be honest, that's something that I think that's kind of what we're building right now. So I'm quite optimistic. The thing is going to change a lot of things for society in general because a lot of this drug will just be different.
42:31It's a beautiful vision. Thomas, thank you so much for joining us today. We really enjoyed this chat. Thanks. It was a pleasure. show."
From the publisher
Thomas Wolf, co-founder and Chief Science Officer of Hugging Face, explains how his company is applying the same community-driven approach that made transformers accessible to everyone to the emerging field of robotics. Thomas discusses LeRobot, Hugging Face's ambitious project to democratize robotics through open-source tools, datasets, and affordable hardware. He shares his vision for turning millions of software developers into roboticists, the challenges of data scarcity in robotics versus language models, and why he believes we're at the same inflection point for physical AI that we were for LLMs just a few years ago.
Hosted by: Sonya Huang and Pat Grady, Sequoia Capital




