Open Source Self-Driving with Comma AI

16 Apr 2026 · 46 min · 22 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Comma AI’s OpenPilot open-source self-driving/ADAS stack and how it delivers highway autonomy via an add-on device, trained with simulation and a “world model,” plus what’s needed for broader robotics (controls, RL, continual learning).

Guests

Harold Schaefer, CTO at Comma AI; has worked there for nine years and has focused on OpenPilot. Hosts: Daniel Whitenack (CEO, Prediction Guard) and Chris Benson (principal AI and autonomy research engineer).

Key claims

OpenPilot is the most popular open-source self-driving stack (and among the most popular robotics projects on GitHub). Runtime is on-device: a neural network takes camera video and outputs longitudinal acceleration and road curvature, then a car CAN API layer commands steering/gas/brake. OpenPilot training uses simulation with photorealistic, input-responsive diffusion-based video; imitation learning alone fails without recovery training. Comma aims for “shipping intermediaries,” not all-or-nothing autonomy.

Notable examples

>50% of user miles driven on highway with system driving; Tesla FSD and Waymo discussed as benchmarks; green-light detection improves ~2x with a 10x larger model (and ~100x compute via external GPU).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Introducing Harold Schaefer

0:45 to 1:44

Discussion of Harold Schaefer's background and role at Comma AI.

“I'm CEO at Prediction Guard, and I'm joined as always by my co-host, Chris Benson, who is a principal AI and autonomy research engineer.”

The Evolution of Comma AI

1:44 to 3:04

Harold shares the journey of Comma AI and its products over the years.

“So Kama makes this device, like you said, that you can install in cars and gives them autonomy features that they hadn't had before.”

Self-Driving Landscape Over the Years

3:04 to 4:30

Comparison of self-driving technology from the inception of Comma AI to now.

“And also like, I guess, on the commercial and closed source side and the open source side, what did that look like kind of then compared to now if you look back on that journey?”

Attraction to Self-Driving Technology

4:30 to 5:51

Harold discusses why he was drawn to the self-driving field.

“We have over 50 % of miles driven with people that have our system are driven by, you know, the system and not the human.”

Understanding Autonomous Systems Architecture

5:51 to 7:49

Harold explains the components of self-driving systems like OpenPilot.

“I'm curious, as you were getting into this, what was it about this particular problem that attracted you?”

Functioning of OpenPilot

7:49 to 10:50

Detailed breakdown of how the OpenPilot system operates in vehicles.

“like OpenPilot is maybe a piece of that and fulfills a role.”

Intermediary Solutions in Self-Driving

10:50 to 14:02

Discussion on how Comma AI addresses self-driving capabilities in existing vehicles.

“I don't know if there's anything I missed there or any questions we can discuss there.”

Developing a Scalable Driving Agent

14:02 to 15:10

Learn how to create an intelligent, scalable driving agent that can be functional across different vehicles.

“we create some kind of end-to-end solution that's scalable and can be the most intelligent driving agent that can drive better than a human.”

End-to-End Solutions in Autonomous Driving

15:10 to 17:28

Discover the differences between end-to-end models and traditional methods in autonomous driving.

“And I think that's just generally a bad approach.”

The Role of Simulation in Training Models

17:28 to 20:49

Understand the importance of simulation for training AI models in autonomous driving.

“back end, but they're also trying to shift fully towards end to end.”
Show all 22 chapters

World Models and Real-Time Planning

20:49 to 24:26

Learn about the implementation of world models in simulation and their role in autonomous driving.

“That might be a first place to start because even before training the policy model or the model that you end up wanting to use, right?”

Advancements in Edge Computing for AI

24:26 to 28:00

Explore the evolution of edge computing hardware for autonomous vehicle applications.

“You also mentioned having a data center.”

Model Improvements and User Experience

28:00 to 29:04

Learn about the enhancements to self-driving models and user interactions with the system.

“I mean, so we just started working on this, but our tests currently indicate that it recognizes green lights in nuanced situations twice as much if we have a 10x bigger model.”

User Experience with Kama Installation

29:04 to 31:05

Discover the installation process and user experience when utilizing Kama in cars.

“And how do I like typically if I, for example, imagine I'm driving a car with lane assist, right?”

Exploration of Use Cases Beyond Driving

31:05 to 33:18

Uncover potential applications of self-driving technology in various domains.

“I can, I can put it over here and over here too.”

Challenges in Robotics and Machine Learning

33:18 to 34:46

Understand the ongoing challenges in robotics, particularly low-level controls.

“principle a similar thing as moving a steering wheel and learn about it in the same way.”

OpenPilot Development Decisions

34:46 to 36:53

Learn about the motivations behind creating OpenPilot and its open-source nature.

“It will do some weird internal logic that we sometimes don't fully understand.”

Choosing Programming Languages for Development

36:53 to 39:36

Explore how decisions are made regarding the use of Python versus C++ in development.

“When it comes to inter-process communication, I think OpenPilot does this better than anyone else, including ROS.”

Addressing Controls and Reinforcement Learning

39:36 to 41:58

Delve into the unsolved problems in controls and the role of reinforcement learning.

“Ultimately, most of the development happens in Python.”

Continuing the Discussion on Control Systems

42:02 to 42:48

Learn about the complexities of continuous learning in self-driving technology.

“If you inflate your tires, it will affect how open pilot drives.”

Future Aspirations for Autonomous Technology

42:48 to 43:59

Explore potential developments and personal visions for future technologies.

“Love the fact that it's open source so we can really talk about it in depth.”

The Desire for Practical Automation Tools

43:59 to 44:58

Understand the speaker's desire for practical, open-source household automation.

“Sure, but it's not going to be that visionary.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:01Welcome to the Practical AI Podcast, where we break down the real-world applications of artificial intelligence and how it's shaping the way we live, work, and create. Our goal is to help make AI technology practical, productive, and accessible to everyone. Whether you're a developer, business leader, or just curious about the tech behind the buzz, you're in the right place. Be sure to connect with us on LinkedIn, X, or Blue Sky to stay up to date with episode drops, behind-the-scenes content, and AI insights. You can learn more at practicalai.fm. Now, on to the show.

0:41Welcome to another episode of the Practical AI podcast. This is Daniel Whitenack. I'm CEO at Prediction Guard, and I'm joined as always by my co-host, Chris Benson, who is a principal AI and autonomy research engineer. How are you doing, Chris? I'm doing very well today. How's it going? It's going great. I was commenting to our guests today just before we started recording that earlier this year, I was in the car with one of our engineers, shout out to Ed, and he's like, hey, have you heard about this cool thing we're driving in the car right and he's like hey there's this cool thing you can put in your car and make it like a like a ai ai assisted driving car without it being like a specific self-driving car and uh and so he he uh forwarded me the information about comma and i'm really excited today to welcome harold schaefer who is cto at comma ai welcome harold Thank you.

1:44Thank you for having me. Excited to be here. Yeah, yeah. Well, obviously, I kind of alluded to some of what you're involved with, but maybe could you give us just a little bit of background about yourself and Kama and kind of how you ended up in this spot of working on some of the things that you're working on? Sure. So Kama makes this device, like you said, that you can install in cars and gives them autonomy features that they hadn't had before. You know, things like auto steer and better ACC. So on the highway, you kind of get some level of autonomy. And, you know, the software that runs that is OpenPilot.

2:25And that's a completely open source autonomy stack for cars. By far the most popular open source self-driving stack online. And I think it's currently even the most popular robotics project on GitHub. So that's kind of where we're at. I've been working on this a really long time. I've been at Karma for nine years now. So it's basically been my entire life. I don't think there's much to say about my professional life that's not related to Karma. But I came to the US like 10 years ago, graduated and started working here. And so I've been working on OpenPilot and this type of stuff ever since. So what was when you started that journey, what was kind of the state of both like self-driving autonomy when you started things?

3:13And also like, I guess, on the commercial and closed source side and the open source side, what did that look like kind of then compared to now if you look back on that journey? So when I joined, we didn't have a product. It was just a project. And you could kind of install the software if you went through the hassle of installing like a beefy laptop in your car and installing all the power. Or you could like retrofit a phone that you'd have to do all that yourself. So that was the state the project was in when I joined. The company was pretty young at that time. George, the founder, had I think worked on it for a little over a year, maybe two years at the time.

3:52And that's the state it was in. It was a usable ADAS system, but the product side was really not that in a greater state. So that's where we were. In terms of open source, I think there is no real genuine open source ADAS product that's useful in any way other than us. That was true then. That's true now. But the commercial side has obviously changed massively in that time. 2017, when I joined, you know, highway autonomy was bad, maybe usable in some cases, but most people probably wouldn't use it because it just makes too many mistakes. It's too uncomfortable. That's obviously not true at all anymore.

4:30You know, OpenPilot is really good. We have over 50 % of miles driven with people that have our system are driven by, you know, the system and not the human. Obviously, there's Tesla FSD, which is, you know, at an even higher level of what kind of things it can do. And you can get over 90 % engagement there if you really wanted to. And you've got Waymo, which is like supervised robo taxi. That's like an actual product. You know, unclear if they're a money making product, but it is a real product that people can use. And obviously, none of this stuff existed in 2016. As for how I got into it, I think it's worth mentioning because this is kind of the journey that I just talked about is I remember when I was, I must have been really quite young at the time, but I saw the Chris Ermsson TED talk about Waymo where they had that we had a closing statement that his kids were growing up and he was hoping they would never need a driver's license.

5:26and I think there must have been like, I don't know, less than 10 at the time and they have a driver's license now. So I think that's the timeline we're talking about is from that prediction about Waymo being able to displace the need for people to have a driver's license. That's obviously didn't manifest, but we did make a lot of progress and there are some cool stuff on the market now, even if it doesn't mean that all driving is autonomous. I'm curious, as you were getting into this, what was it about this particular problem that attracted you? What was it that caught your imagination enough to grab you for such a long period of time as well and still hold you today?

6:07What was it versus all the other things out there that you might have dived into? I mean, something that George was saying a lot of the time, and I think I really resonated with me and is still true, is that self-driving was the most interesting applied robotics problem, period. It was a place where you could make products that were essentially immediately useful. You know, we don't have 100 % reliable autonomy. You need supervision for it to be useful. But people buy our products because it adds value to their lives. This is actually quite unique in robotics. There are not many use cases where this is the case.

6:38You can have some kind of hyper-specific industrial robots are obviously useful. robot vacuum cleaners. I'm a big fan of robot vacuum cleaners. There's robot lawnmowers. And that's roughly it. There's not that many places where you can do like applied AI, applied robotics in the real world. And that's why I thought self-driving was so cool. And then the other thing is I just really like the open source nature that we have with OpenPilot that we're trying to promote. I think an open source future is just generally better for everyone. Well, yeah, I'm particularly interested in this conversation because, you know, a lot of times on the show we have really interesting people on, but sometimes they're in great discussions, but sometimes they're constrained with in terms of what level of detail they're able to talk about kind of architecture and things like that naturally, because they're maybe working on certain proprietary things.

7:33And with having so much in the open source world with Kama and OpenPilot, I'm kind of interested if maybe from a high level perspective for those out there that are listening that might not have an idea of like what sort of architecture and kind of main components are part of a self-driving or autonomous system. like OpenPilot is maybe a piece of that and fulfills a role. You have the devices. Could you just give us like a mental model for how to think about how these pieces fit together? What's where and what components are needed to make the system work, I guess? So we're talking about ours in particular, right?

8:23Or just in general? Sure, yeah. So ours in particular, we have a device that you can install in the car. The device has compute and it has some cameras. and then some other sensors like GPS and IMU that are not necessarily that needed, but they're in there. That device then runs some machine learning models by looking at the road. That tells you where roughly to drive. And concretely, that is a longitudinal acceleration and a curvature of the road. So that's like an angle of the steering wheel. I'll talk about how that works internally later. But in runtime, this is really all that happens is there is just a machine learning model that takes in the video input and outputs those actions of acceleration and curvature.

9:07And then that goes to some API that, you know, we kind of develop that interfaces with all these different car models. So we reverse engineer the canvas of the car. And so we can understand what messages we need to send to command, steering, gas and brake, and do all the auxiliary stuff of like the engagement state, you know, all those kind of things that are needed to do autonomy in a car. So that's basically what's happening at runtime. We just have a device. It runs some models. It outputs the actions to take. And then there's a car API layer that, you know, is reverse engineered for every type of car we support that sends messages on the canvas.

9:44to do the certain actions. Then those models themselves that we train, so we've been kind of talking about end-to-end training for a long time. So they're end-to-end in the sense that they take in video and they output actions. There's no intermediate space. So, you know, our cones detector, traffic lights detected, that all doesn't exist. And very high level, our training stack looks like we have hundreds of millions of miles of humans driving. We sample some of that. we teach the model okay if a human's in this situation this is the most likely trajectory they're going to take if you train that directly you get a system that doesn't really work this is like a well known machine learning problem that it's not necessarily one of them understood but you can't just do imitation learning and expect things to work in the real world you need to expose the model during training to mistakes and show how to recover from those mistakes And we do that by training in a simulator where we can introduce, you know, going off the center of the lane line and then we can supervise, OK, this is how you would recover from this.

10:46That's very high level how things are trained and how things work. I don't know if there's anything I missed there or any questions we can discuss there. Now, I was just wondering, just on that note, the open pilot project, that is part of that, I guess, the software stack that interfaces with the models and executes kind of the policy and interacts with the car API? Or does that fulfill a different role? Because it is kind of more broadly a project related to kind of autonomy and robotics generally. Or am I misunderstanding? Yes. So, I mean, the goal of OpenPilot is to be general robotics, but we're clearly focused on driving right now.

11:29But a lot of things that are written in OpenPilot are things like, OK, well, there's a UI and then there's a localizer, which is written in classical code that takes in all these input sensors and makes the best estimate of the current motion of the device. You know, that's robotics for driving. Then there's, like you said, there's this whole layer of interfacing with the car. So that's also part of OpenPilot is how you communicate with the car and manage like the state machines. And then in general, it's also an operating system, right? So there's many different threads and processes running and those intercommunicate and those need to get managed.

12:01And so that's OpenPilot. But a lot of the decision making about what happens and how to control the car essentially happens inside the neural network. It's pretty cool. I'm curious, I think the space that you've chosen in terms of addressing the need and, you know, where there is, you know, vehicles without autonomy capability altogether, traditional, below you and then kind of above you, there are the fully integrated, you know, vehicles. You know, we talk about things like Tesla's and their competitors, where they were the whole, you know, the hardware software is all working together for the whole vehicle.

12:37And that gives you another level of capability. How do you target the level of capability that that Kama is addressing in terms of when you're doing this? this, you know, what is kind of an add-on to an existing vehicle that doesn't have any autonomy capability. Like, how do you think about the problem of what you can add value to by bringing autonomy capability? And what also, like, what's the constraint? What's too much without having a fully integrated platform from the get-go, you know, that you're manufacturing out of a factory? So just to be clear, our mission is to solve self-driving cars while shipping intermediaries.

13:20And now that, you know, we're kind of, we're kind of evolving that to solve robotics. So that is what we see everything as we're trying to build solutions that, you know, genuinely do solve the AI and the applied robotics problem in a real and genuine way. That's the starting point of how we think about it. It's not that we think like, oh, we don't really start from the product or the user side. But given that starting point, we then think, okay, if we want to make progress on the robotics problem, on self-driving cars, how do we have intermediaries that make sense? We don't want to be doing research in a lab and saying like, oh, we're never going to release everything until it's done.

13:58We want to make progress and at the meantime, be able to ship useful features. So we really focus on how do we create some kind of end-to-end solution that's scalable and can be the most intelligent driving agent that can drive better than a human. But then we just kind of see, okay, given this system, can we apply this in a way that's useful to people? And then we kind of figure out, you know, which cars does this work well enough on that can be supported? Because you were talking about constraints. Like there are some cars that have certain constraints that make them unusable for open pilot.

14:31And so then that's not as interesting. Does that kind of answer your question? I'm not sure. I think so. Yeah, definitely. And I'm glad you expanded on that. I may have had too narrow a vision in my understanding of what you were addressing there. So I like the way that you are kind of iteratively solving that problem in the large. Yeah, so I mean, the distinction is, I think we're trying to make incremental steps towards some late stage solution that also have incremental usefulness to people. And I think that's kind of where we differ from a lot of other people. I think a lot of other people are like, you know, everything or nothing kind of approach.

15:10And I think that's just generally a bad approach. which I think we want to make money along the way. We want to pay our own bills. We run our own data center. And, you know, those limitations obviously mean that we run with 100x less compute in our data center than, you know, Waymo or Tesla. And that obviously has some side effects. But I think in the long run, this isn't really a big deal. As long as we're making progress towards some solution and shipping useful products, you know, I think the long-term future looks bright. And I'm not sure, you mentioned a couple times this kind of idea of end-to-end, which you mentioned was, you know, part of what you've been talking about for some time.

15:47Some listeners might not totally get kind of the implications of that. So if you could maybe talk through like an end-to-end model, as in your case, and how it might compare or contrast to other approaches that maybe have been used in autonomy that wouldn't be considered end-to-end solutions? Sure. So, yeah, end-to-end, it kind of is a matter of interpretation. When I say end-to-end, what I mean is we have examples of humans driving competently by just recording them. That data contains the information about how to drive, and we want to let a machine learning model distill that information and learn how to drive.

16:30So directly taking in just raw sensor data and outputting a policy that looks like that or some really good subset of that. In contrast, something that would not be end-to-end at all, for example, is something like some kind of segmentation, semantic segmentation network that detects where all the lane lines are, detects where all the traffic lights are, then builds this huge grid. And then you write some algorithms like, okay, well, don't touch lane lines, don't hit kids and don't crash. And then that goes through some kind of optimization, you know, that's again hand-tuned by someone and produces some trajectory.

17:04we've been focused on the end-to-end approach for a long time you know end-to-end has generally made really fast progress over the last several years so basically everyone's interested in end-to-end to some extent what people have actually shipped is more some kind of mixture now Waymo is I think slightly less end-to-end they're doing end-to-end research but they still have a lot of classical detection I think that's relevant Tesla I think has a little bit on the back end, but they're also trying to shift fully towards end to end. But yeah, I think that's kind of the state of things. As you look at the kind of the trajectory that you're currently on and both, you know, in terms of where you've come from and kind of where you're at now and into the near term future with the recognition that the vision is to solve, you know, autonomous driving kind of in the large to paraphrase you where do you see yourself now on that and kind of and also like in the way that you're approaching it why are you at where you're at now um you know you know on that path how are you envisioning your approach to this uh to a solution uh compared to the waymos and the teslas and the others of the world and the fact that they have different uh different approaches there.

18:22Yeah, I mean, the difference is a lot of it's just constraints, right? The reason we're so passionate about this end-to-end thing is because it requires less human effort to leverage the same levels of capability. Waymo and Tesla for a long time were able to get capability by doing things like having humans hand label things in scenes, by, you know, piping back data that showed a lot of uncertainty and then humans would label them. They had like this whole data engine thing. You know, they all have, they've used many different approaches to patch the holes in what is, I think, everyone's kind of vision of this end-to-end thing.

18:57Whereas, you know, we're constrained. We don't have that sort of, we want to be profitable. We don't have that sort of money. So we've been focusing on this idea since the beginning because we think this is the right long-term strategy. I think most people see this as the right long-term strategy. But I think because we're so focused on it in just strictly the strategy sense, We're just, you know, slightly more sophisticated, but, you know, in the capability sense, I think not quite. I think so one big innovation that we have, and I think we're the first people to do this at this point, is our models are trained in simulation, but they're not just they're trained in a machine learning simulation.

19:30So, you know, those like video generations you see like Sora, all those types of things, you know, we train our models now in a diffusion simulator. So the videos are all generated. Again, Waymo and Tesla are exploring these things, but they haven't been as focused on it. So I think they've not quite shipped this yet. And they also have a higher bar to meet before they can ship that. But for us, this is the most efficient way to make progress. So I think we're just we're slightly closer to this kind of envision of the strategy. Could you dig into that a little bit? Because I think that is a really interesting point.

20:01And I think some people have maybe on the hype side of things heard mention of, you know, a world model or or other things like that. You know, you talk about your world model in the most recent release blog post and this fact of how, you know, the agent that you've released is maybe the first to your knowledge that is kind of fully trained in this learn simulation. Could you could you just pick apart that a little bit for, you know, maybe more on the practical side for for our listeners in the sense of like, what what what do you mean by a world model? How is that problem actually, like what are the challenges associated with that problem actually creating that world model?

20:49That might be a first place to start because even before training the policy model or the model that you end up wanting to use, right? Somehow you have to, if you're going to train it in that simulation, you have to have a really good simulation. So, yeah. Yeah, so, I mean, I talked about in the beginning, this is just an assumption that we have to make. is that it's impossible to just do imitation learning. You have to do some tricks on top of that to get something to recover from mistakes and not drift out of the lane line. And everyone has kind of different strategies on how to deal with that.

21:23Like I said, I think we think the long-term strategy is you just need a really good simulator. And so our solution has always involved training things in simulation. It's just our previous simulators were like classical. They would estimate depth and then reproject depth. This has artifacts. You make some assumptions that don't hold. So you want a really good simulator that can capture the world completely. And obviously, if we think about end-to-end again, the end-to-end way to do is just tell a machine learning model, okay, try to simulate the world. The difficulty here is, okay, first of all, you need to make video that looks somewhat realistic because otherwise you have all these artifacts that can be exploited.

22:00And I mean, it's only in the last couple of years we've gotten anywhere close to that. And then the other big challenge, which is where we differ from most video generation models is if you want to make like a robotic simulator with this approach, you need it to be accurate in terms of responding to inputs. If you tell your simulator, okay, the car turns left 10 degrees, the simulator actually has to then produce video that reflects that left turn 10 degrees. It's not enough for the video to just look realistic if it doesn't respond accurately to the inputs. Let's say those are the two challenges in making a simulator.

22:37It has to look photorealistic, has to be obviously as diverse as the real world, and then has to respond accurate to inputs. So that's the challenges we were working on.

22:48The photorealistics kind of solved for us, right? A lot of people are working on this. And then I think where we kind of had to, where our stuff kind of deviates from all the published papers and stuff is trying to make it respond accurately. I'm curious, as you implement the notion of a worldview, how are you approaching it? Is it more of a training mechanism or is it something where you're doing real-time planning, you know, in route? How does it fit into the overall architecture and how are you guys using it? The world model? Yeah, yeah. Yeah, so the world model acts as a, in our case, it acts as a simulator.

23:26So you can just initiate, instantiate it on some real video. And then basically you take control in simulation and you give to the world model actions like turn left 10 degrees and then the world model will produce the next image. So that's one aspect of how it's used. Additionally, the world model also supervises the recoveries during training. So the world model that's producing these images also sees the future. Like there is some real future attached to this scenario that we're making. And so we can then introduce deviations and the world model who sees the future can also say, here is a likely trajectory to get from this deviated state that we are now to this future.

24:10Sorry, I'm trying to explain this in a way that's not extremely confusing, but it is just extremely confusing. but that is how our training training stack works is that the world model simulator produces the images and supervises the recovery trajectories that then are eventually what gets into the car and on that point of kind of the trajectory to what gets in the car one of the things that might be helpful to highlight like you mentioned like we mentioned different models the world model the model that's on the car the kind of harness around that, if you will. Am I correct? You also mentioned having a data center.

24:52Am I correct that that data center and centralized infrastructure is mainly geared towards that training and simulation? And then a lot of the real-time inference for individual cars and decision-making, does that happen on device? or? Yeah, exactly. So everything at runtime is strictly on the device. The device doesn't need any connectivity other than updates or to send data back so we can use it to train. And the data center, we don't run any user-facing services there. It's just the training data center that runs all these different training things. And yeah, just to be clear, so there are two main models involved, which is the world model that acts as the simulator that we train inside of.

25:40And then there's the agent policy, which is in comparison, a really tiny model. That's the one then that trains inside the simulation and gets shipped to the devices. Yeah, that's interesting. What have you, I guess, devices have obviously advanced. So like in addition to some of these things, like your ability to create more realistic imagery for like the simulations and that sort of thing, obviously hardware or maybe capabilities of running models within kind of edge environments has advanced over the years since you started this. Could you give us a little bit of a picture of that world and maybe what's different now versus when you started this, both in terms of what you need to do to run this sort of model in that kind of edge environment, maybe the tooling or ease of doing that now versus years ago?

26:34Sure. Actually, on our product and specifically, there's not been as much progress as in other places. We still run a relatively old chip. Because our device sits on the windshield, there's quite a limitation to how much heat we can generate. So we haven't followed really the trend in how much these other systems have increased their compute power. That's something we're now trying to solve by we're going to sell an external GPU that you can plug in that can, you know be under your seat or something that way we can kind of match the compute that things like FSD and Waymo use right now we use probably less than a hundred x a hundredth of of what what an FSD computer has there's definitely been increase in efficiency so you can you those those cars have better chips than they used to but a big change has just been that they put in more power hungry chips than they used to I think there's been more of a recognition that it's worth spending watts there.

27:26I'm curious as you're looking at that as a possibility in terms of upgrading the hardware, the processing capability on board, what is an immediate term type of capability that you would add to your stack? So compared to like, if I'm looking at the Coma website and seeing the demos that are listed there, if you could like have a quick hit on it by having extra compute, What are the things that you guys are talking about, at least in the open, that you can share with the public that would be good for that? I mean, so we just started working on this, but our tests currently indicate that it recognizes green lights in nuanced situations twice as much if we have a 10x bigger model.

Read the full transcript

28:15And an external compute GPU like this would give us a 100x bigger model in the limit. So that's roughly the type of improvements I think you can expect. And that's also roughly, I think, the scaling that exists, which is, you know, we're a hundredth of the compute of an FSD computer. An FSD computer is more capable, but on the highway, you're not even really going to notice the difference. You really need, you need exponential increase in compute to have like these marginally noticeable gains. But I think, you know, at 100x with the external GPU, that's roughly the things that we're talking about.

28:48is like twice as reliable at detecting, you know, lights in weird nuanced situations, or maybe if we optimize it a bit more. And could you talk us through, I guess, more of the user experience side within the car? So if I'm installing Kama for like, what does that experience look like for me? And how do I like typically if I, for example, imagine I'm driving a car with lane assist, right? I know it's on maybe via an icon and, you know, I feel the steering wheel move in a Tesla sort of autopilot, etc. There's a different kind of experience, right? What does the user experience look like from this standpoint in terms of the kind of putting the device in and what happens as I utilize the system?

29:41Yeah, so to install it in most cars, there's just by the rear view mirror, there's a canvas connection. So you take the trim cover off, you plug in our thing and you stick it to the windshield. And that's pretty much it. I mean, that is it. So then it's connected to your car. And then as for engaging, we just use the CAN signals of the engage button of the cruise control of the car. So you would essentially engage cruise control. You would get feedback from our device. there's a UI, there's sounds that you get, but that's how you interface with it. And then, you know, like I said, we are extremely reliable on the highway.

30:17Over 50 % of miles of our users are driven on the highway, driven by the system. And the thing that we're really focused on now is we want to get really smooth, like red light behavior, like in the city driving really smooth. We've had that in like beta mode for a couple of years, but we're trying to get that to a point where it's like so reliable and so comfortable that you prefer to your own driving, which is the state it's in on the highway. I'm curious, as you are kind of taking this and kind of super, you know, giving superpowers to your existing car in terms of what it can do, how are you thinking about, as you talked kind of at the beginning of the conversation about the larger problem of solving, you know, for autonomy and self-driving, what are some other use cases that you think this would apply to fairly easily without, you know, without a big jump or a big development ever, something where you could say, okay, this works here.

31:12I can, I can put it over here and over here too. Do you have any thoughts in mind? And is the company, um, have any intention of, of, of looking at alternative use cases as, uh, kind of opening up new lines of business or anything? Yeah. I mean, we're, we're open to this and I'll give you a very concrete example of something that we worked on. We use this world model simulator so that you can simulate scenes. If you had a simulator where you had to manually put in traffic lights and assets and say the timing of the traffic light, that obviously wouldn't expand to anything except a road. We're also interested in indoor robotics.

31:45And I think the first challenge in indoor robotics that's basically not been solved in some kind of machine learning way is indoor navigation. The most basic thing you would want your indoor robot to do is to just drive around, figure out where it can go in the house, what the map of the house is. And, you know, your Roomba or, I mean, Roborock, whatever is the best now, could do this, but not in some kind of end-to-end machine learning way. You would have to hand code a LIDAR SLAM algorithm or a Vision SLAM and probably still have to hand do some stuff that's specific to a house. So as like a short project, we translated all of our stuff to, you know, indoor and just had a robot drive around indoor and try to path plan indoor.

32:24And like that kind of worked. let's say the machine learning state is not good enough yet to make that reliable. But that's kind of what we're imagining is, you know, as we create these things more generically, these will transfer easier to, for example, indoor navigation. The next thing, of course, is action. In this case, I'm just talking about driving. But at some point, you want your robot to do something other than just drive. And then you're talking about things like, you know, let's say moving an arm or grasping something or moving your head. this would require some like really high level machine learning approach that treats the driving actions of curvature and longitudinal acceleration the exact same way it would treat like the movement of an arm.

33:04So that's the direction that we're thinking. Machine learning is we are not at all in a state in robotics where this is close. But I think that's kind of our long term view is how do we create a machine learning system that can understand that moving an arm is in principle a similar thing as moving a steering wheel and learn about it in the same way. Right now, that's far away. What do you view as outside of maybe larger simulations or more training of the same types of models or the same architecture? What are the kind of, I guess, outside of the box things that are the unsolved problems that when you're kind of thinking about the solution space that you're working in to get to some of the larger vision that you're talking about.

33:55But what are some of the main, I guess, the main research areas that are still unsolved that are most relevant to this type of problem? Yeah, I'd say there's three things. There's controls, RL, and continual learning. I think those three things are necessary for like this kind of end vision of robotics. And they currently don't work at all. None of those three things work at all. And could you break down a little bit like what you mean by each of those things? Yeah. So controls is one that we particularly deal with much more than other companies because we don't have any control over what, you know, the car, how the car responds to requests of steering and gas and brake.

34:38And in general, the cars that we support respond very poorly. You will ask, you know, put stork on the steering wheel and it will do it delayed. It will not do it the way you ask it. It will do some weird internal logic that we sometimes don't fully understand. So we deal with very crappy controls, essentially. And we solve those problems with classical control solutions. Machine learning, we've tried this so many times. We have open challenges about this. And as far as I understand, no one in the research community has made significant progress on this either. When it comes to low-level controls, machine learning, there's no good solutions.

35:12I think you need something that looks like RL, probably, to solve that. To give you an example of what we do is we learn the tire stiffness of all the cars that it runs on and it has to learn it live on every car as well as like the friction coefficient of the tires. Without that, we can't get good control and that stuff is all, you know, classical optimization. No machine learning there. So that's the controls thing. I don't know. Any questions about any more to elaborate on that? That makes sense to me. It does. I'm kind of curious. I actually want to go back for a moment to something we were talking about early in the conversation that's been on my mind to ask.

35:47And there wasn't a good moment before now, so I'm just going to dive back. And that is when you guys, you know, first, you know, decided to, you know, write your own software in the form of OpenPilot, I'm curious, a two-part question is, why did you elect to do that from scratch at the time? What was that motivation that you had on that versus like using like ROS, the robotic operating system these days, There's, you know, obviously Rust 2 is replacing it, you know, or some other alternative that are out there. And then as a tag on, if you will, what was the what what drove the decision to open source it versus keeping it proprietary up front when you're in those early days and stuff?

36:32Because I see that you are you guys, you know, you take a lot of pride in being an open source company and that we certainly are very supportive of that. we'd like that. But I'm curious what your motivations were on writing it and making it open source. So why not use, some of these decisions predate me, so I can't really say exactly what the thought process was. But in general, OpenPilot is extremely efficient. When it comes to inter-process communication, I think OpenPilot does this better than anyone else, including ROS. I don't know what exactly the details are. I think there's some stuff about but we have zero copy messaging and stuff like that that just makes it more efficient.

37:11I mean, we run on a not that new phone chip, right? So these things do matter. Yeah, I can't say that much more about ROS than that. As for the open source aspect, it is pretty important for it to be open source because there are many cars supported and the community helps port them, right? If we had a closed source stack that interfaced with the cars. I mean, this would be extremely difficult for anyone to add support for a new car, which is a pretty important part about the ecosystem. So that was kind of a requirement from day one. Then there's parts that don't need to be open source for the whole kind of ecosystem to work, like the machine learning model running and all that.

37:51So that's just more from a principled point of view. I like open source. Everyone that works here likes open source. I think if you buy a device and you don't get to control what runs on it and you don't get to know what runs on it, I think you should question whether you really own the device or you're in some weird contract with someone. So I think just philosophically, we're very much pro this. And we've generally become more open source over time. Like people complain that our trading stack is not open source. That's not something we're against. We want to open source more stuff. We just open sourced a lot of the data center management stuff that we have, which is all stuff that we need as well.

38:25And I have one other quick question I wanted to throw in. And I'm looking at the OpenPilot GitHub, you know, at the bottom where it kind of gives the percentages of different languages and stuff. As an approximation, I notice you guys are kind of rough with other things thrown in. You're really two-thirds Python, one-third C++. I'm wondering if you, as people, one of the things that we've been doing a lot lately on the show is we've been hitting different autonomy stories and kind of sharing a little bit more about autonomy as a major use case for AI. I'm curious, how do you, like obviously Python being the primary language these days for AI and stuff, but how do you differentiate what needs to be in Python versus what needs to be in C++ these days or in other, you know, obviously there are now alternative high-performance languages out there, But like, how do you think about that architecturally as a CTO in terms of saying, this is where we want to play for this particular type of function?

39:29Do you have a methodology or a philosophy around that? I mean, I think just everything that can be in Python should be in Python. I think that's the answer. Ultimately, most of the development happens in Python. If you do debugging, that's most likely in Python. The machine learning training is in Python. the more other things are in Python the easier it is to go from experiment to shipping and that's kind of what we want to optimize some things cannot be in Python for various reasons you know there's a whole layer that runs with interfacing on a car on a separate chip that has safety implications that's written according to specific standards that's written in you know in C so there's not a Python option there.

40:15Sometimes there's performance reasons why something can't be in Python, but very rarely, I'd say, especially since most of the compute happens in neural networks anyway. So yeah. And just circling back, making sure we get kind of those other two elements. You had mentioned these three kind of open challenges, I guess, the first around controls. I believe the second was RL. I'm blanking on the third you mentioned. Could you hit those last two? Yeah, so I mean, they're all very related, I think, those three things. So I talk about controls, and I think the solution to controls is some kind of RL.

40:52I think what's special about controls is that imitation learning doesn't really work at all. And imitation learning is basically how we got all of the machine learning progress that we've gotten, right? You learn on some large corpus of tokens, and now all of a sudden, your model is smart somehow. But when you're talking about tight feedback loop stuff, imitation learning doesn't work. there you need RL and by and large RL just doesn't work you know there's there's some types of RL that they use in the LLM stuff that from what I can what I've read is is not isn't exactly the type of RL that will work for controls and now we've got like we've got RL strategies that seem to work for the humanoid robots as far as I understand like a lot of the cool tricks are some type of RL in a very constrained very accurate simulation environment but yeah it's RL is just not in a state where we can say, okay, have this reward function to optimize, which in our case is, you know, don't oscillate the steering wheel and also do what you're asked to do.

41:49A very simple reward function is not trivial to optimize that in a noisy real-world environment. So that's one thing. Yeah, keep going. I'm sorry. I was just interested. Keep going. I kind of cut you off right there. I know, that's okay. But yeah, so that controls are kind of related, and continual learning comes in there too, because like I said, there's many things that make controls a problem that needs to be learned live on that car, on that drive even. If you inflate your tires, it will affect how open pilot drives. And you need to learn that. We do continue learning now with a lot of classical optimization, but that ideally would be a smart neural network that can understand all those things.

42:30If it starts raining and you start losing traction, as a human driver, this is something you notice and adjust to. But modern machine learning strategies can't do that. So as we, this has been really, really cool. Love having conversations like this. And thank you for sharing your knowledge. Love the fact that it's open source so we can really talk about it in depth. As you are kind of going from, you know, we've had a, we've, we've got a mostly talked about kind of where you're at now and some of the thing, opportunities that it might be near term and stuff. but I'm kind of curious to put you in a different frame.

43:07If you're, you know, the workday is over, you're kind of doing whatever you do to chill out at the end of the day. What is in your head about like places to go with this kind of technology going forward? And I don't mean things that necessarily Kama is going to do right now or even things that you have on, you know, on the plan, on the roadmap. But more specifically, like when you're just thinking about that would be something I would like to evolve to, to get to at some point, maybe without a known path. What are some of those things that you could see in your capacity as a CTO of an autonomy company?

43:48What are the types of things that excite you that maybe the rest of us aren't as familiar with or we haven't thought about that that you'd like to see this grow into? Any of those kind of dreams and aspirations you can share as we close? Sure, but it's not going to be that visionary. I think it's a practical AI podcast. No pressure. I'm very pragmatic as well. I think I want very simple things, which is, you know, I really like my dishwasher. I really like my vacuum cleaner. I think they make my life a lot easier. And I want more tools like that, that use some kind of technology to make my daily chores and stuff that's just annoying, you know, like driving and stuff easier and those things will happen i hope they happen as soon as they can um these are really really hard problems and uh and when they do happen i i want them all to be open source such that you can own them you can control them uh you if there's spyware in your house you can delete it uh so yeah let's say that's the general vision i just want simple robotics projects uh products that uh that make your life easier in ways that it's tedious and that aren't owned by a big corporation.

44:58Good answer. Yeah. Yeah. I think that's a, that's a very appropriate and good way to close out, um, today. Uh, Harold is, is really nice to have you on the show. Um, we look forward to having you, you back some time to talk about all, all the cool things that, that will happen. I'm sure this year and in future, future, uh, roadmap for, for comma. Thank you so much for the work here and, um, hope to talk to you again soon. Cool. Yeah. I hope so too. Thank you very much for having me.

45:28All right, that's our show for this week. If you haven't checked out our website, head to practicalai.fm and be sure to connect with us on LinkedIn, X, or Blue Sky. You'll see us posting insights related to the latest AI developments, and we would love for you to join the conversation. Thanks to our partner, Prediction Guard, for providing operational support for the show. Check them out at predictionguard.com. Also, thanks to Breakmaster Cylinder for the beats and to you for listening. That's all for now. but you'll hear from us again next week.

From the publisher

Autonomous driving is not just a big tech or closed-source game, it's becoming accessible through open innovation and real-world deployment. Dan and Chris sit down with Harald Schäfer, CTO at Comma AI, to explore how OpenPilot is bringing self-driving to everyday vehicles using open source AI. We dive into the intersection of machine learning, robotics, and simulation, including how world models are enabling training at scale and shaping the future of autonomy.

Featuring:

Links:

More from Practical AI

All 157 episodes
Open Source Self-Driving with Comma AIPractical AI · 46 min
Listen in VO