999: What's Left to Build When Software Is Free, with Chip Huyen

9 Jun 2026 · 1 h 16 min · 27 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Episode topic: The future of software and AI engineering when “the cost of building software is heading to zero,” and what coders should do next—system design, efficient LLM application architecture, evaluation, and moving beyond software into physical AI/robots.

Guest background

Chip Huyen is an AI engineer, entrepreneur, and two-time mega bestselling author. She founded Claypot (acquired by Voltron Data, associated with Wes McKinney). Her book Designing Machine Learning Systems was translated into 10 languages. Her newer book AI Engineering was the most popular O’Reilly item in 2025.

Key claims

  1. System thinking/design remains valuable even as coding becomes automated.
  2. AI engineering differs from ML engineering: many AI products use LLMs “as a service,” requiring system-level components (prompting, routing, guardrails, memory, retrieval) rather than training models from scratch.
  3. Fine-tuning is often a last defense; inference optimization and model/version churn make it hard to keep fine-tunes competitive.
  4. Start simple: prompting first, then RAG, then fine-tuning if needed.
  5. Evaluations-driven development and careful “LLM-as-judge” guideline writing are essential.
  6. Physical AI is the next frontier; digital agents’ skills transfer, but physical world modeling and safe control are major gaps.

Notable examples

  • Customer support chatbot: classifiers route requests to FAQ vs expensive LLM responses.
  • GPT-4 enabled Chip to label data efficiently.
  • Web search cost issue: repeated URL revisits due to overlapping search queries; caching opportunities.
  • “Mythos” model reportedly broke long-horizon human-eval benchmarks (16+ hour tasks).
  • Micro-tool: GitHub repo crawler/ranker that drew ~300k views in a week.
  • Robotics: Unitree humanoids (G1 vs H2) and kung-fu demos using pre-captured motions, with claims of arbitrary motion generation soon.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Zero Cost of Software

0:00 to 0:36

Explore the implications of the declining cost of building software.

“The cost of building software is heading to zero.”

In-Person Greetings

0:41 to 1:17

Jon and Chip share their excitement about meeting in person.

“This episode of Super Data Science is made possible by Anthropic, Cisco, Excel Data, and Notion.”

Reflections on Previous Podcast

1:17 to 2:36

Chip reflects on past experiences and the changes in AI since her last appearance.

“It was pre-GBT4, so I mean, like, the world has changed, right?”

The Evolution of Designing Machine Learning Systems

2:36 to 6:06

Discussion on the relevance of system design in AI and Chip's experiences teaching.

“Are you trying to get me to promote my book?”

Understanding AI Engineering vs. ML Engineering

6:06 to 8:03

Chip explains the distinctions between AI engineering and ML engineering.

“So your new book is called AI Engineering.”

Leveraging AI Models in Engineering

8:03 to 9:27

Discussion on the practical aspects of using existing AI models and techniques.

“like how to provide model with the right context.”

Fine-Tuning vs. Simplicity

9:27 to 12:15

Chip shares insights on fine-tuning models and the importance of starting simple.

“So you can have systems that combine many, many different components.”

The Start Simple Approach

12:15 to 14:01

Chip discusses her philosophy on starting with simple solutions in AI engineering.

“It was like, it's a lot of fun, like thinking about like how to make systems faster and cheaper, more efficient.”

The Importance of a Start Simple Approach

14:01 to 18:50

Learn why starting simple with AI applications can lead to better understanding and debugging.

“You know, we spend all this month like, yeah, dealing with all the issues that you're describing there, you know.”

Web Search Challenges in AI

18:51 to 19:50

Explore the complexities and costs associated with web searching in AI applications.

“The industry has focused on scaling AI vertically, bigger models, more compute.”
Show all 27 chapters

Inference Efficiency and Cost Management

19:51 to 23:58

Discover strategies for optimizing inference in AI models to manage costs effectively.

“the$200 a month tier just to get access to deep research.”

User Feedback and System Evaluation

23:59 to 28:00

Understand the significance of user feedback in improving AI systems and ensuring their effectiveness.

“So I think there's a part of the actual reasoning, actually a lot cheaper, but a lot of them are more like tool use tokens, right?”

Evaluating AI Systems: The Importance of Feedback

28:00 to 29:38

Explore the challenges and strategies for evaluating AI model outputs and the significance of user and agent feedback.

“So another feedback is also good for tuning model.”

Evaluations-Driven Development in AI

29:38 to 31:18

Learn about the transition from test-driven development to evaluations-driven development in AI engineering.

“Like it used to be the case that you could have tests in software development where you were like, if it isn't character by character, this output, there's a problem.”

Crafting Effective Guidelines for AI Judges

31:18 to 33:01

Uncover the complexities of creating guidelines for AI models to evaluate responses effectively.

“So I think it's like, it's important to not just assume that AI can do everything.”

Using AI for Data Labelling and Synthesis

33:01 to 35:19

Discuss how AI can assist in labelling data and synthesizing training data for improved model performance.

“a good understanding of what your instructions are doing.”

Micro Tools: Building with AI

35:19 to 39:20

Delve into the ease of building micro tools using AI and the implications of replicating software functionalities.

“of time on like tool use and like web search and stuff to make my things better.”

The Future of AI in Physical Systems

39:20 to 42:01

Explore the potential of AI in hardware and physical systems amidst the rapidly evolving software landscape.

“Yeah, so I do think that it's like, it's an interesting feeling.”

The Future of AI in the Physical World

42:01 to 45:01

Explore the challenges and opportunities of AI operating in physical environments.

“And, yeah, so the cost of building software is zero.”

World Models and AI Development

45:01 to 47:15

Learn about the concept of world models and their importance in AI training.

“It can know that, okay, if I want to do this, I should be able to, like, make up, come up with a plan of actions that can help me achieve that task without causing the consequences I don't want to cause, right?”

Challenges in Physical AI and Robotics

47:15 to 50:53

Discuss the key challenges faced by AI in robotics and physical interactions.

“They had to make AI capable of understanding the world, coming with a plan of actions.”

Future-Proofing in the Age of AI

50:53 to 56:01

Discover strategies to remain relevant in a rapidly evolving AI landscape.

“For our listeners who can't raise a billion dollars in VC money in the first round, what recommendations do you have?”

The Challenge of Terminal Interfaces

56:01 to 58:00

Explore the complexities of terminal use and the need for user-friendly alternatives.

“And then it was like, okay, why is this so hard to use?”

AI Interaction with the Physical World

58:01 to 1:00:07

Discuss the challenges AI faces in interacting with the real world, illustrated by a food-delivery robot example.

“Well, I'm just wondering what, this has all been very interesting, but what is this all, what is this all, are you saying that there's lots of lots of problems that we can still solve?”

Audience Questions and AI Perspectives

1:00:08 to 1:04:02

Engage with audience questions about AI, including the concept of AGI and its implications.

“So, like, actually some cities have this, like, streetlight API.”

Chip's Insights on AI and Reading

1:04:03 to 1:09:44

Chip shares thoughts on AI, literary influences, and book recommendations.

“There's a paper I did an episode a few years ago, a Five Minute Friday episode, on the five levels of AGI.”

The Future of Software Building

1:10:03 to 1:10:58

Explore how AI and software accessibility are redefining value in tech.

“a model from scratch by collecting data, training and deploying, and AI engineering, which treats powerful models as a service you call instantly, making time to market dramatically faster.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Jon Krohn:The cost of building software is heading to zero. So what happens to everyone whose career was built on writing it? Welcome to episode number 999 of the Super Data Science Podcast. I'm your host, Jon Krohn. My returning guest today is the sensational AI engineer, entrepreneur, and two-time mega bestselling author Chip Huyen. Indeed, her most recent book, AI Engineering, was the most popular book within the O 'Reilly platform last year. In this episode, we dig into the interesting and practical content from that book, as well as how coders can future-proof their careers, her newfound fascination with physical AI systems like robots, and much more.

0:36Jon Krohn:We're so lucky to have Chip on the show. I hope you enjoy this enthralling episode. This episode of Super Data Science is made possible by Anthropic, Cisco, Excel Data, and Notion. Chip Huyan, welcome to the Super Data Science Podcast. I can't believe I'm with you here in the flesh. How's it going today? It's crazy. I feel it's good to see you in person, like very three-dimensional. I am three-dimensional. It's true. I can't believe it either. So we're live together in San Francisco, which is fun. Your hometown, you've done so much here. Last time you were on the podcast, which was episode 661, if people want to listen to that, invaluable gems from you in that episode, as usual.

1:17It was pre-GBT4, so I mean, like, the world has changed, right?

1:21Jon Krohn:Yeah, a couple weeks later. And the reason why I know that is because I was mentioning to you how you were in 661 and you said, who was 666? Yeah. And that was GPT-4, I remember, off by heart. So a few weeks later, GPT-4 came out. And yeah, we're in a very different world. Lots of things have happened to you since then. At that time, you were the founder of a startup called Claypot. It was acquired by Voltron Data, I think. Yeah, exactly. Wes McKinney's company, so really cool company to be acquired by. And the foundation model wave hit. And then you spent a couple of years writing a new book.

1:53Jon Krohn:We're going to talk about that new book, which was the most read item in the O 'Reilly platform last year. So of all their videos, of all their courses, of all their books, your book, your new book was the most popular item, which is really exciting. But before we get to that new book, when you were last on the show, you had recently published your first English language book, which was Designing Machine Learning Systems. It's subsequently been translated into 10 other languages as well. But you had an interesting point that you made to me before we started recording, which is that in the time since that book came out, it's actually become even more relevant, which is unusual in tech.

2:36Jon Krohn:Tell us about that. Are you trying to get me to promote my book? No, because we were just talking about this outside, how designing machine learning systems is one of the things now, when you wrote that book, people in our profession would still have been writing their own code. And now we don't really do that anymore, especially in the last few months. And so we were talking about how designing machine learning systems is still a really valuable skill to have if you're an AI engineer or a data scientist or a software developer. Yeah, I think system thinking or system design has always been very, very cool to me.

3:13And I think with AI, automating a lot of skills has become even more important to think about things in system because I do believe that like the more narrow the skills are, the easier they are to be automated. So I started like thinking about system design back in like 2018, 2019. I think the pain came from college. from college. So I was teaching this course on TensorFlow. And the first iteration went pretty okay. And then they asked me to do that again. And I realized it's like in the meantime, TensorFlow has updated, right? And TensorFlow 2.0 came out. And that means that if I wanted to do the course again, I would have to redo 80 % of the tutorial.

3:55And I was like, okay, I do not want to do something like that. How do I focus on things that's like do not change very often, which forced me to like just look into like, what changed fast and what don't change? And like just asking a lot of people, like what, like people's much smarter than me and what are the heuristic for like fighting out things that like truly matter over time. And I realized it's like one thing that I landed on was like system thinking, system design. And at that time, a lot of companies also starting having like this system design thinking, like system design interviews, but they were not very popular.

4:27So it's just like, but a few very progressive tech companies had them. So I just tried to collect information. And they wrote this notebook of machining system design. And I put it on GitHub. And it got quite popular. And back then, if you put the search query, like machining system design on Google, they will put two results. One from Neil Lawrence, which is an amazing professor in the UK. And another is just GitHub. And one of the reasons that Ripple got popular was that in the end, I have a bunch of questions, a bunch of systems to think about. Like, okay, if we want to build a system that do this, how do we go about it?

5:10And I think people found it useful to practice for interview questions. And so after that, so it's 2019, and then I taught a course at Stanford on machine and system design, and then it just came about. So the lecture notes became the book Designing Machine System. I couldn't get the title Machine System Design because someone else apparently, like, we got it with 2022. But yeah, so I think over time people do get used to the idea that a technical book doesn't have to be just pure code. I remember when the first book came out, a lot of people were like, okay, this is not a technical book because there were no code snippets.

5:50Jon Krohn:Yeah. So I think we have more and more over time get used the idea that like, OK, an engineering job is not only about writing code, like sort of about thinking about problems and how to find the solutions that make sense over time. Yeah. Invaluable book. Really appreciate it, Chip. Let's talk about the new bestseller. So your new book is called AI Engineering. And seriously, it's crazy to have a book. There's so many books, hundreds, maybe thousands of books that get released in the O 'Reilly platform every year and to have the most popular book of 2025. That's wild. And in that book, you were emphatic that AI engineering is distinct from machine learning engineering.

6:33Jon Krohn:So we probably have a lot of listeners. We have data scientists. We have machine learning engineers. We have AI engineers who listen. And some people probably conflate those things two together. You think about ML engineering and AI engineering is the same thing. But they're different, right? I think they're different enough for me to have two books on two different topics. But I do think that for a lot of jobs, it actually can involve both. So it's not about, OK, I want to become an AI engineer, or I want to become a machine engineer as well. I'm like, OK, what problems are you trying to solve?

7:06And what is the best way of assuming the problem? And a lot of time, a solution will involve both machine learning, like classical machine learning, and then a generative AI model. So one example, so in the past, machine learning engineer was more about building a model from scratch. If you want to build a regular system for Amazon or Spotify or fraud detection model, you have to collect your own data and then train the model. And then you deploy it as part of a system. Whereas with Jative AI, you can just start with a demo. You can just have access to the most amazing AI model, host somebody's cloud, and then you can just use it as a feature.

7:50So, like, the time to market is extremely, extremely fast. So, it's a model as a service. And to do that, right, I think there's a whole new technique. So, now we have, like, instructions, like how to write good instructions. I still think that writing prompt is an extremely important skill, like how to provide model with the right context. I do things like context and memory system, actually very, very related. They're both involved with, like, yeah, I think I like, I spend a lot of time thinking about memory system design. like how to get the model to like both retain information and then access that information efficiently over time.

8:23Yeah, so I think there's a lot of techniques around and guardrails and then like also like how to combine it into a system. So let's say that you have like customer support chatbot, right? And a lot of time a request might not even be relevant to your system, right? Like if you, you probably don't want your chatbot can't do an argument over who wins the elections. So you probably have some kind of filter, some kind of classifier to detect whether this request is relevant to your product. Then you might also think about cost saving. Maybe not only requests need to be sent to an expensive model.

9:03So you might, okay, maybe this question is about how to reset my passwords. there's an FAQ somewhere you can send the user the link instead of having to generate an expensive response, step-by-step instructions on how to change the password. So you can have classifiers say, okay, this request should be routed here, another request should be routed here. So that classifier is like classical machine learning, and you might want to build that in-house. So you can have systems that combine many, many different components. And as a person who is trying to solve that problem, you might want to understand what solutions are possible and what is the best choice for you.

9:43Jon Krohn:Yeah, so you're actually bridging your old book with your new book there, where this idea of designing the systems appropriately is critical to having, for whatever your application is, like that customer service example you're giving there, you're blending together those more classical machine learning concepts of the classifier to say, identify whether this is a relevant conversation or not, or to send something from the FAQ. And so that kind of classifier model is more like the machine learning engineering, the ML engineering, where you might need to label. Or today, we usually don't need to manually label because we can use a large language model to do it, which is a super nice thing to be able to do these days.

10:23Jon Krohn:And we wouldn't have been able to do that well back a couple of years ago when you were on the show. It's basically since GPT-4. That was the release that I started using GPT-4 to label my data for me and be much more efficient at the machine learning engineering. But then on top of that, we blend in the AI engineering, which I don't know if you agree with this definition or not, but it seems to me in my head, obviously there is some AI engineering out there that is like AI research where people at frontier labs or in places are creating foundation models and training them from scratch. But for the most part, for most of our listeners, and probably a lot of what your book is talking about is more about calling existing models, maybe fine-tuning them, but you're leveraging LLM APIs for the most part in AI engineering, right?

11:10Yeah, so for a lot of people using AI in their products nowadays, AI is mostly a service. And I do think this is the questions I still brought up like fine-tuning. I think that's a whole another debate on like how, like when you should start fine-tuning your model. So I'm on more of the skeptical side with firetuning because I do think that firetuning can actually solve a lot of problems. And definitely in a lot of cases, firetuning is necessary, but it's usually the last line of defense. I think there are many, many different techniques that we can try out before moving to firetuning. Because firetuning itself is not hard.

11:51It's actually like we have a lot of off-the-shelf frameworks that make firetuning pretty straightforward. However, the challenge is like we have a model and then now what, right? And I have a fine tool model. We have to deploy it. We have to serve it. And like inference optimization is not easy. It's not straightforward.

12:10Jon Krohn:You have a whole chapter in your book on that, actually. That is actually one of my favorite chapters. It was like, it's a lot of fun, like thinking about like how to make systems faster and cheaper, more efficient. But it's definitely not for everyone. Yeah, I think that. And also models, also evolved over time. I touched a lot of companies, and they told me that their development cadence is about two to three months. And that matches the cadence of when there's a new model hours that have a significant step change in functionality. So maybe every two to three months, a new model just blows everything out of the water.

12:47And a lot of like a lot of like techniques that you use to like address the weaknesses of like older generation models are no longer relevant.

12:57Jon Krohn:Totally. Yeah. So so like if you do a fine tune model, right, like you need to understand whether you your fine tuned models can keep up with like the base models that be increasingly more powerful. Yeah, 100 percent. I mean, that was definitely something we used to really pride ourselves in a startup that I had co-founded. I'm no longer at, but in that startup, we were like, oh, you know, we actually need to train and deploy our own models. We were downloading Lama model weights and fine tuning them, you know, as a small Lama model, one that you could fit on a single GPU, like a single H100. And we needed to do that, quote unquote, because for the kind of performance that we wanted, we couldn't get the results from a single GPT-4 call.

13:40Jon Krohn:Like it would have required a whole bunch of calls to GPT-4 to get the result that we wanted. But we could use GPT-4 to create a great training data set. And then we could fine tune this LLAMA model to be able to do a task. But of course, then six months later with like, you know, GPT-4.1 or something, you know, all of their numbering was weird in the fours. And I can't remember how like with 4.0, something, a new model release would have come out. And then it does everything. You know, we spend all this month like, yeah, dealing with all the issues that you're describing there, you know. You look very excited talking about that time in your life.

14:14Well, yeah.

14:16Jon Krohn:Was that the last time that I was fine-tuning AI models? Maybe. But yeah, you actually, you talk in your book about a start simple approach where you focus on prompting first. Yeah. And then techniques like retrieval augmented generation rag and then fine-tuning if you need to. So how can you convince our listeners or how do you convince people that you talk to that this start simple approach, prompting before rag, rag before fine tuning is the way to go? So it's not necessarily like this has to be before that. It's more of like looking into what you want to build and see like where is that failing and come up with a solution, like the simplest solutions that address the failures.

14:58So I am a big fan of start simple, but mostly I'm a lazy person. Like I don't want to do things that are too complicated if I don't have to. And also like when you start to complex in the beginning, it just make it harder to understand the system and debug. Because like if things fail, we don't know, okay, is this like this component is failing or another component is failing? So I just want to understand just like how everything works together first and slowly build on top of things that I already have a good understanding of. So yeah, so the first thing is like prompting. I do think that prompting can actually take you a very, very long, take you pretty, pretty far.

15:34So I think that's like for a lot of, I view a lot of like applications now. And prompting actually helps like tremendously. And I think like the reason I can see from like half-assed prompt versus like a prompt I really spend a lot of time thinking about is huge. And one of the things I do notice about prompt writing is that over time, prompts tend to get more complex, like very, very long. Even the prompts that I write, they can go into thousands and thousands of tokens, right? And then one day I was like, okay, throw the whole prompt into my AI and say, okay, analyze my prompt. And then it pointed out to a bunch of things.

16:16It's like, okay, in this part of the prompt, you said that I should do this. But is that part of the problem you say, no, no, you shouldn't do that, right? So I think that's because over time, I just keep on addressing, using new examples. Because that's my concerted with older examples. And it's a prompt I wrote myself. And I realized when I talked to a bunch of other companies, when multiple people get involved in writing the same prompt, that happens a lot. The prompt gets extremely long. And sometimes I ask an engineer, like, okay, has anyone on your team, like, just spend, like, just sit down and read the prompt hand to hand?

16:53And, like, it's almost, like, never.

16:55Jon Krohn:Yeah, yeah, yeah. So I feel like a lot of times the performance is not quite, yeah, like, I feel like people can squeeze out more performance just from, like, running better prompts. And, of course, there's a question of, like, people talking about, like, giving the model, like, knowledge. So one huge failure mode is that, like, the model does not have the information to answer the questions, right? 100%. And if it doesn't have the information, it's going to make something up. I think nowadays, models are getting much better at saying that it doesn't know when it doesn't know something. But when you don't give it the right information, it's more likely that it's going to hallucinate.

17:30So, of course, you have to provide it with the right context. And one kind of context information you can provide is giving it a document. So first of all, if you want to analyze super data science podcast, I think I might be able to provide a bunch of transcripts from the previous podcast. So I think the models can reference to the transcript. And it can be very helpful. And as a guide context, it's like providing it for tools, like web search. So the models can do web search and find different information. So that's extremely, extremely important for a lot of cases. But web search, I'm not sure that you have talked to a lot of people who do like using AI with web search.

18:16But web search.

18:17Jon Krohn:It hasn't been a dedicated topic. It's extremely expensive. Oh, yeah. It's like painfully expensive. Like it's making me scared to like run my model. In terms of token consumption. Yes. Yes. I think there's still a lot of room for like improvement and how to make web search more efficient. So quick reality check for anyone building with AI agents. Your agents can discover each other. They can pass messages. They can coordinate on tasks, but here's what they can't do. They can't think together. When your agent figures out how to handle a complex workflow, that knowledge stays isolated. The industry has focused on scaling AI vertically, bigger models, more compute.

18:56Jon Krohn:Those breakthroughs matter, but intelligence also scales horizontally. Agents sharing knowledge across a network, coordinating on common intent, reasoning together. The infrastructure for that second horizontal axis doesn't exist yet. Outshift by Cisco is formalizing it. They call it the internet of cognition. They're publishing the architecture and building reference implementations. Read Scaling Out Superintelligence. We've got a link to that in the show notes. Then check out episode number 961 in it. Dr. Vijoy Pandey, the head of Outshift by Cisco, walks through how horizontal scaling of intelligence works and why it matters.

19:33Jon Krohn:It also, it depends so much on exactly how the provider you're using implements that web search, like how the agent does it. So for example, I don't actually use OpenAI Deep Research anymore, but a year ago I was using it all the time. And Deep Research from OpenAI, even though I was paying, you had to pay at that time, I don't know what it is now, but at that time I had to pay for the$200 a month tier just to get access to deep research. And then it would only, it would only check a few references. It would only do a few web searches. Whereas in contrast, a few months later, you know, so now about a year ago when Claude started allowing search, it would like spin up like a hundred agents instead of like six to look up like a hundred articles.

20:13Jon Krohn:And I was like, wow. And that's obviously using. Do you look at the cost, like the API cost? I mean, it's still, for me, it's very manageable. Like I, like, yeah, it's, you know, it can end up being for a single search, you could maybe spend like tens of cents in some cases. Yeah. So, so it was like, so, so I'm doing some, of course, like, I feel like, um, um, a lot of people get, look into a British market and I try to see if AI can like, I don't know, do a British market, right? Like, can change a British market. Yeah. Yeah. So, so I think, I think, I think I had that face and was just a curious, um, so, so I asked like a bunch of like AI agents to like to spin up.

20:51like, okay, given this market, do a bunch of research about it and make predictions on the outcome of this market. And I was just like, holy moly, each of the requests cost me like a dollar. Sometimes I call it was like$2. That is insane. So I look into the search result, and I feel like, I think I saw for some of the requests, it visits like a thousand URLs, like a thousand web pages. And so it was like, where are all these web pages? So I started just analyzing them. And I I found out that out of this 1 ,000 web pages, only 20 of them are unique. So on this, I just keep revisiting the web page over and over again.

21:31Jon Krohn:Interesting. And I was like, why is that? So what happened is that when you give it a request, say, hey, do a deep research about it, it comes up with different search queries. So for example, tell me about Super Data Science Podcast. It might come say, okay, Super Data Science Postcard 206, Super Data Science Postcard guest. So they had different queries, and these different queries might return the same URLs. So they might revisit the URLs over and over again. So, yeah. It sounds like an opportunity for caching. You could have thought. But I thought, so I'm looking through it. It's because a lot of it, it also, like, how data on the internet is structured.

22:12So, like, Google search, right, is structured based on, like, data chunks. So if you put in a query, it wouldn't find the website, but you surface the part of the website that is most related to the query. So that means that when you do a search, you don't retrieve the whole webpage. You retrieve the part of the webpage that's relevant to you. So I think it's a very involved process. And there's a whole part about, okay, you have the information, how you decide the freshness information. right maybe like for for something that's very much news related you do not want something that's from like two years ago but if something that's like technical like okay like what is an embedding right so maybe the result from like five years ago is fine so you look at like a system prompt of like a lot of this um of this as services like chat gpt or cloud it will see that they have like a really really big sections just trying to get the model to like do web search and like okay like like the freshness of data.

Read the full transcript

23:14OK, if the Skype query is like this, then maybe you can use the data from a week ago. If not, it can. Yeah, it's just like it is very, it's still very much manual, like the way I see how things are being defined. But anyway, so by the way, when I mentioned the system prompts, like the last time I looked at the system prompt of this model was about four months ago. So things may have changed. So I'm not sure how they are doing web search nowadays.

23:41Jon Krohn:Yeah, I'm sure it's a fast evolving space, especially if it's consuming tons of tokens. All of these frontier labs are trying to keep their compute down as much as they can to get the same quality of result at a lower cost. So there's got to be optimizations going on all the time. I think so. It's looking at how people are spending money on their agent. So I think there's a part of the actual reasoning, actually a lot cheaper, but a lot of them are more like tool use tokens, right? Like web search, file search, like one big category of tool use. The other category of tool use is like I see like productivity.

24:20Tool use, like when you connect it to like Gmail, Slack, Asana and like they're going to have to like do like tool calls and then analyze the results of my tool calls. And then of course it's also always like valuations. Yeah, some people were telling me that they were paying like 20 % of their token costs. Some neighbors? Oh, no, some company.

24:43Jon Krohn:Oh, some company. Some neighbor. You know, my neighbor is not into AI now. Oh, really? In San Francisco, you have neighbors that aren't into AI. That's amazing. Yeah. Fantastic. I'm going to fast forward a little bit in your book because earlier in the conversation, we were talking about how interested you are in inference. Yeah. So I want to give you some time on that and explain to our listeners why that's so exciting for you. It gets its own chapter in the book. And for our folks who are watching their token bills balloon, what are the highest leverage moves that they can do with their own LLMs in production in order to be more efficient?

25:21Honestly, I feel like inference. I'm interested in it because I'm a nerd. But I think for a lot of people, I think if you don't control the model, then you can't really do much to optimize that. Yeah.

25:33Jon Krohn:Yeah, if it's your own, I guess if it's your own, then you're not so worried about token costs themselves. I guess you're maybe just worried about how often you have the model up or how many GPUs you need to have running. Maybe we're starting to get into something that's really for people that have a large number of models on it, a large number of GPUs on it anytime doing inference. I think one thing people can do for token cost if they're worried about is that I try to look into what owns the usage and what kind of usage that can be offloaded to smaller models. Not everything needs to be sold by very big, very massive, expensive models.

26:16So if something that is small, like, for example, like going back to the classic example of customer support, like if you can see a certain category of queries that can be answered just by mapping to FAQ, then maybe you don't need to send them to big models. Yeah.

26:35Jon Krohn:So bringing that back again to the idea of great systems design for your AI system and bridging that with your book, which actually we're going to move on to, I just have one last question for you related to your books, and then we're going to get onto what you're excited about right now. But your final chapter in your AI engineering book is about closing the loop with user feedback. And it seems like a good user feedback design is critical to having a great AI system in production. Do you have any thoughts for us on that? So I think feedback, user feedback, serve many purposes, right? One thing is that if you serve people, if you charge people money for any service, you kind of want to know how happy they are about the service, right?

27:27So, of course, you need feedback. Feedback is also very good to uncover the area for improvement, right? So, I think if you look at feedback, you need to understand. So, it has maybe like for this certain demographic, somehow the users are not extremely happy. And you want to understand why. So we try to look into the data as detailed as possible, like slice and dice by users, by use case, by locations, such as you understand different patterns in there. So another feedback is also good for tuning model. So a lot of evaluations to... So user feedback is also one type of feedback, right? But the ultimate goal is so that we have a way to automate.

28:16automatically evaluate our systems in production.

28:20Jon Krohn:Agent feedback. How do our agents feel about our AI system? You really need to ask the agents. Exactly. Yeah, so I think it's like, okay, so let's say you build this very complex and very amazing systems, and your engineers sway on their life that, okay, this is amazing. But you're like, yeah, I trust you, but maybe you should just put something to see whether it's actually working, right? So you kind of want to build an evaluation system to see if the responses make sense. And a lot of people using LM as a judge, right? They use another AI model to look at the response and evaluate. And one thing that is actually really, really hard that I see that a lot of people spend so much time on is how to write the guideline to have the AI models evaluate, like output score that makes sense, right?

29:12So I think we discovered that pretty early on, there were a lot of case studies when somebody was like, okay, we spend 80 % of our development time just to write guideline for our AI judge, like I am as a judge. It's painful. Yeah.

29:26Jon Krohn:You have two chapters in your book on evaluation, in your AI engineering book on evaluation. And in it, you describe things as getting slippery as soon as the outputs that we're evaluating are generative. because it's very difficult to come up with a really rigorous test. Like it used to be the case that you could have tests in software development where you were like, if it isn't character by character, this output, there's a problem. And now obviously with Gen AI, you can't do that anymore because you'd have such a wide range of responses. So that's kind of what we're touching on here, right?

29:58Jon Krohn:You need to spend a lot of time figuring out what your evals are to ensure that you're getting the kind of result that you want. Yeah, I think like bringing back to what you just mentioned, in the past, I saw where we had test-driven development, which is like a very interesting program, extremely useful for a lot of use cases. For AI, I think the simple calling is like evaluations-driven development. When you like, there are only development applications that you can't measure the output for, right? Because I think like there's some questions sometimes when I give talk to a company, I ask people, like, okay, which one is worse?

30:37Like, having a system in production that, okay, like not having an AI system in production or having a system that nobody knows whether it's working or not. Yeah, so I think it's happened quite a couple of times. It was like, okay, we deploy this AI system. And some people was like, okay, like, we save so much money, we make so much money. And I was like, okay, how do you know? And so we're like, we guess. Yeah, so it's tricky. So, yeah, so going back to the guideline for the AI judges, it's painful. So, for example, like, how do you even explain to it, like, okay, this is a good response, this is a bad response, and why is that a good response?

31:17Like, a lot of people, like, we look at something, we can tell that it's bad or it's good, but we don't really want to sit down and write down our reasoning. and but like those um and i think like one of the most tests i have with um a lot of people is it's like okay after you write out the guideline for the model use yourself go and follow that guideline and go through some examples and see like okay like whether you're good right i like or and then i all i give it to the co-worker and like beg the co-workers like just follow it

31:45Jon Krohn:so you're saying that in today's day and age humans can still have taste yeah that is more useful than an LLM. That is something. Wow. So I think it's like, it's important to not just assume that AI can do everything. Like AI can follow instructions, but if the instructions are like trash, it's not going to be good, right? And I think something you talked about earlier in the episode, there's one of the absolutely most important things to having an LLM being able to do what you want or an agent to be able to do what you want as an individual or in an organization and an enterprise, whatever, it's having the right context.

32:24Jon Krohn:And so you might think you could have, you know, your AI system that's doing evals. You can't just trust the outputs because what if, you know, you think that you wrote a great prompt, you think that you have the evals set up properly, but if you don't go through and read them individually, there could be some key piece of missing context, maybe some product requirement that came from your users or the client. And only you as, you know, the stakeholder, this human who's been in all those meetings really knows what's important and can give the final sign off. Yeah. I think that like is, it's very important to like have a good understanding of what your instructions are doing.

33:04So, so a lot of time you don't need to read it line by line, but like just be aware that there's something you need to pay attention Like, for example, like I threw it into like another model and like just analyze for me finding contradictions in my own prompt. And it does that pretty well. And also have explained a lot of weird behavior that I felt. So, so for example, like so, so another applications when I use AI for like to do to do like deep research and generate a summary of the research. and then all the summaries somehow always focus on very specific things very weird and was like why so it was at first like okay is this is the ai models like bias why are you so insistent on like this new york office of this company like one of them and then it's it turned out like in some like weird part of the prompt i give an example of like okay if you if somebody asks about a company find its offices, maybe check whether it's in big city or New York and stuff.

34:03And it totally overindexed on that. So it was like, OK, and so it removed that example. And it works much better. So you just be aware of what you're asking the model to do. And I think it's one of the things that also useful for having good guideline is that you can also use it to synthesize data later on if you ever want to need more, fight more and more data. For sure. For sure.

34:28Jon Krohn:Yeah, that's, I mean, that's what we were talking about. Oh, yeah. I mean, you can be using, we didn't actually talk about synthetic data specifically. So we were talking about, I was talking about earlier in your episode, I'm taking up way too much of your time, but I was talking about how when GPT-4 came out, that was the first time that I could, that I felt confident enough about an LLM that I could be using it to label data. Yeah. But you could also use GPT-4 and any of those kinds of frontier models since to actually be synthesizing not just the labels but the training data as well yeah and you've got to be very careful in doing that because you can end up you might not get the breadth of the sample space that you want so you have to be really careful how you see that training but it can be very effective yeah like i personally don't fight tune models uh anymore i think i did try like early on but then at some point was like okay this model just keep getting a lot better yeah so i spent a lot more of time on like tool use and like web search and stuff to make my things better.

35:27But yeah, so I actually don't use a lot of data. But I do use AI to label a lot of stuff. Because a lot of my applications require AI to label different stuff, categorizing things, taxonomy, and that evaluations, obviously. Yeah.

35:44Jon Krohn:All right. I think it's time to go on to, you know, we've been talking about what you're excited about now. It's not fine tuning. So let's move on to what you're really excited about. It's actually, it's going beyond software. It's going into hardware, into physical systems. And so at CES this year, Jensen declared the chat GPT moment for physical AI is here. Is it? That's what he says. And so I'm here to ask you about that. Is that just hype? Or do you think that this is a really exciting time for you, maybe for our listeners, to think about getting into physical systems? I'm not sure I have the street cred to argue with Jensen at this point.

36:32Jon Krohn:I think you're the next. It's the two of you. You are the two people that the world looks up to the most on this kind of stuff. No, no, that's crazy. But yeah, so something is like, I'm not sure that you have gone through that, but I feel like a lot of my friends and I, we went through this, like we call existential crisis, right? And we're like, holy shit, like what do we do now? AI is coming for our job. So we thought a lot about, not just like what to build, not just like how to build things, but like what to build. Because I do think it's like AI is getting incredibly good at a lot of things.

37:05and building things is actually not, the process of building things is not that hard anymore. So let's say you have a great idea, right? You write a spec for it. You can ask AI to review those specs, improve those specs, and then it inputs the spec into AI and it can create some amazing website applications which is super cool. But then you build something and then what, right? So it has an experience of like, so it has a side project about two, three months ago and I put it online. And it's like, it's not like mind blowing anything. It's like something chill, something like, it was useful for me.

37:42Like I think one of the things I really love about AI nowadays is it just like allows me to build very good micro tool. It makes my life so much easier. Sure. It's very easy to build.

37:51Jon Krohn:Can you share some of your favorite micro tools that you built with us? So the things I did was like good AI list. So like every day you just like crawl, like crawl GitHub and it's fine based on like a bunch of keywords I give it and it finds all the repos related to this keyword. And then it analyzes repos and tells me what's interesting about it. And then it's categorized and also rank them. So it's helping to resurface things that might be relevant to me. So I put it out there. And it got, I think, 300 ,000 views in a week. And the next day, someone emailed me. It's like, oh, I love what you did.

38:30So I used AI to replicate exactly that. and here's that and it was just like i'm not sure how i feel about it like you know where i'm flatter but it was just like is that like copy like yeah i don't know like what is that right so it just made me so realize it's like anything that anything that exists today software can be copy like replicated when there was the leak of the cloud code repo about a month

38:57Jon Krohn:ago at the time of us recording, somebody, obviously, you can't, that's copyrighted, you can't publish exactly the same code, but somebody recreated it in Rust or something, use Claude code to recreate the Claude code code base in Rust and then publish that in GitHub. And so, yeah, it's pretty wild what you can replicate so easily today. Yeah, so I do think that it's like, it's an interesting feeling. Because I feel like, you know, on the one hand, AI mix is, it's like, it's very exciting because it allows me to build anything I want. And not everything, of course, there are some things that's harder to like build than other, right?

39:39I'm not going to be like Google search in like a webcam. But AI is also getting like, if you look at the complexity, there's a level of complexity of tasks that I can do. Like that level of complexity is like going way, way up. It's insane.

39:52Jon Krohn:With the Mythos release. So you know the people at Meter, M-E-T-R? Yeah. And so they have, I'll put a link. It's a very cool organization. Very cool organization. I should have someone from them on the show. If you know anyone, let's talk. And so they, I love, I always, every talk that I do these days, I always have the latest meter chart because they take the latest and greatest frontier model. And then they benchmark it against how well it can replace a human on a software development or a machine learning task. And Mythos broke their evaluations because they didn't have reliable tests that take humans longer than 16 hours.

40:32Jon Krohn:Imagine how long it would take. It was very easy when meters started going and they're benchmarking, okay, let's find some tasks that it takes a human a few seconds. Let's find a task that it takes a human a few minutes or a few hours. but now that they have to be coming up with tests that take 16 hours, 24 hours, it's going to be no time before it's 50 hours, 100 hours. How do you even find and pay humans to do that as a benchmark? That is scary. And exciting. Yeah. So Mythos broke it because it was the previous front, like top performing model was Opus 4.6 in the meter evals. And that was averaging, like it was able to do on average 50 % accuracy on a task that would take a human about eight hours.

41:12Yeah.

41:12Jon Krohn:And then Mythos, they're like, okay, they put the dot at 16 hours, but they also put this disclaimer that like, but we don't actually have good evals past 16 hours. So it might be even much more. Mythos might be able to handle on average a 30-hour task. We just don't even know. It's crazy. Yeah. Wow. Yeah. So yeah, things are improving fast. Yeah. Exactly. So is that what you're making? Yeah. Improving for machines. Yeah. So, yeah, back to the land of the humans. So, it just made me think about, okay, so it was saying that the cost of building software now is approaching zero, right? It's just code generation is cheap and nobody really cares about, like, in the past, people compared how many lines of code you write today, right?

42:01It's quite meaningless nowadays. And, yeah, so the cost of building software is zero. So, what does it mean? like it doesn't mean that the value of building software is also zero right because if the cost of it is zero then obviously um but yeah so so so maybe so think about like okay so what is the next frontier um and i think that a lot of people uh things that um we still have a lot of open ended problems in software i'm not saying that like is is also i think like but i think it's like we are on a trajectory where things are just being sold at a much much faster rate right And it's so hard to predict how long things are going to take.

42:41So a lot of people start looking into the physical world. How do we get AI to be able to perform in the physical world? And I think it's like, even from my perspective, and I could be wrong, is that I see the AI agents in the digital world actually have a lot in command with AI in the physical world. So I think if you look at the, so I was looking, so when we were working on AI agent, we were trying to create a mental model of what it's like. And I think it's like an agent is something that is interact with the environment, right? It can perceive the environment and act upon it. And then I get feedback from the environment, right?

43:21So in the digital world, like a lot, so the agents work in the digital environment. They can work in a browser. It can work in like a computer, like a terminal. It can work in a VS code. And the actions can be read, write, web browsing, like send queries, right? So a lot of that is like a digital agent. In the physical world, the environment can be the road, right? The agent can be a car with actions like turn left, turn right, break, and things like that, like accelerate. And on the actions, so you can actually map them. But the challenge with doing physical world is that the physical worlds are not very well described.

43:59So in digital world, ideally, not everything is like that, but in a lot of APIs, you have pretty good documentations. You have descriptions of the functions. If you send a request like this, you receive responses in this shape. You have status codes. You have arrow descriptions. It's pretty interesting. Whereas in physical world, we don't really have a good description of like, okay, if you squeeze this much force into the X, it's going to break. It's going to break in that direction, right? So like as humans, we learn to operate in the real world because we have observed over time. Like we learned just like, okay, if we, this, we kind of like some kind of haptic intuitions.

44:41Like we don't, we don't press too hard something. Like we don't step on a child because we know that the child is going to get hurt. Don't step on a child.

44:48Jon Krohn:Yeah, this is very important. Yeah, we learn that. Yeah. Everyone at home, remember. So, yeah. So, the physical world does not have this. So, the AI, it can maybe reason, right? It can know that, okay, if I want to do this, I should be able to, like, make up, come up with a plan of actions that can help me achieve that task without causing the consequences I don't want to cause, right? So, I do things that's like AI is getting really good at reasoning. It can come up with a task and come up with a task and then it can come up, sorry, it's given a task, it can come up with a plan to solve the task if it has a good understanding of the physical environment.

45:33And I think there's a lot of initiative around the world model, like how to build a model that encode physically accurate information about the world so that AI can operate in it.

45:45Jon Krohn:World models, that's the word. Yeah. Yeah, so we've had big fund races. I think Fei-Fei Li was the first with her World Labs. Yann Le Ke now with AMI Labs, a billion dollar first VC round, which is wild. David, I'm forgetting his last name. There's also a Google. David Ha. No, David Silver. David Silver. From Google DeepMind in London. He raised like a billion dollars for his world model company too. Who are you saying? David Ha? Oh, David Ha and Smith Schuber. were the authors of the proof of concept world models, I think, in 2018. And it was one of the first. Like, world models is not a new concept, right?

46:27I think we've been trying to model the world for, like, a very long time. But the ideas of, like, using neural network to, like, build that, I think, is more modern. Yeah, the world models.

46:38Jon Krohn:Yeah, it's really exciting. And so there are hard problems in robotics. Simulation to real-world transfer, scarce action data, sub-second latency, The cost of being physically wrong, like stepping on a kid. So it seems like that is similar to the kinds of real-time machine learning problems that you were obsessed with at ClayPod. So I think a lot of – so I think there's a different part of it. So I think, like, there are many challenges with AI in the physical world. So one of that is the AI part, right? They had to make AI capable of understanding the world, coming with a plan of actions. So, like, the hardware part, right?

47:23So, I think I actually realized this talk by the Unitree CEO. Do you know Unitree?

47:29Jon Krohn:Unitree? Unitree. Evening Tree? Unitree. Unitree. Yeah. Oh, Unity. Oh, no, Unitree. Unitree is a Chinese company. Oh, no, I don't know that. Yeah, no, I think they're one of the very cool robotics companies. Yeah, so they just fight for IPO, by the way. and one of the very rare robotic companies that claim to be profitable. So I'm actually very excited about the IPO. So the CEO has a great talk when he talks about robotic intelligence and he distinguished between two parts, right? One is like reasoning and the other is movement. So what that means is that like reasoning is like, okay, so robot look at the thing, given a task and I think about how to achieve the task, which is actually very similar like reasoning for like digital agent with the caveat that it has to like the physical agent have to have a good understand of the physical world.

48:24The other part is movement. And what that means is that like the robot may come up with a great plan and say, okay, to do that, let's just walk over there and like open the door. But then if it's just like go one step and fall over because it tripped on the wire or something,

48:38Jon Krohn:you would think it's stupid, right? So it's very important to have the robot to do like be able to do a lot of movement. I think that part is doing a lot of increasing performance in the last few years. So I think he was talking about how they had this really, really cool robot doing kung fu. So they have a lot of fleet of robots, and they're just performing kung fu. And it's like six feet tall or something? Like human-sized? So they have two humanoid. One is a G1, and the other is a H2. so the H2 is very imposing it's like 6 foot something but it's actually very hard to work with I don't think anyone I know is working with them actually I have one friend who has an H2 and he was like oh do you want to take it and the reason is that they got one but then they found it's too heavy they wanted to just give you a robot but then I don't know what to do with it what am I going to do with a 100 something pound robot it's heavy, I cannot carry it and it is falling on you So you're not ideal.

49:40It's not great. But it's a G1 scooter. It's like five foot or something. Oh, that's too pretty big. It's a lot easier to work with. So yeah, so they have this demo of robots doing kung fu. And instead, it's like, OK, to do that demonstration, they program the robots to do 20 movements, right? So like, OK, do this, I don't know, punch or something, whatever. And all of that is pre-captured. And then the robots just use different motion to combine them to do the demo. But then he was talking about like, we actually get, he believes that in the next six months, they can also do like instant, like arbitrary motion generations.

50:20So the robots can do like different motions without having being like pre-programmed, like pre-captured. So I think it's very exciting. So I think like, okay, there's two parts of robot intelligence, right? One is the reasoning, which I think this AI is getting like pretty good at. We're understanding that a lot of people with a lot of VC money is trying to solve. And there's the motion parts that I think is like getting very exciting. So, yeah, so I feel like all the pieces are moving, I hope, in the right direction.

50:47Jon Krohn:Very exciting. I can't wait to see what you do with this. Obviously, it's stealth right now, but it's going to be exciting. For our listeners who can't raise a billion dollars in VC money in the first round, what recommendations do you have? You were kind of talking about this earlier, this kind of existential crisis, which I feel as well. You know, it ranges from across everything I do. You know, I had, I did a PhD in AI and that gave me a real moat around my career. You know, other people couldn't create a machine learning classifier or, you know, understand problems with labeling data or these kinds of things.

51:20Jon Krohn:But now all of those kinds of things, a machine can do, you know, no problem. For the podcasting too. I mean, it lowers the barrier to entry. You know, you don't have to be a great writer to come up with great topics, you know, to script episodes. so you know there's there's a lot of people in a lot of industries who would feel like the mode is going away from them i'm wondering so obviously focusing on physical systems is one way to you know create a bit of a moat for yourself because you know hardware r &d cycles are going to be longer than software and there's lots of different ways that hardware can specialize and anthropic or open ai aren't going to next month just all of a sudden have a robot that does that too So there's a mode in physical AI.

52:01Jon Krohn:So maybe that is just the answer, but are there any other ways? Yeah, what other ways do you recommend? I guess you also have the systems design idea right at the beginning of this episode. What other tips do you have for our listeners on how they can try to future-proof themselves a little bit in this time? Well, nothing. this is really funny because um i i have um i have a friend who is an economist like he's one of the smartest people i know and he does like consult a lot of governments and coming up with like technical uh policy tech policies on like um yeah how how how to get their nations like stay up to date in the air era and he was straight up telling some of them it's like yeah there's nothing you can do it's like you're you don't have enough budget for it like just don't do anything so it was like yeah some some people do have a very very pessimistic views of the world um but um i think like i'm more on the optimistic side i do things that's like ai can solve a lot of problems but i also think that like there will never stop being problems for me like for us to solve so for one thing like doesn't matter like how many ai models are there or how good AI models are, I would never stop being angry at people on the internet.

53:20There would always be people that piss me off. There would always be customer services I'm unhappy with. There would always be things just like collaborations. It's not quite straightforward. So recently, there's a founder, and I really like the founder. He's very smart. So he came to me and he pitched this idea of another Asian orchestration framework. And then he told me that all the problems that he has seen with a lot of companies is that there's not enough communications between product and engineering, which is very classic. And he was like, okay, and my agent orchestration platform is going to solve that.

53:58And I was like, I don't think that's a technical problem. I think usually when product and engineering people don't talk with each other, that requires people's solutions. It usually, you don't solve that by like, okay, here's another tool. We magically make product people and engineer people get along. So there are a lot of people problems that it's not quite like so. So I think there's several categories of tasks. I think it's like, are not entirely completely clear for us to like how to solve. Like one is like human AI collaborations, right? So a lot of AI tools nowadays kind of built upon legacy systems and things about how we interact with old software systems.

54:47So just an example of coding tools. So originally, we have a lot of coding tools that are just part of VS Code because VS Code existed. And then we have a lot of coding tools as part of the terminal because terminal has always existed. But then this made me think, wait a second. why are they why are like editing like note taking apps like VS Code and Terminals why do they need both of them like why you know like why I mean I'm just going back like why is it different stuff and another thing is why is Terminal so hard to use so a lot of engine for me like I have used Terminals like for I know in school and stuff and for work like I use it but I'm not crazily happy with it

55:33Jon Krohn:a big Vim person. No, okay. So I have friends who are crazy Vim person. Like I have friends who just don't use VS Code or like PyCharm or anything. They just like straight up code in their terminal. So you think it's a lot faster with all the like key, right? So in theory, you could use terminals as a... As an IDE. Yeah, as an IDE. Yeah, but so, but like, yeah, okay. So we know that terminals exist, but like because of like coding tools like Cloud Code, a lot of product people or like people who never used terminal before are suddenly exposed to terminals. And then it was like, okay, why is this so hard to use?

56:07You know, it's like, it's very painful. And I think it was like, okay, terminals, maybe that terminals are hard to use on purpose because terminals are actually very powerful. Like terminals basically give you access to a control plane for you to like control the computer. You could easily like remove like RF, like everything, right? Like, yeah, like you can just do that. So maybe like you make it hard to use so that only people who are willing to get used to it use it so they are less likely to make mistakes. So, yeah, so I think maybe, but it's also like a legacy thing, and I think, like, why don't we have something, like, in between?

56:43Like, yeah, we have something that can be both very, like, can give you, like, access to the file system, access to the computer, like a control plane, a terminal, but also easy to use, like an IDE. And a thing like that is, I'm going to talk about it, And actually, OpenAI and Anthropocene introduced a bunch of services like desktop apps, right? Like the Codex desktop app, which is basically the same idea of like, okay, very easy to use interface, but give you access to a lot of things the ways that a terminal can. So I think this is evolving. And also there's a bunch of how to access your agent when the computer is not working.

57:20You probably have seen people complaining about like, okay, like why I have to keep my computer open all the time because my coding sessions, like my cloud coding codecs are doing their things. Yeah, yeah. Yeah, so I usually just get into an Uber and it's just like, can my computer open? I look like a freaking nerd. But I was just like, it's just making me think like, there's no reason my computer should be open, right? Yeah, because it's totally run as a cloud. And then we need like something that can access through the phone. I think like people, a bunch of people are building it. Like, okay, now you've won your things on the phone.

57:52Now you need some kind of sandbox, right? because like how do you share context between the phone and the computer? So anyway, I'm going to like very much, it was like genres like, shut up, she's talking with you all.

58:05Jon Krohn:Well, I'm just wondering what, this has all been very interesting, but what is this all, what is this all, are you saying that there's lots of lots of problems that we can still solve? Yeah, so I think it's like, I use it, not a great example, obviously, but I was saying it's like, it's one thing is like, we don't quite have a good understanding like what is the optimal way for humans to use AI. So human AI interface is one big thing, right? Another category of problem is it's like... I see. You know what I'm saying? She's not done. It's just like just the first. It's one of my first in the 20 points.

58:38No, this is really important. This is good. Yeah. Another thing, if you will, let me say it, is how AI interacts with the world. So we have a lot of techniques to make AI good at using tools, right? But I think for AI to interact well with the world, we do not just want to improve AI. We can still make the world more AI-ready. Right? How do you, for websites, for apps, we can make better documentation, better APIs that agents can call, better security, making, OK, this kind of actions are dangerous. So maybe you should have less permission to AI and stuff like that. But how about physical world?

59:19So I'm not sure you've seen this very cute video of a food-delivering robot. A foot what robot? Food-delivering robot.

59:29Jon Krohn:Food-delivering, right, right, right. Yeah, so it's a very tiny robot. Not a foot-delivering robot. That would be weird. Do you have extra feet to deliver? I'll take six feet, please.

59:41So the robots are very cute. And then the robot just couldn't cross the street. So the robot had to ask the pedestrian, like, hey, can you press a button for me so that it turn green so that I can't cross? And the pedestrian was like, what the heck is going on? So I actually talked to someone who worked at one of those robot food delivery companies. And he told me, like, the hardest part is just, like, how to get the robot in track of the world. So, like, actually some cities have this, like, streetlight API. so that the robot could connect to the Streetlight API so it can turn, it can press the button via the API.

1:00:21Jon Krohn:Oh, because it can't press the button to say that I want to walk. Yeah. So I'm just saying that part of how you could provide tooling to make the wall easier for AI to operate in. Sure, sure, sure. Yeah. Nice. Well, these were lots of great ideas. We do need to start wrapping up the episode a little bit because our next guest has actually arrived, a friend of yours. So we're going to, you're going to chat to him in a moment, but we do have actually for you, we have some audience questions. And so I want to get at least one or two of those in because we got a huge response to you coming on the show.

1:00:58Jon Krohn:It's maybe, you know, one of the biggest responses we've ever had about a guest coming on. All right. So our first audience question. So I posted on LinkedIn that Chip would be on the show. We had this tons of questions come in. And my first question, there were way too many to ask, but I'm just going to ask a few that I think are some of my favorites that might interest the audience the most. So Rahichia Valpuri, who is a business systems analyst for CIBC, a big Canadian bank. She's in a suburb of Toronto called North York. It's a nice city. Yeah. I'm from Toronto. Did you know that? Oh, that's why you're so nice.

1:01:35Jon Krohn:and um so rahisha says that she loves your writing so much so that when she's reading your books i kid you not there are moments when she finds herself smiling and she wonders if you're working on a new book oof so i've been playing with this idea of like writing a novel from an ai perspective with an ai narrator so it's kind of like try to explain how ai works from the ai perspective It's a crazy idea. I'm not sure I wouldn't do it. But I do want to get back into crib writing. I think my physical AI could be very interesting as well, but that also evolving quite a bit. So I'm trying to finish one blog post about physical AI.

1:02:16Hopefully it's going to be out by the time the episode is air.

1:02:21Jon Krohn:Nice, yeah. That should be easy to find from your website, which we'll mention the URL of in a moment. And then, yeah, some of the questions we actually addressed in the episode. So someone named Jing Xu, who is a, she is a very frequent listener and she makes lots of posts tagging me. And so I really appreciate that. Keep it up, Jing. Thank you. It's always great to get so much interaction from you. So she works as a director of data science at Elevance Health in Chicago. And she had lots of questions about how Vibe Coding, so Gen AI models, how that has changed. What is expected of AI engineers?

1:02:56Jon Krohn:And I think we talked about that a lot in this episode. So that was kind of even your very long answer near the end of this episode was basically about opportunities for us. You are free to cut it off. Okay. No, I can't. I can't do that. So, yeah, I think that we've kind of covered a lot of the really big questions. Okay, here's one very last one. So this is from Brian Willett, who's a retired research scientist. Very easy question for you, Chip. When will AGI be achieved? Oh, is that like – so it's the best definition. of AGI, do you follow the lawsuit with OpenAI? I think the contract between Microsoft and OpenAI is contingent on some definition of AGI.

1:03:37They already evoked that. So by the definitions we are...

1:03:43Jon Krohn:Legal definitions, we have AGI. I mean, it depends on whose legal team, right? But I think by the contract, I think Microsoft evoked that we are in AGI. Wow. I guess we'll see what happens there. We'll let the lawyers decide. Do you feel better or worse that we are in AGI? I mean, it's a difficult thing. There's a paper I did an episode a few years ago, a Five Minute Friday episode, on the five levels of AGI. And I think this is an important thing. And I think there's lots of areas where we do have amazing capabilities now in text-to-text, text-to-code, code-to-text. You know, in a lot of ways, in a lot of domains, you know, we have now exceeded the average person's capability on those kinds of tasks.

1:04:29Jon Krohn:But there's all kinds of things you brought up, like, you know, squishing an egg or stepping on a child. There's all kinds of things that AI systems don't need to improve on. So, you know, yeah, difficult to define exactly, but I guess we'll see what the lawyers say. So, yeah, so that brings us to the end of this great episode. It's been so nice to be here with you in person. And before I let you go, I know that you're famously a voracious reader of books. I'm a good influencer. I'm not sure that means anything to you. Yeah, you're a big, I mean, you're a huge AI influencer. And now you're going to be a big literary influencer because I understand you have an Instagram channel dedicated to your books.

1:05:07Jon Krohn:Not that you write, just the ones that you read. So I really like books. I like, I like, I like reading. So I have this Instagram account I created recently, like Chips Lip. which is really hard. I just realized they're very hard to pronounce. But like Chips Library, short for it. Chips Library. Yeah, I was showing my friend, my account the other day, and he was so disappointed. He was like, this is the most overhyped book account I've ever seen in my life. Because I feel like, okay, I finished reading a book and I want to post a review. And then when I get busy, so I just don't post a review.

1:05:39So yeah, but I try. So yeah, so I like, so I try to focus on books. So I think of as both fun to read and teach me something new.

1:05:52Jon Krohn:Fantastic. Do you have a new book recommendation for us? Since when you were last on the show, we did that one remotely and you were on your walking pad. And I had to ask you before we started recording, I had to ask you to stop walking on the walking pad because like the camera was moving, you could hear it. But behind you on your walking pad at your standing desk at home, you have a big bookshelf with lots of books. You pulled a number of books off the shelf to tell us. Do you have anything new that people need to read? Oh, so last year, I think I realized a book like Apple in China. So it's about the Apple supply chain in China.

1:06:27It's fascinating. It shares a lot about like how both Apple strategy and how China, just Chinese government thinks about technology supply chain as a comparative advantage. And I think like it's actually placed out quite interesting as in robotics nowadays. it's just really hard to find an American company that is as fast moving as a Chinese company in terms of robotic hardware. So I try to buy some robots, right? And I go to a bunch of American companies and a lot of them just don't have robots to sell. Whereas if you go to a Chinese company, you're just like, okay, which one do you want? Here's all the options.

1:07:06It's like, whoa, it's amazing. So it's very interesting. that book is I still like very much into like like urban design because I feel like city is a system and I'm thinking about how I approach that and there's a book that I found very amusing recently it's a book on parking I think it's like paved paradise I know parking right I think do you know that like on average there like you have to allocate six parking spots per car in the city

1:07:36Jon Krohn:really I know it's quite crazy that's wild Yeah, yeah, yeah. Because you have one other apartment, one other work, shopping mall or something like that. Oh, I see. So it's kind of interesting. And so I tried to read about it because I do think the self-driving cars are going to change a lot of that. Sure. Because a lot of cities nowadays are designed around parking space. Like having been to LA, it's extremely spread out because it needs a lot of parking space. And which also makes it worse because there are a lot of parking space, a lot of empty space. So now you got things just further apart.

1:08:10So more people need cars. It's just like kind of like vicious cycle cycle. So I think like self-driving cars are going to make things so much different because people don't need to park anymore.

1:08:19Jon Krohn:For sure. So I think I find that book's very, it's amusing. It's interesting. Yeah. Just like a bunch of like other books like that. Yeah. Nice. I'm sure we could go on and on. Pave paradise. Put up a parking lot. We will end on my little jingle. how should people follow you after the episode? You have over 300 ,000 people following you on LinkedIn. It's incredible. So you can follow Chip obviously on LinkedIn. I wish it means something, but yes, thank you. It definitely means something. I'll tell you that for sure. Anywhere else that people should be following you? We got Chips Lib on Instagram.

1:08:57Jon Krohn:We've got your LinkedIn account. It's not an account. If you follow it expecting to see the latest Asian tech news, You want to be so disappointed. But yeah, so like, I think that's just classic Twitter, LinkedIn. I'm trying to start a sub stack. I've been trying for like two years. I have a coming soon and people just keep tagging me like, you said coming soon a year ago. I know. When? I know. All right. Well, we'll see maybe next time you're on the show. We'll have your sub stack to talk about as well. It'll be vibrant by then. thank you so much chip for meeting with me in person in san francisco it's been such a great episode such an honor to meet you in person and yeah hope it's not too long before you're speaking to my listeners again yeah no thank you so much for having me again and congratulations for like 10 years a thousand episodes i know this is amazing i know yeah the very next episode will be episode a thousand thank you for being episode 999 what a special spot yeah no this is amazing Thank you.

1:09:56Jon Krohn:Wow, what an episode with Chip Hu Yen in it. She covered the sharp line between machine learning engineering, which meant building a model from scratch by collecting data, training and deploying, and AI engineering, which treats powerful models as a service you call instantly, making time to market dramatically faster. She talked about her start simple philosophy, meaning reaching for prompting first, RAG, retrieval augmented generation next, and fine tuning only as a last resort. She talked about how the hard part of physical AI isn't reasoning, but the fact that the real world has no documentation.

1:10:28Jon Krohn:And one fix is making the world more AI ready. Like cities exposing a streetlight API so a delivery robot can change the light itself instead of begging a pedestrian to do so. And she talked about how AI is making the act of building software nearly free. And in that new paradigm, the value of software hasn't dropped to zero. Instead, the durable problems worth solving are increasingly people problems and physical world problems that no model can simply copy. As always, you can get all the show notes, including the transcript for this episode, the video recording, any materials mentioned on the show, the URLs for chips, social media profiles, as well as my own at superdatascience.com slash 999.

1:11:11Jon Krohn:That's fun to say. Thanks, of course, to everyone on the Super Data Science podcast team. our podcast manager Sonja Breivich, media editor Mario Pombo, partnerships manager Natalie Zajski, researcher Serge Massis and founder Kirill Aromenko. Thanks to all of them for producing another fantastic episode for us today for enabling that super team to create this free podcast for you. We are deeply grateful to our sponsors. You can support the show by checking out our sponsors links, which you can find in the show notes. And if you yourself are interested in sponsoring an episode, you can get the details on how by making your way to johnkrone.com slash podcast.

1:11:48Jon Krohn:Otherwise, please help us out. There's lots of other ways you can do that by sharing this episode with folks who would love to listen to it, reviewing it on your favorite podcasting app or commenting on YouTube, subscribing if you're not already a subscriber. But most importantly, I hope you'll just keep on tuning in. I'm so grateful to have you listening and I hope I can continue to make episodes you love for years and years to come. Join us for a very special episode episode 1000 coming up in a few days um it'll have both kyrill the original host and founder of the show as well as me and lots of people dropping in from all over the world to ask us questions uh including listeners just like you all right until next time keep on rocking it out there and i'm looking forward to enjoying another round of the super data science podcast with you very soon

1:12:43Thank you.

From the publisher

Chip Huyen joins host Jon Krohn for this milestone episode 999 to talk about her record-breaking book "AI Engineering" the most-read title on the O'Reilly platform last year and how the AI landscape has shifted since her last appearance. Chip breaks down what separates AI engineering from machine learning engineering, makes the case for a "start simple" workflow, gets candid about the real costs of running LLMs in production, and shares why she's now fascinated by physical AI, robotics, and world models and why the durable problems worth solving are increasingly human ones. Jon Krohn guides the conversation from the practical content of the book through to where the field is heading next.

Additional materials: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://www.superdatascience.com/999⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠

Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information.

In this episode you will learn:

(06:48) What separates AI engineering from machine learning engineering

(14:44) The “start simple” approach: prompting, then RAG, then fine-tuning

(18:19) Why web search is so painfully expensive in production

(35:11) Is the “ChatGPT moment” for physical AI really here?

(52:21) Why the durable problems left to solve are people problems

More from Super Data Science: ML & AI Podcast with Jon Krohn

All 130 episodes
999: What's Left to Build When Software Is Free, with Chip HuyenSuper Data Science: ML & AI Podcast with Jon Krohn · 1 h 16 min
Listen in VO