Al Engineering 101 with Chip Huyen (Nvidia, Stanford, Netflix)

23 Oct 2025 · 1 h 23 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Lenny's Podcast: Product | Growth | Career

Episode Summary

AI Engineering 101 with Chip Huyen

Episode Overview In this episode of Lenny's Podcast, host Lenny Rachitsky interviews Chip Huyen, a core developer at Nvidia's Nemo platform, former AI researcher at Netflix, and author of the widely-read book *AI Engineering*. Chip shares deep insights into building effective AI applications, discussing the differences between AI training techniques, the importance of data quality, and the common pitfalls companies face while adopting AI tools.

Key Topics Discussed

  1. Understanding AI Apps
  2. What People Think vs. Reality:
  3. Common misconceptions about what improves AI applications.
  4. Emphasis on user feedback, data quality, and workflow optimization over chasing the latest AI technologies.
  1. Training Phases in AI
  2. Pre-training vs. Post-training:
  3. Pre-training involves developing a base model on large datasets.
  4. Post-training (fine-tuning) is crucial for adapting models to specific use cases.
  5. Fine-tuning should be a last resort after addressing data quality and user needs.
  1. Reinforcement Learning from Human Feedback (RLHF)
  2. Explanation of RLHF as a method for improving AI responses based on human evaluations.
  3. The importance of effective feedback loops to train models for better performance.
  1. Data Quality and Preparation
  2. High-quality data is critical for AI performance, often more so than the choice of database.
  3. Discussion on how retrieval-augmented generation (RAG) helps improve context in AI responses and the significance of data organization.
  1. Challenges and Opportunities in AI Adoption
  2. Companies often struggle to find effective use cases for AI tools.
  3. The conversation explores how different levels of engineers (senior vs. junior) respond to AI tools and how this affects productivity.
  1. Future of AI Engineering
  2. Discussion on how AI roles are evolving, with a distinction between ML engineers (building models) and AI engineers (using models).
  3. Speculation on the next wave of AI advancements, focusing on multimodal capabilities and the importance of system thinking in engineering roles.

Key Takeaways

  • User Engagement: Directly interacting with users and gathering feedback is essential for improving AI applications.
  • Data is King: Prioritize data quality and preparation over the pursuit of the latest technologies.
  • Adaptability: Organizations need to remain flexible and willing to restructure teams around AI tools and capabilities.
  • Emotional Intelligence in Storytelling: Relating to users' emotional journeys can significantly enhance product engagement and effectiveness.

Quotes from Chip Huyen

  • "A lot of companies are building AI products. A lot of companies are not having a good time building AI products."
  • "Fine-tuning should be your last resort."
  • "Most AI problems are actually UX issues."

Resources Mentioned

  • [*AI Engineering* by Chip Huyen](https://www.amazon.com/AI-Engineering-Building-Applications-Foundation/dp/1098166302)
  • Various articles related to AI evaluation and engineering discussed throughout the episode.

Where to Find Chip Huyen

  • [Twitter (X)](https://x.com/chipro)
  • [LinkedIn](https://www.linkedin.com/in/chiphuyen/)
  • [Website](https://huyenchip.com/)
  • [Substack](https://substack.com/@chiphuyen)

Where to Find Lenny Rachitsky

  • [Newsletter](https://www.lennysnewsletter.com)
  • [Twitter (X)](https://twitter.com/lennysan)
  • [LinkedIn](https://www.linkedin.com/in/lennyrachitsky/)

Conclusion Chip Huyen's deep understanding of AI and engineering offers listeners a valuable perspective on the evolving landscape of technology and its application in product development. The episode serves as a practical guide for those looking to improve their AI strategies and build more effective applications.

---

Feel free to reach out to Chip or Lenny through the provided links for further insights or discussions related to the podcast and its content!

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Our question that gets asked a lot and a lot is how do we keep up to date with the latest AI news? Why do you need to keep up your dick with the latest AI news? If you talk to the users and understand what they want, what they don't want, look into the feedback, then you can actually improve the application way, way, way more. A lot of companies are building AI products. A lot of companies are not having a good time building AI products. We are in an ideal crisis. Now we have all these really cool tools you have to do everything from scratch. It can have your design, it can have your write code, it can have your website.

0:27So in theory, we should see a lot more. But at the same time, people are somehow stuck. They don't know what to build. All this AI hype, the data is actually showing most companies try it, doesn't do a lot, they stop. What do you think is the gap here? It's really hard to measure productivity. So I do ask people to ask their managers, would you rather have a good effort on the team, very expensive coding agent subscriptions, or you get an extra head count? Almost everyone, the managers, would say head count. But if you ask VP level or someone who manages a lot of teams, they could say one AI assistant.

0:58Because as managers, you are still growing. So for you, having one extra head count is big. Whereas for executive, maybe you have more business metrics that you care about. So you actually think about what actually drive productivity metrics for you. Today, my guest is Chip Huen. Unlike a lot of people who share insights into building great AI products and where things are heading, Chip has built multiple successful AI products, platforms, tools. Chip was a core developer on NVIDIA's Nemo platform, an AI researcher at Netflix. She taught machine learning at Stanford. She's also a two-time founder and the author of two of the most popular books in the world of AI, including her most recent book called AI Engineering, which has been the most read book on the O 'Reilly platform since its launch.

1:41She's also gotten to work with a lot of enterprises on their AI strategies. And so she gets to see what's actually happening on the ground inside a lot of different companies. In our conversation, Chip explains a lot of the basics, like what exactly does pre-training and post-training look like? What is RAG? What is reinforcement learning? what is RLHF. We also get into everything she's learned about how to build great AI products, including what people think it takes and what it actually takes. We talk about the most common pitfalls that companies run into, where she's seeing the most productivity gains, and so much more.

2:12This episode is quite technical, more technical than most conversations I've had, and is meant for anyone looking for a more in-depth conversation about AI. If you enjoy this podcast, don't forget to subscribe and follow it in your favorite podcasting app or YouTube. And if you become an annual subscriber of my newsletter, You get a year free of 16 incredible products, including Devon, Lovable, Replit, Bolt, N8N, Linear, Superhuman, Descript, Whisperflow, Gamma, Perflexity, Warp, Granola, Magic Patterns, Raycast, JPRD, and Mobbin. Head on over to Lenny's Newsletter.com and click Product Pass. With that, I bring you Chip Huen, after a short word from our sponsors.

2:47This episode is brought to you by D-Scout. Design teams today are expected to move fast, but also to get it right. That's where dScout comes in. dScout is the all-in-one research platform built for modern product and design teams. Whether you're running usability tests, interviews, surveys, or in-the-wild fieldwork, dScout makes it easy to connect with real users and get real insights fast. You can even test your Figma prototypes directly inside the platform. No juggling tools, no chasing ghost participants. And with the industry's most trusted panel, plus AI-powered analysis, your team gets clarity and confidence to build better without slowing down.

3:26So if you're ready to streamline your research, speed of decisions, and design with impact, head to dscout.com to learn more. That's d-s-c-o-u-t dot com. The answers you need to move confidently. Did you know that I have a whole team that helps me with my podcast and with my newsletter? I want everyone on that team to be super happy and thrive in their roles. JustWorks knows that your employees are more than just your employees. They're your people. My team is spread out across Colorado, Australia, Nepal, West Africa, and San Francisco. My life would be so incredibly complicated to hire people internationally, to pay people on time and in their local currencies, and to answer their HR questions 24-7.

4:06But with JustWorks, it's super easy. Whether you're setting up your own automated payroll, offering premium benefits, or hiring internationally, JustWorks offers simple software and 24-7 human support from small business experts for you and your people. They do your human resources right so that you can do right by your people. JustWorks for your people.

4:30Chip, thank you so much for being here and welcome to the podcast. Hi, Lenny. I've been a big fan of the podcast for a while, so I'm really excited to be here. Thank you for having me. I want to start with this table slash chart that you shared on LinkedIn a while ago that went super viral. And I think it went super viral because it hit a nerve with a lot of people. And let me just read this and we'll show this on YouTube for people that are watching. So it's this very simple table. You share it of what people think will improve AI apps and what actually improves AI apps. What people think will improve AI apps.

5:00Staying up to date with the latest AI news. Adopting the newest agentic framework. Agonizing about vector databases to use. Constantly evaluating what model is smarter. fine-tuning a model. And then you have what actually improves AI apps, talking to users, building more reliable platforms, preparing better data, optimizing end-to-end workflows, writing better prompts. Why do you think this is such a nerve with people? And just if you have to boil it down, what do you think people are missing about building successful AI apps? What I'm saying that got asked a lot and a lot is that, how do we keep up to date with the latest AI news.

5:35And I'm like, why do you need to keep up to date with the latest AI news? I know it's not very counterintuitive, but there's so much news out there. A lot of people also ask me questions like, how do I choose between two different technologies? Like maybe like recently like MCB versus like agents, right? Like protocol. And it was like, which one is better or like this or that? And I think it's a serious question you should ask them. It's like, first, like if how much of the improvement would you get like from like optimal solutions versus non-optimal solutions right and sometimes they were like actually it's not much right and i was like okay if it's not much improvement why do you want to spend so much time debating something that doesn't uh makes a much difference to your performance and another question they asked is like if you adopted a new technology like how hard it would be to switch that out to another and sometimes they were like I think it would be a lot of work switching it out.

6:31And I was just like, let's say here's a new technology. It hasn't been tested by a lot of people. And if you adopt it, it would be stuck with it forever. Do you actually want to adopt it? Maybe you want to think twice about overcommit to new technologies that hasn't been better tested. I love your broader advice is just simple. To build successful apps, talk to users, build better data, write better prompts, optimize the user experience versus just like, what is the latest and greatest? What's the best model to use right now? What's happening in AI? Let me follow this thread of this idea of fine tuning and basically post training.

7:10There's all these terms that people hear in AI. And I think this is going to be a really good opportunity for people to learn what we're actually talking about. Since you actually do these things, you build these things, you work with companies doing these things. And there's a few terms I want to sprinkle in through the conversation. But let's start with this one what what's the simplest way for someone to understand what is the difference between pre-training and post-training and then just how fine-tuning fits into that just what fine-tuning actually is to disclaimer i don't have like phone visibility into like on what like this big secretive like frontier labs are doing uh but right from what i heard right so so i think it's like one is um like supervised fine tuning we have demonstration data and you have like a bunch of like experts, like, okay, here's a prompt, right?

7:51And here is what the answer should be like. And you just train it, like, on, like, to, like, simulate, like, emulate what the human expert could be like. And that's also, like, what a lot of people with, like, so open source models are doing, as they do it by distillation. So instead of having human experts should, like, write really good sounding, great answers to, like, prompts, they get, like, very popular, famous, good models. to like generate a response to it and like getting this trained smaller model to emulate so so sometimes you see people just like so that's because like some i really appreciate open source community by the way but like going from like having been with you train a model so i can emulate a existing good model it's very different from like being actually trained a good model like an output for existing good model so it's a big step there uh so yeah so like we have my supervised fire tuning.

8:45And another thing that's like very big, I'm not sure you have guests talking about it already, but like reinforcement learning is like everywhere. Let's pause on that because I would definitely want to spend time on that. And that's such a cool topic that's merging more and more in my conversations. But just to even summarize the things you just shared, which I think is really, really important stuff. So the idea here is a model, essentially this algorithm piece of code that someone writes and say the frontier models are feeding it just like the entire internet of content. And basically it's trying to test itself on predicting across all that data the next word.

9:18Essentially, token is the correct way of thinking about it, but a simpler way to think about it is the next word in text. And as it gets it wrong, it adjusts these things called weights, essentially. Is that a simple way to think about it, even though that's just very surface level? So I think of language modeling as a way of encoding statistical information about a language. So let's say that we both speak English. So we kind of get a sense of what is more statistically likely. If I say my favorite color is, then you would say, okay, that should be another color. The word blue would be much more likely to appear than the word table.

9:59Because statistically, blue is more likely to get my favorite color is. So it's a way of encoding statistical information. so like when language modeling ministry and a large amount of data like it got it see a lot of languages a lot of domains so it can tell like okay you may say this sentence then it user do the prompts and it would come like with the next uh most likely token uh so by the way it's not a new idea actually a video so the idea comes very very old like from the 1951 papers um um the english entropy i think it's like claus shannon it's a great paper and i think it represents a story i really like is from did you read Sherlock Holmes by the way uh yeah I read a few Sherlock Holmes books yeah yeah so so this is a story of like when Sherlock Holmes was using this statistical information to like have shown a case so he was getting um so this is this story uh there is somebody left a message uh with a lot of like stick figures so Sherlock Holmes was like okay he knows that in English the most common letter is e then the most common stick figure must be And then he goes, he starts like that, he originally saw the code.

11:08So I think that's language. So in a way, it's like simple language modeling, right? But instead of like at a word level, he does it at like character level. And token is something in between, right? A token is not quite a word, but it's bigger than a character. So let's say we say token because it helps us like, what helps us reduce vocabulary because with character, it's the smallest amount of vocabulary, right? So Amphibet has a 26 character, but words can have millions and millions, right? Whereas tokens, you can be able to get the sweet spot between the two. So let's say that we have the new word, like, how to say, like, podcasting, right?

11:50Let's say it's a new word, but you can divide it into podcast and ink. So people understand, okay, podcast, we know the meaning. We know that ink is like a verb, like gerund, whatever it is. So we even know the word like podcasting. So that's why the token comes in. But yeah, that's like the pre-tuning is basically like encoding statistical information of language to help you predict what is most likely. I think that most likely is a more simple way of doing it because it's more like building a distribution of like, okay, so the next token could be like more like 90 % of the time it could be like a color.

12:26Like 10 % of the time could be something else, right? So it's basically distribution. So language would like pick. Like depending on your sampling strategy, like do you want it to always pick the most likely token or do you want it to pick something more creative? You know, so I can make sampling strategy. I think it's something extremely important. It can have you boost performance in a huge way and very, very underrated. Okay, awesome. So essentially, a model is just code with this whole set of weights, essentially the statistical model that has learned to predict what comes next after certain words and phrases.

13:03Yeah. And then post-training and fine-tuning specifically is doing that same thing. So pre-training, you get like GPT-5. Fine-tuning is someone taking GPT-5 and doing the same sort of thing, adjusting these weights a little bit for specific use cases on data that they find is necessary to do their very specific use case. Is that a simple way to think about it? Yeah, I think like width is like functions, right? So let's say just like you have maybe has a functions of like, maybe Lenny's height is maybe like one X, like one X plus something like two X, like one and plus something is a width, right?

13:40So you change it until you fit the correct data, which is like my height and your height, right? So you can think of the weight as just like a weight like they function. So you like chain adjust the weight so they can fit the data, which is the training data. Awesome. Okay. So we're talking about pre-training, post-training, fine-tuning. Is there anything else here that's important to share about just like what this is exactly, what people need to understand about these parts of training? So the vast majority of time, we don't touch on like pre-training model. Like as users, we don't use it at all.

14:11Right. It's already done for us. Yeah. so so i think my action is a bit of fun like uh process like when my friend training model is like trying to play with their pre-training model and they're horrendous they're like saying things it's like oh my god it's like yeah it's crazy um so so it's very interesting to look at like how much of like post training can change the motor behavior um yeah and i think that's where like a lot of time is that a lot of people are spending energy on nowadays they function a lap is on like post-training because pre-training, I think, so pre-training have been used to like increase the general capacity of a model, capabilities of a model.

14:51And it depends on, it means a lot of data and like model size, like to increase, to increase the model capabilities. And at some point, we are actually like have kind of maxed out on the internet data, right? And people like text data, I think a lot of people are doing like with other data, like audios and videos. and everyone's trying to think of like, what is the new source of data? But we're like post-training, but like middle class of like, is this more of like, everyone can have very similar pre-training data. It's like post-training is where they make a big difference nowadays. This is a good segue to, you talked about supervised learning versus unsupervised learning.

15:24I love, we're getting into this, by the way. This is super interesting. So you're talking about labeled data. Basically supervised learning is AI learning on data that somebody has already labeled and told it. Here's correct versus incorrect. For example, this is spam versus not spam. This is a good short story. This is not a good short story. We've had the CEOs of a lot of these companies that do this for labs, Mercore and Scale, Handshake, there's Micro, there's a few others. So is that essentially what these companies are doing for labs, giving them labeled data, high quality data to train on?

15:57It is in a way, but I think it's more like a product of big equations. So there are a lot more different components than that. So that's why I was talking about reinforcement learning. I'm not sure if your CEO that you interviewed bring up that term. So the idea is that you want people to like, so let's say you have a model, give the model like a prompt, right? And it produce an output, right? You want to reinforce, encourage the model to produce an output that is better. So now it comes to how do we know that the answer is good or bad? So usually people realize on signals. So one way to get a first one, good or bad, is human feedback.

16:40You have two responses. You can say, okay, this one is better than the other. and we do that is because like as humans we tend to it's very hard to give like concrete score but it's easier to do comparisons right like if you ask me okay give this song a score i'm not a musician like and don't know like how hard it is like it's like yeah i don't know like what like how 10 i'm gonna do six you know and like if you ask me again a month from now on a completely forgotten so okay maybe now seven only four i don't know but then if you ask me okay here are two songs and which one would you prefer to play for the birthday party i was like okay i can't So like comparison is a lot easier.

17:17So you have human feedback and then you use this human feedback to treat a reward model. So you might tell which, and then the reward model will help you like, okay, the model produces response. And the reward model can score. Is this good or bad? And you try to buy us who are producing better model, the better responses. Another way is that you can instead of using a human, you can use AI, right? Like look at the response, yes, good or bad, right? Or in terms of things that people are very big on nowadays, like verifiable rewards, which is like natural. So basically they give it a math problem and a math solution.

17:51Like it's a model output of solutions. It's like, okay, it's an expected response to be in a 4A2 and it doesn't provide 4A2, then it's wrong, right? It's not a good response. So yeah, so like a lot of time people are using this human labor, like human laborers to produce expert questions and expected answers in the ways that design systems that are verifiable so that the models can be trained on. Okay, I'm really glad you went there. This is essentially RLHF, reinforcement learning with human feedback, which is exactly what I wanted to also talk about, right? Yeah, so I think in general it's a way of learning.

18:33It's like training is controversial learning, and whether it's learning from human feedback or AI feedback or like very terrible rewards. I think they say, you say it's just different way of like clipping signals. Awesome. Yeah, that's, we had that CEO of Anthropic on the podcast and he talked about their version of RLHF, which is AI-driven reinforcement learning. I love the way you phrased it where you basically, you want to help the model, you want to reinforce correct behavior and correct answers. And this is the method to do it, whether it's, say, an engineer seeing an output from a model being like, no, here's how I would code it differently.

19:06and then training. And it's training a different model that the original model works with to tell it, am I correct or not correct? Is that right? Yeah, I think that's a way of looking into it. And I think that's a space is so exciting nowadays because there's so many like domain expert tasks that the model, like that model developers want model to do well on, right? Let's say you're like accountant, right? Like maybe once you use a model, you have accounting tasks. So I need a lot of like accounting data, like examples from my accountant. So I need to hire a lot of them to like do it. Or if you want to do physics problems, I want to do, I don't know, like legal questions and stuff.

19:46Or like engineering questions. Or like somebody was telling me they want to do like using like coding for to solve scientific problems. And not just like coding to build product, which is another different whole realm of things. And I also like using very specific toolings. like uh yeah like i'm not sure what apps you use but maybe like a full dating app or like quickbooks or like google excel like they're very specific like tool specific um expert expertise so you want the models which you learn so they need a lot of like human experts in this area to like create data to treat them um and it's a massive thing it's like people because uh everyone wants a lot of data and like one slaps at like unlimited budget uh but uh whether i think this is also like a little bit of low-key, interesting economics.

20:31I'm not sure you've talked to the guests about. I thought it's very interesting to think about because it's very lopsided, right? Because they're only very small numbers of frontier labs, right? And they want a lot of data. And there's a massive amount of startups or companies that are providing data. So you can see these companies, this startup doing data labeling, maybe they have massive AR. But you ask them, okay, so how many customers do you have? and they could be like a very small numbers. I'm not sure you saw you smiling. Yeah, we chatted about that. Yeah, so I'm like a bit like make me uneasy.

21:08I have like a company that's growing like crazy, but it's like heavily dependent on like two or three companies. And at the same time, if I was this company from TLS, what would be the right economical things for me to do? Right now, I want a lot of startups. I want to have a lot of providers so I can pick and choose. And then these providers can also like to compete each other to lower the price. And it's so dependent on me. It would sound to me regardless. So I feel like, yeah, so this economics, the whole economics is very interesting to me. And I'm curious to see how it plays out. What I'm hearing is you're bearish on the future of these data labeling companies.

21:46Because as you said, they don't have a lot of leverage over pricing because they have so few customers. And there's so many people getting into the space. So basically, even though there's some of the fastest growing companies in the world, you're feeling like there's a challenge up ahead. I'm going to share some barriers on it. I think I'm curious because I think things have had a way of work out in ways I don't expect. So I think that maybe these companies, they have a lot of data. Maybe they wouldn't be able to use that to like have some insight that helps them stay ahead of the curve. You know, so I don't know.

22:22A very fair answer. Okay, while we're on this topic, I want to chat about evals, which is a very recurring topic in this podcast. This is the other piece of data content these companies share that AI labs really need. Can you just talk about what an eval is, the simplest way to understand it, and then how this helps models get smarter? So I think if people approach eval, I think they're like two very different problems. One is an app builder, right? And like, can I say have an app that do like maybe a chatbot. Very simple. And that's the first thing that came to my mind. And I want you to know if a chatbot is good or bad, right?

23:00So, I need to come a little way with like evaluate the chatbot. Another thing is, I think of this as a task-specific event design. So, let's say I'm a model developer and I want you to make my model better at code writing, right? And I was like, okay, but how do I even measure code writing, right? So I even need someone to like, okay, understand creative writing and think about like what makes good story, like what makes a story good. And then design the whole data set and that criteria to evaluate creative writing. So yeah, so I think there's that. I think it's like more like evolve design. That is very interesting.

23:38Come out with criteria, come out with guidelines, how to do it. And then also like train people, like how to do it effectively. So I guess, in a case, I think aval is really, really fun because it's extremely creative. I was looking at like different avals people built and I was like, wow, like, it's not dry at all. It's just like super, super, super fun. We had a whole podcast on evals with Hamel and Shreya. And that's exactly what they talked about. It's just it's actually really fun to create evals for companies especially. So let's still dig into that one a little bit more. There's this kind of debate online that I don't know how big of a deal this debate is, but it feels like people spend a lot of time thinking about this, this idea of do we need evals for AI products?

24:23Some of the best companies say they don't really do evals. They just go on vibes. They're just like, is this working well? Can I feel it or not? What's your take on just the importance of building evals and the skill of evals for AI apps, not the model companies? You don't have to be like absolutely perfect at things to win. You just need to be like good enough and being consistent about it. okay this is not the philosophy i follow but like i have worked with enough companies to see that play out so when i say that white company don't need eva right let's say you are like an executive right and you want to have a new use case so here's a use case you started out with built and it's like it works well right the customers are somewhat happy you don't have the exact metric for it but like so traffic keeps increasing like people seem happy people keep buying stuff right and now here's our engineer coming like okay we need eva for it and so and it's like it was like okay how much effort do we need to put into eval and they were like okay uh maybe like two engineers as much as much and they could maybe would improve that and was like okay so how much expected gain can i get from it and the engineer would be like oh maybe you could improve it from like 80 to like 82 85 right and i was like okay but if you take that two engineers and be able to launch a new feature then it could give me like so much more like improvement right so so i think is like one of them is like eva sometimes people think of eva is like okay this is good enough just don't touch it like if you do spend a lot of energy on eva it would like only incremental improvement where it spends the energy on like another use case and maybe it's good enough that you because the vibe check it right so so i do things it's like maybe like that's a deep bit it's about um i do things that's like a lot of time people just like get things to the place when it's like okay good enough people run but and then but of course it's like there's a lot of risk associated with it because if you don't have a clear metric you have good visibility to how the applications or models are performing it might do something very dumb or it can cause you like I know something like crazy can happen so so yeah so um so so I do think evolve is very very important if you have if you operate a scale and where like failures can have like catastrophic consequences then you do need to be very tyrannical about like what you put in front of the users understand different failure modes like what could go wrong and also maybe in a space when that like it's a feature as a product is as a competitive advantage right you want to be the best at it so you want to have like a very strong understanding of like where you are and like where you are with the competitors but it's just something that's like more like a low-key okay it's like something it's like okay that's not the core or like it helps with our users then maybe you don't need to be so, so obsessed or theoretical about it.

27:09It's like, okay, that's good enough for now. And if it fails, then it fails. Like, okay, I know it's like, it's so terrifying, but like, yeah. Yeah, I think it's only about like the question of like return investment. I'm a big fan of Eval. I love reading Eval. And I say it's like, I understand why some people would choose to not focus on Eval right away and choose like bringing on new functionalities instead. Awesome. That is a really pragmatic answer. What I'm hearing is evals are great, very important, especially if you're operating at scale. But pick your battles. You don't need to write evals for every little feature.

Read the full transcript

27:42Something that Hamel and Shreya shared is that people need just like, I don't know, five or seven evals for the most important elements of their product. Is that what you see or do you see a lot more in production that people build and need? I don't think of like just a fixed number on like the evals. Like what was the goal to eval, right? The goal of Ava is to guide the product development. So like you see Ava, because I think I'm a big fan of Ava, is that it helps you uncover opportunities where the brokers are doing well. So sometimes we've seen it very often. It was like, okay, we look at Ava and we realize it's like, okay, it performed really poorly on this like specific segment of users.

28:21And then we're looking to it's like, okay, what's wrong with it? And it turns out it's like we just like don't have a good messaging to it. So, like, maybe we should, like, just focus on the taste of building polio and can improve significantly. Yeah, so I kind of like the number of evolve is really depends. Like, we have seen products with, like, hundreds of different metrics, right? Like, people are going crazy. This is because, like, that product is, like, general, right? It has different, it has, like, one evolve for, like, I don't know, like, verbosity, it has, like, one evolve for, like, user-sensitive data.

28:51And, like, another is, like, for length. But like has a number of like, okay, let's just be a great example, concrete example, like deep research. So you have the application, you have like build a model to like do deep research for you, right? Like, okay, like have a prompt and I may say, okay, do me a comprehensive research on all Lenny's podcasts and help me like sort of like propose, like show me, report on what kind of topics he's interested in, what kind of videos get the most views or like what topics that he's missing on that he should be covering, right? Like, have this kind of, like, prompt.

29:26Then how do you evaluate the result, right? I don't think there's, like, one, like, matrix that would help. Maybe it's just, like, maybe you have, like, 100. I think somebody has a benchmark and it's a get, like, 100 experts, like, write a bunch of prompts and they go through, like, all the answers on AI and, like, do it. And it's, like, it's extremely costly and slow, right? But if you might have something else, for example, like, one way I was thinking about it, I was talking to a friend about it, And one way is like, how do you produce the result of the summary, right? At first, you need to do like gather information.

30:00And to gather information, you need to do a lot of search queries. You like gather, grab the search results. And then from the search results, you like aggregate. And then maybe say, okay, I'm still missing on this. You have to do another route or like another route. And at the end, you have the summary. So every step of the way, you need evaluations, right? you don't need to the end gen so maybe for the search query in my first thing about like okay now i write five search queries am i looking to like how good are these search queries like do they like as they like similar to each other because in the five search queries are very similar like okay let me podcast let me podcast uh last month let me podcast like two months ago right it's not it's not very very exciting but like if the quality is a podcast like the keywords are like more um more diverse right and then look at the results of the of the search query and say you enter the search query like Lenny Postcard data labeling and then they come up with like 10 pages, 10 results and then you come up with like Lenny Podcast on, I don't know, like Frontier Labs and have like 10 results.

31:05I mean, look at a different web page, like how much of them overlapping. Are we doing both like the breadth, like getting a lot of page, but also like do we have depth and also like relevance because we come up with the search queries that completely irrelevant to the original problem. So I feel like every aspect of it, it would need a way of evaluating, right? So I don't think it's just like, how many evals should I get? But like, how many evals should, do I need to get a good coverage, a high confidence in my application's performance? And so to help me understand like where it is not performing well so that I can fix it.

31:42Awesome. And I'm hearing also just, especially for the very core use case, like the most common path people take in your product is where you want to focus. Yeah, so, yeah. Okay, let me, there's one more term I want to cover and I want to go in a somewhat different direction. RAG, people see this term a lot, RAG. What does it mean? So RAG stands for Retrieval Augmented Generations and also not a specific to JDIP AI. So the idea is just like for a lot of questions, we need context to answer so i think it came pretty oh i think it's from the paper 2017. so someone was like um so they realized it's like for a bunch of like benchmark when the question answering benchmarks they realized it's like okay if we give the model informations about the questions then the answer can be much much better so what they do with us is try to retrieve information from wikipedia so for for question topics it's like we trip that and then put into the context and like answer it is much better so i feel like it sounds like a no-brainer right i mean like obviously so so i think that's what racket as a simplest sense is just like providing the model with a relevant context so so that it can answer the questions and that's where like things get like uh really uh more interesting because traditionally when it started out a rack is mostly like text um so so we talk about like a lot of ways like how to prepare data so that the model can retrieve effectively let's say there's like not everything is a wikipedia page right like wikipedia page is pretty content and like you know okay everything about it is about a topic but a lot of times have documents like it threw me a lot right and like they have a weird way of like structures and documents let's say that um you have documents about lenny uh podcast right and in the in the future in the beginning documents like from now on podcast wouldn't refer to lenny's podcast right so let's say somebody in the future is like okay tell me about lenny right many is worth and because as the rest of the document does not have the term many you just don't know uh you might not retrieve it and the document is long enough to each chunk into a different part so like the second part happens doesn't have the the word many so you cannot retrieve so you have to find a way to process data so that makes sure it's like it can retrieve the information that's relevant to the query even though it might not immediately like obvious that is related.

34:02So people come up with only things like contextual retrieval, like giving a chunk of the data that's relevant, like maybe in a summary, metadata, so that it knows. Or some people use it as a hypothetical question. It's very interesting. For even the chunk of documents, I must get a bunch of questions that the chunks can help answer. So that when they have a query, it's like, okay, does it match any of the hypothetical questions? so it can fetch it. So it's a very interesting approach. Okay, so maybe before I go to the next thing, I just want to say this, data preparations for Rack is extremely important.

34:39And I would say this in a lot of the companies that I have seen, that's the biggest performance in their Rack solutions coming from better data preparations, not agonizing over what better databases to use. Because better database, of course, it's very important to care about things like latency or if you have very specific access patterns, like read heavy or write heavy. Of course, it's like it matters. But in terms of like pure quality answers, right? I think the data preparation is just like hands out. When you say data preparation, what's an example to make that real and concrete for us to understand?

35:13So like one way is like mentioned as in like you have like chunks of data. So we can think about like how big of each chunk should be, right? Because if it's like sort of thing about like if the context you want to maximize, maybe you can, It's a very simple example, right? You want to retrieve like a thousand words, right? So if a chunk of data is too long, then so if a data chunk is long, then it's more likely to contain more relevant metadata. So you can retrieve more. But if it's too long, like then you have a thousand words and the chunk is like a thousand words, you can reach one chunk. So it's not very useful.

35:49But if it's too short, then you can retrieve more relevant information. Like, oh, it can retrieve a wider range of, like, documents and chunks. But at the same time, a chunk is too small. It contains relevant information. So we have, like, very nice, like, chunk design, like how big a chunk should be. You add, like, contextual information, like summary, metadata, hypothetical questions. Somebody was telling me, like, a very big performance they got is that from rewriting their data in the question-answering format. So, like, instead of having, like, so they have a podcast, right, instead of just chunking the podcast, You just reframe, rewrite it into here's a question, here's answers, and produce a lot of them.

36:29It can use AI for that as well. So that's one example of data processing. A lot of examples I see is for people using AI to have specific tool news and documentations. And we write documentation. A lot of documentation today is written for human reading. And AI reading is different. Because it's different because humans, we have common sense. and we kind of know what it is. So one thing is all like human for human experts that have the context that AI doesn't quite have. So somebody told me that like what's a big change they have is like, let's say that you have a function, a document, documentation for this, maybe it's a library.

37:08And the library says, okay, the output of this one is like maybe talking for like, I don't know, some crazy term, maybe there's some temperature or something under graph. It should be like one, zero or minus one. And as a human expert, maybe understand the scale, like what one in the scale mean. But like for AI, it just really doesn't understand what that means. So I actually have like another annotation layer for AI. It's like, okay, what temperatures equal one means like that. It's not like it's an absolute temperature. It's like associated with the scale over there. So it's just having all this data processing to make it easier for AI to retrieve the relevant information to answer the questions.

37:45This episode is brought to you by Persona, the verified identity platform helping organizations onboard users fight fraud and build trust. We talk a lot on this podcast about the amazing advances in AI, but this can be a double-edged sword. For every wow moment, there are fraudsters using the same tech to wreak havoc, laundering money, taking over employee identities, and impersonating businesses. Persona helps combat these threats with automated user, business, and employee verification. Whether you're looking to catch candidate fraud, meet age restrictions, or keep your platform safe, Persona helps you verify users in a way that's tailored to your specific needs.

38:23Best of all, Persona makes it easy to know who you're dealing with without adding friction for good users. This is why leading platforms like Etsy, LinkedIn, Square, and Lyft trust Persona to secure their platform. Persona is also offering my listeners 500 free services per month for one full year. Just head to withpersona.com slash Lenny to get started. That's withpersona.com slash Lenny. Thanks again to Persona for sponsoring this episode. Awesome. Okay. So you've talked a bit about how you work with companies on these sorts of things, on their AI strategies, on their AI products, how they build, which tools they build, all these things.

39:02I want to spend a little time here because a lot of companies are building AI products. A lot of companies are not having a good time building AI products. Let me ask a few questions along these lines of what you've learned working with companies that are doing this well. One is just, I guess, in terms of AI tool adoption and adoption in general within companies, there's also this talk recently of just like all this AI hype. The data is actually showing most companies try it. It doesn't do a lot. They stop. And so there's all this just like maybe this isn't going anywhere. So in terms of just adoption of tools and AI within companies, what are you seeing there?

39:32For Gen AI in company, I think there's two types of Gen AI tooling that I have seen. One is two internal productivity. right like have coding tools like chatbot um uh internal knowledge like a lot of big enterprises have some kind of like a wrapper around like um models so like but like with access like maybe some different kind of racks so we should i think we talk about data um okay like text-based rack i haven't talked about like agentic rack or like having to like multi-model rack yet but it's like yes it's a whole very exciting area around that um yeah so like basically to allow the employee to like access internal document.

40:13Some are going to ask like, okay, I'm having a baby. What could be the maternal or paternal policy, right? Or like, am I having these operations? Could the health benefit like cover that? Or like, I want to like interview, I want to like refer my friend. What could be the process for that? So a lot of this like having chatbot, internal chatbot to help with internal operations. and another thing another category is more like customer facing so or like partner facing so what I guess about chatbot is a big one if a hotel chain you might have like a booking chatbot which is like somehow massive like a lot of booking chatbot because I guess it's I do have this theory of like a lot of applications companies pursue because they can't measure the concrete outcome and I feel like booking or like sales chatbot is very clear right there was a conversion rate right now with a chatbot with human operators and what could be conversion rate with a chatbot and it's something somehow i think it's like very clear outcomes and companies are easier to buy into this um these solutions so a lot of companies have that like customer uh facing chatbot uh so yeah so so that is um another category of tool um and exit um I think for customers or external facing tools, because people are driven to choose applications with clear outcomes.

41:39So the questions of adopting them is really based on whether they see the outcome or not. Of course, it's not perfect because sometimes the outcome can be bad, not because the idea or like the application's idea itself is bad. It's just because the process of building it is not that great. Yeah, so it's tricky. For the internal adoptions of like tooling, so internal productivity, that's where it gets tricky. I would say like a lot of companies, what they think of AI strategy. I think AI trees usually have two key aspects, right? It's like use cases. And the second is talent. You might have great data for great use cases, but you don't have talent and you cannot do it.

42:22So a lot of times in the beginning with Gen.AI, and sometimes I'm really admired a lot of companies for that. It's like, okay, we need our employees to be very Gen.AI aware, like very AI literate, right? So what they do is they start maybe adopting a bunch of tools for the team to use. They have a lot of upscaling workshops. They encourage learning. And it's a really, really good thing. And they're also willing to spend a lot of money into adopting, giving people Chachapiti subscriptions, cursor subscriptions, Cloud Code subscriptions to get the employees to be more AI literate. And that's the thing is like a lot of the security in the country may say, okay, we spent a ton of money on this tooling, but then we don't see, because you can see the usage.

43:10And it's like, but people don't seem to use them as much. And what is the issue? So, yeah, so I think that is tricky, yeah. What do you think is the issue? Is it just they're not, they're like, they don't know how to use them? Like, what do you think is the gap here? Do you think we'll get to a place of just like, wow, work is completely different because of AI for a lot of companies? The main thing is like it's really hard to measure productivity gain. So I taught you a lot of people on the website. First of all, on the same page is coding, right? A lot of companies are not using coding agents or coding, I see the coding.

43:46And I was asking, I was like, do you think that it helps with your productivity? And a lot of times the question is very hand-wavy. It's like, okay, I feel like it's better, right? And I say, okay, because we have more PRs, we see more code, and then immediate correctness. Okay, but of course, code, number of live code is not a good metric for that, right? So it's really, really tricky. And it's something funny. So I do ask people to ask their managers because I work with like each of the VP level, so they have like multiple teams under them. So I ask them like, okay, do you ask a manager? like okay would you rather have access would you rather have give everyone on the team like very expensive coding agent subscriptions or you get an extra head count right let's say it's like maybe like and almost everyone could say the managers could say head count but if you ask VP level or like someone who manages a lot of teams they could say it's like they could want AI system assist tools and the reason is that people say like okay because as managers right because you are still growing like you're not as a level when you manage hundreds of thousands of people so for you like having one HR headcount is like it's big so you want that not for productivity reasons but because you just want to have more people working for you whereas for executive you care more about like the um the maybe you have more like business metrics that you care about so so you actually think about like what actually drive drive productivity metrics for you.

45:21So yeah, so it's tricky. And I think that's like the question of like productivity is not, I'm not sure it's like fundamentally is the sub-more productive, but it's just like we don't have a good way of measuring productivity improvement. Another thing is also very wily. And I think that people do tell me that they notice different buckets of employees like different reactions to ai assistive tools like first of all i keep going by chip coding because it's boring it's big and it's like easier to my reason somehow um so this is like um i have different reports like one team would tell me that like um one of the people tell me okay amongst all his engineers he thinks it's like senior engineers would get the most output like would be more productive because it's like okay so that person very interesting so so he actually divided his team to like three buckets but he didn't tell them obviously he was like okay here's more like currently like um best performing average performing and lowest performing and then there's a randomized trial so like they give like half of each of each group like access to like to like cursor and then he was noticed like over time i was like okay the something funny like the group that get the biggest performance boost like in his opinions like was very close in his team There's a biggest boom boost like the senior engineer, so it's the highest performing.

46:42So the highest performing engineer get the biggest boost out of it. And then the second group is just like the average performing. So his opinion is like, okay, the highest performing engineers is so normal, proactive. They will say no, I just solve problems. So AI helps them solve problems better. Whereas the people who are the lowest performing, they only don't care much about work, right? So it's easier to just go on autopilot, get it to generate bad code and just do it. And they always just don't know how to do it. Another company, however, they told me that actually senior engineers are the ones most resistant to using AI as a tooling because they said it's like, okay, but AI, because they are more opinionated and they have very high standards.

47:27It was like, okay, but AI code, generate code just sucks. So just like very, very resistant in using it. So I don't know. I haven't quite been able to reconcile very different reports on that yet. This is so interesting. So just to make sure I'm hearing the story. So there's a company I work with that did a three bucket test with their engineering team where they created three sorts of groups. The highest performing engineers, mid performing engineers, lowest performing engineers and gave some of them. So they gave some of them access to, say, Cursor. Was it Cursor or what did they give them access to?

48:02It was Cursor. I think by the time it was Cursor. Okay, cool. I didn't work with them. It's more like a friend company. Okay, it's a friend's company. So did they give like half of the higher performing engineers Cursor and half not? Or how did they do the split? Yeah, so like they give like half of the entire company, but like half for each bucket. Yeah. And then they observe the difference in like productivity. I see. Yeah. So how did they even do that? They're just like, okay, you get Cursor, you don't get Cursor. How did they do that? That's so interesting. Yeah, I didn't get into the mechanics of it, but I was at rest back here for doing a randomized trial.

48:32That is so cool. Yeah. Okay, wow. How large was this engineering team? Was it like hundreds of people? It's not that large. It's about like maybe 30 to 40. Yeah. 30 to 40. Okay. Yeah. Wow. Okay, so they found that the highest performing engineers had the most benefit from using AI tools, and then behind them was the middle tier engineers and the worst performers were the lowest performers. Yeah, but it's still not the same everywhere. Like some companies, yeah, different. Right, this other example you shared of just senior engineers in this one example are most resistant to changing the way they work, which I get.

49:11I do feel like the most valuable people right now, other than ML researchers and AI researchers like yourself, are senior engineers Because it feels like junior engineers are just like so much of this is now done by AI, but an engineer that knows what they're doing, that understands how things work at a large scale with AI tools, just basically like infinite junior engineers doing their bidding, feels like an extremely valuable and powerful asset. Yeah, I definitely really appreciate, as you see companies, we appreciate engineers who have a good understanding of the whole systems and be able to have good problem-solving skills, thinking holistically instead of locally.

49:53locally. What other companies have seen as the way they work, as they told me, is that they work completely different now. So they actually restructured engineering org so that they get more senior engineers to be more in the peer review because they get sort of writing guidelines on what is a good engineering practices, what is the process would be like, maybe like, okay, so they write a lot of processes on how to work well. And then they have more junior engineer just like produce code and and like submit PR but senior engineer more in the reviewing case so I think it might be prepared for the future so another company actually told me something very similar so they can't prepare in the future once they only need a very small group of like very very strong engineers to like create like processes and like reviewing code to get into production but like get like AI or like junior engineers to like produce code but then the question becomes just like how does one become a very strong right that's right that's right that's the problem yeah so so i don't know what's the process i was thinking about like yeah um no one's thinking about it's just it's a problem we won't have any more in 10 20 years there'll be no more engineers because no one's hiring junior engineers although i could make the case junior engineers people just getting into computer science right now are just native ai native and in theory you could argue they will become really good really fast if they're curious, aren't just, you know, delegating learning and thinking to AI, but learning how to actually, using it to learn how to code well and architect correctly.

51:29Like you could argue they will be the most successful engineers in the future. I do think that what I mentioned is loading to architect. I think I grouped that in like system thinking. I do think it's a very important skill because I think AI can help automate a lot of like disjointed skills but like knowing how Chile utilizes skills together to solve a problem is very it's it's it's hard so there's a webinar between um Mira Sami was my one of my favorite professors he was a chair of the curriculum at the CS department at Stanford so he spent a lot of time thinking about CS educations right like what what should students learn nowadays in the era of like ai coding and then the other person is like andrew which is of course it's like a legend in the ai space and nera sami president like sami said something very interesting is like he said like a lot more things that cs is about coding but it's not like coding is just a means to an end like cs is about system thinking like using like coding to zone actual problem and problems something will never go away because like what like ai can automate more stuff the problem just gets bigger but it's a process of understanding what caused the issue and like how to like design step-by-step solution to it will always be there um so i think an example of um of like i actually have a lot of issues with like ai for like um in the way of like it's debugging so i'm not sure you use a lot of coding but like a something i've noticed and also seen from my friends it's like it is pretty good when you have very clear well defined tasks maybe write documentations fix these specific features or like build an app from scratch right like doesn't have to interact with a large existing code base but it is something like a little bit more complicated maybe require and try with other components and stuff it's usually like not that good and and for example I was using AI to like use to deploy applications and it was testing out a new posting service I I was not familiar with.

53:30It was like, okay, like usually they form me. So what AI does give me is like confidence to try new tool. Like before what AI is like trying new tools, it has really not documentations for the beginning, but AI was like, okay, just try it out and learn. So I was testing out this new hosting service and it kept getting a bug. It was like very, very annoying. And it was like, okay, I asked a card code, like fix it. And it kept changing the way, like maybe change the environment variable, fix the code, maybe not change from the function to this function. maybe change the language maybe it doesn't process javascript well i don't know whatever and it didn't work and it was like okay that's it i'm just going to read the documented document uh documentation myself and see what's wrong and it turns out it's like i'm on another tier like the feature i want did not is not available in this tier right so i feel like okay so the issue with cloud cores is trying to focus on fixing things from a very a different component versus issues from a different component.

54:27So I think of like, okay, understanding how different components work together and where the source of issue might come from. You need to give a holistic view of it. And it's made me think, it's like, okay, how do we teach AI, like system thinking, like that? I think I have all the human experts, like having, like, right, like very much, people call the scaffold. It's just like, okay, for this kind of problem, look into this, look into that, look into that and then stuff. So I think that could be one way. But that's what made me think, is like, how do we teach humans? My system thinking. Yeah, so I think it's a very interesting skill.

55:02I do think it's very important. That's exactly the same insight Brett Taylor shared on the podcast. He's the co-founder of Sierra. He created Google Maps. He was CEO of Salesforce, Quip, a few other things. And I asked him just like, should people learn to code? And his point is exactly what you said, which is taking computer science classes is not about learning Java and Python. It's learning how systems work and how code operates and how software works broadly, not just here's like a function to do a thing. One thing that I wanted to help people understand, you wrote this book called AI Engineering, which is essentially helping people understand this new genre of engineer.

55:41And you have this really simple way of thinking about the difference between an ML engineer and an AI engineer, which has a really good corollary to product managers now of just like an AI product manager versus a non-AI product manager? The way you describe it and fill in what I'm missing is just ML engineers build models themselves. AI engineers use existing models to build products. Anything you want to add there? One thing I really dislike about writing books is that you have to define like this. And I think it's like no definitions would be perfect because they always really edge cases. But yeah, in general, I think it's like AI as a service, like more as a service, like when somebody build the models for you and the base model performances are pretty strong so so it's like it's enable people to just like okay now i want to integrate ai into my product i don't need to learn work great and design it even though knowing that could really help uh but yeah it's like it makes the entry barrier really low for people who want to use ai to build product and at the same time ai co-abilities are like so strong because like it's also like increased like the possibilities, like the type applications that AI can be used for.

56:52So I think like, yeah, so it was entry barriers like super low and there's a demand for like AI applications like a lot bigger. So I feel this is very, very exciting. It opens up like a whole new world of possibilities. Yeah, it's like now you don't have the time, you don't even have to spend time building this AI brain. Now you can just use it to do stuff. Such an unlock. Okay, maybe just a final question. you get to see a lot of what's working, what's not working, where things are heading. I'm curious just if you had to think about in the next two or three years, just where things are heading, what do you think, how do you think building products will be different?

57:30How do you think companies working will be different if you had to think of maybe the biggest change we expect to see in the next few years in terms of how companies work? I think in a lot of organizations, they don't move that fast, right? But at the same time, they just move faster than I expected. Because again, I think it's like bias, like I don't work with a dinosaur company. I think a lot of executives who come to me are like very forward-looking. So maybe for me, I'm very biased towards organizations, it's like move fast. So, yeah, so I think one big change I see is like in organizational structure.

58:11I think there's a lot of value placed in like, so before we had a lot of disjointed team. We had very clear engineering team, product team. But then there's a question of who should write EVA, who should own the metrics. And it turns out EVA is not a separate problem, it's a system problem, right? Because you need to look into different components, how they interest each other, you need to use the behaviors because you need to know what users care about so that you can so that you can like write eval like reflect what users care about so so all of that like you can sort it from like you know get to different component architectures uh place guardrails and stuff so it's just engineering but understanding users is like what product right so so because of like a lot of things and eval extremely important so like the kind of bring product team and like engineering team even like marketing team like user acquisition like very close to each other so so yeah since in a way so if you go structuring so there's more communications between like previously very distinct functions another thing is just like also see as teams um of course i think about like what can be automated in the next few years and what cannot be automated and i've seen that people already like shedding like actually it's a little bit like scary to think about it but i also think it's like the team the webtopic it's like okay this is between you and me, but we haven't got rid of these functions.

59:33For a lot of things like previously outsourced, for example, traditionally, it's a business outsourcing this core to them and can be done with more systematized. So with that, you can actually use AI to automate a lot of that. And so there's a separation between more of what is the value of junior engineers or senior engineers, how to restructure engineering for that. So, yeah, so I knew definitely things that is one thing to success organization. People are just moving pieces around and thinking about use cases, whether you want to spin out new use cases and who would lead the new effort. Yeah, that is one big change.

1:00:18Another thing in terms of AI, I think there's, I'm not sure how true this is. I guess I'm also like on the camp of like thinking that it has merit. It's a camp of like, okay, base models, we have probably like not quite maxed out, but we want, we are unlikely to see like really, really strong, like crazily strong model. So like you remember like when we have like GPT, right? And then GPT-2, which is a big step up, like an auto-money to you, like better than like GPT. and then GPT-3, which is like much bigger than GPT-4, much, much bigger. And of course, I'm in GPT-5, but like is GPT-5 like that scale of like much bigger, like a step jump compared to like the previous?

1:01:04I think it's debatable, right? So I think it's like we had this point where like the base model performance improvement is not going to be like mind-blowing as it was in the last three years. so I think it's like a lot of like improvements we're going to see in the post training phase, in the application building phase and yeah so I think that's where I feel I would see a lot of improvement there. I'm also very interested in multi-modality so we've seen a lot of text based but I think there's a lot of audio, videos, use cases that is very very exciting and I think audio is not quite as song as one thing because i do work with like uh with like a couple of like voice startups and when you talk to think about voice it's an entirely different beast uh so let's say you have chatbot right we go from a text chatbot to voice chatbot it's like the concerns are completely different because now with voice chatbot right we need to think about like latency because like multiple steps uh first like have like text like voice to text text to text and text question and just text answer and then like and then text to voice answer right so we have like multiple hops and like latency become very important and there's a question like what does make you sound natural so for example like people think of like um in in ai and humans like when humans talk to each other like if i say i say you try to interrupt me and say um chip that's right i would like pause and i try to hear you out right but sometimes i just even may just say say some word, like acknowledge when I, mm-hmm, mm-hmm, that I shouldn't stop, I just continue.

1:02:46So the question of like phone interruption, like whether it's like, should I stop or not? Like it's a big thing of what perceived as like natural conversations. And that's also regulations, right? Because like a lot of times people want to build AI chatbot, voice chatbots that sound like humans, try to like trick users into thinking that they're talking to humans. But also, right, maybe potential regulations saying like, okay, you have to disclose to users when to talk if the bot is talking to you is human or AI. So I think just like, just a whole space, I think it's not quite as small as you think, but it's not quite like an AI foundation model problem, right?

1:03:28Because like a human interruption detection is actually a classical motioning problem. It's a different framing that like you can build classifier for that. or like the question of like let us see it actually is a massive engineering challenge not an AI challenge of course there can be an AI challenge because people are trying to build a voice-to-voice model so instead of having like having to first like transcribe the voice from me into text and then get a model to generate text answer and get another model to like turn from text to speech you can send you voice your voice directly so that is something we're working on but it's like very hard um yeah so so yeah so like even audio I think of it is like It's easier than video, right?

1:04:08Because video have like both image and voice. It's already like pretty hard. So I think it's a lot of challenges in that space. That was an awesome list of things. Let me mirror them back real quick. So what you're predicting in the next few years, things that will change in the way we work. And these actually resonate with so many conversations I've had on this podcast. So it's just kind of doubling down on where things are heading. One is the blurring of lines between different functions instead of just like design engineering. Everyone's going to be doing a lot of different things now. Two is just more of work being automated with agents and all these AI tools and just, in theory, productivity going up.

1:04:47Third is shifting from pre-training models to post-training, fine-tuning and things like that. Because to your point, models maybe are slowing down in how smart they're getting. Although I'll point folks to the chat with the co-founder of Anthropic. He made a really good point here. He's like, we're really bad at understanding what exponentials feel like. we're in the middle of that and also models are being released more often so the difference between them we may not notice because they're just happening more often versus gpt3 came out like a year i don't know after before after gpt2 so uh maybe true maybe not and then the fourth point you made is this idea of multimodal investing in multimodal experiences i cannot wait for chat gpt voice mode to get better at interruption like exactly what you're saying i'm just like talking to it and then someone makes a little sound it's like okay and then you have to and then it's like and then it stops talking.

1:05:35It's so annoying. I'm shocked that we don't have better voice assistant at home yet. I think I have been testing out a bunch. I keep hoping, oh my God, Zach would be the one. And then I don't know how many of them I just had to get away because they're not that good. I think it's coming. I hear it's coming. Anthropik's working with someone that I don't know if it's launched or not yet. Yeah. I was certainly wanting to bring back to what you mentioned about the your guest from Anthropik mentioned about the performance improvement. I think there's a big change. I think like this difference between a model-based capability.

1:06:08So I'm talking about like the pre-trained model, right? Versus the perceived performance. So let's say it's like a machine thought about like, are you familiar with the term test-time compute? I don't think so. Help us understand. So the idea is it's like, okay, like you have a fixed amount of compute, right? So you're going to spend a lot of compute on pre-training or training the model. pre-training and then I've spent a lot of some compute on like 5-tuning and the ratio of like pre-training to post-training compute is like crazy very different even different now um and also like since then it has to spend compute uh on like generating inference when I have a trained and 5-tuning model and now you want to like survey to users so I might type of questions or prompt and it's like generate like do inference like and that requires a compute and like you say people discussion of like uh should i spend more compute on like pre-tuning or fire tuning or inference right because like inference and people found out like test time compute so like spending more compute on inference is like called like test time like uh compute uh like as a strategy of like just allocating more resources compute resource to generate uh inference when i shouldn't bring better performance and how does that do it like let's say um let's say you have a math questions right and maybe instead of just jerking one answer i can just like four different answers and say okay whichever is uh the best according to some standard uh or like okay i have four answers and then maybe like three of them say 482 and one of them said like 20.

1:07:38okay three of them in the in agreement so the answer should be 482 right so like just people shouldn't generate a bunch of it or another thing is like a lot of time like reasoning uh thinking it's just like people should like generate more thinking tokens and spend more time thinking before showing the final answers. It's like require more compute but it's like give it more better performance. So yeah, so I think it's like from the user perspective, right? Like when the model spend more time exploring different potential answers, thinking longer, it can give you much better final answers. But the base model itself does not change.

1:08:15Does that make sense? Yes, that does. Absolutely. That is a good corollary to Ben Mann's point. Yeah. Chip, we covered a lot of ground. I've gone through everything I was hoping to learn and more. Before we get to our very exciting lightning round, is there anything else that you wanted to share? Anything else you want to leave listeners with? So I do work with a few companies that does these things of like, they want employees to come up with ideas. So there's a big debate on like, what is a better way for a strategy? Should it be top down or like bottom up? right should like executive come up with like one or two like killer use case and like everyone like allocate resource to that or like should you give engineers and pms and smart people like come up with ideas and it makes this a mixture of both so some companies it was like okay we hire a bunch of smart people like let's see like what they come up with and they organize like more than hackathons or like internal challenge to get people to build product and one thing that um I noticed it's like a lot of people just like don't know what you built.

1:09:16And it shocked me. Like why? I feel like we are in some kind of like an ideal crisis, right? Now we have all these really cool tools to have you like do everything from scratch. I can have you like design. It can have you like write code. You can have your website. So in theory, we should see a lot more. But at the same time, people are like somehow stuck. Like they don't know what to build. And I think it's like maybe it's a lot of had to do with like maybe like society expectations. because we have gone through we have gone into this phase of specializations. People are very highly specialized and people are supposed to do focus on one thing really well instead of being a big picture.

1:09:54And we don't have a big picture of you. It's hard to come up with ideas of what you build. So I know when I work with this company on this hackathon, we do work out how to come up with a guideline and how to come up with ideas. And usually what we think of is like, okay, one tip is like go look from the last week right like for a week just like pay attention to what you do and what frustrates you and when something frustrates you think like is there anything we can do is there like can it be done a different way so it's not frustrating and you can talk like people can swap to accept subnobes or teams and if you see they come on frustrations maybe just something you can think about is just to build something about that so yeah so i feel like um just like notice like how we work uh thinking of like ways like constantly ask questions like how can be better and then i just build something to like address the frustrations i think it's a good way to just like learn and adopt ai i think people have felt exactly what you're describing every time they open up one of these vibe coding tools whether you could just describe anything you want i'm like i don't know what do i want and and i love this very tactical piece of advice just like what frustrates you just pay attention to where you're frustrated for example i just built a very cool little vibe coded app.

1:11:03I was working on a newsletter post inside Google Docs. And I pasted all these images into the Google Doc from screenshots and stuff. And then I forgot, oh, yeah, you can't take images out of Google Docs. It's like this Hotel California experience where you can paste stuff into it. Very hard to get images back out. So I just went to all the vibe coded tools and just build an app that I can give you a Google Doc URL and let me download all the images automatically. And it worked amazingly well. And I made it really cute. And I'll link to it in the show notes. Oh, I would love to see that. I'm very bullish on using AI, just create like micro tools.

1:11:36It's just something that's like make your life a bit easier. And 100%, I feel like that's one of the main ways people are using these tools, just like a little niche problem they have. With that, Chip, we've reached our very exciting lightning round. I've got five questions for you. Are you ready? Yeah, always. I don't know. It depends on how hard the questions are. They're very consistent across every guest. So I imagine you've heard them before. First question, what are two or three books that you find yourself recommending most to other people? Oof, I'm really terrified of, like, book recommendations because I feel like what books or people should read really depends on what they want and where they're in life and where they want to get you.

1:12:18But I guess several books I do think have really changed the way I think and see the world. So one thing is a selfish gene. It's like you understand. It actually changed. it actually helped me with the question like whether I want to have kids or not because it's like understanding more of like yeah a lot of our functions of where we operate is the functions of our genes and genes want to do one thing was like to procreate so yes in a little way but it's like so but also propose another thing is like so everyone wants to live forever right and maybe is not like consciously but subconsciously we do we do want that and and i said two ways like one is like by genes like genes one's just like once i continue forever but also there are two ideas um i think there's something going to mean uh it's like being able if you have some ideas out there and then it's like last for a long time there's one way to like live on i know it's like it's a little bit like um abstract but i thought it's very interesting the other books i really really like is like from like um the book from uh singaporean um previous um i think he's as a called as the father of singapore i don't know like um i don't know what's the title it but like he did so he was the one who led singapore from uh he's changed uh singapore from a third world country to a fourth world country within 25 years and i have never seen any country leaders spend so much effort into like putting down his thought on like how to build a country uh like like that But yeah, they talk a lot about public policy, how to create policies and encourage people to do the right things that is good for the nations.

1:13:55And also talking about foreign affairs, foreign policies, the liberation of the country with others. So it's a really good book to think about. For me, it's a system thinking, but it's a different kind of system, which is a country, which a lot of us don't get a chance to ever experience in our life. So it's good to learn about that. What was the name of that second book? it's called like from third to first world fashion i think we have it somewhere here yeah there it is very show and tell that's that's awesome i definitely want to read that that's a really good tip i've heard a lot about just the impact he's had and i've seen all these videos on twitter just his really wise insights into how to build a thriving society and clearly it worked like how does he have time to write such a tech book it's like insane that is claude please summarize i'm just joking uh by the way selfish teen i also absolutely love that book that is such a good choice it's such an under the radar kind of book that really changed the way i see the world as well uh so really good pick okay next question do you have a favorite recent movie or tv show you really enjoyed so i watch a lot of movie and tv shows as a research uh because i i working on my first novel and i recently uh sold it so i'm interested in like what makes it's a drama it's not a science fiction or uh anything that like take people usually read so it's very like i know it's a very um out of the left field and like very um so it's like reading watching tv to see like what kind of stories become popular trying to understand the truth and stuff like that so i'm not sure if the audience are like well what's one what's one that taught you something about writing i think it's like uh yamshi palace it's a chinese tv show cool okay i haven't heard that one on the podcast before okay cool next question do you have a life motto that you often think about come back to when you're dealing with something hard whether it's in work or in life this sounds very nihilist i think to say it's like in the end nothing really matters uh usually I usually think of like in the grand scheme of things like in a billion years, nothing will like no one will ever be there.

1:16:02I think, okay, someone will argue with me about that. So I go to things like, so my theory is like in a billion years, like none of us will never exist. So like whatever like messy things, like crazy things we do or like how bad we do it. I mean, no one will remember, we'll be there should remember it. And I think in a way it's like, it sounds scary, but it's very liberating because it just allows me to say, okay, let's just try things out. right like why does it matter and this is a story of like recently um so I have a some family member who passed away recently and I was talking to my dad because I couldn't be home for that I was asking my dad like okay say anything I can do to make the person like oh something like comfort so that anything you can get the person and my dad was just like what can he possibly want at this moment like it just made me feel like at the end of life like there's nothing that can bring you like material can bring you joy there's no like money, no product nothing and in a way it's being like okay what really do I really care about at the end of the day so I guess it's like I think about it it's like okay maybe I fail it, maybe I don't get that contract, maybe do things like but at the end of life like I don't think that actually really matters so in a way it's like it's kind of liberating I know you said it might be nihilistic, this is what Steve Jobs shared too in one of his most famous speeches It's just, we will all die someday.

1:17:21So don't take things so seriously. And it is freeing, absolutely. It just makes you appreciate every moment, every day you have. Just like, yeah, let's just do something hard and scary. Okay, final question. You talked about how you're writing a novel. Most people in tech have never written something creative and fiction. What's just like one thing you learned in the process about how to write better stories, better fiction? A lot of time when we read, we get tripped up by some small things. So I think I want to do creative writing because I just want to become a better writer. And it tells us maybe try a different audience to help me become better at anticipating what this different type of audience would want to hear and what they care about.

1:18:05So it's a way for me to get up. So I think about writing, even any kind of content creation, it's about predicting the user's reactions. The next token. Just kidding. Yeah. So like you do a podcast, it's like, okay, what kind of things that the users could find engaging, right? And I find it's like a little bit like, and a lot of companies, like you have like launch a product, you have a narrative coming out. So, okay, what kind, how do we position this product in a way that like users would want, right? So I feel like I have done technical writing for a while. And I felt like I have had some experience like trying to predict what engineers would want to hear.

1:18:43or like care about but then I don't have an experience like this completely different type of audience so that's what I want you like career writing writing a story and that's why I was doing a lot of research on the question I mean going research I actually enjoy a lot like watching a lot of dramas I just see like what it will like so one thing that I care about is just like I think a lot is like emotional journey it was from an editor right so like when we write something like we care about like how users would feel like across the story like we want something in the beginning right we want something just like we need to have a hook so that people continue reading but we also don't want too much of like drama because we'll get like too tired right like uh because like the emotion exhausted like because it's like you're being like emotionally manipulated like a lot of time so it's like emotional emotional journey maybe have like some some climax or like some something more chill or like maybe like and so care about another the things I didn't realize like for me for technical writing you entirely focus on the content like the argument is very impersonal right like it like for example like people like ml compilers like doesn't matter if they like the person telling them about compiler or not right because it's just like objective like like but like for novel people care about like character likability so so like in the first version is my story and makes the characters like a little bit more like very logical, very rational, and just does everything just like very rationally.

1:20:11And then the feedback I got is I have a very good friend and he was like, he's an amazing person, he's a great person. And he was like, shit, I'll be honest with you, I hate that person. So it doesn't matter as a story, it's just like the person is so unlikable, just like he doesn't punch you. So in the second version, it makes that character more likable. Like how she makes that character more likable is that you put in some vulnerability. like sometimes like okay maybe this person like has setback because sometimes we can relate to it so you know in a lot of ways like it's very interesting it's like a lot of it it's like yeah a lot of it is about like understanding the emotional bits like how the users feel not just about the story but also about the characters that is so interesting wow I learned a lot more there than I thought that was awesome really good example Chip two final questions where can folks find you online if they want to reach out and maybe work with you or maybe even just share the stuff that you offer if folks want to reach out?

1:21:05And then how can listeners be useful to you? I'm like, I'm on social media, LinkedIn, Twitter. I do post a lot, but I keep telling myself that I should do more because I kind of like the competition with readers. So I'm actually about to start a subspec. So I have like a placeholder for subspec right now and I'm thinking of doing it for more system thinking because I think it's a very interesting skill. And so like thinking of doing a YouTube channel on book reviews and basically books that help you think better. So I think it's the first book I'm going to review is probably like this book because it's like my favorite book growing up.

1:21:42And I have been like keep on reading it. So, yeah. So how can it be helpful? Like send me books that you like, books that help you have changed the way you think or change you the way you do anything. So I would appreciate it. Amazing. I'm excited to read that book. Chip, thank you so much for being here. Thank you so much, Lenny, for having me. Bye, everyone. Thank you so much for listening. If you found this valuable, you can subscribe to the show on Apple Podcasts, Spotify, or your favorite podcast app. Also, please consider giving us a rating or leaving a review, as that really helps other listeners find the podcast.

1:22:25You can find all past episodes or learn more about the show at lennyspodcast.com. See you in the next episode!

From the publisher

Chip Huyen is a core developer on Nvidia’s Nemo platform, a former AI researcher at Netflix, and taught machine learning at Stanford. She’s a two-time founder and the author of two widely read books on AI, including AI Engineering, which has been the most-read book on the O’Reilly platform since its launch. Unlike many AI commentators, Chip has built multiple successful AI products and platforms and works directly with enterprises on their AI strategies, giving her unique visibility into what’s actually happening inside companies building AI products.

We discuss:

1. What people think makes AI apps better vs. what actually makes AI apps better

2. What pre-training vs. post-training is, and why fine-tuning should be your last resort

3. How RLHF (reinforcement learning from human feedback) actually works

4. Why data quality matters more than which vector database you choose

5. Why high performers are seeing the most gains from AI coding tools

6. Why most AI problems are actually UX issues

—

Brought to you by:

Dscout—The UX platform to capture insights at every stage: from ideation to production: https://www.dscout.com/

Justworks—The all-in-one HR solution for managing your small business with confidence: https://www.justworks.com

Persona—A global leader in digital identity verification: https://withpersona.com/lenny

—

Where to find Chip Huyen:

• X: https://x.com/chipro

• LinkedIn: https://www.linkedin.com/in/chiphuyen/

• Website: https://huyenchip.com/

• Substack: https://substack.com/@chiphuyen

—

Where to find Lenny:

• Newsletter: https://www.lennysnewsletter.com

• X: https://twitter.com/lennysan

• LinkedIn: https://www.linkedin.com/in/lennyrachitsky/

—

In this episode, we cover:

(00:00) Introduction to Chip Huyen

(04:28) Chip’s viral LinkedIn post

(07:05) Understanding AI training: pre-training vs. post-training

(08:50) Language modeling explained

(13:55) The importance of post-training

(15:20) Reinforcement learning and human feedback

(22:23) The importance of evals in AI development

(31:55) Retrieval augmented generation (RAG) explained

(38:50) Challenges in AI tool adoption

(43:19) Challenges in measuring productivity

(45:20) The three-bucket test

(49:10) The future of engineering roles

(55:31) ML Engineers vs. AI engineers

(57:12) Looking forward: the impact of AI

(01:05:48) Model capabilities vs. perceived performance

(01:08:23) Lightning round and final thoughts

—

Referenced:

• Chip’s LinkedIn post on what actually improves AI apps: https://www.linkedin.com/posts/chiphuyen_aiapplications-aiengineering-activity-7358971409227792384-y0mf/

• Prediction and Entropy of Printed English: https://www.princeton.edu/~wbialek/rome/refs/shannon_51.pdf

• Why experts writing AI evals is creating the fastest-growing companies in history | Brendan Foody (CEO of Mercor): https://www.lennysnewsletter.com/p/experts-writing-ai-evals-brendan-foody

•Inside the expert network training every frontier AI model | Garrett Lord (Handshake CEO): https://www.lennysnewsletter.com/p/inside-handshake-garrett-lord

• First interview with Scale AI’s CEO: $14B Meta deal, what’s working in enterprise AI, and what frontier labs are building next | Jason Droege: https://www.lennysnewsletter.com/p/first-interview-with-scale-ais-ceo-jason-droege

• Anthropic’s CPO on what comes next | Mike Krieger (co-founder of Instagram): https://www.lennysnewsletter.com/p/anthropics-cpo-heres-what-comes-next

• Why AI evals are the hottest new skill for product builders | Hamel Husain & Shreya Shankar (creators of the #1 eval course): https://www.lennysnewsletter.com/p/why-ai-evals-are-the-hottest-new-skill

• The rise of Cursor: The $300M ARR AI tool that engineers can’t stop using | Michael Truell (co-founder and CEO): https://www.lennysnewsletter.com/p/the-rise-of-cursor-michael-truell

• Stanford webinar—How AI Is Changing Coding and Education, Andrew Ng & Mehran Sahami: https://www.youtube.com/watch?v=J91_npj0Nfw

• He saved OpenAI, invented the “Like” button, and built Google Maps: Bret Taylor on the future of careers, coding, agents, and more: https://www.lennysnewsletter.com/p/he-saved-openai-bret-taylor

• Anthropic co-founder on quitting OpenAI, AGI predictions, $100M talent wars, 20% unemployment, and the nightmare scenarios keeping him up at night | Ben Mann: https://www.lennysnewsletter.com/p/anthropic-co-founder-benjamin-mann

• Lenny’s vibe-coded app made on Lovable: https://gdoc-images-grab.lovable.app/

• Story of Yanxi Palace: https://www.imdb.com/title/tt8865016/

• Steve Jobs’s quote: https://www.goodreads.com/quotes/427317-remembering-that-i-ll-be-dead-soon-is-the-most-important

—

Recommended books:

• The Complete Sherlock Holmes: https://www.amazon.com/Complete-Sherlock-Holmes-Volumes/dp/0553328255

• AI Engineering: Building Applications with Foundation Models: https://www.amazon.com/AI-Engineering-Building-Applications-Foundation/dp/1098166302

• The Selfish Gene: https://www.amazon.com/Selfish-Gene-Anniversary-Introduction/dp/0199291152

• From Third World to First: The Singapore Story: 1965-2000: https://www.amazon.com/Third-World-First-Singapore-1965-2000/dp/0060197765

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email podcast@lennyrachitsky.com.

Lenny may be an investor in the companies discussed.



To hear more, visit www.lennysnewsletter.com

More from Lenny's Podcast: Product | Career | Growth

All 287 episodes
Al Engineering 101 with Chip Huyen (Nvidia, Stanford, Netflix)Lenny's Podcast: Product | Career | Growth · 1 h 23 min
Listen in VO