AI Engineering Pitfalls with Chip Huyen - #715

21 Jan 2025 · 58 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Episode Summary: AI Engineering Pitfalls with Chip Huyen - #715

Podcast Information

  • Title: The TWIML AI Podcast
  • Host: Sam Charrington
  • Guest: Chip Huyen, independent AI researcher and author of "AI Engineering"
  • Episode Date: Not specified
  • Listen to Full Episode: [TWIML AI Podcast Episode #715](https://twimlai.com/go/715)

Episode Overview In this episode, Sam Charrington interviews Chip Huyen, who discusses her new book "AI Engineering." The conversation delves into:

  • Definition and distinction of AI engineering from traditional machine learning engineering.
  • Common pitfalls in AI system engineering.
  • Strategies to avoid these pitfalls.
  • The role of AI agents and their limitations.
  • Importance of planning and effective tool utilization in AI systems.
  • Evaluation processes for AI systems.
  • The influence of open-source models and the potential of synthetic data.
  • Predictions for advancements in AI engineering.

Key Concepts and Discussions

  1. AI Engineering Defined
  2. AI Engineering vs. Traditional ML Engineering
  3. AI engineering encompasses new methodologies and tools distinct from classical ML approaches.
  4. Foundational principles from ML engineering can still apply, but new challenges arise with AI models, especially generative AI.
  1. Common Pitfalls in AI Engineering
  2. Complexity Management
  3. Beginners often jump into complex frameworks without understanding the fundamentals.
  4. It's advised to start simple, experiment with models directly, and gradually introduce abstractions.
  • Evaluation Challenges
  • Many fail to establish clear evaluation metrics for AI outputs, leading to reliance on vague standards.
  • Effective evaluation processes require well-defined criteria, human oversight, and systematic approaches.
  1. Understanding AI Agents
  2. Definition and Functionality
  3. Agents are defined as entities that perceive their environment and act upon it, beyond simple interactions seen in base LLMs (large language models).
  4. Current agents may still lack the sophistication expected for autonomous operation.
  • Limitations and Future Directions
  • There is skepticism about whether agents are simply rebranded LLMs with tools.
  • The effectiveness of agents hinges on their planning abilities, problem-solving capabilities, and the range of tools at their disposal.
  1. Evaluation in AI Systems
  2. Importance of Structured Evaluation
  3. Regularly inspecting data and actively involving human review can improve model reliability.
  4. Developing concrete benchmarks and understanding the AI's capabilities are critical to successful deployment.
  • Human Oversight
  • Continuous human involvement is essential to validate AI outputs and ensure they align with user expectations.
  1. The Role of Open-Source and Synthetic Data
  2. Open-Source Advances
  3. The increase in powerful open-source models provides opportunities for experimentation and innovation without reliance on corporate models.
  4. Developers can deploy open-source models, enabling better governance and customization.
  • Synthetic Data Generation
  • The use of synthetic data is becoming prevalent, with approaches to validate the quality of generated data being crucial.
  • Techniques such as back-translating code or using AI to generate documentation can help ensure data integrity.
  1. Future Predictions
  2. Evolving Landscape of AI Engineering
  3. Anticipation of advancements in governance, compliance, and improved tools for building agents is expressed.
  4. Acknowledgment that compute resources will continue to be a key factor in AI development.
  • Long-Term Outlook
  • The field is expected to see more practical applications and clearer regulations for AI deployment, fostering greater confidence among enterprises.

Key Takeaways

  • AI engineering is a rapidly evolving field requiring a blend of traditional skills and new methodologies.
  • Avoiding complexity and focusing on robust evaluation processes can significantly enhance AI project success.
  • Open-source models and synthetic data are reshaping the landscape, offering new avenues for exploration and application.
  • Continuous learning and adaptation are crucial as the AI field progresses.

Conclusion This episode presents valuable insights from Chip Huyen regarding the challenges and future of AI engineering. The discussion emphasizes the need for a thoughtful approach to evaluate and implement AI systems effectively while navigating the complexities associated with generative AI and agents.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00If you have a question, finding the answer, if the answer exists, is a lot easier analysis with AI. But I do think what is hard is like coming up with the right questions. So that made me realize it's like, okay, as a writer, as a technical communicator, that I need to be able to like think through things and come up with like, ask the right questions.

0:33All right, everyone, welcome to another episode of the TwiML AI podcast. I am your host, Sam Sherrington, and today I'm joined by Chip Nguyen. Chip is an independent AI researcher. Before we get going, be sure to hit that subscribe button wherever you're listening to today's show. Chip, welcome to the podcast. Hello, Sam. Hi, everyone. Thank you so much for having me here. Oh, by the way, I'm like an independent researcher in general. AI is just like one of the things I'm interested in. Awesome. Awesome. Well, it is great to connect. It's been a little while. I think a couple of years since we last spoke.

1:10And obviously, a lot has happened in the world of AI since then. Maybe you can catch us up on what you've been up to. I think you were just starting Claypot at the time when we spoke. You were getting into real time and some other things. What's been new for you for the past couple of years? Yeah. Yeah. So we started Claypot AI and then we sold the company last year and then stayed with the acquirer for a while. It's actually like for things to settle. And this year, and after that's why I'm being unemployed for this year and I'm pretty really unhappily sorry pretty happily unemployed not like i guess it's better than like unhappily employed uh but it's been pretty good um yeah and i yes i think this is here i feel like it's been the last few years have been pretty intense especially uh with like running a startup and then with all this craziness about AI.

2:09So this year I just decided, let's just try to pursue my interests and see why they pick me and try not to have a lot of deadlines and expectations and goals. Wow, I do sound a bit lazy, don't I? Not at all. And I don't think anyone that follows you would call you lazy. You just had a book come out from that book. You just posted a really interesting post about agents and we'll be digging into that topic. Maybe we can start, though, with independent researcher. Tell me, do you structure that in any way? You were specific to call out that it's not AI. What are some of the other things that you're researching?

2:52I think AI as a tool, and it's definitely a very fascinating tool. It's really hard to avoid AI nowadays. It's just like I have this test, like if I go to a party, how long until somebody brings up AI? This is like in the first five minutes. Like it's really impossible to escape it. So definitely AI has to be a part of it. But I'm interested in like problem solving. So there are several problems I'm interested in. And I just want to see like how to solve these problems. And if AI is part of the solution, it's great. But if not, I'm very happy. I would like to be very happy to look into other directions too.

3:30What are some of those key problems that you're interested in? um so so one thing is just like i think that's the way people read and write and learn is going to be so different so when i was writing the book uh my book i read a lot of um papers like a lot like i think my uh my book has about reference alone like about like over a thousand links and of them is like 400 somethings are like archive papers so that as only the papers i referenced in the book and I read a lot more that were not referenced so and it's a process of reading these papers I realized as a way I read has changed so for example in the past right I would like read uh so abstract and read the I read the I read the conclusions discussions and then I skimmed through the paper but now it's just like I felt myself like with AI I become a lot like the time it takes for me to read a paper like it's a lot less like I could read the abstract still but then most of it I would throw into ChatGPT or Cloud, and then I start asking questions.

4:32So one thing I usually ask is, especially for me, I understand a new technology better in comparison to what is the old technologies. So if there's a new technique coming out, I would like, okay, but how is it different? How is this improvement? And why does it work compared to the last techniques? So I would put two papers or three papers into ChatGPT and ask to explain to me how is that an improvement? like why from the previous ones and i was doing it i was just like okay if i don't read things like end-to-end anymore then why do i write things from end-to-end you know like why do i write the whole thing people like that so it does it does make me think about like what kind of writing is harder to automate and then i realized something is that like first um if you have of questions finding the answer if the answer exists it's a lot easier nowadays with ai like yeah first of all if you ask it to explain like what what is x you ask chachibidi it can come up with like a thousand of different definitions like you can explain to you as if you were five you can explain to you as if you were like i don't know a dancer or like i don't know it's really good at coming up with like explanations or definition for you but i do think what is hard is like coming up with the right questions.

5:48So that made me realize it's like, okay, as a writer, as a technical communicator, that I need to be able to like think through things and come up with like, ask the right questions. So I'm less pressured in like asking, like trying to like define things or explain things. I mean, like I still want to do a good job at it, but I mean, I'm a lot more concerned with like thinking through things, like structure things, asking the right questions. So that is one of the interests you asked me. But yeah, I do. I'm interested in learning, education, maybe information consumption in general. That makes me think a little bit about this joke that we've heard or seen come up a lot recently, which is people spending a lot of time using ChatGPT to help them take their little blurb and expand it into something formal.

6:44and then people on the other side are taking that formal thing and expanding it back into little blurbs or bullets. How do you see that evolving communication? I do think that it will change a lot now in the way we communicate or the way we interact with each other. So the way of communications... So one example is just like, you know, search engine speak. It's just like people have started learning how to type. gets gets a like ask just engineer google so so he gets the right thing and perhaps start learning to like talk like that with like chat gpd one is it's very interesting so i have i have a friend who who has like a quarantine a covid baby right so the so the kids like spend the first like two years being very isolated because of parents uh a nanny and alexa and she when she was free like i noticed that she talked a little bit like with people the way she could talk with alexa she'll be like mom get me this dad get me this right it's just like so so i felt like we can see that so so i have no doubt that's like we interact with mobile ai wouldn't change the way like in san francisco have you seen that people like have already been studying uh self-driving cars um it's because like you you you know you know that like self-driving would always stop for you to cross the street so when you see several cars it was like okay i'm just gonna cross now i don't i don't have to wait for the for the for the crosswalk right or like sometimes i was just like let's annoy i think i'm not saying i did it i'm not saying some of my friends did it i just like to see certain regards like let's let's annoy the center in god like given that there's no one inside but i feel like um i think with definitely change uh one thing i'm actually really curious about is that um um i've seen like a few companies that uh provide um ai companionship and I think I was very so I have I was able to see some of that data and it's a little bit crazy like how much time people spend on like chatting with AI like it's really insane if you go first of all I think you can pre-see some of it if you go to like some subreddit of like character AI or like some others and it's just like people spend a lot of time talking with AI and it just made me feel like let's imagine like you are younger today right you were like and and then i maybe like not maybe you know in the teens or like twins like pretty young and you you spend a lot of time talking with this ai and it was so used to like ai like just accommodating you like just saying exactly what you want to hear like always be there for you then how was that experience compared to talking to humans like why would i bother to talk to humans who i cannot program who i cannot like just say things and don't want to hear.

9:35So it just made me feel like, hmm, how is this going to be in the future? But anyway, I'm not sure. I don't remember what our question was. I just went to Iraq. Yeah, no, you know, we were talking about the question was how you think our interaction with AI and more specifically, I was asking about our use of AI tools would change communication, and kind of reflecting on that loop that you used in reviewing the papers, how that has shifted and how your writing has shifted.

10:13Did you, to what degree beyond using the chat GPT and Claw to summarize papers, what other AI tools did you use in the writing of the book? Oof, I use a lot of it. Not a lot of tools. As I say, I use some tools a lot of times. So I actually, like, did, I actually use Jachipiti and Cloud. So, and also use, like, something like, okay, I don't watch, I feel like it's weird mentioning those names because it sounds like promotions. I do use a few other tools, just, like, heavy editing, like, catching, like, you know, just, like, general writing assistant. And I just, like, recently, like, just exported Jachipiti.

10:55and like during the so i found out like during the period of writing the book i i had about like 3 000 conversations with judge pd so a lot of them is not just like summarize paper comparing paper but also like uh do the analysis um so yeah so i just trying to understand or like help me like create a graph or like just brainstorm like ways of like how to explain the concept some validating ideas um i use cloth a lot because one of the interests i have right now is write a math novel which i really want to do but i felt like there's pretty no market for it uh but but i just like like math i think math is very beautiful and people keep asking me like hey is there like any like so so i did competitive math for a while when i was a child in 10 years and I think it's like people keep asking me hey is competitive math useful I do study like number theory and stuff I do use it nowadays and I feel like a lot of it probably not but I do think it's like mathematical thinking it's like really really important and it's just once you write like a novel to like just like how beautiful math is and I did you like cloud and chat GPT just like to like help me like think through the character development is that realistic like is that probable and like coming up with like examples because sometimes you know it's like it's pretty they're both pretty good at coming up with examples and in the process I found out it's like it's pretty useful for like writing a book on AI as well because I learned a lot about like what they are good at and what they are not good at and also like just in general just helping understanding have giving more ideas so it gave me a bunch of ideas about different benchmarking and design like how to evaluate different capabilities of these different models because for one example right like creep writing um i was talking to someone at like um anthropic and that person told me it's like oh claude is really good at creep writing and i'm just like how do you know do you have like internal benchmark for it and this person was like no people on twitter told me so and i'm like hmm even if people were really good at it really smart right and apparently like this kind of thing is like quite important to them i still don't have benchmarks to measure things like that then how like it must be like really hard and that's made me think about like how we can design like benchmarks to evaluate things like script writing storytelling um yeah so i have been like thinking about like a different uh approaches to like benchmark design so that's one of the things i want to pursue in my like leisure time or like unemployed period.

13:38Nice, nice. Well, let's jump into the book a little bit. The title of the book is AI Engineering. The book is mostly about generative AI as opposed to, you know, other traditional machine learning, for example. Maybe start by why do you call it AI engineering as opposed to something around generative AI or, you know, any number of things that you cover in the book, what does AI engineering mean to you? Yeah, no, that's a good question. And I did agonize about it a lot, about the name, quite a bit. So I think it's like, first it's like, should we stimulate machine engineering? Because machine engineering has been around for a while.

14:23And I do think that's why there's a lot of overlap. And a lot of the foundational things that we know from AI engineering can be applied. AI, like, foundation models do come with a lot of new things that I do think is, like, we need to talk about those new things. So we do need something to separate AI engineering, like this new thing from, like, traditional ML engineering. And there were a lot of terms flying around. Like, some of them were talking, like, LM ops, some of them were like AI ops, right? And I was like, I didn't really want to use the word that end with ops because I do think it's, like, a lot, there's just a lot less ops ratio.

15:00like operations about this right it's a lot more if everything's behind the service yeah yeah yeah um and i think there's a lot more creativity like uh more problems it's like it's a lot more i'm not saying operations not fun i'm just saying this can be a lot more fun um so so i did ask a bunch of people so at that time i i knew quite a few people who are already building like really cool applications with jet of ai and i asked them like so so how would you call what you're doing and they were like engineering so after you got like 20 people saying that i'm saying okay i'm just gonna go with what people call it you know i'm not gonna pick my own you know so interesting interesting interesting um and one of the topics that you cover in the book is agents you know before we dig into some of the way you're thinking about agents nowadays i'm curious about your meta thinking about agents and that is um you know i see a lot of debate raging on the the social medias about whether agents are an are a thing or whether agents are just kind of rebranded llms like llms plus tools is you know now we're calling that agents um how you know assuming you kind of have thought about that a bit how do you think about that um and do you think kind of referring back to your comment earlier about asking the right question do you think that's an interesting question or is it you know just kind of throw away social media fodder and there are more interesting questions to ask about this and what are those questions i think it's it's an interesting question and i really want to go into this but can we just i just want to end the thought on like the engineering versus ML engineering and then we go into because I feel like we're going to be spending a lot of time talking about it.

16:54So yeah, this is on the note of like engineering and ML engineering. So I don't think they are like mutually exclusive. So I do think it's like in the vast majority of Jive AI system I have seen, it's just not like a Jive AI or traditional ML, a classical model. It's like it's both. So for example, you build like a customer support chatbot, right? Like so you have a request from like a customer before sending the request to like the ai model um you might want to have like a intent classification to see what the request is about maybe there's something like easy you can answer with an faq yes or router so that you can build with like a classical machining model or like as a scholar right after you get a response back from a from ai you may want to say hey is that safe does it contain pii right does it reveal anything you and say, or like whether it's good enough or not.

17:48So you can also build those scores with classical ML. So I do think it's like, it's very easy. They do go hand in hand. So in a lot of organizations, I see that they have the same team for engineering and engineering. And in some organizations, they separate them out. But I do think it's like for anyone who wants to build like solutions to real world problems, maybe like complex problems. it's very likely that that's a person we need both like traditional ML engineering engineering and I think like the terminologies I don't think we do I don't yeah I don't think we need to hang get hung up on the terminologies because even the same role title or different companies can mean very different thing even the same company has the same title can I do very different things yeah so I think just to close out that thought and okay next we talk about agents right And in contrast, while we don't need to get caught up in terminology about ML engineering versus AI engineering, it sounds like you do think the terminology around agents is interesting.

18:57So, do you know if I had this feeling, it's like you like something, but just hate the way it's called?

19:05yeah so so like i just traveled to vietnam recently but uh with my friend and he was like i realized this dish like it's like very tasty like over the charcoal vietnamese snack but then people try to sell the tourists and call it like vietnamese pizza i was like he was like why would you call it something so tasty vietnamese pizza like totally not that so um so so anyway so it's about the naming uh but um but yeah so um i think a lot of people don't like the word agent and a lot of people is like oh like you said it's like what is agent is a new thing or just like lm plus tools um for me um i try not so i try to avoid debate on on naming because i feel like So one thing I really hate about being a writer, I love writing, but one thing I really hate to write is definitions.

20:01Because I feel like no matter how I define something, someone's going to get upset. Because it will never be good enough. So I did have a section on agents in my book and it went with agents because that's what a lot of people call it. So I know a lot of people don't like it, but it's only something that people call it. And agent by itself is not a new term. And I have a lot of friends who just really hate the term. And there are already so many things with the name agent. Like I have reinforcement learning agents. You have user agents. You have a lot of different kind of agents, right? What is this agent?

20:39So I did go through a lot of reading AI books, some classic books on AI and see how they define agent. And I think definitions I quite like is that agent is anything that can interact, that can perceive the environment and interact with the environment. So that means... Yes. Yeah. So it's a great book, by the way. Which is written mid-90s. Yeah. To underscore the idea that this is not a new concept. Yeah. No, no, definitely. So an agent itself, right? So it has to perceive the environment. So it means that agent has to operate in an environment. So first of all, a coding agent, it operates inside, maybe coding maybe like in VS code or like in terminal, right?

21:28Or like if it's Charged GPT is an agent that operates maybe on the internet, right? It can like browse the internet. It can also like answer a question that you interact with it through the internet. So, and then to like interact with the environment, you will need like be able to perform a set of actions. So a lot of things I've interacted with nowadays are only agents. Like Charged GPT is certainly an agent, right? The action they can do is like execute code, browsing web, user calculator, like even generate image or retrieve image. So there's a lot of things JARP can do already in similar as Cloud.

22:03So agent is already here. But the capability of an agent, a lot of it is not what people hope for yet. Maybe people think it's like, oh, agent is like, it sounds cool and all. like you want to have something autonomous, I can have you do a lot of things. But maybe the promise is like the actual capability is not quite where people want it to be yet. But the question is, it's like how far or how big is the gap? And I'm pretty bullish on the improving capability. So I do think that the capability of an AI power agent depends on shooting. One is a set of tools you're going to get access to. And the second is ability to leverage those tools to solve problems.

22:50And the second aspect, we can go in planning, right? Like given a task from a user, how does the patient think? How does the agent think about using those tools to solve the task? May come up with a sequence of actions, like a roadmap to get the goal and maybe determine whether the goal has been accomplished, right? So it's a tool. So I think a lot of agents, like what we use today, LM, we think of them not as agents because they have a limited set of tools. But they have tools, right? But I do think the more tools we give AI models, the more powerful they're going to become. However, using tools well requires a lot of capability.

23:37It's a lot easier to use two or three tools well. it's a lot harder to use 3 ,000 tools well. So I have some particular benchmark, and I want to see how well a model can use a different number of tools. And in my personal, not peer review, I just want to put it out there, it's a personal benchmark. I feel like AI model fails very, very quickly with the number of tools increases. So the planning capability of AI has to improve. people to use their tools well and I solve the task well. Can I ask you a question about this? Your non-peer-reviewed research study on tool use? Have you looked at, like it strikes me that an interesting factor would be like the semantic distance between the tools.

24:32Like if there are a bunch of tools that are closely related to another Or, you know, inches to millimeters, millimeters to inches, inches to feet, that kind of thing. And then like a recipe tool, like maybe there's more confusion around the things that are closely related and, you know, less confusion around things that are different. Or maybe it's the opposite. I don't know. I'm wondering if you have any thoughts on that based on what you've played around with so far. I think it's a very good research question. so in my benchmark a lot of tools are like not quite similar there's a range of different tools but I do think this is a school like a very subset of papers I've read so it's like this go into like what kind of not as a kind of tool like there's a higher abstraction a higher level of abstraction than that so let's say it's like okay let's say like you write code, right?

25:34Write code and it's an abstraction. And then lower, you're going to write Python code, write C code, or writing codes, right? So when you ask AI to generate a plan, you can either ask to generate a very high-level plan, a very granular plan, right? It's just like planning quarterly planning, you can do monthly planning or weekly planning. So I think it's a different study on how good the generated plans can be like the different level of like randomity and also there's a question of like um you can ask some model to generate plan using the function name directly or should you generate something like in natural language and then have a translator from this natural language into the into the into the function names so so i do think there's like there's a lot of studies around that and i felt like a lot of assignments around agent is great uh but i still like feel like a lot of the excitement nowadays is about like hey how to build an agent very quickly right like hey this is framework to build agents in like five minutes and i'm always like okay build an agent and then what you're like but what do we use that agent for right so so i feel like i think like maybe we can have some of that energy directed toward like understanding like how do we properly structure these agents or like use that you know like yeah understanding how it uses tool like how to, so planning language or like just how to get the agent to do things like more correctly.

27:04It's going back to this, the initial question around agents versus LLMs and tools, or, you know, does agents equal LLMs and tools? I think that, you know, it is true that a lot of that is kind of semantics and not all that interesting, but an interesting potential aspect of that is if, you know, you want to explain to someone who knows LLMs and knows tools, you know, what opportunities there are in thinking about some broader universe of agents. And if that requires, you know, new skills or new things to learn or, you know, new approaches. How do you think about it from that perspective? Meaning, you know, if you've got someone who, you know, knows LLMs and tools pretty well, and it's coming to you and says, you know, is there anything that I need to learn beyond that to work with agents or to take advantage of, you know, agentic systems?

28:10you know what are the things that you talk about so i think the question of like whether agents is just m plus tool is the same as the questions whether a car is just a battery plus wheels you know like yeah you can have like on this component but is that enough for you to like want to buy it so there are a lot of other things about it first one has to like first like package into a product that people want to buy or like have it like be with your own problem like efficiently. So first we have an LMP can execute a lot of tools. And imagine you just keep on calling execute, like calling maybe open a million web pages and send out like 10 ,000 API columns.

28:54It can do that very well. But then what problems will that LMP plus tool can solve for you? How easy is that for the users to be able to specify a task and then get back the result? How can that be integrated seamlessly into the workflow? and then how do you as an Asian developer get all of this feedback from users and continue to improve the system? So I do think that is a big question. So definitely, like I say, LM is the brand. So I do all the planning and reflections and the tools. It's just how it can execute and perform actions to accomplish a task. But there's a whole ecosystem around this.

29:36You have to do system design. You have to think about and maybe supplement it with memory, think about how to evaluate the agent, different failure modes, like put guardrails around. Because the more tools the agent has access to, the more dangerous in a way it can be. Anyone says it's dangerous because it can scare a lot of people. But this thing is, we just need a higher level of security. If you have an LAM, just can I say read email versus an LAM that can send email. or if we have an LIM that can like beat you, like tell you the bank balance versus an LIM that can like transfer money out of your account.

30:18So like this is more action you can give the LIMs or more guardrails. And the more we can think through like how to catch the failure, make sure that things don't go wrong. And I do think that's a lot of work around that. And it's way more than just like LIM plus tools. When you see folks approaching, you know, getting started with agents or even kind of this broader picture of AI engineering, like what are the things that get people caught up? What are the mistakes that you see people making kind of over and over again? Yeah, that's an interesting question. I feel I have a lot of thoughts about it.

30:58They have to be careful not to make it sound like I think everyone's an idiot. Yeah.

31:06So I do think that engineering as a field, engineering field, is still pretty new. And with anything new, we're still learning how to do things right. We're still coming up with understand best practices and be around the right level of traction. So I clearly think it's a lot of like making mistakes is inevitable, right? And that's how we learn. So there's certain mistakes or common pitfalls, I see. so first of all one thing is that a lot of time people start way too complex too early right like first of all you see some new framework coming out let's try it out let's build an agent right let's multi-agent yeah and it's very easy to get caught up in this like the layers instead of like thinking about what you actually want to do so first of all before like using a framework that's like how abstracts away on the API columns, like provide with like some pre-built prompts.

32:07Maybe just try out, like call the model directly, write their own prompts, you know, and just try to see how things go. And because like, actually during the process, during research for my book, I actually went into the code bases, a lot of like popular frameworks, and I saw a bunch of them built in prompts, and a lot of them have typos. because we need to understand that framework developers can make mistakes as well. And a lot of times using the right tools can help, but using the tool too early can introduce unnecessary and necessary

32:49complexities and hideaway mistakes. So I think it's good to just understand how things work underneath the hood really, really well before introducing different levels of like tools to make it easier um is that might change the future though okay i do think like over time if we have like much trust in frameworks or like we have best practices and then we can trust us like these frameworks and are reliable and do what expected then i think it's like it's okay to start more more complex because like nobody nowadays actually they write assembly code anymore right we we do go like absolute level attraction abstraction over time Another example is a lot of people when they do Rack, Retrieval Augmented Generations, people start with vector search, using very complex databases and embedded embeddings and stuff.

33:37Whereas you can do something a lot similar, just like Turnbay, Retrieval, or something very old school system like BM25. It's actually really, really hard to beat. another thing is just like that I get a lot of like debate about is like how much AI knowledge do you need to start building applications so in a way like with engineering there is a lot less ML like before if you wanted to build an AI application you would need to build some models yourself but nowadays you just access this build model right and it's an incredible thing because now like anyone Like, it's, like, a really low entry barrier for people to, like, access AI and use AI for their needs.

34:23And I do think that you can do a lot, like, starting out with it. But I do think it's, like, to use it well or, like, to, like, avoid headaches. Then some understanding of how AI works and neighborhood can help you go a long way. So, like, first, like, it can help you explain some, like, weird behaviors of AI. For example, understanding how AI sample their responses can help you maybe understand more about why hallucinations happen or why the AI changes its responses. Or it can help you come up with strategy to improve the model responses by changing different ways of the model sample responses.

35:07Or maybe you can learn from the past. so for example like hallucination again there's like a sub-fill of NLBs that can study something like why the models output things that are not consistent with what's given to it for example like the question like hey I wanted to I actually like summarize this book why the summary not consistent with the book but it comes up with things that's not in the book right so that's filled with natural language inference like textual entailment. So in textual entailment, you have a hypothesis, right? And they have a statement. And you want to determine whether this statement is consistent or can be derived from the hypothesis.

35:54First of all, it has a hypothesis of like, all fruits are tasty. And if you say like, apple is tasty, then you can say, okay, this statement is consistent with this, because apple is a fruit. But this saying is like married like apple, then it's not. You can't really derive. It's mutual. It's not related to this. So it's very useful in the context of LOM. So, for example, you give it some documents about your company policies, and the answer is like, can it be derived from this document? Is this consistent, or is this contradictory, or is this something you cannot sufficiently derive? from it.

Read the full transcript

36:34So I think it's just having some understanding, basic knowledge of ML theory fundamental can be very, very helpful. I think the opposite question is also kind of interesting now in that code generation has made coding so accessible. I think it's tempting to ask, how far can I get without really understanding engineering? But I also find that to really use code gen tools and get over some of the humps and barriers that you're presented when you use those tools you have to understand at least a little bit about abstractions and like common refactoring patterns and things like that or else you end up in a corner um so it's it's like it's i guess it's both the ai and the engineering and they're kind of inseparable in the sense that uh you need both of them if you're going to build anything more than kind of a toy.

37:33Do you agree with that? Yeah. And a lot of it it also needs to understand the problem you're trying to solve in your users. Yeah. As you brought up the question about coding, I think a really interesting question that come up is it's like, so I guess quite good at coding. And people wonder is there still any point in learning how to code? Or will there still be software engineers in the future? You probably saw that question right i'm not sure if you like thought um like thought about the answers like well what the answer would be there um but i was watching this uh a webinar from stanford from the chair of curriculum for the cs department and like um andrew at stanford about like the future of cs education and something they said i really really like is that like the job of software engineering as a Java software engineer is not just writing code.

38:31I said, I can write a code itself. It's actually pretty manual and like not that exciting. But so there's a task of like understanding the problems, coming up with solutions. Like it's a lot more exciting. So I think it's like just because AI can like write code well, doesn't mean that it can solve problems well. And I think it really depends on us when we decide like what kind of things that we want to do. Do you want to be a software engineer who just knows how to write code well or someone who can build beautiful or wonderful software solutions to address real-world problems? And it might need to start with us understanding the problems that we want to solve.

39:11So you were talking about mistakes that folks are making. Are there others on your list? okay so one things I see is that like not spending enough time crafting the guideline for what you want the model to do so a lot of time we interact with AI but meaning a spec going in or what you expect your outputs to be or all of the above yeah all of the above like the instructions so we want AI to act in a certain way like to exhibit certain behaviors and And we need to tell AI what we want to do. And so for example, I can give it like examples of like, hey, like you wanted to respond like this is a good response.

39:54This is a bad response. So like do more with a good response and do less with bad responses, right? And a lot of times it's really hard to define what a good response would be like. So a very interesting case study from LinkedIn. So they were building a chatbot to help candidates prepare for a job, right? and the candidate can ask the question of like, am I a good fit for this job? And they thought I was like, correct, it's not the same as good. So the bot my response was like, you're a terrible fit. So it's correct, but it's not helpful for the candidate, right? So they need to understand more on like, what does the candidate want to hear?

40:40Like, what is it going to need? So can it like, maybe they need pointers on how to improve their profile so they can be a good fit. Or another example from a company I can't name because it's confidential. So they were trying to build a bot to help summarize meetings, transcripts. And at the beginning, they thought, okay, we need to think about how long this summary should be. Should it be three sentences? Should it be ten sentences? They thought of what the users wanted to know, but actually not that. The users actually care. They mostly care about what action items they have, what is timely, urgent.

41:23So a lot of it is just like you need to understand what users want. And put that into good instructions and guidelines, and then AI can help you with that. It makes me think about eval as a challenge also, because if you aren't able to clearly articulate, you know, what is a good response versus what is a bad response, then you're not going to be able to formalize. That's like a prerequisite to formalizing that into a set of evaluation criteria and tooling. And that's just incredibly hard because I'm not showing you like, I think it's like even I sometimes like if I wanted to do something, I'm just like, oh, my God.

42:02and have to explain very well. Like, let's say an example, like you want to grade an essay, right? It's like, hey, this is a good essay. This is a bad essay. It's not that simple. We need to come up with a lot of criteria, like very examples as well. And whenever I think about this as a problem, it makes me think about Andre Karpathy's, you know, software 2.0 thing. Like the reason why ML was interesting was because like we couldn't write the rules to teach, you know, to get from a cat face to a class based on the features of the pixels, right? And the same is true in evaluating text responses.

42:42Those rules are just really hard to write. So we need to use tooling in order to, I guess, that both speaks to the interest in AI as a judge, but also the need to put a lot of thought into how you as a person think about the quality assessment of whatever your output is. Definitely. I do think that maybe another bit of a pitfall is that people don't spend enough time thinking about evaluation as they should. I do think that evaluation is the biggest bottleneck to an adoption nowadays. Are there, you know, things that you've seen people doing that are working well? Is there a silver bullet, you know, for evaluation beyond the kind of thinking hard about the problem that we've been talking about so far?

43:41or is there a you know a path you know a next level of um detail or if you were instructing the the engineers as a planner like how to think about kind of walking through the evaluation process like you know what's step one what's step two step one uh for me personally i so you know what is get tedious so I can see why people don't do it but I just spend a lot of time just like looking at the data so first of all like maybe for me like every day I would look at like certain numbers of like samples so like don't be afraid to involve human in the loop and like just looking at the data I think there's some interesting tweet I think Greg Brockman was saying it's like manually like looking at data is one of the most like low predicts but like high value activities you can do.

44:40Yeah. Also like try to like systemize things because yes, eyeballing is important. Like just looking at the data, but also like try to come in systematic ways. Like don't just like rely too much on vibe check. You need to be able to like benchmark the progress over time, send me some concrete numbers. Yeah. Being so like understand your tools you're using for evaluate. So first of all, a lot of people nowadays use AI as a judge, but like an LM judge is not like static, right? Like it depends on the model, it can change the prompt. And in some, a lot of teams I've seen is it's like some, some teams build the AI judge, like running the prompt and some other teams use the judge.

45:26And I would say it's like, so a team, like someone comes to me and say, hey, we have this score of like 90 % on faithfulness, for example. and I'm just like, so what is a prompt you use to score such faithfulness? And they were like, I don't know because somebody else built it. And I was like, hmm, like you have too much. I feel like AI is great, but sometimes we have too much trust in it, you know? Like we just need to like understand more on how things work. Related to your earlier comment about diving into tools and getting abstracted away from how things are actually working. Yeah. I'm sure maybe it can change the future as the field evolves and we have more reliable tools.

46:08But I think we're still at the phase where people are still learning and things are still getting into shape. So being a little bit more paranoid is probably a good thing right now. We talked about how you wanted to avoid anything related to ops for the title of the book. but you spent a lot of time thinking about MLOps and, you know, ML systems in earlier times. Like, how are you seeing folks operationalize LLM systems and agentic systems? Are there patterns forming yet? I do like operations. I have a lot of respects for operations people. They are pretty like my favorite people, one of my favorite people to hang out with.

46:54yeah so I think that we can definitely learn a lot from like the operation world like especially like the SRE talks about like how to make things reliable and trust me like reliability is one of those things that we really need to get better with AI so like a lot of those like how they do things how they like come up with metrics so first of all like two metrics that I thought was like very important is like first of all like time to detections like if something fell like how long do you detect it how long until you can detect it and the second is like time to address issue okay i'm pretty um butchering the name and my brain is like we just came back from like asia and like i'm still jet lagged uh but like how long it takes for you to like after you take the issue like how long does it take for you to like fix the issue so i think it's like a lot of AI system today, like AI, like, especially in JNBAI, they fail silently.

47:54Like, the responses can sound like, okay, coherent, but they can be like, like, information hallucinating or like containing some, like, weird things. So, and people might just like, and especially something so scary about, like, factual inconsistency, is that, like, me, for example, me as a user, right, I ask AI a question because I don't know it. So, like, when it gives back to me something that's wrong, I don't know if it's wrong, right? I might even trust it because that's precisely why I asked it in the first place. So I do think it's like having a more consistent way, systematic way to define these different kind of failures, how to detect them.

48:36And I started getting more rigorous about trying to measure and address them. It's extremely important. Are you seeing... a lot of activity around open source models or open weights models, LAMA and the like local execution of models are the folks that you are working with and talking to using a lot of those or primarily using the models provided by the larger companies? I think that definitely changed a lot last year that I see more and more people using open source models um i definitely like um i think like open sourcing is great for the community and we do need to think about like incentives right like say if you have the best model out there in the world like why would you open sources so i think it's like open source make a lot of sense from like users perspective or from the world peace level perspective, but we need to think about the incentive for the people to want to make things open source.

49:49And I think that's something we are not quite that good at. So right now, we do rely on big corporations to open source, but at the same time, it's incredible that teams are doing it, but at the same time, there's no guarantee that they will continue doing it. So I do think it's open source, especially with the introductions of very powerful open source models recently. it does really help the community accelerate development. We see a lot of people more willing to try out use AI because you don't have to like, because now they can deploy this model into their own systems so they can do governance better, all that stuff.

50:29Yeah, so I do think it's a big boon to the community and industry. I hope that we can continue with that in the future. How about synthetic data? Are you seeing a lot of synthetic data applications in the context of agents and alums? Definitely, like, huge. I think, like, if you're interested in synthetic data, I really like this paper. I like the Lama 3 paper. I think it's an amazing paper. You should just read it. Like, if you have, I'm sure a lot of people already read it. But, like, if you haven't, like, do check it out. It has some really good sections on data synthesis. So, like, they were able to use AI to, like, synthesize, like, I believe it's 3.7 million samples for instruction file tuning.

51:14And one thing I really like about this paper is that not only it explains how they synthesize data, but also how to verify data. Because now there's a challenge, hey, we can generate any data, but how do you know that this data is good? How do you know that? Because you know the importance of data quality. If you train a model with bad data, the model is going to be worse. so so it i realized that's like on a data verification so they have a really clever a few very clever techniques and that's what i love about like working uh as uh in ai nowadays because it's really creative like you can cover it like a lot of like it's like i know it's like it's like it's like so clear there's so much creativity in there so first of all they a lot of the example inputting code because they naturally with code you can like have a good way to validate whether the code runs or not and whether it's opposite everything so they have like code generation, they also have like back translations with coding.

52:09So for example, if they start with like a problem, like a description of what the code is going to be like, and then they generate the code, and then they use LM to like generate documentation for generate code. And as I compare the original document with the AI generate documentation, and I can see, okay, if it's good, if it's matching, then the code is probably good. So like a lot of like creativity like that in data verification. yeah I'm very bullish on cinematic data well we're speaking at the beginning of 2025 what are your predictions for the year what are you excited about and where do you think the field is going to be if we were to talk at the end of 2025 or beginning of 2026 well one thing I'm excited about is more dramas for open AI it's like quite exciting

53:03um what um what else um i'm do see i want to see more applications uh from um from the community from the industry i think it's like last year we seen i kept hearing that a lot of like one of the biggest challenge for industry like for enterprise to adopt applications is like security governance right compliance so I hope that we see more like law around AI when we become evolving so we have more clarity so like people can feel confident in like deploying these applications I do see an increased capabilities of AI using tools like maybe what do you call like agent maybe they have more framework around like how to build agents and how to evaluate them I also like expect a lot of like I'm so like a never-ending question.

53:59It's like compute. So I'm very excited to see more progress there. Yeah. Compute progress can be slow. Are you expecting big shifts in the upcoming year? Big shift? Like what kind of shifts? How do you define big? Is that the question? Yeah. So I feel like NVIDIA stocks will likely continue to fluctuate and giving people who own it like a lot of like heart attacks. I'm not jealous. I think it's interesting. I would love to see more of like chips, designs, architectures. I think GPUs have been around for a long time. Could be very exciting. Maybe I see more. Yeah, I think like the computer is interesting.

54:52I think a friend was asking me is like, do you ever think that we will reach the point where compute is no longer the bottleneck? And it's interesting questions because like, will we ever reach the point like when compute is just not like important, you know? Like I feel like if we can produce more compute, but then there will always be more applications. Like we have the whole thing of like audio, speech, images, generation, videos, and they're going to consume a lot of like compute. so I felt like compute would continue to be important for a while that's why everyone's talking about quantum yeah that is extremely interesting I think that's a lot longer out than people are thinking though that's what I said about AI five years ago right yeah I think the word is interesting I do feel very fortunate to be living in this in this time period of time i felt like i was thinking i'm like i was looking back like all the other period of time like is it any period of time would rather live in so i probably know like i mean like as a woman like for the first like you know like before the 1900s pretty sucks to be a woman so i don't know like that's definitely out and then like isn't okay i feel like i'm gonna go into this pretty not not not a worthwhile discussion right here Well, Chip, it's been wonderful catching up and we can take the rest of the time for a discussion offline.

56:24But it was great talking a little bit about the book and what you've been up to recently and hearing some of your thoughts on agents. Yeah, it's been great chatting with you. Thank you so much for having me here. And I'm sorry if I ramble on too much, feel free to cut those. Not at all. But yeah, thank you. Thanks so much, Chip.

From the publisher

Today, we're joined by Chip Huyen, independent researcher and writer to discuss her new book, “AI Engineering.” We dig into the definition of AI engineering, its key differences from traditional machine learning engineering, the common pitfalls encountered in engineering AI systems, and strategies to overcome them. We also explore how Chip defines AI agents, their current limitations and capabilities, and the critical role of effective planning and tool utilization in these systems. Additionally, Chip shares insights on the importance of evaluation in AI systems, highlighting the need for systematic processes, human oversight, and rigorous metrics and benchmarks. Finally, we touch on the impact of open-source models, the potential of synthetic data, and Chip’s predictions for the year ahead.

The complete show notes for this episode can be found at https://twimlai.com/go/715.

More from The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

All 156 episodes
AI Engineering Pitfalls with Chip Huyen - #715The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) · 58 min
Listen in VO