Evals, error analysis, and better prompts: A systematic approach to improving your AI products | Hamel Husain (ML engineer)

13 Oct 2025 · 55 min · 18 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Systematic improvement of AI products using trace logging, error analysis (open coding), evals, and targeted prompt/tool fixes; then scaling fixes via unit tests and LLM-judge evals with human-validated binary outcomes.

Guest background

Hamel Hussain, ML engineer. Works with AI product teams and clients; teaches a course and uses Claude/GitHub-based workflows for writing and business operations.

Key claims

  1. Start with real user data by logging traces (multi-turn chat, tool calls, retrieval).
  2. Use error analysis: open-code notes on a random sample of traces, stop at the most upstream error, then categorize and count.
  3. Don’t rely on vague “helpfulness” dashboards; use task-specific yes/no evals and validate the judge against human labels (“agreement with hand labels”).
  4. Fix prompts/tools based on eval failure modes; fine-tuning is optional and often unnecessary early.

Notable examples

  • Nurture Boss (virtual leasing assistant): user intent unclear (“what’s up to four month rent”) leading to wrong responses (specials vs lease terms).
  • Prioritized issues: transfer/handoff failures, tour scheduling/rescheduling loops, missing follow-ups, incorrect info.
  • Prompt bug: missing “today’s date” caused scheduling for “tomorrow” to fail; deleting two incorrect UUID-related words improved tool calling.
  • Unit test example: ensure tool outputs never leak user UUIDs.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding Quality in AI Products

0:00 to 1:09

Learn about the importance of data in improving AI product quality.

“What are the fundamental concepts folks need to know of getting to higher quality products?”

Product Management in AI

3:07 to 4:26

Explore how product management is evolving with AI technologies.

“have been building products for a very long time.”

Fundamentals of AI Product Quality

4:26 to 5:11

Learn foundational concepts for achieving high-quality AI products.

“So the fundamentals really come down to the most important thing is looking at data.”

Exploring User Interactions with AI

5:11 to 11:35

Examine real user interactions to understand AI response challenges.

“I think one of the most transformational skills I learned as a young baby chicken product manager was being able to write SQL and actually do my own data analysis and exploration.”

Systematic Error Analysis in AI

11:35 to 14:00

Discover systematic approaches to error analysis in AI products.

Systematic Approach to Error Analysis

14:00 to 21:18

Learn about the systematic solution of error analysis in AI products.

“is the SQL query I write to get like the first prompt?”

Implementing Error Analysis for Practical Solutions

22:25 to 28:00

Explore how to implement error analysis for improving AI systems and writing evaluations.

“you have to read till you hit a snag, right?”

Creating Effective Test Cases for AI Outputs

28:00 to 29:00

Learn about the importance of creating test cases and synthetic data for AI systems.

“Yeah, because they can show up by accident.”

Understanding LLM Evaluation Metrics

29:00 to 31:40

Explore how to effectively evaluate LLM outputs and the pitfalls of generic scoring systems.

“Okay, we already covered logging traces.”

Establishing Trust in LLM Assessments

31:40 to 33:50

Discover the significance of binary outcomes and hand labeling data for reliable evaluations.

“And so the way that you make sure you can trust these automated LLM evals is to measure sort of agreement with these hand labels.”
Show all 18 chapters

Improving AI System Prompts and Instructions

33:50 to 36:50

Learn strategies for refining prompts and handling common errors in AI systems.

“notes and critiquing things that you can then like start refining the LLM judge.”

Analyzing Agent Performance and Handoffs

36:50 to 39:25

Understand how to analyze agent performance and improve conversion through error tracking.

“In the ReChat case, we had to do fine-tuning to get the extra mile.”

The Importance of Data-Driven Decisions in AI

39:25 to 41:08

Realize the value of data-driven decisions to enhance AI product quality and user experience.

“It is very interesting, like as a product manager, you can get really far with AI assisted notebooks.”

Leveraging AI Tools for Business Efficiency

41:08 to 42:00

Discover various AI tools that can streamline communication and proposal generation in business.

“So I think this has been super illuminating in terms of helping people like me that are building AI products.”

Leveraging AI for Proposals and Course Creation

42:00 to 45:04

Learn how AI can streamline proposal writing and course material preparation.

“so it's basically like um an example of consulting proposals it's um you know i'm so it's kind of funny, I have a skill level partner of Palantir, is expert at generative AI, blah, blah.”

Integrating AI in Educational Workflows

45:04 to 47:58

Discover how AI can enhance educational workflows and content extraction.

“I can put in the slides all at once and have a lot of examples and I give it to it and it produces this.”

Role of Subject Matter Experts in AI Projects

47:58 to 50:56

Explore the importance of SMEs in product development and user experience.

“Okay, we might have to have you back to go through this thing in detail.”

Optimizing AI for Writing Tasks

50:56 to 53:30

Understand techniques for refining AI outputs in writing to maintain personal style.

“essentially what you're doing when you're annotating and doing this error analysis.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00What are the fundamental concepts folks need to know of getting to higher quality products?

0:04Hamel Husain:The most important thing is looking at data. Looking at data has always been a thing, even before AI. There's just a little bit of a twist on it for AI, but really the same thing applies. When you see a real user input like this, you actually look at what users are prompting your AI with. You realize it's very vague. Absolutely. That's the whole interesting bit. Once you see that people are talking like that, you might actually want to simulate stuff that looks like that. because that's what the real distribution of the data or that's what the real world looks like. I'm sure our listeners expect some like magical system that does this automatically.

0:37And you're like, no, man, just spend three hours of your afternoon, go through, read some of these chats, look at some of them with your human eyes, put one sentence notes on all of them and then run a quick categorization exercise and get to work. And you see this have actual real impact on quality and reducing these errors.

0:55Hamel Husain:Yeah, it has an immense quality. It's so powerful that some of my clients are so happy with just this process that they're like, that's great, Hamill, we're done. And I'm like, no, wait, we can do more.

1:09Welcome back to How I AI. I'm Claire Vo, product leader and AI obsessive here on a mission to help you build better with these new tools. Today, I have such an educational episode for people like me that are building AI products. We have Hamill Hussain, who's going to demystify debugging errors in your AI product, writing good evals, and show us how he runs his entire business using Claude and a GitHub repo. Let's get to it. This episode is brought to you by GoFundMe Giving Funds, the zero-fee DAF. I want to tell you about a new product GoFundMe has launched called Giving Funds, a smarter, easier way to give, especially during tax season, which is basically here.

1:52GoFundMe Giving Funds is the DAF, or Donor Advice Fund, from the world's number one giving platform, trusted by 200 million people. It's basically your own mini foundation, without the lawyers or admin costs. You contribute money or appreciated assets, get the tax deduction right away, potentially reduce capital gains, and then decide later where to donate from 1.4 million nonprofits. There are zero admin or asset fees. And while the money sits there, you can invest and grow it tax-free. So you have more to give later. All from one simple hub with one clean tax receipt. Lock in your deduction now and decide where to give later.

2:34Perfect for tax season. Join the GoFundMe community of 200 million and start saving money on your tax bill, all while helping the causes you care about the most. Start your giving fund today in just minutes at gofundme.com slash howiai. We'll even cover the DAF pay fees if you transfer your existing DAF over. That's gofundme.com slash howiai to start your giving fund. Hamill, I'm really excited for this particular episode because I have been building products for a very long time. And this has been one of a few times in my career where the how and what of products that I'm building are so different than what I've built in the past.

3:26They're technically different. They're different from a user experience perspective. And then they have these non-deterministic models on the back end that I'm somehow, as a product leader, responsible for making output high-quality, consistent, reliable, interesting user experiences. And it's such a challenging problem. And what I love about what you're going to show us today is how to approach that systematically, that quality of product building in an AI world systematically, and how you use different techniques to get AI products, which are new to all of us, from good to great.

4:06Hamel Husain:Yeah, I'm happy to be here. Excited to talk about it. So, you know, this is such a new thing for product managers. I'm curious if you could start with the fundamentals. What are the fundamental concepts or things that you think folks building AI products really need to know about the process of getting to higher quality products? And then I know you're going to show us a couple examples of how to do that. So the fundamentals really come down to the most important thing is looking at data. and I believe from working with many product managers in the past is looking at data has always been a thing like even before AI you know like I'm pretty sure that product managers that can like write a little bit of sequel are okay with spreadsheets looking at numbers looking at metrics you know that feels like it's kind of table stakes for being a good product manager nowadays and so there's just a little bit of a twist on it for AI but really the same thing applies and it's just like, okay, how do you do that for AI?

5:07Hamel Husain:And that's what we teach and that's what I'm gonna show you today. Great, and I cannot agree more. I think one of the most transformational skills I learned as a young baby chicken product manager was being able to write SQL and actually do my own data analysis and exploration. But I think the surface area is so broad now with AI and the data is different. So why don't you show us what we should be looking at when we're building these AI products? Yeah. So let me share my screen a bit. Let me give you some background first. So this is one of my clients. The name of the company is called Nurture Boss.

5:44Hamel Husain:And as you can see, it's an AI assistant for apartment managers or property managers. and really like you know you can kind of get an idea from their website which i'm showing right now you know it's a virtual leasing assistant so you know they help with the whole top of funnel of like helping set up appointments helping prospective residents like find their apartments setting up appointments questions about rent so on and so forth kind of like trying to reduce the toil of property managers still having humans in the loop and so when they came to me they had already prototyped something out you know kind of vibe checking it just like everyone does and put everything together but they wanted to know like okay how do we actually make it work well because the ai fails in weird ways and it doesn't always do the right thing but it feels like okay every time you fix a prompt we're not really sure like maybe we're breaking something else or is it really improving things as a whole we don't really know we're just guessing we're just kind of like looking at it and is getting vibes.

6:54Hamel Husain:And that is a very uncomfortable feeling of trying to scale a product. Okay, so the first thing that I'll jump right into is this idea of traces. So traces are this concept from engineering, but it doesn't have to be scary. It's basically like, and it's very topical for AI because with AI, usually have many different events, are, especially like for a chat bot, you have multi-turn conversations where you're going back and forth with an AI. There might be retrieval of information. They might be calling some tools, external tools, internal tools, so on and so forth. And so you want to log these traces.

7:34Hamel Husain:And there's many different ways to go about it, but just to kind of show you exactly what happened at Nurture Boss, let's go into what that looks like. So this is a platform called Braintrust. There's a lot of them. This is one called Phoenix, which is like the same exact data in here. It doesn't really matter. You can see like they're both the same, right? Like, so what we have here, let me just go into a single trace. So this is what I would call a trace. I can make this bigger so you can see in a full screen. And you can see what an AI interaction looks like in this product. So you have, okay, the system prompt, you are an AI assistant working as a leasing team member at some apartment.

8:21Hamel Husain:These are all fictitious because these have all been scrubbed for PII stuff. You know, your primary role is to respond to text messages. So this is receiving text messages. Okay. And you have a whole host of rules like respond, you know, provide accurate information, answer any question, for residents, do the following, provide this website, for example, if you had to ask for a rental application, provide this on and so forth, all these rules, right? And this is a real user saying, hello, there's what's up to four month rent. I don't even know what that means. I got you. I got you. Let me read it.

8:58Hello. Hello there. What's up two, four month I thought I had it I thought I had it

9:08Hamel Husain:It's unclear but okay, I mean like it's fine, this is real, this is the real world these are real traces and then there's a there's a tool call here get communities information it's calling this tool, this internal tool and the tool call result comes back with this information and this is all hidden from the user The user is not seeing this tool call result. He's like, okay, here's information you can use about the community, blah, blah, blah. He's not even sure this is the right tool call. We'll get to that in a moment. And then the assistant goes, hello, we are currently offering up to... So this is like back to the user.

9:48Hamel Husain:This is what the AI responds to the user with. Hello, we are currently offering up to eight weeks rent-free as a special promotion. Please note, the applicable lease specials and concessions can vary, blah, blah, blah. Okay. So like, is this, and I have a cheat sheet for myself about what is actually right and wrong. Okay. So like the comment here is the user is probably asking about lease terms and stuff like that, not about specials. So like, it's not really clear, like this is the right, this is not like what we want. And this is so realistic, right? Like everyone has experienced AI. Like this is like it's kind of, it's being helpful, but it's not really doing what you want to.

10:32Hamel Husain:And it's actually pretty challenging because it's not really clear what the user wants. You could go in a lot of different directions of this. You know, when I'm testing my own AI, this is such an eye-opening example. Because when I'm testing my own AI, I ask it good questions and I spell correctly and I'm very clear. But when you see a real user input like this, you actually look at what users are prompting your AI with. You realize it's very vague. They say stuff like, what's up? the question, there's no clear question. And so I really do think looking at real user data kind of can get a developer or PM out of their own mind on how they think users are going to interact with the system.

11:13Hamel Husain:Absolutely. It's very critical that you do this. And so now you might not have this data. And I just jumped right into a real example just to set things off. And we can go into all these different rabbit holes or like what if you don't have data and stuff i just want to like ground it and like okay so set the stage like this is kind of one foundation is like you have to have data there's different ways to get it one is you can log it from your real system and you have these things to look at another way is like okay you can have synthetic data where you sort of generate with an llm you can generate questions like this you know hello what's you know it might be hard to generate stuff that looks like that because I don't even know we don't know what it means and probably an LLM won't generate stuff like that but that's the whole interesting bit it's like once you see that people are talking like that you might actually want to simulate stuff that looks like that because if that's what if that's the real distribution of the data or that's what the real world looks like you might want to challenge your LLM or your AI system appropriately okay so let's step back here so you have the system it's doing stuff it's like there's stuff like this happening we can look at another trace if you want just to kind of get an idea and this is you know this is not pre-scripted i didn't memorize what's going on these uh traces we're just looking at them naturally so this is something this is another uh apartment complex meadowbrook apartments same idea so we won't read the whole system prompt again okay so we'll scroll down here let's get to what the user is asking walk in tor so this must be another text message situation and the assistant says our team tries their best to accommodate walk-ins me get you now that's hilarious like i don't know what why is the lm that's surprising like why is it saying me get you to someone who can help maybe he's trying to mimic the uh the uh the user somehow um and then it does like uh yes and then okay great so it seems like this one maybe is okay um let's see what we end up annotating uh yeah we said this one is okay there's there's some metadata down here about our labels which we'll talk about next but yeah you can so you can see like this is a real system there's many different things that can happen here so the question becomes like okay so we talked about this like writing sql and data but like how do you take that same mindset to this like what do you even do with this right you have this like crazy like interactions like how do you analyze this without go without getting stuck because like this seems like um intractable right at first no i i was just thinking i was like what is the SQL query I write to get like the first prompt?

14:06And like, how do you query for give me all the first prompts that include typos, like give me all the first prompts that are ambiguous questions, it just feels almost insurmountable. And then, you know, you showed us two examples, and it's two of probably 1000s and 1000s and 1000s. So going through it manually is probably not super scalable. So I'm curious, what is the systematic kind of solution here? Okay.

14:31Hamel Husain:So the systematic solution is something called error analysis. So error analysis just means it's kind of a counterintuitive process that's extremely effective and it's dumb, but it's accessible to everybody and it works. and it's not something that I made up. It's been around in machine learning for a really long time because actually machine learning has the same problem like before like generative AI. Like we had these stochastic systems that can do like a whole number of things and like how do you actually like analyze that and like figure out like what's going wrong and improve it. So error analysis has two steps.

15:12Hamel Husain:The first step is writing notes and it's called open coding and it's basically like journaling what is wrong. So if we go back to like that other trace that we saw, so let me just go back to it, like the first one, we would step into this trace and we would say, okay, like every observability tool has their own, let's say, different ways to take notes. You know, I already have a note in here. Assistant should have asked follow-up questions about, you know about the question what's up with four month rent because it's unclear user intent and this is writing notes about what is going on okay and you do that for like 100 traces randomly sample 100 traces and you do that you and you stop at the most upstream error you find so you read this and you see what's going on and you're like hmm okay the user intent seems like we didn't do a good job of like clarifying what the hell that they're they need yeah and so i think that's the most upstream problem in this sequence of events so i'm going to go ahead and just write that as a note yeah and and you say focus on the most upstream problem because you presume that if you can get early intent early kind of clarity correctness right the rest of the system is more likely to be correct downstream.

16:33Hamel Husain:Yeah, because it's causal in nature. So as we have the sequence of events, whether it's like user prompts, tool calls, retrieval for rag, whatever it may be, any error at any point along the chain, you know, like will cause downstream problems. And so to simplify our lives for this purposes of error analysis, it's a heuristic, you know, eventually you do want to care about the different errors and different downstream, But when you're starting out, just focus on the upstream error because we're trying to make it tractable. And this is like the way that you're going to get results fast. So basically what you do is you go through and you collect a bunch of notes.

17:14Hamel Husain:And then what you do is you can take these notes and you can like download them or whatever. And you can categorize those notes. And you can even put these notes into like chat GBT. It's like, hey, here's all my notes. Like, can you bucket these into categories? And you kind of have to go back and forth with it a little bit. Like, hey, these are my notes. These are the categories. I think like you're missing a category, whatever. Now with Nurture Boss, what we ended up doing is we actually made one of the things that we highly recommend a lot of people think about is to make your own custom annotation tool.

17:53Hamel Husain:Like there's, you see this is here in BrainTrust and it's also here in Arise Phoenix. they're very similar you can see this is a very similar looking ui and you have they even called it error analysis here and you can like add your notes like you know whatever and you can save those notes and same thing if you're going to be looking at a lot of data you don't want to slow yourself down and you want to be able to have like very human readable sort of you know output and sometimes like this markdown stuff is like not that readable and you want to make sure that okay like it makes sense to you and you can fly through it as fast as possible so um you know it's really easy to vibe code this stuff um because ultimately what you're doing is like showing data so when the in the nurture boss situation so as you might have gathered like they have multiple channels that customers can contact them on they have like text message which would be which we saw they have email they have a chat bot on the website so on and so forth so they just wanted something they could like navigate faster just like vibe coded essentially i mean they have the person we were developers but you know we're using ai in our process and do this very fast is okay like what channel is the trace from and then like some other filters about like hey did we already annotate this or not and then just kind of have some statistics at the top you know this is like what the annotation like looks like it's kind of very similar but just like dialed into what we wanted and like you know we just took notes and then what for nurture boss what we did is okay we had an automated process that would summarize like categorize those notes into like what are the biggest issues and then we would just something very simple like counting counting is always powerful as you know as a product manager you can go into a system the c clear you experience like writing sql queries like you know how powerful counting is counting remains powerful and so you can uh count these issues right so so like okay for nurture boss i don't know if you can see my screen if it's too small i'm trying to zoom in more yeah yeah that's great is okay what are the most what are the biggest issues after doing that error analysis exercise which only took you know a few hours yeah it's like okay um we're having a lot of transfer and handoff issues we're trying to transfer the you the customer to a human we're having a lot of tour scheduling issues so like they're trying to schedule a tour but like a rescheduled tours in this case we found that like someone's asking to reschedule there is no rescheduled tour but like the ai doesn't know that it just keeps scheduling more tours which is bad um you know uh follow-up so you know ai not following up when the user has a question you know sometimes incorrect information provided okay so like you see like these are kind of the count and now we have now we're not lost now we know what we should be working on we know okay you know what we should fix this like transfer handoff issue in this tour scheduling issue.

21:08Hamel Husain:We have confidence. Like, you know what? Like, we're not paralyzed anymore. We know, okay, this is what we need to fix it on our AI. This episode is brought to you by Persona, the B2B identity platform helping product, fraud, and trust and safety teams protect what they're building in an AI-first world. In 2024, bot traffic officially surpassed human activity online. And with AI agents projected to drive nearly 90 % of all traffic by the end of the decade, it's clear that most of the internet won't be human for much longer. That's why trust and safety matters more than ever. Whether you're building a next-gen AI product or launching a new digital platform, Persona helps ensure it's real humans, not bots or bad actors, accessing your tools.

21:55With Persona's building blocks, you can verify users, fight fraud, and meet compliance requirements, all through identity flows tailored to your product and risk needs. You may have already seen Persona in action if you verified your LinkedIn profile or signed up for an Etsy account. It powers identity for the Internet's most trusted platforms, and now it can power yours, too. Visit withpersona.com slash howiai to learn more. I love this. Just to recap, so you're taking these traces of these real conversations and, you know, you don't even have to read all of it. you have to read till you hit a snag, right?

22:34To hit an obvious sort of like incorrect or high friction part of the experience. You have Vibe coded an app that makes it really easy for the team generally to go in, annotate these, rate them sort of like good quality, bad quality, automatically categorize them, count them. And then you have a prioritized list and you're like, here are the problems that I need to go solve. And what I love about this is, You know, I'm sure our listeners expect some like magical system that does this automatically. And you're like, no, man, just spend three hours of your afternoon, go through, read some of these chats, look at some of them with your human eyes, put one sentence notes on all of them, and then run a quick categorization exercise and get to work.

23:20And you see this have actual real impact on quality and reducing these errors.

23:26Hamel Husain:Yeah, it has an immense quality. it's so powerful that some of my clients are so happy with just this process that they're like that's great hamel we're done and i'm like no wait like we can do more um you know you've paid for more like you know whatever they know this is so great like i just feel like i i know what to do and so they find so much value in this like process that and it is like very important this is something that no one talks about like people when you talk about evals like well how do you write an eval? What eval do you do? What tool should you use? Before you get into all that stuff, you need to have some grounding in like what eval you should even write because there's infinite eval.

24:11Hamel Husain:So like in this case, we would write, we wrote an eval about tour scheduling issues and we wrote an eval about transfer handoff issues. And we felt really good about that because we knew that like that is a real problem. And we knew how to write the eval because like we saw that error. And we knew how to find data to test that eval because again, we already tagged it and we saw that error, which is exactly the way you want to do it. Yeah. And what I also like about this is it does take the burden off your users. I mean, so many people try to collect this data by like putting a little thumbs up and thumbs down or little comments.

24:43Like I even have that on parts of my product. And yes, it is useful, but it only gives you a sliver of the kind of self-identified errors in the app. And users are highly tolerant of systems. And so sometimes those errors just don't get escalated by user, they'll either abandon or they'll just work through too many steps to get to the outcome that they want, they'll have a quality experience. And so I'm just taking the burden on yourself and saying you're responsible for looking at the data, you can create simple ways to categorize it. And then you have a prioritized list. Now, if your client is willing to go the next step and do something about this and write evals and fix prompts.

25:26What are your kind of next steps here? What's another example of where we're from here?

25:31Hamel Husain:I just want to talk about this for a minute. Like, okay, so this particular technique is so powerful and not that many people know about it. You know, so I actually recently did a training with open ai showing the people at open ai like you know how this works for domain specific evals um if you want to learn more about like this we had jacob the founder of nurture boss like walk through like this whole process in like two minutes so you can find it on this on this page if you like um okay so to get to your question like what do you do now okay so you have uh like you know you've done your error analysis you have like prioritized these things so like now what do you do so now you get into uh writing the evals so now you have to decide like what kind of evals do you want there's different kinds of evals so there's reference based evals which is like you know what the right answer is and maybe you can write some code you don't need like an lm to do the eval for you.

26:35Hamel Husain:Or if it's more subjective in nature, then, you know, maybe like this transfer handoff issue, maybe it's more subjective in nature, then you need an LLM judge. And so what you can do is you can start to write those evals. And so I have this blog post here about evals in general. So there's this diagram. It's really hard to put this whole thing into a diagram, honestly, but because, you know, it can be, it's kind of, it's not, it's a non-linear process. But really what you want to do is, okay, we already covered like logging traces and there's two different kinds of, but there's different kinds of evaluators or evaluations.

27:21Hamel Husain:There's like kind of like unit tests, which is like, well, I would say like code-based evals. And then there's like models. So like LLMs, you know, code-based evals. So like, you know, So for example, what kinds of things that would be good for code-based evals? Like, okay, if you have like user IDs showing up in the response or something like that, okay, you can test for that in code. I have to say you're saving my life here because I was thinking, what is one of these unit tests I need to write? And that is exactly one of them, which is my tool calls need UU IDs and users definitely do not. So that's a great example of one for anybody that's writing a chatbot that does a lot of kind of tool calling.

28:02Hamel Husain:Yeah, because they can show up by accident. Like you might have the UID in the system prompt, and it inadvertently shows up in output for some reason or another, and you don't want that. Okay, you want to write these tests. No matter what kinds of tests you write, you want to create test cases, and sometimes you can gather those from your traces. Sometimes you might want to generate synthetic data. and so um you know this is like a prompt for a different real estate agent assistant called reach out which is for residential real estate um and this is kind of like a simplified version of your prompt right 50 different instructions that a real estate agent can give to their assistant it creates contacts on their crm contact details can include name phone email whatever and basically you know it can generate synthetic inputs to a system that then you can then log traces from.

28:59Hamel Husain:I'm going to jump around a little bit, so we'll kind of come back to that. Okay, we already covered logging traces. This is another custom log annotation thing, yet again, because we really emphasize this, that it's really important to remove all friction doing this, so I won't linger on this too much. And basically, you know one kind of thing you want to do is like okay if you're using LLM as a judge or anything else what you want to do is so one thing that's usually skipped when we talk about LLM as a judge is like people just using LLM as a judge off the shelf like they're like writing a prompt they're saying okay judge it and then reporting that let me actually go to a different blog post that is a little bit better for LM judge, which is this one.

29:54Hamel Husain:Okay, so LM as a judge. So you often see sometimes in LM eval land, like a dashboard that looks like this. Helpfulness, truthfulness, conciseness score, tone, whatever. What the hell does that mean? Does anyone know what that means? Nobody knows. No one understands concretely. Like if the helpfulness score is 4.2 and it goes to 4.7, like, do you really know, like, what's wrong, what changes? No. And so there's a lot of guidance in how to create an LLM as a judge. It's probably too much for this podcast to, like, tell you all of the things. And this blog post is quite long, like, enumerating, like, how to do it correctly.

30:39Hamel Husain:But the main things that you need to keep in mind is, like, one, you need to have binary outputs like is it good or bad for a specific problem so for like you know the handoff problem for nurture boss like okay was there a problem or not and you want specific evaluators for specific problems number two is like you want to you need to hand label some data which you already kind of do an error analysis and you want to compare the judge to the hand label data so that you can trust the judge. The last thing you want to do is like throw up a judge on the dashboard like this and then like people don't know if they can trust it.

Read the full transcript

31:18Hamel Husain:And the worst thing you do as a product manager is like start showing people evals and then at some point the people's perception of the product or their experience of the product doesn't doesn't match the eval. So like hey like it's broken but the evals are showing that it's good And that's the moment people lose trust in you. And then it's going to be really hard to regain that trust. And so the way that you make sure you can trust these automated LLM evals is to measure sort of agreement with these hand labels. Yep. So what I'm hearing from you in terms of LLM as a judge is these general buckets with arbitrary ratings against them, not useful and will often work against you.

32:08You want to write specific binary outcome evals for specific tasks. So you want a set of evals that are like, does this get scheduled correctly? Yes or no. And so you're making a list of evals that the LLM as a judge is evaluating that gives you a pass, fail or yes, no, true, false, binary outcome. Very simple. And then you're doing the additional layer of work of validating that the eval itself is valid by actually looking at that outcome and saying, do I actually agree with this LLM as a judge evaluation of the quality of this output? And those steps together are going to give you a much more comprehensive view of how your product's performing.

32:54And then that second layer of human evaluation, it's going to give you more confidence that either your LLMS judge is good and is evaluating your outputs correctly, or you actually need to tune that judge itself to get to higher quality evaluations. Is that kind of a summary of what you did as well?

33:14Hamel Husain:And the thing that's really important is like it's really difficult to write any LM judge prompt if you don't do this because the research shows and there's some research that my co-instructor for the course that I'm teaching. There's a paper called Who Validates the Validators? and the research shows that people are really bad at writing specifications or requirements until they need to react to what an LLM is doing to clarify and help them externalize what they're what they want and it's like only going through this process of sort of okay writing detailed notes and critiquing things that you can then like start refining the LLM judge.

33:58Great. And so we've covered sort of traces and errors, annotation. You have kind of how to build unit tests that are automated tests. Of course, you're looking at it manually. You're doing LLMS judge the correct way. Now tell me, I've identified all these problems. I have these evals that give me data. How do I write a good prompt? Like, are there some techniques or, Or, you know, what do I do? Are there things that you found consistently in the next step of improving your system instructions, improving your tools, where you actually have to go solve these problems are effective? Yeah.

34:39Hamel Husain:So when you get to, like, the errors that you have, so, like, you know, you're going to use these evals and you're going to deploy it at scale, okay? It's like you're not looking at all your data. You're looking at a sample of data. and you're going to score your LM as a judge against like a sample of label data. And you're going to deploy that at scale. And you're going to like look at where are there errors. And it's pretty like, you know, you have to make a judgment call on like, how do you improve your system based on the errors you're finding? Like, is it a retrieval problem? Is it a prompting issue?

35:19Hamel Husain:Is it, should you be putting more examples in the prompt? and you know this is not really a silver bullet there i would say um you know retrieval is its own sort of beast it tends to like retrieval tends to be the achilles heel of a lot of ai products um you know where things tend to go wrong but sometimes yeah it's just like especially in the beginning you're going to find a lot of low-hanging fruits like for example in nurture boss the system prompt didn't contain today's date so when the person said hey can you do a schedule for tomorrow ai had no idea what like we don't know what tomorrow is but didn't didn't tell the user that right we just guessed so like you know that's really obvious so there'll be like obvious things you can fix and then there's like lesser obvious things you can fix you could try like prompt engineering so there's a spectrum of like okay prompt engineering all the way to like fine-tuning.

36:18Hamel Husain:Most people shouldn't get into fine-tuning. I will say that if you do all this eval stuff, fine-tuning is basically free because you have all this infrastructure set up to do all these measurements and curate data, like high signal data that is difficult. And that difficult data, those difficult examples where your AI is not getting right, That's exactly the stuff you want to fine-tune on. That's the very high-value stuff for fine-tuning. Fine-tuning is not so hard. In the ReChat case, we had to do fine-tuning to get the extra mile. But in most cases, it's prompt engineering. There's no magic prompt engineering tricks.

37:00Hamel Husain:I would say there's a lot of experimentation that you should engage in. One of the things that I found so interesting as an AI builder that comes from a software engineering background is now I have a natural language surface for bugs in terms of my system instructions and prompts. And I had this experience recently on ChatPRD where we were really having a hard time with tool calling. Like one of our tools just was intermittently not being called no matter what the user would say. And it was really hard to pin down. And we have this monster system prompt and I went through and there was like two words in the prompt that were just incorrect.

37:37They were incorrect. it was about UUIDs, but it was like incorrect. And as soon as I deleted those two words, which had just been, you know, typed in by somebody and pushed in the repo, our quality of that tool calling shot right up. And so I just have to, you know, we have to, as product people, as engineers, start thinking of the full surface area of our product. And it's not the construction of the agent or the chatbot itself. It really goes down into what words are going in and out of your system. And it's a complicated surface area to debug and keep track of because it's unstructured, but it's super high impact in my experience.

38:14Hamel Husain:Yeah, definitely. You know, when it comes to tool calls, actually, let me show you one thing that always comes up is people wonder, like, how do you evaluate agents? Because like, you know, there's so many different handoffs. Like, how do you actually like, do it in real life? so let me see if I can share that okay so I'm sharing like um the book that we give students in our class um but let me go to the table of contents so there's all these different areas we'll kind of skim towards the agent part of it so um there's like analytical tools you can use for everything you know for agents you can build these transition matrices so going from one step to the other, where are the errors located?

39:05Hamel Husain:In like what agent handoffs? Or what steps are being handed off to what other steps? So like in this case, okay, we have this like generate SQL to execution SQL. That's where a lot of this like errors are happening. And then you can like, then you can narrow it down. So as you get more advanced into evals, it's a very deep subject. There's a lot of analytical tools you can use to kind of go about things. It is very interesting, like as a product manager, you can get really far with AI assisted notebooks. Yeah, what I was going to say about this from a product manager perspective is this is really put from the frame of errors and evals, but even just analytics for agentic systems, figuring out what your users are trying to do.

39:54I haven't thought of this idea of actually mapping out the different conversation to tool or tool to tool handoffs. And even if all of this was working effectively, a product manager's ability to see the data of its agent's behavior from a tool to tool handoff perspective and really identify like where are users trying to get value out of the system also can do things like drive roadmap ideas, right? If you're seeing, okay, people are just writing SQL, executing SQL, like we need to dig into what other things around that could we build for users that are interesting. So I like it from the error perspective.

40:30I also like it just from the product discovery perspective.

40:34Hamel Husain:Yeah, definitely. That's very true. Yeah, I like that perspective. Okay, so you've shown us how to... The other thing that I like that you've shown us is that there's no way to do this than just do it. But people want these tricks. They want some hack. They want some off-the-shelf solution. And you're saying, honestly, look at the data. Build yourself a solution if you have to. Validate it yourself. Do the hard work. And if you do the hard work, you can actually create these leaps in product quality and experience. But right now, you just got to look at the data and make some decisions and make things better.

41:12So I think this has been super illuminating in terms of helping people like me that are building AI products. make them higher quality let's spend just a couple minutes on a totally different topic which you are running this business you're running a course you are clearly an expert in ai what tools are in your stack for kind of running your day-to-day life or at least your business life yeah so i do

41:35Hamel Husain:a lot of writing and i do a lot of communication with clients and you know i also want to reduce my own toil and so um let me share my screen again yeah it's probably easiest to show you claude project so i have all these claude projects um so okay i have like one for copywriting i have a legal assistant i have consulting proposals consulting proposals is pretty interesting so it's basically like um an example of consulting proposals it's um you know i'm so it's kind of funny, I have a skill level partner of Palantir, is expert at generative AI, blah, blah. And, you know, I give it some instructions on the other, like, let's say proposals I have.

42:21Hamel Husain:And, you know, I have like this prompt, you know, whatever, get to the point, writing short sentences, whatever. And basically I have a lot of examples. And basically anytime I have a intake call with a client who wants a proposal, I give this the transcript and then it's made it's basically almost ready it's like just need it takes me about a minute to to kind of edit it and get it going so that's that's proposals you know i have one for the course which is like you know a lot of context about my course which is like the entire book i have an faq that's very like extensive that i've published um there's all the transcripts all the discord messages, office hours, you know, and again, my prompt is like, hey, your job is to help course instructors to create standalone interesting epic cues.

43:13Hamel Husain:These are, this is like a writing prompt that I have everywhere. Do not add filler words. Don't repeat yourself. Get to the point. Yeah, yeah, yeah. It's very, you have to really, you know, and so, okay, like, yeah, it's just, you know, this stuff here um you know so there's like one for the course there's um you know there's one to help me create these things called lightning lessons which is basically like you know this lead magnet um so there's all kinds of stuff like this um i see you and i share a general counsel here oh okay with claude ai oh yeah right exactly yeah there you go um so there's that and also have like my own software that I have.

44:02Hamel Husain:So I have, let me see if I can find it. I mean, I'm not really advertising it, but I have YouTube chapter creation. And I basically have this thing that will create blog posts out of YouTube videos. So let me show you an example. So this one, basically what I do is I take a YouTube video and it becomes an annotated presentation. So you don't have to watch the video. Like you can just, especially if the video has slides, what it'll do is screenshot all the slides and then have a summary under each slide about what was said. So you can consume like a one hour presentation and like, you know, whatever, five minutes.

44:46Hamel Husain:And that's really good because like, you know, I have, I teach a lot and I have a lot of content. And so I distribute notes. So all of that. So like a lot of that stuff, educational stuff is part of my workflow. So and that this is used like this uses Gemini. Essentially what it does is it pulls the transcript. It pulls the video. I can put in the slides all at once and have a lot of examples and I give it to it and it produces this. Yeah, I've heard this in a couple of podcasts that we've done recently that folks really like Gemini for video information. in Jest seems to be the fan favorite for taking basically YouTube videos or other video content and turning it into text or other applications that you can extract from that.

45:30So try the Gemini models for that, folks.

45:33Hamel Husain:Yeah, it's absolutely brilliant. It's amazing. Cool. Okay, so you have cloud projects for every little part of your business. I love the proposal workflow. It's something that we folks that do enterprise sales could probably make some use out of. I'm about to start doing blog posts on all the How I AI podcasts. So maybe I will download your repo and give that a little spin. And then you're using Gemini models to extract out content and share it as templates. And then you have, oh, look at these prompts. We've got a GitHub with prompts. Yeah, so I give GitHub with prompts. This one is private. But just to give you an idea, conceptually, it's basically a monorepo of everything.

46:13Hamel Husain:the reason that is is because I like to have clawed code open hands you name it and basically what I say is because all these things are all interrelated right like a lot of these projects so like you know this is my my blog is in here this is my blog for example this is that that like YouTube thing I just showed you this Hamill project this is like something else that fetches discord this is about copywriting proposals whatever and i just point ai at this repo and you know there's like claude rules in here that says like okay what is this repo about and like where do you find stuff like okay you know this is like if you need to like for writing you should look here um you know so on and so forth so my friend you have buried the lead here because we could have done an entire episode on just this repo.

47:07What this makes me think of is, you know, five years ago, there was this big like note taking second brain, where do you put all your information so you can have access to it forever? And I see this and my little engineering brain goes, obviously it should go in a repo and it should be a combination of data sources, notes, articles things that I've written things that I like and prompts and tools to actually do something with that so you have given me a personal project that I'm going to go work on in the next couple days because I think this is this is how I as somebody who lives with cursor or cloud code as sort of co-pilots for everything I do this is how I would want to organize my data and my prompts to

47:51Hamel Husain:be able to do something with it yeah I don't want to be locked in right like to any one provider and so this is how I do that. Amazing. Okay, we might have to have you back to go through this thing in detail. This has been so great. I have two lightning round questions for you and then I will get you out of here. I know you're a busy guy. My first question is, a lot of what you showed us requires someone, a person to go through with their human eyes, read things and evaluate. And I'm curious, whose role do you think this is? Is this the product manager's role? Is it the engineer's role? Is it the subject matter expert's role?

48:26Who does this?

48:27Hamel Husain:I think the subject matter expert is very central. A lot of times the product manager is the subject matter expert in SME in a lot of organizations. They're kind of the person that everyone looks to for the taste of like, hey, this is what should be happening with the user. So I would say a lot of times it is the product manager that should be doing that annotation. Now, when it gets into the analysis, it's really interesting. it would be good if a product manager, like the more you can do, the better, just like the SQL and the stuff that you know about. At some point, you probably need a data scientist when it gets advanced.

49:07Hamel Husain:But the more you learn, the better, and vice versa. The more data scientists learn more product skills, it's going to be better. It's hard to predict. There's always this tension or this kind of, okay, can we collapse roles? Can we collapse the product role in this data scientist type AI role? I'm not sure. It's yet to be seen. I don't think so. There's a lot of surface area, actually. There's something called AI engineer. There's AI product manager. And there's also still this data scientist aspect. So those three roles are still operating on this problem. And there's definitely a lot of surface area for all of them, especially as you scale?

49:54The one other thing that I would call out or my hope is in addition to sort of like the technical building teams who are sort of proxies in my mind for the subject matter experts. So a lot of times the product manager is a proxy for like the leasing agent in this example. They understand that user. They understand what high quality is. But, you know, I would really love to see folks that are in operational or more functional roles come in and actually contribute to the quality of the products because you know what makes a good user experience. You know what makes a good leasing agent. You know how they should speak and what they should do.

50:28And I think there is an opportunity for folks to lean in and bring that expertise to bear in a way that scales across a company. That if you're willing and brave to do it, I think product teams would welcome in kind of like non-technical colleagues into this process to add some more kind of user empathy and subject matter expertise.

50:48Hamel Husain:Yeah, definitely. Yeah. The more you can bring like the actual required taste in the product sense into the process, the more that, yeah, because that's essentially what you're doing when you're annotating and doing this error analysis. And the error analysis is the foundation for everything. Yep. Okay. And then my final question, ask everybody, I know you're very structured and you'll tell me you'll look at the data and then figure out exactly what to say. But you have to admit sometimes AI is very frustrating and doesn't do what you want it to do. Do you have any back pocket prompting techniques you use?

51:21Do you yell? Are you all caps? What's your strategy?

51:25Hamel Husain:AI has frustrated me the most is writing. Because like writing, I don't want the writing to sound like AI. And it's hard. That's the last thing you want in certain situations for your writing to sound like AI. And not that AI is like wrong. It's just that, yeah, you want to make sure your like flavor is coming across. And so, um, so one thing, one thing is like, okay, I showed you my writing prompt a little bit of it. I can share it with you separately also is like provide lots of examples, but then also take it step by step. So for writing, what I do is have it write an outline and then I have it write the first one or two sections and edit it very carefully.

52:06Hamel Husain:Now, one tip is use something like AI studio that allows you to edit the output of what the LLM is giving you. That's really important because like what that ends up doing is it creates examples for the LLM in kind of right there. Yeah. And so, yeah, you want to edit the output and you know, yeah, something like a notebook or AI studio, There's not too many things that let you edit the output. But once you do that, once you do that hard work of those examples, especially the thing you're trying to write now, then it starts to work really well. Yeah, it was one of the most important things that I built into my AI product was every asset that gets generated has a real-time editor for the user to update.

52:54And then those updates go back into the model. Because I just think if the central value proposition of your product is writing, which mine is, it's one of the hardest stylistic challenges I've seen AI struggle with. It all sounds like slop. Like I can identify AI writing from a mile away. And so, yeah, I found this like incremental optimization, first outline, then draft, then edit, then refine. Process takes a while. There's some latency in the experience, but it ends up netting higher quality. And then just like use it as a draft, edit it, get the system, get the system to be better. So that's really, really great feedback.

53:31Hamel Husain:Is this for chat PRD? Is this for chat PRD? Yep. Very cool. Yeah, you know, I have high standards for writing too, so it was important to me. Well, this was so great. Where can we find you and how can we be helpful? Yeah, haml.dev is my website. You can also find me, Haml Hussain, on Twitter. And yeah, I'm teaching a course on Maven, as you know, about evals that go into all these subjects very deeply. but yeah that's where to find me great yeah and for our listeners that don't know Lenny's list is on Maven including a how I AI section that I think features your course so you can check it out there thank you so much for the time it was super educational very practical I'm going to take these tips right away and go improve my own product have a great day yeah thank you for having me on thanks so much for watching if you enjoyed this show please like and subscribe here on YouTube, or even better, leave us a comment with your thoughts.

54:28You can also find this podcast on Apple Podcasts, Spotify, or your favorite podcast app. Please consider leaving us a rating and review, which will help others find the show. You can see all our episodes and learn more about the show at howiaipod.com. See you next time.

From the publisher

Hamel Husain, an AI consultant and educator, shares his systematic approach to improving AI product quality through error analysis, evaluation frameworks, and prompt engineering. In this episode, he demonstrates how product teams can move beyond “vibe checking” their AI systems to implement data-driven quality improvement processes that identify and fix the most common errors. Using real examples from client work with Nurture Boss (an AI assistant for property managers), Hamel walks through practical techniques that product managers can implement immediately to dramatically improve their AI products.


What you’ll learn:

1. A step-by-step error analysis framework that helps identify and categorize the most common AI failures in your product

2. How to create custom annotation systems that make reviewing AI conversations faster and more insightful

3. Why binary evaluations (pass/fail) are more useful than arbitrary quality scores for measuring AI performance

4. Techniques for validating your LLM judges to ensure they align with human quality expectations

5. A practical approach to prioritizing fixes based on frequency counting rather than intuition

6. Why looking at real user conversations (not just ideal test cases) is critical for understanding AI product failures

7. How to build a comprehensive quality system that spans from manual review to automated evaluation

—

Brought to you by:

GoFundMe Giving Funds—One account. Zero hassle: https://gofundme.com/howiai

Persona—Trusted identity verification for any use case: https://withpersona.com/lp/howiai

—

Where to find Hamel Husain:

Website: https://hamel.dev/

Twitter: https://twitter.com/HamelHusain

Course: https://maven.com/parlance-labs/evals

GitHub: https://github.com/hamelsmu

—

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

—

In this episode, we cover:

(00:00) Introduction to Hamel Husain

(03:05) The fundamentals: why data analysis is critical for AI products

(06:58) Understanding traces and examining real user interactions

(13:35) Error analysis: a systematic approach to finding AI failures

(17:40) Creating custom annotation systems for faster review

(22:23) The impact of this process

(25:15) Different types of evaluations

(29:30) LLM-as-a-Judge

(33:58) Improving prompts and system instructions

(38:15) Analyzing agent workflows

(40:38) Hamel’s personal AI tools and workflows

(48:02) Lighting round and final thoughts

—

Tools referenced:

• Claude: https://claude.ai/

• Braintrust: https://www.braintrust.dev/docs/start

• Phoenix: https://phoenix.arize.com/

• AI Studio: https://aistudio.google.com/

• ChatGPT: https://chat.openai.com/

• Gemini: https://gemini.google.com/

—

Other references:

• Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences: https://dl.acm.org/doi/10.1145/3654777.3676450

• Nurture Boss: https://nurtureboss.io

• Rechat: https://rechat.com/

• Your AI Product Needs Evals: https://hamel.dev/blog/posts/evals/

• A Field Guide to Rapidly Improving AI Products: https://hamel.dev/blog/posts/field-guide/

• Creating a LLM-as-a-Judge That Drives Business Results: https://hamel.dev/blog/posts/llm-judge/

• Lenny’s List on Maven: https://maven.com/lenny

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

More from How I AI

All 103 episodes
Evals, error analysis, and better prompts: A systematic approach to improving your AI productsHow I AI · 55 min
Listen in VO