Building Real-World LLM Products with Fine-Tuning and More with Hamel Husain - #694

23 Jul 2024 · 1 h 20 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

TWIML AI Podcast Episode #694 Summary

Episode Overview Title: Building Real-World LLM Products with Fine-Tuning and More with Hamel Husain Host: Sam Charrington Guest: Hamel Husain, Founder of Parlance Labs Description: The episode delves into the practicalities of developing real-world applications using large language models (LLMs), focusing on the fine-tuning process and the challenges developers face when transitioning from demos to functional applications.

Key Topics Discussed

Introduction to LLM Applications

  • Novel Applications of LLMs: Hamel discusses various applications of LLMs, emphasizing the importance of thoughtful user interfaces that integrate humans into the workflow effectively.
  • User Interaction Design: A case study on ReChat, a real estate CRM, illustrates how a rich user interface can enhance the user experience by allowing users to refine their searches dynamically instead of merely interacting through chatbots.

Fine-Tuning LLMs

  • Understanding Fine-Tuning:
  • Fine-tuning is presented as a straightforward process with efficient tooling available.
  • The conversation shifts to the contexts in which fine-tuning should be applied, emphasizing that it is not always necessary, especially for general-purpose chatbots.
  • Fine-tuning is most useful for narrow, domain-specific tasks.
  • When to Fine-Tune:
  • Suggested use cases include:
  • Converting natural language to domain-specific languages (e.g., text to SQL).
  • Enhancing performance in niche applications with specific requirements.
  • Expectations and Caveats:
  • Developers should manage expectations regarding the performance boost from fine-tuning smaller models versus larger ones.
  • Fine-tuning requires ongoing management and curation of data, which can become resource-intensive.

Tools and Techniques for Fine-Tuning

  • Recommended Tools:
  • Axolotl: A tool that simplifies the fine-tuning process, providing configuration files and community-shared settings.
  • LoRA (Low-Rank Adaptation): A method to fine-tune models with fewer parameters, making it feasible to run on consumer-grade hardware.
  • Inference Frameworks:
  • Discussion on frameworks like NVIDIA Triton and VLLM that facilitate model inference and optimization.

Evaluation Strategies

  • Importance of Evals:
  • Hamel underscores the need for systematic evaluation processes to ensure LLM applications function as intended.
  • Evals should be integrated into the development process, akin to unit testing in software engineering.
  • Key Evaluation Strategies:
  • Data Instrumentation: Collect usage data to identify failure modes and areas for improvement.
  • Assertions and Tests: Write specific assertions based on observed failure patterns to create a more robust model.
  • User-Centric Testing: Test cases should reflect real-world usage scenarios rather than relying solely on generic evaluation tools.
  • Iterative Improvement:
  • Continuous evaluation allows for incremental improvements in LLM applications, enabling teams to refine their products systematically.

Conclusion The episode emphasizes that building effective LLM applications is a multifaceted process that requires thoughtful design, strategic fine-tuning, and rigorous evaluation. Hamel Husain's insights underscore the importance of focusing on user experience while leveraging the latest tools and techniques for optimizing LLMs.

Key Takeaways

  • Thoughtful integration of LLMs into user workflows enhances user experience.
  • Fine-tuning is most effective for narrow, specialized tasks; expectations should be managed accordingly.
  • Systematic evaluation and testing are crucial for developing reliable LLM applications.
  • Tools like Axolotl and methods such as LoRA facilitate easier fine-tuning processes.

For complete show notes, visit [TWIML AI Podcast Show Notes](https://twimlai.com/go/694).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00So evals, I think, okay, so evals is the most important thing. In the consulting work that I do, people call me when their AI doesn't work. It's like, hey, like we've kind of prototyped something. We're excited about it, but it's like not really working. People don't know how to get unstuck. And so that's where this notion of evals comes in. You want a systematic way to test the efficacy of your system. It's really the only way you can build AI. So it's like integral to the building process of AI products.

0:40all right everyone welcome to another episode of the twimble ai podcast i am of course your host sam charrington and today i'm joined by hamel hussein hamel is founder of parlance labs before we get going be sure to take a moment to hit that subscribe button wherever you're listening to today's show hamel welcome back to the pod yeah thanks for having me back it's been a while great chatting it has been a while likewise likewise uh about two three years end of something like that i think something like that yeah yeah um you were deep into ml ops i think we spent a lot of time talking about that we spent a lot of time talking about uh notebooks and literate programming and fast AI tooling and a bunch of interesting topics.

1:31But I wanted to get you back on the show because you've been quite prolific on LLMs and the way folks can take advantage of fine tuning and eval. Lots of interesting stuff I wanted to dig in with you. Excellent. Yeah. Looking forward to that. Were you at GitHub the last time we spoke? Yeah, I think I was at GitHub. GitHub, we're just leaving and you went to work with... Outer Bounds. Outer Bounds for a bit. What have you been up to since? Yeah, I decided to do my own consulting company just because it's such an exciting time to be working with large language models. And it's a really exciting time to see lots of different use cases with this new technology and work with lots of different companies.

2:24And so I've been advising lots of companies and helping them build AI products. Awesome, awesome. Are there any use cases that jump out at you as novel, under the radar, without kind of divulging your client's IP or anything like that? I sometimes think that, yes, LLMs are very exciting, but a lot of the use cases are very similar. And I'm always looking out for kind of novel use cases where it's like surprising that that would actually work. Does anything jump out at you when you hear that? You know, I wouldn't say there's any novel or use cases that surprise me anymore. I think I have gotten over the stage of being surprised.

3:14But I will say that the things that seem to work best are user interfaces that have humans in a loop in a very thoughtful way. So instead of just having, let's say, a chatbot, having some kind of interface that allows folks to interact more richly with the software. So for example, one of my clients is this real estate CRM company called ReChat. And they have a chat interface. however whenever you interact with the chat interface it pulls up different widgets so if you try to search for a listing in the chat interface it'll pull up sort of a like a listing you know you say okay i want to find properties that are you know greater than two million or less than two million dollars with such and such and the zip code or whatever your search is it You'll pull up a little widget with the results that you can refine.

4:15So it actually brings the user interface to the, or it's a very thoughtful sort of interaction between the software and a human rather than this like naive, let's just push the chat interface, you know, to the user. And so like things like that make a lot of sense. Because it allows people to work with AI when they should be working with AI and not work with AI when it doesn't make sense to work with AI. And like, you know, kind of inject the AI in the workflow in a very thoughtful way. And it takes a lot of work to do that. From a design perspective, a lot of people just want chat GPT for their software.

5:03Right. That's kind of the lazy thing to do. The thing that takes a lot more work is like, hey, let me thoughtfully figure out how to integrate it into the user's workflow. So wherever that happens. I appreciate that. Yeah. I really appreciate that take. You know, I see as well vendors trying to rebrand everything as agents. And I don't necessarily want to be forced into a dialog box every time I want to interact with a piece of software. Like we spent a long time refining the user experience for different types of software and it all doesn't look, you know, the same way. And I'd hate to think that we have to throw all that away in order to incorporate the intelligence that we can get from an LLM.

5:50uh what you're what you're describing reminds me of a i saw i've read some blog posts a while ago that coined this as generative ui that was trying to like to some degree dynamically um kind of push ui elements into a chat framework have you seen that kind of take root more broadly among the folks that you're you're working with yeah and what i suspect it's like there's there's many purposes of this kind of ux it's like one is okay obviously to help people do their work a lot faster but also like to educate the user so that you know you want to use the same ui elements throughout your product for example so when someone sees a listing finder they they know when next time they see a listing finder in another place even without ai they're how to use it because they've already seen it before they've kind of been walked like walk through it and people can choose to like do whatever they want like sometimes you don't need ai sometimes you just want to you know look around or whatever so like it's it's an important like onboarding thing um and so there's some deep kind of design thinking that i think needs to be done on a lot of these AI products.

7:15The thing that I wanted to really dig into with you, or at least a really good starting place, is fine-tuning. You recently ran what I think you initially intended to be a course on fine-tuning, and it turned into a virtual conference with a ton of speakers and looked like a lot of really interesting and good information. But that course, I think, grew out of your belief that people think of fine tuning as this thing that's like really complex and hard to do and get right, but it actually doesn't necessarily need to be all that complex. Is that right? No, it doesn't need to be complex. And so the thing that I really like about LLM fine tuning is that it's very easy from like the mechanics of it.

8:09You know, it's actually easy to actually apply the fine tuning itself. So unlike other kinds of deep learning that has come before it, where you kind of have to fiddle a lot with hyper parameters and architectures and stuff like that, when you narrow it down to large language models, there's like really good tooling. And like, it's also fairly easy to fine tune. from a machine learning perspective, like especially if you're doing something like a LoRa, it's kind of hard to screw up the fine tuning. You kind of have to try to screw it up. If you can just - I think that's probably going to be surprising for a lot of folks.

8:45But before we dig into the details there, we probably should take a step back and like answer, you know, the bigger question, which is like, when should you be thinking about fine tuning? Because I think, you know, at some point the narrative was like everyone wanted to fine tune and like the people in the know were like don't do it like you know for sure yeah 99.99 percent unless you're google or facebook or whatever you shouldn't you probably shouldn't be fine tuning and you know you're saying now it's you know it's easy you know it's not as easy like part of it is easy part of it the whole thing yeah okay so part of it is part of it is easy you know so i think it still leaves this space of like does that mean everyone should be doing it now or like who should be doing it and when and also you know what should the expectation be you know when you fine-tune you know versus use other techniques like what are you what is fine tuning really getting you so really i'm wanting to like take a step back and introspect your you know thinking about fine tuning in its role in place so most people should probably shouldn't fine tune especially when you're starting out you should try to use a model off the shelf if you can use open ai use anthropic and start building stuff um you don't want to get nerd sniped by this whole exercise of fine tuning because like you need to like have a reason to fine tune.

10:23And like you need to have some failure mode that you are experiencing that necessitates fine tuning. When you fine tune, okay, there's many different answers for this. So one answer is you want your own model of some kind. And that model is usually an open weights model. Let's say, you know, llama, mistral, whatever. and that open weights model off the shelf doesn't quite do what you want it to do or it's not powerful enough you know to accomplish whatever you want to do on your domain and in the situation where you so like yeah in that situation maybe fine-tuning could be the right fit for you with the caveat of, you know, fine tuning really shines when you have like really narrow use cases.

11:23Like we're trying to have a model do something very specific or the LM do something very specific. If your LM is just a general purpose chat bot that can do anything for anybody, then that's not, you know, fine tuning is going to be a lot more difficult. and it's going to be a lot harder to evaluate that. And so, you know, I would say, okay, like, you know, if you're like a narrow scoped problem that's doing a very specific thing, then fine tuning is more likely to work. The canonical use case that comes up all the time with fine tuning is like text to SQL. Exactly. That's the best. That's one of the best use cases.

12:07In fact, one of my clients is Honeycomb. Honeycomb has an observability platform and they have a query language called HQL, a Honeycomb query language. So it's kind of like, it's not SQL, it's a domain specific language. So they have, you know, a natural language to HQL sort of product. But that is such a narrow task that fine tuning works really well for them. So like there's a lot of dimensions to like when fine tuning would work for you. One is like maybe you want some kind of data privacy. You want your own private model and you want to make that model a little bit better. And if you have like that narrow scope use case, okay, that might be like some sweet spot.

12:50Another reason why you might think about fine tuning also has to do with open weights models is if you want to try to have a small model, like a very, very small, very cost effective model that you can deploy in your infrastructure in a flexible way. If it's like, maybe even if it's small enough, you can try to deploy it, you know, on mobile devices and things like that. So that's another reason where you might consider fine tuning is like, especially with those like smaller models, you might want to try to specialize them to get them to do really well on that task to kind of, you know, raise their performance level to try to approach the ones of the bigger models that are not fine-tuned on your domain.

13:42But there's a trade-off, like fine-tuning isn't free. Even though fine-tuning is easy from the perspective of like, there's a lot of tools out there. I mean, you can even, you know, fine-tune OpenAI with their API, which is a really straightforward process to go through. Once you start fine-tuning, now you have to manage that whole process. There's a whole process that you have to think about of continuously training that model, evaluating that model for the purposes of fine-tuning, curating data, you know, debugging things. And once you kind of commit to fine-tuning also, like, you sort of start to specialize your model a bit on that domain data.

14:30So like you have to constantly sort of manage that because as your product is evolving and, you know, changing over time, you have to co-evolve your fine-tuning with that. So it's not free. It's actually, it can become very expensive in that sense. So that's why the narrow use cases mitigate that. Like if you have a really narrow use case, then like, okay, Especially if it's stable, it's not changing that much like the Honeycomb query example. It's like, okay, Honeycomb query language is not really changing at all. And if they want to tweak it over time, it'll be like some incremental tweaks. It's not like this insanely shifting product.

15:13Whereas like, okay, if you want a general purpose chat bot to do anything on Salesforce, then I don't know. because maybe Salesforce is changing and adding new things constantly. It's a really big service area and all this stuff. So that's kind of some things to think about. Of course, you can try to tackle that problem. You mentioned something in there about powerful enough and even this idea of using a smaller model that without fine-tuning wouldn't have the capability that you need. what should expectations there be like are you you know does fine-tuning allow you to take a you know a three billion parameter model and get you know gpg4o level performance like how should someone think about what you mean when you say make a smaller model more capable Yeah.

16:15So you can think of it like this, like the honeycomb example, we're going from natural language to honeycomb query. You can almost think of it as like, that's almost trivial for GPT-4 in a way. That's, you know, GPT-4 is like, kind of maybe overkill for that in one sense. That's like one side of the equation. And then another side of the equation is like, hey, it's such a very, it's a very, very, very scoped task. and so you might not need all of this like reasoning capability and so you can raise the bar so in this situation we were able to like fine-tune a 7 billion parameter model to outperform like gpt 3.5 we didn't benchmark against gpt4 but it's pretty good and so that's a that's a big win if now, for example, they can offer that to their enterprise clients and they don't have to go through this complicated data privacy situation and SOC 2 compliance and all this stuff where they have to shift data around and all that.

17:23So you can approach the quality of bigger models, but you have to keep in mind of this narrow tasks thing. And so a lot of people don't realize, like the narrow task thing is a big sort of elephant in the room because people are not good at scoping narrow tasks. Right, right. It's not like, you know, you have this list of potential use cases for fine tuning that is narrow domain or you want a smaller model or you want privacy. narrow domain is kind of a precondition that you need to meet in order to like if you don't have a narrow domain then you're probably going to have a problem you know even if you just want a smaller model like is that so so there's a spectrum yeah there's a spectrum so i'm i'm talking about like most people when i say this i'm like generalizing and saying like hey like you should have a narrow domain um you know that's a sweet spot like when you have a narrow domain and you want you want data privacy um and you want a smaller model and you want the freedom to like you know use it wherever then you have like lots of reasons to do fine-tuning you have lots of reasons and like one reason may not be defensible enough depending on what happens like if it's just data privacy you're like oh like i don't trust open ai um you know it might become a day where you do trust open ai so like you know um now there's a spectrum between like fine tuning and continued pre-training you know so yeah i'm right on that yeah i think they're used interchangeably quite a lot kind of yeah so the idea is like they're really like okay pre-training and fine tuning are really the same procedure like you're you're content you're continue the training process with more data the idea of like continued pre-training is in addition to like domain specific data you like mix in this more generalized data like from the web and so on and so forth the idea being you You don't want the model to lose its like general capabilities.

19:47You don't want it to over-specialize in your domain. Okay. You want it to like retain those skills. Don't want to get rusty just, you know, on those things. Got it. Okay. And so there's some like really competitive fine tunes out there that people have done exactly this. and organizations like news research have put out a lot of fine tunes like this where they've trained on lots of different kinds of heterogeneous data sets from different domains like at the same time like they mixed it together and they've curated the data sets they clean them up to to like raise the performance generally of these models on a bat against a battery of different benchmarks.

20:34So that is a thing. And you may, you maybe want to do that, but, you know, from a commercial perspective, I would say for the most part, you know, you want to fine tune on a narrow task. I think it's maybe important to highlight what you just said from the lens of, you know, let's say you've, you've checked all the boxes in terms of, you know, your use cases and the applicability for using fine tuning. I think the conventional thinking is you use an off the shelf model that someone provides you or you fine tune it yourself, but you, you also kind of highlighted the idea that there are other people that are out there publishing fine tune models.

21:21And so if you can find one of those that meet your needs, you might not have to do the fine tuning yourself. yeah so i would say like one try to not do work yourself wherever possible i think that is like a good principle in life it's like one use open ai if you can use anthropic if you can um and go as far as you can with prompting you know like don't fine-tune and try to do things like including examples in the prompt so that's referred to as few shot prompting it's also like a little bit more advanced techniques like dynamically including examples in the prompt that are related to the user input so it's kind of like if you think about rag for examples you know things like that like try to push that as far as you can before going to fine-tuning and then also even if you do need to go with an open weights model try to find an open weights model that is already fine-tuned on that domain and then you might want to just start from if you do want to fine-tune maybe you start with those models like the one that is as close to your domain as possible you just maybe get a bit of a head start instead of just going from just starting from the blank slate or starting from like the base model and are you just looking on hugging face for those models or is there some other repository that people yeah looking at hugging face and like roughly bookmarking things on Twitter and hoping that I remember, you know, um, there's a, there's a leaderboard.

22:52There's some leaderboards. You have to take all these leaderboards to a grain of salt. Um, cause benchmarks are kind of noisy and people, you know, there's a lot of data leakage from the benchmarks into the training. Um, so you have to kind of ask around and dig into it to see like what you can trust. You kind of have to use it yourself also. Um, but yeah, like starting from there and i'll say like in a shrinking number of cases it is the case that you know as models get much better um you know like the business case for fine tuning keeps like shrinking more and more but it's still there like the you know uh it's still there and you know you can it doesn't need to be complicated so in the easiest way to fine tune is like use open AI.

23:45You can fine tune open AI models. Um, and you're essentially just giving it a list of examples there. Yeah. You just give it a list of examples and you can start there. Um, it's a, it's a good way to try it at least like get a sense and, you know, you can kind of iterate from there and see kind of what works for you. Um, but it is really fun and beyond fun. It is very valuable to a business.

24:15It's a slam dunk. When you have a small model, a narrow use case, and that small model really knocks it out of the park, and you can just hand it to a company, it's like, this is your model, and they can just deploy it. That can often unlock lots of value. So it's kind of like a spectrum. it's not like a good let's say flow chart i have but it's kind of like as many you know when you start to stack up all of these reasons yeah the more you have then the more it kind of tilts in favor of fine-tuning but yeah the data privacy one i suspect i mean you know in the early days of cloud people were very suspect of putting their data on the cloud and now no one really bats an eye on like you know uh about the cloud i suspect the same thing might happen with large language models um so i don't know if the data privacy concern is going to you know endure but certainly the other ones might um and so yeah we'll see

25:30so you figured out your kind of business case for doing fine tuning and now Hamill's here to tell you it's not as hard as you were told. What does that process look like? What is the first thing that you need to do when you are committed to going down that path and you're past the doing it via an API level? Yeah. So there's some really good tools out there for the help you fine-tune. My favorite tool in sort of open source land is Axolotl. And I really like Axolotl. Yeah, it's called Axolotl. And it's made by this guy, Wing, Wing Lian. Pretty, pretty cool guy. He, you know, so basically what it is, it's like these config files.

26:25And I hate config. I normally really, really hate config-driven development. Like I hate anything that is like you fill out a config file and then that's how you program. Like there's nothing I hate more, except for this. Like I hate everything. Except for Axolotl. Yeah. Like I hate Kubernetes. I hate GitHub Actions, like all that stuff. I just want to throw it out the window. But like, it's a very useful feature for kind of this like fine tuning large language models because these config files, like they kind of like traded in the community. Like, oh, how did you train whatever like Mistral 7B on this data set to get that outcome.

27:06They're like, oh, let me send you my config file. That's all you need to know. And what's the scope of the config file? Yeah, so the config file has what the data set is that you're going to be fine-tuning over. A lot of times you might be referencing a data set that's on Hugging Face, but it can be a local data set. It'll tell you what the prompt template is. so that's like a really huge thing like the prompt template is like a it's like the thing you have to like really worry about that in terms of like string formatting how you format it is it is it compatible with the base model if your base model has already been you know fine-tuned with a template you want to kind of preserve that as you keep tuning and just to take a step back well a couple of interjections.

28:01The first is if you can't find Axolotl, we will put a link into the show notes page so that you can just click it and find Axolotl. But to take a step back, like in terms of the prompt template and fine tuning, you're fine tuning, you know, on your data. So presumably you've got some, you know, data set of examples that you want that are representative of where you want to improve the model. And, you know, are those typically in, and this is kind of getting to the prompt template, are they typically in like a question and answer format or, you know, is it better to keep them more neutral so that you can then impose a prompt template?

28:48Like what should the dataset look like? So the dataset needs to look exactly like what is going to happen in production. so if a user is asking questions then it needs to look like questions and answers if it's going to be multi-turn conversations then it needs to be multi-turn conversations if it's going to be text to SQL it needs to be just text to SQL and it needs to match exactly what's going to happen at inference time so whatever the prompt template is if it's like you know if you have some like bespoke prompt templating syntax like angle bracket sam forward slash at the beginning of every message then you have to make sure that that is the exact same kind of templating you use at inference time um and that i see is like a really big foot gun a lot of people that is the thing that is the nightmare of of everyone doing this meaning they think they've collected you know this very valuable data set is x but it's really why like why doesn't it work oh guys i left out a space here i left out this colon here in this like place you know because like the model was trained with one i was like oh like i i'm pretty sure the model was trained with like prompt template x but then when i go to inference it i'm like using a different one and it's like well why isn't it working and so there's many such cases of of like yeah i'm missing a space or whatever And sometimes we can really throw it off.

30:23Is it abstracted from the way you've collected the examples? Is the idea that it's at a higher level than the examples themselves? Yeah, that's a good question. So the prompt template is just like, so it's all a string at the end of the day that you feed into. It's all a string. The string gets tokenized. and so the question is like how is the string constructed before it gets tokenized and so when the user asks a question you know they're not saying like they don't have like you know they don't type type in this is the question section and then put their question they're just typing the question that should be the right user interface but then when you send it to the large language model, you need to put some, like, you need to give it some idea of what the hell is going on.

31:19Kind of like a human. I don't want to anthropomorphize things too much, but you want to say, hey, like, you want to give the language model some instructions, like in the QA, in the QA example, you want to say, hey, there's going to be a question, and the question is going to have, is going to be inside question tags. And then I want you to produce an answer inside answer tags. and so you put some guidelines that help the language model understand what the hell it is in this blob of all these strings like what is going on like what is the question what is the answer what is the function call what is the function call result and all that formatting then is your prompt template which is different from your examples are kept separate for you know the convenience of reformatting typically you're from the examples themselves a lot of times it's kept separate yeah then axolotl is kept separate and you can choose which one to apply we like kind of you say okay i want like the chat ml format or i want the alpaca format or whatever and you know those are different whatever templates um which one to use is kind of up to you like different templates have better support or they have like some templates are like very simple.

32:41They don't even have sections for certain things. They might not have do a good job of multi-turn conversations. They don't have an affordance in the templating specification, so you kind of have to make one up, and then you might decide, you know what? I was going to use a different template that actually has thought about multi-turn conversation. Stuff like that.

33:09But, you know, when I'm fine-tuning, so like just to give you some insight, because when I fine-tune, I'm doing it for these very narrow use cases, usually. I don't like to use a model that has any templating at all. Or I just do my own templating, like a little bit of templating or like zero templating. Because I'm usually going from like text directly to something.

33:34So, yeah. I like to start with base models and not instruction tune models. Instruction tune models are already is like, you know, instruction tune just means you're trained a model to, you know, have question and answer, which implies that there's a chat template or there's a prompt template that you have to worry about. So I just don't want to deal with that mess. I just try to use the base model and just come out with my own like simplified template that's specific to the domain.

34:05and is the is that true even if the domain is chat or conversation like is yeah you know the domain is chat or conversation is helpful to have a template i mean it's helpful to like think about it and like yeah sometimes you want to just keep training on the instruction tune model but if you I try not to. So, okay.

34:30For, so chatbots, in a lot of situations, it's, you know, it's usually like not narrow. It's like a correlate, like there's this, you know, when you put a chatbot in front of a user, you have to work really hard to scope it. There's this like inherent user expectation, like, oh, this is like magic AI chat GPT. And so usually what ends up happening is, at least when I'm involved, is I say, okay, let's not do that. Let's kind of pick a more narrow use case to do fine-tuning. I will say, though, I have fine-tuned models like that before, and it's totally fine. But I would just say for the most part, yeah, I don't.

35:18Usually not fine-tuning for chat. Usually fine-tuning for a very specific purpose. of like doing some kind of data transformation or extraction or something like that. Okay. I think we were talking about like what is in the exo-auto config. Oh yeah. So it's just like, it's your data. It's sort of like the prompt template. And it's all these like hyper parameters, like batch size, learning rate, learning rate schedule, all that stuff. And what's useful about it is there's like a whole. That sounds like all the hard stuff from deep learning that you just said. You don't have to worry about it because you can just say, okay, if you want to train whatever, some model, you can just like look at the config that everyone is using.

36:05Because you're not building the config from scratch. You're just getting the config. Okay, got it. Yeah, that's the magic part. You don't have to worry about all this like fiddly crazy stuff. You're just like, okay, look, I'm going to copy and paste this config. Man, I even understand. And the presumption then is that those types of hyperparameters are not, you know, deeply sensitive to, you know, the distribution of my particular data. It's more maybe correlated to the model that you're fine tuning. That's a good question. Because otherwise, you know, I've got the hyperparameter problem. So, okay, like it can be sensitive, but it's not, especially when you're fine-tuning like a LoRa, which is in a lot of cases people are, these like adapters, it's not incredibly sensitive.

Read the full transcript

36:58It's kind of like, if you remember Random Forest, it was like kind of hard to screw it up. You just like point in at your data and it kind of felt like it was working. sure there was some hyper parameters but generally people are like oh yeah this is like very easy like just call dot fit and it's like working and like you know it was less fiddly than deep learning yeah for sure like 100 is the same feeling now with training large language models it's very like it just works like you get something that at least you get something that seems to work. And then you can always make it better, of course, by like, you know, kind of varying these hyper parameters.

37:43But you're not like, oh, it doesn't work at all. Like, I'm just lost. Like, what the hell, you know, kind of, you know, it's not like that. And like starting with a known config is a great starting point. You know, you're on this journey. You've decided that you're going to fine tune. You found Axolotl via this podcast or wherever you found Axolotl, you've got your config file. Are those config files just in the Axolotl? Yeah, they're in the repo. Or do you have to go find something? Yeah, they're in the repo. They're in like an examples folder, yeah. So you've got the config files from the repo.

38:15Do you need to like deeply understand Laura, have done the math, have read the paper? Or is it just a kind of infrastructure that you don't have to think about? No, you don't have to do it. You don't have to think about it. What you do have to worry about is like installing Axolotl, dealing with CUDA errors and all that stuff so um you know because you said you wanted to run it yourself so yeah you said you want all those problems like yeah so you're gonna have that that problem has not gone anywhere um you know there's a docker container so you can use that and that seems to be fairly stable if you can get it to find your gpus you know how it is like i can i can say yeah it's gonna be okay but like i mean that's not i would be lying if i told you that so yeah you can just uh you can use it to you know fine-tune a model on your data so i mean like where's the thing that we're waving our hands over in this whole conversation is like oh yeah like you have a you have your data and like we shouldn't wave our hands over that it's like we've people been waving their hands over that for ever ever since data scientists were coined as a profession.

39:30And the thing that's really interesting is if you already, if you start with not fine tuning, like if you're using OpenAI or whatever, and you're - And you instrument it. You already, yeah, you instrument it a bit. Hopefully, yeah. Then that can become your source of data, which is really fun. So like you, and you have - I guess the question is like, how many examples do you need to effectively fine tune? Yeah. How dependent is that on your use case? Yeah. I mean, so honestly, I don't have the answer to this, uh, of like how, what's the minimum number. It doesn't feel like it's that much. Um, I just kind of get as many as I can within like the time budget I have in the compute budget I have.

40:19So like, you know, you have to make sure, okay, like one, so you have to make sure you're instrumenting your system. So you're capturing all of this large language model invocations. And then you have to have a way of like curating and filtering that data. You don't want to have, you're going to waste a lot of compute if you have like bad garbage data that, you know, used to train your model. But once you curate that, yeah, I can get really good example, like really good results in like a thousand examples and a lot of times what you can do is you can even synthetically generate data based on examples so basically you know so I've done that before I've like expanded data significantly by using a very powerful model like GPT-4 to like expand my data set so yeah there's a lot of tricks that you can do to generate that data and kind of bootstrap yourself into getting data, which is the key difference between classical machine learning is like, oh, I don't have data.

41:26I need to go label everything. So I guess we're stuck. There's like lots of ways to get unstuck. You've mentioned LoRa a couple of times and positioned it as an option. If you're going to fine tune and use LoRa, then it becomes easier. Well, first of all, I'd love for you to spend a few minutes talking about LoRa and what it is, but also what is it an alternative to or what are your other options when you're kind of filling out your config file? Yeah, so it's basically, I think it stands for low-rank adapters. Low-rank adaptation or something like that. Yeah, something like that. It's basically like, so you're not doing a full fine-tune of all the weights in your model.

42:10You're just fine-tuning a very small, like this adapter. and it just makes it a lot easier in many ways. It reduces the amount of GPU VRAM that you need, which makes it a lot more tractable on consumer GPUs or other kinds of GPUs. And it's also, yeah, it tends to, I guess, It's like some much like smoother loss surface. It's like a lot, it's a, you know, it's not that as many parameters. And it's, it's a nice way to get started at least in terms of like training these adapters. And the, actually it can be fun to, there's a lot of tricks you can use with these adapters. So like these inference frameworks allow you to like load the base model and then hot swap adapters.

43:11Okay. So, which is really cool because you can, you don't have to incur the cost of loading an entire model. Just load the adapter for that, you know, forward pass and then, you know, go on your day. And that, incidentally, if I'm not mistaken, is a framework that folks like, you know, Google and Amazon use to allow customers to fine tune their own models. they, you know, are using LoRa and they have, you know, the customer is basically building their own adapter and then they're like hot swapping the adapters out at inference time. I have no idea. It's my understanding. Yeah. I saw some advertisement of like Apple, the what they're doing on device.

44:00It just seemed like from their explanation that they're doing something like that. but you know i don't know exactly what they're doing but it seems like very similar yeah so that's what it is it's just a way to it's like a parameter efficient kind of way to fine-tune you mentioned inference frameworks what are the inference frameworks that you are thinking of and what's their role yeah vlm and nvidia there's like the nvidia triton is nvidia triton which is like the front end the back end they have like tensor rt llm um i like vllm because it's a lot it's like really easy to use and straightforward um it's like kind of a good trade-off if you're not wanting to squeeze tremendous amount of performance out of your models but um have you know it's like a pretty easy to use framework for inference so i would say those are the two that i kind of use is either the nvidia stack the trt and you know the trt lm with a some kind of inference server um or vlm and what is uh like what specifically is it doing is it like um i don't know you remember bento ml like is it analogous to this where you've got like an inference server and distribution and load balancing and all that kind of stuff or is it just like I glossed over it too much so okay sorry so inference servers a lot of time there's like a front end and a back end and what does that mean so like the back end will do things like model compilation um you know so like quantization, model compilation, different kinds of optimization to the model to make it run faster.

46:01And then the front end... So meaning your model may be distributed as some weights, some set of weights in like a safe tensors format, but the inference process is probably going to be using some other form or artifact in memory. And the backend that you're referring to the inference or is going to get as close to that as can be cached and like start there at for an inference just for performance reasons that's kind of what i'm hearing i have no idea exactly what is going on concretely like i don't i don't deeply know what compilation does i have a general idea but i don't want to i just like actually treat it just like magic i'm like okay this is gonna make my this is gonna make my uh model run faster it's gonna so i'm just going to apply the steps um that they're telling me to apply like you know yeah like there's a lot that you can you know yeah you know do this for people like and not have to know those details that it speaks to I think a level of maturity in the tool chain and like abstraction that, you know, wasn't here, you know, a year or two ago and probably was the reason why everyone said don't even bother trying to fine tune unless you're Facebook or Google, because you had to invent all this stuff yourself from the ground up.

47:35Sure. I mean, it helps to know a little bit like, okay, it's helpful to know what quantization is so that when you do, and you want to measure everything. So it's okay to, like, generally speaking, you want to try stuff. There's a whole host of different options. And a lot of these inference servers and backends have lots of knobs, lots of hyperparameters that you can tune. And honestly, like anyone that I talk to in industry, like, they don't understand all of it. And they just try stuff. Like, they try a lot of stuff. They fiddle with it. um it really depends on like you know your like throughput versus latency requirements like how to actually tune these things to do what you want them to do and so the reason why like it's okay you know why i'm okay with sort of treating some of it like magic is because i end up measuring it all you know and then kind of reasoning about what's happening it's helpful to be somewhat familiar with the terminology so you can kind of know like why certain things are happening but i would say like you don't need to be an expert in all of it to start using it um you can you can definitely measure things as well and kind of like you know you don't need to like know all the ins and outs of like model compilation and exactly how it works you can You can do it and see like, okay, what's happening to my performance, latency, so on and so forth.

49:08So yeah. So, you know, that part is not non-trivial. Like hosting your own model is actually non-trivial. There are a bunch of startups out there that will host your model like fireworks or replicate modal, things like that. but yeah you have to think about like those costs and you know how that makes sense given your scale and so kind of that's why it goes back to this whole like in the beginning just use off-the-shelf model because those off-the-shelf models like actually they're they're very like attractive in their costs like they don't cost that much especially at a small scale and they have somewhat low latency and they're available most of the time.

49:58So it's a great place to begin. It's actually very difficult to compete with that. So yeah, I think I strayed far from what you're asking. But yeah, like inference, and I'm not an inference expert. There's people, and there's people that came into the course that were like inference experts. Actually, this like whole debate about fine tuning, I had two guest speakers. One was Emmanuel Amison from Anthropic. The title of his talk was Why Fine-Tuning is Dead. And then there's another talk from Kyle Corbett, who is CEO of OpenPipe. And he was talking about... He was also saying, like, you should usually not fine-tune, but he had some other perspective of like okay in these situations um you know you it is extremely helpful helpful to fine-tune and open pipe is a uh is a tool that helps you fine-tune based on your logs so based on you know how we're talking about collecting those logs for fine-tuning um so i would highly recommend watching both of those.

51:08Those are like pretty rendered. Are the videos posted now or do you need to register for the course? No, no, you don't have to register. All those. So by the way, the whole course is going to be open. Most of it is open. They're just going through. We're processing the videos. You know how it is. It takes time to, whatever. So yeah, I can give you the links for those and you can put those in the show notes. Okay, awesome, awesome. Awesome. So the question one level up on the stack was, what were your alternatives in terms of LoRa versus? Was the answer to that VLLM and other inference servers?

51:47I feel like that's mixing up different levels of abstraction. So you can do inference with LoRa. The adapters, you can also merge the adapter back into the model so that you no longer have an adapter. or you can choose to keep it separate. And the inference frameworks are all good about, okay, like if you have an adapter, let us know. But you can, yeah, you can just, it used to be the case that like when it was earlier on, like not everyone had supported it, support for adapters and the inference side of things. So you kind of had to merge it. But now you don't, so. Are you typically choosing between LoRa and some other alternative when you're filling out your config file and like what are you thinking about as you're making that choice or yeah so okay i don't go i don't try to optimize it too much like i try to start with the simplest thing first so i'll usually try like laura and try to fine tune it on the like a like a consumer grade gpu or workstation grade gpu not server grade just to see what happens and if it's good enough if it's pretty good i'll just you know call it a day and say like okay this is like really good not try to like you know hyper optimize things you can sure you can do a full fine tune um but you know like yeah full fine tunes can give you better results um but i've had really good luck with using those adapters to be honest this like worked really well so i haven't really had to go i haven't really had to do full fine tune and can you do this reasonably on apple silicon like just on your laptop oh interesting i haven't tried that so i actually try to keep pleasantly surprised by what you can do now uh on apple silicon after you know years of just not having any gpu options yeah on a mac i know that it's like people are inferencing models a lot on their Apple Silicon.

53:53I haven't tried to fine tune. I know there's some frameworks out there, things like MLX or something that may be. Metal or something, metal extensions, I think. Yeah, something like that. Again, like there's too many things going on. A lot of moving parts. That I just tried to simplify a workflow. I don't want to deal with that. Yeah. So I just try to remove like these variables that can complicate things for me. So I just stay on NVIDIA. Because like, yeah, this is like cleaning the data and like, designing the application as hard, hard as it is. And it sounds like another, uh, kind of linchpin in all this is, uh, evaluation and measurement.

54:36And so that you can, uh, in, in the case of this conversation, like feel good about all the things that you don't necessarily know, like the inner workings of, because you know, the impact that they're having on your data and your models performance. So like maybe let's segue into the way you think about evaluation measurement for LLM apps. So evals, I think, okay, so evals is the most important thing. So like I think like when it comes to large language models, that's all I talk about nowadays because like in the consulting work that I do. So people call me when their AI doesn't work. It's like, hey, like we've kind of prototyped something.

55:19it um you know we're excited about it but it's like not really working because it demos well but it doesn't work like in the real world or it's not working as a lot of failure modes and we don't know what to do it seems like a lot of the game right now is like i can get a really nice demo but how do i get it to production actually i mean we've been saying this for a really long time and ml and ai yeah yeah but it's especially the case in large language models because you can get to demo insanely fast, like within a day for sure. And so, but even like the very polished products, like, okay, people don't know how to get unstuck.

55:57They're like, okay, I use like OpenAI. I've built like some interesting product around it. It solves a problem. But like the AI just like fails too much to do something or makes too many mistakes. and so the key like you know and then like what people do to start off with is they do prompt engineering is you know they see a failure mode of the something going wrong and they update the prompt but then like really fast it becomes very unwieldy and you end up making a change to a prompt and actually causes some other error and you kind of get stuck. And so that's where this notion of evals comes in is you want a systematic way to test the efficacy of your system so that you can make a change and you can see immediately like, okay, what, you know, did you fix your error?

56:57Did you create new ones? Things like that. And it's like very analogous to software engineering in a way. You would never write software without unit tests. It would be very difficult beyond a certain point. It would be unsustainable. Same thing with large language models. You need some kind of tests. And it's really the only way you can build AI.

57:21So it's integral to the building process of AI products. It's not even an optional step. the interesting thing is there's not too much information out there around how to do it there's a lot of tools um so you know there's a lot of people uh you know there's a lot of different startups and things like that and the thing that trips people up is they focus a lot on the tools but not the process of doing the thing. So the tool can help you a little bit in the sense like make the process easier but if you don't have a good process on like how to think about evaluations and do the evaluations then it's not going to help you.

58:09And so with evaluations really it's, you know, a process of looking at lots of data. First of all, removing all the friction you can in looking at data. So making sure that you have a way of looking at all these logs, these LLM logs, in a way that makes sense for you, like in a domain-specific way. Then that allows you to look through the data really fast and navigate the data and also navigate it in a domain-specific way that makes sense for your business. But then also, you know, writing tests, different kinds of tests. So not going straight to LLM as a judge. So a lot of people, when they think about evals, they're like, okay, one failure mode is they go to tools.

59:00And the tools have kind of this box of evals off the shelf. Like here's a toxicity eval. And here's a conciseness eval. And here's a helpfulness eval. The problem is that all that stuff usually is not helpful for you and your use case. generic what what you yes totally generic and um and so what you what you need to do is like write lots of different assertions lots of other kinds of tests so like you want to make sure that you write um as many assertions you as you can before getting to lm as a judge and then when you do go to LM as a judge, you need to make sure that that judge, that you agree with it.

59:50So the question, the problem with LM as a judge is the question particularly becomes like, how do I trust it? How do I trust this AI to score my data or to like, do I know? And so you can imply, you can apply like a systematic way to reason about that which is like have a human also score or you know make a judgment about that data and then measure kind of like how much in agreement those two things are and then try to align the judge over time you know with the human so like that's like in a broad sense kind of a process of like looking at data, writing tests, you know, and doing that over and over again, and then making changes to your system.

1:00:45And that also ties into fine-tuning because that is the work. Also, if you do that properly, if you have, like, a good, if you have good instrumentation, good way of looking at data, good way of, like, labeling, curating data, and you have all of these tests that can also filter out bad data, you can curate data extremely fast and you basically accomplish all the work of fine-tuning. So it does tie into fine-tuning, but it's more general than that. It's like, even if you're not fine-tuning at all, it's really the only way that you can systematically improve your AI. and you can almost like if you're not doing any evals at all you can almost guarantee or you can you know pretty much guarantee that you can improve your ai by doing evals because you have a way to test things and improve things and there's like different techniques on how to improve things but even things like you know even things like these automated prompt optimization frameworks.

1:01:59Like DSPY? Yeah, like DSPY or, you know, DSPY is probably the most kind of well-known kind of tool in that genre. You know, if you have really good tests and metrics, then that is going to, then you can actually do something with those prompt optimization things. If you don't really have a good handle on that and it's not really dialed in, then yeah, it's going to not be as great. So once you do evals, you unlock a whole host of things that are available to you. In addition to just being, just helping you make progress, you unlock all these other tools, whether it's fine tuning, whether it's prompt optimization, or really anything else.

1:02:50And yeah, so that's the thing that people struggle with. I would say a lot of people that pick up the phone and call me and they're struggling with their AI, you know, one of the first questions is like, hey, tell me a little bit about how you're testing your system and tell me about what we're measuring. And a lot of times there is nothing. A lot of times it's just, hey, you're just looking at it, like as a user. So, yeah, I think it's the most important thing, at least at the moment, in terms of what is the bottleneck in terms of, you know, moving the community forward in AI. Yeah. Yeah. Do you distinguish between kind of frameworks and tools when you think about testing?

1:03:49Meaning you said, you know, a lot of people kind of jump at the tools first and, you know, get stuck behind what the tools offer. Or is the alternative to that, like building everything from the ground up from scratch, custom to your use case? Or are there some medium level frameworks that you can turn to to help organize things? So what I recommend people do is actually try to use what you have first. So use your existing unit testing framework. Use whatever CI that you have. if you're going to like log the results of your test log the results in your test into whatever analytic system you have it's not going to be great but start there and do the process like two or three times and then you have a good foundation in which to think about tools the tools can certainly process that you're doing two or three times like yeah like is to like go through it to like write the test observe like the results of the test try to make your system better try to you know to bang your head against it a little bit until you feel like you're at the try to you know make some progress on your system like oh i have this error let me try to make it better and then like fix the error observe the error rate going down validate that like okay you kind of have tests that make sense and then like will click in your head like okay this is what this this is a good way to think about my system this is what I'm trying to measure this is what's important to me you know I need to look at this kind of data this is the kind of failures that I even have you know this is a process like who is going to look at the data who is going to write the test you know all these different moving parts like you kind of have an idea and then when you approach a tool, you can actually use the tool.

1:05:49You can like appreciate the person. Yeah, you can appreciate the tool and you can evaluate these different tools and like, you know, for your domain. Um, and then you're like way likely, more likely to be successful rather than, yeah, if you're just going straight to a tool, it's really hard for people to, you know, they just do whatever the tool says on the front page of the documentation. They don't know. So, yeah, that's kind of what I would recommend. But there's different kinds of, there's an open source tool. So there's something that JJ Allaire has been working on. But JJ Allaire, he created RStudio.

1:06:31ColdFusion and all that stuff back in the day. All kinds of crazy stuff. So he's recently created something called Inspect, which is an open source evaluation framework. So yeah, he has one that's open source. There's not too many open source things. There's another one called OpenLLMetry. Okay. But I think even OpenLLMetry is vendor affiliated. So I think JJ's is the only one that I know of that's completely non-vendor affiliated. But then everything else is vendor stuff. so as I said like first yeah let's like do it yourself based on the tools you already have maybe if you don't want to do vendor stuff for some reason maybe you're just by yourself and you're not like wanting to do yeah you can try this open source thing and then from there maybe try one of the vendors or just skip the open source thing go straight to a vendor if you want so the key insight is

1:07:36bang your head against a problem in your own words and in your own terms before you jump to some tool to define the problem for you because you won't know what you're kind of missing in that loop until you do. Yeah, absolutely. And it's not that bad. It's not as hard as it seems. It's actually, it's a good exercise to go through yourself. It's not like you're going to create all this technical debt or something by doing it once with the tools that you already have. There's a lot of people that just stick with that, by the way. It's like user tools they already have and they like kind of like it.

1:08:13But you know, yeah. So it is good to start there. And do you find that it's obvious to folks what to measure or to what degree is it use case specific versus kind of standard? um you know just by virtue of the fact that it's an llm system yeah it's usually very all domain specific okay um like when we look at the data i think with enough coaching um you know 100 of people that i've come across have been able to learn how to do it very effectively there's some skills involved in learning how to look at data and sift through it basically teaching people to be very skeptical and saying like hey something's broken like let's be skeptical of the whole system you know let's look at your data very carefully a lot of times it takes pair programming and pair debugging but I find that when people people usually get it really fast they're like okay you know what this is it's kind of like it's kind of similar to the scientific method.

1:09:35You know, it's like a way of thinking like, Hey, I need to prove to myself like why there, this is like black box in a sense of this large language model. Don't necessarily understand like exactly how it's working, but you can run some experiments. You can measure some, you can take, you can take some measurements and you can change some things and you can, you can observe the changes to those measurements and you can think about that. And I think people kind of like a lot of people have that inherent skill of thinking about that or like working that way. And it's like once it clicks in their mind, people are able to do it.

1:10:14Yeah. Maybe as a follow up to the question about kind of the domain specific nature of these measurements, like what are some concrete examples of like, you know, a project and the important things to measure, you know, that became apparent to measure in that project relative to, you know, other types of projects. Yeah, so, okay, like the real estate CRM example, the one I was talking about earlier, the company ReChat. So, like, their errors are quite boring in the sense, like, okay, you have a UUID that's showing up in the output that's contained in the system prompt. And it's contained in the system prompt because it's trying, you need it to use, you need it for a function call later on.

1:11:00You might need to perform it later on. For whatever reason, that user UUID has been reproduced in the final output, and that's not good. Or certain sentences are being repeated twice. are in ReChat's case, because they're rendering these like UIs, these widgets, not just text, they have these special delimiters they're using to indicate like what the metadata is. Where to insert the widget. Yeah, where to insert the widget. And things like, okay, is the delimiter, you know, is there something wrong with the way the delimiter is placed? You know, there's all kinds of assertions like that.

1:11:54And so at one level, like there's, you know, there's query of response and you can like, you, you know, probably want to be tracking everything at this query and response level. But it sounds like also you're almost like featurizing that or applying assertions to that and coming up with flags that identify specific errors and storing that information. is that the idea definitely yeah the idea is like and the reason i'm harping on assertion so much is a lot of times people think they can't write assertion they're like hey it's fine i just it's very subjective and then that is the case sometimes but oftentimes i will look at data very closely and i'm like i find kind of these patterns of failure modes that can be expressed as assertions.

1:12:45It only comes to like looking at the data very carefully. And that's kind of the first line of defense. Say, hey, you should try to solve all these dumb failure modes that you can test with assertions before moving on to things that are more subjective or require like higher level of reasoning that then you have to go on to use like an LLM as a judge or something like that. Interesting, which kind of goes back to the like bang your head against it. Like, well, collect the data and bang your head against it because I'm imagining like you collect the data, you, you know, bang your head against it and you'll start to identify these common patterns of things that can go wrong.

1:13:25And then you can both longitudinally track them. But if you know that anytime like a bad delimiter, you know, is, you know, appears in your response is bad, You can just like put up a guardrail against it in the, you know, in the application so that you're not showing the bad thing to your users or you're taking some action, regenerating it. Or if you can fix it, fix it manually or whatever. Or like there's, you know, once you start to, you know, once you get the process down of like building a mental model around the things that go wrong or expressing that in terms of assertions, there are lots of places that those assertions can manifest or be used.

1:14:07That's a great point. And so this is like why evals are not really a separate thing. It's very much just part of building. And you're not, you can't build without it. and kind of as you evolve the testing, you don't want to, let's say, silo the assertions in your unit testing framework, for example. You want to keep it in some other location where you can both use it in your unit testing framework, but also you can use it in your LLM invocation pipeline. So you can do things like test for that and then do some self-healing, depending on how expensive that specific test is. So the assertion is cheap, so you can do it.

1:14:47and so like all that stuff is you know it's it kind of very much blends into the process of building the ai very cool you know another example like the honey the honeycomb example the texas equal it's like not really texas equal it texted that domain specific language but the assertion there is like is the query syntactically correct yeah and not only syntactically correct like in the abstract sense of like is it syntactically correct given the user schema and things like that so that you push the assertion as fat as far as you can go in a way that makes sense like in a naive you could say is it a valid honeycomb query okay that's one but like if you think about it further you can you gotta you say okay like let's think about the schema as well and so um it's always the first line of defense yeah interesting i I came into the conversation thinking that like the conversation around evals would be like very, I don't know, like maybe dialogue centric or RAG centric or, you know, precision recall, that kind of stuff, which comes up in a lot of conversations around RAG.

1:16:03But this is much more of like a software engineering. Yeah. I mean, you do want to measure the various components of your system. like RAG, and you definitely want to measure retrieval. Like, actually, RAG is a fantastic place to apply evaluations because retrieval is a very mature field, especially information retrieval. So you can apply a lot of the tools that we've learned over the last few decades to RAG in terms of your ability to retrieve the right information. There's a lot of metrics. There's a lot of science out there around like in tools. But even before you go there, it's good to have these like more simple kind of holistic evaluations of your system.

1:16:54And then you can start to like measure these different components and stuff like that. Awesome. Awesome. I think that's probably the most important thing is to like look at your data and it's really to look at your data. It's easier said than done. We've been saying that. You've probably been hearing that in your podcast for like, you know, as long as the podcast has been alive. It's very easy to say no one does it. It's why I have a consulting business because no one's looking at the data. If people start looking at their data, then I, you know, I could do something else. But they're not. And I think it's a little bit cheeky because there is some skill involved, actually, that we can take for granted as data scientists.

1:17:39Like there's a skill in like how to sift through your data really fast, having a nose for what's wrong, how to like kind of analyze your data or like look through it with the lens of the machine learning system and kind of intuiting like why something might be happening. So it's like, it's funny for me to say like, hey, look at your data, but there's a lot there. But, you know, starting to practice it is very critical to this whole building AI stuff. You know, it makes me relate back to a comment that you made earlier in the context of fine tuning. Like we're far from like the, you know, the hard days of training deep learning models.

1:18:26And, you know, even before that was like old school data science. But like you're kind of pointing to this idea that it's, you know, either it's a cycle or it's all connected or it's all the same thing. like it's all about looking at your data and understanding how to intuit your way around the relationship between data and some objective that you have. Yeah, definitely. Yeah. Very cool. Well, Hamill, thanks so much for jumping on and sharing a bit about, you know, what you've been learning and seeing with regard to eval and fine tuning and building systems around around LLMs. It's great to speak and catch up.

1:19:10Yeah. Thank you.

From the publisher

Today, we're joined by Hamel Husain, founder of Parlance Labs, to discuss the ins and outs of building real-world products using large language models (LLMs). We kick things off discussing novel applications of LLMs and how to think about modern AI user experiences. We then dig into the key challenge faced by LLM developers—how to iterate from a snazzy demo or proof-of-concept to a working LLM-based application. We discuss the pros, cons, and role of fine-tuning LLMs and dig into when to use this technique. We cover the fine-tuning process, common pitfalls in evaluation—such as relying too heavily on generic tools and missing the nuances of specific use cases, open-source LLM fine-tuning tools like Axolotl, the use of LoRA adapters, and more. Hamel also shares insights on model optimization and inference frameworks and how developers should approach these tools. Finally, we dig into how to use systematic evaluation techniques to guide the improvement of your LLM application, the importance of data generation and curation, and the parallels to traditional software engineering practices.

The complete show notes for this episode can be found at https://twimlai.com/go/694.

More from The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

All 156 episodes
Building Real-World LLM Products with Fine-Tuning and More with Hamel Husain - #694The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) · 1 h 20 min
Listen in VO