In short
Practical AI Podcast Episode Notes
Episode Title
Fine-tuning vs RAG Episode Description In this episode, hosts Daniel Whitenack and Chris Benson are joined by Demetrios Brinkmann from the MLOps Community to discuss the nuances between fine-tuning models and retrieval augmented generation (RAG). The conversation also touches on OpenAI Enterprise, insights from a recent MLOps Community survey, and the orchestration and evaluation of generative AI workloads.
---
Key Guests
- Demetrios Brinkmann - [Twitter](https://x.com/Dpbrinkm)
- Chris Benson - [Website](https://chrisbenson.com) | [GitHub](https://github.com/chrisbenson) | [LinkedIn](https://www.linkedin.com/in/chrisbenson) | [Twitter](https://x.com/chrisbenson)
- Daniel Whitenack - [Website](https://www.datadan.io/) | [GitHub](https://github.com/dwhitena) | [Twitter](https://x.com/dwhitena)
---
Episode Highlights
- Introduction and Context
- Hosts: Daniel Whitenack and Chris Benson introduce the episode with a focus on productive AI implementations.
- Guest Introduction: Demetrios Brinkmann is welcomed back, highlighting his role in the MLOps community and recent events.
- MLOps Community Updates
- Recent Events:
- Discusses recent LLMs in production conferences with significant speaker participation and in-person meetups globally.
- Highlights the vibrant growth of the MLOps community, noting events in major cities and the establishment of new local chapters.
- Shift in AI Discussions
- Use Cases & Practical Applications: The conversation shifts to how discussions have moved from theoretical to practical implementations. Users are now sharing concrete experiences and workflows.
- Survey Insights: Demetrios mentions a recent survey conducted by the MLOps community to gather insights on LLM usage.
- Fine-tuning vs. Retrieval Augmented Generation (RAG)
- Fine-tuning Misconceptions:
- Demetrios emphasizes a common misunderstanding of what fine-tuning can achieve, particularly with LLMs versus diffusion models.
- Fine-tuning is often misapplied to tasks where RAG would be more appropriate.
- RAG Advantages:
- RAG provides a more efficient way to adapt models for specific tasks without the extensive resources required for fine-tuning.
- He suggests that using retrieval systems along with few-shot prompting can yield better results than fine-tuning.
- Challenges and Considerations
- Evaluation of Models:
- Discussion on the difficulties in evaluating LLMs and the inadequacies of standard benchmarks.
- Emphasis on the importance of context and specific use cases when evaluating model performance.
- Common Pitfalls:
- Many companies assume fine-tuning will magically improve their interaction with LLMs without understanding its complexities.
- Highlighted confusion around what constitutes effective fine-tuning versus simple retraining.
- Upcoming Courses and Events
- New Course Announcement:
- Demetrios mentions an upcoming course focused on RAG, designed to quickly teach participants how to set up effective LLM systems.
- Hackathon Insights:
- Feedback from a recent hackathon indicates that practical, hands-on experience is invaluable for understanding how to build effective AI systems.
- OpenAI and Enterprise Solutions
- Enterprise Skepticism:
- There is ongoing skepticism regarding the use of OpenAI's enterprise solutions, especially concerning data privacy and security.
- Discussion on how larger organizations are slower to adopt new technologies due to compliance and security concerns.
- Positive Trends in AI
- Access and Innovation:
- Demetrios expresses excitement over the rapid democratization of AI technology, allowing more product managers and teams to implement AI solutions quickly.
- LLMs in Production:
- The potential for LLMs to streamline processes and enhance productivity in various sectors is emphasized.
---
Conclusion
- Final Thoughts: The episode wraps up with a call for listeners to stay engaged with community learning opportunities and to participate in ongoing discussions about practical AI applications.
Related Links
- [MLOps Community](https://mlops.community/)
- [LLM Survey Report](https://mlops.community/surveys/llm/)
- [LLMs in Production Event - Part III](https://home.mlops.community/public/events/llms-in-production-part-iii-2023-10-03)
---
Episode Sponsors
- [Fastly](https://fastly.com/?utm_source=changelog) – Powering fast, secure, and scalable digital experiences.
- [Fly.io](https://fly.io/changelog) – Deploy apps and databases close to users, with no operations required.
- [Typesense](https://cloud.typesense.org/?utm_source=changelog) – Fast, globally distributed Search-as-a-Service.
---
Listen to the Episode For more insights and discussions about AI, subscribe to Practical AI on your favorite podcast platform!
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:06Welcome to Practical AI. If you work in artificial intelligence, aspire to, or are curious how AI-related technologies are changing the world, this is the show for you. Thank you to our partners at Fastly for shipping all of our pods super fast to wherever you listen. Check them out at Fastly.com. And to our friends at Fly, deploy your app servers and database close to your users. No ops required. Learn more at fly.io.
0:42Welcome to another episode of Practical AI. This is Daniel Whitenack. I'm the founder at Prediction Guard, and I'm joined as always by my co-host, Chris Benson, who is a tech strategist at Lockheed Martin. And today, Chris, I don't know if you've been listening to the Changelog, our sister podcast. they've been doing these like change log and friends episodes where it's not necessarily like a guest interview, but it's like, Hey, let's invite one of our friends on and just talk about cool stuff. And I feel like we've a little bit got like practical AI and friends today because we're joined by Demetrius from ML Ops community, which, you know, you're involved in events and podcasts and reports and surveys.
1:29and basically you run the whole AI world. So, you know, welcome. And we're glad to have you back as our friend. He's like the deep state, you know? He's the deep state. He's like running everything behind the scenes. I am honored to be considered a friend, first off. I just want to say that because I appreciate the amazing stuff that you all are doing here. And of course, whenever I get the opportunity to come and chat with you folks, I am going to jump at it. Excellent. I know that last time I saw you, I think, was at one of the recent LLMs in production event, which was super fun. Give us a little sense of what has the past few months looked like in the ML Ops community and some of the events that you've been doing.
2:22It seems like there's so much. I think you had also in-person things. So give us a sense what's happening. Yeah, I appreciate you calling that out. And of course, I appreciate you presenting at the LLMs in production conference. That was a blast. Yeah, it's a good time. That event itself, we had two days and each day had three tracks. So there was two full tracks and then one workshop track. And there was over 82 speakers in the whole event. and man, that was a lot of work. Yeah, I bet so. Add on to that. So that was the virtual part of it. But add on to that, we had in-person parts of it. So in Berlin, there was a hackathon that we did.
3:10And in Amsterdam, we had a meetup. And then in London, we had a watch party. In San Francisco, we had a meetup that happened and then a hackathon right after and then a workshop. So there was all kinds of craziness that was going on. And that was just for the LLM in production event. Later on, I mean, we're in now 37 cities around the globe. So wherever you are, you probably have an MLOps community meetup around you unless we're in like North Africa and South Africa. And then also even in Lagos, Nigeria, we're in Australia, which is always awesome because I keep threatening to go out there and just pop up at a meetup.
3:57Yeah, so that's the in-person stuff. There's people that have just gotten super excited about it and they decided to start a community chapter, which is the unbelievable power of community. I am blown away by it every time someone approaches me and says, hey, I want to do something in my city. Yeah, that's so cool. And I've been able to attend a couple in-person events this year related to AI stuff. And I will in the fall as well. I'm curious to know, from my perspective, it's cool to see things, I feel, transition a little bit from kind of hypothetical stuff, a lot of discussion, to people talking about like, hey, we did this.
4:46This is how we implemented our workflow. Hey, have you tried this? And so I, of course, love those types of conversations. Do you get a similar vibe? Or how does the community in terms of generative AI and LLMs, how does it seem different now than even like six months ago or something like that? Oh, yeah. That's so good because it does feel like there are use cases that are becoming very clear on what LLMs shine in and what they're not good at. And then there's also the stack that's forming. And we did a survey, probably. That's another thing that we did in the community on top of all the fun other stuff.
5:27We did a survey and we surveyed people that are actually using LLMs or even not using LLMs. And we asked them why you're not using them. And we went through a bunch of stuff or why you are using them. What are some big pain points? And it was becoming very clear that there are certain use cases that people are using LLMs for. And we can get into that in a minute. And then there's also this stack that is forming. And the stack was probably the most interesting. I know you guys mentioned it with the A16Z article and Rajko is one of the authors of that. And he helped me with the report when I wrote it too.
6:05So he's behind the scenes, like moving the puppets, the puppeteer. The cool thing with the stack is you kind of have, if I break it down for those who haven't seen the diagram that we put together, it's like you have the foundational model and then you have some kind of vector database which is like the hero the champion in this whole llm scene you've got if you need to do some fine tuning or model building you have that component to it but we can get into that i'm very opinionated about the fine tuning part and i'll tell you why in a bit and then you have stuff like developer sdks and this is you know like your llama index or lane chains.
6:47And then on top of that, you have like the monitoring or experiment tracking, prompt tracking, those kind of things like a port key, a promptimize, prompt layer. There's all kinds that are coming out. And so what we didn't have in that moment that I think are starting to emerge more now and I'm really excited about is like, how are people actually evaluating these models and is it coming with your different tools or are you doing extra stuff on top of it so that was the inspiration behind a whole nother survey that we're doing right now on evaluation yeah i can say personally i'm doing extra stuff on top that's my like short and very short answer, which is, of course, much more involved.
7:44But yeah, I don't know. What is your sense of that? Just intuition-wise, am I out of the norm or in the norm with that? No, you are completely in the norm. And I think the hardest part, and this is why we wanted to do a survey around it, is because this is one thing that is super unclear. And nobody really knows if they're doing it right and they don't really know what the best practices are. And so you also don't really know what you're evaluating. Are you just evaluating the model? One thing's for sure. All these benchmarks are complete bullshit. That we all know, right? That is very clear. How do you really feel about it?
8:29Yeah. And they are interesting in certain ways, but the types of evaluations that are going on there. So like, I think they serve a place maybe, but they don't translate into like, okay, now I have this use case, right? If I take the model on the top of that leaderboard, I am very much not guaranteed to have like the quote, the best results for my use case. And I think that's what's confusing to a lot of people. 100 % that's exactly it is that these models and the use case that you have who knows how it's going to match up against one another right and then it's not only that but how are you monitoring or evaluating for toxicity or the ability for it to do the one thing that you care about I mean I don't care if it's and also it just kind of feels like a lot of marketing at the end of the day when you see the newest model comes out and everybody loves SOTA.
9:37SOTA. This is, you know, state of the art. SOTA beats ChatGPT on all these different metrics. And I just kind of laugh because it feels like I've desensitized to that these days. Yeah, I think also it's kind of often kind of funny to me that even like ChatGPT is being used as a static baseline for these things when it's not even like chat GPT isn't a model, right? It's a product that has layers on top of it for, you know, that handles all sorts of things. And so that's another misconception that I've seen is like, well, is it really fair to compare a model's output to the output of a product that has a lot of kind of functionality built around it and with it.
10:28Yeah. And that also kind of lets people know when that's drawn out, it lets people know that, hey, the LLM here is not your application. There's this whole layer on top of it, which I know you were talking about retrieval based augmentation or retrieval augmented generation as we are gearing up for this episode. There's of course, like whatever your opinion about prompt engineering is, there is an engineering element to how you call these models and chain things together. And like you say, evaluate things, validate things, filter things for whether it be toxicity or factuality or, or whatever it is.
11:10So there's just so much around that, that's not the LLM that I think people confuse those concepts a lot of times. You know, just as a little aside here, listening to you guys talking about this, and you guys are experts at this stuff, and I'm just thinking about all the poor people out there who are listening and maybe aren't at your level. It's a tough thing to try to figure out how to navigate this. When you think about it, you guys are debating this and are not completely in alignment yourself. I'm having a lot of empathy for people in the audience who are going, how the hell am I supposed to do this?
11:48Well, and then alignments, the other thing. Yeah. Oh, yeah. Another buzzword. Tick it off on the buzzword bingo. There you go. There you go. Well, it's funny because when you are trying to figure out your use case and how to get the best performance out of an LLM, there are things that you go through, right? There's like almost stages and you figure out the debugging is quite difficult. as you were saying, is it the prompt that's giving me the problems? Or is it that something in my retrieval or the way that I'm creating these vector embeddings, are those the problems? Where exactly is my problem and how do I isolate that so that I can make the whole system better?
12:35And that is, again, it goes back to evaluation and evaluating the whole system. How do you look at what you're doing as a whole, as opposed to just like, oh, cool, there's this model. And if I go to a hosted version of it and I ask it if the earth is flat, it tells me yes or no. And most of the time they say yes, which is crazy. It's like, oh, you've been trained on one too many flat earth subreddits. Or they've just seen a lot of answers that are positive. so it seems probable yeah yeah i as you're talking i think one of the things that i've realized is in the a16z kind of stack they call this layer orchestration which i think a lot of these tools are amazing that fit into that layer and the things that they're doing you know lang chain llama index we've mentioned sometimes though it's just like how rapidly the field is advancing.
13:32If you import whatever chain from LangChain and then whatever model and then whatever vector database and then you put it all together, you run the thing and then you get an empty string output. You're saying, where did it go wrong? And how do you divide and conquer, debug that? So I think a lot of what I've just in my own applications that I'm working on for clients and other things. A lot of times, to be honest, for me, it's a lot simpler to write out my chain of my LLM reasoning in just regular Python logic, add like whatever exception handling I want, like make the call to the vector database like manually and like create some logic around that.
14:22So I feel that that sort of Python DIY side, it's like less convenient, but I usually end up getting there still. You know, again, it's part of the maturity of this field, I guess, and partly like how things are advancing quickly. You know, they'll have a fix for that. You know, in the spirit of you may have heard like, you know, when you have a problem with Facebook, what's the answer? more Facebook. Well, in that spirit, I'm sure Meta will put out debug llama for you in no time, and it will solve all of your problems right there. It's funny too, because it feels like that just forces you to stay simple.
15:02Yes. Again, like going back to the KISS principle and realizing, you know what, maybe I'm trying to over-engineer this and I can go far with, maybe I don't even need a vector database, which is kind of blasphemy. How dare you? You did call it a hero and a champion a few minutes ago, I want to point out. Yeah. I mean, it is wild that you can go so far. I talk to startups almost every day in my day job, which is not the MLOps community. And they have real revenue. And a lot of them, I'm always asking them like, oh, so how are you doing this behind the scenes? And a lot of them aren't even using vector databases.
15:45But they're like cashing in and they have a product that's working and it's working at scale. And so there is this misconception, I think, sometimes that we need all the bells and whistles. And this stack that I just was talking about and saying that the vector databases are the champions that potentially for your use case, like, do you really need it? Because it does add that complexity. And so going back to the tried and true principle, like just keep it simple.
16:25You mentioned that you had some strong opinions related to retrieval and fine tuning. I think this is the time for the hot take. officially declaring this practical ai and friends episode so it's a safe space to declare your hot take safe space i love it well i know i just think that i hear a lot about fine tuning and i don't know if people who throw around the idea of fine tuning something really understand what you fine tune a model for, especially like LLMs. If it's a stable diffusion model, that's a whole different story. And I think sometimes these diffusion models, they give us the wrong idea of what fine tuning in LLM will do.
17:18So undoubtedly you saw the rise of like Lenza or Photo AI or all of these fine tuning. Basically what these companies were doing in the background is they You're running some kind of diffusion model and you would upload your selfie or a picture of your dog and you would be able to bring it into the world of AI art. And if you take that concept over to LLMs, you think, oh, well, if I just fine tune an LLM on all of my emails, then the LLM will know how to write emails like me. But it's not like that. There's the misconception that it's not like equal in that regard. Fine-tuning, you don't fine-tune something so that it can understand you more and you can call it out and say, now write like Demetrius.
18:14Because what you want to fine-tune for, for that case, let's just be clear, that's where retrieval augmented generation shines. Because you just say, hey, here's a database or a vector database of all of Demetrius' his emails and the most you can do some few shot prompting and say, write like this. Here's like five styles of Demetrius writing a response to this. So make a sixth one and you're golden. You don't need to go through like burning a lot of cash on GPUs and GPUs are scarce these days to fine tune some model that may or may not work after you fine tuned it. So I think it's worth calling out like when you should be fine tuning things because I don't I don't want to say like never fine tune it I just want to say like I've seen a lot of people talking about it and also a lot of companies starting that will tote how easy it is to fine tune and how you should be fine tuning and use your company's data to fine tune it's like fine tuning I think the best way that I heard it talked about was in a recent MLOps community podcast that I had with Shaul S he created ragas which is like an evaluation framework for rugs that's awesome by the way go check it out but he's also very big in fine tuning and he is also one of the main people that does the open instruct project which is a whole open source llm project and so he was mentioning how you want to fine tune when you have some new function or some way some new output something that you need to teach the LLM to do that it doesn't necessarily know how to do.
19:55So a perfect example of this is OpenAI's functions. So like ChatGPT or GPT-4 functions. This makes it much easier for you to get very, very clean data or data that's outputted in a certain way, right? And it gives you this structured data. And that is a whole reason that you would want to fine tune something. Or the other example I think is Lama, the code Lama. But I think where code Lama falls down is that if the original base model doesn't see a lot of examples of code, then no matter how much you fine tune it, you're not going to get a good coding model out of it. And some people who are using code llama and loving it, like, come tell me because I haven't found anybody yet.
20:47So that's my rant about fine tuning versus rags. And I have the last thing that I will say is like, if you're looking at fine tuning and you've been brainwashed by the society out there or the community at large, and you think you need to do it, like, let's just remember why ML was so hard before LLMs. Collecting data is not so easy. Labeling data, cleaning data, all of that stuff is quite difficult. And you need to do that if you're going to fine tune. I just want to say, I want to remind everyone that Daniel likes to clean data. That's something I miss. I feel like because of LLMs, I'm not doing it quite as much, which is a hole in my life.
21:36I'm determined to remind our audience of that fact. It's just kind of horrifying. So I kind of remind everybody once a year. It's therapeutic. There you go. I think along with some of what you've expressed, Demetrius, that there's a general misconception about the data that you would use to fine-tune an LLM. So I've been in countless conversations where the idea is, oh, well, we're wanting to do question answering on our documents or something like that. And we would love to fine-tune a model on our internal company documents, and then it's going to be better at question answering. So what I generally tell those people is like, hey, when you're fine-tuning one of these models just on that raw unstructured text, in the best case scenario, what you're creating is a better autocomplete model to autocomplete your type of documents, right?
22:43What you're not doing is creating a better question answering model. Those are two way separate things, right? So if you wanted to fine tune a model like that, like you're saying, there may be cases where that's useful. Maybe it's a domain thing. Maybe it's a very like you want very structured type of output answers like you're talking about. But in that case, what you want to do, you actually don't want to fine tune just on that raw text data. You want to create your own set of instruction prompts likely that are fed in with various questions and you have the answers and you have everything like set up and you have thousands and thousands of these examples, which that I think when you frame it that way, then people are like, oh, so it's like a lot of work to create that sort of data set.
23:36And yeah, like turns out it still is, right? Oh, that's so funny. That is so well put too. I love that because that's my rub on fine tuning is that people don't realize how difficult it is and how much effort and how much work goes into it. and that's kind of why Mosaic sold for a billion, you know, because it's hard. And so companies that actually are doing it, that Mosaic was ever able to convince they needed to do it, they got a lot of money for it. But that's a whole different story. Yeah. And maybe for those that are less familiar in our audience. Maybe they've interacted with LLMs. They understand the concept of fine tuning.
24:24Maybe they're not as familiar with retrieval augmented generation. I know that one of the things you mentioned is the ML Ops community, you're doing a whole course on this, it sounds like. So could you just give the kind of high level pitch for retrieval augmented generation, what does that mean to you as a person who is obviously a promoter of this approach and something that people should try as they're getting into these technologies? Maybe that general pitch, and then we can talk about maybe a few more specifics that you'd like to highlight. Thank you for bringing that up. I mean, this is the first time that we're actually doing a course, so it is a little nerve-wracking and at the same time, super exciting.
25:12And so I'm just going to clarify that it is not me leading the course. We got an expert, the expert of experts who is Raul, who created the course for us. He is an engineer in San Francisco that's been doing this for a long time and he's been doing it at some serious scale. So he knows what's up and he goes and does everything from like the Kubernetes clusters to the end prompts and then monitoring the whole system. So again, I'm a big thinking about things in systems, but when it comes to like retrieval augmented generation, let's just go back to what is the hero of our story in this case. And it's not the vector database.
26:01I would say that like question answer systems are basically the hello world of working with LLMs these days. That's kind of what it feels like to me. I don't about you guys if you have that feeling too where you want to get your feet wet with llms you probably are going to do some kind of like talk to your data or hey if i ask this a question what responses do i get and in those use cases the best way to set up and architect your system is through like retrieval augmented generation again going back to what i was saying earlier not absolutely needed because I've seen viable businesses built without it but if you are at that point where you're like oh well this will add to my tool belt you know it's like another tool in my toolkit then what we created was a course that goes over and we wanted it to be something that you don't need to spend like six weeks on we wanted it to be that you could level up really quickly.
27:05And so you go through creating a data pipeline and pre-processing that data, then ingesting that data into a vector database. And then you can semantic search for the answers from the questions that you're getting from the end user. And then it compiles a response using an LLM. So that's like the basics of the course. And the reason we wanted to do this was because of a hackathon that we did. Actually, funny enough, right before the LLMs in production conference, we did a hackathon in San Francisco, and it was all about how bulletproof your LLM stack is. And so we created a bunch of questions and we gave everyone that was part of the hackathon all the data from the MLOps community Slack, which this Slack has been around since 2020.
27:58and there's like 17 ,000 people and it's very active. Like all the channels are going off every day. It is very, very active. I can't remember how many megabytes there were, but for text data, it was a lot. It was, people were saying like, oh my God, that's some serious amount of text. So everyone had access to that. And then we would rate everyone's stack on how accurate the answers were from the questions and answers that we gave them. So we basically asked people to build these different QA bots or chat bots, if you would call it that. And then at the end of the hackathon, we gave them 100 questions and we saw how accurate were these responses.
28:44And so the questions were some things that were from Slack or they were just random questions about ML and MLOps. and the best ones were the ones that would give you a accurate answer and then cite oh but you know this is another way of looking at it and here's a thread on it and so it would go back to the slack thread so all that being said yeah we're excited about the course we're excited to do that i mean there's all kinds of cool stuff that i want to do in the community and the only way that i'm able to do it is that people in the community are participating in this i said raul is the guy that created this course.
29:24We've got all kinds of other courses in the mix from other community members. So people that are experts on what they feel like they know best, they can propose topics and then we're just putting it on our learning platform. And where do people go to find said learning platform? Because I need to point a couple of my coworkers to it. Well, if you just want rants from me about how you shouldn't fine tune, then you can find us on the MLOps Community Podcast. But the easy one is we have learn.mlops.community. And that should get you to the learning section of our website. And yeah, as I said, it's exciting.
30:10It's a little bit nerve wracking because we plan on doing two styles. This one that we just released is go at your own pace. You get it and then you can go through the lessons and hopefully you can get your company to pay for it because it's within the learning budget and it actually is for work, right? So that should be an easy sell. But if your boss is not interested in paying for it, just DM me and I will give you some amazing copy that you can send to your boss, a nice little email that will hopefully convince them to change their mind. But the other pieces that we're going to start doing like cohort-based courses, and that is interesting on another level because we've got the whole MLOps community and we've got everyone that is part of the courses.
31:04They can go into a special Slack channel. They can be with the teachers and the teacher's assistants and all that fun stuff. I mean, it's not like we're breaking any new ground here. Courses are kind of a tried and true method of learning. So this is just us having fun with it. Yeah, that's awesome. And speaking of learning, you already mentioned the surveys that you've been doing. And there's a previous survey that it's already published. I know you're you're working on the next survey that will come out. but people should look at this survey. We'll definitely link it in our show notes. And there's a lot of really interesting stuff in the survey and some of the highlights.
31:46I'll just call out a few of the highlights here and then I'd be curious to know what stood out to you, Demetrios or Chris, whichever one. So some of the highlights just at a high level company use cases like text generation and summarization are useful, but participants are going deeper and exploring a lot of other ways to use LLMs like data enrichment and data labeling augmentation, generation for subject matter experts and other things. It's still unclear. The use of large language models in organization is still unclear due to some high costs and unclear ROI. They talk about hallucinations, the speed of inference with LLMs being potentially a blocker for certain types of use cases.
32:32You talk about the infrastructure. We already talked about that stack a little bit. Some of the things around augmentation and consistency of models. So those are just some of the highlights from the survey, which people can go into a lot more detail on. But I guess one question I would have is, as you were building this report, first off, is the whole report generated using an LLM? Because I didn't see that highlighted in the report. That would be the most meta thing ever, huh? Not the company meta, just the old term for meta. So that is classic. Yeah. And honestly, I tried really hard. But one thing that I did, which I'm going to preface this with, I am not Gardner.
Read the full transcript
33:23And so I did not know how to create reports before I did this. Now I've learned since I learned what was painful. And I spent months with this data and just tearing it apart, trying to figure out what are some clear signals here? And you can find the signals, but the reason that it was so difficult and that it was a blessing and a curse was every single question, instead of having answers given to you, it was just a free form text box. And so this is the report that we created, but the actual raw data, it's linked in the report and anybody can see it. And so I'm a big fan of that because I have my biases.
34:08I wrote this report, but the report is for the people that don't have the time or want to go and spend hours upon hours looking through all the raw data. Fine tuning a model on it. Yeah. They're like, give me the TLDR and I'm good. And so that's kind of what I set out to do. And I also wanted to see what are some big things that stood out to me, but not only me, every time, I think the reason it took me so long to release this was that I would ask a cohort of friends to review it and they would give me feedback and I would incorporate that feedback and then be like, okay, I think I can send it off to the designer now and let's get it going.
34:50And next thing you know, I would ask another cohort of friends to review it or people in the community and boom, I would have all kinds of new feedback and, oh, but you know, I noticed that there was this and people were talking about that. So I did that probably six or seven times. And that gave me the confidence that it's still biased, but it's not like as crazy biased as I think it would have been if I just put it out myself. And I got a lot of input from other people. So if anything, it's biased from a lot of different people. And so there's that piece and what we're doing with again like the evaluation survey I want to do the same exact thing I learned my lesson so it's not only free text form answers that you have now I kind of put a lot of multiple choice and check all boxes that apply and then also the other at the end So hopefully that will help me be able to do more Excel, fancy math on it and formulas.
35:54Because even with LLMs, I was even trying all these new ones that, you know, it was use ChatGPT in your Google Sheets. It didn't work. I spent days trying to ask it questions and it just did not work, man. And I ended up getting kind of frustrated because I spent more time trying to get the LLM to give me some kind of insight than if I just spent the time with the data and got the insight. And I think everybody who has played around with LLMs has probably had that experience once or twice where it's like, I've been prompting this for a really long time. I wonder if I just sat down and wrote the report or if I sat down and tried to think of things on my own and create something.
36:40I could have just done it in the amount of time that I've been prompt tuning this. But anyway. I get a question as I'm looking at it. Like, obviously, it's a population of people like us that are answering the questions. Because, you know, right off the bat, it's like, how many of you are using LLMs in your company? And it's 61%, you know, which is a pretty high number right there. but I'm curious do you go through and like what constitutes using an LLM can it be as simple as pulling up a prompt for chat GPT and you know posing prompts to it or does it need to be like putting in your company data right there you go oh my god everyone owns it now yeah oh that's so funny I would love to talk about that for a minute too but The, yeah, it was, there was only a few people that were just like, oh yeah, I'm just using chat GPT and I'm going directly to open AI's website.
37:36That was not the majority. I think it was like one person, if I remember correctly. And don't quote me on that because it's been probably two months since we put this out. And so I can't remember anything since we put it out, But it's more people that are trying to set up systems with LLMs. And so these systems may be the API calls to OpenAI, or they may be hosting their own open source LLMs. Right. And what's unique or what questions are you intrigued to find the answers to in the upcoming survey? Maybe there's some commonalities that you'd like to see you know, carry through the surveys. But what are you most curious about going forward into this next round?
38:27I'm really fascinated by how many people are using open source versus using open AI. And one of the most hated visuals of the whole report is this one on like page 10. And it's talking about who's using OpenAI and what size they are. And people were like, this does not explain anything. What you're trying to say and the visual do not match up at all. And so again, I am not Gardner. I am one random guy who has never written a report, never did a survey. But for some reason, I felt compelled to do this three months ago, five months ago now. And I almost regret it. But seeing the final product, I'm very happy that I did it.
39:13Yeah, you should be proud. There's always things to look back on. It's a nice report. It's a good one. Exactly. And we were able to move fast on it. And so I'm sure Gardner is going to put something out soon. But that's the beauty of the community that we can move a little bit faster. Anyway, back to this visual that people did not like. And I mean, like a lot of people did not like it. So the whole idea was, are you using OpenAI? and we found that there's a bit of a correlation between people that are in super small startups like zero to 50 zero i mean one to 50 and so it's an autonomous startup yeah exactly it's so starting it's so startupy so anyway the one to 50 range they're not using open ai and the 1 ,000 plus are not using OpenAI.
40:15But the 500 to 1 ,000, and the numbers that I'm throwing out here, I didn't preface this because I got into the startup thing and the zero to 50, the amount of employees that you have at your company. So if it's one to 50, then you're not using OpenAI. If it's 1 ,000 plus, you're not using it. At least that's the preliminary data that we saw. And if you're in the middle, then you are. And so we had some theories about this and it was like, hmm, I wonder if it's because if you're a startup, you think that you can create a moat by not using it. And maybe your whole business is around using some kind of LLM and creating some kind of difference than OpenAI.
41:04And then if you're a larger company, A, this was before the enterprise scam that they've got going on, but that's, again, we can get into that in a minute. And so if you're a larger company, you probably have resources to figure it out yourself and you don't necessarily need to use OpenAI. And you're probably less comfortable with your data going outside of your walled garden. I think it's the latter. Yeah. If you're in the middle, it's like, let's just go as fast as we can. And so I want to see if that theory, if that holds up in the next one. Yeah. If you're over a thousand people, you have a legal department and that's what's inhibiting.
41:47Good point, Chris. But now, I mean, Chris, you tell me, man, like, do you think that this enterprise play is going to work out? Do you think people are going to trust them? In terms of open AI's enterprise. I do think it will work out from a business standpoint for them. I get dangerously close to some conflict of interest here, so I'm going to pass on this one. You're going to plead the fifth. I'm pleading the fifth. I'm pleading the fifth. So I'm backing away from this question. It sounds like you're not on the enterprise hype train, Demetrios. I work for a company that's a thousand plus is what I'm saying.
42:30So I'm going to back away from that question. He's got a legal department. We have a legal department and some of them might listen to this podcast. Exactly. I just wonder, I mean, I've been memeing with friends about it. And I think it's kind of funny how they do say, they specifically call out, we're not going to use your data to train any of our models. We're not going to know about any of your data. and I think that the biggest question is like, oh yeah, because they have the best track record of doing what they say. There is a healthy skepticism, I would say, among large companies on that, definitely.
43:10Yeah, well, also, I don't know, this is kind of avoiding the question a little bit, but I think that it is related in that you brought up the leaderboards earlier, Demetrius, And I think any company that like goes all in, it's almost like a new version of like vendor lock in like we used to talk about where like now you have like model family lock in where, hey, you know, these models, they're good. Like I'm no doubt, you know, GPT models really good. Are they going to be the models that are going to be best for your use case, either in terms of like output or in terms of the other things that are highlighted in your survey, right?
43:58Like latency and resources and like a lot of these practicalities and how you can control them and all this stuff. You know what they need, don't you? Yeah, exactly. And so I think there is an element here of like, hey, do I want to go all in on a single model family? Or is my strategy play more to have a bit more of model agnostic approach where I can pivot between different models for different uses? Maybe fine tune when I need to. But even if I don't fine tune, I have a lot of potential options to use and in a privacy conserving way. So yeah, I think that that's another element. However that works out, it will work out.
44:42But I think there's this kind of side element here, which is how the model landscape is evolving versus a single model family is evolving, which is good to highlight. So true. And actually, you know, it's funny. like I don't want to say that there is not an immense amount of value in chat GPT and GPT-4 because one thing has become very clear after interviewing a ton of people who are using large language models in production they are able to get up and running and proving value with their LLM so quickly. And I was just talking to Thibaut, I'm going to have to check how I pronounce his name. He's French and it's spelled very different than how it's pronounced.
45:35And he's running the LLMs at AngelList. And I asked him, hey, so are you worried about that vendor lock-in type thing because you're only on OpenAI? Have you messed around with even just Anthropic or Cohere? And I thought his answer was fascinating because he told me, look, you know what? There are so many other pieces of surface area that I would like to cover, so many other features that I would like to implement with these LLMs and be able to use in our product that if I'm stuck on one feature and trying to figure out what the best model is, then that's going to slow me down. I just want to go and get as many features plugged in as possible.
46:24And I know that ChatGPT works really well. And so I'm just going to go as hard as I can and incorporate these features because I have a laundry list of them that I want to do. And then once I get all of that out of the way, then I can start going back and saying, OK, let's figure out, should we bring a model in-house? Should we use Claude or something else? You know, it's really interesting to hear you say that because that's such a startup mentality. You know, we just have to run really fast and get as much done as we possibly can in the shortest possible time. And then you get to that large organization thing and they're worried about the lock-in.
47:02They're worried about where their data is going and they go much slower, you know, as a result of that. It's almost like an inverted, you know, approach based on size of the company and maturity. And that's what I was trying to show in this horribly positioned visual graph that you see on page 10 of the report. So we've finally got to the conclusion that the graph, we all agree on the point that the graph is showing. And to leave the graph out of it, I think that that's good. Yeah. Don't look at a graph. I'm sorry, he's laughing because we can all see each other even though this is audio only and he's laughing at himself.
47:40Not to belabor the point, but I think you all are exactly right. And I think I see this because almost every lead that's coming into Prediction Guard, just in terms of where people are at, regardless of whether they're a good fit for what we're doing or not. But almost every lead that's coming in, it's almost laughably predictable that they say, hey, we've prototyped out something very quick with open AI that shows like there's huge value here. Now what do we do? It's almost every conversation is starting like that. So yeah, I think that there's even like a temp, you know, we can make your graph, I think we should make your graph like sideways and add like a temporal element, make it 3D over time.
48:28People like that more, where like, you know, at the beginning of a project, I think a lot of people are doing that, whether they're authorized to do it or not in their organization. And then they get to that point where like, how do we scale that up, especially if we're in this larger organization environment? Yeah, it's like, oh, I got to go present this to the C-suite. We got to erase any use of sending data outside of our company. We can't tell anybody about that. Yeah, exactly. As we're coming close to the end here of our and friends episode, which I hope is only the second of many times we'll get to hear from you on the show, Demetrius.
49:09as you're looking to the next, I don't even feel like we could go to the next year. As you're looking to the next couple months of AI life, what are you hyped about? Just generally across the industry, what positive trends are you seeing that give you hope for where things are headed? Let's see. That is a great question. Leading question, but great question. It's got to be the positive side, huh? We got to stay. End on a positive note. You can dip though before you get there if you want to though, just for fun. Because I want to hear what he has to say. Well, for those that are just listening, I am wearing tie dye.
49:50So it is all peace and love here. And I am excited because right now, anyone who wants to mess around with machine learning and AI, they can. I have seen so many scenarios where a product person has said, you know what, let's try and throw some AI with this. And they've been able to create an enormous amount of value for their company by adding some features that just call ChatGPT or Claude or whatever. And that for me is really enticing because the barrier to entry has just been destroyed. You know, it almost was like the last couple of years, the first couple of years of the MLOps community.
50:41We were, I mean, we still are, there's still a lot of traditional machine learning where I put it in quotation marks because it's not that old, but there's a lot of really hard stuff happening with quote quote unquote, traditional machine learning. But with the advent of LLMs, a lot of that has become really easy. And so all these NLP tasks that were really hard up until 2022, they're not as hard. And so you're seeing the creativity of people being able to put LLMs into their products. I love that. That is something that now I didn't realize how much I enjoyed product and the idea of speaking with product people until ChatGPT came out because now I'm like, oh, I want to talk to more product people.
51:39The product owners are really great people to talk to because they have these wild ideas and they know how to figure out if this is actually a success or not. So there's that piece and dovetailing on that. So we're having another LLMs in production conference on October 3rd. I don't know if I could told you guys this, but definitely come and I will explain why I'll try and sell it to you as much as possible right now. But I wanted to create as many talks from product people as possible. And some of the stuff that we are talking about are like how to build an economical LLM solution and how to prioritize LLM use cases and how to put LLMs into your product.
52:27So those are very much for the product owners and the product engineers. And it's because of that. It's really like I got really excited now that this whole space has been opened up to the product owners. And so can I tell you what you can expect in the conference? Of course. Sure. I'm just going to tell you why it is the greatest conference on the internet right now, because where else, and Daniel can attest to this, all right? Live music interludes. Yes. Where else can you prompt me, not an LLM, you can put in the chat what you want me to sing about and I will sing it in real time just improvising on my guitar and last time we had a whole song about catastrophic forgetting and for those who do not know that is where actually I'm not even sure I fully understand what happens but basically like when you fine-tune going it's all that damn fine-tuning when you fine-tune sometimes a model will forget something because it's new data is replaced with, uh, it replaces the old data, but the catastrophic forgetting song was a hit.
53:42You've also got some semi-illegal betting going on during the breaks and you win swag. Uh, and then semi-illegal betting. There's a gray area, Chris, It depends on what country you're in. All right. And so that you also can expect. I mean, I'm just not sure that I've seen any conference that goes into the amount of technical details that we go into. And I want to highlight this piece. We can end with this. it is really hard, but it is very important for me to have a fully diverse field of speakers. And so I cannot tell you how much work it is. And it frustrates me now that I've looked at other conferences and I see it's almost like, oh, these organizers were lazy, you know, because there are amazing people out there from underrepresented groups that are doing some incredible stuff, but you almost have to look a little more because they're not necessarily on the conference circuit.
54:57They're busy shipping. Like, let's be honest, they're not out there talking about it. They're actually just doing it. And so I've had to look really hard, but I am very excited about the speakers that we have and the diversity of our speakers. And I think that that's probably like, out of all the things I'm proudest about, that's probably what I'm the most proud about. That's awesome. Yeah, I can't wait. And you said it was October 3rd. Is that right? October 3rd. Awesome. I'm going to be so we have one sponsor for this event. And they rented a whole studio in Amsterdam and Amsterdam is like four hours from where I live.
55:41And so hopefully, you know, everything goes all right. I don't like eat any mushrooms or anything or smoke too much weed and not show up for the actual event. But in case that does happen, we have shirts. Did you see the shirts, Daniel? I think I've only seen the ones I hallucinate more than chat GPT. That is it. I may live that in Amsterdam. You never know. Oh, how can I get one of those? Yeah, I'll share the link with you. We've got special. They only pop up for sale during the conferences. So you can't get them now, but hopefully about a week before the conference starts, we'll start selling them again.
56:22And yeah, that's it. That's what I've been up to. Not too much, you know, just trying to stay relevant. Keeping things chill. yeah yeah that's awesome well thanks so much uh demetrios for joining us uh i hope that we see you again here in some number of months don't uh don't stay away too long and um we'll look forward to to hearing about the results of the survey the the events and i'm sure all the like 15 new things that you're doing next time around that you're not doing this time around totally dude I got so many ideas I have so many I mean yeah I really appreciate you guys letting me come on here and rant about fine-tuning and talk about the cool stuff we're doing in the MLOps community and I love what you all are doing so thank you it was fun thanks we'll see you soon
57:22thank you for listening to Practical AI Your next step is to subscribe now, if you haven't already. And if you're a longtime listener of the show, help us reach more people by sharing Practical AI with your friends and colleagues. Thanks once again to Fastly and Fly for partnering with us to bring you all Change Talk podcasts. Check out what they're up to at Fastly.com and Fly.io. And to our Beat Freakin' residents, Breakmaster Cylinder, for continuously cranking out the best beats in the biz. That's all for now. We'll talk to you again next time. Thank you.
From the publisher
In this episode we welcome back our good friend Demetrios from the MLOps Community to discuss fine-tuning vs. retrieval augmented generation. Along the way, we also chat about OpenAI Enterprise, results from the MLOps Community LLM survey, and the orchestration and evaluation of generative AI workloads.
Changelog++ members save 1 minute on this episode because they made the ads disappear. Join today!
Sponsors:
- Fastly – Our bandwidth partner. Fastly powers fast, secure, and scalable digital experiences. Move beyond your content delivery network to their powerful edge cloud platform. Learn more at fastly.com
- Fly.io – The home of Changelog.com — Deploy your apps and databases close to your users. In minutes you can run your Ruby, Go, Node, Deno, Python, or Elixir app (and databases!) all over the world. No ops required. Learn more at fly.io/changelog and check out the speedrun in their docs.
- Typesense – Lightning fast, globally distributed Search-as-a-Service that runs in memory. You literally can’t get any faster!
Featuring:
- Demetrios Brinkmann – X
- Chris Benson – Website, GitHub, LinkedIn, X
- Daniel Whitenack – Website, GitHub, X
Show Notes:
Something missing or broken? PRs welcome!




