In short
Practical AI Podcast: Episode Summary - RAG Continues to Rise
Episode Overview In this episode of *Practical AI*, hosts Daniel Whitenack and Chris Benson engage in a lively discussion with Demetrios Brinkmann, a notable figure in the ML Ops community. The trio delves into the findings of the latest MLOps Community survey, highlighting trends in AI, particularly the emergence and rise of Retrieval-Augmented Generation (RAG) technologies. The episode also previews the upcoming AI Quality Conference.
Key Participants
- Demetrios Brinkmann - ML Ops Community Contributor
- Chris Benson - Principal AI Research Engineer at Lockheed Martin
- Daniel Whitenack - Founder and CEO at Prediction Guard
Episode Highlights
RAG vs Fine-tuning
- Demetrios notes a shift in preference from fine-tuning generative models to utilizing RAG, which enhances generative models by incorporating external data retrieval.
- The conversation highlights the emerging trend of "RAG as a Service," suggesting a growing market for RAG implementations.
Survey Insights
- The MLOps Community recently conducted a survey with significant participation (322 respondents), revealing:
- 45% utilized existing budgets for AI initiatives, while 43% allocated new budgets.
- Companies are exploring various use cases for AI, particularly in generative AI and RAG technologies.
- The majority of respondents identified themselves as intermediate users of RAG, with only 6% claiming to be on the cutting edge of RAG innovation.
Maturity of AI Workloads
- Discussion around the need for traditional ML workloads versus emerging generative AI use cases.
- There's recognition that while traditional ML tasks (e.g., fraud detection) persist, generative AI is increasingly employed for tasks such as transcription and code generation.
Challenges in Data Handling and Evaluation
- Participants discussed difficulties in data evaluation, citing that 72% of ground truth labels were manually marked by humans.
- Many organizations face challenges in rapidly iterating evaluations due to variable API latencies and the complexity of managing data sets.
The Future of AI Architectures
- There's speculation about the potential obsolescence of transformer models, with the panel discussing alternative architectures like neuromorphic computing.
- They ponder the limitations of transformer-based models and the need for innovative solutions to overcome current challenges in AI.
Community Fragmentation
- The hosts noted the fragmented nature of AI communities, where different groups focus on various tools and methods, leading to a lack of general best practices across the industry.
AI Quality Conference Preview
- The episode concluded with an enthusiastic preview of the AI Quality Conference, scheduled for June 25, highlighting interesting speakers and interactive sessions aimed at enhancing AI quality in production environments.
Key Takeaways
- RAG technology is gaining traction: Companies are increasingly moving towards RAG implementations, recognizing its advantages over traditional fine-tuning methods.
- Budget allocations are increasing: Organizations are willing to invest more in AI, signaling a growing maturity and acceptance of these technologies.
- Evaluation processes need improvement: The time-consuming nature of current evaluation methods is a significant challenge that needs addressing.
- Future architectures may evolve: As the AI field matures, there's a potential shift away from transformers towards more efficient architectures.
- Community collaboration is essential: Bridging the gap between different AI communities could facilitate the sharing of best practices and accelerate innovation.
Conclusion This engaging episode serves as a valuable resource for anyone interested in the evolving landscape of AI, particularly the rise of RAG technology and its implications for the future of machine learning and data handling. The insights shared by Demetrios, Chris, and Daniel underscore the importance of continuous learning and adaptation in the AI field.
For more information, check out the full episode and join the community discussion at [Practical AI](https://practicalai.fm).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:05Welcome to Practical AI. If you work in artificial intelligence, aspire to, or are curious how AI-related tech is changing the world, this is the show for you. Thank you to our partners at Fly.io, the home of changelog.com. Fly transforms containers into micro VMs that run on their hardware in 30 plus regions on six continents. So you can launch your app near your users. Learn more at Fly.io.
0:43Welcome to another episode of the Practical AI Podcast. This is Daniel Whitenack. I am founder and CEO at Prediction Guard. I'm joined as always by my co-host, Chris Benson, who is a principal AI research engineer at Lockheed Martin. How are you doing, Chris? Doing good, Daniel. How's it going today, man? It's going great. I just landed in Boston and you were texting me and you're like, hey, Dimitrios from the ML Ops community wants to hop on and record an episode. And I was like, I've got to get out of this train station. So I just found the nearest stop and got out and I don't have my normal setup.
1:22So I probably sound weird, but I was like, these are the best times when we get to have our friend Dimitrios on. How's it going, man? You guys can't get away from me. I blackmailed your boss into letting me come back on here. Standing invitation. Absolutely. What's up, man? How you doing? I'm so excited because I try not to abuse this standing invitation. I had a great time the last time I was on here. We had a bunch of laughs and a lot of people reached out to me because we had this rag versus fine tuning conversation. And I think things have kind of like the pieces have fallen, the cookies have crumbled.
2:01It feels like fine tuning is not as popular as it was. I don't know what you all are seeing out there, but rags are the go to in these days. Yeah, that's what I'm saying. Everyone's ragging each other all over the place. And I think, yeah, along with that, there's sort of like in my mind has developed this category of like, so there was rag, which of course you're augmenting the generative model with data, but not in the way that people typically think of fine tuning, you're just doing retrieval. But I think also these other sort of workflows around calling external tools or the neurosymbolic stuff that we're seeing of combining a rules-based algorithm or function or traditional quote unquote machine learning algorithm with generative models, maybe to get the inputs and then like connections to databases these all these different ways like it kind of seems like people are figuring out that generative models are great at being kind of assistants and automators but not necessarily like predictors right or or other kind of functions too like analytics types of things and that that sort of stuff i've noticed there's been a new term coined at least for me, rag as a service.
3:17Have you guys run across that? Rag as a service is now like a thing. Is that acronym just rag?
3:28Don't even try. You're just going to strain the vocal cords if you do that, man. Yeah. Rag as a service. We'll rag you. If you bring your thing, we're going to rag you. We're going to rag you all over the place. Well, I think what you're saying, Daniel, it really speaks to something that I've been seeing two, which is the maturity in the last, whatever, six months, it's become very clear that there's traditional ML workloads and use cases that are kind of going to always be traditional ML workloads. You think about like your fraud detection models or like the term prediction or the recommender system even.
4:04And then you have your generative AI workloads or use cases. And that's something like these like transcription or you have the LLMs, which are doing all sorts of stuff. But rags are probably the biggest ones in that where you get that copilot or code generation, I think, is a huge one. There's not like such a big overlap where you're saying, OK, the generative AI use cases or the generative AI models are going to dethrone the traditional models. I agree with that. But yeah, definitely different use cases. I think one thing we had our first Gen AI Mastery webinar with the podcast. When was that, Chris?
4:49A couple of weeks ago or something like that. We were talking about text to SQL specifically and that, you know, analytics like SQL is really good at doing analytics, right? Especially descriptive analytics and aggregations and all of that. And it just doesn't make any sense for you to take a big table and somehow figure out how to dump its contents in a prompt and have a model reason over it. Because it's probably going to get it wrong anyway. But also there's this existing tool which can be called, which is really awesome at doing those things. I love the idea and the exploration that's happening right now on how can we merge both of these worlds and how can we see what different parts work well together and which combination of the traditional ML plus the generative AI can go together.
5:43And I know that you do a series of surveys with the ML Ops community. I think the last time that we talked, we were talking about some of your survey work and some of the interesting findings, but I think you've gone through other iterations of this, right? So are you seeing interesting things pop up as the community around this technology matures? So 100%. And I just come on here when I got survey insights to share, I guess. That's what I'll be known for. Whenever we have a nice survey of AI, I'll come and share it with you all. But this one is cool because this time, so we did an evaluation survey and we launched it when we had our virtual conference, which had a huge turnout.
6:32It was over two days spanned. Which was awesome, by the way. It was great. You were part of one of the past ones, right? And Daniel, you had an awesome spot. But the two-day span, we tried something different and we said, what if we do two days? But since it's virtual, nobody has to fly anywhere. So instead of you trying to watch a live stream for eight hours a day on a Thursday and then a Friday, why don't we just do two Thursdays in a row? and then you don't have to feel like wow i just had 20 30 of my week eaten up by that eight hour live stream or 16 hour live stream now you can tune in tune out on a thursday of your choosing but we launched it there and the response has been amazing because normally maybe we'll get like 100 150 people that will fill out the survey this time we had 322 is the number of these which is super cool to see.
7:28And so let me give you some of the clear insights. One, there's budget being allocated towards AI these days. I don't think that's going to surprise anybody. The fascinating part is that like 45 % of the respondents said that they're using existing budget. And then a whole like whopping 43 % here said, no, we're using a whole new budget. So you've got exploration happening in generative AI like never before. But when it comes to that, one other takeaway has been that like the MLOps, AI, ML engineers, they're really trying to figure out what the biggest leverage use cases are and how they can explain that.
8:21And I think what we're seeing is there's a lot of companies that are open to the exploration right now. And they're open to letting people say, all right, cool. What is the most valuable for our teams and our company? Is it a chatbot that is an internal chatbot? Is it an external chatbot? What does that actually look like? What is the use case? It's kind of funny. We actually talked about that a little bit last week, Daniel and I did, in terms of trying to get non-technical people engaged in it. And I think that there are organizations all over the world right now that are doing exactly that, you know, to your point on the result that you're seeing.
9:00There's a lot of effort and a lot of money being thrown at how do we start doing that. And it's all really like the shining star here was just rags. Obviously, it's very clear everybody's using rags and the participants self-identified as being intermediate in rags. That was the majority. So we had like 31 % saying that we have some experience with LLMs and rags. And then you only had 6 % saying we are at the frontier of LLMs and rag model innovation.
9:50What's up, friends? There's a new book out there called The Hacker Mindset. This is a productivity cheat code to unlock new levels of success in your career, in your creative pursuits, and in your personal growth. This book is about leveraging the principles of white hat hacking and applying those skills to the broader world. It's available for pre-order right now and it's not your typical productivity guide. This is written by Garrett G., a seasoned white hat hacker with over 20 years of experience. This book reveals the secrets of hacking and how you can apply those skills to overcome obstacles and achieve your goals.
10:29So don't miss your chance to get ahead and get this book, The Hacker Mindset. You can pre-order your copy today at thehackermindset.com. Be among the first of many to tap into this power of hacking for your success. Join the movement and embrace a new way of thinking. Again, that's thehackermindset.com.
11:00So let me ask you guys a question. You get some opinions going here on that. Do you think, you know, with all of these assistants and chatbots and all the other kind of, you know, focus things in Gen.AI, with those using RAG, do you think those are for more kind of general use cases having some domain knowledge in? And do you think that maybe going back to the last time we had you on the show, when we were talking about fine tuning, do you think that kind of highly specialized fields will still stick with fine tuning instead of rag? Do you think, in other words, the degree of expertise required, if you will, to get a job done productively, do you think that makes a difference on whether or not you go rag or fine tuning?
11:45Or do you think it has nothing to do with that? I would love to hear what Daniel has to say in a minute, but I've heard something said about like fine tuning is for form. So if you're trying to get a different form on the output, or for example, if you're trying to get functions and the whole basically like GPT functions are a perfect example of this. So if you're trying to get a homegrown model to do that type of thing, then fine tuning makes sense. Otherwise, it's not necessarily the good call because unless you're using a very small model, I think is the thing there that, okay, what's the trade-off that you're doing?
12:31And the other thing that is probably worth talking about when it comes to the difficulty levels that I've seen is people that want to go and do use a small model, a small like domain specific model that is distilled or it is very fine tuned and distilled and it's on their own in their own infrastructure and in like they need a whole team to support that. That's like hard mode. You're playing on hard mode and you contrast that with just a GPT-4 call. That's a whole different level of the game that you're playing. And so I kind of look at it that way. How much are you willing to trade off when you're trying to figure out, is this the way forward?
13:19I would agree with that. I think the only thing I would add is there is still, at least as far as I've seen with our enterprise customers, an inclination that they need to fine tune. That's still a kind of general like, oh, we need to do that at some point. And I think once they kind of solve a few of their use cases without that, they kind of disillusion themselves of that notion in many cases. But the good thing, like you're saying, Demetrius, is like you can probably prove to yourself without a huge amount of effort using an easy to use API, whether or not you'd need fine tuning and do that in like a day versus like just immediately going to jump to fine tuning and how do we get GPUs and how are we like what model server are we going to use and all of that stuff.
14:10Like you say, that gets very much into hard mode very quickly and you don't sort of need that to validate your use case often. And even if you do fine tuning, sometimes it may be down the line when you've been running the pre-trained model for quite some time and you actually have a good prompt data set to fine-tune with because most people don't start out with that either. That's a huge point too and this is probably like the biggest unlock that happened is that okay now that anybody can use the OpenAI API you can quickly see if there's value in that crazy idea you have and then you can go down the line like, all right, now I'm going to use an open source model, which is turning the knob to something harder.
14:57Or maybe it's not even using an open source model. Maybe it's just like, can we get the same results with a smaller model from OpenAI? So instead of GPT-4, we're going with 3.5 Turbo. Can we then go to an open source model and get the same results? I almost look at it like a spectrum of how difficult you want to make your life and how much upkeep you're going to need and all of that. But as with everything, there are benefits if you go to that very small model, if you need it. And you just have to really play out and see if you actually are going to need it. Speaking of another survey, just to make it an increasingly survey-driven show, I don't know if you saw the...
15:41Survey says. Yeah, survey says. The Andresen Horowitz post that they did another sort of survey, which was kind of interesting. And part of what they drew out, I forget the range of participants that participated in the survey. We can link it in our show notes. I was just posted March 21st. We're recording this maybe a little more than a week later. And it was about enterprise. So 16 changes to the way enterprises are building and buying generative AI. And one of the things that they specifically highlight there was enterprises are moving towards a multi-model future. And specifically a multi-model future driven at least partially by open models.
16:29And so I think the other kind of interesting trend that you're seeing is they have like graph with like how many model providers are people using per company or whatever. And you see like three, four, five, like in many cases, and also a high adoption of open models. And I think what they're trying to draw out is like some of that is maybe driven by security, privacy things. But I think also it's driven by like control and flexibility. And once people start realizing, I also find it still a pretty big misunderstanding that people have that all of these models sort of behave the same. And in reality, pretty much every model has a character of its own and a specific behavior.
17:15And even just switching from open AI to an open model for like text to SQL, for example, a model that doesn't do other things well, but does that really well can prove to be really useful. Or maybe it's a specific language thing, like we're doing a Mandarin chat right now with one of our customers. And so whether it's language or whether it's task, people are, I think, finding out that their future is multi-model or multi-model provider or whatever, mainly because of that behavior thing, but also because they can have some control over when they use this model or when they use that model and kind of create the mix that's right for them.
18:00And that's kind of a way, it's like a route around fine tuning in some ways, because you can kind of assemble these reasoning chains, even with multiple models involved that do very specialized tasks. And that can kind of help you avoid spinning up and running up your own fine tune that does this very, you know, unique reasoning workflow. You can kind of bring all the experts and bring all the expert models and to help you. I can't help but wonder, you know, like that's a built in capability that you have at Prediction guard, but I think that there's a maturity issue there with a lot of organizations on getting to that point.
18:40And so what I'm seeing is very mature organizations are going exactly to that, having multi-models capabilities, and they have the ability to distinguish between which models they should use for which circumstances. I suspect, and maybe there's some survey data on this, but I suspect that that's still a fairly small group of even enterprise organizations that have gotten to that level of understanding of what they can do. And I think there's a spectrum falling off from there that the bulk of the world is in right now in terms of trying to figure out how to make it work for them. Let me twist some statistics to play in your favor.
19:18Hold on. I'm going to crunch some numbers. I love it. Live. You're probably not using an AI model to do it. No, I'm going to weave that narrative that you've got, Chris, and I'm going to go with it with the survey data as you're asking for. No, but actually the biggest question as you guys are talking about this that goes on in my head is, you know what engineers really do not like, and I think it makes them very anxious, is having a single point of failure. And so if you are relying on OpenAI's API and you have a lot riding on that, where does that all go if the CEO gets ousted? And so I imagine that a lot of people thought twice after there was that big drama that happened and people started thinking, you know what?
20:09Maybe we should try and have a few redundancy options just in case. Now you do have to have a bit more maturity to say, okay, I can't use the same prompt as I use always. So I have to have this prompt suite or prompt tests or prompt templates. And I think that's another thing that's happened since the last time we talked when we had the conference to try and get people to, well, prompt ops is one, agent ops is another one, but I created a song called Prompt Templates and I'll put it in the, maybe we could play it real fast.
21:34As someone who's in kind of the defense industry, I'm for agent ops because it sounds bad, doesn't it? Oh, I thought you were going to say you were for prompt songs, which would bring smiles to people in the defense industry. That's what I was hoping to let you do. Exactly. Missed out on that one. But he is, I don't know if people can see this, but he's got a wonderful shirt on. That is for sure. And being in the defense industry, I don't know how you can get away with that. Work from home. Work from home is how I get away with it. And for those that aren't watching his shirt, it says, I hallucinate more than chat GPT, which is a classic shirt.
22:14Which Demetrius sent to me, I have to say. So thank you very much. I love this shirt. I love that you wear it. That is what I'm very happy with. But the other thing that I wanted to mention about the survey, and then we can move on and keep talking about other stuff of the day, topical issues. But the data that people use and the data with which we evaluate the output, it seems like people just don't know what's going on there. We haven't figured that out yet. There's no consensus. It's not really clear. And like the classic data sets or the classic evaluation pieces that you use, they don't really hold up.
22:53So everyone's got to be having their own data that they've created and that they're testing against the output. But it's really hard to do that at scale, right? And it's really expensive also. So that's what we saw in the evaluation data or the evaluation survey data is that you've got to handpick these. You've got to match them up. And it's human curated. The testing data sets that you create. We had 42 % are using data that they've created as their data sets to evaluate if the model is working or not. Yeah, that's crazy. And you mentioned expensive. It could be monetary, but it could also be a sort of iteration time too.
23:39Typically, like when we were back when we were creating machine learning models, which lots of people still do because it's the thing driving basically all the predictive stuff. But, you know, you could run your model and evaluate it. Maybe it took a few seconds, but, you know, a couple minutes. But here, when you're thinking about running, especially against an API that has like variable latency or maybe the calls, each execution of your prompt chain is taking like 15 seconds. And even if you want to run that over like 100, 200, 300 example reference inputs, all of a sudden your iterations become really, really slow.
24:24That's something I've noticed that people kind of struggle with is really making their evaluation quick enough that they can iterate, feel like they can try a lot of things. even if they have like a big budget to try a lot of things this kind of iteration time is really frustrating also because maybe there's other people that are involved that aren't technical right and they don't want to think about like concurrency and python right they just want to go into an interface and like try some stuff and yeah so you've got all these things mixed together which make it a bit of chaos and in many cases that's so true the iteration speed the time the, I mean, we see here that this is crazy.
25:1172 % of ground truth labels were manually labeled by humans. And so to have to go and do that, and then also how often are you doing that? How, like there's so many questions and so many unknowns for what the best practice is. That's one thing that came up on the challenges is that it's just like a lot of people called out something that we synthesized into lack of guidance. Like nobody's saying this is the best practice. This is what we've seen works really well for us because maybe some people say, well, this worked well for us sometimes and you can try it and see if it works well for you too.
25:51That's kind of the state of the industry right now. There's something I want to tie back into something else that we've been talking about lately and that is the fragmented nature of the community. And it's another thing Daniel and I have talked about recently is that we do have communities, but we have multiple communities. And in many cases, not in your case, but in many cases, they're very platform dependent, vendor specific. And it makes it, compared to a lot of programming languages, it makes it harder for people to come in and find specific best practices. So I'm actually not at all surprised to hear the survey kind of playing that out.
26:26I think that that's kind of a natural fallout of the challenges that we're having with community in general. So in my understanding this correctly, it's like because a lot of the communities are being built around certain tools, you have the best practices for those tools, but not necessarily for the industry. You can't generalize those best practices. Yeah. And I think also the different channels through which people are communicating kind of naturally develop their own bias. I don't mean bias in a bad way necessarily, but just the bias towards like emphasizing certain things, right? Like you get into the news research community.
27:05We had a great conversation about that. And like people are talking about, oh, like we're doing all this like activation hacking and representation engineering. And like that's but that's like not really talked about. Like if you're over here in the Llama Index Discord or LanceDB Discord or like whatever. And some of that's driven by the focus on what those tools do, but also like where people are coming from. Right. And more of the indie hacker building app sort of stuff or the rigorous like academic. side or the like enterprise, like I really just want to get something into production side. There's like all these different slants people are coming from.
27:47That is so fascinating to think about like how each of these communities has their main focus. And since it is, there is so much surface area and there's so many areas that you can go, different areas to explore, that each community is exploring their own area. And if you go into that, you can tap into what people are talking about in that area versus if you go into another community, you get, oh, well, what's going on here? What's the focus of this community? This is a different outcome from what we've seen. If you step out specifically of kind of the AI, ML world, and you look at more just computer science, computer programming, communities out there, there's usually kind of a place to go and you kind of learn the same sets of skills and values around that.
28:41And that's a little bit different from this. It's been one of the challenges, I think, that the AI ML world has struggled with a little bit. So I'm not at all, like I said, I think your survey captured that essence. Thank you for sharing that with me because I'm going to steal it and I'm going to say it a bunch. hopefully you don't mind you didn't put a trademark say it all you want it's a great insight i've seen it just in the ml ops community right we have people that are really trying to productionize ai and so what people in there are talking about is really like pragmatic and practical how can i get this being used in my company so that i can either save money or make money.
29:25Like money is the ultimate metric there. If you go into, as you were mentioning, like these different communities, if you go into the Lama Index community, there's a lot of talk of rags. And actually we had Jerry on in the conference and he showed this slide that I thought was so incredibly done. It wasn't him that did it. I can't remember the person who created it, but it was like the 11 ways that rags fail. And so it had all these different ways that you need to be aware of. And I think one that's coming to light that people are seeing is so important is how you need to get that retrieval evaluation correct.
Read the full transcript
30:03Because if you're not retrieving the right thing from the VectorDB, then it doesn't matter what you give or what the output of the LLM is. If you give it some kind of crap, then it's not going to give you anything there. And And the other piece that I think is fascinating is that, like, how do you make sure that all this data that you've got in the VectorDB is up to date? And so we've talked about this a bunch. And again, this is in the MLOps community. We're very industry focused. And how can we make sure that we are productionizing this? So in a production scenario, you've got your HR chatbot that is using a RAG system.
30:43and you say, all right, cool, we've updated the vacation policy. So we went from a European vacation policy to an American vacation policy. And you've got Daniel over here saying, all right, HR chatbot, like how many days of vacation do I have? How do you make sure that everywhere in the vector database, it now is updated to the American vacation policy? And so, okay, cool. In the vector database, maybe you say, you know what, we were able to scrub everything, or we just pull from the most recent documents. But then you were a good engineer and you made sure to pull in a bunch of different data sources.
31:21So in Slack, turns out that you're grabbing some data from that and people talk about how it still is the European vacation policy. And now Daniel's been quoted of having 30 days of vacation when really he only has two. That's unfortunate. Yeah. Actually, this is a conversation we just had the other day with a customer, because also, at least some of these databases, they have, depending what you go with, if you go, you know, with a plugin to an existing database, maybe there's kind of more traditional updating and upserting sort of functionality. But some of these is just like, put a document in, get it out, delete it.
32:04There has to be a layer of logic on top of these that actually help you do some of that. So in their case, it was like, oh, we want to take in all the articles that we've had on our web on this website, and that's going to be it. And then they're like, well, what if we update those? Do we just like blow everything away and redo it? And I think my ended up my answer, like with the amount that they had, I was like, probably like if you can have something running in the background, honestly, that's probably the safest thing for you and it's going to take like a couple hours or something but then at least you the going into the making sure everything's synced up is and in that case like they could just version the files of the embedded database but yeah it's an interesting set of problems it is a fun one and also you know what i'd love to explore too is the idea of like RBAC or role-based access control.
33:02How are you seeing people go through that and do it well? Because that feels like another one that can be really misused. So for RAG is one thing. For text to SQL, some of that maybe can be kind of nice because if you're embedding some function in an application that already has RBAC on the database, then you could use that credential and hopefully that carries through. But for the vector database side, we've interacted with people that have maybe like an internal chat and an external chat where the external chat is a subset or should use a subset of the documents from your internal chat. So in that case, you sort of have like two, it's bifurcated rather easily.
33:51And, you know, that's like somewhat easy to deal with because you could just have like two tables or two collections, whatever that is in the vector database and kind of merge the retrieval or use them selectively in certain ways. But as soon as then you have like many, many different roles or even user specific things, I don't know, like many vector databases that would be, however you manage that would be transparent to that vector database. So you'd have to somehow manage the metadata associated with it. But there may be certain people we'll have to follow up. Chris, we haven't had a Muta on for a while, but they're always thinking about these role-based access to really sensitive and private data.
34:34I'm sure there's people doing advanced things, but in terms of the main tooling that people are just grabbing off of the shelf, a lot of that logic is just absent. Exactly. Yeah, I want to hear if anybody is doing RBAC and they've figured it out. that's one thing i'm fascinated with because it is a very again going to the community that i i run with that's something that productionizing kind of it comes hand in hand with that yeah and it could also have to do with the guardrails that you put around the large language model calls because if it's like a public facing chat or something like that that you may want to filter out pii or like prompt injections may be a very important thing versus like internally, ideally you trust people as long as you know how the data is flowing.
35:29Like there might not be as many restrictions in terms of what can go in or who's accessing things and that sort of thing. But yeah, it's interesting.
35:53This is a Changelog News Break. Pierre-Carl Longleus announcing the release of Common Corpus on Hugging Face. Quote, contrary to what most large AI companies claim, the release of Common Corpus aims to show it is possible to train large language models on fully open and reproducible corpus without using copyright content. This is only an initial part of what we have collected so far, in part due to the lengthy process of copyright duration verification. In the following weeks and months, we'll continue to publish many additional datasets also coming from other open sources, such as open data or open science, end quote.
36:34Here is more info about this massive dataset. Common Corpus is the largest public domain dataset released for training LLMs. Common Corpus includes 500 billion words from a wide diversity of cultural heritage initiatives. Common Corpus is multilingual and the largest corpus to date in English, French, Dutch, Spanish, German, and Italian. Common Corpus shows it is possible to train fully open LLMs on sources without copyright concerns. You just heard one of our five top stories from Monday's Changelog News. Subscribe to the podcast to get all of the week's top stories and pop your email address in at changelog.com slash news to also receive our free companion email with even more developer news worth your attention.
37:22Once again, that's changelog.com slash news.
37:30so in addition to this sort of evaluation stuff we've spent a lot of time talking about data and evaluation and retrieval what about on the on the model side do you think we'll ever escape the world of transformers mitros so this is something i've been thinking about a ton man and i've got some thoughts on this that like is everything that we're doing now in AI a band-aid because transformers just aren't the right tool for the job have you guys thought about like one big workaround yeah exactly am I crazy to think that I don't think so actually I was talking to one of our customers about this like they have so much like logic around double checking the outputs of models or formatting the outputs of models.
38:24And I'm talking hundreds and hundreds and hundreds of lines of code, thousands of lines of code, I don't know, written all around this sort of workaround. And it's because they're using general purpose model that you sort of have to massage into how you want it to behave. Is it a little bit ironic that you use a rag to clean up the problems with transformers? Is that what we're saying here? Oh, I get it. The rag. What we need is the Lysol wipes. There you go. But I often wonder, are we having to over-engineer this because the core of the problem, it's like we're trying to put a Band-Aid on something instead of going and fixing the root of the problem.
39:12And right now it feels like there's nothing out there that can even stand a chance against the transformer architecture. So, of course, we can't say, well, I would rather use XYZ. But I just get the feeling like when we think about AI in 2024, or like the chat GPT AI era, we're probably going to be laughing at the whole idea of transformers. If in 10 years we're looking at that, it's going to be like, yeah, okay. Transformers were great, but they were a stepping stone. I know that there's quite a bit of research going on in general about doing different types of architectures. I know that there's a number of organizations that have been testing alternatives to transformers in the last couple of years, but I don't think anyone's gotten there.
40:07Or if they have, then you should reach out to us and let us know so that we can be talking about it here on these podcasts. So I think there's a lot of folks out there that are really wondering what's next because we're essentially taking one superset of architecture and we're doing everything we could possibly do with it. And every big step forward in the last few years has been around what else can we do with this architecture. So at some point, I agree with you, Demetrius, something's going to give and we've got to try some new approaches in there. Yeah, that's what it feels like to me. It's just like, what's the next step?
40:44And I would love to also hear from whoever, if there's something that feels like it's promising. It's really exciting to me. I don't know enough about that. That's very much the research community that I don't get to spend a lot of time in. And I'm sure there's a bunch of false flags and people get excited about something. And then it turns out that after you throw a bunch of GPU at it, it doesn't work out like we thought it would or like we saw a promise, but it didn't actually work out when it holds up to scale. So I understand that right now we're in the era of transformers. I wonder how long we're going to stay in this era.
41:24Not only around specific architectures in that capacity, but almost new approaches. For the first time in a while, neuromorphic computing is really rising again as a topic of interest. And it's not there yet. You're talking about architectures, both on the hardware and the software side, that are not specific to either transformers or even GPUs underlying it. But it's been interesting to see the maturity that's developing. You talked about the exposure to research. Even for me, that's the same case is that you have all the pure researchers out there, but now we're starting to see them expand out in lots of ways and trying completely different approaches.
42:06And I'm pretty excited that we're going to start seeing some interesting results over the next few years as people are looking for alternatives across both hardware and software architectures. I think we're pretty close to a turning point. Can you break down real fast? What was that big word you just used? Anthromorphic? What was that? I can't even say it. I got tongue tied. Neuromorphic computing, I think is what you're talking about. Neuromorphic computing. That is a big word. What does that even mean? I don't know. I got to Google that real fast. So, and I am the last person on the face of the earth that should be trying to explain neuromorphic computing.
42:44I put you on the spot. But having been exposed to that, the short version is almost like, you know, in the earlier days of AI, and people would say, you know, in the marketing, people would talk about, oh, mimicking, you know, the neocortex, you know, the human brain and stuff like that. And we all kind of, as this GPU and transformer-based architectures, we're like, well, it's not really like the human brain. Well, the neuromorphic architecture is actually that. It's the legitimately like, how does that, the architecture of a brain, and I'm saying this like, there's probably neuromorphic computing scientists out there listening to me now going, oh my God, somebody take his mic away.
43:27That's a terrible explanation. But in my fairly primitive understanding, that's kind of where it is. How do neurons really work in real life and how do you do compute artificially in that capacity? But I know that there's definite interest in doing that. I know Daniel has a relationship with Intel through PredictionGuard. And I know Intel has an interest in that field. I think they're one of the leaders in it. I Googled it, Intel's all over the first page, or I perplexity that it was all cited from Intel. That is very true. I would hesitate to say it out loud because I'm probably wrong, but they may very well be the global leader in that space right now.
44:09Yeah, makes sense. Well, that is awesome. I'm glad that you taught me about that. I appreciate you for teaching me. Neuromorphic. Now I can say it properly and everything. Well, you know what? now that you say that, we're going to have to have a show on neuromorphic computing coming up pretty soon. Yeah, exactly. Let's get down into it. I want to listen to that for sure. We'll dive into that. Daniel can reach out to his contacts there. Oh, that's classic. Well, dude, thank you very much for coming on as we wind up here. It's always a pleasure. Anyone who has been listening to the show long knows that you join us regularly on the show.
44:46It's always special for us. We have a great time with you. So thanks for coming on today. We will get the show notes for the survey and some of the other topics that you brought up today so people can join. And folks, if you haven't gotten into the MLOps community podcast that Demetrius hosts, you definitely need to check that out. It is an awesome podcast, highly recommended by both myself and Daniel. So hope people join you over there. Oh, and can I also plug, we're going to have an in-person conference and I'm really excited about that. A little bit shaking in my boots because June 25th, it's going to be our first in-person conference ever.
45:26And it's going to be all about AI quality. And we've got some super cool speakers coming. We managed to get the CTO of Cruise to come and talk about what they've done since their little mishap in regards to making sure that their AI is quality. we've also i mean there's so many great people you can go to ai quality conference.com and we'll throw the link there in the show notes too i'm very excited for it but the speakers are going to be awesome the attendees are going to be amazing i think what i'm most excited for though is that we're going to have all kinds of fun random stuff you can imagine it's going to be a conference but it's probably going to be more like a festival i may have people riding around in tricycles, giving out coffee, or we'll have a little DJ area, or a jam band breakout room, a bunch of Legos hanging around.
46:22I don't know yet. So if anybody has any ideas on how we can make it absolutely unforgettable, I would love to hear about that too. And I'm going to throw out one last plug for you, is that when you say that, I believe you, because I know that you've heard me say this when we were off the air, but just in case anyone doesn't know this, Demetrios is the funniest guy in the entire AI world and does hilarious things. If you don't follow him on social media, you are missing some really great, great content. Anyway, just wanted to say that people should show up at the conference just to see what you're doing.
47:01If no other reason, even aside from the cool content you have, they'll enjoy it. Thanks for coming back. There's going to be great speakers. You're going to learn a ton. but there's also going to be some really random stuff that you're going to be like, what is going on here? And hopefully you really enjoy it because that's kind of what I'm going for. Okay. Well, thanks a lot, man. I'll talk to you next time. Likewise. See ya.
47:31All right. That is Practical AI for this week. Subscribe now. Now, if you haven't already, head to practicalai.fm for all the ways. And join our free Slack team where you can hang out with Daniel, Chris, and the entire ChangeLog community. Sign up today at practicalai.fm slash community. Thanks again to our partners at fly.io, to our beat-freaking-residents, Breakmaster Cylinder, and to you for listening. We appreciate you spending time with us. That's all for now. We'll talk to you again next time.
From the publisher
Daniel & Chris delight in conversation with “the funniest guy in AI”, Demetrios Brinkmann. Together they explore the results of the MLOps Community’s latest survey. They also preview the upcoming AI Quality Conference.
Changelog++ members save 4 minutes on this episode because they made the ads disappear. Join today!
Sponsors:
- The Hacker Mindset – “The Hacker Mindset” written by Garrett Gee, a seasoned white hat hacker with over 20 years of experience, is available for pre-order now. This book reveals the secrets of white hat hacking and how you can apply them to overcome obstacles and achieve your goals. In a world where hacking often gets a bad rap, this book shows you the white hat side – the side focused on innovation, problem-solving, and ethical principles.
- Changelog News – A podcast+newsletter combo that’s brief, entertaining & always on-point. Subscribe today.
- Fly.io – The home of Changelog.com — Deploy your apps and databases close to your users. In minutes you can run your Ruby, Go, Node, Deno, Python, or Elixir app (and databases!) all over the world. No ops required. Learn more at fly.io/changelog and check out the speedrun in their docs.
Featuring:
- Demetrios Brinkmann – X
- Chris Benson – Website, GitHub, LinkedIn, X
- Daniel Whitenack – Website, GitHub, X
Show Notes:
- MLOps Community
- AI Quality Conference
- Evaluation Survey
- RAG failover talk from Jerry Lui
- Prompt Templates the Song
Something missing or broken? PRs welcome!




