Apple Intelligence & Advanced RAG

25 Jun 2024 · 45 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Practical AI Podcast - Episode Summary: Apple Intelligence & Advanced RAG

Podcast Overview

  • Title: Practical AI
  • Description: Focuses on making Artificial Intelligence practical, productive, and accessible. Engages technology professionals, business people, students, and expert guests in discussions about AI, Machine Learning, Deep Learning, and more.

Episode Details

  • Episode Title: Apple Intelligence & Advanced RAG
  • Description: Daniel Whitenack and Chris Benson discuss the current state of AI in enterprises, explore Apple's recent AI announcements, and delve into the topic of Advanced Retrieval-Augmented Generation (RAG).

Key Discussions

  1. State of AI in Enterprises
  2. General Sentiment:
  3. Daniel and Chris express optimism about the developments in AI but acknowledge the challenges faced in real-world implementations.
  4. There is a shift from hype to practical applications as organizations navigate AI adoption.
  • Adoption Reality:
  • Many companies are cautious about implementing generative AI due to costs and data privacy concerns.
  • Organizations are considering a mix of open-source models and commercial API calls to optimize performance and control.
  1. Apple Intelligence Announcement
  2. Initial Impressions:
  3. The hosts share mixed feelings about Apple's recent AI initiatives, recognizing the company's shift from being a leader in innovation to being slower in their AI developments compared to competitors.
  • Key Features:
  • Apple's approach emphasizes AI as a feature integrated into devices rather than the product itself. This positions them differently from competitors focused on AI as a standalone product.
  • Users will have the option to integrate OpenAI's GPT capabilities with Apple's ecosystem while retaining control over their data.
  1. Advanced Retrieval-Augmented Generation (RAG)
  2. Introduction to RAG:
  3. RAG uses external data to enrich the context provided to language models, enhancing their responses beyond their inherent capabilities.
  • Common Pitfalls:
  • Many organizations get stuck using naive RAG approaches, which can limit effectiveness. The episode encourages moving beyond basic implementations.

Advanced Techniques Discussed

  • Context Enrichment:
  • Use surrounding chunks of data (previous and next) along with the relevant piece to improve context.
  • Two-Level Search:
  • First, summarize documents to find which document is most relevant, then perform a second retrieval within that document.
  • Hybrid Search Approach:
  • Combining traditional keyword search with vector-based search to improve candidate relevance.
  • Re-ranking:
  • Utilizing models to rescore and reorder candidate documents after initial retrieval.
  • Query Transformation:
  • Modifying user queries for better retrieval outcomes.

Key Takeaways

  • Navigating AI Implementation:
  • AI adoption is complex, and organizations must weigh the benefits of generative models against privacy and operational concerns.
  • Companies are increasingly using mixed models (open-source and proprietary) for productivity and control.
  • Apple's Strategy:
  • Apple's emphasis on integrating AI features into their products offers a fresh perspective against the backdrop of competitors focusing on AI as the main product.
  • RAG Enhancements:
  • Understanding and implementing advanced techniques for RAG can significantly improve the effectiveness of AI systems.
  • Organizations need to continuously explore improvements to avoid stagnation in their AI capabilities.

Featured Guests

  • Chris Benson: Principal AI Research Engineer at Lockheed Martin
  • Daniel Whitenack: Founder and CEO at Prediction Guard

Show Notes & Resources

  • Links to explore:
  • [Apple Intelligence](https://www.apple.com/apple-intelligence)
  • [Apple's AI Features Announcement](https://www.apple.com/newsroom/2024/06/introducing-apple-intelligence-for-iphone-ipad-and-mac)
  • [Hybrid Search Techniques](https://blog.lancedb.com/hybrid-search-combining-bm25-and-semantic-search-for-better-results-with-lan-1358038fe7e6)
  • [Advanced RAG with HyDE](https://blog.lancedb.com/advanced-rag-precise-zero-shot-dense-retrieval-with-hyde-0946c54dfdcb)

Conclusion The episode encapsulates the current landscape of AI in enterprises, Apple’s strategic approach to integrating AI, and the practicalities of enhancing generative AI systems through advanced RAG techniques. It encourages listeners to explore beyond basic implementations to harness the full potential of AI technologies.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:05Welcome to Practical AI. If you work in artificial intelligence, aspire to, or are curious how AI-related tech is changing the world, this is the show for you. Thank you to our partners at Fly.io, the home of changelog.com. Fly transforms containers into micro VMs that run on their hardware in 30 plus regions on six continents. So you can launch your app near your users. Learn more at Fly.io.

0:46Hey, friends. Do you remember the day that ChatGPT launched? I do. Well, not the exact day, but the time frame. It felt like the LLM was this magical tool out of the box. However, the more you use it, the more you realize that's just not the case. And as AI developers yourself, you know. The technology is brilliant, but it's prone to issues like hallucination. And that's not good, but there's still hope. Feed the LLM reliable current data, ground it in the right data and context. Then it can meet the right connections and give the right answers. The team at Neo4j has been exploring how to get results by pairing LLMs with knowledge graphs and vector search.

1:28and to hear all about this check out the podcast episode about LLMs and knowledge graphs at graphstuff.fm they share their tips on retrieval methods prompt engineering and more do not miss it you'll also find a link in the show notes again graphstuff.fm that's G-R-A-P-H stuff S-T-U-F-F dot F-M

2:06Welcome to another fully connected episode of the Practical AI podcast. In these fully connected episodes, Chris and I keep you connected to a bunch of different things that are happening in the AI community and try to plug you in with some learning resources to help you level up your machine learning game. I'm Daniel Whitenack. I'm founder and CEO at Prediction Guard, where we're safeguarding private AI models. I'm joined, as always, by my co-host, Chris Benson, who is a principal AI research engineer at Lockheed Martin. How are you doing, Chris? Doing great, Daniel. So many things happening in the news.

2:44I was looking forward to a chance for us to finally... We've hit specific topics a lot lately, and I'm hoping we have a chance to jump in and just talk about all the stuff. How's your mind at these days in relation to AI? We haven't done a sort of general check-in on, both of us are probably, I think, know each other well enough that probably we're both fairly hopeful about things, looking forward and seeing many good things. But yeah, generally, how has this year looked for you in relation to how you thought it might look and or in your own work with this technology. How's your view of how the technology is shaping up changed or stayed the same?

3:28In the large, I think some of the things that we have kind of predicted, you know, are happening. A lot of the developments are somewhat predictable. Like if you do something with imagery, you'll probably get there with video and things like that. And we talked about, we made some predictions last year about that. And I think those types of things are playing out more or less in a broad scale, kind of how we would have expected. Would you agree with that in general? Yeah, yeah. I would say definitely multi-modality wise. Yes, we talked about that a lot. What about, I guess, both of us either are at or have interacted with friends of ours or colleagues at a variety of enterprise organizations.

4:09What do you think is the reality on the ground in terms of adoption of AI versus what's all in the news and the hype and that sort of thing in terms of the practicalities and actual adoption rate of generative AI versus kind of the things that we've always had with us, the machine learning, the data science types of projects. Yeah, I think it's interesting. There's a lot of reality checking happening this year, especially in the last few months. Everyone's been hit with so much gen AI marketing and just all the hype and everything with it. But we're also expected to get things done, you know, at work.

4:51And so trying to finally get past the hype and get stuff, which requires a lot of hard decisions to make. And so, you know, companies kind of going, well, do I, you know, I'm looking at the cost of one of the big providers, you know, with OpenAI being kind of the leader in those. And do I want to pay for that for everything? And thus, do I also want to send my data out? How can we do things with smaller models or other large language models that are open source and the across multiple companies as i'm watching people make these choices and we're having conversations about it offline and stuff there's an agony associated with trying to to navigate correctly and not end up in a bad position for your organization and stuff like that yeah and so i'm seeing a lot of well we're going to use some open source for this and we're going to use some API calls to commercial stuff for that and we're going to use some smaller models over here and how are we going to put them all together and we've talked about these issues across a bunch of different conversations even the last guest we had we had this conversation but I think people are really challenged with making it all work as I talk to people at different companies I don't necessarily see everyone doing it the same.

6:06There's enough variability to where we haven't arrived at the world of best practices yet, in my view. Yeah. One of the things that you highlighted is the sort of multi-model future people kind of spreading out their workloads across multiple model providers and open models. I think that's something that seems to be only increasing and will continue. You know, with all of this emphasis on large language models, generative AI, have you got a sense from data scientists that you're interacting with and or others that, you know, the actual day to day of data science teams is shifting or they're still just training their support vector machines and whatever time series forecasting and whatever those things might be?

6:56I think the thing, you know, to that point, they're shifting, but there's also the field is exploded out in the number of positions to support. And, you know, way back when I was young, which was, you know, back when the dinosaurs were roaming the earth, you had like software developers and they kind of had to do everything, you know, which was, is very reminiscent of how AI has been in previous years. And in the last couple of years, we've seen, you know, an explosion of, we first saw machine learning engineers, you know, beyond data science. And then you kept adding each title and position. And there are now there's UX people in AI concerns.

7:32And it feels very much like software exploded from the 1980s when I was a kid into young adulthood for me in the 90s and 2000s. And now, you know, we're seeing that same. It's very similar to me. I look at it and I go, I bit, you know, deja vu for me. so it's a maturing of the industry and people are starting to figure it out I think there's a recognition finally outside of the marketing and hype machine which goes hard and constant always I think for the worker bees like me there's a realization that it's part of software and we've talked about this for a long time that that needed to happen and so a little bit less hype and more about what can the models do and how do I combine them and what sizes and it's putting the jigsaw puzzle together of what makes value for a particular organization.

8:23And that's been interesting for me to be part of that in my own organization and help us navigate through the morass. And every other organization I'm talking to is doing the same. So, yeah. Yeah. I was just sort of curious about boots on the ground. What's changing day to day for data scientists? It seems like one of the things that you're indicating is it's more that roles and teams are expanding. They are. Yeah, versus the data science teams that have existed sort of cease to exist in their form, you know, creating scikit-learn models and start moving over to Gen AI, which is probably not the case.

9:03I was just looking at Google trends of terms, which is always a fun thing to look at. And I was looking at Gen.ai versus Scikit-Learn, and there's sort of a, you know, Scikit-Learn still quite impressive search. You can see this surge of interest kind of in the data science hype period, at least as far as I can tell. But then there's also been a surge since kind of 2022 and on, and, you know, that's gone down a little bit. So I don't know how much you can draw from that, but the data science team still lives as far as I can tell. I don't think it's going to die because people are also, to your point, they're also realizing the limitations and constraints of Gen.AI and what types of things it doesn't do well.

9:55And so people are, I think, a bit smarter about it in 2024 versus 2023 and definitely 2022. It seems like the wisdom is finally kind of spreading out and people will kind of go, instead of just saying Gen AI is going to solve everything, they're recognizing it for what it is and the capabilities it has. and they're starting to say, this is a good use case for it, but we need to pair it maybe with a reinforcement learning model. And they're starting to remember, oh yeah, there's all these other capabilities, which we were quite enamored with until Gen AI came along, and they're still really good technologies to use.

10:33And so trying to start recognizing what shouldn't be an AI at all and combining those together in unique value propositions for their organizations is the thing I'm seeing. One other point that I've noticed also for the first time this year is that companies like the software side and the AI side are finally really coming together operationally instead of being very stuck apart, which is one of the problems we've talked about on the show many times. and I'm seeing like agile methodologies play out that had been on the software side for years in these organizations and they're now including the AI and data science teams and how they're like, if they're, you know, I'm just making up one of the, like if they're using safe or scrum or whatever they're using, they're starting to account for that and it feels more real life to me.

11:24It feels like, ah, we're finally getting to a point of maturity and recognizing that all the pieces need to come to play or we need to be efficient in how we do that. So that's been my kind of enterprise observation of the last few months. Yeah. And I don't know, again, Google search trends only give you so much, but it seems like the main trend with data science as a function, at least according to searching, just sort of keeps going up pretty steadily, even though there's a switching of technologies. It would be nice, though. I know that we talked quite a while back in a number of episodes about being a full stack data scientist.

12:05And I know recently we had some of that discussion around full stack AI agent development. But that sort of idea that there would be more integration of that software side into data science teams and vice versa is something that maybe this is a push that's kind of materializing some of that. You know, it's interesting. The term full stack is so loaded in terms of how people perceive it. And it's a bigger thing if you're in a very small organization. And what it means there is, thank goodness I've got somebody who can handle all these things that we have gaps on because we don't have enough resource to go buy somebody in all these different areas.

12:44And so it's more meaningful in a smaller midsize organization based on the nature of the organization. You're going to see it a lot more job specific and role specific in the enterprise, which in my view is a good thing because you don't want to just put full stack this full stack into everyone because in a large enough organization that doesn't help with your efficiencies and stuff. But team wise, there could be more integration. There could be. And I think integration is really important. And so I think this is the first year where I have a little sense of actually seeing that out in the workplace.

13:38complex AI pipelines super fast. You can easily create AI pipelines using their node-based editor, iterate and deploy faster and more reliably than coding by hand without sacrificing control. Deployment is easy. Pipelines are live API endpoints. They eliminate the need for constant code redeployment and debugging by deploying complex AI pipelines as API endpoints. Team collaboration is easy too. Plum's declarative node-based editor enables you to build quickly while empowering non-technical roles to iterate on what you've done without breaking it. You can build advanced AI features, get structured output every time, transform data, and leverage validated JSON schema to create reliable, high-quality structured output.

14:25So Plum is built for builders. Early-stage product teams are using Plum to go from idea to validation in record time. To get started, go to useplumb.com. That's Plumb with a B, as in plumber, to request access today. That's U-S-E-P-L-U-M-B.com. Again, useplumb.com.

15:05Well, Chris, we have got to Apple Intelligence. Yes. This last cycle that we went through, everybody's got their AI play now. And I guess, of course, Apple had been in AI in one way or another. So it's not like they were totally absent. But we got the announcement about Apple Intelligence. So what do you think? First impressions? excited confused a mix what's your impression i'm always skeptical about everything when it comes out because of the hype machine as you know uh but as an apple user i'm looking forward to it i want to see what they do i buy a certain degree into the apple ecosystem but i also am not a hundred percent invested in every way uh the way some folks are i use google as well and i and i you know various things.

15:57So they were very slow. Apple's received a lot of criticism the last couple of years because once upon a time, having been perceived as the Steve Jobs-esque leader of we're the ones that bring you completely new ideas that are going to change your world, like the iPhone when it was released originally, they have definitely not been fulfilling that role. They've been slow. Having said that, having thrown the criticism out first, they have certainly, I would say the announcement seems differentiated in that Apple is kind of putting it out. They're a product-focused company and they've made these AI announcements that clearly position AI as a feature and not the product itself, which a lot of the other big companies are, it's almost AI is the product that they're trying to do.

16:45And so with the announcements at WWDC 2024, which is their developer conference every year, often called DubDub by insiders there. They are talking about AI in the context of the devices and of the tasking that their users are doing. And so I actually like that. I like, as you know, you and I both get absolutely inundated with reach out from startups and companies always promoting and hyping their AI product. And at least to see Apple talking about it being feature enhancement rather than the thing itself is good. It's a little bit of fresh air on that one. Yeah, I think that there's like when you have a little button that summarize or rewrite or whatever that button is, that's very much how I like the sort of first wave of pre chat GPT AI features, a lot of them that came out where, yeah, you just see the suggested text or you'd see something that makes sense.

17:43And UX wise, that's probably like you say, very fitting with Apple and their approach to things. I know that there was definitely some shade thrown by some, including Elon Musk, about the reliance on open AI in Apple intelligence. Did you see any of that? That's Elon Musk being Elon Musk. I think part of it is just a gambit for the spotlight at any moment, inserting himself into any spotlight. I won't go into the Elon. If you're listening, Elon, you can come on our show and steal the spotlight. You're welcome. although it might be an interesting interesting conversation there yeah he's he's a he's the same age i am more or less and uh yeah and so yeah i always kind of go i'm just trying to imagine when he does some of the things but anyway back to the apple bit without derailing on that one elon came out with the specific criticism of okay you're going to send everything off to the GPT API and that's a huge privacy breach and stuff like that.

18:48But Apple had already clearly in the announcement said every user on a per use basis will be given the option of do you want Siri will say, do you want to send this off to GPT for an answer? And the user on a per use basis can say, yes, I want to do that or no, I don't. And they were very explicit about that up front. So I'm like, you know, Elon, just if you're going to use an iPhone or an iPad, just say no, just say no and stop because that way you still have control. And that seems like a reasonable thing. I use the the open AI app all the time on my iPhone. And it's one of those things that's open all the time.

19:26And that is good enough for many cases. but there are times when I would certainly like to integrate that capability into my other activities on the iPhone in a more integrated way. And this gives me that opportunity. So Elon was saying only have the open AI app. And as a user myself, I say, no, give me the option. Sometimes I'm just going to have the open AI app there, but other times let me integrate it with my other activities. Apple's going to give me the choice on whether I want to do it. I'm happy with that. all good. So he's not speaking for me in that capacity. Yeah. It might be interesting to talk just for a second about whether it's Elon or, you know, there's certainly in a less public or memeing way, a lot of people out there that do have concerns about this sort of closed model providers and some people that are still blocked by using these.

20:23I'm wondering from your perspective, I have a few of my own thoughts, but like from a practitioner's perspective, just to kind of make it practical as we are on practical AI. What are those sort of tradeoffs, I guess, as you see it now with closed model providers or kind of using open models or some version of hosted open models in kind of the enterprise or development scenario versus kind of your personal device? Certainly, there's a direct-to-consumer type of angle to what we've discussed so far. But in terms of the practitioner themselves, we've touched on this occasionally, but I think it's probably good to continue touching on it occasionally because things are changing over time and changing very rapidly.

21:12So yeah, from your perspective, do you have any thoughts on that? Sure. I think that is an issue that every large organization is navigating because you have a certain amount of funding to support your operations. And almost everybody has some, you know, tie into one or more of the large commercial APIs. And it's a different context from a personal user. Like I talked about, you know, I'm paying my open AI monthly fee, and I use that all the time for a variety of different tasks. But in the enterprise, it's a bit different. There may be that capability, but I'm also seeing enterprises that are really concerned about their data going out, about their information going out.

21:54If it's not their own, it's their customers that they have. By using a public API that you're paying for where that data goes outside of your control, there is a huge concern and risk, not only about the immediate privacy concerns, but also about the liability and the legal concerns around that. Because most organizations have a mixture of their own and other organizations, a whole bunch of partnership agreements in large companies. Maybe you're okay with your data going out, but of the 50 partners that you have, it might be that 36 of them aren't too keen about their data that they have an agreement with you holding goes out as part of your data to a third party.

22:34So that makes it pretty challenging to use third-party APIs in a manner that everyone is comfortable with. So I'm really seeing a lot of open source models being internally hosted. there is still a lag because Google, OpenAI, Anthropic, they keep pushing the boundaries on what they're offering. And the open source community is not typically, you know, all the way to the, it's not just the model, but the services built around the model that makes it easy to use. So you have to kind of recreate that or use existing open source capabilities that are out there. And that requires effort. And there is funding and money spent on a good bit of money spent on that.

23:12But I would say that of the ring of people that I hang out with across multiple companies, that I'm seeing more of the internal hosting of models with the effort of trying to stay on top of current releases and monitor that as being the more widespread way. And there's a recognition that those models may not give you quite as good of an answer if it's a very expansive thing you're prompting on as a GPT model would. But that's okay in a lot of cases. It can get you through. And if you have multiple models to choose from and combine, then you can usually be very productive without kind of violating all those concerns that I enumerated.

23:50So it's both. But for me, I'm seeing more people turn in where now I run in a national security world. And so we might be a little bit more conservative about that across the various defense companies and stuff like that. And so acknowledging that there's a bit of a, you know, an array of possibilities there. Yeah, one of the things that maybe has shifted a little bit in my mind since the last time I've been considering this question and we've talked is, I would say there's all of the sort of privacy, data misuse, data leaving your network, all of that, I think, is a big piece of it. there's been kind of a developing mindset that I've kind of picked up on, which is slightly different.

24:34And I started to pick this up from, I've mentioned this a couple of times because it's been helpful for me, A16Z's recent surveying and reports. But the fact that, yes, there's a privacy element, but a lot of times organizations are using open models because of the control element. And it took me a while to, I think, fully parse through the implications of that. And I think some of what it gets down to is when you're connecting to one of these closed systems and there's hosted open models in a variety of places you could get. So when I say hosted closed system, I mean, like you literally don't know what's happening behind that API or how the model is called and that sort of thing.

25:22Those are productized AI systems, which means that they're making opinionated choices about how to improve the performance of that product surrounding the model, right? Yes. And that actually can be an amazing thing, right? Like open AI functionality is spectacular, like without doubt. These other systems, Anthropic, others, really spectacular functionality. But there's this element of it where they've made some opinionated choices for you about how to process the data that you're putting into that system before and after it hits the model. And so there's a lot more going in. And I think you see this come out very much in, for example, the stuff that happened with Gemini, where you put in your prompt to generate an image of American founding fathers.

26:16And there's clearly, you know, however that worked, a modification of your prompt or extra instructions to bias that output to look a certain way. Sure. If you're interested, go look it up. There's lots of interesting pictures. And to be fair, they've rectified that situation as far as I know. But when you have that sort of decision made, you don't have full control. Like it's not just your prompt going into the model and you kind of choosing how to govern or bias that or process user inputs or do your prompt templating. And so it can be really good, but it can be sort of frustrating at that level where you get like 80 to 90 % of the way towards what you want.

26:58And then for some reason, you just can't figure out why you can't get that last bit or you can't figure out why this error is happening or there's latency types of fluctuations or whatever those things might be. It could be bias in the output. So I think that that like opinionated product ties thing, like it's both a good and a bad. And depending on your scenario, that may actually be what you need, right? Like I'm not going to worry about these things. I trust the way that sort of these things are being handled internally in a system like this. And I'm guessing that will be fine for many people.

27:34But then there's people that want to build kind of these competitive AI features into what they're creating as a company. And they want full control to figure out like, you know, to build those in exactly the way they want to make sure that they can test those in exactly the way they want and to have that control element. I think that's way more crystallized in my mind than it was previously. I think that's a fantastic insight right there. And I think most people miss that because with the hype machine going, we have a habit of talking about the models themselves all the time and, you know, kind of as product.

28:10And therefore there's so much that these companies that are putting out these as a service are doing. There's so many humans involved that you never see. And yes, that can really make it much better in some ways because they're kind of shortcutting what the model may not be able to do on its own. They are shortcutting and greasing the skids to make you get what you want. But at the same time, anytime there's a human involved, you're going to have the bias as well. And they're trying to make it safe and controlled and not have some sort of thing that ends up in the news in a very negative way. And that puts constraints around it, just as what you knew.

28:46So yeah, I think it's really key that we look at it as not only a model, but model plus the services around it, whether you're building them or whether someone's building them for you.

29:09What's up, friends? I love Backblaze. I'm happy to have them as a sponsor. Backblaze makes backing up and accessing your data astonishingly easy. This is a service I personally use. Go to backblaze.com slash practical AI. You get unlimited cloud backups for Macs, PCs, businesses for just$99 a year. You can easily protect business data through a centrally managed admin, protect all the data on your machines automatically, easily deploy across multiple workstations with various deployment options. You can add on enterprise control, including granular access permissions, advanced single sign on group management controls and compliance support.

Read the full transcript

29:50They even offer multiple restore options, including rapid recovery in the event of data loss or ransomware. That sucks. You can access your backed up data from anywhere in the world using their web app or their iOS or Android app. You can even restore by mail. They'll give you a hard drive with all your data shipped to your door. You buy a hard drive restore, send the hard drive back within 30 days and get a full refund. Get one year file retention and version history. Over 55 billion with a B files restored for customers so far. Visit backblaze.com slash practical AI so they know where you came from and continue to support the show.

30:28This is a service obviously recommended by me, but also by New York Times, Inc. Magazine, Macworld, PCworld, LifeWire, Wired, Tom's Guide, 9to5Mac, and just so many more. You receive a fully featured or no risk trial at backblaze.com slash practical AI. Again, they're supporting the show. Go there, play with it, start protecting yourself from potential bad times. Start today.

31:13All right, Chris. Yes. Well, a lot of times in these fully connected episodes, we do try to kind of bring some learning element to the forefront as people are exploring these topics. One of the ones that has been coming up a lot for me, which we've talked a lot about on the show, is RAG or retrieval augmented generation. but we've only sort of talked about it at a surface level and at this sort of naive rag level, which might misrepresent sort of some of what people are doing with this approach under the hood. And this would like generally kind of be framed. I think if you search advanced rag, you'll find a whole bunch of articles.

31:57And really what's happening is there's a naive approach to this RAG type of workflow, which can get you to some really amazing results really quickly. But then when you have to kind of fine tune, improve that system, load in more data, use documents that are closely related one to another, various types of documents, there's a lot to dig into in terms of fine tuning that system and fine tuning both how the retrieval and the generation works. And there's a whole variety of sort of workflows that have been developed by the community that can help you improve your RAG setup. So that's kind of one of the things that I wanted to bring up here and maybe talk through a couple of those.

32:43I know our friend Demetrios, we talked to him. He's got some opinions about RAG versus fine tuning. That's one thing you'll hear. But yeah, I would love to dig into that if it sounds interesting to you. Absolutely. And I think you're the right person to do this, given that you're diving into this stuff all the time. Yeah, well, certainly this is RAG pipelines are probably the first thing that people are building with generative AI. And the idea is not too complicated. The idea is, let's say I have a bunch of documents that contain information relevant to questions or queries that I might have.

33:24instead of just asking an LLM to give me an answer or to do something which would rely on the probabilities of that model in generating its own text and what the data that was trained on you could get any sorts of answers out of that LLM rather than just relying on that I'm going to inject on the fly some of that external data that I have into the prompt to help answer the question or the query, something like that. So we do this all the time with ChatGPT and other systems, right? When I say, summarize this email for me and then I paste in an email. That's how I'm injecting data into the prompt to the model right at the time that I need it to run.

34:09So there's no fine tuning of the model here. It's just a strategic insertion of data when I'm prompting the model. And often this happens in like, oh, I have a bunch of developer documentation or onboarding materials for my company or a wiki or a bunch of webinars or a bunch of podcasts or whatever, and I want to answer questions out of that material, then these can be loaded into a vector database, which allows you to do the retrieval part. So to find the relevant chunk of information that's required to answer the question, and then you take that relevant chunk, insert it into a prompt, and then respond or let the LLM generate based on that given context.

34:54So that's sort of the naive RAG approach where you sort of have a user query, you find a single chunk of information in some repository of information, insert it into the prompt as context, and hopefully get an answer, which it is naive, but it surprisingly works amazingly well in many cases, right? It does. It's interesting as we get in from kind of the naive to the two advanced ideas in it. And you also just mentioned for a second fine tuning along the way. It is definitely the first step. I think it's easy to implement in general for people, which is why it's the first step. But I also think before you go on, I think a lot of organizations are getting stuck on naive rag and just kind of stopping there.

35:41And I've noticed that. So keep going. Yeah, yeah. I think you're totally right. And I think this is why I wanted to bring up this topic, because some will hit and they'll get like sort of OK performance out of their RAG system. But then they don't realize that there's more options to improve that system. Yeah, I've seen a lot of people thinking it solves it. Yeah. Like all we need is an LLM and then we're just going to give it the data for RAG that we're going to inject into it and we're done. And that's, I'm hoping, I'm hoping we can break some of that perspective over the next few minutes. Yeah.

36:14And the question would be like, well, okay, if you're getting some good answers and some not good answers, for example, what do you do to improve your RAG system? and that's where there's a whole variety of things to explore and like I say if you're really interested in this I'd recommend searching for advanced rag I'd recommend looking at the llama index blog the lance db blog there's a lot of really good content out there to help you parse through this but but let me kind of inspire with a few snippets of things that you could keep in mine. So the first I think is around context enrichment. This is a very simple thing that you can do where let's say you have a hundred documents and you split them up into little chunks, which you embed in the vector database and you search against those little chunks to find the relevant thing that might help you answer the question.

37:13Well, depending on how you chunked up that information, it might not give you all the context you need to answer. It might be in like the previous chunk, it might be in the next one, it might be in the one that you found, or it might be in a combination. And so this sort of idea of context enrichment might be that you just find that chunk that's relevant. And then instead of inserting just that chunk, insert that chunk plus the one before it and the one after it, for example, just expand it a little bit, enrich it a little bit. Another sort of common thing is maybe you want to pull the three most relevant chunks rather than the one most relevant and add more context there.

37:55So there's more that you can add in more than just a single chunk. The other sort of related methodology here before we get into maybe the more fancy stuff is actually doing a two-level search over your data. So if you think about it, let's say that I have, again, 100 documents. And there might be similar content across those documents. They might overlap in certain cases, but they're different documents. Well, if you take and summarize with an LLM each of those documents or pages of those documents, and then you also chunk it up into smaller chunks that you eventually want to use for your RAG, you could first search on the summary, which would kind of point you to the right document that's going to answer your question, and then do a second phase of retrieval within that document itself to pull out the relevant section, right?

38:53This helps you kind of hone in on the right document that you're using. So those are two fairly easy to implement in terms of how you set up your vector database and how you do your querying, but they can provide a boost. Now, there's more complicated things. You can get those to those in a second, but do those make sense? It does make sense. Yes. I'm just curious, two second question. When people are going, you know, in your first case there where they're just going for the answer and not adding the chunks around it, do you think that's kind of bias from traditional database operations where you find the answer and that's it?

39:30Or is that in my space? I was just wondering, as you were saying that, why people might be limiting themselves in that way? Yeah, I think it's a perception problem. It's just like there's so many examples out there of getting started with RAG. And that's all you kind of see, unless you kind of really dig in. So maybe it's just a perception issue. It is also like maybe that sort of holdover from how you would retrieve things from in a traditional database sense. Those might both factor in. There's another kind of hybrid or two level type of search that happens, and this is implemented in several different vector databases, even natively now, because it can be quite useful, is actually doing two levels of searching.

40:19But the first, which is a traditional full text search or keyword search, and then a vector comparison, rather than just relying on the vector comparison. So you kind of hone in on the full text kind of keywords first and then do a vector comparison. And you could even ensemble these in various ways and use one for re-ranking or ordering versus the other one. There's a variety of ways to implement this, but this would be kind of generally categorized as hybrid ways of searching, I think is most frequently termed. So there's the context enrichment, there's the hierarchical search or index retrieval.

41:01That's the kind of summary, then chunk. And then there's like the hybrid search, which would be actually using two different search methodologies. And notice all of this has to do with the retrieval part for the most part that we're talking about here, not mostly the LLM side, although you could use an LLM to generate the summaries for the hierarchical approach. So it's interesting that those TF-IDF, keyword searching, full text search sort of things are coming up again. So back to our original way we started this episode, the data science pieces still survive in many ways. It's still relevant there.

41:41And I don't think that's changing. Yeah. The last two that I'll highlight, one that comes up a lot that people will use is a method called re-ranking. So there's actually models out there known as cross encoders. And what happens is you might do a first level vector search to get a smaller number of candidate documents and then use maybe a more expensive model based approach to actually rescore the candidates that you pulled and reorder them. Hence the name re-ranking, reorder them or filter them to the most relevant document. So that's kind of the re-ranking approach. There's a couple of really interesting ones where you use an LLM in the loop.

42:26One of those called Hide, LanceDB has a good blog post about this, uses LLMs to generate sort of hypothetical documents that should answer this question. And then you kind of use those hypothetical documents in the retrieval. there's people that do also query transformation so they actually take the query in this kind of fits our previous discussion about modifying a prompt except now maybe you're in control of it where you take that prompt in and you actually regenerate the query such that it's more favorable to the retrieval task yep that was a lot and i know it was quick but i think it might be good for people to hear that and just kind of see that there's a much wider picture of these advanced rag techniques.

43:14And I didn't even get a chance to sort of get through all of them. People are exploring a lot of things, but that I think paints a much more rich picture of what can happen in these rag pipelines versus just that naive approach. Thank you very much for kind of bringing this to attention. I think it would be well advised for people to recognize they're kind of getting to first base with the typical rag approach and that's working for them in some cases quite well, but these tools are out there now where it's not so hard to then go on and move past that. But I'm seeing a lot of people get stuck there.

43:49So thank you for kind of covering that territory and giving people an opportunity if they're not familiar with it to maybe dive into this. Yeah, definitely. Well, it's been a fun one, Chris, to bring back some data science discussions into our podcast. And yeah, excited to see what's coming over the next couple of weeks we can catch up on soon. Absolutely. Talk to you later, David.

44:18All right. That is Practical AI for this week. Subscribe now. If you haven't already, head to practicalai.fm for all the ways. And join our free Slack team where you can hang out with Daniel, Chris, and the entire ChangeLog community. Sign up today at practicalai.fm slash community. Thanks again to our partners at fly.io, to our Beat Freakin' Residence, Breakmaster Cylinder, and to you for listening. We appreciate you spending time with us. That's all for now. We'll talk to you again next time.

From the publisher

Daniel & Chris engage in an impromptu discussion of the state of AI in the enterprise. Then they dive into the recent Apple Intelligence announcement to explore its implications. Finally, Daniel leads a deep dive into a new topic - Advanced RAG - covering everything you need to know to be practical & productive.

Join the discussion

Changelog++ members save 6 minutes on this episode because they made the ads disappear. Join today!

Sponsors:

  • Neo4j – Is your code getting dragged down by JOINs and long query times? The problem might be your database…Try simplifying the complex with graphs. Stop asking relational databases to do more than they were made for. Graphs work well for use cases with lots of data connections like supply chain, fraud detection, real-time analytics, and genAI. With Neo4j, you can code in your favorite programming language and against any driver. Plus, it’s easy to integrate into your tech stack. 
  • Plumb – Low-code AI pipeline builder that helps you build complex AI pipelines fast. Easily create AI pipelines using their node-based editor. Iterate and deploy faster and more reliably than coding by hand, without sacrificing control. 
  • Backblaze – Unlimited cloud backup for Macs, PCs, and businesses for just $99/year. Easily protect business data through a centrally managed admin. Protect all the data on your machines automatically. Easy to deploy across multiple workstations with various deployment options. 

Featuring:

Show Notes:

Something missing or broken? PRs welcome!

More from Practical AI

All 157 episodes
Apple Intelligence & Advanced RAGPractical AI · 45 min
Listen in VO