In short
Software Engineering Daily: DeepMind’s RAG System
Episode Overview This podcast episode features Animesh Chatterji, a software engineer, and Ivan Solovyev, a product manager at Google DeepMind. They discuss the complexities and advantages of the Retrieval-Augmented Generation (RAG) system, focusing on the newly released File Search tool integrated into the Gemini API.
---
Key Concepts
What is RAG?
- Retrieval-Augmented Generation (RAG) is a method that combines retrieval of relevant documents from a knowledge base and generating responses using advanced language models.
- RAG systems are foundational for building production AI systems but can be complex and costly to implement.
Challenges in Deploying RAG
- Managing vector databases.
- Developing effective chunking strategies.
- Implementing embedding models and indexing infrastructure.
- Keeping up with evolving techniques and best practices.
The File Search Tool
- A fully managed RAG system within the Gemini API that abstracts the retrieval process.
- Allows for easy uploading of documents, code, etc., to create embeddings and query knowledge bases.
- Designed with an emphasis on simplicity and accessibility for developers.
Benefits of File Search
- Simplicity: Reduces complexity in configuration and setup; no need for database management.
- Pricing Transparency:
- Charges for indexing (initial upload and embedding processing).
- Charges for tokens used in queries.
- Eliminates fees for storage and additional components.
---
Discussion Highlights
Evolution of RAG
- RAG has maintained its relevance despite the emergence of new AI models and techniques.
- Improvements in context size for LLMs (Large Language Models) enhance RAG's functionality.
- RAG is especially beneficial for enterprise-level use cases with large datasets, allowing efficient processing without complex infrastructure.
Enhancements in Embedding Models
- Recent advancements in embedding models improve retrieval quality and understanding of context across multiple languages.
- Techniques such as "refrag" involve embedding chunks dynamically to enhance retrieval accuracy.
RAG vs. Advanced AI Agents
- RAG is not a competitor to AI agents but rather a complementary tool that enhances their capabilities.
- AI agents benefit from RAG's ability to retrieve relevant content from specified corpora.
---
Technical Details
Retrieval Mechanics
- The File Search tool uses semantic vector-based search.
- Queries are processed by embedding the user input and comparing it against indexed chunks.
- Developers can manage document versions and maintain relevancy through API calls.
Handling Updates
- New versions of documents can be indexed while retaining the previous versions.
- Developers have the choice to manage old content manually, ensuring relevance remains intact.
Future Directions
- Plans for incorporating multi-modal capabilities (text, images, and videos).
- Continued efforts to improve retrieval latency and quality while expanding support for structured data.
---
Use Cases
- Beam: A game development platform using File Search to assist new developers by providing contextual documentation and code examples.
- Performance Metrics: Retrieval latency is in line with model latency (around a few seconds) with a retrieval accuracy rate reaching up to 85%.
---
Conclusion The episode concludes with the excitement surrounding the File Search tool's adoption and its potential to streamline RAG implementations for developers. Animesh and Ivan emphasize the importance of feedback and the ongoing improvements to the system.
---
Additional Resources
- Full episode: [DeepMind’s RAG System with Animesh Chatterji and Ivan Solovyev](https://softwareengineeringdaily.com/2026/03/12/deepminds-rag-system-with-animesh-chatterji-and-ivan-solovyev/)
- Explore the Gemini API for more about RAG systems.
---
This markdown file captures the essence of the podcast episode, outlining its critical discussions and insights into the evolving landscape of AI and RAG systems.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOIntroduction to Retrieval Augmented Generation (RAG)
0:00 to 0:28
Learn about the foundational role of RAG in production AI systems.
“Retrieval Augmented Generation, or RAG, has become a foundational approach to building production AI systems.”
Google DeepMind's File Search Tool
0:28 to 0:56
Discover how the File Search tool simplifies RAG implementation for developers.
“a fully managed RAG system built directly into the Gemini API.”
The Problem Addressed by File Search
1:26 to 2:15
Understand what problems the File Search tool is designed to solve.
“This episode is hosted by Sean Falconer.”
Simplicity and Accessibility of RAG
2:15 to 3:35
Learn how the File Search tool prioritizes ease of use and pricing transparency.
“So, you know, we're talking about this product, which you mentioned to your file, the Fire Search tool for Gemini APIs.”
Evolution and Relevance of RAG
3:35 to 5:45
Explore the evolving landscape of RAG and its enduring significance.
“If you compare it to what's available on the market, first of all, the pricing is usually fairly complex.”
Improving RAG Techniques and Approaches
5:45 to 7:33
Discuss advancements in RAG techniques and their real-world applications.
“Rack been there from the very beginning.”
RAG vs. Other Tools
7:33 to 9:27
Understand the relationship between RAG and other AI tools in the market.
“Also, the embedding models themselves have improved, right?”
Chunking and Embedding in RAG
9:27 to 11:24
Learn about the chunking and embedding processes in RAG systems.
“I think that it's kind of a misunderstanding of what RAG is to frame it in that way.”
Files Supported by File Search
11:24 to 14:00
Discover the types of files that can be indexed through File Search.
“What is that spot in terms of the number of chunks to return?”
Indexing and Multi-Modal Support
14:00 to 16:50
Learn about the challenges and strategies for indexing various types of files and incorporating multi-modal data into the system.
“You are processing very well-structured data, specific tables, specific graphs that our system does not yet recognize well.”
Show all 19 chapters
Document Updates and Indexing Process
19:19 to 22:35
Understand how document updates work within the indexing system and the importance of managing document versions.
“Fidelity is an equal opportunity employer.”
Citations and Source Mapping
22:35 to 25:25
Learn about the mechanism behind citing sources in generated responses and how to map them back to original chunks.
“Is that something that you see as a future direction for this?”
Beam's Use Case and Retrieval Performance
25:25 to 28:00
Explore how Beam uses file search to assist new developers and the performance metrics of the retrieval system.
“I think their use case is pretty neat and simple at the same time.”
Improving Model Quality in RAG Systems
28:00 to 29:00
Learn about the importance of model quality and retrieval techniques in RAG systems.
“Those apply to any tool that we have trained Gemini with.”
Challenges of Fine-Tuning Models
29:00 to 30:40
Discover the limitations and challenges of fine-tuning models for specific use cases.
“In fact, I mean, that just adds more complexity.”
Recent Innovations in Embedding Models
30:40 to 32:20
Explore recent innovations in embedding models that enhance performance and retrieval quality.
“With that said, we do see people using fine-tuning in specific use cases.”
Multimodal Capabilities and Future Improvements
32:20 to 34:30
Understand the advancements in multimodal capabilities and their implications for RAG systems.
“So I do believe that a lot of this additional configuration that is happening around those models will go away and you will be able to just embed the thing, hit the search and get the results that are really relevant.”
Migrating to File Search Tools
34:30 to 36:20
Learn how to migrate to file search tools while leveraging existing infrastructure.
“related queries, but there are certain languages where certainly we could do more.”
Getting Started with File Search
36:20 to 40:00
Find out how to start using file search tools and the resources available for developers.
“loading a small portion of their database and just running evals, comparing both systems.”
Transcript
Automatic transcript. May contain errors.0:00Animesh Chatterji:Retrieval Augmented Generation, or RAG, has become a foundational approach to building production AI systems. However, deploying RAG in practice can be complex and costly. Developers typically have to manage vector databases, chunking strategies, embedding models, and indexing infrastructure. Designing effective RAG systems is also a moving target, as techniques and best practices evolve in step with rapidly advancing language models. Google DeepMine recently released the File Search tool, a fully managed RAG system built directly into the Gemini API. File Search abstracts away the retrieval pipeline, allowing developers to upload documents, code, and other text data, automatically generate embeddings, and query their knowledge base.
0:47Animesh Chatterji:We wanted to understand how the DeepMine team designed a general-purpose RAG system that maintains high retrieval quality. Animesh Chatterjit is a software engineer at Google DeepMind, and Ivan Solovyev is a product manager at DeepMind, and they worked on the file search tool. They joined the podcast with Sean Falconer to discuss the evolution of RAG, why simplicity and pricing transparency matter, how embedding models have improved retrieval quality, the trade-offs between configurability and ease of use, and what's next for multimodal retrieval across text, images, and beyond. This episode is hosted by Sean Falconer.
1:28Animesh Chatterji:Check the show notes for more information on Sean's work and where to find him.
1:45Animesh Chatterji:Yvonne and Animesh, welcome to the show.
1:48Ivan Solovyev:Hi, Sean. It's a pleasure to be here.
1:50Animesh Chatterji:Awesome. Well, why don't we, you know, we have two guests today, just so everyone can kind of learn whose voice is who. Why don't we start off with, Ivan, we'll start with you. You know, who are you and what do you do? Yeah, my name is Ivan. I'm product manager for File Search on Gemini API. Great.
2:07Ivan Solovyev:And Animesh, you? Hi, I'm Animesh. I'm the engineering lead on File Search. Awesome.
2:13Animesh Chatterji:Well, thanks both for being here. So, you know, we're talking about this product, which you mentioned to your file, the Fire Search tool for Gemini APIs. And before we get too deep into, I think, sort of some general things around like AI, RAG, agents, and so on, can you talk a little bit about what the File Search tool is and what problem it tries to address? Maybe, Ivan, you can take that. Yeah, absolutely. So File Search tool, first of all, is an integrated Rack solution that makes it super easy for you to take lots and lots of data, text, PDFs, codes, whatever you have, upload it into Gemini and start asking questions about your data.
2:51Animesh Chatterji:There are plenty of Rack pipelines available on the market, like we have Vertex, Rack Engine, and there are other providers who do support this feature. So it's nothing new. And what we focused on in file search in particular is accessibility and simplicity of use. We made some opinionated decisions. We removed a lot of complexity in terms of configuration, setup. You don't need to set up your database. You don't need to set up your infrastructure. The tool is just there. You just upload your data and you're going to use it right away. So we believe this simplicity is something that can help a lot of developers get started and overcome the complexity of setting up their own pipeline.
3:34Animesh Chatterji:The other big aspect that we're actually proud of is how we price the whole product. If you compare it to what's available on the market, first of all, the pricing is usually fairly complex. There are multiple components that are coming into the picture. You are paying for storage, you are paying for inference, you are paying for indexing, yada, yada, yada. What we did was we decided to simplify the whole model. We removed most of the things that you are paying for. And we focused on two simple aspects. So first of all, you are paying for indexing. So whenever you upload the file, we do need to do a lot of complex processing.
4:07Animesh Chatterji:We need to do embeddings. So you'll pay for that. And after that, whenever you do a query to Gemini, you're just paying for tokens. Obviously, there's going to be some addition from file search, adding data into the context. But that's it. You're not paying for storage. You're not paying for anything else. Why make that change around pricing? Is it primarily to just really try to simplify things for the users of this? Why was it needed to kind of, I don't know, buck the trend of how perhaps people used to have been paying for RAG in the past? Yeah, I think in Gemini API and AI Studio in general, we are aiming a lot for simplicity.
4:42Animesh Chatterji:And we do hear a lot of feedback from developers that it's hard to deal with lots and lots of products. It's hard to deal with different billing models. and be like cycles and how the whole cost is being calculated. So we do see it as a decent improvement over other products. And the price is actually much cheaper. So it's a big competitive advantage. Okay. Can you talk a little bit about, I guess, like the evolution of RAG? Obviously, like RAG was kind of the buzzword of the moment, I would say, like a couple of years ago. Now there's also, I think, been some things out in the zeitgeist of like, do we need RAG still?
5:21Animesh Chatterji:we have agents and RAG is dead. Do you hear all this kind of stuff? Like, I guess, like, where do you stand on that? And also, can you talk a little bit about sort of the history and how maybe things are approached to even RAG and the way that we use it has changed during that time? Let me talk through what I think, like where we are with RAG and maybe Animesh can chime in on the history of development of this feature. In terms of where we are, I think RAG is a fundamental capability. Rack been there from the very beginning. Whenever those models get really popular in use, it was always a staple whenever you wanted to process the data.
5:58Animesh Chatterji:And hype cycles were going up and down. We see this with different features related to LLMs. But I feel that the Rack was always there, and it was always useful to some extent. In the latest years, we saw improvements in the context size, obviously available to LLMs. And this does help a lot with use cases with limited data sets. And we do see much better quality whenever you try to do simple retrieval tasks on a small data set that fits into the context. And we usually do recommend to use that approach. However, whenever you start doing any enterprise use cases, whenever you have a huge code base, whenever you have a large file sets, like any legal documentation, anything like that, having RAC becomes very, very beneficial.
6:44Animesh Chatterji:First of all, you can work with the whole database without building the complicated pipeline or infrastructure to actually juggle the data in and out of the context. You can work with the whole data set. The costs become much better with Rack. If you put everything into the context, it becomes expensive very, very fast. And especially if you're using the higher tier models like Pro models, it's getting expensive. And with file search or other RAC solutions, you are able to actually use this cost. And for the large databases and the large enterprise use cases, this actually adds up fairly, fairly quickly.
7:22Animesh Chatterji:And Animesh, in terms of history, has our techniques and approach to RAC changed over the last couple of years? Has that evolved as well?
7:30Ivan Solovyev:Yes, I think the use case has evolved. so i think to your question whether the long context models have straightened the proposition of rag i would say that in fact they have encouraged rag to cover even more use cases like there are a lot of edo use cases where people want to upload documents for their entire semester and there is now even more focus on making rag efficient because even the long context model like retrievals are working fine we still see sometimes this context rot or lost in the middle syndrome where models are not great at retrieving data which is kind of in the middle of the context so now there are new techniques on making sure that how can we improve even the chunks that we are feeding into the model from rag so there was a recent paper last year which is called a refrag where what they are trying to do is instead of passing the chunks as it is to the model they are trying to embed the chunks and give all these embeddings to the model and let the model decide which of these embeddings may seem more interesting and only expand those cases so So yeah, from the kind of initial vanilla rag where we kind of find everything, give it to the model, we are trying to make that more smart by figuring out which chunks to give.
8:39Ivan Solovyev:Also, the embedding models themselves have improved, right? The way we are able to represent data that has significantly improved. So now we are better at understanding the context. We are doing better in languages other than English. So internationalization has also picked up. So there are these different parts using which we can say that RAG as a product is improving.
9:02Animesh Chatterji:And would you say that it's kind of like the wrong, I don't know, like question or view to take that around, like tool use versus RAG. Are they really competitors or are they more like these are collaborators in some sense?
9:14Ivan Solovyev:I mean, RAG is in some sense, a tool that we give it to the model, right? In case this needs more information from your specific private corpus, this is the way to go. But yeah, I wouldn't really see them as competitors.
9:26Animesh Chatterji:Yeah, absolutely. I mean, I agree as well. I think that it's kind of a misunderstanding of what RAG is to frame it in that way. But I think it is something that you see out there in the wilds of, I don't know, Twitter sphere and so forth. But that's a little bit of the wild west of AI in some sense.
9:43Ivan Solovyev:Yeah. In fact, I would say RAG is coming in popular in different ways now. We recently are kind of public previewing the personalization, which enables the model to give more context about your persona. And the way to enable is, again, something like RAG, where you figure out relevant chunks and give it to the model. And then it can understand your persona better and answer queries in that context.
10:07Animesh Chatterji:Can you talk a little bit about how that works in terms of being able to determine what are the right chunks to feed in the model? And how do you reduce the error rate of essentially identifying incorrect chunks?
10:19Ivan Solovyev:I think what we do is when you provide the data we chunk it and we embed it using the latest Gemini embedding model. And then we basically index it internally. And then at the time of the query, when the user provides the query, we again embed it using the same embedding model and then try to figure out the relevant chunks from that corpus that the user has uploaded. And then we have some knobs on which embedding model to use or how many chunks do we want to retrieve and pass it to the model. And we have run a bunch of evals to kind of find the sweet spot in terms of latency of how many chunks we want to retrieve like versus the quality we see and there are some knobs that the users can provide in terms of like how do they want to chunk the data but mostly it's the default settings that we have iterated upon we have tweaked the SI to make sure that the model actually triggers this tool when it actually feels that it is necessary to get more context and it's not unnecessarily triggering so yeah the entire suite of tools we have kind of evolved and evolved through to make sure that we provide the right default settings.
11:23Ivan Solovyev:And there are some capabilities that the users can override. What is that spot in terms of the number of chunks to return? So I think it's like in low double digits right now. And we have kind of kept it open. We don't document how many chunks we want. But yeah, it's not too many at this point of time.
11:45Animesh Chatterji:Is there some use case dependency on that? Or can you actually have sort of more of a universal approach to this? So right now we are going with the one solution fits all approach because we want to keep it simple.
11:58Ivan Solovyev:As and when we hear use cases of customers who feel the need that they want more of these chunks retrieved, it's easy to expose that as an option in the API. We don't want to do that right now. But if needed, we could do that. if the threshold at which you want to retrieve the chunks or the number of chunks you want to feed to the model, those are all things we could tweak around.
12:18Animesh Chatterji:Ivan, you were adding something? Yeah, so far what we saw from the partners integration is that the default configuration actually fits most of the use cases. We do have people doing search of illegal documents. We do have people doing searches over their code databases to provide relevant guidance for code completion and such. And in all of those cases, somewhere around five chunks returned in the response from file search was doing fairly well for them. And then in terms of, you know, I think historically, if you look at how people have approached RAG, there's a lot of people who want to really exert a lot of control over things like, you know, chunk overlap, chunk size, various settings.
12:59Animesh Chatterji:So I guess by abstracting away a lot of that retrieval pipeline, like how do you sort of balance that? Is it that you're targeting a specific type of use case or specific type of user? Or have you really figured out the secret sauce of the right collection of those things that's just going to work for people out of the box? I think most of the quality actually comes from the embedding model. So it's like you should think about this as like 80 % of quality is embeddings, 20 % is your configuration. So as long as we have the best embedding models, which we believe we do, the rest is less relevant to the quality of the outcome.
13:36Animesh Chatterji:So we do believe that for most people, playing with those configurations will not yield significant improvement, and the time is better spent in building their own pipelines. So that's where we focus on. At the same time, we never say, like, don't use any other RAC pipelines. We actually say, like, file search is the simplest tool. You should try the first thing. You should work for the majority of people. But if you really need the configurability, let's say your use case is very, very complex. You are processing very well-structured data, specific tables, specific graphs that our system does not yet recognize well.
14:13Animesh Chatterji:In that case, you may want to adjust all the little knobs that are coming with the more complicated pipelines. And what kind of files are you capable of indexing? I would love to say all of them, but we are indexing text files mostly. So PDFs, docs, code files, anything with text. We are currently doing OCR on images. So we're not fully ignoring images within PDFs and other files, but we are going through the OCR system. We're extracting text out of them and putting that into the context as well. And we are actually working on getting the multi-model support in as well. So we want to support native image processing, video processing, and at some point native audio as well.
14:59Animesh Chatterji:Gemini models are pretty good at reasoning on top of image and video data. So we want to have this ritual capability to actually find the relevant images and put them into context so the Gemini model can see them and act on them. So even if you're processing text, there's lots of different types of text files. You could have code could be a text file. You could have markdown. You can have documents that have not just images, but like tables and so forth. So are you able to dynamically figure out like the chunking strategy on behalf of essentially the user? Or does it matter? Like, do you have to use a different strategy for say, basically breaking up code to be able to find the relevant chunks versus something like, I don't know, legal document?
15:43Ivan Solovyev:So far, we have not done anything majorly different across these different types of documents. And based on the customer feedback so far, things like code are working fine. And we see that in some cases where there are like graphs or tables, we have sometimes seen that the chunking strategy, like the default chunking strategy doesn't work. We are working on techniques to make sure we represent this data in a more structured way and we can provide it to the model without breaking that structured context. But yeah, that is something being in the works.
16:13Animesh Chatterji:Yeah. And in a lot of cases, it is about chunking. But if you look at the structured data that is not just plain text, like parsing tables and graphs, that's where we see some regressions in terms of quality. But the way we address this is not through different chunking. It's mostly through pre-processing the data, making sure that the columns and rows in the table are aligned well when the data is represented to the model as text. So this kind of pre-processing is, I think, more important to get the quality right. And I guess it also going back to like what you're saying, sort of 80 % the embedding model.
16:48Animesh Chatterji:So there's also this reliance on, I guess, if you can use the embedding model to like truly represent the semantics of what it is that you're creating the embedding from, then you're going to get a higher quality search result.
17:00Ivan Solovyev:Yes. That's the fact that you are overlapping chunks. So potentially you would be retrieving multiple chunks which have the overlapping parts and then together they'll kind of recreate the whole context that is needed. Right.
17:11Animesh Chatterji:Why is there always a meeting bot in your Zoom call? Blame Recall.ai. Recall.ai powers the meeting bots and desktop recording apps behind products like Cluely, HubSpot, and ClickUp. They handle the hard infrastructure work, capturing clean recordings, transcripts, and metadata across Zoom, Google Meet, Microsoft Teams, in-person meetings, and more. So developers don't have to build it themselves. If you're building a meeting notetaker or anything involving conversation data, Recall.ai is the API for meeting recording. Get started today with$100 in free credits at Recall.ai slash software. In mobile application security, good enough is a risk.
17:54Animesh Chatterji:GuardSquare uses advanced, multi-layered code hardening techniques and automated runtime application self-protection and mobile application security testing. combined with real-time threat monitoring to deliver the highest level of mobile app security. Discover how GuardSquare brings all these together to provide mobile app security for your Android and iOS apps without compromise at www.guardsquare.com. You know Fidelity is a financial services leader, but did you know that Inside Fidelity is a community of technologists working together to shape the future of finance and tech? Fidelity is always investing in tomorrow, from emerging tech to cutting-edge tools that will transform what comes next.
18:39Animesh Chatterji:Their technologists are encouraged to keep learning so they can expand their skill sets, explore new ground, and stay ahead of this rapidly evolving industry. And right now, Fidelity is hiring technologists to join their team. Fidelity technologists get the best of both worlds, startup energy that's grounded in the stability of a financial institution. That means support, resources, and amazing benefits. Bring your skills to a culture where you're empowered to dream big and build the tech that drives an organization and makes a real impact on people's lives. Find out more at tech.fidelitycareers.com.
19:16Animesh Chatterji:That's tech.fidelitycareers.com. Fidelity is an equal opportunity employer. So you're abstracting away the vector database and the indexing that you're doing, but how How does this work with updates? I think that's historically been a challenge. If I process a document and then later that document changes, or maybe a website is maybe even a better example where a website is going to change from time to time, but I've already indexed that particular page and then I need to re-index it. How does that update process work?
Read the full transcript
19:45Ivan Solovyev:There are two parts to this update. One is basically you calling our API to ingest those documents. So we try to make sure that we are highly parallelized in terms of our ingestion latency. so we pretty much can panelize at a chunk level and ensure that all of those are ingested into the database and then google has the spanner which is also exposed externally as the cloud spanner which provides very strong consistency guarantees so once you pretty much write the data it's almost instantaneously available to be indexed and we leverage that capability of spanner to make sure that we can pretty much read our writes as soon as they are available.
20:25Ivan Solovyev:That significantly reduces the delay in reading the indexes and reading the embeddings.
20:31Animesh Chatterji:Do you have to, like if I've already indexed a particular page though, or a document, and then I'm re-indexing it, do I have to blow away the initial indexing indexes in order to re-index it? Or is there essentially the equivalent of like an upstart in the vector world?
20:47Ivan Solovyev:So essentially you have the corpus, you can add your new document to that corpus, which would just mean that the new chunks are indexed. The rest of the index remains as it is. In our world, you are not updating the document, you are inserting a new version of the document and we will be chunking that and indexing it.
21:05Animesh Chatterji:But if the old version is there, do you run into this potential risk that when you're pulling back relevant chunks, you could pull back relevant chunks that are no longer actually relevant because the fundamentals of the document has changed?
21:17Ivan Solovyev:Yeah, so that capability we provide to the developers in terms of the corpus management or the document management APIs. So if they want, they could delete the earlier document. But from our perspective, it's difficult for us to figure out whether it's the new version of the same document or not. We are not doing data at our end. So it's up to the developers to kind of remove the old content if they think that's not relevant anymore.
21:39Animesh Chatterji:I see.
21:40Ivan Solovyev:Okay.
21:40Animesh Chatterji:And then is the search that's going on, is it purely vector-based? Is there a hybrid element to this?
21:46Ivan Solovyev:Right now, it's purely semantic search, which is vector-based. We have had some requests of users wanting a keyword-based search, and that is something we're considering adding to the roadmap. Given the indexing capabilities that Spanner offers, we think it's like a natural extension to the offering, and yet something which should not add too much complexity to our system.
22:07Animesh Chatterji:We also looked at the GraphRack systems in the past, but for now, it, to me, at least feels a little bit more complex for the product that we're trying to build. So we haven't found the right way to simply integrate it into the system yet. Yeah, I mean, I think that you see a lot in these like more complex rag pipeline scenarios where they're using a combination of vector search. There might be a knowledge graph or ontology or something like that to also ground the results in some semantic understanding. Is that something that you see as a future direction for this? Or would it be more, you would use this in combination perhaps with a separate system that would handle that piece of it?
22:46Animesh Chatterji:So far, what we saw from our customers is that the current setup is working well for them. And I think we will not overcomplicate it just yet. To answer your question directly, I feel that we better have two separate systems that can complement each other. And as you need to grow, you can implement both the services. In a more complex use cases, I would say we also have, I mentioned the Vertex Rack system, which is built on top of the, it's not quite Gemini API, but it's a very similar Gemini API inside the Vertex. So for anything that it requires a lot more complexity, configuration, maybe swapping out the databases or adding these additional systems on top, we can always guide customers to use more complex solution.
23:35Animesh Chatterji:They really need to, and we can focus on the simplicity and getting started. Okay. How do citations work? How do you map, I guess, sort of the generated token back to the specific source chunk? Yeah.
23:48Ivan Solovyev:So right now it's the models are trained to cite the responses. So when they generate the responses, they actually cite every sentence that they have used the original corpus to generate from. And then it's just a matter of post-processing that response, removing those citations and adding that separately as grounding data. Yeah. So essentially it's models generating citations to the data that they refer to.
24:13Animesh Chatterji:Yeah. So the model is trained to essentially figure out or to provide a reference back to where that text or what the source text was. And then you have to map that source text, I guess, back to the database chunk in the original source in order to inject the link or something like that, that refers to the citation. Is that right?
24:32Ivan Solovyev:yeah so basically when like the flow is something like this when the model realizes that it needs to use file search it will emit a query that i want this query to be answered by the file search tool you run the query and you give the responses each response in some sense is indexed uniquely so if the model is receiving five chunks of data it knows that each of them is a different index and this index can vary per turn as well so now when the model responds it kind of sites, the exact unique index using which we can kind of figure out which chunk it was referring to, and then figure out which document it was part of and add more metadata about that.
25:09Animesh Chatterji:Okay. The blog post that talks about the product, in that you cover, there's a company Beam, which is a AI-driven game generation platform that's using this. Can you talk a little bit about how are they using this product in their, I guess, to solve problems in their world? I think their use case is pretty neat and simple at the same time. So they have lots and lots of new developers coming to the platform who want to build games with AI, and they don't necessarily experience developers. They are mostly learning. And AI helps Beam to educate and help those developers to create their first game.
25:51Animesh Chatterji:So the way they are using file search is they have a huge code base that is their engine plus the documentation on top of that, that is talking about how each component is used, how animations are happening, how scripts are implemented, et cetera, et cetera. It's a rather big data set. So what they do is they put it all into the file search, they index it, and whenever the user starts experimenting with the agent and the agents that supports it, they will naturally ask questions how to do specific things. And through file search, they can very quickly pull all the relevant documentation into the context and actually present to the developer that, hey, you probably want to use this module.
26:31Animesh Chatterji:Here is how it works. Here is the old documentation. And they've been able to close this education loop for their customers, receive great feedback. What's the performance on the retrieval look like? Performance in terms of retrieval quality or latency? Let's start with latency. Latency is somewhat in line with the model latency, a couple of seconds for the retrieval. In terms of quality, it will depend on the use case. If I recall correctly, we saw up to like 85%, depending on the use case of the retrieval, like correct hits in terms of documents redrift. As a user of this, given that, I mean, any reg system, it's going to be very difficult to get like 100 % accuracy on retired documents.
27:17Animesh Chatterji:But what are some of the things that, you know, approaches people take to help increase the accuracy?
27:24Ivan Solovyev:I think there would be like a few things, right? One would naturally be the embedding model that Ivan talked about and kind of called out the importance. The second is your retrieval strategy. Like sometimes you would want to optimize quality for latency, whether you want to kind of go through your entire database, find all relevant chunks, or kind of figure out the first few relevant chunks and give it to the model. So that would kind of be the other aspect on like trading of latency versus quality. And third, I think is just the model training on triggering the search only when relevant and also not hallucinating the answers.
27:59Ivan Solovyev:Those are kind of orthogonal to file search. Those apply to any tool that we have trained Gemini with. But I would say those are kind of the three aspects, the mailing quality or retrieval quality and just the model quality, which probably is of utmost importance.
28:13Animesh Chatterji:Yeah, that's that's mostly on our end. And that's what we are working on in terms of improving the quality. For developers, what I saw is some developers actually implement the post-processing. So they would implement file search calls in a separate, in a sub-agent or a separate flow. And they will do filtering on top of the returned results. So they will call Gemini one more time. They will have a prompt that is doing the verification of the results in comparison to the context that the model already has. and they will cull out the results that don't really fit the convex-on-victorization and improve the quality of the output that way.
28:49Animesh Chatterji:Is there value in using a Reranker model? Have you found in your own experiments that it actually improves? It's worth, I guess, the investment of introducing something to Rerank the results returned from the vector search?
29:03Ivan Solovyev:Not so much. In fact, I mean, that just adds more complexity. And if we kind of extrapolate the question that we were talking about a bit earlier about whether even we do we need drag i think if you kind of take that logic and apply it here that once you give the relevant chunks and as long as your context is not blowing up too much i think letting the model figure out what is relevant probably is better yeah in terms of like if you are retrieving too many chunks then we have some threshold on what is the quality score of this chunk right and then we have some threshold below which we don't return those to the model but that's not like re-ranking between chunks is just like a vanilla cutoff beyond which we don't get any more chunks.
29:43Ivan Solovyev:But between the time, in fact, we have not seen any advantage of providing a ranked order to the model.
29:51Animesh Chatterji:What about in terms of people who look at fine-tuning and bending models for specific use cases? Has there been recent results around actually improving retrieval or is it, again, just sort of overcomplicating the whole process? We have a general recommendation in GDM, and I think we made this a year or so back, that people shouldn't do fine-tuning in most of the cases. The speed of progress of the models is so much faster than what individual smaller labs can do in terms of fine-tuning that it's almost irrelevant. And by the time you actually have a fine-tuned model, and it will probably perform better for a use case for like a month or two, we're going to have the next 003 embedding model.
30:36Animesh Chatterji:It's going to be better across the board, like 15 % on all the benchmarks. And fine-tuning won't be that relevant anymore. With that said, we do see people using fine-tuning in specific use cases. I think it does make sense if your use case is very, very niche and you own a very particular data set, which you don't expect Google or anyone else to pay attention to anytime in the future. That may yield good results. But as I said, so far, what we've seen is fine tuning becomes relevant within six months. Yeah, I think that's fair. I've found a similar result working with business on my end. And I think that one of the challenges as well is that if you do go through the process of fine tuning, then even if you are getting better results for six months or whatever, then a new model comes along.
31:24Animesh Chatterji:You've adjusted the weight. So how do you apply that to a new version of the model? Yeah, you need to start over. That's pretty much from scratch. You have to fine tune the model again. Yeah, it gets kind of expensive. As the models improved, do you think that a lot of the things that we've historically done with RAG to try to drive out performance ends up, basically, we don't have to make things quite as complicated because we can rely on better performance from the model and the model is kind of absorbing a lot of the complexity? Yeah, absolutely do think so. So we did see an amazing progress for Gemini models in the last year.
31:58Animesh Chatterji:And not just Gemini models. You look across the board, Anthropic, OpenAI did great work in improving their text LLMs. First of all, these improvements do convert into embeddings models as well. And separately, we are working on improving the embeddings model more and more. And we're going to see in the next year or two, we're going to see significant improvements in terms of retrieval quality, in terms of use case complexity that those models can handle. So I do believe that a lot of this additional configuration that is happening around those models will go away and you will be able to just embed the thing, hit the search and get the results that are really relevant.
32:35Animesh Chatterji:And what are some of the things that have happened over the last year or so that have made the embedding models better? Like what are the particular innovations that have happened there to really drive up performance?
32:47Ivan Solovyev:I mean, we have added the multimodal embedding support now, which would really improve the quality of understanding of things beyond text. So that is one thing. And I think it's in public preview right now. So we are hoping to do a GA launch for that. The other thing which kind of launched, I think, in the last version of our embedding model or the one before that, I don't remember, is we started representing embeddings as this Matryoshka representation, which basically means the embedding vector that is generated. the front part of that vector has more context about the thing being embedded than the latter part which makes it very easy for the end user to just truncate the embedding so in case like the embedding is it's a 3k dimension embedding vector and you don't want to store as much size you could just truncate it at any point and it would still give you an accurate enough representation of the entity so some of those things have been really helpful and some of those we can actually explore like we were talking about the knobs that we could give to the user that could be another knob in future we could give to the users if they want to reduce the size of their storage by using a truncated embedding instead of the full embedding at the cost of some quality so based on their use case users can choose to pick one of the two
33:58Animesh Chatterji:knobs what do you see as some of the like hard problems in this space that are yet to be solved when it comes to rag in particular i think even please add on but like multimodal yes
34:10Ivan Solovyev:I'm sure we'll kind of add on more capabilities on the multimodal side. That is one thing. We talked about chunking. That's still an area I feel we can get more benefit out of by kind of capturing the structure better in certain kind of use cases. The multilingual or like the internationalization, that is another aspect. I feel we are getting great at solving English related queries, but there are certain languages where certainly we could do more. And as the user base expands to countries across the world we have more of these internationalization use cases so that is another aspect where i feel we can certainly improve upon yeah just in general
34:47Animesh Chatterji:i think getting the quality higher hit rates better retrieval is something that we always pursue multi-modal is very interesting aspect text is working great for a lot of use cases but there is a lot more multi-model use cases that people are thinking of right now. Looking for images, looking through video opens up a lot more consumer products that can be built on top of that. And the file search and any record drill here is even more beneficial than for text because of the size of the data that you are feeding into the model. So if you can reduce that as much as possible through file search, that'd be very interesting.
35:27Animesh Chatterji:One thing, just bringing us back to the file search tool. In terms of people who've already invested in a particular stack to do RAG, whether it's a combination of vector database, maybe a particular framework, they have their chunky strategy, what is it that they need to do if they wanted to migrate essentially over to the file search tool approach? Well, migration itself, I hope, is fairly easy and straightforward. You upload your data. But the reality is, I think, what I would recommend them to start with is probably use the embedding model that we provide. It is available as a standalone service.
36:04Animesh Chatterji:So if they have their own pipeline and they want to run the evals and experiment with the Gemini infrastructure, I would recommend using the embeddings model first with part of their data and just comparing the results. And the next step would be using the file search, loading a small portion of their database and just running evals, comparing both systems. And then what is the status of a file search tool in terms of its availability? Is this early access? Is this preview? Or what can you share around the timelines where people can start to get their hands on this more generally? Yeah, absolutely.
36:38Animesh Chatterji:File search is actually generally available for our 2.5 model family and recently our 3 model family. So 2.5 models are in GA. So it's both models and the two combination are generally available. Flash 3 and Pro 3 models are still in preview, but file searches as an API is generally available there as well. Okay. And then what is next besides moving all the stuff to getting it in the hands of more users? What can you share about some of the things that you're thinking in terms of additional problems to attack or things that you want to continue to make investments around making the product really, really easy to use?
37:19Yeah.
37:19Animesh Chatterji:So as we mentioned, Multimodal support is the big push we're doing. We want to invest into a better understanding of structured data. And we keep collecting these examples from our developers in terms of tables and graphs and whatnot that they're trying to process. So I think that's going to improve the quality and the applicability of this a lot. And then latency and being able to work with a much bigger data sets. We do limit at one terabyte right now for the highest tier, but the latency go down quite a bit. if you start consuming all of the water. So we want to invest in that as well and improve the retrieval latency.
37:59Animesh Chatterji:Is that one terabyte of total storage? Is that right? Yes, that's one terabyte of total storage across all of your file stores. Is there a limitation around the size of a single file that can be processed? Other than, I guess, it needs to be smaller than a terabyte.
38:14Ivan Solovyev:It's 100 MB right now. So that is what we have kept on the file. And for making sure that the retrieval latencies are kind of acceptable and in the good range, we recommend users to keep their individual corpus to about 20 GB. So their total data across corpuses can be one GB, but each corpus or each file search store, as we call it, should be relevant to 20 GB. And then at the query time, you can provide multiple of these file search store IDs. So we can find out those queries in parallel, but the size of a individual file search store performs good till the limit of more 20 GB.
38:49Animesh Chatterji:Okay. And then if I want to play with this, Like, how do I get started? The simplest way, I think, would be to go to AI Studio and play with the file search applets. I think the link is available in the blog post that we published. And the other way is to start hitting the Gemini API.
39:06Ivan Solovyev:Okay, great. We have two samples that would make it really easy for somebody to just start playing around, just upload their data using the upload API and hit the Gemini API.
39:16Animesh Chatterji:Oh, and our wipe coding environments in AI Studio also fully supports the file search. So you can just prompt the model to generate the code that uses file search and it will do it for you. Okay, well, even easier. Is there anything else you would like to share as we start to wrap up? I just want to say that it's been really exciting to see the adoption that we're getting for file search. We actually received quite a lot of great feedback from developers and quite a lot of excitement. And it's been really nice to see how this works for their use cases. Fantastic. Well, Ivan and Emesh, thank you so much for being here.
39:53Animesh Chatterji:Thank you for having us. Cheers.
40:03Thank you.
From the publisher
Retrieval-augmented generation, or RAG, has become a foundational approach to building production AI systems. However, deploying RAG in practice can be complex and costly. Developers typically have to manage vector databases, chunking strategies, embedding models, and indexing infrastructure. Designing effective RAG systems is also a moving target, as techniques and best practices evolve in step
The post DeepMind’s RAG System with Animesh Chatterji and Ivan Solovyev appeared first on Software Engineering Daily.
