In short
Google AI: Release Notes - Episode Summary
Episode Title
Deep Dive into Long Context
Episode Description In this episode, host Logan Kilpatrick interviews Nikolay Savinov from Google DeepMind to explore the significance of long context models and their interplay with Retrieval Augmented Generation (RAG). The discussion sheds light on advancements in large context windows and future directions in AI.
---
Key Chapters
- Introduction & Defining Tokens (0:52)
- Importance of Context Windows (5:27)
- RAG vs. Long Context (9:53)
- Scaling Beyond 2 Million Tokens (14:19)
- Long Context Improvements Since 1.5 Pro Release (18:41)
- Challenges of Attending to Whole Context (23:26)
- Evaluating Long Context: Beyond Needle-in-a-Haystack (28:37)
- Integrating Long Context Research (33:41)
- Reasoning and Long Outputs (34:57)
- Tips for Using Long Context (40:54)
- Future of Long Context: Near-Perfect Recall and Cost Reduction (48:51)
- The Role of Infrastructure (54:42)
- Long Context and Agents (56:15)
---
Episode Highlights
Understanding Tokens
- Definition: A token is roughly equivalent to less than one word in text, which can include words, parts of words, or punctuation.
- Purpose: Tokens are essential as they enable faster generation compared to character-level processing.
Context Windows
- What is a Context Window?: The context window comprises tokens fed into an LLM (Language Model), including prompts and previous interactions.
- Memories:
- In-Weight Memory: Knowledge learned during pre-training.
- In-Context Memory: User-supplied information for real-time updates and personalization.
RAG vs. Long Context
- Retrieval Augmented Generation (RAG): An engineering technique that retrieves relevant information from a knowledge corpus before involving the LLM.
- Synergy: Long context can enhance RAG's effectiveness by allowing for more extensive relevant data retrieval.
Scaling Challenges
- Current Limits: The discussion revolves around the technical and cost-related limitations of scaling beyond 1-2 million tokens.
- Quality vs. Cost: While larger context sizes improve recall, they come with exponentially higher costs.
Advances Since Last Release
- Quality Improvements: Significant advancements in the processing capabilities of long context models since the 1.5 Pro release.
- Benchmarking: The 2.5 Pro model shows improved performance over prior models and competitors.
Evaluating Long Context
- Needle-in-a-Haystack Tests: Previously a solved problem; more complex tasks with hard distractors present new challenges for evaluation.
- Real-World Applications: Emphasis on practical benchmarks that test LLMs’ capabilities to synthesize information across large inputs.
Future Directions
- Near-Perfect Recall: Expectations of achieving high-quality retrieval and cost-effective scaling within the next few years.
- 10 Million Context Windows: Potential future developments could unlock new dimensions in coding and data processing applications.
Developer Best Practices
- Use Context Caching: To improve response speed and reduce costs when reusing the same context.
- Combine with RAG: Especially beneficial in cases requiring retrieval of multiple pieces of information.
- Avoid Irrelevant Context: To optimize model performance.
Education and Collaboration
- Research Integration: Importance of collaboration between various teams while maintaining ownership of specific capabilities.
- Long Context in Agent Frameworks: Agents can function as both consumers and suppliers of long context, automatically sourcing information when needed.
---
Conclusion The episode provides a comprehensive overview of the advancements in long context models, their relationship with retrieval techniques, and the future potential of these technologies in AI-driven applications. The insights shared by Nikolay Savinov underline the importance of continual research and development in enhancing AI capabilities, particularly in handling large datasets and complex tasks.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00I'm really impressed by the work of our inference team. I've got a bunch of spicy rag versus long context questions for you. You can rely on context caching to make it both cheaper and faster to answer. What's the limitation of continuing to scale up beyond one to two million? This thing is going to be incredible for coding applications. We will have lots more exciting long context stuff to share with folks.
0:51Welcome back to Release Notes, everyone. How's it going? Today, we're joined by Nikolai Savinov, who's a staff research scientist at Google DeepMind and one of the co-leads for long contacts pre-training. Nikolai, how are you? Yeah, hi. Thanks for inviting me. Let's start off with, at the most foundational level, and we'll sort of build up from that, what is a token and how folks should think about that? So the way you should think about a token, it's basically slightly less than one word in case of text. So a token could be a word, part of a word, or it could be things like punctuation, like commas, full stops, etc.
1:39And for images and audio, it's slightly different, but for text, just think of it as slightly less than one word. Yeah. And why do we need tokens? Like, humans generally are sort of familiar with characters. Why does AI and LLMs have this special concept of a token? What's, like, what does it actually enable? Well, this is a great question, and actually many researchers ask this question themselves. So there were actually quite some papers trying to get rid of tokens and just rely on character-level generation. But the thing is, while there are some benefits of doing that, there are also some drawbacks.
2:23And the most important drawback is, well, the generation is going to be slower. Because you generate roughly one token at a time. And if you are generating a word in one go, it's going to be much faster than generating every character separately. separately. So those efforts, they, I would say they didn't really succeed and we are still using tokens. Yeah. For folks who haven't spent a bunch of time thinking about tokens, there's a bunch of good, uh, Andre Karpathy videos and tweets and stuff like that of how like tokenizers are the root of all like weirdness and complexity and LLMs, like all these weird edge cases that you run into.
3:04It's like most of them are rooted in the fact that the model is not looking at things from a character level, it's looking at it from a token level. And actually, like, the pertinent example, which folks love to go to these days, is, like, counting the characters in a single word. Like, how many R's are there in strawberry is, like, a weird problem to solve. My understanding is because tokenizers, like, break the word into different parts, it's not actually looking at the word at, like, the individual character level. Is that an apt description? Yeah, I think that's a pretty good description of the problem.
3:38And one thing you should realize is that those models, due to tokenization, they view the world very differently from how humans view the world. And when you see a strawberry, you see a sequence of letters. But what the model says, it could be even one token. and then you ask like, hey, count number of R letters in this token. But this is a pretty hard, it's pretty hard to get this knowledge from pre-training because you would need to associate the R letter token that you encountered somewhere on the web with the word strawberry, which is also one token. So if you think about the mental load of doing that, it's not such a trivial task, I would say.
4:33Although obviously when the model can't do it, we start complaining, hey, like if it's AGI, how come it can't count number of R letters in strawberry? That's like a child could do that. Yeah, it is super weird. And actually another interesting thing is that if you watch some of the Carpathi videos, then there are a lot of problems with the white space. So this is an interesting point because normally most tokens are prefixed with the white space. And then some really weird effects might happen because you might encounter problems on the boundaries when you think you are concatenating something, but this concatenation is very unusual for the model to see.
5:27Interesting. That is super interesting. I think this actually takes me to just like generally talking about context windows and I think there's a lot of discussion or obviously we're talking about long context which sort of assumes you know what a context window is but can you give the lay of the land of like how folks should thinking about like what a context window actually is? Why do I, as a user of LLMs or somebody who's building with AI models, why do I need to care about the context window? So context window is, those are basically exactly this context tokens that we are feeding into LLM.
6:02And it could be the current prompt or the previous interactions with the user. It could be the files that the user uploaded, like videos or PDFs. And when you supply context into the model, the model actually has knowledge from two sources. So one source is what I would call in-weight or pre-training memory. So this is a knowledge that, well, the LLM was trained on a slice of the internet, and it learned something from there. it doesn't need additional knowledge to be supplied into context to remember some of those facts. So there is already, even without context, there is some kind of memory present in the model.
6:50But another kind of memory is this explicit in-context memory that you are supplying to the model. And it's pretty important to understand the distinction between those two, because in-context memory is much, much easier to modify and update than in-weight memory. So for some kinds of knowledge, in-weight memory might be just fine. Like if you need to memorize some simple facts that the objects fall down and not up. This is a very basic common fact. It's fine if this knowledge comes from pre-training. but there are some facts which are true at the time of pre-training but then they become obsolete at the time of inference and you would need to update those facts somehow and the context provides you a mechanism for to do this update and it's not only about the up-to-date knowledge there are also different kinds of knowledge like private information The network doesn't know anything about you personally, and it can't read your mind.
8:02So if you want it to be really helpful for you, you should be able to supply your private information into context, and then it will be able to personalize. Without this personalization, it's going to give you generic answers it would give to any human instead of answers tailored to you. And the final category of knowledge which need to be inserted in context is some rare facts. So basically some knowledge which was encountered very sparingly on the internet. And I must say, I suspect this category of knowledge, it might go extinct with time. Maybe future models will just learn the whole slice of the internet by heart.
8:50and we will not need to worry about those. But the reality at this point is that if something is mentioned once or twice on the whole internet, the models are actually unlikely to remember those facts and they are going to hallucinate the answers. So you might want to insert those explicitly into context. And the kind of trade-off we are dealing with is for the short context models, you have limited ability to provide additional context. Basically, you would have a competition between knowledge sources. And if the context is really large, then you can be less picky about what you insert and you can have higher recall and coverage of relevant knowledge.
9:41And if you have higher coverage in context, That means, well, you're going to alleviate all those problems with in-weight memory. Yeah, I think there are so many angles to push on. That was a great description. One of the follow-ups from this is we talked about sort of in-weight memory. We talked about in-context memory or just in-context in general. The sort of third class is around how to bring context in that, like, through RAG systems, retrieval augmented generation. Can you sort of give like a high level description of RAG? And then I've got a bunch of spicy RAG versus long context questions for you.
10:25Yeah, sure. So what RAG does is, well, it's a simple engineering technique. It's an additional step before you pack the information into LLM context. So imagine you have a knowledge corpus and you chunk this knowledge corpus into small textual chunks. And then you use some special embedding model to turn every chunk into a real valued vector. And then based on those real valued vectors, if you get the query at the test time, you can embed the query as well. and then you can compare this real valued vector for query to those chunks from the corpus. And for the chunks which are close to the query, you're gonna say, hey, like I found something relevant, so I'm gonna pack those chunks into context and now I'm running LLM on this.
11:27So that's how RAG works. And this is maybe a silly question. RAG, my sense has always been like, lets you obviously there's like very hard limits on context that you can pass to the model we have 1 million we have 2 million that's awesome but like actually if you look at like internet scale you know wikipedia has you know many trillions of tokens or whatever or maybe not trillions maybe billions of tokens whatever it is um why is like rag as this notion of like bringing in the right context to the model not just like baked into the model itself like is it just that to the point of the conversation, the model just not working well for like, it's just like the wrong research direction to go in or like, why don't you think we build that mechanism in?
12:12Because my face value perspective is like, it seems like that would like kind of be useful. The model could just like do rag. And if I could pass a billion tokens and then let the model sort of figure out heuristically or through the, you know, whatever mechanism, what the right tokens are or is that just like a problem somewhere else in the stack that should be solved and the model shouldn't have to think about that? Well, one thing I want to say is that after we released 1.5 Pro model, there were a lot of debates on social media like is RAG becoming obsolete? And well, from my perspective, not really like, say, enterprise knowledge bases, they constitute billions of tokens and not millions.
12:58And so for this use case, for this scale, you still need RAC. What I think is going to happen in practice is that it's not like RAC is going to be eliminated right now, but rather long context and RAC are going to work together. and the benefit of long context for RAC is that you will be able to retrieve more relevant needles from the context by using RAC and by doing that you're going to increase the recall of the useful information so if previously you were setting some rather conservative threshold and cutting out many potentially irrelevant chunks then now you're going to say hey I have a long context so So I'm going to be more generous.
13:47So I'm going to pull more facts. And so I think there's a pretty good synergy between those. And the real limitation is the latency requirements of your application. So if you need real-time interactions, then, well, you'll have to use shorter context. But if you can afford to wait a little bit more, then, yeah, you're going to use long context just because you can increase the recall by doing that. Why is 1 million just like a marketing number or like is there something like intrinsic that like after a million or 2 million like is there actually like something technically happening around the like million token mark for from a long context perspective or is it literally just we found a number that sounds good and then made the technology work from a from a research perspective?
14:39Well, when I started working on Long Context, the competition at the time, I think it was about 128k or maybe 200k tokens at most. So I was thinking how to set the goals for Long Context projects. And well, it was at the time, it was a small, small part of Gemini. And I originally And I thought, well, I mean, just matching competitors doesn't sound very exciting. So I thought, let's set an ambitious bar. So I thought, well, one million is an ambitious enough step forward. It was like compared to 200K, that's like 5X. And very soon after we released one million, we also actually shipped two million, which was about 10X larger.
15:34and I guess one order of magnitude larger than the previous state of the art that's that's a good goal that's what makes it exciting for people to work on yeah I love that and how my my follow-up spicer question from that is like we shipped 1 million we shipped 2 million rapidly after that like what's the limitation of like continuing to scale up beyond 1 to 2 million Is it like from like a serving perspective, it's too costly or too expensive? Or is it just like the architecture that makes one to two million work, like just like fundamentally breaks down when you go larger than that? Or like how how come we haven't seen the frontier for long context continue to push?
16:19Yeah. So when we released the 1.5 Pro model, we actually ran some inference tests at 10 million. And we got some quality numbers as well, and for, say, single-needle retrieval, it was almost perfect for the whole 10 million context. We could have shipped this model, but it's pretty expensive to run this inference. So I guess we weren't sure if people are ready to pay a lot of money for this. So we started, you know, more, you know, with something more reasonable. like in terms of the price but uh i think in terms of quality that's also a good question because uh it's it was so expensive to run we didn't run many tests and so just you know bringing up this server again it's also quite costly so unless we want to you know ship it to a lot of customers right now it's you know we don't have chips to do that yeah do you think that that will continue to hold that like the it's like this like matt i don't know if it's like an exponential increase in capacity that's needed as we do more long contact stuff but like do you have an intuition that like that will like do we need like fundamental breakthroughs from a research perspective for that to change to make it so that we can actually keep scaling up beyond or is it like one to two million is going to be what we stick with and if you want more than that do rag and then like be really smart about bringing context in and out of the of the context window from the model perspective so my feeling is that we actually need more innovations so it's not just a matter of brute force scaling to actually have close to perfect 10 million context we need to learn more innovations but then in terms of the rack and which paradigm will be more powerful going into the future.
18:21I think that the cost of those models is going to decrease over time. And we're going to try to pack more and more context retrieved with drag into those models. And because the quality is also going to increase, then it's going to be more and more beneficial to do that. Yeah, that makes sense. Can you take us back to when we originally landed long context? My understanding of the story is for 1.5 Pro, we didn't, it wasn't like it had been built for a long context to begin with. I think you had obviously tried to kick off that work stream with others inside of DeepMind. And it ended up just being that the pace of research progress was super fast.
19:07And I think my loose understanding of the story is we had the breakthroughs, we realized it worked, and then it was shortly thereafter, it ended up actually landing in the model side. or like what was the like timeline from the the effort starting to like actually landing it into a model that was available externally to the world oh i think that was uh that was indeed pretty quick and just to clarify we didn't really like we were wishing to to go long and achieve say one million or two million context but we kind of didn't expect uh ourselves to get there that fast and when it actually happened then we thought like hey like this is this is really impressive like we actually made some strides on this task so now we need to ship it and then we actually managed to assemble a pretty awesome team very quickly and the team worked really hard.
20:10Like, to be honest, in my life, I've never seen people working this hard. I was really impressed. I love that. That's awesome. And that was for the original 1.5 Pro series. We landed it for 1.5 Flash as well. We now have it for 2.0 Flash. We have 2.5 Pro. And I think the, can you sort of give us the lay of the land of like, what's, what's been happening from from a long context perspective, from that original launch when we know long context is possible, we released the technical report for 1.5 Pro, which showed the needle in the haystack results, a bunch of stuff like that, to today, where I think a lot of what's actually making this 2.5 Pro model blow people's mind is actually how strong it is at long context, which has been awesome for coding use cases and stuff like that.
21:01So what's happened in the long context world from original launch to today? Yeah, so I think the biggest improvement was actually the quality. And we made strides both on the quality at, say, 128k context and also at 1 million context. So if we look at the benchmark results for 2.5 Pro model, we observe that it's better compared to many strong baselines, like GPT 4.5, CLOT 3.7, and also O3 MiniHi, and some of the DeepSeq models. So the quality, like, to actually compare it to those models, we had to run the evals at 128k so that they all comparable and we saw quite a big improvement for 2.5 pro and now in terms of 1 million context we compared it with 1.5 pro and we also saw some significant advantages this is maybe a weird question but like does the quality like ebb and flow at different context sizes?
22:17Like, do you see like, is it like, like almost like linear quality, like on a, you know, 100 ,000 token input versus like 128 ,000 or like a 50 ,000 versus 100? Like, is it like pretty consistent across? Or is there like weird, like, I'm trying to imagine, maybe the, it all generalizes when you make it into the final model, and there's no there's no difference. But like, is there any like, nuance in that perspective? Have we done evals that show anything like that? Yeah, internally we looked at some of those evals. I guess maybe your question goes into these effects that people observed in the past, like a very popular one was lost in the middle effect.
22:57And to answer your question, the lost in the middle effect where you have a deep in the middle of the context, we don't really observe this with our models. But what we do observe is that if it's a hard task, not like a single needle, but some task with the hard distractors, then the quality slightly decreases with the increasing context. And that's something we want to improve on. Yeah. And just for my own mental model, does like when I think about putting 100 ,000 tokens into the context window of the model, Should the should like from a developer perspective or a user of who's like actually using the long context functionality.
23:43Should I assume that the model is like actually attending to all of the different context? I know it could definitely do the like the one needle. It can pull that out. But like, is it actually reasoning over all those tokens like in the in the context window? I have just a bad mental model of what's happening behind the scenes when you have that much context in the context window of the model. Yeah, I think that's a good question. So one thing you need to keep in mind is that attention, in principle, has a bit of a drawback because there's a competition happening between tokens. so if one token gets more attention then other tokens will get less attention the thing is if you have hard distractors then one of the distractors might look very similar to the information that you are looking for and it might attract a lot of attention and now the piece of information that you are actually looking for it's going to receive less attention and the more tokens you have the harder competition becomes so it it depends on the hardness of uh destructors and also also on the context size yeah this is another silly follow-up question but like is the is the amount of attention always fixed like there's no like is it possible to have more attention or is it just like whatever it's like a value of one and then there's like you know spread across all of the tokens in the context window and like so the more tokens you have like literally the less attention there is there's no way for that to change normally that's that's the case that um like the whole pool of attention is um is limited yeah from that example you gave about like the like hard distractors like causing the model to do a lot more work and sort of split the attention has the t has your team explored or other teams on the applied side explored like like pre-filtering mechanisms, any of that type of stuff.
25:50So you want long contacts to work really well in production. It sounds like the best outcome is you have very dissimilar data that's in the context window. If there's a lot of similar data, and you're asking a question that could be relative to all of it, you'd expect the performance to be worse in general in that use case. So is that just something that developers or the world needs to figure out from that perspective? Or do you have any suggestions of how folks should approach that problem for me as a researcher i think it's it would kind of be a move in the wrong direction i think we should work more on improving the quality and robustness instead of coming up with some hacks for filtering one practical recommendation though is of course uh try not to include totally irrelevant context like if you know that something is not useful then what's the goal of including it into the context because in the very minimum it's going to be more expensive so why would you do it yeah that's it's interesting because i feel like um in some sense that like goes against the like core way that people use long time like i think the examples i see online it's just like people being like oh just take all this random data and throw it in the context window of the model and have it sort of figure out what's useful for me So you'd almost expect the model to do, given how important it sounds like it is to remove some of that stuff, the model to do the pre-filtering itself almost to only include the relevant.
27:22Because I think that, not that humans are lazy, but I feel like that's been one of the selling points. It's like, I don't need to think about what data I'm putting into the context window. So do you think there's a world where the model, it's like a multi-part system or something like that, where the model is actually doing some version of the model is doing some of that? eliminating the extraneous data based on what the user's query is and then like making sure when the context actually goes to the model it's a little bit easier over time as the models get better quality and they get cheaper you'll just you'll not need to think about this anymore i'm just talking about the like the current reality is like if you want to make a good use of use of it right now then well let's be realistic uh just don't put irrelevant context but Also, I agree with your point that if you spend too much time manually filtering it or like hand crafting which things to put into context, that's annoying.
28:21So I guess there should be a good balance between those. Yeah. I think the point of context is to simplify your life and make it more automatic, not to make it more time consuming or, you know, make you spend time hand crafting things. Yeah. I've got to follow up on this around evals and the evals that you're thinking about from a long context quality perspective. Needle in a Haystack, obviously, that was the original one that we put into the 1.5 technical report. And for folks who aren't familiar, Needle in a Haystack is just asking the model to find one piece of context in 2 million, 1 million, 10 million tokens of context.
29:00The models are extremely good at this. How do you think about the other set of long context? Like, is there like a set of like standard, I think like folks like generally, I feel like needle and haystack gets talked about a little bit, but are there like another set of like standard benchmarks that you're thinking about from a long context perspective? So let's see, I think the evaluation is pretty much the cornerstone of the LLM research. And especially if you have a large team, evaluation provides a way for the whole team to align and push in the common direction. So the same applies to long context.
29:39If you want to make progress, you need to have great evaluations. Now, single needle in a haystack, it's a solved problem, especially with easy distractors. So if it's like Paul Graham's essays and you put a phrase, here is my magic number for the city of Barcelona is 37. Give me the magic number for the city of Barcelona. This is really a solved problem. But now the frontier of capabilities is handling hard distractors. If you, for example, packed your whole context with a magic number for CTX is Y, and you pack, say, the whole million contexts with these key value pairs, that's a much harder task because then distractors are actually looking very similar to what you want to retrieve.
30:39another thing which is hard for llms is retrieving multiple needles so i i feel like these two things the hardness of destructors and multiple needles they are the frontier but also there are there are additional considerations for the evals one consideration you might have is oh well like those new in the haystack evals even with hard destructors they're pretty artificial so maybe I want something more realistic. This is a valid argument, but the thing you need to keep in mind is that once you increase the realism of the eval, you might actually lose the ability to measure the core long context capability.
31:21For example, if you are asking a question to a very large codebase, and the question is basically can be answered by just one file in this codebase, and then the task is to actually implement something complicated, then you're not really going to be exercising the long context capability. Instead, you are going to be exercising the coding capability. And then it will give you a wrong signal for hill climbing. It will basically hill climb on coding instead of long context. Yeah. So that's one consideration. Another consideration is that something which people call retrieval versus synthesis evolves.
31:58So theoretically, if you need to just retrieve one needle from the haystack, that can be solved by Rack as well. But the tasks that we should really be interested in are the tasks which integrate information over the whole context. And for example, So, well, summarization is one such task, and Rack would have a hard time dealing with this. But now these tasks, it sounds nice and the right direction to go, but they're actually not so easy to use for automatic evaluation. For example, the metrics for summarization, we know that they are like metrics like rouge, etc. They are imperfect. And if you're doing hill climbing, then you're better off using something which is more,
32:56how do I say, less gameable metrics. And just a quick follow-up, what makes them less useful? Like, for summarization as an example, is it just that it's more subjective of what a good summary is versus what isn't it? It doesn't have a ground truth, source of truth? Or what makes that use case hard? Yeah, those of us are going to be pretty noisier because there will be a relatively low agreement even between the human raters. Of course, this is not to give an impression that we shouldn't work on summarization and we shouldn't measure summarization. These are important tasks. I'm just talking about my personal preferences as a researcher is to hill climb on something which has a very strong signal.
33:42Yeah, that makes sense. How do you see sort of as long context, especially for Gemini, is just like a core part of the capability story that we're telling the world. It's like a core differentiator for Gemini. And yet at the same time, it feels like the long context has always been like an independent work stream of like everything isn't long context. Do you think there's a world where like, you know, we have, there's a ton of other teams hill climbing on a bunch of other random stuff, factuality, whatever it is, reasoning, et cetera, et cetera. Do you think the directional, from a research perspective, from a modeling perspective, is that long context is just fused into every other work stream?
34:21Or do you think there's still, it needs to be an independent work stream because it's just fundamentally different in how you get the model to do useful stuff with long context versus reasoning as a corollary example, perhaps? So I guess my answer will be twofold. First of all, I find it helpful to have an owner for every important capability. But second, I think it's important for the work stream to also provide tools for people outside of this work stream to contribute. Yeah, that makes a ton of sense. I have another follow up around reasoning stuff, and I'm curious how the interplay between reasoning and long context.
35:07We had Jack Ray on and we were both at dinner with Jack last night, and we were talking about random reasoning stuff. Have you been surprised by how much it feels like, and you can correct me if this is wrong, the reasoning capability actually makes long context much more useful? Is that just a normal expected outcome just because the model is spending more time thinking? Or is there some inherent deep connection between reasoning capabilities and long context to make it much more effective? I would say there's a deeper connection. And the connection is that if the next token prediction task improves with increasing context length, then you can interpret this in two ways.
Read the full transcript
35:54One way is to say, hey, I'm going to load more context into the input and predictions for my short answer are going to improve as well. But another way to look at this is say, hey, well, the output tokens, they are very similar to input tokens. So if you allow the model to feed the output into its own input, then it kind of becomes like input. So theoretically, if you have a very strong long context capability, it should also help you with reasoning. Another argument is that long context is pretty important for the reasoning because if you are just going to make a decision by generating one token, even if the answer is binary and it's totally fine to generate just one token, it might be preferable to first generate a thinking trace.
36:54And the reason is simply the architectural. Like if you need to make many logical jumps through the context when making a prediction, then you are limited by the network depth. Because that's roughly the number of attention layers. That's what's going to limit you in terms of the jumps through the context. So you are limited. But now if you imagine that you are feeding the output into the input, then you are not limited anymore. Basically, you can write into your own memory and you can perform much harder tasks than you could by just utilizing the network depth. That's super interesting. You and I have also both related to this reasoning plus long contact story.
37:43You and I have both been pushing for a long time to try to get long output landed into the models. And I think developers want this. I see pings all the time. I'm going to start sending them to you now so that you have to answer this question. But lots of people saying, hey, we want longer than 8 ,000 output tokens. We sort of have this to a certain extent now with reasoning token or with the reasoning models. They have 65 ,000 output tokens with the caveat that a large portion of those output tokens is actually for the model to do the thinking itself versus generating some like final response to the user.
38:14how connected are like the long long context input versus like long context output capabilities like is there any interplay between those two things i feel like for a lot of the like core use case i think that people want is like you know dump in a million tokens and then like refactor that million tokens um do you think we'll get to a world where like those two things are actually like the same capability do you look at them as the same capability or is it like two like completely fundamentally different things from a research perspective? No, I don't think they are fundamentally different. I think the important thing to understand is that straight out of pre-training, there isn't really any limitation from the model side to generate a lot of tokens.
39:01You can just put, say, half a million and tell it, I don't know, copy this half a million tokens and it will actually do it. and we actually tried it, it works. But this capability, it requires very careful handling in the post-training. And the reason why it requires a careful handling is because in the post-training, you also, you have this special end of sequence token. And if your SFT data is short, then what's going to happen is the model is going to see this end of sequence token pretty early in the in a sequence and that's just going to learn like hey like you you're always showing me this token within context x so yeah i'm going to generate this token within context x and stop generation that's what you are teaching me this is actually an alignment problem But one point I want to make is that I feel like reasoning is just one kind of long output tasks.
40:13And for example, translation is another kind. And reasoning, it has a very special format. It packs the reasoning trace into some delimiters. And model actually knows that we're asking it to do the reasoning in there. But for translation, the whole output, not just reasoning trace, is going to be long. This is another kind of capability that we want the model to encourage to produce. So it's just a matter of properly aligning the model, and we are actually working on long output. I'm excited. People want it very badly. I think that gets to a bunch of a broader point around just like how developers should be thinking about best practices for long context and also for RAG potentially as well.
41:06Do you have a general sense of, and I know you gave a bunch of feedback on our long context developer documentation, so we have some of this stuff sort of documented already, but what's your general sense of what the suggestions are for developers as they're thinking about how to most effectively use long context? So I think suggestion number one is try to rely heavily on context caching. So let me explain the concept of context caching. The first time you supply a long context to the model and you're asking a question, it's going to take longer and it's going to cost more. While if you're asking the second question after the first one on the same context, then you can rely on context caching to make it both cheaper and faster to answer.
41:54That's one of the features that we are currently providing for some of the models. And so, yeah, try to rely heavily on this thing. Try to cache the files that the user uploaded into context because it's not only faster to process, but it's going to cost you on average four times less for the input token price. And just to give an example of this, like the most common, and you correct me if this is wrong or not the same mental model that you have, but like the most common application where this ends up being really useful is like the like chat with my docs or like chat with PDF or like chat with my data type of applications where the actual original input context to your point is the same.
42:39And that's one of the, again, correct me if my mental model is wrong, like that's one of the requirements of using context caching is that the original context you supply has to be the same. If for some reason that input context was changing on a request by request basis, context caching doesn't actually end up being that effective because you're paying to store some set of original input context that has to persist from like a user request by user request basis. Yeah, I guess the answer is yes to both. It's important for cases where you want to chat to a collection of your documents, or like some large video, you want to ask some questions on it, or a code base.
43:24And you are correct to mention that this knowledge shouldn't change. or if it changes, then the best way for it to change is at the very end. Because then what we're going to do under the hood is we're going to find the prefix which matches the cached prefix and we're just going to throw away the rest. And sometimes developers ask a question like, where should we put the question? Before the context or after the context? Well, this is the answer. like you want to put it after the context because if you want to rely on caching and profit from cost saving, then that's the place to put it because if you put it at the beginning and if you are intending to put all your questions at the beginning, then your caching is going to start from scratch.
44:19Yeah, that's awesome. That's helpful. Other tips? Anything else besides context caching that folks should be thinking about from a developer perspective? One thing we already touched on, and that's combination with Rack. So if you need to go into billions of tokens of context, then you need to combine with Rack. But also in some applications where you need to retrieve multiple needles, it might still be beneficial to combine with Rack, even if you need much shorter contexts. Another thing which we already discussed is that, well, don't pack the context with irrelevant stuff. It's going to affect this multi-needle retrieval.
45:03Another interesting thing is we touched on the interaction between in-weight and in-context memory. So one thing I must mention is that if you want to update your in-weight knowledge using in-context memory, then the network will necessarily get two kinds of knowledge to rely on. So there might be a contradiction between those two. And I think it's beneficial to resolve this contradiction explicitly by careful prompting. So for example, you might start your question with saying, based on the information above, et cetera. And when you say this, based on the information above, you give a hint to the model that it actually has to rely on in-context memory instead of in-weight memory.
45:55So it resolves this ambiguity for the model. I love that. That's a great suggestion. And your sort of comment about this tension between in-weight versus not. And again, we talked a little bit about this, but how do you think about, from a developer perspective, like the fine tuning angle of this? And the only thing that's maybe more controversial than like, is, you know, long context going to kill rag is, you know, should people be fine tuning at all? And like Simon Willinson has a bunch of threads about this. It's like, does anyone actually fine tune models? Does it end up helping them? How do you think about this from like, would it be useful to do fine tuning and long context for like a similar corpus of knowledge or like does the fine tuning piece potentially lead to like better general outcomes for fine tuning?
46:42How do you think about that interplay? Yeah, so let me maybe elaborate on how fine tuning could actually be used on the knowledge corpus. So what people sometimes do is, well, they get additional knowledge. Let's say you have a big enterprise knowledge corpus, say a billion of tokens. And, well, you could continue training the network just like we're doing with pre-training. so you could apply language modeling loss and you can ask the model to learn how to predict the next token on this knowledge corpus. But you should keep in mind that this way of integrating information, it actually works, but it has limitations.
47:32And one limitation is because you're actually going to train a network instead of just supply the context, you should be prepared for various problems, like you will need to tune hyperparameters, you will need to know when to stop the training, you'll have to deal with the overfitting. Some people who actually try to do that, they reported increased hallucinations from using this process. And they hinted that maybe it's not the best way to supply knowledge information into the network. But obviously, it's also like this technique also has advantages. In particular, it's going to be pretty cheap and fast at inference time because, well, the knowledge is in the weights, so you're just sampling.
48:26But there are also some privacy implications because now the knowledge is cemented into the weights of the network. And if you actually want to update this knowledge, then you're back to the original problem. Like, this knowledge is not easy to update. Like, it's in the weights. So how are you going to do it? You will have to, again, supply this knowledge through the context. Yeah, I think it's such an interesting trade-off problem from a developer perspective about, like, how rapidly you want to be able to update the information. I think the cost piece of it is, like, it's not cheap to just, like, keep paying to, like, I feel like RAG is actually, like, pretty reasonable.
49:07You're paying for a vector database, which I feel like there's a lot of offerings and that's reasonably efficient to do at scale. But I think like continuously fine tuning new models is like oftentimes potentially not not cheap, which is super. Yeah. A lot of interesting dimensions to take into account. I'm curious about the sort of long term direction from a fine tuning or not from a fine tuning from a long context perspective. Like what can folks look forward to in the next like three years for long context from maybe an experience perspective? But like will people, will we even talk about long context in three years?
49:41Will it just be like the model does this thing and I don't need to care about it and like it just works? Or yeah, how are you thinking about this? So I'll make a few predictions. What I think is going to happen first is the quality of the current one or two million context is going to increase dramatically. And we are going to max out pretty much all the retrieval-like tasks quite soon. and the reason i think it's going to be the first step is because well you could say like hey but why don't we extend the context why stop at one million or two million but the point is that the current million context it's not close to perfect yet and while it's not close to perfect there's a question why do you want to extend it because what I think is going to happen is when we achieve close to perfect million context, then it's going to unlock totally incredible applications.
50:52Like something we could never imagine would happen, like the abilities to process information and connect the dots, it will increase dramatically. This thing, it already can simultaneously take in more information than a human can. Like, I don't know, go watch a one hour video and then immediately after that answer some particular question on that video. Like at what second someone is dropping a piece of paper. You can't really do that very precisely as a human. So what I think is going to happen is these superhuman abilities, they are going to be more pervasive. Like the better long context we have, the more capabilities that we could never imagine are going to be unlocked.
51:46So that's going to be a step number one. The quality is going to increase and we're going to get nearly perfect retrieval. After that, what's going to happen is the cost of long context is going to decrease. And I think it will take maybe a little bit more time, but it's going to happen. And as the cost decreases, the longer context also gets unlocked. So I think reasonably soon we will see the 10 million context window, which is like a commodity. like it will basically be normal for the providers to give 10 million context window, which is currently not the case. When this happens, that's going to be a deal breaker for some applications like coding because I think for one or two million, you can only fit some, I don't know, somewhere between small and medium-sized code base in the context.
52:51but 10 million actually unlocks large coding projects to be included in the context completely. And by that point, we'll have the innovations which enable near-perfect recall for the entire context. This thing is going to be incredible for coding applications, because the way humans are coding, well, you need to hold in memory as much as possible to be effective as a coder, and you need to jump between the files all the time. And you always have this narrow attention span, but LLMs are going to circumvent this problem completely. They're going to hold all this information in their memory at once.
53:47And they are going to reproduce any part of this information precisely. Not only that, they will also be able to really connect the dots. They will find the connections between the files. And so they will be very effective coders. I imagine we will very soon get superhuman coding AI systems. They will be totally unrivaled. And they will basically become the new tool for every coder in the world. And so when this 10 million happens, that's a second step. And going to say 100 million, well, it's more debatable. I think it's going to happen. I don't know how soon it's going to come. And I also think we will probably need more deep learning innovations to achieve this.
54:43Yeah, I love that. One sort of quick follow up across all three of those dimensions is like, how much from your mind is this like hardware story or like the infrastructure story relative to like the model story? Like there's obviously a lot of work that has to happen to like actually serve long contacts at scale, which is why it costs more money to do long contacts, etc, etc. Do you think about this from a research perspective or is it like, hey, the hardware is sort of going to take care of itself, the TPUs will do their job and I can just focus on the research side of things? Oh, well, yeah.
55:16I mean, just having the chips is not enough. You also need very talented inference engineers. and I'm actually, I'm really impressed by the work of our inference team, what they pulled off with the million contacts that was incredible. And without such strong inference engineers, I don't think we would have delivered one or two million contacts to customers. So this is a pretty big inference engineering investment as well. And no, I don't think it's going to resolve itself. Yeah, our inference engineers are always working hard because we always want long contacts on these models and it's not easy to make it happen.
56:14How do you think about the sort of interplay of a bunch of these agentic use cases with long contacts? Do you think is it like a fundamental or enabler of different agent experiences than you could have before? Or like what's the interplay between those two dynamics? Well, this is an interesting question. I think agents can be considered both consumers and suppliers for long contacts. So let me explain this. So the agents to operate effectively, they need to keep track of the last state, like the previous actions that they took, the observations that they made, etc. And of course, the current state as well.
56:55So to keep all these previous interactions in memory, you need longer context. So that's where longer context is helping agents. That's where the agents are the consumers of long context. But there's also another orthogonal perspective is that agents are actually suppliers of long context as well. And this is because packing long context by hand is incredibly tedious. Like if you have to upload all the documents that you want by hand every time, or like upload a video, or I don't know, copy paste some content somewhere from the web. This is really tedious. You don't want to do that. But you want the model to do it automatically.
57:48And one way to achieve this is through the genetic tool calls. So if the model can decide on its own, like, hey, at this point, I'm going to fetch some more information, and then it's going to just pack the context on its own. And so, yeah, in that sense, agents are the suppliers of long context. Yeah, that's such a great example. My two cents and I've had many conversations with folks about this. I think this is actually one of the main limitations of how people interact with AI systems is like your example of like, it's tedious, like it's so tedious, like the worst part about doing anything with AI is like, I have to go and find all the context that might be relevant for the model and like personally bring that context in.
58:36And in many cases, like the context is like already on my screen or on my computer. Like I have the context somewhere, but it's like I have to do all the heavy lifting. So I'm excited for like a, we should, you know, we should build some like long context agent system that just like goes and gets your context from everywhere. I think that would be super, super interesting. And I feel like solves a very fundamental problem, not for, not only for developers, but like from a end user of AI systems perspective, like I wish the models could just go and fetch my context and I didn't have to do it all.
59:10Yeah, MCP for the win. I love that. Nikolai, this was an awesome conversation. Thank you for taking the time. I'm glad we got to do this in person and appreciate all the hard work from you and the Long Context teams. And hopefully we'll have lots more exciting Long Context stuff to share with folks in the future. Yeah, thanks for inviting me. It was fun to have this conversation. Yeah, I love it.
From the publisher
Explore the synergy between long context models and Retrieval Augmented Generation (RAG) in this episode of Release Notes. Join Google DeepMind's Nikolay Savinov as he discusses the importance of large context windows, how they enable Al agents, and what's next in the field.
Chapters:
0:52 Introduction & defining tokens
5:27 Context window importance
9:53 RAG vs. Long Context
14:19 Scaling beyond 2 million tokens
18:41 Long context improvements since 1.5 Pro release
23:26 Difficulty of attending to the whole context
28:37 Evaluating long context: beyond needle-in-a-haystack
33:41 Integrating long context research
34:57 Reasoning and long outputs
40:54 Tips for using long context
48:51 The future of long context: near-perfect recall and cost reduction
54:42 The role of infrastructure
56:15 Long-context and agents
