#299 Jacob Buckman: Why the Future of AI Won't Be Built on Transformers

9 Nov 2025 · 57 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Summary: Eye On A.I. Episode #299 - Jacob Buckman: Why the Future of AI Won't Be Built on Transformers

Episode Overview In this episode of *Eye on A.I.*, host Craig S. Smith interviews Jacob Buckman, co-founder of Manifest AI. They delve into new architectures that surpass transformers and tackle issues like memory, context retention, and learning in modern AI models. The discussion revolves around the innovative Power Retention architecture developed by Manifest AI, which aims to improve the handling of long contexts in AI systems.

Key Themes and Concepts

  1. Current Challenges with Transformers
  2. Transformers struggle with long context due to their quadratic scalability costs.
  3. As the context window increases, the computational expense grows quadratically, making it inefficient for long inputs.
  1. Power Retention Architecture
  2. Power Retention combines principles from recurrent models and attention mechanisms to create a more efficient architecture.
  3. It allows for linear growth in computational costs as context increases, making it scalable for longer contexts without loss of performance.
  1. State Space in AI Models
  2. The episode explores the distinction between state and weight in AI models.
  3. *State*: Represents the current context based on inputs.
  4. *Weights*: Fixed parameters updated during training that encapsulate learned knowledge.
  5. Power Retention models maintain a larger state size, allowing for better performance over long contexts.
  1. Applications and Implications
  2. The architecture could revolutionize applications in areas like information retrieval and multi-agent systems.
  3. Potential uses in scientific research, allowing models to synthesize vast amounts of information without losing knowledge over time.
  4. The possibility of continual learning through evolving contexts rather than relying solely on weight updates.
  1. Innovations in Model Training
  2. Buckman discusses post-training and retraining methods which allow existing transformer models to transition into Power Retention models without discarding pre-trained weights.
  3. This approach provides the flexibility and efficiency needed for modern AI applications.

Key Takeaways

  • Long Context Handling: Power Retention could mitigate the limitations of existing models by allowing efficient processing of lengthy inputs, leading to improved performance in various applications.
  • Shift in Learning Paradigms: A move towards updating models via context instead of weights may reduce catastrophic forgetting and enhance continual learning capabilities.
  • Collaborative Development: The project is open-source, encouraging community involvement in refining AI architecture and applications.

Future Directions

  • The discussion concludes with an excitement for further developments in AI architectures that leverage the principles of Power Retention.
  • Buckman invites collaboration from listeners with relevant datasets or enterprise applications to explore the potential of this new architecture further.

Conclusion The episode highlights significant shifts in AI architecture through the lens of Power Retention, suggesting a future where models can efficiently learn, remember, and synthesize vast amounts of information across diverse contexts. The conversation emphasizes the importance of community contributions towards advancing AI technologies and fostering innovations that could redefine interactions with AI.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Mamba and state-based models in general are what I would consider retention models. These models that have this sort of duality where they can be expressed either as recurrent models or as attention models. And once you have this, this model that can be expressed in both of these forms, then it also unlocks a secret third form of the model, which people call the chunked formulation, where you basically combine all the best parts of recurrence and attention to get an architecture that both has tractable growth in cost, It has linear growth in cost as you increase the context, but also is able to really saturate the GPU.

0:37I think that many tasks right now are better suited for in-context learning. Because current models have either very short context or, because of their architectures, a limited ability to actually use the thing in their context, it actually doesn't seem that most people are able to get enough value out of pure in-context learning. And so they have to essentially resort to fine tuning as a way of getting the information into the model. Build the future of multi-agent software with Agency. That's A-G-N-T-C-Y. Now an open source Linux foundation project, Agency is building the Internet of Agents, a collaborative layer where AI agents can discover, connect, and work across any framework.

1:28All the pieces engineers need to deploy multi-agent systems now belong to everyone who builds on agency, including robust identity and access management that ensures every agent is authenticated and trusted before interacting. Agency also provides open, standardized tools for agent discovery, seamless protocols for agent-to-agent communication, and modular components for scalable workflows. collaborate with developers from cisco dell technologies google cloud oracle red hat and more than 75 other supporting companies to build next generation ai infrastructure together agency is dropping code specs and services no strings attached visit agency.org to contribute that's A-G-N-T-C-Y dot O-R-G.

2:37Hey, I'm Jacob. I studied computer science at Carnegie Mellon and then went on to do a PhD focused on AI at Mila in Montreal. And I'm one of the founders of Manifest AI. We're an AI research lab, independent based in New York City, focused on pushing the frontiers of the current paradigm. Understanding that the stack that has taken us this far is not going to be the only pieces needed to get us all the way to the AGI future that everyone is hoping for and looking at the whole field from first principles to make the fundamental breakthroughs needed in order to get to that point. Just a quick question.

3:15How are you funded in that you're a lab without a product at the moment? We ultimately do plan to release products. Once we've made these breakthroughs, we're going to then build technology, build products around the technology that we develop. We're backed by Decibel and True, which are two highly technical VCs based in San Francisco. And we're looking forward to, you know, now that we've made a lot of progress on the architecture side, starting to build our first products around this technology. So we're going to talk about your new architecture that goes beyond transformer models and solves in particular the scaling problem with transformers as you scale the context window scales quadratically because every token has to record its position relative to other tokens.

4:18And as you scale, that becomes untenable or just needs a tremendous amount of compute. Your concept of power retention, which, as you were saying, is a combination of power attention and recurrence, not necessarily recurring neural networks. was that the concept that you started manifest with or is it something that developed once you had started the lab no the specific architecture is something that we developed only after really understanding what the sort of fundamentals of the problem were we started off with just a vision for the future of the field and an understanding of the first breakthroughs that would be needed at minimum to get there.

5:09And ultimately, the angle was based around the specific sort of like technical requirements to solve the problem. But that was sort of downstream. You know, the process of doing research is about understanding and discovering these things. So the insight that we started Manifest with was about scaling loss. And in particular, that certain scaling laws were sort of well-situated by the current paradigm. In particular, the parameter scaling laws is one example, where as you increase the parameter size, the model's performance improves, and the cost to do both training and inference on the model grows linearly with respect to the parameters.

5:49Linear growth is very tractable. This is what has allowed us to push the scale this far. Yes, it's expensive, but you get good returns on the investment because the growth in cost is consistent. But there's a different dimension of scaling. the input size, where this is not the case. And right now with transformers, as you scale up transformers to larger and larger inputs, you end up paying, as you mentioned earlier, a quadratic cost on the size of the input. So if we're talking LLMs, language models, the input is a big prompt, what's called the context window. It's all the text that the model is ingesting, synthesizing in order to produce a response.

6:29And as you increase that context window, The cost to train these models grows quadratically instead of linearly. And it's basically unique among the possible axes of scaling that one might consider in this fact. And if we think that the neural networks of the future are going to be able to process and synthesize data from very large inputs, then just some quick napkin math will make you realize it's completely intractable to be processing these inputs at a quadratic cost. and understanding this and basically making the first sort of like insights towards what a solution would look like was where we were at when we started Manifest.

7:08And since then, it's been about two and a half years and it's been a journey of essentially, both from an empirical and theoretical perspective, understanding exactly what it means to have, where the quadratic cost comes from, why it's important, how it affects learning, how it affects training time, how it interacts with the hardware, right? we're training these models on GPUs. So you need to have something that is very hardware efficient as well. And power retention is the outcome of all of this research. It's not dissimilar in that you're dealing with a state space model, is that right? Then from Mamba, which I've done an episode on earlier, did you start by looking at Mamba or just were state space models attractive to you for other reasons?

7:57Yeah, Mamba and state-based models in general are what I would consider retention models. They're models, actually Mamba 1, less so, but Mamba 2 for sure, are these models that have this sort of duality where they can be expressed either as recurrent models or as attention models. And once you have this, this model that can be expressed in both of these forms, then it also unlocks a secret third form of the model, which people call the chunked formulation, where you basically combine all the best parts of recurrence and attention to get an architecture that both has tractable growth in cost, has linear growth in cost as you increase the context, but also is able to really saturate the GPU, which is what attention is able to do.

8:42And traditional recurrent models, like recurrent neural networks, like LSTMs from the previous generation of recurrent neural networks, were unable to do. So that's sort of the first piece of the puzzle was, and this wasn't a piece that was done only by us, this was done by many groups in parallel, including the Mamba group, to just realize that there exists this sort of like other family of recurrent models that have this attention form duality that enables their really efficient implementation. But that's not enough. The issue that we found with Mamba, and actually with basically all other modern RNNs or modern retention models that we were looking at as we started research is that the state size is too small.

9:23And in particular, the state weight ratio is highly skewed in the direction of the weights. In other words, the state is too small for the weight size of the model. And this is not true for transformers. Transformers, the state of a transformer, it's typically not thought of as a state in the sense of an RNN, but it actually is a very nice identification to understand that the state of a transformer is the KD cache. It's this sort of record of all the tokens that it's seen up until now. And if you consider the KD cache of a transformer to be the state, well, you realize that the KD cache is massive at long contexts.

10:01So transformers actually have a huge, huge state compared to RNNs, even modern RNNs. And this is the downfall of RNNs because state size has scaling laws just like parameter size does. Having a larger state translates into better performance. And so what you end up seeing with a lot of these retention architectures, including Mamba and including Gated DeltaNet and Rockov and various other subquadratic architectures that you might encounter, is that the speed-ups that you get, which you do get at long contexts, come into play at exactly the same point that performance relative to transformers takes a hit.

10:44In other words, by the time you're training on context long enough for them to be meaningfully fast, the performance, because the state is so small, is far worse than transformers anyways. So they end up only being sort of compute optimal at a very narrow range of flop budgets, like very, very small flop budgets. And you basically never want to actually train an architecture like this at scale. What power retention does differently is that the retention part is generic to all state-space models, including, you know, Mamba and whatever else. But the power part, that's the unique part. Power retention is a retention model that has an extra adjustable axis of scale.

11:29You can adjust the state size independently of the parameter count. And this allows us to exploit state scaling as an axis of scale, just like we've already been exploiting parameter scaling. And so we're able to construct models that are retention models. They are RNNs, but they have state sizes that are as large as we want them to be. And this allows us to basically pick the optimal state size always and have actual compute optimal performance for a very, very large range of floppy. Can we back up a little bit for the benefit of listeners that are not familiar with a lot of these terms? First of all, state space models, my understanding is that they record the state of the model at any point in time.

12:25Tell me if I'm wrong, but the weights of all the nodes and a network and all of that is recorded as a state. And that state then goes through a transition as the model learns or operates in inference. Is that right? Can you describe that a little more precisely? Yeah, absolutely. So I think separating between the state and the weight is an important conceptual distinction here. There's sort of two things that inform a model's decision. When the model sort of gives you back a response, there's two things that affect it. One thing is whatever question you asked it. So if you ask a model a question, that is the context.

13:17All the tokens of the question get turned into context, and that context gets summarized into a state. The state is just sort of like a list of numbers that summarize all the context seen so far. And that state obviously impacts what the model's reply is going to be. Then, separately, there's a different source of information. This is basically the entire Internet. And now the entire internet affects the response of the model via the training process. Training is essentially about compressing the knowledge on the internet into the weights, which are just a list of numbers describing all the knowledge the model has about everything.

13:54So it's these two things, the state, which summarizes the specific information in the question, and the weights, which capture all of the knowledge the model has on the internet that combine to produce the response. And state-space models, they maintain a state, and as they see more tokens, for example, as your conversation continues and you go back and forth discussing some topic, they keep updating the state. I see. And the state is the context side. It's not the model weight side. So, yeah. Okay. And then recurrence. Can you do the same with recurrence? Yeah, they're the same. What I just described is the same exact concept of recurrence.

14:38Recurrence is just a function that you can apply over and over again. And that is just exactly what I described with a state. It's a function that you, the input is the state and the new token scene is the thing you're adding to the context. And then the output is the new state and the updated context. So yeah, instance case models are essentially just a modern rebranding of recurrent neural networks. And the power in power retention is, describe that again. The power in power retention refers to a sort of esoteric mathematical operation called the symmetric power. It's basically just a mathematical operation you can do to a list of numbers to turn them into a larger list of numbers with some exploitable symmetries.

15:27And this is basically the key insight is that you can use this operation to make a state larger and larger without introducing any extra parameters. So the details of how it's implemented, while very interesting, I think are not crucial to the larger picture. The important thing is it's just a lever, essentially, a hyperparameter that can be set to adjust the state size of a recurrent neural network. And what's crucial is that it's a very hardware-friendly one. So you can adjust it in a way that doesn't hurt the ability to exploit a GPU. Yeah. And in fact, GPUs have been tooled to work well with state-space models.

16:13Is that right? Yeah. State-space models, one of the main reasons why state-space models are sort of the most powerful family of RNNs is that you're able to get really good performance from GPUs on state-space models specifically. This is not true for all RNNs. For classic RNNs such as LSTM or GRU architectures, there's, for some deep reasons, it's actually basically impossible to get really, really good performance on a GPU. Just the way that GPUs are constructed, they're kind of good at just one thing big matrix multiplications it's not strictly true but first order approximation that's basically what they're good at and there's no way to express uh arbitrary rnm as a big matrix multiplication but for state space models or you know the germany is usually retention models retention models are a special family of rnm for which you can implement them as a couple really big matrix multiplications and therefore get really good hardware utilization.

17:21Through post-training and through inference, the model weights don't change. Is that right? I mean, they're fixed by the training. You operate, you perform inference from them, but the knowledge of the model doesn't change through operation in a transformer model. In this model, do the weights change over time? No, it's the exact same setup. Conceptually, the weights are almost by definition. The weights are basically defined as the things that get updated during training, and the state is in some sense defined as the thing that gets updated as you show it more tokens of context during inference.

18:14So this is basically like the way that these two concepts are partitioned. So no, the weights do not change during training. There might be some other research that blurs the boundaries here, but by and large, this is the set. Okay, so you've now built a model, this power retention model. And what's the name of the model? So we've currently released one model trained with one model whose architecture is based on power retention called Power Coder. And it's a coding assistant at the 3 billion parameter scale, so still fairly small. But what's really cool about it is it's sort of a proof of concept for this idea.

18:54We call it metamorphosis. You start off with a transformer, a pre-trained transformer. And by just doing a little bit of retraining, what you end up with is a power retention model of equivalent performance. What this means is you can actually get the benefits of power retention, such as fast linear cost inference and cheap post-training on large contexts, while still getting the benefits of the pre-trained transformer weights that, for example, large model providers will release. And so you're not actually forced to throw all that very expensive, very useful weight training away. You can actually exploit it even when you're using power retention.

19:35So the point of PowerCoder is sort of as just a proof of concept of this. It's a usable, good code model, very fast inference. You can download it from HuggingFace and play around with it right now. But I think what we're really hoping to do with this is showcase what it looks like to start from a really stalled transformer pre-trained weights and convert it into equivalently performant, but much faster and more flexible power retention there. And that's interesting. When you say start with a transformer model with pre-trained weights, will you use one of the open source models or do you build your own GPT and then work on it from there?

20:22No, we get to start with the open source models. So, for example, we could have started with LAMA. We didn't in this case, but we could have started with, for example, the LAMA 70 billion weights. This is a good, solid transformer model. It has its architecture, it has its weights, and you can use it for inference. And what you can do then, you just take a reasonably small number of GPUs, not nothing, but on the order of dozens of GPUs, not thousands of GPUs, and spend maybe six hours retraining Llama into a power retention variant. And all that means is you start with the regular LLAMA weights and the regular LLAMA architecture, go into the LLAMA architecture code, and remove the call to attention, and put in a call to power retention.

21:10Just that one line swap. And then just start training it. And what you'll see is within just a couple hours of retraining, we'll get performance back to where it originally was when we were just using the LLAMA model off the shelf. But now, of course, it's a power retention model. It's no longer a transformer. And that training, what do you call that? Post-training or fine-tuning? Yeah, we typically call it retraining. Fine-tuning often has a connotation of only training some of the weights. And we find that this works best when you train all of the weights. So we just call it retraining. Okay.

21:50But it doesn't affect the size of the model or the number of parameters. Correct. It's the exact same number for others. Yeah. I mean, what I'm sort of leading to is, and you're talking about how, you know, you're looking for new architectures to get us beyond transformers and toward the Holy Grail. And one of the blockers with transformers, aside from the scaling context window, is catastrophic forgetting. that you can't grow the model's knowledge beyond its original training. You can fine-tune, tweak a little bit, but if you try and do a full retraining or full, you lose a lot of the original knowledge.

22:44Does the power retention model have the same problem? Yeah, it doesn't address that aspect of things. That is just a property of the weight updates. Green descent is making local updates to the weights. And so that means if you make enough local updates to enough of the weights, you'll be in a very different sort of part of function space than where you started out. And I think these days, it's actually not such a problem. People have gotten very good at doing what they call like LoRa, like low rank updates. Or this is what I was referring to earlier with the fine tuning. Essentially, people have gotten good at starting with a base model, a pre-trained model, and only updating a few of the weights or only making small updates to the weights in order to specialize the model on some particular task or data set.

23:36And so doing, they avoid destroying the knowledge in the rest of the network. They're just adding a little bit of specialization for this one particular task that they care the most about. doing real meaningful additional pre-training, like adding a huge amount of extra knowledge to the network, that is still seemingly out of reach of standard techniques. You need to have the original data set still around. And as you do continued training, you would need to basically continue training on the original data set to preserve that knowledge. And then you can start to mix in some new data to obtain that knowledge.

24:14But if you just try to train only on the new data without also in parallel continuing to train on the original data distribution, you will typically forget a lot. But that only happens when you're doing a lot of training on all of the weights. But in your model, if you can increase the context window, I don't know, is it indefinitely? I mean, provided you have enough compute, you can add knowledge in the context window for the model to use an inference, right? Yeah, exactly. I think that many tasks right now are better suited for in-context learning, which is the process you just described, and fine-tuning, which is the process I described a moment ago.

24:58But because current models have either very short context or because of their architectures, limited ability to actually use the thing in their context, it actually doesn't seem that most people are able to get enough value out of pure in-context learning. And so they have to essentially resort to fine tuning as a way of getting the information into the model. But one of those strengths of power retention is that a model trained end-to-end with power retention on long context still has a tractable cost. And these models are not suffering from the problem of only part of the context being genuinely useful.

25:39They're able to get value out of the entirety of their context. So we think that when you start really seeing these power retention models trained at scale on varying interesting data sets, we'll be able to get most of the value that we currently get from fine-tuning just off of pure in-context learning. Just show it some examples in the context at inference time of what you want the model to do and watch it sort of like learn on the fly how to do that correctly. Yeah, that's amazing. And what are some of the applications that you see with this? I'm just thinking people use RAG and all these different pre-training to get around hallucinations and have models make inferences based on proprietary data.

26:31But if you can put it all in the context window and then work with the model, you aren't faced with those engineering problems. Is that right? Yeah, I think it makes things a lot, a lot easier. RAG is essentially a technique for selecting the right tokens to put into the context. So the reason you need to do RAG is because you have to make basically hard decisions about what information do you actually want the model to look at, to synthesize when making its decision. And the larger the context window gets, the less you're forced to rely on RAG because the more you can just dump everything into the context.

27:11and let the model decide for itself what information is most relevant to the decision that it's about to make. Now, this doesn't mean that it will necessarily be able to memorize everything still, so there will probably still be a place for RAG or RAG-like methods going forward, but it will basically relieve a lot of the burden of having to design these essentially ad hoc external similarity functions or search heuristics in order to decide what stuff to put into the context. One of the things that I'm the most excited about actually within this space is information retrieval agents. So rather than having a rad heuristic to decide what information gets thrown to the agent, another way to handle really, really vast or private information sources that the model needs to look at before making its decision is to let the model use something like tool calls to decide for itself which pieces of information it wants to obtain.

28:05Now, the training process becomes one that you could do with, for example, reinforcement learning, of letting the agent learn, given a particular query, what are the right things to search? Given what I found in my search, what is the right thing to search next? What is the next sources of information I should pursue in order to get a complete picture of the response that I'm trying to give? And this is something that requires really, really long context length, because you want to have this entire trajectory be remembered by the agent instead of having to break it up into little pieces where the agent sort of like forgets what's going on in between.

28:41And I think once we move on to that sort of paradigm of information retrieval and we start asking questions to agents that can end to end represent their entire history and access vast troves of private information in order to come up with a good response, will really feel like we've stopped the issues with hallucination. They just won't come up anymore. Because like a human, the model will spend some time doing research before coming to a conclusion, and will actually be able to remember everything that it encounters during that whole research process before giving your response. Build the future of multi-agent software with Agency.

29:21That's A-G-N-T-C-Y. Now an open source Linux foundation project, Agency is building the internet of agents, a collaborative layer where AI agents can discover, connect, and work across any framework. All the pieces engineers need to deploy multi-agent systems now belong to everyone who builds on agency, including robust identity and access management that ensures every agent is authenticated and trusted before interacting. Agency also provides open, standardized tools for agent discovery, seamless protocols for agent-to-agent communication, and modular components for scalable workflows. collaborate with developers from cisco dell technologies google cloud oracle red hat and more than 75 other supporting companies to build next generation ai infrastructure together agency is dropping code specs and services no strings attached visit agency.org to contribute that's agntcy.org yeah but what about the problem of you know the longer the context window the less performant the model and the model gets lost or in the context or you know some of the knowledge in the context it doesn't pay attention to.

31:16What about that issue? I don't know what the word is for that issue. This is something that we attribute mostly to that self-same quadratic cost of long context, but indirectly. So let me explain. Given that model providers want to offer their users long context models, and it's intractably expensive to train a true transformer on long contexts for any meaningful amount of time. What a lot of model providers do is come up with a sort of in-between solution. What they do is, instead of using a true transformer, they use what's called a windowed transformer or a sparse transformer, something that only attends to a small local subset of context rather than actually attending to the entire context window.

Read the full transcript

32:09Oftentimes they'll do what's called hybrid sparse attention, where some layers will be full attention, but the vast majority of the layers will just be this local windowed attention or this sparse attention. And the result is that it's not true that the entire context is being processed via attention. The effective context size is actually much smaller, and because there's these weird sparse attention systems where the attention is sort of heuristically distributed across the input context, there's just going to be some sort of like hot spots and some dry patches throughout the context layer. This is part of the problem.

32:46And then this is compounded by the fact that instead of training on long context data, most of these models are trained mostly on short context data for the vast majority of their train time. And again, this just comes down to the fact that training on long context is really, really expensive. And so the model providers do it as little as they can get away with. And the result is a model that has some limited long context ability. And of course, it gets advertised as a model with context length X, right? Because, you know, the users are typically not running detailed evaluations, and they're not privy to the details of the training or of the architecture that might make somebody raise an eyebrow as to whether this is actually a model that's going to be effective at long context.

33:30They just, you know, take the model provider's word for it. They say it's a million context. Okay, it must be a million context. and then they try to use it on tasks that are a million tokens of useful context long and are surprised when the models can't actually effectively process pieces of the context, right? That it is actually not able to use the whole thing as effectively as able to use the first, say, 32K tokens of context. And this little, like, you know, shell game bait and switch, I think, is ubiquitous. Like, I'm not pointing fingers at any one person in particular, and it's also very understandable, But it is sort of a dirty secret of the industry that nobody's models are actually long context transformers.

34:10The long context models are, they're something else. And what's beautiful about power retention is that it will actually enable end-to-end training at long contexts of every layer of a neural network for the entire training duration. And in the models that we've trained so far, we've never seen any degradation at any points in the context from power retention. but when we train similarly sized sparse attention or windowed attention or hybrid attention models we do actually see similar degradation to what we saw in like for large scale models so we're guessing that this sort of difference is going to persist as we continue to scale things up and once the largest models are power retention based there won't be any downsides to using more types of context.

34:58Wow. And yeah, that has huge implications. I mean, you were talking about agent information retrieval and building up the context for inference, but could it also apply to, you know, there's been a lot of talk about, you know, as the reasoning models have developed about doing original scientific research. If you have essentially an unlimited context, could you load all of, say, one domain scientific knowledge and then ask the model to reason over that or come up with novel hypotheses or something like that? I believe so. And this is one of the reasons why at Manifest, we were so focused on unlocking this problem of large input sizes, like input size scaling.

36:07It's because we view that many of the most critical and important applications of AI, including, of course, scientific research and making discoveries about the physical world, are going to be heavily reliant on synthesizing huge, huge amounts of information in order to reach a conclusion. Like a human scientist, not only are they reading a vast literature, they're also thinking constantly about their research domain, engaging in conversations for hours and hours, for days, for weeks, for years, for an entire career, right? Sometimes it takes 10 years of consistent work on a problem, thinking about it, talking about it, reading about it, reading what other people have done in similar fields in order to reach an actual insightful conclusion.

36:54And this is just something that's not going to be achievable by models whose context length is so constrained. Maybe they can take a little bit of information, the amount of information they can store in their context, and think really hard about it and come to something interesting about that. But the context required to say something deep about the real world is simply too vast, in my opinion, to be expressed in a very short context to reason over. So we're going to need models that can actually genuinely synthesize massive, massive amounts of information before reaching a conclusion in order to unlock these high-value applications.

37:32Yeah. And in terms of continual learning, I mean, it's the catastrophic forgetting, the fixed parameter size. Does this have any implications in solving some of those issues? I know you said it doesn't with catastrophic forgetting. But I'm just thinking if the context window can grow and evolve, in effect, the context window becomes kind of another model, another knowledge base that's talking back and forth with the inference model. Can you talk? Am I off base there? No, you're right on point. I think that people who are very concerned right now about things like catastrophic forgetting are concerned because at the moment there's only one sort of good lever for injecting new knowledge into a model, and it's the weights.

38:36and the reason it's the weights is because the context is just not long enough to inject very much knowledge. You have a giant data set of valuable information. Well, you're not going to be able to put it all into the context and so all you can do is put it into the weights. Putting it into the weights means you're battling against catastrophic forgetting because now the weight updates could destroy pre-existing knowledge in the model. And so now that seems like a big pressing concern. But it's only a concern because we chose to put the new knowledge into the model via the weights. If we instead had decided to put the new knowledge into the model via the state, via the context, there'd be no issue with the fact that there's catastrophic forgetting on the weights because we wouldn't need to update the weights in this web.

39:27And that's what I believe the future direction should be, that we just focus on updating the knowledge of the model via the context. I think something that a lot of people, sort of an analogy that a lot of people use when thinking about AI. They think about the weights, like our brain, and weight updates, like learning things, experiences in the real world, translating into new thoughts in our heads. And I actually think this is completely wrong. I think the proper analogy is that us living our lives, interacting with the world and absorbing thoughts into our heads, I think that is state updates.

40:08I think the whole life that I have lived up till now has been one long sequence of context and my brain, the electrical signals in my brain that control my thought patterns, is basically the state that I have arrived at as the result of all the context I've seen. weight updates had nothing to do with it. Where do weight updates live? In my opinion, the proper analogy here is evolution. The human genome has been shaped by many years, many, many years, of course, of evolution that controls not just like us and our ancestors, but like our even like primordial ancestors and how they see and interact with the world.

40:50And evolution, like grading the sentence, is basically this process of local search. And the outcome of that is we have a damn good brain. And this brain, just like a trained neural network, is capable of implicitly storing a huge amount of knowledge. And most of that knowledge is focused on how to process incoming context to update the state to a correct new state. And I think that is where the field is going to head. We're going to stop worrying about updating the weights with new information. We're going to say, okay, the weights are good. And we're going to think about how can we put the right new information into the state via the context.

41:28Okay. But one thing I don't understand there. So the context can grow and you can have agents adding to the context. The model performs inference based on its original parameters but the new knowledge that's created isn't stored in the model it's stored in the context stored in the state is how i would describe it yeah but but the state is the context the state is derived from the context yeah so there's sort of like uh you start with the context which is like text for example right and you turn that text into the state the state is like a list of numbers, a vector of numbers that's stored internally in the model.

42:20Now, there's actually a perfect analogy here on the other side, which is you have a data set. A data set, for example, all the text on the internet, that's text, right? So the data set is the analogy for the context. And what do you do with this data set? Well, you turn it into weights. Weights are a list of numbers stored on the model that basically tell it how to behave. So the state is a list of numbers just like the weights, and the context is some text just like the data set. Where is the state stored? I guess that's my struggle. It's stored in the same place as the weights, just a file on disk or in the RAM of a GPU.

42:57It's just a list of numbers that the model uses every time it makes a prediction. And so that state evolves with new knowledge, and the model will always have access to that new knowledge. It doesn't disappear the way now if you load a bunch of stuff into the context window of a model and then ask for inference. it performs the inference but when you're done with the operation the model is still where it was and the the context is still uh outside the model yeah exactly this is a product like the reason why these interfaces are set up like this is because contact lengths are very limited so typically when you're interacting with say an ai chat bot right you'll open a new chat and this new chat will be a model whose weights are the pre-trained weights that you're working with and whose state is basically the empty state, or usually the chatbot will come with a small prompt that says, you are a helpful and friendly AI assistant, blah, blah, blah, right?

44:12So that context is injected into the model, updating its state to some initial state. And then you start having a conversation. You say something, the model replies, you say something else, the model replies again, and with each response on both sides, the model state continues to be updated. So now, halfway through your conversation, you have the same weights as you started with, but a state that has been updated to reflect all the content of the conversation that you've just had. And then, finally, the AI gives you the correct answer to your question, you're satisfied, and you end the chat. Now you start a new chat, and in this new chat, you're back to square one on the state.

44:54right it's thrown out the state of the old chat but this is because current models are not designed to be able to constantly update their state at fixed cost current chat bots are based on transformers and transformers unlike the retention models that we develop also mamba rnns like lstms etc transformers are different from all of those because transformers do not keep a fixed size state. The transformer state constantly gets more and more expensive as more of the conversation is had. And as a result, these products are designed to force you to end the chat after a certain amount of time. And so this is how people are used to interacting with AI chatbots and basically what they're familiar with.

45:40And the dynamic this produces is one akin to a consultant, right? When you go to an AI chatbot, you're basically going to somebody who is smart, but has no knowledge of you or your problem, or very, very limited knowledge of you and no knowledge of the problem that you're about to ask them. But what we want is a dynamic that's more like a butler, somebody who has known you your whole life, raised you from a boy, and understands all of your needs, everything you like, all of your dislikes, all the problems you've had in the past, and what the solutions you ended up coming to were, somebody who can understand and preempt you and doesn't need to be re-given a whole bunch of context for every question that you want to ask them because it's all just been retained.

46:24It's just been one long session. There's no opportunity to start a new chat with a new model. You wouldn't want to because the model that you've been talking to already knows so much about you and is actually capable of using that information to better answer your future queries. This is something that just currently doesn't exist. And it is downstream of the fundamental design of the architectures powering the models that people currently use. And once these models switch over to the retention framework, the dynamic will completely change. And so continual learning is then possible. Continual learning in the sense of updating the state with each new piece of information.

47:02Yes. Yeah, but updating the state is also updating its knowledge. And would that be a model that's updating its state while interacting with different users? or will it be an instance of a model that updates its state over its use with a particular user or a particular team or particular enterprise? It's an interesting product design question. Off the top of my head, I would guess that probably the most useful generically modality here is to just have it be updated per user. So each user has one particular copy of the AI with a state that records all interactions previously with that user. However, I can think of definite applications for the other ones.

48:00For example, maybe at a company, it would be useful to have one AI that is the company's AI, and every single person at the company talks to the same AI. This doesn't mean it can sort of only reply to one person at a time, but it also means that it's constantly absorbing knowledge from all the different employees, from all of its conversations. So it can use something that it learned in a conversation that you had with AI to give me a good answer. And at a company like that organization where everyone is driving towards the same goal, I think this could be a very powerful sort of information sharing tool.

48:35It's serving the same role as a project manager. Everybody tells the project manager what they're working on and what their blockers are. And the project manager collects all this information and then sort of broadcasts it out to the people who need it. And if we imagine a single AI whose context is constantly being updated by conversations with not just one, but everyone on the project, you can imagine that AI itself serving a similar role in that it coordinates the team towards the objective. So I think this is a big and very interesting space of product design that just gets unlocked once you have the fundamental capabilities available.

49:11Yeah. Yeah. I mean, it sounds potentially revolutionary. I mean, is that how you feel? Yeah. I'm very excited to have been able to make this much progress on what we feel is probably the most important project in the field. This is the breakthrough that needs to happen. We feel that we've made it. And I'm excited to see how the next couple of years of AI progress pan out as a result. Yeah. And the idea of this evolving state that's accumulating new knowledge, that could be applied. I mean, the state could be a physical state, right? It could be applied to robotics. Yeah, absolutely. That's a domain that we've worked with less.

49:57I think, you know, once you start touching the physical world, things get, they get complicated. But yeah, absolutely. Just as with the sort of like brain analogy that I described a moment ago, I think the very promising way to have robots that can handle sort of like arbitrary or unexpected terrains is to be maintaining a single unified state for the entire continuity of this robot's existence. it can get used to its own limbs right like even if you have multiple uh if you have even if you manufacture multiple instances of the same robot each one is going to be slightly different from the others and each one as uh it progresses across its lifetime is going to accumulate you know like wear and tear on its limbs it's going to have particular idiosyncrasies that the other versions of the robot don't have and crucially that it wasn't trained to have at the beginning of its lifetime.

50:49And I think in order for these robots to basically be able to adapt to this damage, you know, like a human who injures their foot doesn't lose their ability to walk. They just walk with a limp a bit more slowly, but they can still get around, right? And if it gets really bad, maybe they have to get a crutch. And now they have what announced to a completely different walking pattern that they have to learn, but they can learn it and they can learn it because they have their entire memory of lifetime walking experiences, their whole physical intuition up till now as a starting point. And then they're able to learn and adapt on the fly to the new sensations that they get as they keep moving around.

51:25And I think this is going to be something that is going to be very important for the sort of like robotic control applications of the future or to make them really work. Yeah. And the inputs for the context and for the state, it doesn't have to be text. I mean, I'm really interested in the whole Jan LeCun's world models and, you know, where the model is learning from direct experience. I mean, not direct, but through video. I mean, eventually, if it's embodied, it would be direct. Could this apply to this power retention model? Yeah, the underlying tech is completely generic. The way that we get text into the model in the first place is via this thing called embedding, right?

52:17You basically turn each token into a vector of numbers. And you don't have to turn only tokens into vectors of numbers. We see even with multimodal models today, you can also turn an image into a vector of numbers. You can turn a frame of a video into a vector of numbers, a snippet of audio, right? All of these things are just signals that can be converted into numbers and fed into the model. And once they're in the model, that's when the power retention sort of layer gets applied. It's about basically combining different sources of information in a way that's both computationally efficient and very powerful.

52:53And this is going to apply generically across all modalities. I do think there's something very special about text, which is unlike basically every other modality, text is from human brain to human brain. It's a communicative medium instead of a natural medium, something that can be understood instead of just perceived. And as a result, the signal to noise ratio in text is way higher than basically any other modality. If you look at the pixels of an image, just the vast majority of the information contained in those pixels is sort of surface statistics or just like general signals, correlations between things in the environment.

53:35It's not sort of dense knowledge in the way that a paragraph of text would be. Bit for bit, text is just the best source of knowledge. And as a result, it's the way that we can get the most intelligent behavior on the lowest compute budget, because we're spending basically all of that compute budget extracting intelligence, rather than spending a lot of it parsing through, you know, like shapes and colors and other sort of less complex features, less semantically rich features. So I do think that text is going to inevitably be a crucial component of any future model, but I don't think it's going to be limited to text, not by a long shot.

54:16And so you built a proof of concept, you published the paper, done a preprint. The paper, when is the paper? Is it going to be at NeurIPS? Not at NURPS, but probably at the conference after. Okay. And what are you doing? Are you building a larger model now? We're continuing to do these retraining at larger and larger scales. And we're also starting to work with partners in particular domains. I'm looking to apply power retention to various data sets where we think it would be particularly well-suited. And it is sets where long context processing can add a lot of value. And your hope is that this architecture will be taken up.

55:04You're publishing. This is going to be open source. Is that right? It already is. You can just pip install retention and you'll have our whole toolkit of power retention. You'll also have our fast implementation of flash retention and a couple other goodies in there as well. It's all open source on GitHub. The models and model weights are available on Hugging Face. So yeah, we just want people to get these things in their hands, play with them, train big models, do fast inference, and really just see how powerful it can be. It's something that we believe is that no individual lab can do all the work in validating new architecture or something that deserves to be done by the community.

55:53and when did it go up on Hugging Face? a couple weeks ago, I think this was two weeks ago and have you had, is there been a strong reaction or are people still trying to figure out what it is? yeah, people seem to like it they've been playing with it again, this is all more on the proof of concept side still, but it's just something that people can have in their hands and see the power of it so yeah, we've gotten a good amount of traction so far with the initial launch and I'm excited to see how things pick up in the next couple of months to a year as well. Yeah, well, I'm excited too and I'm going to be following it closely.

56:32Is there anything I haven't asked, Jacob, that you think listeners should hear? No, I think you mostly covered it. Yeah, I would encourage any listeners who have interesting long context data sets to please reach out, especially if you have cool enterprise applications to go along with them because we'd love to work with you, training models, seeing how power retention can really add value in lots of different domains.

From the publisher

This episode is sponsored by AGNTCY. Unlock agents at scale with an open Internet of Agents. 

 

Visit https://agntcy.org/ and add your support.

 


Why do today's LLMs forget key details over long context, and what would it take to give them real memory that scales?

 

In this episode of Eye on AI, host Craig Smith explores Manifest AI's Power Retention architecture and how it rethinks memory, context, and learning for modern models. We look at why transformers struggle with long inputs, how state space and retention models keep context at linear cost, and how scaling state size unlocks reliable recall across lengthy conversations, code, and documents. We also cover practical paths to retrofit existing transformer models, how in context learning can replace frequent fine tuning, and what this means for teams building agents and RAG systems.

 

Learn how product leaders and researchers measure true long context quality, which pitfalls to avoid when extending context windows, and which metrics matter most for success, including recall consistency, answer fidelity, task completion, CSAT, and cost per resolution. You will also hear how to design per user memory, set governance that prevents regressions, evaluate LLM as judge with human review, and plan a secure rollout that improves retrieval, multi step workflows, and agent reliability across chat, email, and voice.



Stay Updated:

Craig Smith on X:https://x.com/craigss 

Eye on A.I. on X: https://x.com/EyeOn_AI 



 

More from Eye On A.I.

All 266 episodes
#299 Jacob Buckman: Why the Future of AI Won't Be Built on TransformersEye On A.I. · 57 min
Listen in VO