In short
Engram co-founders Dan Biderman and Jessy Lin discuss memory and continual learning for AI systems that must adapt to new, private, team-specific context. They argue that today’s “test-time” context engineering (long prompts, RAG, tool use) is insufficient for learning evolving company knowledge efficiently, and that key information should be internalized into model weights via continual training.
Guest backgrounds
Dan Biderman comes from neuroscience/AI interests (memory/continual learning in biological constraints). Jessy Lin is a co-founder of Ngram with a focus on continual learning and memory architectures.
Key claims
Models should “always be training” in practice; memory should be internalized selectively (not everything); workspace-level continual learning can reduce inference tokens by 2 orders of magnitude (up to 100x) versus re-reading context; RAG vs weight-updates is an unsolved “internalize vs externalize” problem.
Notable examples
company workspaces in Notion/Microsoft/Harvey; hypothetical “math Olympiad in a week” training vs manual cataloging; reducing repeated token costs by making the model “know” what an employee would recall without searching.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe Challenge of Model Generalization
0:00 to 0:41
Explore how models learn to generalize from training data and private context.
“What about pre-training or even post-training makes it possible for the models to generalize in these magical emergent ways and controlling that process so that a company has a set of private data?”
Engram's Approach to Continuous Learning
1:00 to 1:40
Understand Engram's unique perspective on continuous learning and memory.
“So I think like models today, obviously, you know, a lot of things.”
Context Engineering in AI
1:40 to 3:06
Delve into the significance of context engineering for model performance.
“How do you make the models learn new things and bake them deeply into the weights of the model?”
Engram's Architecture and Implementation
3:06 to 5:29
Learn about the technical aspects of Engram's architecture and product.
“One is that the amount of tokens we will all collectively, individually generate is going to be in the tens of millions of tokens per day soon.”
Trade-offs of Using Bespoke Models
5:29 to 8:06
Examine the trade-offs between bespoke and generalized models in AI systems.
“continuously on the things that people care about.”
The Balance of Memory and Learning
8:06 to 11:19
Discuss the balance between fact memorization and conceptual learning in AI.
“The models are not fully great for them.”
High-Stakes Training in AI
11:19 to 14:00
Consider the importance of high-stakes training scenarios for AI development.
“we would have, you know, databases as its own curriculum and we would have algorithms and the databases is like facts about the world and capitals of whatever, store them or query them.”
Integrating AI Learning Strategies
14:00 to 18:09
Explore various strategies for integrating AI learning in high-stakes domains.
“How fast can it handle these kind of tail extreme things, the same ones that we dream about at night?”
The Challenge of Continual Learning
18:10 to 23:57
Discuss the challenges and breakthroughs needed in continual learning for AI.
“They throw it over the fence to the product team who then prompts or contacts engineers new product surfaces on top of the core models.”
Memory and Its Role in AI Performance
23:58 to 28:00
Understand how memory functions and affects AI model performance and efficiency.
“And what we're saying is like, look, if you're really bitterless and pilled, what you want to do is you want to think, how can I burn more compute?”
Show all 20 chapters
Inference Costs and Knowledge Retrieval
28:00 to 28:32
Learn about the challenges of inference costs in AI and how models should internally manage knowledge.
“And we think models should be the same as well.”
Caching Mechanisms and Model Limitations
28:32 to 30:10
Explore how caching mechanisms impact AI retrieval systems and the limitations of current methodologies.
“And that should be something the model just knows.”
KV Caches and Memory Efficiency
30:10 to 31:38
Discuss the inefficiencies of KV caches in AI and the potential for memory compression techniques.
“KV cache is a monstrosity of the current way of doing things that, you know, think about it.”
Future of Continual Learning in AI
31:38 to 33:16
Examine the prospects of continual learning in AI and its implications for model intelligence.
“I think everybody's waiting to see that, you know, and no matter how sophisticated the context engineering approaches are these days, they're not getting there.”
Evolution of AI Capabilities
33:16 to 34:09
Reflect on significant milestones in AI, including GitHub Copilot and ChatGPT, and their impact.
“Just a couple, like, this is just for fun, like rapid fire questions going off.”
Memory Wallets and Digital Skills Transfer
34:09 to 36:08
Discuss the concept of memory wallets in AI and how they could facilitate skill transfer across jobs.
“And people are working on completely new ways of doing things now.”
Language vs. Vision in AI Progress
36:08 to 38:43
Analyze the surprising dominance of language models over vision models in AI advancements.
“I started a PhD in the SaaS firm in 2007 at Stanford.”
Theoretical Perspectives on Language and Vision
38:43 to 40:58
Explore a theoretical analysis of why language has outperformed vision in AI development.
“Now I'm going to tell you the crackpot theory.”
Knowledge Work and Brain Processing
40:58 to 42:02
Consider the implications of AI in knowledge work and how human cognition relates to AI capabilities.
“And then LLMs are, we're just a really, really smart architecture that's better suited for language than for vision.”
Exploring the Future of AI and Memory
42:02 to 44:15
Discover how personalized models and neural memories may shape the future of AI.
“I internalize just, you know, important things like my emotions to you.”
Transcript
Automatic transcript. May contain errors.0:00Dan Biderman:What about pre-training or even post-training makes it possible for the models to generalize in these magical emergent ways and controlling that process so that a company has a set of private data? How do we make the models learn that just as well as the models know the capital of France or how to write Python? So I think it's a really fun problem to think about.
0:41Dan Biderman:Welcome to Training Data. We are delighted to have Don Biedermann and Jesse Lin, co-founders of Ngram today. Ngram is a neolab focused on memory and continual learning, and two of the hottest topics in all of AI research today. And Sean and I are delighted to dig in on those topics with you today.
0:57Jessy Lin:Awesome. Happy to be here. Great.
0:59Dan Biderman:So maybe to kick off, the Engram website says, we don't see the world through the lens of pre-training or post-training. Our models are always training. What does that mean? So I think like models today, obviously, you know, a lot of things. They're incredibly smart. But we kind of think the bottleneck for making these models more useful these days is not really raw intelligence, but understanding like new and evolving context. So whether it's like, you know, a new task that you're doing or a particular context for, you know, like a job or something like this. How do you bake that into the model weights the same way that, you know, pre-training and post-training bakes that into the model weights very deeply?
1:39Dan Biderman:And this is kind of why we think of ourselves as working on these fundamental problems of memory and continual learning, which are really two sides of the same coin. How do you make the models learn new things and bake them deeply into the weights of the model? And is your premise then that memory as a separate database or separate, you know, thing that you've shoved into the context window is not true memory and it's not true continual learning? I think all of these tools will kind of come together. So these days, like the way that people are solving these problems is with context engineering.
2:07Dan Biderman:So you take like a huge prompt, maybe you like keep talking to the model over many, many turns and hours and, you know, reorganize the context to better understand like what you're trying to do. And we think like these kinds of things like tool use, context engineering will play a part. But I think an under-leverage tool these days is using the same kind of training pipeline or framework or kind of workflow that the frontier labs are using to make these models really good at frontier math or code, but applying that to every kind of domain, every kind of context that you have, like let's say in a company.
2:40Yeah.
2:40Jessy Lin:And to me, it's like as an individual, taking notes and having sticky notes is a very valuable thing. we should never discard this. But whenever we get back to business the next day, we always have some sort of trace of memory in our brain, some new intuition about how things should be and where should we look. So these two things should come together. And current solutions are more kind of externalized memory. And this has two issues. One is that the amount of tokens we will all collectively, individually generate is going to be in the tens of millions of tokens per day soon. So just keeping it and searching through it is going to be, and rereading it's going to be pretty expensive, but it's going to also be pretty hard, pretty confusing for the models, unless we have major, major breakthroughs.
3:23Dan Biderman:Tens of billions of tokens for Sean.
3:26Jessy Lin:That's good. Depends on the day.
3:29Dan Biderman:Could you maybe tell us a little bit about the Engram architecture, the Engram product, and how it works? Yeah, I mean, at a high level, I think what we're trying to do is take any context. Like there's all these different workspaces, let's say. So we're working with partners like Notion and Microsoft and Harvey that have these places where people are doing a lot of work over a long period of time. There's all this context, both in terms of like, you know, documents that you've already written as a team, as well as like now people are interacting with these agents more and more in these products.
4:00Dan Biderman:We're having conversations, giving them feedback and figuring out how to have a model that deeply understands that context. So not just reading the files at test time, but really understanding it the way that an employee that's worked at your company for years has. So you kind of understand at a high level, oh, these are the initiatives across the company. This is the way that we do things. You've studied like how to run the hiring pipeline or how to do this kind of thing within the company and can operate just as well as like anybody else can in the company. And so what we're doing is training per team models within these workspaces that deeply understand those contexts and can improve with time on the things that people care about.
4:43Dan Biderman:So the way that we do this at like a technical level maybe is training these into weights. So we do a lot of like adapter fine tuning. So adapters of many types, like I think people have looked into this for decades at this point, like whether it's lauras or prefixes or, you know, sparse architectures. I think like all of these tools are at our disposal and then figuring out what the right data is. So how do you turn any kind of raw like document or interaction into useful training signal for the model? So, again, we have like a variety of tools now, like supervised fine tuning, you know, RL, you know, on policy distillation, like all of these things that, you know, the field has kind of developed and trying to fit these pieces together into a model that learns continuously on the things that people care about.
5:31Dan Biderman:about.
5:31Jessy Lin:Yeah. And it's not a bet that tools are not there. Our models always work under the assumption that some knowledge is externalized, some tools are always there. But what you need to do is you need to figure out, and that's the hard task, is what needs to be internalized and what can be externalized. And even for stuff that's externalized, many individuals and companies have their own bespoke tools and ways of doing things. Not everyone has the same bash CLI tools that the frontier models are training on and how to get the models to better understand your bespoke setup, I think, is its own interesting thing.
6:03Dan Biderman:And so is the premise then that my Notion agent will be a custom agent that is LoRa fine-tuned or, you know, some way with an adapter tuned so that it's constantly learning on new content that's added into my Notion workspace? Is that the premise?
6:19Jessy Lin:Yeah, and they're working with many models and they're the early users of all the Frontier models and they're probably going to keep doing that. Does this approach
6:27Dan Biderman:work on the frontier models or are the closed frontier models?
6:30Jessy Lin:We need white box access to the weights. So we can partner with companies that have closed source weights and do this with them. But it's easiest for us to do it with open source models. But any model that's a transformer model, we can do our thing to it.
6:49Dan Biderman:And what's the trade-off then when people are comparing the before and after using you? Is it that they're no longer sending so much context? And so the trade-off is like you burn more compute up front to learn your company's way of doing things into the weights. And then you're sending less context to the model on every inference pass. Is that the rough trade-off?
7:10Jessy Lin:Yeah, that's one thing. The fact that you don't have to research things and reread things and the fact that you don't have to write monstrous system prompts, that can give you two orders of magnitude reduction in token inference consumption. It's not like 50%. It can be 100x fewer tokens because many things, especially things that relate to people and teams and organization and priorities, these are things that you can't really find in one document unless you really have it, really regimented and document everything. These kinds of things, the model can kind of implicitly learn by training on some of the data and answer within 100 tokens, what the best frontier models would consume 100 ,000 tokens doing.
7:51Jessy Lin:So these kinds of examples are interesting. And also the quality. There are tasks that are not supernatural for the current generation of the models. And we kind of think there's going to be consistently this gap of like three to six months ahead where there's certain things that are bespoke that people are just exploring. The models are not fully great for them. The models will at some point be great for them. But if you can autonomously learn in a very lightweight way, it will give value in that time in terms of capabilities.
8:19Dan Biderman:Why train on the workspace level versus the individual level, for example?
8:22Jessy Lin:Either is fine for us. It's just easier to start with, you know, teams of people have, you know, are more, you know, disciplined in how they collect contacts and in the amount of contacts they have over years. And it's easy for us to start there. But every person's computer and every person's phone one day is a useful target for our technologies. And in fact, it will be very interesting to go there. We just think the big deposits of information are now in teams of people collaborating in knowledge work.
8:51Dan Biderman:Is it a feature or a bug that there is so much fact memorization basically built into the large language models? And there's a school of thoughts that the models just rote memorizing the fact that the capital of France is Paris is actually a bad thing. And what we would prefer for the models to do is, you know, abstractly learn the concepts of countries and capital cities, but not to memorize all these facts in the weights. And so I'm curious what you think about disentangling memorization versus learning how it's done in the models today and then how you're thinking of approaching it. Yeah, I think it's a really interesting question.
9:27Dan Biderman:Like, to some extent, you kind of need to remember stuff in order to, like, compose them into more complex concepts. I think the thing that's kind of missing is figuring out what's important to remember. And I think even now, when you think about like learning new knowledge, if you look at a lot of these academic benchmarks, it's like, how can we learn very specific facts like, you know, the length of a bridge in this like African country? And that's not something that you really want the models to devote capacity for. And it's not something that we devote capacity to. So I think if you look at human memory, I mean, you can say a lot more about this, but like it's lossy because part of the feature of intelligence is compressing what's important and separating that from what's not important.
10:07Dan Biderman:And so I think like you can't really separate fact learning from like non-fact learning or skill learning as some people would like to think. Like if you take a model and like some people have done this with models where you like strip out, you know, like all the facts and just have it like the pure core or something like this. It's very unnatural as a model. It doesn't know basic things. And you kind of need that. But I think why do you need that? Like, why can't you look up facts and then just have I think if you look at like how the models think if you need to recall basic facts in order to like take the next step in your thinking, you can't get very far.
10:45Dan Biderman:Maybe that's like a high level intuition, but it's part of like the reason why we think training is really important in order to like think more and more complex and deep thoughts about things. You kind of need to internalize something so that you can compose them into more abstract concepts.
11:00Jessy Lin:And there have been efforts before that were hard to scale to try and disentangle the two and pre-train the models in a way that allows it to retrieve and search for things and not internalize them. The recipe we know to hill climb on collectively right now is this fact pre-training step. And I think the magic of, or the mystery of this approach is that, you know, traditionally in CS, we would have, you know, databases as its own curriculum and we would have algorithms and the databases is like facts about the world and capitals of whatever, store them or query them. There's also algorithms of how do you efficiently manipulate information and get some answers in a sample efficient way.
11:43Jessy Lin:And I think the magic of deep learning is that these two things are now mushed together and we need all these smart people and anthropic interpretability to try and break them apart. And I think a lot of what we're seeing now in the adoption of AI into the economy is that these things are gradually separating again, where companies have their own context and they really handle them with care and engineer them with care. And there's a generic model that's completely a stranger through these contexts. And the model is operating on them. But for us, it's clear that there needs to be a certain convergence, at least with some cadence, where the facts and the stories and the details are getting mixed into the model.
12:21Jessy Lin:It has disadvantages as well, because if you have to, you know, capitals of countries are, you know, they can change, but it's not very frequent. But there's many other facts that are changing all the time. And just imprinting them into weights is a challenging thing to do.
12:35Dan Biderman:I see. So you're saying it's a false dichotomy that's trying to separate algorithms from databases here. What really matters is like how to distinguish what's important to remember versus what's not important. Exactly. And it's an open question. I guess it's part of how we dream. And are you guys taking any inspiration from that in terms of ranking? Very, very loosely, I think. Just the idea that that's kind of a phase that's missing, maybe, where you take a context and you deeply internalize it. Right now, it's like everything happens at test time. You look at the context that the user gives you and you do some thinking on the fly.
13:08Dan Biderman:But again, you can't get very far or you can get so far maybe. And you make mistakes along the way. How do you digest that back into the model so that next time you do it, you do it the right way and make even more progress?
13:21Jessy Lin:And what are dreams? Dreams are pretty crazy things to say we want to build an AI that's like our dreams sounds a little bit like a nut thing to do. There's not a lot of coherence there. But what's interesting there is like what happens in our dreams. We see things, we talk to ourselves and we experiment with the affordances of what can we do and can't we do in the world and social situations. And, you know, any any it's heavily biased towards social stuff. Right. So for us, too, with things we're building is, you know, we give the models the time to then go back, retreat from the actual interaction and experiment with its affordances.
13:57Jessy Lin:What can it do in an environment? What does it know? How fast can it handle these kind of tail extreme things, the same ones that we dream about at night? You guys come from academic backgrounds. What's a canonical example that motivates this problem or that's a win so far? Yeah, I have one example. Maybe Jesse can give another one. A hypothetical one, for example, imagine one of the AI labs say OpenAI has to win some math Olympiad in a week time from now. Would they construct a catalog of all the math textbooks and really have people annotate which chapters to get and which graphs to see? Or will they actually collect this, synthesize some training data, launch a training job, see where it lands in five, six days, start evaluating it and stuff like that?
14:48Jessy Lin:So it's obvious for anyone who's trained models that there's a superior way to integrate across the ideas and capabilities and involves this kind of magic of training. And we are clear that this has to happen in those high stake domains of math and coding and cyber and stuff. We just think much of this magic can actually end up in the hands of many more people in interesting ways. Like, why isn't it just the foundation model labs that own the end product here? How do you go between giants?
15:18Dan Biderman:Yeah. So I think the worldview that we have is a bit different from the Frontier Lab worldview, where it's like we want one model that's bigger and bigger, that's more and more intelligent across a variety of domains. Instead, how we see it, we kind of imagine this world where everybody has their own model. A lot of the things that people want to learn are either private, things that will never see the light of day in a post-training data set, or even conflicting, like, oh, the way that I want to do the task is different from how another company or another individual wants to. And I think a lot of these things we're already seeing are hard to train into the models with the same tools that we have used for decades in machine learning, which is you have really clean supervision, you have ground truth reward signals, and you create a nice environment, and you train the model to use the tools to better accomplish this coding task.
16:09Dan Biderman:And instead, a lot of the things that actually happen out in the world are very ambiguous or like it's hard to say like what makes something good and so I think a lot of these things are very specific to individuals and I think very kind of misaligned or not very aligned with how the frontier labs think about the whole training pipeline and what kind of models will exist in the longer term.
16:33Jessy Lin:Yeah and to add to it I think you know what is the peer zero for the frontier labs and some of here are pretty close with them. It's getting to AGI, getting this one generic model that's extremely capable in coding and math and then using it to automate the economy or to solve really hard, you know, long-term problems in cryptography and defense or whatever. And it's pretty clear what needs to happen to push this, you know, more pre-training, bigger models, more data, more RL, more inference time compute, that kind of stuff. That's P0. That's where the majority of expenditure and talent goes. And definitely all of them are thinking about memory and all of them are thinking about continual learning.
17:13Jessy Lin:It's just more of a product kind of effort right now. We think it deserves its own attention. We think breakthroughs need to happen there. And Demis and the Sequoia event about a month ago said pretty clearly that we need new breakthroughs around these topics. And obviously they're thinking about them. We're just focusing exclusively on this. And we think certain things are on incentives of where the data is and who owns the model are pretty interesting. So if you could learn from many humans or organizations at scale without necessarily sending someone work with them shoulder to shoulder, that would be a pretty big unlock.
Read the full transcript
17:49Dan Biderman:And maybe another point on that is like, I think a lot of things need to look different in the world. So one is there needs to be new research breakthroughs. Two is new infrastructure for training small models for everybody rather than one big model, one big run. And then the third, I think, is a different way of combining research and product. So right now, I think there's researchers in these frontier labs. They train the model. They throw it over the fence to the product team who then prompts or contacts engineers new product surfaces on top of the core models. But in this world where the models are always training, I think the inputs that users provide are very intricately tied to what the models learn from, like what the training signal is.
18:31Dan Biderman:And so there needs to be a lot more of a kind of integrated loop between research and product. And so while we're focused on tackling a lot of the core research challenges, and that's our background, I think we're also very focused on how to deploy this as quickly as possible to learn from actual feedback in the real world. What motivated you to work on this problem? I think it's obviously one of the grand challenges in AI. I think everybody's talking about it these days because like the models are so smart. What else is left? You know, it's, I think, learning like at the edges, like learning the remainders of what makes these models useful.
19:05Dan Biderman:It's not just about raw intelligence anymore. It's about like learning new things. And I think it also feels very fundamental because it kind of goes back to really understanding what makes the model so good. So right now the models kind of incidentally know a lot of things from pre-training and we don't really understand why. It's like the internet was just, you know, this gift granted to us where there's like a diverse set of data that contains like all of these different examples of coding and like writing and all these other things. And it just happened that way. And now to figure out how to crack this problem of continual learning, it's about figuring out what about pre-training or even post-training makes it possible for the models to generalize in these magical emergent ways and controlling that process so that, you know, a company has a set of private data.
19:56Dan Biderman:How do we make the models learn that just as well as the models know, like the capital of France or, you know, like how to write Python? So I think it's a really fun problem to think about. And Dan, you came from the neuroscience world, is that right?
20:10Jessy Lin:Yes, yes. So I was initially interested in questions around, you know, consciousness and the human condition and things like that.
20:17Dan Biderman:Are the models conscious?
20:18Jessy Lin:I don't have any advanced thoughts on this more than you would read. I don't think so, but it's important that smart people are thinking about it. I would say I was interested in how humans think, how humans perceive. And as Amos Tversky, the Israeli psychologist, used to say, he's not interested in artificial intelligence. He's interested in natural stupidity. So I would say I started kind of similarly, trying to see how people and animals experience the world. Gradually, you know, my inclinations took me to the stats and AI domains. And there I figured that so many of the same problems of memory and continual learning are really, really urgent.
20:54Jessy Lin:And the kind of solutions we have in the current systems are pretty far from what we have in biology. And I'm not one of these people who would say that the machine should be like, you know, like the animal or the human brain. I don't think so. There's many things computers can do better than us. But human memory has these very different things in it. If you want to store a whole code base or you can use a computer, you don't even need AI on the computer to store everything losslessly and just get it. But the human brain evolved to work in these constraints of information capacity and to have these fuzzy representations that can then be abstracted and form connections and form the next day.
21:33Jessy Lin:Current systems don't really have that beyond the generic pre-training step. And I was really interested in, you know, what are ways to build that in? What are ways to learn from that? This is more of a philosophical question. You know, you mentioned in the brain, there's a bunch of different real estate, different co-processing units, whatever. Modern computer architecture, there's CPUs, GPUs, you know, memory, there's different co-processors. With the bitter lesson, do you think that what's happening is that like LLMs are, you know, converge to say like one coprocessor that's just totally dominant.
22:14Jessy Lin:It's like everything, all compute is going to happen in, you know, the GPU equivalent of like a language model. Or do you think that these models are kind of building a bunch of coprocessors, like, you know, emergently inside the model? Like, you know, and take with memory, Do you think that the models themselves will just build whatever part of the brain equivalent would be that's good at memory? Or do you think there needs to be another standalone architecture?
22:46Dan Biderman:Yeah, is memory an emergent property? Exactly. And almost everything.
22:52Jessy Lin:Is everything that we need in intelligence will just be emergent with better training data and more scaled compute? Yeah, I would say just on a more like a superficial perspective on the current deployment of AI, it's way more than just GPUs and we're seeing all these sandboxes exploding and models operating on other computers trying things. I'm more mean on the model architecture level rather than on the... So other experiments, there have been many previous experiments on different architectures that we contributed to like the state space family and others to try and handle very, very long context more efficiently.
23:26Jessy Lin:The thing with all these methods, it ends up being a trade-off, usually a trade-off between memory and accuracy. And memory not in the behavioral cognitive sense, memory in the computer sense. Instead of having the memory footprint of the transformer attention, which is quadratic in the sequence length, these models have. Some are claiming they have sub-quadratic. Yeah, some are claiming. And some do have it. And some of the best Chinese model have layers that are inspired by those state space architectures and are not quadratic in cost. Thing is, is that in our hands, we find that you always compromise accuracy for this memory.
24:02Jessy Lin:There's no free lunch. And what we're saying is like, look, if you're really bitterless and pilled, what you want to do is you want to think, how can I burn more compute? And how can I burn it on new contexts that I have not seen before? So we're as bitterless and pilled as anyone else. And we are not betting that the overall direction of AGI is going to end anywhere soon. We just think there's more compute to scale. And if I truly want to understand Sean and Sean's work and Sean's context, just like rereading files is not going to make it, especially for a special person like you. We got to train 100 trillion parameters.
24:42Jessy Lin:Cosign.
24:44Dan Biderman:What are you finding that people care most about their models learning? Is it memorizing facts about the organization? Is it remembering like, ah, no, we do CI this way? What are people actually hoping to... And then maybe this feeds into how you do the ranking of memory slots and all that. Yeah. Well, I think if you look at what people are spending their time in the app layer doing these days, it's a lot of just trying to make the model work well for your use case. like, oh, I want the model to like, you know, let's say like design my website with my brand style. Like that's like a, you know, very common example these days, but there's many kinds of different tasks that people do with agents, like learning how to run a workflow or, you know, kind of your particular way of like writing, let's say.
25:32Dan Biderman:So there's many, many kinds of things. And honestly, like, I think when we think about these methods, kind of going back to this distinction between like facts and skills there really is none um i think the methods are kind of
25:44Jessy Lin:agnostic to that yeah to me it's like the the natural thing almost all the app layers are basically you know a frontier model wrapped in a loop with search tools and stuff and what they're all interested in doing with us is finding ways to kind of interface with their data in a way that's you know faster more efficient and also is more contextual so almost all of them, it's like we want to have our firm knowledge be encoded in something that's more efficient that I don't have to research. We want to have the model know in a targeted way who's the person I should triage a thing to. And we're just showing them that with pretty lightweight training, these things can be instinctual to the models.
26:27Jessy Lin:They don't have to have these very involved long REPL loops to solve them. So in a sense, it's a rag killer kind of thing. Again, we can always do rag and we can always retrieve, but that's the thing that people are interested in, interfacing with very large data planes and automating very repetitive things this way.
26:46Dan Biderman:Yeah. And I want to double click on this rag killer thing, and I'm sorry to beat a dead horse. I just don't fully grok it yet. Is the premise that there's some trade-off between doing rag versus updating your model weights? Is it an idea that you should be doing both? Like what types of things should be done in the weights versus what types of things should be externalized to rag?
27:06Jessy Lin:I think it's an unsolved problem. I don't think anyone has answered to it. We're all working on it. It's also the fundamental question of biological memory, what should be internalized versus what not. lot. I do think that things that are like, you know, do you need to internalize the room number in a hotel that you were in like a year ago? Probably no, not in your neural tissue. Probably that's good to write down. But do you need to internalize maybe the password to your home right now? Probably it's useful for the next few years to have that imprinted somewhere. So yeah, how does this translate into like knowledge work and products?
27:44Jessy Lin:This is still something we figure out. And we try to take the approach that we try to use as few heuristics as possible. It's easy to run filters on the data and say, like, I'm going to keep this, discard that, train on this, train on that. But as humans, you know, we watch TikTok and we, you know, get exposed to a lot of garbage. And still the brain is able to learn and not completely go off the rails. And we think models should be the same as well.
28:05Dan Biderman:Yeah. Maybe concretely in the short term, I think a lot of what people are worried about these days is the huge inference costs of running these agents like for days on end. High inference costs a good thing. I mean, consuming tokens for what? Sonia works with fireworks. She really loves Zyan. We love inference.
28:26Jessy Lin:We love inference too.
28:28Dan Biderman:Yeah. So I think it's like in the short term, I think that's the immediate pain point. Like, why are you reading the same files over and over again, you know, even in the same query, but like definitely, you know, across people in the same company, they're running the same queries on the same documents over and over again. And that should be something the model just knows. Like in the same way you ask an employee, they don't, you know, type into the search box like what was I working on yesterday? They just know. But doesn't caching kind of solve that? I think to some extent, yeah. But I think going back to this question of what should be internalized versus what's something you retrieve at test time, I think, again, a lot of it is about building on your knowledge.
29:07Dan Biderman:So if you are always doing RAG, you can't make associations like, oh, I see somebody on the team is doing this kind of research. And I kind of like recall at an abstract level, oh, there's this like related thing that you might want to know about. You didn't even ask about it. Right. But I think like these kinds of associations can only happen in weights because they're not really about, you know, you asked me to search for this. I'm going to search for this.
29:31Jessy Lin:And also, I think the main limitation with retrieval systems in general and in AI specifically is like the problem is not so much what to store and where to put it. the problem is like how to address it, like how to query the thing. Do you know what to look for even? And this involves some sort of intuition that sometimes the models don't have, interestingly enough, they don't know where to look. And especially if you're, you know, limited to the current way of doing things, which is keyword search, it is just easier to scale an RL and least involved in terms of like infra for embeddings and stuff.
30:02Jessy Lin:So yeah, knowing what to search is something that's intuitive and can happen in the weights. And also about caching and inference, Like much of this company started with us taking a deep dive into KV caches and caching. And this is a fascinating thing, right? KV cache is a monstrosity of the current way of doing things that, you know, think about it. A KV cache for a single Wikipedia article for some Taylor Swift or something like this, it will be like 80 gigabytes of HBM memory on the GPU. And an entire LAMA, it's for say a 70B LAMA model. And the entire weights of the model would be about 100 gigabytes.
30:42Jessy Lin:And with some distortion, they remember the entire internet. And how come this thing is so, one thing is so bit efficient. And we have this proof of existence that gradient descent can pack a lot of information in very few numbers. Whereas this KV cache thing, you take a few tens of kilobytes of article and it becomes those 80 gigabytes of brain state. So sure you can cache this, you can load this, you'll have issues with disk to HBM stuff. People are working on it. It's pretty interesting. But what if we can take those 80 gigabytes, spend some compute offline, maybe also in fireworks, but then compress it and make it really, really small so that the thing we load in cache is like a thousand x smaller that would have tremendous implications for how we load things how fast we can do things and what the fidelity of the representation is super interesting yeah what are some of the things that could happen in the next year or two that would be like the chat gpt moment of memory or do you think that that's not how things will play out it's a good
31:49Dan Biderman:question um i don't know i think like the first proof of concept of the thing that people keep talking about with continual learning, which is you have an intern that you can teach things over time and it actually gets better. I think everybody's waiting to see that, you know, and no matter how sophisticated the context engineering approaches are these days, they're not getting there. So I think you need, you know, all of these tools at your disposal to make that happen. But I think it will be something like that, where it's like the model's actually getting smarter. Like, whoa, it's different from yesterday.
32:20Jessy Lin:Yeah. And it's important to say that the chat GPT model was not anticipated. We just read about all the different products, the product directions that certain people had before ChatGP was different. I feel like to me, the example is like, look, if you resign from your job today and your sole mission was to make a model that's better for you, and you would use OpenAI and ThoughtPix and all these frontier models, and you just 24-7 engineer the context right skills, your way to move the needle is very limited as an individual. You'll just be better off waiting for the next version of the model, and you'll take it from there.
32:54Jessy Lin:And we would like to see a future where actually the more time you spend on the thing actually translates to the quality of performance, at least in the things and domains you care about. And this is pretty hard to achieve. And the only reason we think it could be achieved is if you start scaling compute and training on these data without destroying them all, importantly, which is pretty hard. Just a couple, like, this is just for fun, like rapid fire questions going off. It's just memory. When's the last time you reached surprised about something in AI? In any area? When reading about fundraising.
33:30Jessy Lin:A lot of surprises every day. I would say all of us felt, you know, a little bit of a change around the capabilities of the coding agents. That's true. But we've been, you know, dabbling with these things and trying to make them work in more effortful ways before. So it didn't come is a complete surprise. But yeah, I think to me, the main events were GitHub Copilot. That for me was just the main event and chat GPT. And then seeing the agentic stuff, we all anticipated, I think. And different people had different expectations on how far it can go and how long the horizon it can go. But I feel, yeah, we're yet to see something fundamentally different.
34:09Jessy Lin:And people are working on completely new ways of doing things now. But yeah, to me, its model is actually changing in a way that's not harmful and learning new things on the fly that are personally and economically viable. That's interesting.
34:25Dan Biderman:Right now there's this idea of like we're each going to have a token wallet that we're going to bring around to companies or to different apps, different workspaces. Do you think that we're going to end up with like a memory bank, a memory wallet that we're going to move around across the digital world as we go? I think it's an interesting question. I don't know if we've fully figured out what the right kind of like product form factor is in the sense. in a way, even with like ChatGPT memory, let's say, I kind of don't want it to remember across my like personal and work context. Oh, yeah. Like it's like, oh, you know, you might like these sheets because you trained a model on a GPU last week.
35:04Dan Biderman:It's like that's totally irrelevant. And to some extent, it's like because the memory is flawed. But also, I think you do want memory in your, I guess, tools and the products that you use to be separated to have control over that. so I personally think like there needs to be some separation there but I guess to be determined what
35:23Jessy Lin:that might look like yeah and like I think a holy grail is like you go to work and you just burn through all these tokens and you create all this value and somehow you know all the IP and stuff stays with the company but somehow the skills you learned the things you invented your ways of doing things some of them you can take with you as well to your next job in a way that's you know sanitized and not harmful to any other company's IP. So I do think carrying a set of skills will be interesting. We do it in our biology right now, and we just sign NDAs and have ethical rules around it. But I think doing it in the digital world would be pretty interesting and pretty rewarding because it will force each of us to push the frontier and implement AI more deeply in our companies and our individual life and then be rewarded for it.
36:08Jessy Lin:I started a PhD in the SaaS firm in 2007 at Stanford. and AI was boring as hell at the time. It was all statistical learning. And there's basically two areas, like computer vision and NLP. So like vision and language were kind of the two areas. And I think that's still true. In 2012, Alex, that happened. Vision was dominating for six years or whatever. Are you guys surprised that language seems to be, like the language approach seems to be like dominating over vision in progress? question two do you think vision has any chance of coming back how do you think about this yeah
36:48Dan Biderman:I think it is pretty surprising to me I mean some people maybe saw it coming but I think I've always kind of been interested in language as like I don't know I guess like a medium for communication and like so many kind of complex abstract things can be done in language I do think like you know I imagine like in the longer term, language and vision will kind of like combine in this more like unified system where, you know, we kind of like take in inputs from all of these different modalities and like understand them in this abstract way. But yeah. Yeah.
37:24Jessy Lin:To me, like I've never been interested in language. It seemed to me such an advanced capability that, you know, the entire animal kingdom has very different forms of speech and language than how we communicate with ourselves in writing. And I was always, as many other leaders in AI, had this thought that, you know, the natural thing is you have to experience the world, act in it, envision it in action. That will be the key. But then I've, like anyone else, seen the chat GPT moment and went to do some work at Mosaic and stuff like that to learn how the sausage is made on the NLP side. And the thing that's striking is that the language should be pretty hard.
38:06Jessy Lin:Each word has this one hot embedding vector that's as dissimilar to any other word than it is. you know, it's a completely high dimensional space and it's really artificial in a sense. And we learn it with models that are order of magnitude bigger than the best vision models. And still, you know, things work pretty well. I do think there's a lot of juice to be squeezed in an image and video. And I think you guys are doing good investments in this space. But I think the two would keep being interesting in different ways. I mean, that was my lead up. Now I'm going to tell you the crackpot theory.
38:45Jessy Lin:I like and this this podcast is not for me to pontificate for you guys but this is something I've been thinking a lot about and I just you're the right people to share this with I I was pretty shocked that language kind of surpassed vision and I underestimated what was happening with LLMs in like 2018 2019 2020 because I just had this bias towards vision and when I look back on it now like i think what's basically happening is that in biology like vision has a massive fundamental advantage over language in biology and maybe i'm wrong but basically like the bit rate that your brain can process optical data through the eye is and this is my i'm not a biologist this is just kind of my dumb assessment it seems many orders of magnitude greater and And there's a lot of like optical processing that happens like even before you reach, you know, like electrons.
39:49Jessy Lin:And so it's just like the total bit rate that is of training data that's kind of being processed and then making it to your brain. It seems many or so magnitude greater than the audio data where, you know, it's sound waves where sound waves are fundamentally like much slower bit rate than light. Yeah. And then there's almost like an upscaling from the acoustics to electronics, which make it into your brain, where it's like a downscaling from photons to electrons with vision. Whereas in computers today, everything is electronic. So it's kind of like you nerfed vision and you like promoted language where it's like all processing is on the same playing field.
40:37Jessy Lin:It's all electronic. And I just, I think this might, this is like my crazy ass, dumb, non-technical crackpot theory. But I think this might be part of why, just like from an information theory perspective, that like maybe language and vision are on a similar playing field by the time you get to like LLMs. And then LLMs are, we're just a really, really smart architecture that's better suited for language than for vision. um how dumb does this out especially to you don the neuroscientist jesse also has some background in cognitive computational science right so i would say my my point here is like look much of what we're doing in knowledge work we haven't evolved to do right we're sitting on these computers reading these things writing these memos whatever we are not evolved to do this it's new to us our brains are not wired for this still nevertheless it's useful to have lms to do this for us and you know as humans we're heavily vision biased you know other rodents are more olfactory uh biased and i've worked on these things myself before so what's the real estate in the brain that's allocated to vision and you know occipital lobes versus like language areas temporal lobe probably more vision i'll have to check with chat gpt but i think that's the situation you don't know from memory no man i'm externalizing i'm a big rag believer in my personal lifestyle.
42:02Jessy Lin:But I think... In the limit, it's all ragged. I internalize just, you know, important things like my emotions to you. No, just kidding. Sorry.
42:16Jessy Lin:Anyways, yeah. And vision is dominating. When people are training vision language models, they end up... Language ends up dominating the vision content there. But yeah, it's hard to say that because a certain brain is more, you know, biased towards a certain modality doesn't mean necessarily that we're going to more efficiently do it. I do think that efforts on like brain computer interfaces should take this into account. How do you then relay it back to the brain? That's where I think it's really important to think like what real estate do we have there right now? But for knowledge work, it's equally fine if it's text, I think.
42:46Dan Biderman:Last question. If everything was right, what does the world look like in five, 10 years? And then what is Engram's role in it? I think I'm imagining like a world where everyone has their own model that is really different from the other person's model and from the frontier model and all of these kind of serve different purposes. And to have a model that really, you know, I think people often talk about like knowing you, but also like kind of like helping you in the ways that make sense to you personally, whether it's like an individual or a team. I think there's an element of like having different kinds of intelligence everywhere.
43:24Yeah.
43:24Jessy Lin:And to me, actually, it's a variant of the story where like, you know, in neuroscience, we know that memory and navigation are pretty closely related. Same circuits in the brain that, you know, represent landmarks in space are in charge of some, you know, elements of episodic memory and things like this. And for me, I think the company can be the actual LLM interface to the data plane for everyone. So sharing some similarities to great companies like Databricks and Oracle, where we form these memories that happen to be neural memories with models that happen to be personalized and happens to be there's hundreds of millions of them.
44:00Jessy Lin:But they're basically a neural interface to the data plane in a way that's very different from what we know. And it's more efficient. It's more associative. It's not representing the file system as it is. it's representing a brain state of that file system. So that's, for me, a vision.
44:14Dan Biderman:Beautiful vision to end on. Thank you guys so much for coming by to share with you at the building. Awesome. Love it.
44:20Jessy Lin:Thank you guys.
44:42You
From the publisher
Dan Biderman and Jessy Lin, co-founders of Engram, are building a neolab around memory and continual learning, which they call two sides of the same coin. Their contrarian premise: instead of stuffing ever-larger prompts into the context window or bolting on RAG, bake a team's knowledge directly into the model's weights, so it knows your company the way an employee of several years does.
The payoff: matching or beating frontier models while consuming up to 100x fewer tokens. Working with partners like Microsoft, Notion, and Harvey, the team draws on roots in computational neuroscience and state-space architectures to attack what they see as the real bottleneck in AI — not raw intelligence, but memory and continual learning. In contrast to the frontier labs' race toward one ever-bigger model and AGI, Dan and Jessy imagine a world where everyone has their own model — privately trained, always learning, and good at the things you actually care about. The real ChatGPT moment for memory, they argue, is the day your model feels like an intern that genuinely got smarter overnight.
Hosted by Sonya Huang and Shaun Maguire, Sequoia Capital




