In short
a16z Podcast Episode Notes: What's Missing Between LLMs and AGI
Episode Overview In this episode, Vishal Misra returns to discuss his groundbreaking research on Large Language Models (LLMs) and the essential components needed to achieve Artificial General Intelligence (AGI). The conversation is led by Martin Casado, who engages Misra in a detailed exploration of how LLMs function, the implications of their design, and the future direction of AI research.
Key Topics Discussed The Nature of LLMs
- LLMs vs Consciousness: Misra emphasizes that LLMs, like those developed by Anthropic and others, while impressive, do not possess consciousness or an inner monologue.
- Pattern Matching vs Intelligence: LLMs operate through pattern matching and correlation but lack true understanding of cause and effect.
Research Insights
- Mathematical Modeling of LLMs: Misra explains his mathematical model that demonstrates how transformers update predictions in a predictable manner when processing new information.
- Bayesian Inference: The conversation highlights that LLMs exhibit behavior resembling Bayesian updating, learning from new evidence in real-time.
Achieving AGI
- Requirements for AGI: Misra outlines two key requirements for reaching AGI:
- Continual Learning: The ability to learn and adapt post-training.
- Causal Understanding: Transitioning from correlation to causation in model design.
Key Takeaways
- LLMs' Limitations: While LLMs can generate impressive outputs, they fundamentally lack the ability to understand context in a deep, causal way.
- Bayesian Updating: Misra's work indicates that LLMs can use Bayesian principles to update their predictions based on new data, akin to human learning processes but without retaining that knowledge in future interactions.
- Human vs Machine Learning: The distinction between human learning (which involves continuous adaptation and a deeper understanding of causality) and LLMs (which are static post-training) is crucial to understanding the limitations of current AI.
Future Research Directions
- Causal Models: There is a need for constructing causal models that can enable AI systems to simulate and make predictions about the world rather than merely correlating data.
- Plasticity and Learning: Research into mechanisms that allow for ongoing learning and adaptation in AI systems is necessary for advancing towards AGI.
Conclusion This episode showcases the depth of current understanding regarding LLMs and the challenges ahead in the quest for AGI. Misra's research provides a compelling framework for thinking about the evolution of AI, emphasizing the importance of causal reasoning and continual learning.
Resources
- Follow Vishal Misra on X: [@vishalmisra](https://x.com/vishalmisra)
- Follow Martin Casado on X: [@martin_casado](https://x.com/martin_casado)
- Access related content and subscribe to the a16z Podcast: [a16z Podcast](https://a16z.com)
--- This markdown document summarizes critical insights and discussions from the podcast episode, ensuring that readers can grasp the essential themes and advancements in AI research presented by Vishal Misra and Martin Casado.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding AGI and Consciousness
0:45 to 1:30
Discussion on the characteristics needed for AGI and the nature of consciousness in LLMs.
“So he set out to build a mathematical model of how LLMs actually function.”
Exploring Misra's Research on LLMs
1:30 to 2:46
Vishal Misra discusses his work on mathematical models for LLMs and their predictive capabilities.
“This is one of my favorite topics, which is how do LLMs actually work?”
The Matrix Model of LLMs
2:46 to 5:18
Explanation of how LLMs use a matrix model to predict token distributions based on prompts.
“And by the time you talk to all the lawyers at ESPN and productionize it, it took a while.”
In-Context Learning Explained
5:18 to 8:04
Detailed description of in-context learning and how LLMs adapt to new examples in real-time.
“And now the LLM is going to sample this next token and pick synthesis or shake.”
Real-World Applications and Insights
8:04 to 14:00
Discussion on the practical implications of using LLMs in specific applications, including cricket stats.
“is protein shake a subset of protein or is it different?”
Bayesian Inference in LLMs
14:00 to 15:10
Discover how LLMs resemble Bayesian updating processes.
“And finally, when I gave the new query, it was like it had almost 100 % probability of getting the right token.”
The Reaction to Bayesian Characterization
15:10 to 17:10
Explore the backlash against the Bayesian characterization of LLMs.
“And then I still remember the WhatsApp text.”
Developing TokenProbe for Insights
17:10 to 19:30
Learn about the creation of TokenProbe and its educational applications.
“which could let you look not only at the probabilities, but also the entropy of the next token.”
The Bayesian Wind Tunnel Concept
19:30 to 22:50
Understand the concept of a Bayesian wind tunnel and its implications.
“You don't fly it and you test it against all sorts of, you know, aerodynamic pressure.”
How LLMs and Humans Differ in Learning
22:50 to 26:50
Examine the differences between human learning and LLM behavior.
“Give it a task where we know what the answer is.”
Show all 19 chapters
Causation vs. Correlation in Deep Learning
26:50 to 28:00
Explore the limitations of deep learning in causal inference compared to human cognition.
“I mean, I trained it for 150 ,000 steps, and the accuracy was 10 to the power minus three bits.”
Understanding Causation and Deep Learning
28:00 to 29:00
Learn about the limitations of deep learning models in causation and intervention.
“Causal models are the ones that are able to do simulations and interventions.”
Shannon Entropy vs. Kolmogorov Complexity
29:00 to 30:20
Explore the difference between Shannon entropy and Kolmogorov complexity in AI.
“Another example, I think, which will make it clear is the difference between, I'll use these technical terms, Shannon entropy and Kolmogorov complexity.”
Challenges in Continual Learning
30:20 to 31:40
Discuss the difficulties of continual learning in AI and the importance of architecture.
“research directions to kind of improve the state of the art.”
AGI: Path to Causation and Plasticity
31:40 to 38:20
Examine the requirements for achieving AGI through causation and continual learning.
“So, you know, to get to what is called AGI, I think there are two things that need to happen.”
The Case Study of Hamiltonian Cycles
38:20 to 42:00
Analyze a case study involving LLMs and their limits in solving complex problems.
“Plasticity, continual learning properly, and building a causal model from, you know, in a more data-efficient manner.”
Exploring Simulation and Causation
42:00 to 43:14
Discussion on how simulation relates to understanding complexity and causation in intelligence.
“But so that's where I think it's my bias.”
Recent Research and Feedback
43:14 to 44:25
Insights into the reception of recent papers and advancements in Bayesian learning with LLMs.
“But here it's making a difference in the way we view intelligence.”
The Future of LLMs and Causal Mechanisms
44:25 to 45:57
Discussion on the next steps for LLMs and the importance of causal frameworks.
“What's next is, you know, these two parallel tracks.”
Transcript
Automatic transcript. May contain errors.0:00Martin Casado:Anthropic makes great products. Plot code is fantastic. Co-work is fantastic. But they are crazy of silicon doing matrix multiplication. They don't have consciousness. They don't have an inner monologue. You take an NLM and train it on pre-1916 or 1911 physics and see if it can come up with the theory of relativity. If it does, then we have AGI.
0:21Vishal Misra:Just today, by the way, Dario allegedly said that you can't rule out that they're conscious. You can rule out that they're conscious.
0:28Martin Casado:to get to what is called AGI. I think there are two things that need to happen.
0:35Vishal Misra:Five years ago, Vishal Misra got GPT-3 to translate natural language into a domain-specific language it had never seen before. It worked. He had no idea why. So he set out to build a mathematical model of how LLMs actually function. The result? A series of papers showing that transformers update their predictions in a precise, mathematically predictable way. In controlled experiments, the models match the theoretically correct answer almost perfectly. But pattern matching is not intelligence. LLMs learn correlation. They don't build models of cause and effect. To get to AGI, Misra argues, we need the ability to keep learning after training and the move from correlation to causation.
1:18Vishal Misra:Martin Casado speaks with Vishal Misra, professor and vice dean of computing and AI at Columbia University. Vishal, it's great to have you in again. Great to be back. This is one of my favorite topics, which is how do LLMs actually work? And I think that, in my opinion, you've done kind of the best work on this, modeling it out. Thank you. For those that did not see the original one, maybe it's probably worth doing just a quick background on kind of what led you to this point. And then we'll just go into the current work that you've been doing. Five years ago, when GPD3 was first released, I got early access to it.
1:56Martin Casado:And I started playing with it. And I was trying to solve a problem related to querying a Cricut database. And I got GPD3 to do in-context learning, few-shot learning. And it was kind of the first, at least to me, it was the first known implementation of RAG, retrieval augmented generation, which I used to solve this problem of querying, getting GPT-3 to translate natural language into something that could be used to query a database that GPT-3 had no idea about. I had no access to GPT-3's internal, but I was still able to use it to solve that problem. So it worked beautifully. We deployed this in production at ESPN in September 21.
2:40Martin Casado:Wow.
2:41Vishal Misra:You did the first implementation of FRAG in 2021?
2:44Martin Casado:No, no, no. In 2020.
2:45Vishal Misra:2020.
2:46Martin Casado:2020, I got it working. And by the time you talk to all the lawyers at ESPN and productionize it, it took a while. But October 2020, we had, well, I had this architecture working. But after I got it to work, I was amazed that it worked. I wanted to understand how it worked. And I looked at the attention is all your deep papers and all the other sort of deep learning architecture papers, and I couldn't understand why it worked. So then I started getting sort of deep into building a mathematical model.
3:20Vishal Misra:And now you've published a series of papers. The first one that I read is the one where you had kind of your matrix kind of abstraction. So maybe we'll talk about that and then we'll talk about the more recent work. So perhaps we'll just start with the first one, which is you're trying to come up with a mathematical model of how LLM works. And you have, which is very helpful to me. And at the time you were actually trying to figure out how in-context learning was working.
3:42Martin Casado:Yes, yeah.
3:43Vishal Misra:And you came up with an abstraction for LLMs, which is basically a very large matrix, and you use that to describe. So maybe you can kind of walk through that work very quickly.
3:50Martin Casado:Sure, yeah. So what you do is you imagine this huge, gigantic matrix where every row of the matrix corresponds to a prompt. Yeah. And the way these LLMs work is given a prompt, they construct a distribution of probabilities of the next token. Next token is next word. So every LLM has a vocabulary, GPT and its variants have a vocabulary for about 50 ,000 tokens. So given a prompt, it'll come up with a distribution of what the next token should be. And then all these models sample from that distribution. So that's the posterior distribution. That's the posterior distribution. That's how LLMs work.
4:28Martin Casado:And so the idea of this matrix is for every possible combination of tokens, which is a prompt, there's a row. And the columns are a distribution over the vocabulary. So if you have a vocabulary of 50 ,000 possible tokens, it's a distribution over those 50 ,000 tokens. And by distribution, it's just the probability. Just the probability, sorry. Just the probability that the next token should be this versus that. So that's sort of the idea. And when you start viewing it that way, it makes things at least clearer to people like me who want to model it what's happening. So concretely, let's say you have an example that, let's say your prompt is just one word, protein.
5:09Martin Casado:So if you look at the distribution of the next word, next token after that, most of the probabilities would be zero. But you'd have non-zero, non-trivial probabilities on, let's say, two words. One is synthesis. The other is shake. Right. And now the LLM is going to sample this next token and pick synthesis or shake. or you as a human will give the prompt protein shake or protein synthesis. Now, depending on whether you pick synthesis or shake, that row looks very different, right? If you pick protein synthesis, the terms that would have a higher probability would be all concerned with biology, right?
5:53Martin Casado:But if you pick protein shake, it'll all be about gyms and exercise and all bodybuilding stuff. So that synthesis or shake completely changes what comes next. So this is an example of, you can say, Bayesian updating. You start with protein. You have a prior that after protein, this is going to happen. As soon as you get new evidence, then the next term is synthesis or shake. You completely update the distribution. so now you can imagine that the whole the entirety of LLMs is this giant matrix where you have every row protein shake protein synthesis the cat sat on the humpty dumpty blah blah blah now given the vocabulary of these LLMs let's say 50 ,000 and the context window so GPD for instance chat GPD the first version had a context window of 8 ,000 tokens if you look at all possible combinations of 8 ,000 tokens and 50 ,000 vocabulary, the number of rows in this matrix is more than the number of electrons across all galaxies, right?
7:05Martin Casado:So there's no way that these LLMs can represent it exactly. Now, fortunately, this matrix is very sparse. Why? Because an arbitrary combination of these tokens is gibberish. We're never going to use that in real life. Also, the columns are also mainly zero. If you have protein, then you won't have lots of, you know, you won't have arbitrary numbers or arbitrary words after that. It's very sparse both in rows and in columns. So in kind of an abstract way, what all these LLMs are doing is coming up with a compressed representation of this matrix. And when you give a prompt, they try to approximate what the true distribution should have been and try to generate it.
7:51Martin Casado:That's what, in my mind at least, it boils down to.
7:55Vishal Misra:Just from my understanding, so if you have a row of protein and then you have one with protein shake, is protein shake a subset of protein or is it different? It's different.
8:10Martin Casado:It's a continuation from.
8:12Vishal Misra:I see.
8:13Martin Casado:Yeah. No, but I'm saying like the actual posterior distribution, is that a subset? You can say it's a subset, right? if you have protein, then protein shake and protein synthesis are all continuations from protein. So both synthesis and shake have non-zero probabilities. So you can, yeah, you can think of it as somewhat a subset.
8:33Vishal Misra:You use this approach to describe how in-context learning works. And so maybe first describe what in-context learning is and then kind of the conclusion that you came from that.
8:43Martin Casado:So 8-Context Learning is when you show the LLM something it has kind of never seen before. You give it a few examples of this is what it wants, this is what you're trying to do. Then you give a new problem which is related to the example that you've shown. And the LLM learns in real time what it's supposed to do and solves the problem.
9:08Vishal Misra:By the way, the first time I saw this, it absolutely blew my mind. I actually used your DSL when I was first learning about it. So maybe the DSL thing is just crazy if this works at all.
9:20Martin Casado:It's absolutely mind-blowing that it works. And so going back to that cricket problem, because in the mid-90s, I was part of a group that had created this cricket portal called Crickinfo. Cricket is a very stat-rich sport. You think baseball multiplied by 1 ,000, and it's at all kinds of stats. and we had created this online searchable database called Stats Guru, where you could search for anything, any stat related to cricket and has been available since 2000. But because you can query for anything, everything was made available and how do you make something like that available to the general public?
9:58Martin Casado:Well, they're not going to write SQL queries. The next best thing at that time was to create a web form. Unfortunately, everything was crammed into that web form. So as a result, you had like 20 dropdowns, 15 checkboxes, 18 different text fields. It looked like a very complicated, daunting interface. So as a result, even though it could solve or it could answer any query, almost no one used it. A vanishingly small percentage of cricket fans use it because it just looked intimidating. And then ESPN bought that site in 2007. I still know people who run the site. didn't I always told him you know why don't you do something about StartsGuru and in January 2020 the editor-in-chief of Crickinfo Sambit Bal he's a friend so he came to New York and we'd gone out for drinks and again I told him you know why don't you do something about StartsGuru so he looks at me and says why don't you do something about StartsGuru he was joking but that idea kind of stayed with me and when GPT-3 was released I thought maybe I could use StartsGuru, use GPT-3 to create a front 10 for stats guru.
11:08Martin Casado:And so what I did was I designed a DSL, a domain-specific language, which converted queries about cricket stats in natural language into this DSL. And to be clear, you created this. It wasn't like part of any training that was online that
11:25Vishal Misra:like GPT could have seen.
11:27Martin Casado:Nothing GPT could have seen. I created it. I thought, okay, this makes sense. So I designed that DSL and then I did that few short learning thing. So I would, So I created about a database of what I would say of 1500 natural language queries and the DSL corresponding to that query. So when a new query came in, somebody is asking a stats question in English. What I would do is I would go through the natural language queries, do a semantic search, pick the most closely matching top few. and then use that natural language query and its DSL and send that as a prefix. Now, GPT-3, if you recall, had a context window of only 2 ,000 tokens.
12:12Martin Casado:So you had to be very judicious about which examples that you picked. But you pick that and then you send the new query and GPT-3 would complete it in the DSL that I had designed, which until milliseconds ago, it had never seen. And I had no access to internals of GPT-3. I had no access to the weights. But still it worked. So that's how.
12:34Vishal Misra:So it's not obvious to me, given your matrix example of like a prompt and then a distribution, how something like in-context learning would work. And so like, I think your first paper tackled this problem. Right. And so maybe you could walk through your understanding of how LLMs do in-context learning.
12:58Martin Casado:Yeah. So when you think about what in-context learning is, is that as you see evidence, so, you know, in the first paper, what I also did was I took this cricket DSL example and I depicted the next token probabilities of the model as it was shown more and more examples. So the first time you show it this DSL, the natural language and the DSL, the probabilities of the DSL tokens were extremely low. Because ZPD3 had never seen this thing. When it saw the cricket question, in its mind, it was trying to continue it with an English answer. So the probabilities that were high were all English words.
13:50Martin Casado:Once it solved my problem where I had the question and the DSL, the next time I had the question in the next row, the probabilities of the DSL token started going up. With every example, it went up. And finally, when I gave the new query, it was like it had almost 100 % probability of getting the right token. So this is an example of in real time, the model was updating its posterior probability. It was updating its knowledge that, okay, I've seen evidence. This is what I'm supposed to do. Now, this is a colloquial way of saying what Bayesian inference is. Bayesian updating basically is you start with a prior.
14:33Martin Casado:When you see a new evidence, you update your posterior. That's the mathematical division. But in English, it's basically you see something, you see new evidence, you update your belief about what's happening. So it was clear to me that LLMs are doing something which resembles Bayesian updating. So in that first paper, I had this matrix formulation and I showed that, you know, what it's doing. It looks like Bayesian updating. Then we can come to the sort of next series of papers. That's right.
15:05Vishal Misra:So, okay, so, I mean, it seemed pretty conclusive to me at that time. And then you went quiet for a while. And then I still remember the WhatsApp text. You said, Martin, I know exactly how these things are working now. Yeah. And then, listen, you dropped a series of papers that kind of broke the internet. Like, you went super viral on Twitter. I mean, people really noticed. And so I want to get to that in just a second. But before that, I remember when your first paper came out, people would be like, you know these things are definitely not Bayesian like you know anything could be considered to be Bayesian but they're not like why do you think that there was this reaction to like you know there's something new they're not Bayesian I mean I felt like there's almost kind of a backlash just because they're being characterized as Bayesian
15:55Martin Casado:I think this whole world of probability and machine learning that there have been camps of Bayesian and frequentists. And I don't want to get in the middle of that sort of political battle, but Bayesian has become like almost like people had a reaction to that. It's part of that war.
16:15Vishal Misra:I see. So it's like the old Bayesian frequentist type battle. Yeah.
Read the full transcript
16:20Martin Casado:So the people just had, oh, no, you can say anything is Bayesian, right? So I said, okay, maybe they have a point. Maybe what we are seeing is not really Bayesian. How do we prove that it's Bayesian? Right. So then first I have to thank you and Andreessen Horowitz for this. You know, when I said that in my first paper I showed these probabilities, it was because OpenAI had in its charge interface this option to display those probabilities. Then they stopped. So we could not peer inside what's happening. For some reason they stopped. OpenAI, I'm not going to get into the open and close, but they stopped.
17:08Martin Casado:So then we developed our own interface, which could let you look not only at the probabilities, but also the entropy of the next token.
17:16Vishal Misra:Was this on top of an open source model?
17:18Martin Casado:Yeah, yeah. So you can load any sort of open source model, but being in academia, we didn't have access to compute. Thanks to your generous donation, we got the clusters to run what's called TokenProbe. So you can go to tokenprobe.ch.columbia.edu. Is it still running? It's still running. It's still running and people come to it. I use it in my classes to get students to do assignments. They write their own DSLs and they say that it really helps them understand how these LLMs work.
17:50Vishal Misra:So my understanding of LLMs came from TokenProbe. Sit there and just look at the distribution as you filled out a prompt. is actually very, very enlightening. So for those of you that are listening, what's the URL again? Tokenprobe.cs.columbia.edu. Yeah, check it out. It's actually a very, very useful way to actually see how the probability distribution gets updated as you fill out a prompt. Right. But then I cheated. Oh?
18:20Martin Casado:I, you know, it was running, but I also had access to the GPUs that were powering it. And then, along with colleagues at Columbia, and one of them now is at DeepMind, we started to sort of think about how do you really prove that it's Bayesian?
18:42Martin Casado:Can you just explain it?
18:44Vishal Misra:I actually don't know the answer to this. Yeah. It seemed to me you proved it in the first paper. Like, what was missing? Well, in the first paper, we showed it. It was empirical. And you could see. I see, I see. You could see. Not a mathematical book, because it was very obvious to me that. Yeah, it was even obvious to me.
19:00Martin Casado:But to convince, you could say, you know, people who dismiss it over anything can be Bayesian. I see, I see. We had to show it precisely mathematically. Got it. So then we came up with this idea, you know, my colleagues, Naman Agarwal and Siddharth Dalal, the series of papers were written with them. We came up with this idea of a Bayesian wind tunnel. Okay. So what's a wind tunnel? Well, wind tunnel in the aerospace industry is where you test an aircraft in an isolated environment. You don't fly it and you test it against all sorts of, you know, aerodynamic pressure. Then you see what will withstand, what kind of altitude, pressure, blah, blah, blah.
19:41And you don't want to do it up in the air testing.
19:45Martin Casado:So we said, OK, why don't we create an environment where we take these architectures and we tested transformers, Mamba, LSTMs, MLPs, all architectures. We said, why don't we create, take a blank architecture, give it a task where it's impossible for the architecture to memorize what the solution to that task should be. The space is combinatorially impossible for given the number of parameters and we took very small models. so it's difficult enough that they cannot memorize it but it's tractable enough that we know precisely what the the Bayesian posterior should be you can calculate it analytically so we gave these models a bunch of tasks where again we show that it's impossible to memorize we trained these models and we found that the transformer got the precise Bayesian posterior down to 10 to the power minus 3 bits accuracy.
20:50Martin Casado:It was matching the distribution perfectly. So it is actually doing Bayesian in the mathematical sense, given a task where it has to update its belief. Mamba also does it reasonably well. LSTMs can do one of the things. So in the papers, we have a taxonomy of Bayesian tasks. Transformer does everything. Mamba does most of it. LSTMs do only partially, and MLPs fail completely. So is this a reflection of the data that it's trained on,
21:23Vishal Misra:or is it more a reflection of the mechanism?
21:27Martin Casado:It's the mechanism. It's the architecture. The data decides what tasks it learns. Right. So in the first paper, we had these Bayesian wind tunnels, and we showed that, you know, it's doing the job. We had different tasks. In the second paper, we show why it does it. So we look at the transformers, we look at the gradients, and we show how the gradients actually shape this geometry, which enables this Bayesian updating to happen. Then in the third paper, what we did, we took these frontier production LLMs, which have open weights so that we could look inside them. And we did our testing, and we saw that the geometries that we saw in the small models persisted in models which are, you know, hundreds of millions of parameters.
22:17Martin Casado:The same signature existed. The only thing is that because they are trained on all sorts of data, it's a little bit dirty or messy. But you can see the same structure. So the whole idea behind the Bayesian wind tunnel was, unlike these production LLMs, where you don't know what they have been trained on, So you cannot mathematically compute the posterior. So again, how do you prove it? I mean, it looks Bayesian, you know, from the first paper. It looks Bayesian, but, you know, so the wind tunnel sort of solved that problem for us. We said, okay, let's start with a blank architecture. Give it a task where we know what the answer is.
22:55Martin Casado:It cannot memorize it. Let's see what it does.
22:59Vishal Misra:So do you think this provides any sort of, like, indication of how humans think? Or do you think that these things are totally independent? No, no, it does provide.
23:08Martin Casado:So, you know, human beings also update our beliefs as we see new evidence, right? So we do, in some sense, Bayesian updating, but we do something more than that. I'll come to that. But these transformers or even Mamba do this Bayesian updating. And but the difference with humans is you know, we will update our posterior when we see some new evidence. But the way our brains have evolved over hundreds of millions of years is our optimization objective has been don't die and reproduce, right? That's been sort of the driving force and our brains have learned to adjust. And so when we see some danger, there's something rustling in that bush, don't go near.
24:04Martin Casado:we know how to react to that danger we know how to save ourselves we internalize that learning and our brain cells or our synapses remain plastic throughout our lifetime what happens with LLMs is once the training is done those weights are frozen when you're doing an inference for instance in context learning or anything during that conversation okay you're doing Bayesian inference but then you forget the next time a new conversation starts with zero context you don't retain any learning that happened in the previous instance so for instance with the cricket DSL that I was doing every invocation of it was fresh it did not remember the last time I sent a query what the DSL looked like so that's one difference between how humans use sort of Bayesian updating, which is we remain plastic all our lives, whereas LLMs are frozen.
25:12Martin Casado:And there's another sort of difference, which if you want me to get... Tell me, yeah, yeah, yeah. So the other difference is, well, first, you know, our objective is don't die, reproduce. LLM's objective is predict the next token as accurately as possible, right? So all these scary stories that you read about that, oh, the LLM tried to deceive and it tried to prevent itself from being shut down. That's not a function of the architecture. That's a function of the training data. It has been fed, you know, articles on Reddit or Asimo or whatever. I mean, just today, by the way,
25:57Vishal Misra:Dario allegedly said that you can't rule out that they're conscious. You can rule out they're conscious.
26:06Martin Casado:I mean, come on. As I said, you know, Anthropic makes great products. Cloud Code is fantastic. Code Work is fantastic. But they are grains of silicon doing matrix multiplication. They don't have consciousness. They don't have an inner monologue. They're not driven by the same objective function. don't die, reproduce, right? They're driven by don't make a mistake on the next token. And that's driven entirely by the training data. You train the LLM with stories of Asimo or Reddit where, you know, to survive it's going to do this or that. It'll reproduce that. So it's a reflection. It's not a mind.
26:48And the results, just to say it for the 10th time,
26:52Vishal Misra:are perfectly visioned. Perfectly, yeah. To the digit. To the digit, yeah.
26:58Martin Casado:I mean, I trained it for 150 ,000 steps, and the accuracy was 10 to the power minus three bits. I could have trained it for, you know, this happened in half an hour. On the infrastructure that you provided for token proprio. In the background, I could use those APUs to train. So thank you again for that. So, no, human beings, coming back to it, we are Bayesian, but we do something else. you know when I throw this pen at you what will you do?
27:28Vishal Misra:Dodge it Why will you dodge it? To avoid being hit
27:33Martin Casado:Avoid being hit but your head is not doing a Bayesian calculation of okay this pen is coming the probability that it hits me it'll cause this much pain or all that What you're essentially doing in your head is you're doing a simulation You see the pen coming and you know that it'll come and hit me. Your mind simulates and you dodge it, right? So all of deep learning is doing correlations. It's not doing causation. Causal models are the ones that are able to do simulations and interventions. So, you know, Judea Pearl has this whole causal hierarchy where the first hierarchy is association. which is you build these correlation models.
28:25Martin Casado:Deep learning is beautiful. It's extremely powerful. I mean, you see every day, all these models are like amazingly good. They do association. The second is intervention in the hierarchy. Deep learning models do not do that. Third is counterfactual. So both intervention and counterfactual, you can imagine it's some sort of simulation. You build a model of, causal model of what's happening. And then you are able to simulate. So our brains do that. The current architectures don't do that. Another example, I think, which will make it clear is the difference between, I'll use these technical terms, Shannon entropy and Kolmogorov complexity.
29:11Martin Casado:Sure. So if you look at the Shannon entropy of the digits of pi, it's infinite. Sure. It's impossible to predict and learn what digit will come after. so that's the definition of Shannon entropy and Shannon entropy sort of tries to build a correlation it tries to learn the correlation deep learning does the Shannon entropy Golmogorov complexity on the other hand is the length of the shortest program which will reproduce the string that is under question now the program to get the digits of pi are very small Thanks to Ramanujam and others. There are all sorts of really small programs that can reproduce it exactly.
29:58Martin Casado:So the Kolmogorov complexity of pi is very small. Shannon entropy is infinite. I think deep learning is still in the Shannon entropy world. It has not crossed over to the Kolmogorov complexity and the causal world.
30:14Vishal Misra:Wow, interesting. So to what extent do you think this provides us research directions to kind of improve the state of the art. So let me just give you a specific example. You talked about human beings don't actually update, you know, the matrix. They don't kind of update their weights. But right now there's a lot of research on continual learning. So does your work provide some guidance of how you might approach those problems? And in particular, I've always had this question, which is we use so much data and so much compute to create these models. Like, is it even reasonable to think that you can update the weights and actually have a meaningful impact, you know, in real time?
31:00Vishal Misra:I mean, it just seems like you just need so much more data in order to do that. So can you start answering these questions?
31:04Martin Casado:You can start answering some of these questions. And one of the misconceptions that exists today is that scale will solve everything. Scale will not solve everything. You need a different kind of architecture. And this continual learning is a difficult problem. you have to balance the fact that you will learn something new against the risk of catastrophic forgetting. Right. Right? Right. If you update the weights and you forget what was important and what you have already learned, then you are, you know, you're not making progress. Then it'll just be some sort of random chaotic model. So to solve that problem is difficult.
31:41Martin Casado:That's one aspect of it. So, you know, to get to what is called AGI, I think there are two things that need to happen. One is this plasticity, which has to be implemented through container learning. Secondly, we have to move from correlation to causation. That's...
32:02Vishal Misra:How much is this similar to what Jan LeCun talks about? So Jan LeCun... Causality planning, you know, predicting how your action would...
32:13Martin Casado:It is related. You know, he's coming at it from a different angle than the J-perm model. Right. But it is related. The other thing is, you know, the first time I came on this podcast, I mentioned this test of AGI. Yeah. The Einstein test. I don't remember. So I said, you know, you take an LLM and train it on pre-1916 or 1911 physics and see if it can come up with the theory of relativity. Yeah. If it does, then we have AGI. I mean, it's a high bar, but, you know, we should have high bars. It won't. And this is the same test that I think Demis mentioned at the India AI Summit a couple of weeks ago.
32:57Martin Casado:It's created a lot of news. But why is that and how is that related to this idea of Shannon versus Kolmogorov? So, at the time of Einstein, there were a lot of clues that Newtonian mechanics, there was something missing. People knew that Mercury's orbit didn't make sense. There was something off about it. Then there were these experiments done, the Michelson-Morley experiments, where they were trying to figure out
33:33Martin Casado:this medium called the ether through which light travels. And they felt that if, you know, you bounce light in different directions, the speed might change and they could detect a change in the speed of light. They tried several experiments. They had really precise instruments which could measure the speed, and they found nothing. They found that the speed of light did not change at all. Then there was a whole issue of black holes. Then gravitational lensing. So there were a lot of these signs that Newtonian mechanics is not really explaining everything. But until Einstein came up with a new representation of the space-time continuum, we were stuck.
34:24Martin Casado:so if you had a model that just looked at correlations and sees all of this you know all of these pieces of individual evidence and put together it would not have come up with the beautiful equation that Einstein came up with you know I'm forgetting exactly what it is g mu v equals 8 pi t mu v something like that where you know the equation of the space-time continuum, the tensor. So he came up with a new formulation. So he kind of rejected the existing axioms. He came up with a very short Kulmogorov representation of the world. One equation. From that equation, everything else follows. Whether you're talking about gravitational waves or black holes or mercury or how GPS works.
35:20Martin Casado:You know, the GPS that we use every day in our phones, it uses the equation of relativity. So does this end up becoming like, you almost have to ignore the majority of previous data in order to do it,
35:39Vishal Misra:which LLMs can't because they're trained on the majority of previous data. It's like you almost have like this kind of data gravity that's pulling you back. It's like everybody said it's X. There's a little bit of evidence that it's Y, but because everybody said it's X, the alum will always say it's X. It'll always say X. It'll treat that Y as an anomaly. Actually, this is a very nice way to say it, which is like, okay, now I get your Shannon entropy versus Kalamama. One of them is like, the total amount of information there that will always be bound to the total amount of information there, which is what happens right now, where you can actually describe another motion, you can describe everything with a shorter description with the new data, which would be a totally different look, which would be like...
36:29Martin Casado:You need a new representation, right? Yeah.
36:32Vishal Misra:You know, another way that I've always thought about these, I thought you articulated it well the last time we talked about it, which is the universe is this very, very complex space. And then, you know, somehow humans map it into a manifold that's less complex. Yeah. And then that gets kind of written down. And then the LLM, so that's kind of some distribution. Some, you know, it's still a very large space, but it's a bounded space. And the LLM learned that manifold. And then they kind of use, you know, Bayesian inference to move up and down that manifold. But they're kind of bound to that manifold.
37:07Vishal Misra:Yeah. And then, again, I don't want to put words in your mouth. And then, but like, what they can't do is generate a new manifold. A new amount, yeah. Which requires understanding the way that the universe works and then coming up with a new representation of the universe.
37:19Martin Casado:And this is what relativity is, right? Yeah, exactly. Einstein had to create a new manifold. Yeah, yeah, yeah. If you just stuck with the old manifold of the Newtonian physics, then you would see these correlations, but you could not come up with a manifold that explained them. So you need to come up with a new representation. So to me, you know, there are lots of definitions of AGI. You know, Turing test, we have already passed that. You know, performing economically useful work. every day you see, you know, LLMs are doing that. Do we? I don't know. No, I mean, they are.
37:51Vishal Misra:I mean, without human intervention?
37:53Martin Casado:No, no, no. So that's different. But still, you know, it's like a car can run faster than humans, right?
37:59Vishal Misra:Yeah, I mean, that's a very shallow definition. Yeah, so all these definitions. Cars do useful.
38:05Martin Casado:You know, maybe, you know, in six months you'll have Cloud or what a Gemini do without intervention, and coding tasks, which are well-defined, well-scoped. That's possible. But to me, AGI will happen when these two problems get solved. Plasticity, continual learning properly, and building a causal model from, you know, in a more data-efficient manner.
38:34Vishal Misra:We are hearing people now talking about, you know, seeing general, like Donald Knuth, for example, in the last few days, right? They had this aha moment, apparently, that kind of went viral on X. So do you think that that suggests that we're seeing generality? No, no, no.
38:53Martin Casado:So that actually, to me, validates what I've been talking about for a while now.
38:59Vishal Misra:How so?
39:00Martin Casado:So if you read what he did with the help of a colleague, he got the LLMs to solve this particular problem of finding Hamiltonian cycles. odd numbers we wouldn't get into that and he got the llms to keep solving for one odd number after the other right what he also got to do is after it found a solution for a particular value of m he made the llm update its memory with exactly what it learned in solving that problem so the llms tried many different things you know something worked update the memory so that's kind of like hacking together plasticity. It's learning what it has done as we went along.
39:47Martin Casado:Again, it's a hacked version of it. You're not changing the weights. You're just sort of improving the context. But as you learned, and even after that, so this whole space of Hamiltonian cycles and the associated math is well represented in the manifolds that these LLMs have been trained on. You just had to find the right connection. And LLMs, I know, compute, you throw enough compute, they will find the right connection. So, Knuth was able to find the LLMs attempts and eventually it needed him to put together what he saw into a solution. It definitely helped him get to the solution, but he had to create the new sort of manifold to come to the solution.
40:41Martin Casado:The LLMs were after a while stuck, right? You read what he's written. I mean, it just hot off the press, I think two days ago. Two days ago, yeah. Two days ago. But eventually he used the solution and he came up with the proof, right? So it's like, you know, it's like Einstein saw all these evidences then he thought what will explain he came up with a causal model so Knuth and his brain is sort of that's in the karmograph
41:14Vishal Misra:is the human
41:15Martin Casado:and the LLMs are extremely efficient at doing the Shannon part of it it found all the solutions by trying various things and learning more and more
41:26Vishal Misra:clever way to decompose it I'm wondering, like, do you think this, again, I'm going to ask the same question again, which is, do you think this provides some sort of insight on, like, the next problem to tackle? Like, is there a mechanism that will get the Kalmogorov complexity or not? Like, is this?
41:43Martin Casado:It tells us which direction to pursue.
41:47Vishal Misra:But clearly not how to do it.
41:48Martin Casado:Not how to do it. But even Kalmogorov complexity has largely remained sort of a theoretical construct. Yeah, for sure. There's no algorithm. There's no, there haven't been practical implementations of finding the shortest program. We know it exists. You know, you can argue about it. But so that's where I think it's my bias. That's where our energy should be focused, not larger models with more tokens.
42:15Vishal Misra:Can you, can you, can you tie the two things? Like, how does that pair with doing simulation or is that simulation totally orthogonal? No, simulation is, is it related, right? So you think basically you do simulation and somehow that is a step towards doing the complexity?
42:35Martin Casado:The simulator is the program that we create. It may not be the perfect program.
42:41Vishal Misra:Oh, I see.
42:42Martin Casado:But in our heads, we create this simulator that when I'm throwing the pen, you know that it's coming at you. And you duck. So you're not computing the probabilities as it goes. But you have, you know, you build an approximate.
42:55Vishal Misra:physical thing versus we were talking more conceptually.
42:58Martin Casado:Conceptually, but it's... And you think those are the same mechanism? It's the same mechanism. Really? Yeah. You have to build a causal model, right?
43:04Vishal Misra:I see.
43:05Martin Casado:For most things, right? So you have to move from correlation to causation. I mean, we've heard this term, you know, ad infinitum. But here it's making a difference in the way we view intelligence.
43:20Vishal Misra:How have the last three papers been received? I don't know.
43:26Martin Casado:Well, the archive versions will hear.
43:28Vishal Misra:Let me tell you, I mean, a lot of great reception. A lot of people read it. I'm just wondering, like, what kind of feedback that you've got.
43:36Martin Casado:I'm getting good feedback, but I'm an outsider in this field, right?
43:40Vishal Misra:This networking guy. I'm a networking guy.
43:42Martin Casado:Why is he writing about, you know, learning and machine learning and deep learning and vision? But people who have actually taken the time to read those papers, I'm getting really good feedback. there was a recent paper by Google Research which tried to teach LLMs by some sort of RLHF to do Bayesian learning properly. And that's going in this direction. I think people are coming around to the view that okay, LLMs are doing Bayesian learning. I know that some people also looked at the Bayesian Vendanel paper, the archive version, and they reproduced the experiments. They just saw what was written and they did the trading and they saw, yeah, this is actually happening.
44:25Martin Casado:So what's next? What's next is, you know, these two parallel tracks. I hope to make progress there. Plasticity and causal.
44:38Vishal Misra:Because today you've taken an existing mechanism and you've created a formal model how it works.
44:45Martin Casado:Yeah.
44:45Vishal Misra:And so now you're actually interested in improving. and creating a new mechanism. Yeah, yeah. And do you think it's an entirely different architecture? Or do you think LLMs are like part of the solution?
44:56Martin Casado:I think LLMs are definitely part of the solution. I see. But there has to be something more.
45:00Vishal Misra:Another mechanism.
45:01Martin Casado:So, you know, I was not interested in sort of cataloging what all these LLMs can do.
45:06Vishal Misra:Yeah.
45:06Martin Casado:I was more interested in why are they and how are they doing it. Yeah. I think now we have a good grip on the why and how. Yeah. And the next step is to, you know, move them to the next level. Now, I think we have a fairly good understanding of what the limits are. Now, how do you
45:26Vishal Misra:go to the next step? Is there an equivalent kind of theoretical framework for causality that applies here? Like similar to like Bayesian for inference?
45:38Martin Casado:Well, the Judea pulse whole causal hierarchy, I think. I think that's the right one. That's a very good one. The whole do calculus approach. I think it's a good way to think about it. You know, the sort of association, intervention, counterfactuals. It takes you from correlation to actually simulation in a mathematical way.
46:03Vishal Misra:That's great. All right, well, listen, really appreciate you coming. This is awesome. So we had you here for the first paper where you had the empirical results. Then we had you back when you actually have like the formal proof. And hopefully the next time you come back, you will have a proposal for the mechanism that actually provides the next step. Hopefully. All right, come on. We're working on it. Thank you for coming in. Thank you for having me.
46:29Vishal Misra:Thanks for listening to this episode of the A16Z Podcast. If you liked this episode, be sure to like, comment, subscribe, leave us a rating or a review, and share it with your friends and family. For more episodes, go to YouTube, Apple Podcasts, and Spotify. follow us on X at A16Z and subscribe to our sub stack at a16z.substack.com. Thanks again for listening and I'll see you in the next episode.
47:13Vishal Misra:For more details, including a link to our investments, please see a16z.com forward slash disclosures.
From the publisher
Vishal Misra returns to explain his latest research on how LLMs actually work under the hood. He walks through experiments showing that transformers update their predictions in a precise, mathematically predictable way as they process new information, explains why this still doesn't mean they're conscious, and describes what's actually required for AGI: the ability to keep learning after training and the move from pattern matching to understanding cause and effect.
Resources:
Follow Vishal Misra on X: https://x.com/vishalmisra
Follow Martin Casado on X: https://x.com/martin_casado
Stay Updated:
Find a16z on YouTube: YouTube
Find a16z on X
Find a16z on LinkedIn
Listen to the a16z Show on Spotify
Listen to the a16z Show on Apple Podcasts
Follow our host: https://twitter.com/eriktorenberg
Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.
Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
