Columbia CS Professor: Why LLMs Can’t Discover New Science

13 Oct 2025 · 51 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

a16z Podcast Episode Summary: Columbia CS Professor: Why LLMs Can’t Discover New Science

Podcast Overview The a16z Podcast, produced by the Silicon Valley venture capital firm Andreessen Horowitz, discusses technology, culture trends, and future innovations. In this episode, the focus is on Large Language Models (LLMs) and their limitations in scientific discovery, featuring insights from Columbia CS professor Vishal Misra.

---

Episode Details

  • Title: Columbia CS Professor: Why LLMs Can’t Discover New Science
  • Guest: Professor Vishal Misra, Columbia University
  • Host: Martin Casado
  • Release Date: [Insert Release Date Here]

---

Key Themes and Discussions

Limitations of LLMs

  • Inability to Generate New Knowledge:
  • LLMs are trained on existing data and therefore cannot produce groundbreaking theories or concepts (e.g., Einstein's theory of relativity).
  • AGI (Artificial General Intelligence) would require the ability to create new paradigms and scientific theories, something current LLMs cannot achieve.
  • Understanding of Statistical Manifolds:
  • LLMs navigate a geometric manifold based on training data, representing knowledge in a high-dimensional space.
  • Once they stray from this manifold, they generate hallucinations or nonsensical outputs.

Chain-of-Thought Reasoning

  • Mechanism of LLM Reasoning:
  • LLMs can follow chain-of-thought reasoning by breaking problems into manageable steps they have previously encountered.
  • This method reduces prediction entropy, leading to more accurate outputs.

Insights from Vishal Misra's Work

  • Matrix Model:
  • Misra explains a matrix representation where rows correspond to prompts and columns to tokens in the LLM's vocabulary.
  • This model illustrates how LLMs generate distributions for next tokens and highlights their structural limitations.
  • Bayesian Reasoning:
  • Misra's research links LLM operation to Bayesian learning, where context serves as evidence to generate outputs.
  • The complexity of the tasks impacts the LLM's predictive capabilities, emphasizing the need for contextual richness.

Future of LLMs and AGI

  • Plateauing Progress:
  • Misra suggests that while LLMs have improved, significant breakthroughs (akin to the evolution of the iPhone) may plateau unless new architectural approaches are developed.
  • Need for New Architectures:
  • The conversation hints at the necessity for innovations beyond current LLM capabilities to achieve AGI.
  • Suggestions for architectural advancements include integrating multimodal inputs and enabling approximate simulations.

Implications for Human-Like Intelligence

  • Language and Intelligence:
  • The episode touches on whether language development accelerated human intelligence or vice versa, highlighting a fundamental question in cognitive science.

---

Conclusion and Key Takeaways

  • LLMs are powerful tools for navigating existing knowledge but are limited in their ability to discover new scientific theories.
  • Understanding their limitations through formal modeling can help delineate the bounds of what LLMs can achieve.
  • Future advancements in AI may require a paradigm shift in how we construct models and integrate various forms of knowledge and reasoning.

---

Resources and Links

  • Follow Dr. Vishal Misra on X: [@vishalmisra](https://x.com/vishalmisra)
  • Follow Martin Casado on X: [@martin_casado](https://x.com/martin_casado)
  • Visit [a16z](https://a16z.com) for more episodes and resources.

---

Notes The content discussed in this episode is for informational purposes only and should not be interpreted as legal, business, tax, or investment advice.

---

This structured breakdown allows for a comprehensive understanding of the episode while illuminating core discussions and insights shared by the speakers.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Any LLM that was trained on pre-1915 physics would never have come up with a theory of relativity Einstein had to sort of reject the Newtonian physics and come up with a space-time continuum. He completely rewrote the rules. AGI will be when we are able to create new science, new results, new math. When an AGI comes up with a theory of relativity, it has to go beyond what it has been trained on to come up with new paradigms, new science. That's my definition of AGI. Vishal Mishra was trying to fix a broken cricket stats page and accidentally helped spark one of AI's biggest breakthroughs. On this episode of the A16Z podcast, I talk with Vishal and A16Z's Martin Fasato about how that moment led to retrieval augmentation generation and how Vishal's formal models explain what large language models can and can't do.

0:50We discuss why LLMs might be hitting their limits, what real reasoning looks like, and what it would take to go beyond them. Let's get into it. Martin, I knew you wanted to have Vishal on. What do you find so remarkable about him and his contributions that inspired this? Vishal and I actually have very similar backgrounds. We both come from networking. He's a much more accomplished networking guy than I am. That's a high bar given to you in the field. And so we actually view the world in an information theoretic way. It is actually part of networking. And with all this AI stuff, there's so much work trying to create models that can help us understand how these LLMs work.

1:28And in my experience over the last three years, the ones that have most impacted my understanding, and I think have been the most predictive, are the ones that Vishal has come up with. He did a previous one that we're going to talk about called Matrix, is it? Beyond the black box, but yeah. Beyond the black box. Actually, we should put this in the notes for this, but the single best talk I've ever seen on trying to understand how LLMs work is one that Fishall did at MIT, which Hari Balakrishnan pointed me to, and I watched that. So he did that work, and then he's doing more recent work that's actually trying to scope out not only how LLMs reason, but it has some reflections on humans' reason too.

2:05And so I just think he's doing some of the more profound work in trying to understand and come up with models, formal models for how LLMs reason. On that note, you said his most recent work helped you change how humans think. Why don't you flesh that out a little bit? How did it sort of? Well, okay, so can I just try to take a rough sketch at it and then you just tell me how wrong I am? Go right ahead. You're trying to describe how LLMs work. And one thing that you found is that they reduce a very, very complex multidimensional space into basically a geometric manifold that's a reduced state space.

2:43So it's reduced degrees of freedom, but you can actually predict where in the manifold the reasoning can move to, roughly. So you've reduced the dimensionality of the problem to a geometric manifold, and then you can actually formally specify kind of how far you can reason within that manifold. And the articulation is that we, or one of the intuitions is that we as humans do the same thing, is we take this very complex, heavy-tailed, stochastic universe and we reduce it to kind of this geometric manifold, and then when we reason, we just move along that manifold. Yeah, I think you captured it accurately.

3:20That's kind of the spirit of the work. Wait, can I just hear it in your words? Because I'm a VC, so... You're a VC with an H index of what, 60? True. Yeah, so ultimately what all these LLMs are doing, whether the early LLMs or the LLMs that we have today with all sorts of post-training, RLHF, whatever you do, at the end of the day, what they do is they create a distribution for the next token, right? So given a prompt, these LLMs create a distribution for the next token or the next word, and then they pick something from that distribution using some kind of algorithm to predict the next token, pick it, and then keep going.

4:08Now, what happens because of the way we train these LLMs, the architecture of the transformers, and the loss function, the way you put it is, right, it sort of reduces the world into these Bayesian manifolds. And as long as the LLM is going in, sort of traversing through these manifolds, it is confident. And it can produce something which makes sense. The moment it sort of wears away from the manifold, then it starts hallucinating and starts spotting nonsense. Confident nonsense, but nonsense. So it creates these manifolds. And the trick is the distribution that is produced. you can measure the entropy of the distribution.

4:54Entropy the way Shannon describes it. Shannon entropy. Shannon entropy, not thermodynamic entropy. So suppose you have a vocabulary of, let's say, 50 ,000 different tokens, and you have a distribution, next token distribution over these 50 ,000 tokens. So let's say the cat sat on the, right? If that is a prompt, then the distribution will have a high probability for map or hat or table and a very low probability of, let's say, ship or whale or something like that, right? So because of the way it's trained, it has these distributions. Now, the distributions can be low entropy or high entropy. A high entropy distribution means that there are many different ways that the LLM can go with a high enough probability for all those paths.

5:46Low entropy means that there are only a small set of choices for the next token. And the prompts also you can categorize into two kinds of prompts. One prompt is, as you can say, high information entropy. And one prompt is low information entropy. So the way these manifolds work, the LLMs start paying attention to prompts that have high information entropy and low prediction entropy. So what do I mean by that? So when I say I'm going out for dinner, so when I say I'm going out for dinner, that phrase, the LLMs have been trained, they've seen it a lot and there are many different directions I can go with it.

6:39I can say I'm going for dinner, Tonight I'm going to dinner to McDonald's or I'm going to dinner, blah, blah, blah. There are many different. But when I say I'm going to dinner with Martin Cassado, you know, the LLM, now this is information rich. This is sort of a rare phrase. And now the sort of realm of possibilities reduces because Martin is only going to take me to Michelin star restaurants. I'm not going to go to a McDonald's. You get what I'm saying. The moment you add more context, you make the prompt information rich, the prediction entropy reduces. Yep, yep, yep, yep. And another example that I often cite is...

7:22But just quickly, but what is your takeaway? What is your implication on that? Which is, of course, as you're... So, yeah, so you're... Sorry, I forgot how you described it, but so the more precise you are, the more tokens you are, I presume the less options you have for the next token. Is that correct or not correct? Yeah, yeah, essentially. So you're reducing it to a very specific state space when it comes to confidence in an answer. And this is kind of a manifold that you can go on. And then, I mean, do you have kind of a conclusion of what that means for systems or what that means for reasoning?

8:05or is it just a nice way to articulate the bounds of LLMs? No, there is something, I don't know if I should say profound, but there is something about it which tells what these LLMs can or cannot do. So one of the examples that I often tell is, suppose I ask you what is 769 times 1025. You have no idea. You can have some vague idea, given the two numbers, right? And so in your mind, the next token distribution of the answer is going to be diffuse, right? You don't know. You have maybe a vague guess. If you are mathematically very good, maybe your guess is more precise, but it's still going to be diffuse, and it's not going to be the correct answer.

8:54But if you say, can I write it down and do it the way we have learned multiplication tables, now you know exactly what to do next step. Right? You write 769 and then 1025 and then you know exactly. So at each stage of that process your prediction entropy is very low. You know exactly what to do because you have been taught this algorithm. And by invoking this algorithm saying okay I'm not going to just guess the answer but I'm going to do it step by step. Then your prediction and entropy reduces. And you can arrive at an answer which you're confident of and which is correct. And the LLMs are pretty much the same way.

9:42That's why chain of thought works. What happens with chain of thought is you ask the LLM to do something chain of thought. It starts breaking the problem into small steps. These steps it has seen in the past. It has been trained on. Maybe with some different numbers, but the concept it has been trained on. And once it breaks it down, then it's confident. Okay, now I need to do A, B, C, D, and then I arrive at this answer, whatever it is. Let's zoom back out. I want to get into LLMs, but first, Vishal, maybe you can give more context on your background and how that informs your work here. Okay.

10:22So yeah, as Martin said, my background is very similar to his. We come from doing networking. So my PhD thesis, my sort of early work at Columbia has all been in networking. But there's another side of me, another hat that I wear, which is both an entrepreneur and a cricket fan. I was going to say, don't you own a cricket team or something? I'm a minority owner for your local cricket team, the San Francisco Unicorns. That's right. I'm very proud to have you.

10:57So in the 90s, I was one of the people who started this portal called CrickInfo. And CrickInfo, at one point, it was the most popular website in the world. It had more hits than Yahoo. That was before India came on. That's remarkable. And so, you know, we built cricket at the very start with sport. you think baseball multiplied by a thousand. And we had built this free searchable stats database on cricket called Stats Guru. And this has been available on Cricket 4 since 2000. But because you can search for anything, everything was made available on Stats Guru. And you know, you can't expect people to write SQL queries to query everything.

11:48So how do you, how did we do it? Well, it was a web form, you know, where you could formulate your query using that form. And in the back end, that was translated into SQL query, got the results and got it back. But as a result, that because you could do everything, everything was made available. The web form had like 25 different checkboxes, 15 text fields, 18 different dropdowns. The interface was a mess. It was very daunting.

12:18So, and ESPN acquired CrickInfo in the mid-2006, I think, but they still kept the same interface. And that has always sort of nagged me. And so I still know the people who are on ESPN. Wait, wait, what nagged you? Is that CrickInfo did not have informal language that had a web form for doing queries? That web form was terrible.

12:43because of that only the real nerds use of all the things in the world that bother you the fact that an old website was a web form I appreciate your commitment to aesthetic

12:58so I'm still friendly with the people who run ASP and Cricket and for the editor-in-chief whenever he comes to New York we meet up, we go out for a drink and so he was here in 2000 So now the story shifts to how LLMs and me sort of met. So January 2000, right before the pandemic, he was here. And I again said, why did you do something about StatsGuru? And he looks at me and says, why did you do something about StatsGuru? He was kind of joking, but he thought maybe, you know, I had some ways to fix the interface. So anyway, then the pandemic hit, the world stopped. But in July of 2020, the first version of GPT-3 was released.

13:40And I saw someone use GPT-3 to write a SQL query for their own database using natural language. And I thought, can I use this to fix Stats Guru? So I got early access to GPT-3, you know, getting access those days was difficult, but somehow I got it. But soon I realized that, you know, no, I cannot really do it. Because stats grew, the backend databases were so complex. And if you remember, GPD3 had only a 2048 token context window. There was no way in hell I could fix the complexities of that database in that context window. And GPD3 also did not do instruction following at that time. but then in trying to solve this problem I accidentally invented what's now called RAC where based on the natural language query I created a database of natural language queries and the structured queries I created a DSL which then translated into a REST call to stats guru so based on the new query I would look through my set of natural language queries.

15:00I had about 1 ,500 examples, and I would pick the six or seven most relevant ones. And then that and the structured query I would send as a prefix and the new query, and GPT-3 magically completed it, and the accuracy was very high. So that had been running in production since September 2021. You know, about 15 months before chat GPT came, and, you know, the whole revolution, in some sense started and Rack became very popular. I didn't call it Rack, but this is something sort of I accidentally did in trying to solve that problem for QuickInfo. Now, once I built it, you know, I was thrilled that this worked, but I had no idea why it worked.

15:45You know, I stared at that transformer architecture diagram. I read those papers, but I couldn't understand how or why it worked. So then I started in this journey of developing a mathematical model, trying to understand how it worked. So that's been sort of my journey through this world of AI and LLMs because I was trying to solve this cricket problem. Yeah, amazing. And so maybe reflecting back since the release of GPT-3, what has most surprised you about how LLMs have developed? So what has most surprised me, the pace of development. So GPT-3 was, you know, it was a nice parlor trick and you had to jump through hoops to get it to do something useful.

16:34But starting with the, you know, chat GPT was an advance over GPT-3. And then you had all these things like chain of thought, instruction following. GPT-4 really made it polished. And, you know, the pace of development has really surprised me. Now, you know, when I started working with GPT-3, I could sort of see what its limitations were, what I could make it do, what I couldn't make it do. But I never thought of it as, you know, what these LLMs have become for me now and what have become for millions of people around the world. We treat these models as our coworkers, almost like an intern, that, you know, you're constantly chatting with them, brainstorming, making them do all sorts of work, which we couldn't imagine, you know.

17:24Just when ChatGPT was released, it was nice, it could write poems, it could write limericks, it could answer some hallucinogenic questions. But the capabilities that have emerged now, that pace has been very sort of surprising to me. Do you see progress plateauing? Or how do you, either now or in the near future, how do you see it going? I, yes, in some sense, progress is plateauing. It's like the iPhone, you know, when the iPhone came out, wow, what is this thing? And the early iterations, you know, constantly we were amazed by new capabilities. But the last, you know, seven, eight, nine years, it's maybe the camera got a little bit better or, you know, one thing changed here or memory is more.

18:14but there has been no fundamental advance in what it's capable of. You can sort of see a similar thing happening with these LLMs. And this is not true for just one company and one model. You look at what OpenAI is coming up with or what Anthropik, Google, or all these open source Chinese model or Mistral. The capabilities of LLMs has not fundamentally changed. They've become better, right? They've improved, but they have not crossed into a different realm. So this is something that I really appreciate about your work. And so the thing that really struck me is as soon as these things showed up, you actually got busy trying to have a formal model of what they're capable of, which was in stark contrast to what everybody else was doing.

19:12Everybody else was like, AGI, these things are going to, you know, recursively self-improve. Or they'll say, oh, these are just stochastic parrots, which doesn't mean anything. So everybody had rhetoric. And sometimes this rhetoric was fanciful. And sometimes this rhetoric was almost reductionist. Like, oh, it's just a database, which is clearly not true. And the thing that really struck me about your work is you're like, no, let's figure out exactly what's going on. Let's come up with a formal model. And once we have a formal model, we can reason about what that means. And then, you know, in my reading of your work, I kind of break it up in two pieces.

19:48There's the first one where you basically, you came up with this, you know, matrix abstraction. I think it's worth you talking through. And then you took in-context learning as an example and you mapped it to Bayesian reasoning, which to me was incredibly powerful because at the time, nobody knew why in-context learning worked. So I think it'd be great for you to discuss that. Because again, I think it was the first real kind of formal effect on like how are these things working? And then the more recent work that you're working on now is a kind of more generalized version of what is the state space that these models output when it comes to confidence, which is the manifold that we're talking about before.

20:31So I think it would be great if you just described your matrix model and then how you use that to provide some bounds what in-context learning is doing, what's happening. Okay, so let's start with that matrix abstraction. So the idea behind the matrix is you have this gigantic matrix where every row corresponds to a prompt. and then the number of columns of this matrix is the vocabulary of the LLM, the number of tokens it has that it can emit. So for every prompt, this matrix contains the distribution over this vocabulary. So when you say the cat sat on the, you know, the column that corresponds to mat will have a high probability.

21:25Most of them will be zero. But, you know, reasonable continuations will have a non-zero probability. And so you can imagine that there's this gigantic matrix. Now, the size of this matrix is, you know, if we just take just the old first-generation GPT-3 model, which had a context window of 2 ,000 tokens and a vocabulary of 50 ,000 NEXT tokens or 50 ,000 tokens, then the size of it, the number of rows in this matrix is more than the number of atoms across all galaxies that we know of. So clearly we cannot represent it exactly. Now, fortunately, a lot of these rows do not appear in real life. An arbitrary collection of tokens, you are not going to use that as a prompt.

22:24similarly you saw a lot of these rows are absent and a lot of the column values are also zero right when you say the cat sat on the it's unlikely to be followed by the token corresponding to let's say numbers or you know an arbitrary collection of tokens there will be only a very small subset of tokens that can follow a particular prompt so this matrix is very very sparse but even after that sparsity and even after removing the sort of gibberish prompts the size of this matrix is too much for these models to represent even with the trail in parameters so what in an abstract sense what is happening is the models get trained on certain data from the training set and certain a small subset of these rows, you have reasonable values for the next token distribution.

23:27Whenever you give the prompt something new, then it'll try to interpolate with what it has learned and what's there in the new prompt and come up with a new distribution. But it's basically, so it's more than a stochastic parrot. it is sort of based on this subset of the matrix that it has been trained on. So when I say, you know, I'm going out for dinner with Martin tonight. Now, I'm reasonably sure that it has never encountered that phrase in its training data, right? But it has encountered variants of this phrase. and given that I'm going out with Martin, it can produce a Bayesian posterior.

24:19It uses that evidence that Martin is the one that I'm going for dinner with and it'll produce a next token distribution that will focus on the likely places that we are going. So this matrix, because it's represented in a compressed way, yet the models respond to everything, every prompt. How do they do it? Well, they go back to what they've been trained on, interpolate there, and use the prompt as sort of some evidence to compute a new distribution. Right, so the context of the prompt impacts the posterior distribution. Exactly, yeah. Right, and you mapped to Bayesian learning where the context is the new evidence.

25:10New evidence, exactly. So I'll give you, so for instance, the Cricut example that I spoke about earlier. So I created my own DSL, which mapped a natural language query in Cricut to this DSL, which then I can translate into a SQL query or a REST API, whatever. But getting the DSL is important. Now, these LLMs have never seen that DSL. I designed it. Yeah. Right? But yet, after showing a few examples, It learned it. How did it learn it? And this is in the prompt. You didn't, no training, no post-training. 100 % in the prompt, right? So like it's the way to stand it. Yeah, yeah. This was happening in October 2020.

25:55I had no access to internals of OpenAI. I could just access the API. OpenAI had no access to internal structure of Stats Guru or the DSL that I cooked up in my head. Yet after showing it only a few examples, it learned it right away. So that's an example where it has seen DSLs or structures in the past. And now using this evidence that I show, okay, this is what my DSL looks like. Now a new natural language query, it is able to create the right posterior distribution for the tokens. That map to the example that I've seen. Now, the other beautiful thing about this is, this is an example of few-shot learning or in-context learning, right?

26:43But when I give that prompt along with these examples to this LLM, I'm not saying to the LLM, okay, this is an example of few-shot learning, so learn from these examples, right? You just pass this to the LLM as a prompt and it processes it exactly the way it would process any other prompt, which is not an example of in-context learning. So that really means that the underlying mechanism is the same. Whether you give a set of examples and then ask it to complete a talk, a task like an in-context learning, or just give it some prompt for continuation that I'm going out for dinner with Martin tonight.

27:27There's no in-context learning there. But the process with which it's generating or doing this inferencing is exactly the same. And that's what I have been trying to model and come up with a formal model of. What I've found very impressive is you've used this basic model to show a number of things, right? To describe context learning and to map the Bayesian learning. But you did it for another one where you kind of, you've sketched out this almost glib argument on Twitter, on X, where you made this

28:03you made a rough argument for why recursive self-improvement can't happen without additional information. And so maybe just walk through very quickly how like the same model, you can just very quickly show that a model can never recursively self-improve. So, you know, another phrase that we've been using recently is, you know, the output of the LLM is the inductive closure of what it has been trained on. Yeah. So when you say that it can recursively self-improve,

28:44it could mean one of two things. So let's get back to the... Well, actually, you know what's kind of interesting is like often the... Most people agree that if you have one LLM and you just feed the output and the input, like it's not going to do anything. But then often people will say, well, what if you have two LLMs? you have no external information, but you have two LLMs talking to each other. Maybe they can improve each other, and then you can have, like, you know, a takeoff scenario. But again, you even address this, even in the case of, like, N number of LLMs, using kind of the matrix model to show that, like, you just aren't getting any information entry.

29:19Yeah, so you can represent the sort of information contained in these models. And let's go back to that matrix analogy that I have, the matrix abstraction. So like I said, you know, these models represent a subset of the rows, right? Yeah. So a subset of the rows are represented, but some of these rows are able to help fill out some of the missing rows. For instance, you know, if the model knows how to do multiplication doing the step-by-step, then every row that is corresponding to, let's say 769 times 125 or whatever. It can fill out the answer. It can fill out the answer because it has those algorithms sort of embedded in them.

Read the full transcript

30:07You just need to unroll them. So it can sort of self-improve up to a point. But beyond a point, these models can only sort of generate what they have been trained on. So let me give you, I'll give you three examples. Yeah. So any model, any LLM that was trained on pre-1915 physics would never have come up with a theory of relativity. Einstein had to sort of reject the Newtonian physics and come up with this space-time continuum. He completely rewrote the rules, right? So that is an example of, you know, AGI, where you are generating or generating new knowledge. It's not simply unrolling what's already...

31:00It's not computing something. It's actually discovering something fundamental about the universe. Fundamental. And for that, you have to go outside your training set. Similarly, any LLM that was trained or didn't would not have come up with quantum mechanics. That's where particle duality or this whole probabilistic notion or that energy is not continuous but it is quantized. you had to reject Newtonian physics. Or Gettel's incompleteness theorem. He had to go outside the axioms to say that, okay, it is incomplete. So those are examples where you're creating new science or fundamentally new results.

31:39That kind of self-improvement is not possible with these architectures. They can refine these, they can fill out these rows where the answer already exists. Another example, you know, which has received a lot of press these days is these IMO results, International Mathalipiad. You know, whether it's a human solving it or the LLM solving it, they are not inventing new kinds of math. They are able to connect known results in a sequence of steps to come up with the answer. so even the LLMs what they are doing is they are exploring all sorts of solutions in some of these solutions they start going on this path where their next token entropy is low so that's where I say they are in that Bayesian manifold where you have this entropy collapse and by doing those steps you arrive at the answer but you're not inventing new math You're not inventing new axioms or new branches of mathematics.

32:46You're sort of using what you've been trained on to arrive at that answer. So those things LLMs can do, they'll get better at it, of connecting the known dots. But creating new dots, I think we need an architectural advance. Yeah. So Martin was talking earlier about how the discourse was either stochastic parrots or AGI recursive solving. How do you conceive of the AGI discourse or even the concept? What does it mean to the extent that it's useful? How do you think about that? The way I think about it, the way we have tried to formulate it in our papers, it's beyond a stochastic parrot, but it's not AGI.

33:33It's doing Bayesian reasoning over what it has been trained on. It's a lot more sophisticated than just a stochastic parrot. How do you define AGI? Okay, so AGI. So how do I define AGI? So the way I would say that LLMs currently navigate through this known Bayesian manifold, AGI will create new manifolds. So right now, these models navigate. They do not create. AGI will be when we're able to create new science, new results, new math. When an AGI comes up with a theory of relativity, I mean, it's an extremely high bar, but you get what I'm saying. It has to go beyond what it has been trained on to come up with new paradigms, new science.

34:29That's my definition of AGI. Vishal, do you think that based on the work you've done, Can you bound the amount of data, computer, or data or compute that would be needed in order for it to evolve? So one of the problems, if you just take LLMs as they exist, is there was so much data used to create them. To create a new manifold will need a lot more data just because of the basic mechanisms, right? Otherwise, it'll just kind of get kind of consumed into the existing set of data. Have you found any bounds of what would be needed to actually evolve the manifold in a useful way, or do you think we just need a new architecture?

35:14I personally think that we need a new architecture. The more data that we have, the more compute we have, we'll get maybe smoother manifolds. So it's like a map. Yeah, because there's this view that people have. They're like, well, Vishal, this is all good and well, but, you know, I could just take an LLM and I can give it eyes and I can give it ears and I can put it in the world and it'll gain information. And based on that intervention, it'll improve itself. And therefore it can learn new things. But the counterpoint that I've always just intuitively thought to that is the amount of data used to train these things is so large.

35:52How much can you actually evolve that manifold given an incremental? I mean, almost none at all, right? There has to be some other way to generate new manifolds that aren't evolving the existing one. I completely agree. There has to be a new sort of architectural leap that is needed to go from the current, you know, just throwing more data and more compute, you know, it's going to plateau. It's, you know, the iPhone 15, 16, 17. And are there any research directions that are promising in your mind that might help us, you know, go beyond LLM limitations? So, I mean, again, I love LLMs. They are fantastic.

36:33They are going to increase productivity like nobody's business. But I don't think they are the answer. So, you know, Yad Likhan famously says that LLMs are a distraction on the road to AGM. They're a dead end. They're a dead end to AGM. I don't think, I'm not quite in that camp, but I think we need a new architecture to sit on top of LLMs. to reach AGI. You know, a very basic thing, you know what Martin just said, you give them eyes and you give them ears, you make them multimodal, of course they'll become more powerful. But you need a little bit more than that. You know, the way human brains learn with very few examples, that's not the way transformers learn.

37:17And, you know, I'm not saying that we need to create an Einstein or a Gator, but there has to be an architectural leap that is able to create these manifolds. And just throwing new data will not do it. It'll just smoothen out the already existing manifolds. Is that something, so is your goal to actually help like think through new architectures or are you primarily focused on putting formal bounds on existing architectures? A bit of both. I mean, the former goal is the more ambitious one that everybody is chasing. And yeah, I think about that constantly. Are there any new hints that a new architect, or have we started to make any progress on new architectures?

38:03Or is it?

38:09You know, Yarn has been pushing at this JPEG architecture, energy-based architectures. They seem promising. The way I have been sort of thinking about it is, you know, there's this set of benchmarks or the ARC prize. Yeah. That Mike Canoop and François Chalet have. And if you understand why the LLMs are failing on this test, maybe you can sort of reverse engineer a new architecture that'll help you succeed in that, right? And I agree with a lot of what several people say that, you know, language is great, but language is not the answer. You know, when I'm looking at catching a ball that is coming to me, I'm mentally doing that simulation in my head.

39:09I'm not translating it to language to figure out where it'll land. I do that simulation in my head. So, one of the new architectural things is, how do we get these models to do approximate simulations? To test out that idea and whether to proceed or not. So, yeah. Another thing that I've always wondered about is, did we develop as humans, did we develop language because we were intelligent? or because we developed language, we accelerated our intelligence. So I don't know which side of the camp you follow on that question. Well, I mean, what's interesting is you have these anecdotal examples of humans developing languages de novo that have been recorded, right?

40:02It's either the Guatemalan or Nicaraguan sign language, right? Where there's these students that develop their own language without being taught. and so that would suggest that language just follows intelligence. The problem is they're all anecdotal, right? Like who knows if somebody didn't teach them sign language? Like nobody really knows. There is no controls. So this is all these observational studies, and there's so few of them, you have to wonder if it's just kind of sloppy observation. And so I think that the question is still outstanding. Yeah.

40:37So, I mean, language definitely accelerated our intelligence. There's no question about that. But which followed which, we don't know. I view it as a networking problem naturally, which is once you have languages, you can communicate. You can communicate, you can store, you can replicate, yeah. Yeah, exactly. Cool. Again, this is kind of a wonky question, but I think one thing that you've brought to the discourse, and for those that are listening to this, I really think that you should look up Vishal's work and read it. I just think it'll give you a really, really, especially if you have a systems background, like a networking systems bracket, give you a really, really good understanding of kind of the bounds on these.

41:16But like the toolkit that you draw from is like information theory and like more formal. Have you found that the AI community is receptive to this or is it like two different cultures, two different planets trying to communicate and not a lot of common ground? Like how have you found like bringing like the networking view of the world to the AI realm? Some of them are receptive to it, definitely.

41:46But, you know, these large conferences and their reviewing process, it's so random. And the kind of questions they ask, you know, I'm a modeling person. I like to model things. And, you know, I submitted one version of this work to one very famous machine learning or AI conference. and the reviewer said, okay, this is a model, so what?

42:15So there is... That's absolutely remarkable. So you've actually taken a system that nobody understands, we have no models for, you actually provided some model that we can use to analyze it, and that alone wasn't sufficient. They're asking, so where are the large-scale experiments to prove this? I do. Listen, I honestly, I mean, I find there's so much empiricism in like the current, you know, AI community, exactly because we don't understand the systems. You know, it kind of reminds me, I feel like, I feel like systems went the other way, right? It's like we had all of these models, but then we didn't understand how the systems worked.

42:54And then we just like actually did measurement. It feels like ML and the AI stuff is the opposite, which is like, we know we don't understand them. And so we just measure them. But now we're trying to like come up with the models. Yeah, exactly. So it was so easy in some sense to build these artifacts and then just measure them that people have been going around trying to do that. And one term I really dislike is prompt engineering. Why? Engineering used to mean sending a man to the moon or providing five minutes reliability. Prompt engineering is prompt twiddling. you fiddle with a prompt and the barge changes and the inference, the output changes and you have hundreds of papers just doing one experiment on the other changing a prompt this way, that way and writing their observations and as a result lots of these papers are being written are being submitted for review reviewers get busy looking at all this kind of empirical work and my personal taste is first try to understand, model it.

44:06And then you can do the other thing. So like a true theory guy. I don't know about this bit twiddling. Let me ask one more LLM question, which is, are there any benchmarks or real-world tasks that if they occurred, you'd sort of reevaluate and say, hey, maybe LLMs are closer to the path to AGI than I thought?

44:31if the very real world task?

44:40Good question.

44:47You know, which for LLMs or these models, the one domain where you have the most training data is probably coding. And coding is where you can also have the most structure. And yet, anyone who has used these tools, whether it's cursor or whatever, or cloud code, LLMs continue to hallucinate, continue to generate unreasonable code. You know, you have to constantly babysit these models. So the day an LLM can create a large software project without any babysitting is the day I'll be a little bit convinced that it's easier. But again, I don't think it'll be able to create new science. If it does, that's when I'll be convinced.

45:54I think that you can almost take a definitional approach to answer this question, Vishal. Like the problem with these types of questions is if you have billions of dollars and you can collect whatever data you want, you can make a model do anything you want, right? And so like, you know what I'm saying? Like at some level, you've got this entire capital structure, machinery behind these models. So you're like, oh, it can be good at science. Well, sure, you put a billion dollars of solving materials science and collect all this data, you'll be good at material science or whatever it is. And so, but there is a definitional answer, which is, and I'm going to draw from your work, which is there is a manifold that's in there based on the data it's been trading on.

46:35And then the question is, if it ever produces something that's off, like a new manifold, so considering the existing traded data, if it ever does that, if it does something that's outside of that distribution, then clearly we're on a path to learning new things. And if not, then everything is just a computational step from what's already known. Yeah. And I guess the counter to that would be maybe all humans do is work on their own manifold, and Einstein was lucky or something, I guess would be the counter to that. Well, so there's several mini Einstein examples, and yeah, it's creating this new manifold.

47:13I didn't want to use that definitional answer. I thought it might sound too wonky, too mathematical. but essentially if LLMs really created this new manifold then I would be convinced. But so far they have just gotten better at navigating the existing manifold, the existing training set. Which is hugely powerful and is going to change the world. I'm not denying that. I think they are extremely, extremely good at what they can do. But there's a limit to what they can do. So I have one quick question. What's next for you? I mean you've tackled in-context learning, you've got a model for LLMs and I've got a generalized model for their solution space.

47:53What are you thinking about tackling next? In terms of modeling or? Academically, an LLM. Academically, I'm thinking of this, what is the architectural leap that is needed to create this new manifold? and how do we use multimodal data? Awesome. To expand around. Once you figure that out, come back and talk to us. That's right. We'd love that. So, I mean, even with LLMs, in the paper we say that you can improve the inference by following this low or minimum entropy path. So that's a very sort of small step that we are taking. we are building and training models that will do inference based on the entropy path.

48:52By the way, is Model Probe still up? Token Probe. Yeah, yeah. Token Probe is still up. And you can see actually Token Probe is a software that we built and thanks to Martin and A16Z's generosity it's running on your servers and anyone can go and test. And what we have done there is we actually show the entropy. Yeah. It is so enlightening. I recommend anybody listening to this who's interested. Actually, check out TokenProbe. It literally shows you the confidence as you go along. It's remarkable. So in context learning, you create your new DSL and you give it to the prompt and you can see the confidence rising with each new example, the entropy reducing.

49:35And that sort of is a validation of the model. You can see it sort of unfurling right in front of your eyes. The token propane is still adding. Thanks again. Vishal, thanks so much for coming on the podcast. It was a great conversation. It was great fun. Thank you so much again.

50:12at a16z.substack.com. Thanks again for listening, and I'll see you in the next episode. As a reminder, the content here is for informational purposes only, should not be taken as legal business, tax, or investment advice, or be used to evaluate any investment or security, and is not directed at any investors or potential investors in any A16Z fund. Please note that A16Z and its affiliates may also maintain investments in the companies discussed in this podcast. For more details, including a link to our investments, please see a16z.com forward slash disclosures.

From the publisher

From GPT-1 to GPT-5, LLMs have made tremendous progress in modeling human language. But can they go beyond that to make new discoveries and move the needle on scientific progress?

We sat down with distinguished Columbia CS professor Vishal Misra to discuss this, plus why chain-of-thought reasoning works so well, what real AGI would look like, and what actually causes hallucinations.

 

Resources:

Follow Dr. Misra on X: https://x.com/vishalmisra

Follow Martin on X: https://x.com/martin_casado

 

Stay Updated: 

If you enjoyed this episode, be sure to like, subscribe, and share with your friends!

Find a16z on X: https://x.com/a16z

Find a16z on LinkedIn: https://www.linkedin.com/company/a16z

Listen to the a16z Podcast on Spotify: https://open.spotify.com/show/5bC65RDvs3oxnLyqqvkUYX

Listen to the a16z Podcast on Apple Podcasts: https://podcasts.apple.com/us/podcast/a16z-podcast/id842818711

Follow our host: https://x.com/eriktorenberg

Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.

Stay Updated:

Find a16z on X

Find a16z on LinkedIn

Listen to the a16z Podcast on Spotify

Listen to the a16z Podcast on Apple Podcasts

Follow our host: https://twitter.com/eriktorenberg

 

Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.


Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

More from The a16z Show

All 489 episodes
Columbia CS Professor: Why LLMs Can’t Discover New ScienceThe a16z Show · 51 min
Listen in VO