Grokking, Generalization Collapse, and the Dynamics of Training Deep Neural Networks with Charles Martin - #734

5 Jun 2025 · 1 h 25 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

TWIML AI Podcast Episode #734 Summary: Grokking, Generalization Collapse, and the Dynamics of Training Deep Neural Networks with Charles Martin

Episode Description In this episode, host Sam Charrington is joined by Charles Martin, founder of Calculation Consulting. They discuss Weight Watcher, an open-source tool for analyzing and improving Deep Neural Networks (DNNs) using principles from theoretical physics. The conversation explores the concepts of Heavy-Tailed Self-Regularization (HTSR) theory, the learning phases of DNNs, and practical applications in generative AI.

Key Topics Discussed

  1. Charles Martin's Background
  2. AI researcher with a PhD from the University of Chicago.
  3. Experience in industry, consulting, and AI solutions for companies.
  4. Co-developed the Weight Watcher project, an open-source tool with nearly 200,000 downloads.
  1. Weight Watcher Tool
  2. Designed to monitor, train, and fine-tune AI models.
  3. Employs theoretical physics principles to analyze DNNs.
  4. Identifies three learning phases:
  5. Underfitting
  6. Grokking
  7. Generalization Collapse
  8. Uses a "layer quality" metric to evaluate layer performance.
  1. Learning Dynamics in DNNs
  2. The phases of training are akin to baking a layered cake; if the temperature (learning rate) is too high, some layers can overfit while others underfit.
  3. Importance of finding the optimal learning rate to avoid overfitting and ensure effective learning.
  1. Fine-Tuning Challenges
  2. Fine-tuning is often difficult, with many practitioners struggling to get optimal results.
  3. Common issues include changes in data pipelines and lack of good data.
  4. The importance of monitoring and adapting models continuously in production environments.
  1. Grokking and Generalization Collapse
  2. Grokking: A phenomenon where a model memorizes training data before achieving high generalization accuracy—representing a nonlinear learning transition.
  3. Generalization Collapse: Occurs when a model stops generalizing correctly after initially performing well.
  1. HTSR Theory
  2. The foundation of Weight Watcher, combining insights from random matrix theory and renormalization group ideas.
  3. Establishes that optimal model performance correlates with layer quality metrics.
  4. Provides a method to predict when overfitting will occur in a model.
  1. Real-World Applications of Generative AI
  2. Insights shared regarding the practical deployment of generative AI in various sectors.
  3. Importance of continuous retraining and monitoring of models to adapt to changing data environments.
  1. Search and RAG Technologies
  2. Discussion on the relevance of retrieval-augmented generation (RAG) technologies and the challenges faced in implementing them effectively.
  3. Emphasis on the necessity of understanding user clickstream data to improve search relevance.

Key Takeaways

  • The dynamics of training DNNs are complex and require careful tuning to ensure models generalize well without overfitting.
  • The Weight Watcher tool provides valuable insights into the learning behavior of layers within DNNs, enabling practitioners to optimize performance.
  • Fine-tuning and operationalizing AI models in production environments remain significant challenges, often requiring robust monitoring and adaptation strategies.
  • Understanding and leveraging the principles of theoretical physics can provide a novel approach to tackling issues in AI model training.

Conclusion The episode provides a deep dive into advanced concepts in deep learning and machine learning, combining theoretical physics with practical applications in AI. Charles Martin’s insights into model behavior, training dynamics, and the importance of fine-tuning highlight the ongoing challenges and opportunities within the field of AI and machine learning.

For more detailed notes and additional resources, visit [TWIML AI Podcast Episode #734](https://twimlai.com/go/734).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00So think of it like baking a cake. When you bake a cake, you have the temperature, you watch the cake. Imagine a cake has a lot of layers. Well, if the oven's too hot, some layers are going to burn, and the inside's not going to get cooked because you don't get good conduction through the heat. So what you want, even when you bake a cake, you still have to be careful to adjust the temperature right to make sure that the layers all cook it the same way. And it's the same idea in a model.

0:37All right, everyone, welcome to another episode of the TwiML AI podcast. I am your host, Sam Charrington. Today, I'm joined by Charles Martin. Charles is the founder of Calculation Consulting. Before we get going, be sure to take a moment to hit that subscribe button wherever you're listening to today's show. Charles, I am super excited to have you on the podcast. We've been working on this one for quite a while now. Great. Thanks a lot, Sam. I'm really happy to be here. This is a great show. Appreciate it. Appreciate it. You know, to get things started, I'd love to have you share a little bit about your background and what you've been up to recently with our audience.

1:13Sure thing. So, you know, I'm an AI researcher. I did my PhD at University of Chicago. You may know my more famous classmate, John Jumper, who just won the Nobel Prize for Alpha Fold. I then was an NSF postdoc. I my other you may know another classmate of mine, Juergen Schmidhuber, who basically claims to have invented everything else. So I did. So I've been doing this for a very long time. I really got I mean, I've been doing working in industry and I left grad school and postdoc went in the industry. I've been doing consulting work, doing building A.I. machine learning solutions for companies forever.

1:50I did Aardvark was acquired by Google. I did e-how first billion dollar IPO since Google and probably the biggest crash. Um, I was a quant on wall street. I was also been scientific advisor to Larry Page's family. So, and, and during the course of that, I just, about 10 years ago, I decided to get back into, uh, AI research, working with my friend, Michael Mahoney at UC Berkeley, very informally. And we've been working on this project called, uh, uh, the, what I call the weight watcher project, which is a large, which is an open source tool. We have almost 200 ,000 downloads, which is designed to help people who are monitoring, training, and fine tuning their own AI models.

2:26And all of this has been just really a passion project to get back into research using techniques from theoretical physics and chemistry to understand how these models work and how to help my clients. And you've been at this for a while. I think we first got to one another in the deep learning era. Now we're all talking about these large language models. How does this stuff apply to LLMs? Well, it actually works really well. I designed, I did theory uh but as a theorist that you know i work with engineers so i know that theory you know they say you know in theory practice in theory the same and in practice they're different right so i started i actually got into this i had a client in all places of slovenia and we were they were making like fake texts for things like weight loss articles amazon reviews things like that so we were generous this is before even like open ai stuff came out and we're generating this text.

3:23I realized I'm generating like 50 ,000 pieces of text a day. I have no way to know if what I'm generating is any good or not. And I can't hire 500 people to evaluate it. So it would blow my costs. I've got to get a theory. I've got to invent some kind of theory that would let me know whether these AI models are working correctly or not. And I started working on my spare time and working on kind of cracking the books and acting like a scientist again. And that's where it's come up. So it turns out it applies really well to LLMs because they're really big and they have lots and lots of layers. And this is basically a theory which analyzes the weights in the layers.

4:01It's Weight Watcher. So it looks at the layer weight matrices. And so the more layers you have, the bigger the model, the better the theory works. It's kind of interesting to me that even in that, you know, practical context, so working with a client that has a job to be done as opposed to in academia, approach you took was a theoretical approach as opposed to a more, you know, let's call it a statistical course, like the application of evals as a practice, as opposed to, you know, the theoretical take on it. I have, I did all this stuff in the 90s and I have this background in theoretical physics and chemistry.

4:33So, and of course I was a quantum Wall Street. So some of the techniques I'm using actually are quant techniques that we used on Wall Street to predict the markets. So I knew how to apply these techniques to real systems because we use, that's what we do, we do portfolio theory. What's an example of a technique that, and how is it used on Wall Street and how are you using it in your project? So in portfolio, when you're like at a big place, I was at BlackRock, you know, that's the 800 pound gorilla on Wall Street. And we, we have, you know, the group I was in, you couldn't even trade in the group unless you had$200 million, right?

5:06That was the scale. So it was a big portfolio and you have these big portfolios and you have to figure out signal versus noise. Where's the signal in the portfolio versus noise? You're trying to reallocate how much Google, how much Apple, how much GM, how much, you know, that kind of stuff in the portfolio. And so you have to decide signal versus noise. And so you use something called random matrix theory to do that. And you can detect the signal and the noise. And it's important that you not peek at the stock market data. You not peek at things because you'll, you'll overfit yourself to the history.

5:38So there's two things I'm I'm using random matrix theory, which is what we do. And I'm making sure not to look at the data, because if you look at the data, you'll fool yourself, right? You'll think it up. So that idea is actually essentially you can think of Weight Watcher as trying to find the signal from noise in a model by looking at the individual weight matrices, which are like little portfolios. They're like the portfolios you would form when you're doing portfolio theory. And I knew about some of the theory because we do theory when you're a quant. And so it turned out there were some very interesting and interesting scientific properties.

6:11You know, as a scientist, I started studying this stuff. And I go, you know, what I did was I just took the models that people had trained. They did these open source models, right? Today we have Hugging Face. There are a million models on Hugging Face. When I started, there were less than 100 open source models. I started looking at them, right? What do they look like? What do the good ones look like versus the bad ones? Just like when you're a quant, what do good companies look like versus bad companies? And you analyze their properties. And it turned out there was some interesting science behind this.

6:42And one of the guys I worked with, the guy who was the head of futures and equities at BlackRock, was a theoretical physicist. And, you know, he showed me some of the stuff they were doing. You know, we can apply this to AI. And it just turned out to work, you know. And I just kept digging and digging and digging. It just worked better and better and better. And then I discovered there's some deep science in this. So that's where we ended up. And so traditionally in the ML AI world, the way to overcome that, you know, overfitting on the past data is to split it in a test and train and to only look at part of it.

7:16Are you you're not looking at any of it? No, no. In fact, you don't do that in physics. Yeah. Like when you do analysis, you I have a I have like 100 I just 120 long 20 page paper theoretical physics explains how we do things. but it's actually different, right? We actually, what you're doing is you're looking at, it turns out the weight matrices, when you train a model really, really well, the weight matrices have these universal properties. I call them signatures of emergence. And they actually, this idea actually comes from, actually, because I did AI in the 90s, I noticed something about neuroscience.

7:53It turns out these signatures are very similar to the signatures you see when studying spiking neurons. and it turns out spiking neurons exhibit something called parallel structure and they have these sort of universal properties and we knew about them and sort of the original tool is i took these tools from computational neuroscience and i just applied them to the layers of weight matrix and turns out they have the same signals so it turns out that the spatial temporal correlations in the weight matrices in the layers and the neurons are very similar to what you see in actual neurons like that you would culture you take a lab you culture them you grow them in the lab and you look in there it's the same thing and and this work was pioneered actually by um some physicists in the 90s in particular friend of you know some guys that got him jack cowan university of chicago has done a lot of work on this he's sort of invented a lot of this stuff so it turns out this stuff works um and it's actually very simple if you train a model you're training a layer right most people think of training the model as i minimize the training accuracy but you minimize the error right but you minimize the error right accuracy the error either yeah excuse me you You minimize the air, not the accuracy, right?

8:58Maximize the accuracy, right, right, right, right. So think of it like baking a cake. When you bake a cake, you have the temperature, you watch the cake. Imagine a cake has a lot of layers. Well, if the oven's too hot, some layers are going to burn and the inside's not going to get cooked because you don't get good conduction through the heat. So what you want, even when you bake a cake, you still have to be careful to adjust the temperature right to make sure that the layers all cook at the same way. And it's the same idea in a model. is that if you're training, like for example, if you turn the learning rate up too high, it turns out some of the layers overfit.

9:35So there's too much information in them and they just overfit and all of a sudden they stop generalizing. And we have a very interesting paper that I've done with a fellow who came out of just getting his master's degree now, super smart guy. And he was looking at this and he said, well, why don't I take this old rocking problem? And it turns out if you take a simple model and you train it for a very, very long period of time, even if you don't turn the learning rate crazy up high, you'll see that it starts to overfit. And we can detect this. And the signature of overfitting, you know, really the reason I started thinking is because if I were trading in the stock market and I were making an AI model, what's the one thing you don't want to do?

10:15You don't want to overfit to the history. Overfit on the history, yeah. Because, you know, and I was at BlackRock, we saw guys doing this. I'm like, what are you doing? You know, like you can't do this stuff. Well, that's the classic thing. You learn a little bit of machine learning. You say, let me download all of the historical data and train a model on it. Yeah, you go like, here's a guy, PhD from electrical engineering from Stanford. I'm like, what are you guys doing? You know, but this is a classic thing. And so it's very critical to understand when the overfitting occurs. And so, you know, having been a quant on Wall Street, I'm really, really sensitive to this kind of stuff.

10:47and it turns out there's some ideas from physics that are like phase transitions and where you in the stuff that was we knew this stuff happened in the 90s if we understood the theory um the thing is we're physicists the physicists and the computer scientists didn't really talk that much so you know that so we have these theories from physics that we know we can detect this stuff but how do you apply it and that's what we try to figure out how to do and it turns out that yeah you could detect when the model overfits and when it's underfit and you can see it in the layers It's just like you could see this layer is burnt.

11:18It absorbed too much information and absorbed so much information from the training data. It got stuck and it doesn't learn anything. It's just it's overfit. Right. And then there are cases where the layers are underfit. Like the bottle's too big or something's going on. You know, the layers just didn't really learn anything. So we can detect that very easily. So we'll dig into this paper before we do that. But. a contextual context setting question and then a possible rabbit hole. So the context setting question is one application of all this is fine tuning models. And when we were chatting before, you were talking about how fine tuning is really difficult and a lot of folks get it wrong and find it really frustrating.

12:07Other folks I hear from say, oh, fine tuning is so easy right now, compared to, you know, a few years ago, like a, like kind of square that circle for me and talk a little bit about your specific experience with regards to fine tuning. So if you look at what people are doing, there's a report by McKenzie that said maybe 9 % of companies doing AI are fine tuning. So the rest are probably doing prompt engineering. And the question becomes, look, some people, most people do, I'd say out of those, probably of those only 9 % are really doing full fine tuning on very large data sets. Yeah, you could do something along with Laura, which is called low-rank adaptation.

12:47And you can kind of tweak the model. If you want to get a JSON output well, or you're trying to get Markdown output, you just tweak the model a little bit to change its outputs. You can kind of steer it a little bit, right? But if you want to add a huge amount of data to your model, and you want the model to learn from that data while still keeping its own knowledge, it turns out that it's actually quite hard. It's hard for a number of reasons. One, because I've worked in enterprises, it's hard to get good data in the enterprise. And that's part of it. It's just, you know, and I just talked to a client the other a couple weeks ago.

13:24He said, yeah, we had our model running for a year in production. We had to dump it and start over because we didn't realize it was broken because something happened in the data pipeline. The data pipeline stopped. You know, somebody changed a column table or they changed something and it screwed up the model. and you know doing i've worked with projects like work with go daddy and ebay walmart and these things happen all the time you know you're in a production environment the data changes you don't know what changed where that told you you know three months later the model's the model's not working what happened so you know this is the problem with fine tuning or you know because you know you're even training your own model from scratch that the data goes crazy and you don't know it um and that's probably and then the thing is you have to how do you prepare good data sets.

14:09Data sets are their duplicates and there's noise and there's problems. And so in a, in a real environment, a production environment, a big, in a legacy company, this is very hard. So yeah, if you're just training on a very small data set and you can look at the data yourself and curate it, you could probably get the data right. But when you're in a, you know, production environment with, I've worked like, you know, Walmart, millions of examples on the clickstream, you know, from the, from the search engine, you can't curate it yourself. You need, you need tools, which are to look at the model and say, did something go wrong?

14:39It's very hard because you don't, you just don't know. You don't have visibility. The other thing that makes fine tuning hard is just that, you know, you never really know whether you're evaluating the model correctly. There are a hundred different evals and, you know, there's all this controversy. You know, I remember Llama 4, oh, they put Llama 4 out. books they cook the books right they snuck it in it turned out now the and then it turns out that the ellipses guys you know they they've been secretly you know handing over the answers like 25 of the answers to google and that's the cohere paper that came out a couple weeks ago yes yeah the cohere i gave a talk at cohere for ai a few weeks ago yeah so so they've been cooking the books you don't really know what the model can do right so don't say that though because they just raise a ton of money you know uh you know it's what you know they got they're trying to figure out right they're trying to figure out what's going on so look i work with real clients and my clients when my models don't work they don't pay me like i know so i you know i you know it reminds me of the difference you ever see the old movie ghostbusters uh-huh i remember they they're getting ready to leave the university and and uh i think it was um i can't remember which said he goes look man you you've spent your whole life in academics you've never been in the real world they expect results so i'm coming from that perspective like so you know it's yes of course the tool the problem that's happened is that in the industry is that the tooling has gotten easier and with the tooling getting easier there's just there's a lower bar to entry everybody's trying to do things and so there's a much wider variance in what's going on and so we see that i don't see a lot of people, if they are fine tuning, you know, there's a lot of problems, things break.

16:30And the goal of this product was to figure out how to make fine tuning work really well and to detect problems in production. So if you're retraining a model every day, every, you know, I worked in search engine, we train every hour, but you might retrain your model once a month, maybe once a week, once a month, you want to know, you want to make sure things didn't drift and go crazy. And this technology is designed to help you find those kinds of problems that you define problems you can't find using any other technique. But one of the things you said that was pretty interesting was kind of the suggestion that there are different types or degrees of fine-tuning.

17:08Like there's a surface-level fine-tuning and a deeper fine-tuning. I'm going to force-fit your cake analogy. Like you've got a vanilla cake. Do you want just like chocolate icing or do you want chocolate at the bottom layer? Or you want to stick a layer in between. That's exactly right. Yeah. If you're just putting sprinkles on the cake, it's not a big deal. Okay. But if you're trying to stick something in the middle, you know. And I've not really seen like any like concrete, you know, taxonomy or elucidation of like degrees of fine tuning or what characterizes an easy fine tune, what characterizes a hard fine tune.

17:43You have models that are instruction fine tuned. Okay. So instruction fine-tuning, for us, our technology is really, you're doing instruction fine-tuning on 100 ,000 examples or a million examples. We were working, for example, with a group in Poland, and they're trying to do instruction fine-tuning on a model to try to adapt it to the Polish language. And we found there were some funny things going on. They were trying to do, there's a model called solar, which is quite good, and they were trying to adapt the solar technology to their model called Biolick. And the model's not bad, right? But something happened inside the fine-tuning and the model training that somehow whatever they did, the weights and the weight matrices and the layer doesn't, and by the way, this is all published in the city publication.

18:24It doesn't look like solar. Like something went wrong and we couldn't figure out what it was. Like what did you guys do? You know, because you're trying to replicate the engine. It's like you want to keep what you like about the base model, but add some behavior or add some knowledge or add something at the same time. And it sounds like they broke something fundamental. They broke something. And the thing is, you're trying to follow an instruction that somebody else has given you, but you don't really know what's like. You're trying to do something a little different and it's difficult to make the recipe exact.

18:57Right. You change something a little bit. And now, you know, you use the different kind of flour, different kind of oil, a different kind of butter. Something went wrong and you don't know why. And this stuff is so ephemeral. You know, it's so opaque. Like, you know, whatever you do on one data set might not work on another data set for whatever reason. We don't know why. And that's what we see happening is that they tried to replicate the training process exactly. But when you look at the models, well, they're quite different. What did, why did, what happened? And they're quite different in terms of behavior and performance or the, okay, from the perspective of your.

19:34Yes. So one thing that this tool does is it makes clear some set of differences between these two, you know, what they started with and what they ended up with. And the problem is that you don't know what the problems are going to be. It's you go into production, people start doing things. You know, there are a million things you could do. You're trying to, like I said, you're generating fake text, you're answering questions, you're doing things. You don't know how it's going to respond. were they it's already a client of yours and like it was natural to use weight watcher like where they use weight weight using weight watchers from the beginning and so it was clear what no no they came in later what happens they're doing it they train them what did they see then yeah what did they see that said we need help with this the layer weight watcher gives you a layer quality metric so if i have a model with a thousand layers and you have some of these models now have a thousand layers, right?

20:27But you have a thousand layers, maybe a hundred layers in your model. Every layer gets a quality metric. What's the score on the layer? And that score should be somewhere between two and four, two and six. And so in that range too, let's say two to five. Now what you find when you fine tune is you can look at the fine tuning update. So I trained the model. Here's the update. I want to look at the update. If the update should show good quality scores between two and four, two and five. If you look at the instruction fine-tuning updates of Llama, Quinn, Falcon, Mistral, they all show reasonably good layer qualities.

21:04You look at these, even solar, their instruction fine-tuning, for some reason, had a large number of layers that seemed to be underfit. The layer quality was too high. I should say the quality was low in the sense that the metric was above six, maybe nine, maybe 10, maybe 15, way too high, which indicates the layer is almost random. So what happened? Those layers, for some reason, did not learn any information. Or if they learned it, they learned it very weakly. There's only a very small amount of information those layers learned. And that's, you know, you're wasting a lot of compute. You know, you're spending all that.

21:41I mean, they're running on a supercomputer in Poland. So, you know, they're spending all this compute time, energy to figure out what's going on. And that's what the tool is telling you. You know, there's a step before that. What was the thing that they were observing in the training process that caused them to call you? Was it just performance or? Yeah. Yeah. They're just trying to get the performance up. They're trying different things. And again, they don't really know what to do. You know, you read a paper, there's instructions, there's code. but when you run on your data set, it does something different.

22:14So we read this paper, based on everything we read, we should be able to take our data, apply a fine tuning approach to it and see performance like this. But we're seeing performance down here. Can you help us figure out what's going on? Right, and so one of the things people do now is on these big models, they're doing all sorts of, it used to be you take a model, you train it, right? Now people are trying to like, they tried to replicate part of the model. Here's part of the model and we're going to take these layers and replicate them and stick them up here, right? And then we're going to take two models.

22:45We're going to merge them together, right? So they do all these funny things, right? And people say, oh, you can replicate layers. You can merge models together. And it's kind of like, that's a little strange, you know? And it turns out it gives goofy results. And that's the problem, right? And you don't know, did you do it correctly or not? They just don't know. You just don't know because you don't really know, like the recipe doesn't have enough detail and it depends on the ingredients which are, you know, which change. They're not stable. So that's what makes this stuff so hard. And similar to fine-tuning, like if you don't have the exact data, and maybe the hyperparameters were wrong, like maybe the hyperparameters for this data set are different from the hyperparameters for our data set, how do we select the hyperparameters?

23:32How do we select the learning rates for each layer? Should we have dropouts? Should we not have dropouts? Should we have weight decay? Should we not have weight decay? These are the open questions. And it changes from data set to data set. How would you know? And this is the kind of thing that we found that it was very difficult to try to, especially because you're running on a supercomputer. You could run it once, right, on the computer. It runs for a few weeks. You come back, you know, you're not Google. You can't run it a hundred times on their million node cluster. I mentioned that I had a possible digression rabbit hole.

24:06and that is when you were describing the spiking nature of the neural networks. It made me think of like Nementa, the Jeff Hawkins stuff. Like, have you come across that? I've met them at conferences. I like what they're doing. You know, it's, look, the idea of modeling spiking neurons was developed by a guy named Jack. There were sort of two or three people at the time in the late 1960s. And one of them was Jack Cowan. If you follow Schmidhuber on Twitter, you know, he was complaining that Hopfield got the Nobel Prize and he didn't deserve the Nobel Prize. What are you going on Twitter? He's also, you know, what are you going on Twitter saying stuff like this?

24:45You know, I mean, come on, man. You know, I mean, that's something you stay for a bar when you're drunk, not when you're on Twitter. But during the late 60s, people started modeling neurons. You'd go in the lab and you would model the spike electrical activity and try to come up with models for the electrical response or neurochemical response of a neuron. I was a theoretical chemist. That's what we do is what theoretical chemistry does. So it turns out that, you know, people have like models of spiking neurons. And there are people, I mean, but I mean actual neurons like in a lab. You would take the neurons, you grow them in a lab, and you put electrodes in, you watch how they spike, right?

25:23You know, some Elon Musk kind of thing. He's going to try to, you know, how they know. Neuralink, yeah. Yeah, how it works. That started by growing neurons in a lab, right? That's where it comes from. That was like the early work on this. So, you know, there actually is a deep connection between that stuff and AI people. And the models we have today were developed to try to explain these spiking neurons. And they eventually became sort of computer science models and we run them on, you know, GPUs. But there's a lot of theory behind that. And we and there's a lot of experimental observations. And you can observe what I call signatures of emergence.

25:59these signatures. There's a book by a guy named, um, a late physicist per Bach, and he invented a theory called self-organized criticality. And it's this idea that systems self-organize to a critical point, a critical point between order and chaos. And you can see the signature of this critical point between order and chaos inside many physical systems, you know, in particular things like avalanches, is, you know, the snow is falling on the mountain, all of a sudden it collapses. And there's this point right before, there's this point right before the critical point when it collapses, where it's in this sort of semi-stable state of self-organized criticality.

Read the full transcript

26:42That's what's going on inside the brain. That's a theory. It's called the critical brain hypothesis. And so we can see these signatures of criticality, I call them signatures of emergence of AI, inside actual neurons. And it turns out our experimental work shows that the closer the quality, I have this layer quality metric, there's a sweet spot at two. There's a value of two. And when all the layers reach two, we think that the model is perfectly optimized. And the goal of the theory is to try to get your model so that all the layers reach two. And that's what we're trying to do, try to develop technology to do that.

27:20But, you know, here you can, and that's sort of, And we see it in, you know, just really basic experiments on things like grokking and double descent, where you can really flush it all out. But, you know, you have to, you know, we try to study small models and then apply the results to larger models. And then we go back and try to figure out what's going on. And that's what we've been doing. And it's been very successful. There's a lot of open questions. But, you know, the idea of it is, I give people, so that's what the spiking neurons are talking about, is that you can measure these signatures, and you can see them in the actual models you're training.

27:54And you can use them. You can use them, make better models. So you've mentioned grokking a couple of times. What's grokking? Grokking is a phenomenon where if you take a model and you train it for, you know, some small data set, it tends to reproduce the training, it has perfect training accuracy, zero error. And people think that it's somehow memorizing the training set. And then it has almost no test. And the test error is like as big as it could be. It's horrible, right? And then all of a sudden, it very quickly learns how to generalize. And so the training accuracy stays high, but then the test accuracy gets high.

28:36So my test accuracy might get like 85%, you know, 90%. Not super high, but it gets very high for a small data set. That's crocking. And it's something that people, it theoretically is very odd because it's odd that you would get a model that can describe the training data perfectly. You think, well, it must be overfit. It must have memorized the data. And then all of a sudden it's able to generalize. So the kind of non-linearity in the learning is what's interesting about it. Yeah, a phase transition really just goes from not being able to generalize at all to being able to generalize extremely well.

29:13And it happens just sort of suddenly. And that's called, the grok is to understand something. So they call it grokking. Yeah. And then generalization collapse. Generalization collapse, we detect it. What we define is if you continue training, all of a sudden it stops generalizing. It starts going down again. So it still has memorized the training data in some sense that the training accuracy is perfect. But it starts to learn, it generalizes, and all of a sudden it stops generalizing. And it starts going down. And it might go down to like, you know, instead of like, you know, 10 % accuracy would be random.

29:47It might go down to like 50 % accuracy. So it doesn't get it doesn't get it does a little bit of learning, but it's confused. It's like in this state of confusion, deep confusion. Got it. And these terms come from the title of the paper that we've referred to here. Grokking and Generalization Collapse Insights from HTSR Theory. And HTSR Theory. Heavy tailed self-regularization. You got to have an acronym in science. Every theory needs an acronym. Right. So HTSR. That's the theory. Heavy-tailed self-reportation, work I published with Michael Mahoney back in 2021. So, wow, it's been five years. Wow, it's hard to believe it's that long.

30:26So this is the theory behind Weight Watcher. This is why Weight Watcher works. Part of the theory. So what we discovered is that when people think about a model memorizing, they think, oh, the training accuracy is perfect. It must be memorizing. No, no. We actually, and it turns out that, you know, we discovered this other phase of memorizing, which is more like confusion. And so there's a state of memorization and there's a state of confusion and they're very different. And you can detect them using Weight Watcher. Both states appear to reproduce the training accuracy perfectly. Why would confusion reproduce the training data perfectly?

31:06Okay. So what's happening is, what's happening in these models? It turns out that in the state before grok, and we call it the pre-groking phase, training accuracy is perfect. Why can't it generalize? It turns out some of the layers have good quality, but some of the layers have very bad quality. So what's happened is you've learned the training data, but the layer that learned the training data, the layer that needs to learn how to, the layer that's most important for generalizing hasn't converged. So the idea is that it's not that you are memorizing before and then you've switched from like this thing called memorizing to this thing called generalization.

31:56It's more that part of your network is sufficiently memorized or not sufficiently memorized, but sufficiently learned, right? It's converged. And other critical parts have not yet. And that's why you don't. And that's required for the generalization. Yes. Now, there's kind of a, somebody, someone has suggested, I saw a paper recently that said the reason it doesn't learn, there's like a numerical instability in the softmax. And so it just gets numerically unstable. and it just kind of, it's trying to learn, but it just can't find its way. And you have to train it for a long period of time until it eventually jumps out of this, jumps over the hill and learns.

32:42Confusion or what we call, you know, we call overfitting in Weight Watchers. Like a confusion is that now the layers learn too much information. So they're over-converged. They're over-baked, right? They're burned. You know, they're cooked. They're overcooked. So they've learned too much information and now they're stuck. And all they know is what, and they can't, they get confused as to what's going on because they've learned so much that now they can't generalize. So it's a different thing. So in one case, we see the layers, some of the layers converge and going down and some of them are going to stay up here.

33:18And in the other case, they're all way down here. Like, they're all way down. They're all way too low. And what you want is you want them out here in the safe range. And that's what we're seeing. And that's, and the thing that that could happen in a real model in the real world, if you're fine tuning, you know, you're training that you'll see some of the layers are, say they're down here, they're good, but some of the layers are up here. They haven't learned anything yet. They're stuck. And if you, if you go too long, they all fall down. You, you burn the system, you burn it. So I, to me, I like the analogy of baking a cake because it's like, if you turn the oven up or you leave it in the oven too long, it, the whole thing will burn.

33:53Right. It just, it'll cook, it'll over cake. And so something has gone wrong. There's some numerical instability in the solver. Maybe the soft max is off. Maybe there's some goofiness in the training data. Something went wrong that's preventing the model from converging. And we can see that. And that's a different phenomenon than overfitting. And the goal is to try to, how do you fix it? Yeah, yeah. So a couple of things jump out at me as interesting here. One is this idea of like, I don't know, really just kind of thinking about this as layers, I think, and really kind of locking in on that, I think is interesting.

34:31And then that each of these layers can be independently underfit, overfit, or, you know, in the target zone. That's interesting. but then that kind of leads me to and I'm this is kind of a lead into like to what degree have you looked at these things or is this like you know next steps here but like you know when I think about what I think about like data centric AI or like the focus on data curation as a way to get models to perform better, more efficiently. You could also think of it in this context as like, can I construct an incremental training data set that targets the deficiencies of a particular layer or set of layers?

35:24Yeah, yeah, I think so. I mean, that's the kind, we haven't looked deeply at changing the data set as much as changing the learning rates on the layers or making a new kind of regularizer. But I think absolutely that there's something, and because there's numerical instabilities, data is, right. Right. My take on data centric AI is that these guys went into industry and tried to apply their their stuff. And then they realize that it's just a mess. Right. Like, I've been doing this for 25 years. I tell you right now, you know, you're I can't get an SVM working in production. You think you're getting an AI model working?

35:58You have no idea what's going on in these companies. Right. So that to me is like, you know, that's just they didn't understand like their business model, what they were selling. And, you know, it's that kind of thing where trying to figure out, you know, how to get, you know, how to target the data correctly. That's exactly one of the things we'd like to know, for example, is, you know, this is stuff we'd like to get into more if we could, you know, is can we figure out which parts of the data are being learned and which parts are not? Right. And that you should be able to see. You pass the data through and you see which neurons light up and which ones don't.

36:35Those are the kind of things we'd like to do. you know, doing an ablation study on a model to figure out where in your data set is it insufficient, where, what part of the data do you need to fill in? Maybe like part of your data set is not fully filled in. And if you added more data there, maybe even fake data, you know, auto-generated data, you might be able to improve training, stuff like that. Those are things we'd like to understand better. Um, and you know, we're trying to, you know, and of course I'm trying to raise funding actually right now to do that, to develop some of this technology.

37:02Um, and, and I I think it's just, you know, these are things that require just a lot of experimentation to figure out. But we definitely see it now. Like there's a paper that came out just this week from Stanford by Chris Manning about how if you look at some of these large models like Lama or Quinn, that a bunch of the layers are not learning. They're looking at how the residuals flow through the layers, and they can see that some layers are not really converging. Some layers are not learning anything like, yeah, we told you that like three, four years ago. Like it's on the website. Yeah, I wish they'd used the tool.

37:35But, you know, it's a validation of what we've been saying for some time, is that our technology can detect this without seeing the data. And is there a one-to-one relationship between the characteristic that they were talking about and your characteristic? I'm not sure yet. We're starting to just dig into it now. We've seen cases, like we have a paper that came out, where I guess the broader point is like, This sounds like a really interesting way to characterize individual layers of models, but there are probably a ton of different ways to characterize it. Well, yeah, the main thing that we can do that no one else can do is we know the cutoffs.

38:15We know the bottom, we know the lower bound is two and the upper bound is like six. So you can use any model. You can measure the entropy, the distance. You can do whatever you want. You can see, you know, if you take two models. Are those practical bounds or theoretical bounds? Yeah, those are practical bounds, yeah. They work in production. Yeah. It really does work. Are they empirical or are they based on some theoretical analysis, like mathematical analysis? Combination of both. They're based on a combination. There's some new theory coming out. It uses techniques from theoretical physics called renormalization group.

38:49And so it's a combination of theory and experiment going together. You know, we have theories that show there's a boundary at two. We actually have two different metrics for the lower bound. And they have to line up. And when they line up, then we know. So there's some deep theory, like I got a hundred page theoretical physics paper on this thing to really justify it. But you see it empirically as well. And this crocking paper in particular is an important paper for us. We wrote this nice 10 page paper to really show, look, it works perfectly. And for me, the other thing is that I had somebody else do it with the tool.

39:20So it's independently verified. Right. You know, so these are real things you can use in production. I mean, they're real. So the challenge is, as he says, you have to analyze the data, you know, and that's, it's always tough. I mean, from a consulting perspective, look, I did a large project with Walmart. I tried to help them, you know, fix their search engine, you know, and you find out, oh, you know, you need to deploy your product on our Jupyter Notebook serving system. Okay. But you're a consultant. You're not allowed to have access to our Jupyter Notebook serving because we put all our customer data there, including all their credit card numbers.

39:56and you can't have access to that. Okay. You know, what do you want me to do here? You know, that's very common. The same thing, I did a project years ago, almost 20 years ago for France Telecom. And at the end of the, and we had to only use fake data. I'm not an EU citizen. Since I'm not an EU citizen, I can't access personal data. And so what happens in these big companies is that, you know, they'll stuff all the data on Hadoop into one system. And then it turns out they have compliance issues. GDPR, CCPA. You know, you, you, you can't access things. So it's a huge problem. And so part of Weight Watcher is that I wanted to build a tool that I didn't need access to any data because nobody ever lets me look at it anyway.

40:37So. So it's a problem and these are real production problems, you know, uh, and so part of, you know, trying now to convince people, you know, it's, it's an easier sell to give someone a tool and analyze the model. Cause you know, I mean, in a sense, I don't have the tokenizer. So I don't really mean I could kind of reverse engineer the model. It's kind of hard, but getting customer data and trying to look at customer data is much harder. You know, it's a different ask. It's a different ask in the org. There are compliance issues around it. So these are things that, you know, we'd like to try to do as a product level, but it takes some time to do.

41:13Yeah, this is a Herculean effort to try to get around a very pragmatic organizational problem. It's very hard, you know. I always tell people, you know, if you want to do something, you've got to, I have a friend who's, he's figured out a way to like automate all of the marketing in this company. And he's just, he's totally obsessed with prompt engineering. He's got this, and I, and he says, because the last, and they won't let him do a pull, he can't do a pull request. Cause he, you have to get, you have to get someone to sign off on the pull request. And so he built the whole thing in Google sheets, you know, in Google drive.

41:44And it's just this, so this, you know, something an accountant would build, you know? And I said, the last thing you want to do is be able to do a pull request because you'll get into the engineering org. And once you're in the engineering org, you have to follow their best practices. And those best practices may not work for what you're trying to do. They're best practices, they're just not best for you. And not a lot of that, we see a lot of this here is these kinds of tools to run in an org. Part of like I said, why fine tuning is hard. Well, I've had people come to me, we want you to build a model for us, but, and we want the model and we want you to design this model for us, but you're not allowed to ever look at the data for any reason.

42:23Cause it's a compliant, I'm not going to do, how can I constantly get it to work? I don't know how it would work, but you know, so there's those kinds of things are real and they exist in big orgs. And so, you know, looking at data and peaking, we certainly would like to look at and peek at what's going on with the tool. And this is something that we're trying to figure out how to, how to do that in a compliant way. In the paper, you benchmark the results against a few different approaches, activation sparsity, absolute weight entropy, absolute local circuit complexity. Talk a little bit about the prior work.

43:01So in this paper, there are people at Google DeepMind and people at Anthropic, which are trying to find out metrics to figure out what's going on in Grokett. and yeah we picked sort of the four top ones that people have in this in the theory and what we find is that none of them can detect this third phase of anti-grocking none of these metrics they have i mean they can kind of detect grocking like they can see the phase transition but because there's a threshold they don't know what the thresholds are they don't know that anti-grocking or this sort of confusion phase exists the generalization collapse and their technique Like, even if they could detect it, it's not clear what, you know, when it's happening and when it isn't.

43:42So that's sort of the part of the paper is not only can we detect this new phase, but none of the existing proposed metrics can detect it. Only our tool can detect it. So we can just, you know, we pick, you know, the top ones, deep mind and anthropic, right? Those are the big ones, right? And so that, and that's sort of the point is that whatever, you know, even like they have this thing called circuit complexity, where you're trying, you know, this is a big thing that's come out of anthropic recently, the circuits. How do you know the circuit is overfit? Can't tell. Can't tell. And it may see that it has higher complexity.

44:12Or like people do compression. Another thing is, oh, if you compress, the quality of a model is correlated to how much each layer can be compressed. Okay, that's true. No question about it. Question is, what about if you over compress? If you over compress, you overfit. That's the third phase. They can't detect that. So that's what we're able to do with the theory is that we can detect this overfitting phase. And, you know, this is something that I'll give you an example of where it might. And it's not always bad, but let me give you an example where you might actually want to do this. We've looked at models that are like segment anything models, the SAM models from Facebook that allow you to do zero shot vision learning.

44:54In other words, it turns out if you look at those, a lot of the early layers in the SAM models are overfit, according to our theory. So here's my— Early as in base. Yes, closer to—well, they're closer to the data. So you think of the early being that, you know, the closer the data, the earlier the model is closer to the label, the later the layer is. So what I think is happening is that there are these primitive features in all of vision. you know, lines, line segments and little circles and things like that, that our visual system picks up. I think what's happening in the segment anything models is that they look at a large number of natural images and they memorize these abstract features, primitive features in the data.

45:38And that's why they're able to perform good zero shot, zero shot learning, because most natural images, you know, in the natural environment, things are pretty much all the same. You know, a tree is a tree is a tree. And, you know, it could detect what a tree is because it can see, oh, that looks like a tree because it has the same primitive features. You know, a tree in China looks like a tree in the U.S. Although a pine tree is not the same as, you know, whatever, maybe a cherry blossom in Japan is not the same as a pine tree, but its primitive features are close enough that it can segment them and detect it as a tree.

46:10And that's what I think is going on. So it's not necessarily that overfitting is always bad. It might be something good. And so if you're trying to build a zero-shot learning model, you might want to optimize for this kind of overfitting in the early layers. I'll give you an example where it might be bad. We've looked at models like Lama Guard, these guardrail models. Guardrails show the opposite behavior. The layers near the labels seem to be overfit. And so what I think is happening, again, this is all conjectural. I haven't gone and done ablation studies. This is all conjectural. But I think what's happening in the guard models is that it's overfitting to, you know, whatever the examples you're giving to try to build the guardrail.

46:51And which is why they can be, you know, you can get around them. If you can get deeper into the model and find the more abstract thing to tell the model to think about, it's able to go around the guardrail. So it may be that if you're trying to build guardrails for models, you may need to have, and these are the instruction fine-tuned components on top of Llama. So if you're trying to instruction fine tune a model to give it a guardrail, you may want to have the overfitting go very, very deep into the back layers near the data to prevent that kind of backdooring. So we have cases where overfitting is not necessarily bad, but knowing what it is and how to detect it is what we're trying to do with the tool.

47:31So these are examples we've worked out and sort of the goal is, look, we have an open source tool. That's 200 ,000 downloads. I'll give it to you, try it out, and then talk to me about what you're doing. We'll see if we can figure out what's going on, how to make it work for you. That's essentially the Weight Watcher project. And when you compare against these other methods, are you strictly comparing against the ability to predict this thing that you made up? Or are there more like objective metrics that the other labs have published. I'll give you one that was very surprising. Okay. A couple of years ago, I took a look at all the existing base models, you know, Lama and, you know, whatever was back, the CUNY was hot back then, you know, this stuff.

48:20And I compared the average alpha quality metric to the hallucination metric. Okay. And it turns out, according to our theory, the closer you are to optimality, the more the model hallucinates. Interesting. Yeah. That's counterintuitive. Like what? Well, you would think that the thing that hallucinates more is not optimal in some way. No, it's more creative, right? Right, right. People are like, oh, we don't want them all. But other people, hallucination is a feature. In fact, it was so amazing. I gave a TED Talk on it. I mean, it was incredible. Like you think about talking to a child like my three-year-old niece, right?

48:57She'll just make stuff up. It sounds good. I'll make up a little story. It sounds good. I'll tell you, right? That's what they do. That's what children do. They're creative, right? They're exploring still. So these models, they seem to have this hallucination ability seems to be related to the optimal convergence properties. And so that's an example of where we looked at that. I was like, wow, that's really cool. And, you know, maybe it means that certain layers are contributing more to the hallucinations than others, right? And there's a trade-off between, you know, thinking inside the box, thinking outside the box, right?

49:30And I always say, you know, people want models to not hallucinate. In other words, you don't want them to be so creative, right? You don't want them going off into the tangent. Maybe you want them really to memorize the data and not be so good at generalizing. And that tells you something about what the layer should look like. And that's the kind of thing we've been able to figure out just using the theory and comparing to, you know, some things other people are doing. That hopefully, I don't think people are, I don't think they're gaming, maybe they're gaming the hallucination metric, but it'd be the other way, you know, trying to convince people that it's not hallucinating.

50:04But we definitely, that's the kind of stuff we're seeing. And by the way, I think ours is the, that was the only metric of all the metrics. That was the only one that actually correlated. Like the other stuff didn't really, it was just sort of random, like just random stuff, all these other, yeah, the other evals are not really correlated. Like these other evals people come up with, they don't seem to, you know, whatever they're doing is not as correlated with the quality metrics you have as the hallucination. I think hallucination is actually testing something fundamental about the model and how it's trained.

50:37And the other evals seem to be just maybe overfit to the data in some way, you know, very specific to the data you're using. So you test it on one data set, you see one thing, you get different data set, you get something else. Mm-hmm. Mm-hmm. Do you think of this as a mechinterp or mechanistic interpretability, or is that a tool that you're using, or is that an academic thing and you're trying to solve problems? Yeah, yeah, it's an academic. I don't actually read any. There's some people ask me stuff like that. I don't know. What are they doing? I don't look deeply at what they're doing because I know what I want.

51:15I know theoretical physics. I know what I want to do. So like that, that's what we had. That's why it's nice to have someone come and collaborate with me, you know, or, you know, we wrote this crocking paper and, you know, someone else, well, I wouldn't, you know, go out and find all these other metrics and see what they can do. The mechanistic interpretability stuff. I think I, you know, my impression is that it hasn't gone anywhere, right? Like there's like, you know, half a dozen things they've done. And what practical impact has it had on what anybody's doing? I don't know. I maybe internally in the throat, it gets useful, right?

51:50Cause they're doing things specifically for their model. But a lot of it seemed very specific to specific data sets and specific models and not something you can really apply. You know, if I'm fine tuning a model for, you know, Home Depot or someone like that, that's, that's an Atlanta company, Home Depot. How do I, how do I use it? I don't know. So it hasn't, it hasn't been something that we've looked at in any depth. I'm happy to collaborate with anyone doing it. I'm happy to kind of compare what we're doing to them. But we have so many things on our plate that we just haven't looked deeply at it.

52:25So it's the same thing. Sometimes people ask me, well, how is your theory related to like reproducing Colonel Hilbert's basis? Some academic. I don't know. I've never read any of the papers. Colonel is something from physics. What do you guys, I just don't know. I mean, it's, you know, if you're doing it, you tell me, I'll explain to you what I'm doing. I'll explain to you what I'm doing and you tell me how it's similar to what you're doing. I don't know what you're doing. The problem is there's just so much going on. Every day you wake up and, you know, there are a thousand new papers that have been, you know, I think a hundred ML papers being published every day.

52:58So how do you keep up with everything? Right. How do you keep up with everything? I rely on other people to come to me with interesting stuff. You know, I go on, I, if you, you know, that's it. You know, I go on Twitter. I spend a lot of time on Twitter trying to see what are people talking about on Twitter and try to follow good people and listen to podcasts like this and try to keep up. But, you know, it's everything's moving a thousand miles. You know, you're moving at a couple hundred miles an hour. So you're just doing the best you can. I sort of lucked in that, you know, I did this stuff in the 90s.

53:32and all the guys I worked with, they, they went off and the guys who were really sharp went off and became quants and are now running like, you know, the investment arm of Dubai or something like this. Um, uh, you know, or they, they went off and, you know, they, they went off and started companies, but most of the people in the, in the, the physicists, they sort of, people sort of forgot about how you could apply theoretical physics to AI. So I kind of snuck in, you know, through the old, like an, it's like an old back door in the lab that people forgot was there. And I was able to sneak in and do stuff.

54:03If people, I think if enough people remembered all this theoretical physics stuff that the opportunity wouldn't be there because they would have known about it. But I, I sort of, you know, I got kind of lucked out that I'm old enough to remember this stuff. You talked about the analogy in a quant for all of this stuff. Like what's the physics analogy for all of this stuff? Well, you know, physics and quant are very close, Right. So in physics, this whole idea of we talk about self-organized criticality and the emergence, signatures of emergence, there's a technique in physics called renormalization group.

54:38My undergraduate advisor, Ken Wilson, won the Nobel Prize for developing renormalization group. And it's it's actually a really fundamental thing in theoretical physics to understand the properties of the universe. Why do electrons have mass? Why do quarks have mass? Things like this. And how do you describe it? And it turns out that it can be used to describe things like when water boils. Okay. When you have water and you boil it and you see all the bubbles. Phase changes? Yeah. It's a phase change. Yes. It describes the phase change. It's the phase boundary. So renormalization group is the mathematical theory used to describe phase boundaries between, to describe phase changes.

55:14That's the theory you apply for Weight Watcher. And so it turns out that you can think of this. It's kind of convenient that everything's a matrix. Yes. Well, you know, I grew up in the Cold War. I learned all this math. It turned out to be useful. So it turns out like if you think about boiling water, when you watch water boil, and you look at the size of the bubbles, there are all sorts of bubbles. There are little bubbles, medium bubbles, big bubbles, all sorts of bubbles in the water, right? There's not like one size of bubble. That idea is those are the correlations in the system. There's little tiny correlations and there's medium-sized correlations of really big fluctuations.

55:54That's analogous to the information in the layer. So when a layer is learning, it learns little bits of correlations between the training data. It'll learn sort of medium-sized correlations and it learns long correlations between the whole data set, across the entire data set, right? There's correlations across all the data. and they're sort of equally distributed. That's the analogy from physics. The fact that the bubbles, when you boil water at the phase transition between water and a gas, that all those bubbles are basically the same, are different sizes and shapes. They're all circulars. They're all different sizes.

56:34It's the same idea that when a layer is learning the information in the training data, it has to learn all the correlations of all the different sizes. and if it doesn't learn all the correlations, then it can't generalize that well. And if it learns, and that's the idea, and if you cross the boundary, you know, you might be, you might overdo it. Or so you're like, you're going from water to ice, you freeze out. And if you freeze out, you get stuck. And if you freeze out, you're overfit, right? So there's sort of this boundary between ice, water, and gas. And, you know, water is sort of, it can adapt to any situation.

57:16You remember like Bruce Lee said, be like the water. If I pour the water in the vase, it becomes the vase. I pour it in the cup, it becomes the cup. I pour it in the glass, becomes the glass. Be like the water. Being able to generalize is like being like water. If you don't learn enough information, it's like you've overboiled and you're a gas, and there's just no structure. There's nothing there. Just random. And if you learn too much, you freeze and you're like ice and you're frozen and now you can't generalize. That is exactly the analogy. That's a great analogy. Yeah. Because it comes from, that's the physics.

57:50And it's the exact same mathematics and physical theory you use. You know, there's a, I'll give you, like there's a paper that came out today. It was, if you take a reinforcement learning system and you train these models on random reinforcements, like 25 % of the feedback is just random, they still get better, right? How could that be? How could you fine-tune the model on random stuff? That's like RL dropout. Yeah, right. It's because it's like the system is frozen and you heated it up and cooled it down again. That's so much of like all these training recipes is like how do we introduce just enough noise?

58:31Just enough random. Yeah, yeah. Yeah, that's right out of that's right out of physics. The idea is what we call spin glass theory or glass theory that when systems become too brittle, basically they're brittle, right? They're brittle and brittle systems are easy to bake. You think about a metal. If you want to make a metal, that's not probably have to heat it up, cool it down, heat it up, cool it down, heat it up, cool it down, becomes strong. If you freeze it really quickly, it becomes brittle and it will crack. It's the same thing. And in the math and the physics, all the same physics and math.

58:59like all the math used to describe that in the physics is the same stuff I'm using to describe neural networks it just turns out that uh everybody forgot it yeah that's all because you know we all or they all died you know they're all you know they're all right they're all retired right I shouldn't say that should be they're all retired right all the guys are retired didn't have to work anymore but you know there's a saying in science that science progresses when old scientists pass away. Right. Because they, they take, they, they, they stop, they, the old guys stop the new guys from doing anything new.

59:33They don't want anything new. So when they pass away, now you can start publishing, you know, they're no longer there interfering in what you're trying to do. Yeah. But it, it, it turned out that, um, it turns out that this stuff does work. If you spend the physics theories are useful for some things. And, and we're trying to basically, and, and the other thing is it's, to me, it's, it's really important that you have an open source tool. and the work be 100 % reproducible. I need to be able to give the tool to somebody else and they need to be able to run it and try it. Sometimes this stuff works out.

1:00:04Sometimes we don't get the right result. We don't know why. I don't know. What do I know? It doesn't describe, you know, 80 % of the time we know what's going on. But 20 % of the time we see things we don't understand and we're still, you know, picking at it to try to figure it out. But hopefully the tool is still useful to people. And that's the goal of this. So how does the renormalization group stuff and the physics, you know, basis of this lead to HTSR and this whole power law stuff? So in renormalization group, there's this idea of a volume preserving a scale invariant transformation. And that's the idea.

1:00:42There's a scale invariant transformation. Things operate, you have different scales. The physics, if you change the scale of the system, the physics remains the same. just some of the numbers change like the the mass of the electron or the mass of a quark might change depending on the scale see this thing you know the scale okay now is this related to like gauge invariance and that kind of stuff do you know that kind of i do know kind of uh to be technical because i'm using a hard cutoff technically i'd probably my system is probably not gauge invariance so gauge invariance a different kind of invariance this is a scale invariance and if you do renormalization group like i probably need like loop corrections to do the gauge invariance but But yeah, but it turns out that, so when I was formulating the theory, I was trying to understand, I have this HGSR theory, we have this alpha metric, and it seems to be this universal metric.

1:01:29Like every model seems to like to be at two, right? Two seems to be universal. There's this random matrix theory tells you at the bottom of the universality class, there's alpha equals two. There's a little tiny universality class in between. For some reason, two is special. So, you know, I got to figure out a way to derive this from first principles. And in physics, that would be called a critical exponent, a universal critical exponent. Whenever you measure a system near a phase transition, simple systems, simple systems, they exhibit critical exponents. However, the change in, say, the heat capacity is governed by some power law.

1:02:07And it's a critical exponent. And it's the same thing you see in the neuroscience, the self-organized criticality, that these neurons seem to have some sort of critical exponent. they all approach the same, all the data approaches the same exponents. Maybe there's some fluctuations because it's, you know, these are small systems compared to like, you know, boiling a pot of water, you know, with, you know, 10 to the 23rd atoms in it. So this alpha is a critical exponent. So I knew that I need to be able to figure out a way to derive this. And as I'm going through the derivation, sort of in the back of my mind, you know, there's got to be some, you know, critical exponents are typically associated with renormalization group.

1:02:43So when the renormalization group theory applies, you typically expect to see a critical exponent. So we would expect that alpha equal two is that. So in the course of trying to derive the HCSR theory, I realized I have to do this renormalization. In order to make the theory, the math easier, I have to make this assumption about a scale invariant transformation. And then I realized, oh, that's renormalization group transformation. And then I tested it empirically. And it turns out that when you measure alpha equals two, that funny parallel, it turns out you can also measure the scale invariance.

1:03:17You can test whether the system is scale invariant by looking at the trace log of the eigenvalues or what's called the log determinant. So it's a log determinant relation that arises or the trace log. So it's simple. You just compute the eigenvalues using SVD or whatever eigensolver you want. You simply sum up, you take the logarithm of them and you sum them up and you start at the tail and you work backwards. And when, as soon as you get close to zero, boom, that's, that's where the information concentrates. That's the scale invariance. And it turns out in that, and that's the connection. Now I haven't been able to prove that the alpha equals two is in fact the critical exponent for this transformation.

1:03:55I think I know how to do it, but you know, I, I do stuff like that. I'm going to be living out of my car. You know, I need to get away to fund this operation. Right. Right. So, but I'm fairly certain that the alpha equals two is related to this. You can measure them both. And if you, and it turns out you can measure the scale, you can use, it's called the dead X condition because the determinant of the, of the exponent should be one. So in the Weight Watcher, you can measure it and you can see that if you violate the dead X condition, the scale invariant or normalization group condition, if it violates it, you seem to be overfitting.

1:04:28And it turns out like we didn't publish this in the grokking paper because it was too much because it'd be like 50, you know, but it turns out it also works in the grokking paper we published that you can see the dead x condition change. So it turns out that these two things are very related and there's something you can actually measure that the idea, I'll tell you where the idea came from. If you're curious is that we had this sort of side project at BlackRock that was sort of like a side project. We were trying to figure out, can we detect when the market's going to crash? And there was this, you know, cause yeah, cause the market crashed, you know, and we're like, what if we get detected?

1:04:59is there's a theory by a guy named Dieter Sornay, who has this theory about why markets crash. He actually has a book called Why Stock Markets Crash. They published like 20, 25 years ago. And the theory says you can measure the signatures of renormalization group in the stock market. And I was spending our spare time. I was saying we could see them measure. And it's a different technique than what I'm using for Weight Watcher, but it starts by looking for parallel signatures and looking for something called long periodic fluctuations around the parallel signature. I have a blog post on it for Bitcoin.

1:05:30Can you detect when Bitcoin's going to crash? You know, that kind of stuff. So it was like this hot project. Well, you know, it's, I'm not going to trade it. I'll let you trade it. You know, you can protect it, but can you trade it? And it was like a hobby project we were doing. Because I work with, you know, with one of Dieter's classmates at BlackRock. And we were doing this. I knew about the work. And we were sort of doing this in my spare time. Because I knew about this renormalization. It was like one of these like crazy applications of renormalization group. And it turned out, but it was this idea that I could measure the signatures of scale and variance in a physical system.

1:06:04And so I had this idea, there's a way to measure scale and variance in a physical system. And they use it to do things like, can you predict when, and when an avalanche, like when you predict, if you have a crack in a material, can you predict where the material is going to like, is a bridge going to collapse? And so you can look for cracks, the best, basically the distribution of the cracks. and as the distribution of the cracks start following a power law, now you've got a problem. The bridge is probably going to collapse. And then that's sort of like the practical aspect of it. I said, gee, I wonder if I could apply this to neural networks.

1:06:34And it turns out you can. It turns out it works. And we can predict when the crash, in this case, being the overfitting, you know, the generalization collapse. That's the crash. And so it turns out it works. And so that was sort of where the idea came from, just sort of, you know, doing sort of these, you know, When I was at BlackRock, my job was to come up with crazy ideas to predict the stock market. You know, just nutty things, you know. And the point being that only the nutty things are going to work because everyone's tried everything else. So you got to try something no one else has ever tried.

1:07:06Otherwise, you can't predict anything. So I would just come up with this sort of this nutty stuff all the time. This is one of the nutty ideas that we had. And it turns out you apply it to neural networks, it actually does work. So these old theoretical physicists, you know, they were doing, you know, they're smart guys, right? You know, they made the bomb. So, you know, they know what they're doing. It does work. It's just somewhat remarkable to me that, and so that's what's going on. And I'm happy to go through all the math and be a nerd out on as much as you want. But that's sort of the, you know, if you derive all the, I have this long paper.

1:07:38It's about 120 pages long. It's in draft form. Every week, Mike and I, Mahoney, I ask him, try to find some typos in it. Because, you know, it's got like 500, I think it's got like 300, 400 equations. Is it up on archive or something? Is it something that folks can access? Yeah, because I don't want to put it in the archive until we find all the typos, but I have it on a GitHub repo. And it just says draft. And I'm happy to share it. If you find, like last week we found I was missing a trace operator on one of the, you know, an appendix A3. There was no trace. So until we get all the typos out, it's just, we don't have the last time to read this thing.

1:08:11You know, it took me a year probably to write it. But it's something we'll eventually get on the archive, maybe another year. Like the HGSR theory paper, it took us three years to get it published. So that work was done, like it was done in 2018 and we didn't get it published until 2021. It just took that long. So maybe when I retire, I'll get the CETL paper published. But, you know, I'm happy to have anyone read it. And, you know, it's, you know, if you want to, if you like theoretical physics and you want to nerd out and spend, you know, a month, you know, on vacation doing this, uh, probably you'll get divorced if you do that, but, you know, maybe you are divorced to give you something to do.

1:08:52In the, not the HGSR paper, the Gronking paper, one of the noted limitations is that you validated all this on a three-layer MLP. And MNIST is the data set. That, in some ways, is pretty far from how you'd like to use these ideas. Well, you know, it's a trade-off, right? If you think about doing development of the Schrodinger equation, okay? I used to do quantum chemistry. So we run quantum chemistry on big systems, right? But you got to start with the hydrogen atom, right? You got to start, you have to start, if you have to make a theory, you at least be able to do small things. I see it more like the Bohr atom, like what I'm doing.

1:09:33You know, it's, Mike always says that's so arrogant of you. I go, well, the Bohr atom's wrong. The Bohr model is wrong. What do you mean? You know, but it's more like the Bohr model of the atom. You know, we're trying to come up with the Schrodinger equation. Once we come up with it, you know, like when I was in grad school, we would study, you know, like hydrogen diamer or, you know, or nitrogen diamer. And you'd see these interesting properties. And then you have to go and apply it to big systems. So I just don't have the capital. I'm non-anthropic. I don't have any funding for this at all. This has all been a hobby project, right?

1:10:00This is all my spare time. I'm trying not to get divorced, you know, if I keep working on it. But, you know, we only have so much compute resources. And we need to understand these very fundamental things. So my approach to this is we try to study small problems and understand them analytically. And, you know, do analytics here. We have like deep, I mean, the way we watched this idea of the renormalization group, I mean, this really is 300 pages to derive the equations. And in the end, you get one metric, which is one little 10 line subroutine you can put inside the code and you test it. And so my idea is, look, I make an open source tool.

1:10:34I give it to people. You can test it on bigger systems and see if it's useful. And that's the idea. You know, as we tell you, like, well, like we don't understand things like make the Gronk and we'd like to study bigger problems. but you know if you look at an attention model there are a couple things going on there are you know like they have the attention block and yes well does it apply to llms okay well it turns out it seems to work really really well for the internal parts of the attention block the k and q matrices it seems like like even in something like llama where you know half the model seems to be overfit not the k and q's the k and q's line up really really nicely with the theory for some reason, the V matrix, it'll go like this and it just blows up and comes back.

1:11:18Like what happened there? Like, so we know that like, we know it works inside the attention block. Does it work for all the layers of the attention block? And if not, why not? So those are things we're trying to understand better. And it's just really hard. You know, it's, and those models are hard to train, right? So we're trying to look at like what model, maybe BERT is, that's a, that people are fine tuning BERT in production, for example. That, that's still a common thing people do, right? for building classifiers. And so that would be like the level of model we might look at next. All of the theory has been developed on like, this is hydrogen model stuff, right?

1:11:52It, you know, it's the Bohr model, right? MLP unminsed. And I said, and we give the code away, try it on fashion mints, try, you know, we did some experience on CIFAR-10, CIFAR-100. You know, there is a point where, you know, it's, you know, I would love to do it on, you know, big systems, right? But, you know, it's compute, right? It computes, it's expensive. So being able to do observational studies, you know, instead of having to run big experiments, instead of doing like running experiments on small models, which compute is expensive, my approach is sort of like doing meta experiments by doing observational studies on hundreds of models, right?

1:12:33So we have a paper in Nature where we looked at the time, we looked at like 500. At that time, there weren't that, even then, you know, 2020, I think we did it, 2019, 2020, look at 100 open source models. I think it was 500. We looked at 500 open source models and compared how the average metric compares to those 500 models. Today, I have a website on the web. Wait, what's your website? I just have, you know, different models, Llama, Quinn, DeepSeek, you know, Falcon. You know, we just try to look at the big models and write up reports and show them. And that's basically the best we can. So trying to do observational studies.

1:13:09You know, I'd like to study like, you know, you know, 100 ,000 models on Hugging Face. But, you know, that would probably bankrupt me and my VC and I'd probably be sued. But that's what we're trying to do. That's awesome. Awesome. Maybe an interesting question, since you are kind of out in the field working with customers that are trying to put Gen.AI to use any beyond the stuff that we've discussed thus far, you know, hot takes, hard-fought lessons in terms of, you know, making this stuff work? Look, I think that the magic has been the prompt engineering. And that has opened the door for being able to do things when you don't have data.

1:13:54Because data is so hard to get. And training models is hard. And it's hard from a personnel perspective. Like you have to pay a lot of money to hire people to know what they're doing. And if you don't know what you're doing, you can really spin your wheels and not make any progress. and the gen AI stuff has been phenomenal. I don't think anyone really expected in context learning to work the way it does. Now, I see a lot of people trying to do, I've worked a lot in cert. So I see a lot of people trying to do RAG and then I realized RAG is old. RAG is from the nineties. That's called latent semantic analysis.

1:14:29That was invented in Chicago. I've been trying to do RAG for years. It never really works. And that has been, you know, people think, I'll just stick stuff in a vector database and it'll be fine. And you just get all sorts of nonsense, right? I mean, it's not going to, there's no guarantee. So that I think is something people are, and I've done a lot of work in search, most of my, a lot of industry work in search. And you sort of see there's a bit of a naivety about it, that there's a lot of emphasis on getting the search engine to work, the plumbing. So right, getting the vector space, getting the, getting the documents translated into an embedding?

1:15:11And how fast can you do that? Because that's slow and it takes compute. And so you have to pay for that. And what size embedding should you use? And getting the vector database up and maintaining the database. And should you buy a vector database or use an off-the-shelf one? I mean, the original one was developed by Spotify. Something called Annoy. There was an Annoy package over years ago. And then there was the one from Facebook and then one from Microsoft. And now there are all these commercial versions. And what I've seen working with search people is that 90 % of the resources are more go into these, getting the thing operational.

1:15:44And then you get to the relevance and you ask, oh, I can just do RAG and I'll get relevance. It's terrible. You know, it just doesn't really work. And the reason is because it doesn't learn from the clickstream. You want to learn if you're doing a system where you're trying, you're either, you're trying to learn from the clickstream, You got to learn from the clickstream. You just have to train a model on the clickstream. And, you know, that means you either have to put some model on top of RAG. To put that in other words, you talk to enough folks that are doing search and retrieval, like the thing that they're focused on is relevant as opposed to plumbing.

1:16:23And RAG, the thing that they're focused on is, you know, that retrieval step as opposed to the generation. and uh you know what you're highlighting is this idea that um and also you know in traditional search like that was what they were like that was their job and their job was to make sure that when the user searched they get the results that they're looking for and it's not like a one and done it's like you know i've worked with i've worked with guys from google ebay walmart you know the big the big the big engines yeah yeah yeah and i tell you 90 of the you know it's just not relevance is hard and this is the point that i'm getting at yeah and people work on it for years they work over years and and people who try to get into it are very naive and you know they i mean i'm not sure i've ever seen an a b test that i believe right i mean i've been like even like i've seen clients who don't understand like they'll put all this effort in the engineering you have to run an AA test just to measure the variance.

1:17:28They don't do it. Like, okay. I mean, I've worked on internal search. Yeah, I've worked on internal search, semantic search. I mean, all the different variants of search. I invented technology for search. We worked on eHow, first billion dollar represents Google. And, you know, it's, you know, aardvark acquired by Google. Search relevance is widely, widely underestimated. It's overestimated how hard it is. I mean, excuse me, underestimated how hard it is. It's misappreciated for how difficult. And there's very little good academic information. In some sense, search relevance is like trading on the stock market.

1:18:03The people who really know how to do it aren't going to tell you what to do. Because that's where the gold is. And I think that thinking that you can just do rag and expect that to just give you good results is, you know, are you actually, you know, if you have good results? Do you even know what your bounce rates are? I mean, you know, I've seen, I mean, I've seen cases where we, you have search. And so part of the machine learning of Weight Watcher is like, yeah, I'd worked in search so much. I was thinking about, you know, I want to redeploy the search engine every day. I want to retrain the model and I want to make sure it doesn't go bananas.

1:18:34And sort of the motivation was how do I model the thing? I can't do A-B experiments all the time. They're expensive. A-B experiments are expensive. They're hard. They're difficult to interpret. And a lot of data. And a lot of data you have. and people do all sorts of, they'll do things in production for their A-B tests that you really shouldn't be doing. They're running multiple tests at the same time and there's leakage between the experiments. Stuff like this. I'm like, it's useless. Like the noise is, I was trained as a physical scientist, man. I mean, I did quantum physics. I mean, I understand how an interpretive experiment.

1:19:08You know, you can't have, you can't have error bars this big, you know, if you're only looking at that much. And it's sort of like, What are you doing? And, you know, it's just well, and that part of the problem with the rag stuff is that it doesn't take into account the clickstream. And so you have to somehow do that. Meaning, unless you've got either an implicit or explicit feedback mechanism from the user, then you're flying blind. And you see now with AI, this is part of the problem with fine tuning. is really what you'd like to do is fine-tune some sort of adapter on top of the rag to adapt it so it will work well.

1:19:48That thing has to work. Think about a fine... You can't have... You're running inference. If you're in e-commerce, you've got a 200-millisecond, 250-millisecond SLA. You can't... How are you going to run inference on that? I run SVM. The SVM has a... It runs in 30 milliseconds. I could, you know, you know, it just, all the, all the fact, all the, in fact, I can run, you can run XGBoost in 30 milliseconds. You can run the SVM in under 10 milliseconds. All the overhead is based, but you know, it's just latency from the network, but you're trying to run, you're trying to do a model and you're trying to do production search and you're trying to run inference on this huge thing.

1:20:27You know, you've got to boil it down. Even Facebook only uses like simple embedding models. Here's one for you. You know, Facebook doesn't use PyTorch production. They use cafe too. Like, and they have simple, you know, simple models. So I think a lot of what I see sort of is what I'm seeing is that there's just, you know, search is still very, very hard. And, you know, and in rag is not a magic box. Um, and you're seeing now what's really interesting are integrate, you know, people trying to do integrated LLMs that learn how to do search on the fly, right? The search is integrated into the LLM somehow.

1:21:03And I think it's just, it's really, it's just, that's probably the, in terms of being in the field, like trying to get this stuff to work. And then the RAC stuff itself, you know, you have prompt engineering issues. You have to ask the right thing, you have to get the right documents. So I think a lot of that is, you know, there's sort of a bit of a naivety about it. And so if you're trying to train in, And that still is probably one of the biggest challenges I see. And again, and also because you have to, you have to track the click stream all the way through, right? You have to track from the user, know what the user is doing.

1:21:34So I think there are a lot of people wanting to do this kind of stuff. And it, there, you know, there's always a lot of, just a lot of money left on the table, right? To, to get it right. And, and whether this technology can help you, you know, part of the idea is, you know, can you fine tune a model that you can put like a little tiny model, an adapter model you can put on top of the RAG system. You know, if you get sort of the, if you get the prefetching right, right? You get the broad spectrum of it. You get the document set. You can do relevance in real time with this stuff. And that's sort of, you know, some of the motivation to doing this.

1:22:08And I think from the real world, you see, and of course now we're seeing that Google for the first time since eHow really is starting to lose traffic, right? They're losing traffic. And I don't use Google for anything anymore. I mean, I use Gmail, but, you know, maybe Google Docs. But what are you using instead? I use OpenAI. I have O3. If I'm paying for it, I'm going to use it. I use it for everything. If I don't use O3, I'll use Grok. Maybe I'll go to Gemini. But, you know, Grok is, you know, O3 is probably O3 is just slow. It's just slow. Yeah. But I'll go to Grok if I need something quick.

1:22:43You know, Grok is quick. And then I'll go to O3 maybe for, you know, and I started using Codex. Codex is pretty cool. You know, I'm not a cursor guy yet. But I, you know, I try to, I don't know, I like this thing. It's my code. You know, you tell it to fix a problem and it just erased. Codex did this this morning. It erased one of my unit tests. I fixed it. It runs now. Where'd it go? It's completely gone, you know? I like Codex because you can check the pull request. Like, you know, you make sure it doesn't delete everything. But I think that's, you know, but search to me is still like, that's one of the big ones.

1:23:17and one of the main uses for this technology because it really gives you a natural language interface to search. I was almost suggested the other day, I wanted to remake Aardvark. Aardvark was a product that was sold to Google. I was a scientist at Aardvark where you would ask it a question and we'd go out and find someone on the internet to answer the question for you. And I wanted to integrate this into the LLM so when it lies to you or screw something up and you don't know how to fix it because you're way out of your, you know, you're punching way out of your weight class. You don't go find someone to fix the problem that the LLM created.

1:23:52Like find a real person, like find an expert to fix whatever the LLM broke. I thought that would be a good product, right? But, you know, the technology is amazing. I'm very bullish on it, but, you know, it's not a magic, you know, I'm a silver bullet, right? You still have to pay attention to what you're doing, but it is an amazing technology. And that's why I'm sort of all in on it, right? Sort of like, you know, Sergey Brin said, he came out of retirement just to do this. So I'm sort of like, that's what it is. It's just all day long. It's all we do now. And, you know, having worked in AI and machine learning since the 90s, it's unbelievably exciting.

1:24:30But it's also a little overwhelming sometimes, too. Which is why you have these great podcasts. You can try to figure out what's actually going on. Absolutely. Absolutely. Well, Charles, it has been great catching up and digging into what you've been working on. Thanks so much. Hey, thanks for the time. I really appreciate it, Sam. All right. Thank you. All right.

From the publisher

Today, we're joined by Charles Martin, founder of Calculation Consulting, to discuss Weight Watcher, an open-source tool for analyzing and improving Deep Neural Networks (DNNs) based on principles from theoretical physics. We explore the foundations of the Heavy-Tailed Self-Regularization (HTSR) theory that underpins it, which combines random matrix theory and renormalization group ideas to uncover deep insights about model training dynamics. Charles walks us through WeightWatcher’s ability to detect three distinct learning phases—underfitting, grokking, and generalization collapse—and how its signature “layer quality” metric reveals whether individual layers are underfit, overfit, or optimally tuned. Additionally, we dig into the complexities involved in fine-tuning models, the surprising correlation between model optimality and hallucination, the often-underestimated challenges of search relevance, and their implications for RAG. Finally, Charles shares his insights into real-world applications of generative AI and his lessons learned from working in the field.

The complete show notes for this episode can be found at https://twimlai.com/go/734.

More from The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

All 156 episodes
Grokking, Generalization Collapse, and the Dynamics of Training Deep Neural Networks with Charles Martin - #734The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) · 1 h 25 min
Listen in VO