1025: Word Gravity: How Transformers Bend Space, with Dr. Luis Serrano

8 Sep 2026 · 1 h 11 min · 28 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Transformer “word gravity” and curved spacetime analogy; why RAG isn’t an agent; how GRPO works (reinforcement learning behind DeepSeek-style reasoning); plus production agent design and evaluation.

Guest backgrounds

Dr. Luis Serrano, founder of Serrano Academy; visual, intuitive ML explainer on YouTube (>200k subscribers). Previously worked at Apple, Cohere, and a quantum computing startup. Author of Grokking Machine Learning (2nd edition releasing Sept 29, per episode). Also co-runs courses with teams like Ragpack (Jay Alomar, Josh Starmer, etc.) and Neural Maze (agents).

Key claims

Attention can be viewed as “words bending space,” with K/Q acting like directional “pull” and V controlling which properties get moved (not symmetric like gravity). Function words (e.g., “to”, “the”) can show sharp curvature in some layers. RAG is an LLM workflow, not an agent, because agents make decisions (e.g., when to fetch tools/info). Agent evaluation is harder than RAG evaluation because you must score every step/trajectory and verify tool use.

Notable examples

“river bank” embedding analogy; Eddington’s 1919 eclipse experiment mapped to bank/river across transformer layers; canoeing bank (verb can pull bank toward nature but not change noun/part-of-speech).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Reconnecting with Luis Serrano

0:54 to 2:18

Jon and Luis discuss their recent meeting and past experiences, including Luis's work journey.

“This episode of Super Data Science is made possible by Anthropic, Groby, and the Open Data Science Conference.”

The Ragpack and Neural Maze Collaborations

2:18 to 4:38

Luis shares insights about his collaborations with Ragpack and Neural Maze and their course offerings.

“Things here, things there, a lot of things going on.”

Expansion of Luis's Educational Efforts

4:38 to 8:00

Discussion on Luis's growth from 140,000 to over 200,000 subscribers and his approach to content creation.

“And then they knew Martin Grotendorsk and Chris McCormick.”

Second Edition of Grokking Machine Learning

8:00 to 13:25

Luis details the updates and changes in the second edition of his book, Grokking Machine Learning.

“Or have you kind of just been following your whims and kind of like, oh, like Jay and Josh reach out and you're like, cool, like I'd love to do a course with you guys.”

Generative Learning in the New Edition

13:25 to 14:00

Luis emphasizes the importance of including generative learning in the updated Grokking Machine Learning book.

“And you know, surprisingly, like things worked actually really well on scikit-learn.”

Generative Learning in Machine Learning

14:00 to 15:12

Learn about the incorporation of generative learning in machine learning literature.

“You know, like there was the explanation, but then, you know, a few things clicked after that.”

Generative Learning in Machine Learning

15:15 to 15:55

Learn about the incorporation of generative learning in machine learning literature.

“but neither is built for complex, constrained decisions.”

Understanding 'Grok' and Its Significance

16:07 to 18:06

Explore the concept of 'grok' and its importance in understanding complex ideas.

“I'm sure I'll see you in Ontario soon and I can get a signed copy.”

Insights on Explaining XGBoost

18:07 to 20:06

Hear about the journey of understanding and explaining XGBoost in detail.

“you come up, some of the visual, cartoony ways or the grokking ways you come up with, they take a decade to think of.”

Using AI Tools for Learning

20:07 to 22:24

Discuss the role of AI tools like Claude in enhancing learning and understanding.

“Like I, I remember at the beginning, I, you always have that example of King, Queen, man, woman, and there was like the parallelogram.”
Show all 28 chapters

The Value of Human Insight in Learning

22:25 to 24:45

Understand why human insight remains crucial in the age of AI and LLMs.

“Why do you think listeners should still buy books?”

The Structure and Value of Books

24:46 to 25:59

Learn about the advantages of books in delivering structured knowledge.

“You know, it's a very human thing still.”

Transformers and Curved Space-Time

26:00 to 27:35

Delve into the relationship between transformer architectures and concepts of space-time.

“So in addition to your new book coming out, something you've been writing, something you've been working on, something else that you published is actually a new paper.”

Breaking Down Attention Mechanisms

27:36 to 28:00

Explore the intricacies of attention mechanisms in transformers and their applications.

“Then they gave me the formula made even less sense.”

Understanding Word Gravity in Transformers

28:00 to 32:28

Learn how the concept of 'word gravity' explains word relationships in transformers.

“I don't remember which channel was this.”

Gravity Analogy in Neural Networks

33:33 to 40:31

Explore how gravity can be used as an analogy to understand relationships in AI models.

“In that analogy, what plays the role of gravity in the same way that you described the sun bending space so that the earth comes toward it.”

Exploring Function Words and Curvature

40:31 to 42:00

Learn about the unexpected influence of function words on semantic curvature in models.

“Yeah, that was a really interesting experiment because some words are pretty flat.”

Introduction to Agents and Course Insights

42:00 to 43:30

Learn about the concept of agents in AI and the author's recent course on it.

“Well, thanks for talking us through that paper, The Curved Spacetime of Transformer Architectures, Decipio, Diaz-Rodriguez, and Serrano.”

Course Structure and Learning Outcomes

43:30 to 45:54

Explore the structure of the course and the practical projects undertaken.

“They have so much, the Substack subscribers and very, like, they make so much content.”

Predictions on Tools and Agents

45:54 to 47:24

Discussing the predictions made on AI tools and agent use over the years.

“We use Google Cloud and then Cohere for the LLM.”

Importance of Agent Evaluation

49:10 to 52:52

Understanding the challenges and importance of evaluating AI agents.

“who want to get into getting agents into production, get the grokking agents in production course.”

GRPO vs. PPO in Reinforcement Learning

52:52 to 56:00

Learn about the differences between GRPO and PPO and their implications.

“you've also built a whole free course and video series on reinforcement learning for LLM.”

Understanding GRPO and Mathematical Reasoning

56:00 to 58:16

Learn about the nuances of mathematical reasoning in AI and how GRPO optimizes performance.

“It uses compilers for compiling the code.”

Quantum Computing Innovations

58:16 to 59:39

Discover the latest developments in quantum computing and its implications for machine learning.

“And I'll have a link to materials on that, of course, in the show notes as well.”

AI and the International Math Olympiad

59:39 to 1:02:35

Hear insights on the impact of AI on math competitions and the excitement it generates.

“I'll say the two things and we'll have both recordings.”

AI and the International Math Olympiad

1:04:02 to 1:04:34

Hear insights on the impact of AI on math competitions and the excitement it generates.

“Hey, hey, this is your host, John Krohn.”

AI and the International Math Olympiad

1:04:37 to 1:04:51

Hear insights on the impact of AI on math competitions and the excitement it generates.

“We're looking forward to hearing from you.”

Book Recommendations and Closing Thoughts

1:04:51 to 1:08:22

Explore book recommendations and final reflections from the guest and host.

“Across the board, reinforcement learning, agents, curved space-time of Transformers, and of course, Grokking Machine Learning, the second edition, which will be available to listeners very soon.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Jon Krohn:What do planets orbiting the sun have in common with the words inside LLMs? According to my guest's new research, both bend the space around them. Welcome to episode number 1025 of the Super Data Science Podcast. I'm your host, Jon Krohn. Today's exceptional guest is Dr. Luis Serrano, founder of the Serrano Academy, whose YouTube channel has over 200 ,000 subscribers hooked on his visual intuitive explanations of machine learning. Luis previously worked at Apple, Cohere, and a quantum computing startup, and the second edition of his best-selling book, Grokking Machine Learning, is out this very month.

0:38Jon Krohn:In today's episode, Luis explains his mind-bending new paper on how transformer architectures mirror Einstein's curved spacetime, why RAG isn't an agent, and how GRPO, the reinforcement learning technique behind DeepSeq's reasoning breakthrough, really works. Enjoy. This episode of Super Data Science is made possible by Anthropic, Groby, and the Open Data Science Conference. Luis, welcome back to the Super Data Science Podcast. Oh my goodness, so good to see you. How are you doing, man? Good, John. Thank you so much for having me. Big fan of this podcast, so happy to be here. Yeah, we had you on the show a couple years ago, but I actually met you in person about a month ago at the time of recording.

1:17Jon Krohn:We watched the World Cup Final from a bar in downtown Toronto. and boy, did I learn some new curse words in Spanish. Yeah, you got to see me in football mode, which is a different me. More rowdy than when I'm talking about machine learning, for sure. For sure. And I learned your beer preferences. It seems you love wheat beers in particular. I do, do. Yes, yes. It was nice to, yeah, it was nice to meet you in person and get to watch some soccer. Yeah, it was a fun day. So last time you were on the show, It was your final month at what was, at the time, a company that was talked about a lot, Cohear.

1:56Jon Krohn:It felt like Cohear was in the conversation next to OpenAI and Anthropic all the time a couple of years ago. But it's interesting because I don't hear about them as much. And they're a Toronto-based company. Lots of people knew about them. But actually, last time you were on the show, it was your final month at Cohear two years ago. And now you've been fully independent for two years. So what does a week look like for you these days? What are you doing? Oh, yeah. I think it's similar to you. Things here, things there, a lot of things going on. So I think I remember asking you before leaving my job and you gave me an idea and I think it was pretty accurate.

2:32Sometimes super busy, sometimes not so busy. When I've been making courses, two big ones in particular with the Rackpack and with the guys from the Neural Maze. uh and when i'm doing those i'm i'm super busy creating content because i have to create content at a at a rate that is much higher than what i normally create but other than that yeah i've been i've been creating a lot of videos i i sometimes do consulting meetings sometimes teach live at other places uh but yeah basically i i give my uh wild mind a place to sometimes hyper focus on, on things and sometimes just be completely blank. So it's, it's, it's pretty random, but a lot of fun.

3:17Jon Krohn:That's nice. You know, I'm really looking forward to, uh, there's, I don't think there's any way that viewers could tell this, but I'm recording a bunch of episodes this week so that I can enjoy two weeks of vacation coming up in Europe. And it's been, I poof, I think there probably hasn't been a period where I get two weeks without recording a podcast episode since I started doing this six years ago. So that is going to be nice. I think it's important. But back to you, you mentioned the Ragpack there and you mentioned Neural Maze. Tell us about those teams because they're, you know, the Rat Pack, famous Sinatra.

3:54Jon Krohn:Who else is in the original Rat Pack? They took our name. There's an ongoing debate. A lawsuit with the Sinatra estate. The Sinatra. Yeah, no, I, yeah, that name, I was proud of that name because, well, I always wanted to work with other people. That's my main thing. I love working in teams. And I think we fill each other's gaps. You know, I have this sort of conceptual teaching with cartoons and stuff like that. But other people teach with code, more technical ways or more formulaic. And we try to cover like all those bases. So the first one is I started working with some of my best friends in this, like Jay Alomar and Josh Starmer.

4:38And then they knew Martin Grotendorsk and Chris McCormick. And so we started making this group of five. And we were just kind of like hanging out and we were just on Zoom calls thinking of if we could do something together. And then one day we said, okay, well, let's just like launch a course. and we needed a name and we thought about a lot of AI puns. So there were a lot of like, you know, AI Avengers, like this and that. And then the rag pack came up and I think we settled on that one, trying to be fancy. So that was a lot of fun. And we just, actually today we did our second cohort of the reinforcement learning for LLM's course.

5:18We started it live, totally live. And then I also was good friends with Miguel Otero from the Neural Maze. there too was with Antonio, sorry, it was in Yellow Terro. And we always wanted to make a course and their thing is agents, like they're super expert on agents. So we thought let's make one on agents with a pretty intense course, a six week course, we're always going to run it again. But that was a lot of fun. So yeah, I definitely enjoy working with others.

5:47Jon Krohn:Nice, yeah, me too. That's the only way I can work. Don't love doing things on my own. But yeah, so Ragpack, Neural Maze, you guys are churning out fantastic courses when you were talking about the rag pack you mentioned Jay Alomar whom I'd love to have on the show we should figure that out soon he hasn't been on yet but he is another rock star like you he's got some great books, super bestsellers from O 'Reilly the way he understands an LLM is on another level he can just open it and be like look inside it's been very rewarding for me too to learn from. And then Josh Starmer, of course. I mean, yeah, unbelievable.

6:28So yeah, he can make, he can really break down concepts and make them into like a,

6:35Jon Krohn:like a cartoon in a way that's amazing. So I love watching his content. And so he's been on the show. We did almost a two hour long episode, which I never do that anymore. I don't know. I don't know what I was thinking years ago. Sometimes I would just keep going and going and going like the Rogan show. but now I'm like, I'm too tired. So that's episode number 553 if people want to check it out. If the people don't know Josh Starmer, he's like, it's crazy. I mean, we're going to, your channel has also, you know, grown a ton. And we're going to talk about that in a second. But I think he's, of anyone I know, I think he has the biggest YouTube following.

7:11He's coming up on 2 million subscribers on YouTube. Yeah, definitely. You walk with that guy in the street and like people start recognizing him and stuff like that. You go, bam, double bam. Yeah, and all people do, bam.

7:24Jon Krohn:That's what Josh says in his videos for listeners who aren't aware. And yeah, so for lots of introductory statistical or machine learning topics, you can check out Josh's YouTube channel, but you can also check out Luis's channel for a lot of these same ideas. So the Serrano Academy, when we spoke in 2024, you had about 140 ,000 subscribers. Now you're over 200 ,000. You've got a full learning hub. You've got a sub stack. You've got some of the paid cohort courses that you were talking about. Did you have a plan for all of this when you left Cohere a couple of years ago? Or have you kind of just been following your whims and kind of like, oh, like Jay and Josh reach out and you're like, cool, like I'd love to do a course with you guys.

8:08Jon Krohn:Let's do it. Or is there a master plan? I never have a plan, a master plan. I just grade in the center life. Like I just go, okay, if I take one step in this direction, it might be, you know, when I left the job, I kind of, I called it to myself. I called it a sabbatical just to not have pressure. And I started making my own, like more, more YouTube videos and free content and stuff like that. And then basically whatever comes up, that seems like a good idea. I go for it. And then I test it. So very machine learning approach, very RL. but yeah no I I started the the the YouTube channel I I had a lot of problems I mean I had a problem which you may have which is a lot of people were like oh you should make a video on this and I'm like literally the next recommended video is is on that and I feel like there had no no order but I do make the videos with some order in my head like I'm like this follows this and so I put the learning hub bunch of free courses uh in the page the page actually for anybody who's listening, it's literally just serrano.academy.

9:10That's it. And then I organized the courses into what I think would be. And I also had some code labs and stuff that I did for the book or for blog posts. And so I put them all together, like post making all the material. I kind of Frankensteined them into a bunch of courses and the stuff does follow. Like after this, this follows, et cetera. And I, and when people want to learn a particular topic and definitely I point them to there first and I go, yeah, there's like, you may be able to find these videos in different places, but if you actually go there, you get the full thing for a particular topic and just worked well.

9:53I mean, I think when I show it to people, like it works more.

9:56Jon Krohn:It's a smart idea. When I navigated just now to serrano.academy, serrano.academy, and I see everything there. One of the things that I see on the page, in addition to the stuff we've already talked about, like your YouTube channel, we've got your Grokking ML book, which is worth talking about now. Yes. So Grokking ML, it's an iconic book, but the first edition came out five years ago now. Yeah. The second edition is just about to drop. So I'm expecting this episode to be released in early, mid-September, and it looks like the release date for the second edition of Grokking Machine Learning is coming out on September 29th in kind of the US and Canada.

10:36Jon Krohn:And I don't know all the dates around the world, but it's that kind of timeline. So tell us about the new version. I mean, a lot has changed in machine learning since then. We didn't have ChatGPT when the last book came out. So what has survived from the first edition and what did you have to rethink completely? Yeah, so when I started writing the book, there was the only thing that I had in my mind the most important thing is I wanted to do something as evergreen as possible because I didn't want that by the time the book comes out, like there's a new package that kind of changed everything and the stuff is obsolete.

11:13So I went for like the stuff that would be in a first year machine learning course regardless. So the linear regression, that's always going to be neural networks always going to be there. Now they have a lot more stuff, but ChatGPT is a neural network that works with backpropagation. So at some point you have to learn that stuff. And we focused a lot on like the concepts, like I wanted people to have a feel for what's happening inside an algorithm. And we went for a sort of the quintessential algorithms, mostly for supervised machine learning. So any tree-based algorithm, decision trees, boosting, gradient boosting, edge boost, anything regarding neural networks, logistic regression, linear regression.

12:01We had SVMs. Those are not as popular anymore, but I think they're beautiful. And then we had a fair amount of something I find very important, which is kind of how to evaluate the models, right? It's not just train the model, but is it doing well? Do I just add more layers and that's it? Or do I need to check other things? Is it performing well in my data set, outside of my data set metrics? So we did cover a lot of metrics, overfitting, underfitting, all the keywords for a machine learning interview, basically. And we had a whole chapter on using all these together in a data set. So I picked the most popular data set.

12:46The Titanic data set is pretty popular. There's a lot done on it, but we kind of go over it and say, you know, do all the possible things that what goes right, what goes wrong. How do you pick this algorithm, this model versus this other one? So I was very excited about the book and definitely I want it to be to last a while. So five years in machine learning is like dog years, right? It's like 20 years, eight years or something. But definitely every work was needed in several ways. The first thing is the package. I was using scikit-learn, but I used a package that I really liked that was made at Apple called Turi.

13:21I was working at Apple and Turi Create is just the most beautiful package because it's so nice and it does a lot for you. So I wrote half of the book on that. That has been deprecated. So I needed to change that. And I went for scikit-learn. And you know, surprisingly, like things worked actually really well on scikit-learn. And it actually didn't have to do very much. And some things actually were even more intuitive. So that was a big change that I needed to do. I changed it in GitHub, but I wanted to change it inside the book. Another thing that was enhanced was the boosting section. I mean, I did the gradient boosting and XGBoost, but I think ever since I wrote it, it's just more clear in my head.

14:04You know, like there was the explanation, but then, you know, a few things clicked after that. I was like, oh, I need that. And so I did that. But by far, the biggest change is that we needed to add generative learning. So the editors reached out and said, yeah, we'd love to have, obviously we're not going to make it a generative learning book, but if you have a rocking ML and you don't ever mention chat GPT, you're kind of like not covering everything. So we have a new chapter, which is mostly text generation, the architecture of the transformer, the way I see it, right? Like I like to see it visually, everything.

14:40I like to see some kind of visual analogy with little words and things like that. So the whole architecture of the transformer, and we did a bit on image generation. So stable diffusion process, less developed than the transformer parts. Transformers is a big one because it has so many nice kind of foundational things. Like it's a neural network, but it's got the attention mechanism. It's got all these bits some pieces, the Softmax has got the positional encoding. So we basically go through all of them. And yeah, I'm very excited. It's coming out soon. And I'll definitely send you a copy.

15:14Jon Krohn:Machine learning predicts and Gen.ai creates, but neither is built for complex, constrained decisions. That's where mathematical optimization comes in, giving you explainable, trustworthy decisions you can act on with confidence. Girobi is the fastest, most reliable solver organizations rely on for their high stakes decisions. Want to see it in action? Join the 2026 Girobi Decision Intelligence Summit, September 22nd and 23rd in Las Vegas, for training, expert insights, and Gen.AI-enabled accessibility. Discover why 70 % of the world's leading enterprises trust Girobi and start achieving optical outcomes yourself.

15:55Jon Krohn:Head to superdatascience.com slash Garobi for the conference details. That's superdatascience.com slash G-U-R-O-B-I. Fantastic. I would love a copy. I'm sure I'll see you in Ontario soon and I can get a signed copy. That's really what I would like, Luis. If it's not signed, I don't even want it. It'll be signed, definitely. Is it easy to, we actually haven't talked about the word grok. So we should maybe tell our listeners what grok means for those who aren't aware. And yeah, you could maybe give us an example of like, how do you grok a transformer? Yeah, I didn't know what groking was first. You know, they don't teach you that in ESL.

16:44So I didn't know. I didn't know what groking was.

16:47Jon Krohn:I don't think they teach you that in EFL either. Yeah. And so definitely when they said Grok, I knew there was a Groking series because there was a Groking Deep Learning actually by Andrew Trask. I worked with him. He's a great guy. And so when they reached out and said, yeah, let's write the Groking. I knew Groking was a special term. Like I knew Groking was a series of books. And then they explained to me, yeah, it's like understanding something really well. It's kind of like really breaking down a concept and just making it super clear. And so I thought, yeah, that's definitely the one thing that I like to do the most, which is taking a concept and really breaking it down.

17:30So it really made sense when we thought about writing the book. I pretty much do that for everything. I don't really understand stuff if I see it in complicated terms. Like if I see a formula, if I'm relying on the formula, I feel like I haven't understood it. So I need a little story, like a little picture or something. And I feel like so for my own self, like even before I explain stuff publicly, I need to understand my own work and my own courses and my research. And even a university or in high school, I needed a little, I needed to grok the whole thing to understand it. So yeah, it fit right there.

18:06Jon Krohn:When you were on the podcast two years ago, you said that some of the explanations that you come up, some of the visual, cartoony ways or the grokking ways you come up with, they take a decade to think of. So yeah, were there any particular concepts in this second edition of grokking machine learning that, you know, were there some concepts that finally clicked that, you know, five years ago, you weren't sure how to explain it, but now you were like, ah, I've got it. Yeah. I feel like a few of them, for example, like XGBoost, you know, XGBoost was something that I saw as a black box. Like there was this sort of, there's a similarity score that never really clicked to me.

18:48So in the book, I say, you know, there's something called a similarity score. If you have a group of a set of numbers that are very similar, then it's a high score and it's slow if they're not or something like that. And it never really clicked. So I had it there. And then you use that for XGBoost completely. And then later I realized that the similarity score is just a hidden difference between two things. And the difference between, they never tell you that, they just give you the number. It's the sum of things squared divided by their number. So when you take an average of a bunch of numbers, you take their sum and divide it by the number, by n, if there's n numbers.

19:25This is the sum squared divided by n. Well, it makes no sense. Why? When I realized that that's the difference between the variance and the new variance. So like when you have the variance of the original set, you have a lot of variance because the numbers are all over the place. And then you break it into two and then you subtract those variances. So if you break it correctly and you have sets that are more homogeneous, then your difference is high because you went from a high variance to two sets of low variance. And so when I saw that, then I was like, okay, well, I need to completely rewrite XGBoost because that similarity score now needs to be explained as the difference between the original variance and the variance obtained by breaking the set.

20:05So that one was a big one. And then a few other things that came up in the, especially in the new chapter, then I have, uh, you know, the attention mechanism that I like to see it as, as, as a kind of gravity thing of words flying around and then stable diffusion, uh, where I think of clip embeddings, you know, embeddings are something that makes more sense every time to me. Like I, I remember at the beginning, I, you always have that example of King, Queen, man, woman, and there was like the parallelogram. and to me, but eventually when you start seeing embeddings as like every dimension means something, even if you don't know what it means, I feel like that's the clearest way to see an embedding.

20:45It's like if I have a thousand 24 descriptors of a word and those are my thousand 24 dimensions of the embedding, then the parallelograms can happen for free, you know? So I like to see it as that. And obviously I like to see embeddings as words flying around in space and stuff. So yeah, definitely a lot of concepts in that book were sort of as I was writing it they they came up um more clear to me fantastic do you ever find yourself chatting with things like Claude or tools like that to help you understand concepts all the time now yeah now I go to deep seek for a while and deep seek is pretty good Claude Claude too like I just ask some stuff because Claude's like all these models start as started explaining to you the stuff in in the most formal way, right?

21:32So they go, well, this is that. It's like, well, thank you, Wikipedia. And then I just keep asking, asking, asking, and I go, okay, well, I didn't understand this part. Tell me more. And then eventually I start injecting my own examples. So I go, okay, I want you to see attention like this. Now explain this to me. So it definitely, I find that by itself, maybe one day, like I don't discard that possibility, but right now by themselves, they don't rock all the way there, but they definitely take you somewhere and then you help them cross a bridge and then they take you farther and then you help them jump over a wall and then they take you farther.

22:14So I feel like the combination of human machine is taking me farther. So I understand stuff faster because before I had to just bug my friends and ask them, my expert friends and ask them and ask them and ask them until their patience runs out. And I, and they're all very friendly and nice, but, but eventually people have a breaking point, but these models, you can, you can go on forever, you know, where it's going to happen is that you have to upgrade and to the higher tier and then keep asking that, you know, people don't have that property. Why do you think listeners should still buy books?

22:47Jon Krohn:Like, you know, I'm writing a book, Engineering AI Agents, right now. I have some reasons that I guess I could give, but why do you think people should buy Grokking Machine Learning, the second edition, instead of talking to an LLM? I'll definitely get your book. I'm excited about it. Yeah. No, I think humans still, you know, I wouldn't just take a book written by AI, or I wouldn't just rely on AI. I mean, I think the human insight is still important. Even if it's just a human that, like for example, when I look at what I input to a model, like I think I am a way better, bad understander than any model.

23:30And what I mean by that is I know where to get lost. You know, I know exactly what places I get lost as a student and I get more lost than the average student by far. So I think one of my biggest strengths is I know what are the parts that somebody in the room are going to get lost. And I think models overlook that because they just know everything, you know. So I think those kind of skills are what's important. The experience, the human has gone through difficulties at some point. They had an absolute disaster and got fired from a job because they made a huge mistake. The model doesn't have that.

24:12And maybe you can read it from someone, but there's a big chunk that the humans still inputs, especially in education. Education requires a great deal of empathy, a great deal of human connection, a lot of things that the model doesn't have. So at least yet. So I don't think, I still rely a lot on humans. And even though I prompt these models like crazy trying to understand stuff, I mean, I still rely on books. I still rely on YouTube channels and courses because I think education is a group endeavor. You know, it's a very human thing still. Even if it's the most technical subjects, there's a human connection that you feel is someone teaching you or someone learning from you.

Read the full transcript

25:04that I think is hard to replace for a machine.

25:07Jon Krohn:Yeah, I think another one of the great advantages of a book is that it's kind of, it's a convenient collection of everything like at the right time. And you can just know that in a well-written book, you can go from the beginning to the end and you're going to get this great arc. It's going to cover all the key topics you need to know. There was an editorial team. There were people who advised on what should be in the syllabus of the book in the first place. and so I think you typically end up with a really great product and at least not at the time of recording we can't yet completely have the book generated by an LLM maybe by the time this is published maybe we'll generate podcasts too for sure maybe this podcast right now is generated how do we know just the hands, five fingers, five fingers still us All right.

26:03Jon Krohn:So in addition to your new book coming out, something you've been writing, something you've been working on, something else that you published is actually a new paper. So you published last November with, I'm going to try not to butcher their names, Ricardo Di Sipio and Jairo Diaz Rodriguez. Yes. Very good. Yeah. I really put a lot of effort and thought into that. you guys wrote a paper together about the curved space-time of transformer architectures. Yes. And that is pretty mind-blowing. I think we're going to spend a bunch of time on that right now because we'll learn about transformer architectures in a way, but I think we're also going to learn about space-time and relativity and these kinds of concepts.

26:52Yes, this was definitely very exciting. Definitely very exciting to work on. And yeah, definitely for the physicists listening, we use the word relativity, but in a very loose way. It's basically a space-time curvature analogy, a weak analogy of what's happening inside a transformer. But I think it opens the door to what's happening underneath. And the fact that physics-related things start appearing, I found it mind-blowing. The story of that is that in order to understand attention, it never clicked to me with the People say it's like a search table with a query and a key never made sense to me.

27:32Then they said, oh, the words pay attention to other words never made sense to me. Then they gave me the formula made even less sense. It's a soft max of KQ divided by square root of DK times V. Nothing for me. So I started looking at videos and looking at other things and looking at just writing and like watching videos and somewhere, somewhere in the process, somebody, which I bless his soul. I don't remember which channel was this. I think it was a kind of an underrated channel that didn't have a subscribe. but this person is just kind of like made a, like this is a beautiful description where at some point they did a linear combination of the words and they said, this word becomes more like that one.

28:21And the linear combination added to one, the coefficients added to one because there's a soft max. So, you know, you turn your word apple into 70 % of apple and 30 % of orange. If you said the two words consecutively, say orange, apple, you know what I mean? so the word becomes a percentage of itself and and the rest of percentage of another word and to me that's moving in a line right like if you have if I have two points and then I take a percentage of one of the position of one point and the other I'm moving in the line between them and so I thought maybe words are moving in a line and I started rewriting all the equations as in like words are in a position in space because embeddings words are in a position in space and then attention just moves them in a line toward each other.

29:07And I immediately thought, oh my God, that's gravity or magnetism, right? Words pull each other. And I thought that makes a lot of sense because if I'm saying the quintessential example is the river bank. Bank is a bank in the financial sector of the embedding around stocks and bonds. And then you say river bank and the bank just becomes a nature thing. So the word river just pulled it towards itself into the nature region of the embedding, because in the nature region of the embedding lives a river and tree and stream and sea and all that stuff. And so it just pulled it. It infused itself with it, right?

29:48It infused some properties of it. For example, the nature property. It didn't infuse all the properties. There are some that don't, but the nature property, it moved in that. So the moving actually works in different directions. It's not towards it because of the value matrix. But anyway, the fact is, words, I started calling it word gravity. And as I said, any physics words that I say is a very loose analogy. But I started calling it word gravity, word gravity, word gravity. And then one day I gave a talk in the Toronto Machine Learning Summit. And Ricardo Di Sipio, physicist, I also, by the way, I hope I'm pronouncing it well because I don't know Italian.

30:27And Ricardo is a physicist who was sitting in the first row. And then he came to me and said, hey, what you have is actually, you know, when Newton was talking about gravitation and then Einstein came, which is a force between objects, he died and he had no idea why this force happened. And then Einstein came hundreds of years later and he said, it was not a force. The masses are bending the space. Like, you know, the reason you fall towards the earth or the earth falls towards the sun is not because the sun exerts a force, it's because the sun bends space. And all of a sudden, the line in which the earth should be flying in a straight line, it's curved because of the sun and it happens to be curved around the sun.

31:07And that's why we're there, you know? So he said, I think it's the same concept. Like you're talking about word gravity as like a force between words. I think it's more like it's probably a space-time curvature thing. Like it probably words are bending space in a way that they just pull towards each other. Right. And so we started working on that and he worked out the math a lot. He knows the physics a lot more than me. So he actually worked out the geodesics and very much like the matrices that appear in the geodesics appear in are the key query and value matrices. And then we started working with another friend, a higher of the as a professor at York University in data science and statistics to run a lot of experiments.

31:50So these two guys are wonderful. They actually know a lot more about that than me and worked out both the physics and a bunch of experiments that really study this analogy. And so we're very excited actually of this. I mean, it provides an analogy. I think it's more of a visual work. We meant it as a visual work. We've gotten notices of like labs that are working with it for something else. So I'd love to see applications of it. But as of right now, we thought of it as like a fun analogy, like physics analogy of what's happening inside chat GPT or inside the brain of these models.

32:28Jon Krohn:For all you listeners who want to level up your AI career through hands-on learning, ODSC AI West, October 27th to 29th in San Francisco is the place to be. ODSC AI West is my favorite conference and what sets it apart is it's all about doing. You'll gain practical skills by working directly with the latest AI tools and frameworks in immersive hands-on workshops and tutorials led by experts who are actually building and shipping AI. I myself will even be doing a keynote at ODSC AI West this year on how individuals and organizations can thrive in the agentic era. The full program covers where AI is moving now, including AI engineering, AI-powered software development, physical AI, robotics, and data science.

33:06Jon Krohn:Beyond the training, ODSC AI West brings the AI community together with networking events, meetups, the AI expo, and more, giving you the chance to learn, practice, and connect all in one place. Super Data Science listeners can use the code SUPER at checkout on odsc.ai for an additional 15 % off your pass. See you there. Odsc AI West, October 27th to 29th in San Francisco. In that analogy, what plays the role of gravity in the same way that you described the sun bending space so that the earth comes toward it. Yeah. What's gravity in a transformer? Yeah, that's a great question. So it's not the same gravity as in mass one, mass two times R squared in particular, because mass doesn't make sense.

33:55Like if you have a big mass, it attracts every word, right? But here is not the case. For example, river really attracts bank, but it doesn't attract other words, right? Like, you know, if I say river tree, it's still the same tree uh so it's not a matter of how close you are because tree is close to river doesn't get pulled as much as bank which is farther away so it's not it's not that that words independently have an information of how much you pull and they pull everything based on the distance and the mass of the air word it's for every pair there's an actual quantity of pool and it's not symmetric because for example, bank doesn't pull river.

34:37If I say the river and the word bank comes before, it's still a river. You know what I mean? Unless they name something bank, a river of something, then yes. But as of right now, the pool is not symmetric. A word can pull another one, but the other one can't pull. So it's weirder than gravity in many ways. And the formula is not GM1M2 divided by R squared. It's soft max of KQ blah, blah. And it's based on the dot product of the words, but when you multiply them by K and Q. So K and Q, what they do, the key and the query matrices, is they extract all the features of a word that would influence other words for key.

35:18And the query extracts all the features of a word that get influenced by other words. So in river bank, then the key of river extracts the nature part and says, you know what? If a word comes with some nature, the river will pull it. And the query does that for bank. Yeah. And then you multiply those. So it's very dependent on the key and the query matrices that basically transform your space into something else where the dot product says more than just similarity. So it's a similarity. It's a two directional similarity. And then the V matrix comes out and says, you know what? you want to move in this direction, but I actually want to move you in a different direction because some properties are movable and some are not.

36:08I can add nature to the word bank and turn it into a more nature word, but there's a bunch of properties that are not. For example, bank is a noun. I can't change that no matter how much I pull it, right? So for example, in the book, I have the example canoeing bank. So canoeing is a verb and it pulls bank, but it's never going to bring it to become a verb. Right. It's going to bring it towards nature. So out of the all the descriptors of a word, if there are 1024 numbers in your embedding, a bunch of them are movable, a bunch of them are not. So V does that. And interestingly enough, in the analog for relativity, K and Q make a lot of sense, but the V just came out of nowhere.

36:50So it even has extra stuff, but I'm getting ahead of myself.

36:54Jon Krohn:Luis, you're talking about the bank bending towards the river. It reminds me of an analogy that happens in the paper that I found fascinating. And this isn't really that related to data science or AI. It's, I guess, more an astronomical thing, but maybe it'll help people understand the same kind of transformer idea as it was helpful for you. I'd love you to tell us more about Eddington's 1919 eclipse experiment. Yeah, that's one of my favorite experiments, because what happened is you can't really see space-time curvature, right? Like we feel like we live in a giant 3D cube, but we don't know it's bent in different ways.

37:36You know, you can see bending in 2D because we can sort of see the earth being round, but you can't really see the universe being curved if all we have is 3d vision right so um so what eddington did uh was he he thought about a star that was sort of behind the sun and in a certain eclipse and and you would see it not behind but and in another place like you would see it sort of in a place where it's not and the reason is because the light of the star came in and bent thanks to the sun and then got to our eyes, but we could see it right behind. We could see it like it was above the sun. And so we have an analog of that experiment, which is we have the word, and it's always with bank and river, but basically bank, it takes the place of the star and river is the place of the sun.

38:31A fun story. I actually wanted to find the right example where the word would be sun and the word light would be bent, you know, something like maybe there was so much sun that I had to pick up the suitcase and it was light, something like that. Anyway, I couldn't work it out. So if it ever happens, I'll have an experiment with the word sun and star or light or something. But we had to do bank and river. And basically we just looked at space is the embedding, right? And the analog for time is the layers in the neural network. So the first layer is the beginning of time and each layer is a snapshot in time.

39:08And obviously, then that's a problem because time is discrete in our analog. But we can't really mimic the continuity of time. We can only take snapshots. That's a limitation. But actually, we had an... What's the opposite of limitation? We had actually something that was easier, which is that you could see the entire trajectory, whereas Eddington could only see the end of it, right? Like, where do I see the star? He doesn't know where the light went because he can't look at that. We could actually take the position of river and bank in every single layer. So there's a pro and a con. But basically, yeah, I mean, like the word bank flies through the embeddings, the embedding in every snapshot, in every layer.

39:52And it's like bank is here, then it's there, then it's there, then it's there. And eventually it appears in another place at the final layer. and when we added the word river in the sentence then it would basically go around river and the other in a different path so we would see the path of the path of bank is this but when you add river then the path of river is this so we had no choice then to say well river must have must have been the space around it for it to take a different route it's interesting to be able to

40:23Jon Krohn:see that happening through all the layers of the neural network in that transformer. That is very cool. One of the counterintuitive findings of doing that kind of analysis was that I think you found that function words like to, like T-O, which are like, they don't seem like they have much semantic meaning, but they actually showed some of the sharpest curvature in your experiments as opposed to semantically heavy words. What do you make of that. Yeah, that was a really interesting experiment because some words are pretty flat. Like, for example, the word the doesn't get influenced very much. The is the, you know.

41:01On the other hand, two can be influenced a lot because it can mean different things based on what's before, right? Like two can be to something or like, like there's a lot of meanings of it. So in some, in some layers, it really showed a lot of curvature because it was, it was easily influenceable. So that's something interesting we found. We also found that like each layer takes care of certain things. So like sometimes words had a big curvature around them in some layer and in another one, not so much. And that just lets us think that, you know, different layers are taking care of semantics or meaning or things like that.

41:40Jon Krohn:For sure. That's something that we've known about deep neural networks forever, where, yeah, you have different layers that come to represent different kinds of meaning. And so, yeah, in some layers, two doesn't need to change, but in others, it needs to change wildly in order to be able to make sense of some sentence or document or what have you. Very cool. Well, thanks for talking us through that paper, The Curved Spacetime of Transformer Architectures, Decipio, Diaz-Rodriguez, and Serrano. I'll have a link to that in the show notes for sure. let's move on to agents agents are i don't think we've talked about them at all yet in this episode which is pretty amazing because it's hard to go very far without talking about them you actually just wrapped a course called it's your first paid cohort course you can tell us what that means what the cohort course is all about but it's called grokking agents in production and it's with some more names that I'm going to struggle with.

42:46Jon Krohn:I am giving you some. Miguel Otero Pedrido and Antonio Zauros. Zaraoos. Although if I was in Spain, I wouldn't say Zaraoos. I would say Zaraoos. Zaraoos Moreno. There you go. I was saying it wrong, actually. I was saying Miguel Otero Pedrido, and it's Pedrido. Yeah. I didn't even hear a difference between those two. Pedrido versus Pedrido. Okay. So, yeah, even I mess up the Spanish. But yes, yes, I'm very, this course was very fun. It was a six-week course. It was dense. It was really more of an eight-week course. I'm rocking with Miguel Otero and Antonio Zarauz from, they make the Neural Maze, which, by the way, this is an amazing channel.

43:31They have so much, the Substack subscribers and very, like, they make so much content. If anyone wants to learn agents, definitely the Neural Maze, you should check it out. And so I worked with them on this course and I was doing sort of the, it was the same breakdown as I normally do. They have the technical side and the deployment side. And I was doing this sort of conceptual explanations on agents. And it was a lot of fun. I learned a lot about agents. For example, I had a conversation with Miguel where actually we're recording a video because it was such an interesting thing. I used to think that RAG was an agent, you know?

44:16So I would say, you know, RAG is an agent. Whenever you have an LLM doing things, that's an agent. And then I quickly found out, he corrected me and said, no, RAG is not an agent because it's an LLM workflow. So one thing is an LLM that does stuff. And another thing is an agent. An LLM workflow is when you tell the LLM what to do. Like you go, okay, you do this. And if this happens, you do this and then you nod, you close and that's it. And agent is when the LLM makes a decision, right? So when the LLM says, oh, I need to look for more information. I go look for information. Now I feel like I can answer this question, but I'm missing something.

44:57Like when the LLM is in charge, it's an agent. So that's one of the, that's one of the interesting things I learned. And yeah, other than that, it was a lot of new stuff and basically a zero to 100 on agents with deployment.

45:15Jon Krohn:Yeah, lots of the unglamorous parts, containerization, IM, observability, CICD. And it seems like the tagline for the course is demo agents are trivially easy to create, but production agents get hired. Yeah, and that's what the course does. Like it's end to end, like the project was to build a PDF reader that can read everything on your PDF using different techniques, chunking OCR, all that kind of stuff. And I combined them and then answer questions on it. And every, like every module had its own sort of deployment part. We use Google Cloud and then Cohere for the LLM. So it's pretty, yeah, it's pretty interesting.

46:02I learned a lot.

46:02Jon Krohn:In 2024, when you were on this episode previously, you predicted agents and tool use as the next big wave right after multimodality. And you got both of those right. Are there things that happened faster over the past two years than you expected? Or do you think there's parts of agents, multimodality that are still overhyped? I'm so glad my prediction went right. Thank you. You're so kind that if I had done something wrong, you would have erased it, right? But I'm glad it worked out. Yeah, no, I mean, definitely that was the vibe at Cohere when I was leaving. The vibe was tool use, tool use. It wasn't even call agents.

46:37It was tool usage. Like you got the thing to talk. You now have to make it do things. Like it's the next thing because it's the next thing with everything, right? Like with the internet, first it was like text coming in and coming out. And then it started doing things like a bank transaction or something. And it looked weird at a time. Like it looks weird to get the model to do things for you. But that was definitely the trend. No, I mean, I think it went, I didn't predict any speed for it. I knew it was going to be fast that we're going to start getting this to do things. Something that I found interesting is that in my head, the LLM already talks and now you make it do things.

47:24Go open your calendar, go write an email. What I actually didn't occur to me and I found it really interesting. I mean, it was there, but it just clicked, is that you can use agents to make the LLM do better things, right? Like you can have an agent, if you want to write a book, you don't just tell the LLM to write a book. You have a multi-agent system where one writes the syllabus and the other one makes the chapters and the other one checks it and the other one checks the grammar and there's a loop and there's like a manager that decides what part you need more. And so definitely the fact that you can make the LLM talk better by making it do things.

48:03That was a nice addition.

48:07Jon Krohn:regular listeners will already be aware that i'm obsessed with anthropic's fable 5 model and it has taken over my working life i'm writing a technical book that includes latex files mathematical notation python code examples and fable 5 and cloud code handles requests i make across whole chapters with accompanying jupiter notebooks end-to-end work that a few short months ago would have been dozens of separate requests with way more manual fiddling required with fable 5 it just works, essentially like magic, first time. Claude is the AI for problem solvers. It's the collaborator that understands your entire workflow and thinks with you, not for you.

48:44Jon Krohn:Whether you're debugging code at midnight, building a financial model, or strategizing your next business move, Claude extends your thinking to tackle the problems that matter. For problems worth solving, get started with Claude at claude.ai slash superdata. That's claude.ai slash superdata. and check out Claude Pro, which includes access to all of the features mentioned in today's episode. Claude.ai slash superdata. Yeah, worth checking out for folks who want to get into getting agents into production, get the grokking agents in production course. Definitely. You can, of course, an easy place.

49:19Jon Krohn:Of course, I'll have a link to that in the show notes, but again, you can find all of this at serrano.academy in one convenient location. Yeah, definitely check it out. We're going to run another cohort. We had a lot of students. and it was a lot of work, a lot of interaction, a lot of office hours, and then we're definitely going to run it again later in the year. So keep an eye if you want to learn agents. Fantastic. Thank you for all the great courses, Luis. Thank you. Two of your six agent course modules were on evaluation. So things like golden test sets, trajectory scoring, LLM as judge, evaluated deployments.

49:55Jon Krohn:Why is agent evaluation so much harder than say, rag evaluation? Yeah. I mean, I feel like ML evaluation gets harder and harder, right? 10 years ago, it was like evaluating a multiple choice test when it was predictive machine learning, right? Then LLMs came in and it's like evaluating an essay much harder because you have to evaluate conceptual writing. And then one step even harder than that is evaluating doing stuff, right? Like how do you evaluate someone who's doing a job? Like that's even harder than grading an essay. So definitely agent evaluation is one step harder than LLM evaluation.

50:34And the reason is because there's so many things to take care of, right? Like if you only look at the final result, that has to be good, but that doesn't tell the whole story. And then you have to see every single step along the way. Now, every single step is, did you do this step or did you do the step right? Huge difference, but you have to check if the step was done. And sometimes it's hard because let's say something simple in an LLM doing rag, you ask it something and it just answers out of memory and it answers correctly. So you just go good, but no, it didn't, it didn't use rag. And if I'm going to ask it something like, what's the weather today?

51:15It has to use rag, you know, it has to use the tool. So I can't just rely on, did it, did it do the thing right? Right. So basically the modules on that were something as simple as check if all the tools were called, then check if the result was good, then check every tool separately using different ways, you know, root score, just basically comparing words, using an LLM as a judge is also quite interesting. And those two are basically need to be done because there are cases where, yeah, you just check if the right words appear in the answer. But other times you need to check if you need an LLM to check if the answer was, the question was answered.

52:03So there's a lot of little moving pieces that you somehow have to combine into a big evaluation step. But we made a big deal on that because that is less common to see. Like a lot of the work and agencies do it and deploy it. And we're like, no, you have to make sure it works well. I mean, you need to be very responsible on this thing. So we add a lot of material on evaluation, like I do with everything. I think it's very important.

52:32Jon Krohn:Yeah, super important. This is something that we run into a lot with my consulting firm, YCaret. We have enterprise clients, we have government clients, and it is essential for them that if agents are deployed, they can be trusted. And it's these kinds of evals that allow that to happen. So thank you for creating a course like this. In addition to the agents course, you've also built a whole free course and video series on reinforcement learning for LLM. So you cover RLHF, PPO, DPO, GRPO. If we had all the time in the world, I would have you go in and explain all those different things. But so instead, I'm just going to skip to GRPO because that seems to be a reinforcement learning approach that has really mattered a lot, particularly with the DeepSeek explosion from last year.

53:23Yeah. I'm very interested in GRPO and I love it. And yeah, I worked on that. And the video series on GRPO was a lot of fun for me because it was just learning stuff from scratch. And we talk about that in the course a lot, in the course I have with the Rackpack. and it's basically, here's the interesting thing of GRPO versus PPO, right? Basically, GRPO is for reasoning models. It's for making the model not just talk, but do math or write code or do logic, which are much harder. But when you look at GRPO, it's not an actor-critic model. It's really just an actor. Whereas PPO is actor-critic model.

54:08So I'm going to sort of over generalized, but in my head, PPO is more meant to make the model talk well. And GRPO is make the model do math well. And there's one difficulty and one thing that is easier. One thing is that's harder, which makes GRPO what it is. If you want to make the model talk, that's easier than making the model doing math. I'll tell you why in a minute. But if you want to evaluate a model talking, that's harder than evaluating a model doing math. For the simple reason that if you write an essay, that's hard to grade. But if you do a math test whose answers are numbers, that's easy to grade.

54:52Right? So we kind of have a two by two square where we say, okay, talking is easy. Doing math is hard. But evaluating talking is hard and evaluating math is easy. So PPO has an okay actor that talks and a really good evaluator that evaluates. And GRPO has a really strong talker that does math and an evaluator that is not as strong. Obviously, it's strong. But where does it come that the evaluator in GRPO is not as sophisticated as the evaluator in PPO? The evaluating PPO is a model that gets trained. So you're training the talker and the evaluator at the same time. The model talks, the other one evaluates, the one talks, the one evaluates, and they both get better.

55:44On GRPO, you're only training the talker. Your evaluator is fixed. And the evaluator is simple. It's a formula that grades your responses and it doesn't get any better. It just gets trained. It uses LLMs. It uses other stuff. It uses compilers for compiling the code. It uses math engines to check the math. It uses a lot of stuff, but it's not a model that gets trained. So all you have is the actor and the evaluator is sort of this big formula they have in the paper where you have, it's basically a big loss formula. But that's what I find interesting of your PO because you would have thought in order to make the model do math, I need to make everything better, but it's not.

56:29I need to make the model stronger, but the evaluator actually, there's a lot of stuff I can use in the fact that the model is reasoning instead of talking that I can take advantage on. And at the end of the day, I did say another thing that I want to clarify. I said talking is easy and doing math is hard. Obviously, talking is not easy for a model. but what I mean is that probabilistically talking is easier for a model because the model doesn't go for the right answer. The model throws a dart in the vicinity of the answers because it's a probabilistic model and it gets some answer that is close to what you meant to say.

57:11If the perfect answer is here and you throw a dart, you get something close. But in text, that's no problem because if I live in text space and I pick a sentence that's close enough, I'm likely to say the right thing worded differently. You try that in math, that don't work. Because if the answer is two plus two equals four, and I move one millimeter, I get two plus two equals 4.1. And that's not correct. So math is a hostile space where a right answer is surrounded by wrong answers. Code is the same. A right line of code is surrounded by wrong answers where you change a colon into a semicolon.

57:44Whereas text, for the most part, if I land very close to what I wanted to say, I'm okay, right? So in some way, then we have this two by two matrix where we have top left, we have talking is easy. Top right, we have evaluating talking is hard. Bottom left, we have doing math is hard. And bottom right, we have evaluating math is easy. And the top row is PPO and the bottom row is GRPO, basically.

58:11Jon Krohn:That was a really fantastic explanation of GRPO, which I don't think I even said is called group relative policy optimization. That's what GRPO stands for. And I'll have a link to materials on that, of course, in the show notes as well. But fantastic explanation there. And particularly for why it's important, that mathematical reasoning. And it seems like that's a key part of this. Like there was a, GRPO was actually introduced in a paper called DeepSeek Math. Yes, yes, yes, yes, exactly. pushing the limits of mathematical reasoning in open language models. So aligns very nicely with what you're saying there.

58:49Jon Krohn:Luis, you used to work on quantum computers at Zapata Computing. What's been happening in the world of quantum as it relates to ML or AI? Have there been any big innovations in recent years? I think there's been hardware things happening. The computers have been getting bigger. There was this topological qubit thing happening. So I think in terms of that, yes, I don't know. And maybe it's just that I haven't been sort of that up to date, but I don't know in terms of algorithms, at least in machine learning or something. But yeah, I've been out of the world a bit, but definitely hardware advantages have been happening.

59:35And then hopefully we'll have a quantum computer sometime. That I don't know if I can predict, but I'll just, you know what? I'll say the two things and we'll have both recordings. And you're so nice that in two years, you'll put in the podcast the one that was right. Like I'll say that quantum computers are never going to work. And then I'll say we're a year away from our quantum computer and we'll make sure the right video appears. How's that?

59:59Jon Krohn:Nice. Sounds great, Luis. Last technical question for you here. It's actually not that technical, but it's related to a technical topic. So the International Math Olympiad, You're very familiar with it because you won the Broadens Medal in the International Math Olympiad years ago. And in 2025, AI models hit a gold medal standard for the IMO, for the International Math Olympiad. And it seems like in 2026, they did even better. We haven't had full reports out yet. What goes through your mind when you see a machine being able to act at that level on math. Is it just exciting for you? I find it very exciting.

1:00:44I mean, I think there's always a tiny bit of nostalgia saying, oh, I worked so hard to get this and like a machine can do it. But I think that's 0.1 % because I'm 99.9 % very excited to see models being able to solve these problems. These problems at the end, I mean, require a lot of creativity, but a lot of it is using a lot of tools that exist, which is not that different from a game of Go or a game of chess, where it requires tons of creativity. But the fact that you have a finite set of moves, and that a computer can do, can search in a space so humongous, then definitely, I mean, Olympiad mathematics is a game like that.

1:01:29And I actually was really hoping that it would be, that a model would perfect it. Um, I think it doesn't take away from, from the competition in the same way that we still watch people run, even though cars are faster.

1:01:46Jon Krohn:And even in, in, in intellectual pursuits like chess that you already mentioned. Yeah. We still watch chess, you know, goal is still interesting, even though a computer, uh, can be there. So, you know, computers can, can, if we can watch the Olympics with, with machines and it'll be less fun, but they will win. Uh, we could watch two models playing chess, but we still watch two humans playing chess, you know, because I think there's still the fact that we do it with our, with our limited computer in our head and, and we use different, you know, creativity and things like that. I think it's exciting.

1:02:17So I still, I still follow IMO results, very, very excited. And I, and I think it's not going to go away. And I think that's the same thought that I, I talk with a lot of, a lot of IMO people and I see nothing but excitement. I mean, I think, I think it's, it's great to see models that can, that can build on. It's a good benchmark. And I mean, I, I, it's already like research mathematics is already in, in reach. And I think I'm, I'm excited about it because at the end of the day, the, the, the human is going to pass to a higher level of saying, okay, what do I want to, the direction of this field to go?

1:02:55What do I want to, the big problems I want to solve. And for things like, like solving the problems, I mean, mathematics at the end of the day is, is a game we have a finite number of axioms and from the axioms you build lemmas or you build small you know from the lemmas you build theorems and it's it's no different than walking in a very high dimensional world trying to find a needle in a haystack so models models can do that very well and they can they can i mean go has more positions in the board than atoms in the known universe and it's able to navigate that space very well and find the tiniest needle in the biggest haystack.

1:03:38So I'm excited. I'm excited of seeing it solve mathematics. And, you know, people would work, like Terry Tau, for example, is the best mathematician. He's done a lot of work on AI and mathematics. And there's a lot of other famous mathematicians doing that work. Timothy Gowers and other fields of my atlis is also working on that. So yeah, there's definitely interest there.

1:04:02Jon Krohn:Hey, hey, this is your host, John Krohn. In addition to hosting this podcast, did you know that I run an AI consulting firm called Y-Carrot? Yes, that's the letter Y and the deliciously crunchy veggie. At Y-Carrot, our team pairs decades of ML and software engineering experience with recognized expertise in the latest AI techniques, such as all the key generative and agentic approaches. From problem scoping and proof of concept through to high volume production deployments, we've got you covered at every stage of the AI project lifecycle. To learn more, head to ycarat.com. From there, you can click partner with us to give us some context on how we can help.

1:04:40Jon Krohn:We're looking forward to hearing from you. Again, that's ycarat, Y-C-A-R-R-O-T.com.

1:04:51Jon Krohn:Well, fascinating conversation today, Luis. Across the board, reinforcement learning, agents, curved space-time of Transformers, and of course, Grokking Machine Learning, the second edition, which will be available to listeners very soon. In addition to your own books, do you have any book recommendations for us? Yeah, I think I was checking that I gave you someone's last time. So I wanted to not repeat myself, but there's one that actually someone has been in your, on your channel. There's on your podcast, Kathy O 'Neill, Weapons of Math Destruction, Math, like M-A-T-H. That's a really good one.

1:05:32So that, that, that one I recommend. What else is good? I mean, good friends of mine, JLMR and Martin Grotendorf are getting one out on agents pretty soon. So that's, that's one to, to definitely check out.

1:05:45Jon Krohn:It's pretty funny that, am I correct in thinking this when you were giving the Spanish, like the way that someone would say something in Spanish, like in, in, in Spanish from Spain, would they say weapons of math destruction? They would say weapons of math destruction. I think so. Yeah. Maybe that's definitely, oh, that's, that's the perfect way to teach the Spanish set. If you say math versus math. That's how you would say the Z and the C in Spain. Whereas in Latin America, we just say S, Z, and C in the same ways. Yeah, yeah, yeah. Really cool. That's funny. Good to have a laugh with you again.

1:06:27Jon Krohn:If people want to follow you after this show, which I encourage them to do, obviously we know they can go to serrano.academy. Is there anywhere else they should follow you on social media or something like that? Yeah. The page Serrano.academy, the channel Serrano Academy, the YouTube channel is where I put all the stuff I know goes there. LinkedIn. I'm also quite active on LinkedIn. So I'm just Luis Serrano on LinkedIn. Those are the main places. Perfect. We'll have links to all of that in the show notes. Luis, I always enjoy spending time with you. You are a treat. So friendly, so fun, so intelligent, and you tell it like it is.

1:07:05Jon Krohn:So I hope we'll have the honor of having you back on the show again sometime soon. Thank you so much for taking the time out of all the work that you're doing these days to join us. Thank you so much, John. I'm a great fan of your podcast and you're a wonderful person to talk to now. And I know that in person and in virtually. So I definitely look forward to our next conversations. So thank you so much for having me. So it's a treat for me to be here. What a fun episode with Luis Serrano. In it, he detailed his new paper on the curved space-time of transformers, how his team recreated Eddington's famous 1919 eclipse experiment inside a transformer, tracing the word bank through every layer of the network and watching its path curve around the word river.

1:07:47Jon Krohn:He talked about the crucial distinction between an LLM workflow, where you tell the model exactly what to do, and an agent, where the model itself decides what it needs to do next. He talked about why agent evaluation is so much harder than evaluating a chatbot, his elegant explanation of why GRPO powers reasoning models for a probabilistic model, talking is easy but hard to evaluate, while math is hard but easy to evaluate. And he talked about why, as an international math Olympiad medalist himself, he's thrilled, rather than threatened by AI hitting gold medal standard on the IMO, comparing it to how we still watch humans race even though cars are faster.

1:08:25Jon Krohn:As always, you can get all the show notes, including the transcript of this episode, the video recording, any materials mentioned on the show, the URLs for, or the URLs for Luis's social media profiles, as well as my own at superdatascience.com slash 1025. Thanks to everyone on the Super Data Science podcast team for another great episode. We've got our podcast manager, Sonja Brejevic, media editor, Mario Pombo, our partnerships manager, Natalie Zajski, researcher, Serge Macisse, and our founder, Kirill Aromenko. Thanks to all of them for producing such a fascinating and fun episode for us today.

1:09:04Jon Krohn:For enabling that super team to create this free podcast for you, we are deeply grateful to our sponsors. You can support this show by checking out our sponsors links, which you can find in the show notes. And if you'd ever like to sponsor an episode yourself, you can get the details on how by making your way to johnkrone.com slash podcast. Otherwise, please help us out by sharing this episode with other nerds that would enjoy this deep dive that Luis took us on. Review this podcast on your favorite podcasting platform or the YouTube video that you're watching. If you write in an Apple podcast review, that is particularly helpful to us.

1:09:45Jon Krohn:And I will eventually actually read it on the show if you do that. Subscribe, obviously, if you're not already a subscriber. But most importantly, I hope you'll just keep on tuning in. I'm so grateful to have you listening and I hope I can continue to make episodes you love for years and years to come. Until next time, keep on rocking it out there and I'm looking forward to enjoying another round of the Super Data Science Podcast with you very soon.

1:10:22you

From the publisher

In Episode #1025, Dr. Luis Serrano (Founder of Serrano Academy) joins Jon Krohn to explain the paper he co-authored on the curved spacetime of transformer architectures, in which attention stops being a lookup table and becomes something closer to gravity: words bend the space around them, and the embedding of "bank" visibly curves toward "river" as it travels through the layers of the network. In this episode, he recreates Eddington’s 1919 eclipse experiment inside a transformer, draws the line between an LLM workflow and an actual agent, explains why agent evaluation is a step harder than evaluating an essay, and gives the cleanest account of GRPO you will hear.

Additional materials: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://www.superdatascience.com/1025⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠

Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information.

In this episode you will learn:

(00:10:53) What changed, and what survived, between the two editions of Grokking Machine Learning

(00:27:48) Word gravity: how attention pulls "bank" toward "river"

(00:42:12) Why RAG is an LLM workflow rather than an agent

(00:51:23) The two-by-two that explains why GRPO powers reasoning models

More from Super Data Science: ML & AI Podcast with Jon Krohn

All 130 episodes
1025: Word Gravity: How Transformers Bend Space, with Dr. Luis SerranoSuper Data Science: ML & AI Podcast with Jon Krohn · 1 h 11 min
Listen in VO