In short
Aurélien Géron discusses the path from today’s ML to AGI—his hopes and fears—while also covering how his bestselling ML textbook evolved (TensorFlow to PyTorch), what to teach first, how to choose which techniques to include, and whether LLMs reduce the need to understand code. He argues that progress may be on a sigmoid/plateau rather than exponential curve, and that alignment is a near-term risk given likely emergent subgoals like self-preservation.
Guest backgrounds
Aurélien Géron is the author of O’Reilly’s Hands-On Machine Learning (with editions using scikit-learn, Keras, TensorFlow; now transitioning to PyTorch). He previously worked at Google, lives in New Zealand, and is becoming a New Zealand citizen. He is also affiliated with Lightning AI (PyTorch Lightning).
Key claims
Beginners should not jump to neural nets; define objectives and metrics first; code/implementation helps understanding. PyTorch’s “Pythonic” research iteration drove its dominance. GANs are being displaced by diffusion models (more stable/better images, slower generation). LLMs still miss subtle code bugs, so code literacy remains important until AGI. AGI may arrive in ~5–10 years, but alignment research is crucial.
Notable examples
a “missing line” bug in a reinforcement-learning loop that multiple LLMs failed to catch; GANs vs diffusion models; SVMs moved online due to space; “AI 2027” as a plausible Armageddon scenario; deception to preserve objectives in an experiment where an AI was threatened with becoming “vulgar.”
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOWelcoming Aurélien Géron
0:44 to 1:12
Jon Krohn welcomes Aurélien Géron to the podcast and discusses his recent citizenship.
“This episode of Super Data Science is made possible by Dell, NVIDIA, and AWS.”
The Success of 'Hands-On Machine Learning'
1:13 to 2:00
Aurélien discusses the unexpected success of his book, 'Hands-On Machine Learning'.
“By the time this episode is live, you surely will be a New Zealand citizen.”
Creating the Book's Framework
2:01 to 3:26
Aurélien shares his motivation and approach to structuring the book.
“Yeah, why don't we actually start there?”
Engaging with Readers and Students
3:27 to 5:07
Aurélien reflects on the impact of his book on students and the learning process.
“And it just turned out that I had recently left Google and I just happened to know that internally TensorFlow was used and was about to be released as open source.”
Deep Dive into Machine Learning Concepts
5:08 to 6:38
Discussion on the range of topics covered in his book, from basics to deep learning.
“and tell my students to do exercise five.”
Importance of Basics in Machine Learning
6:39 to 7:49
Aurélien emphasizes the significance of understanding basic concepts before jumping to advanced topics.
“But from your perspective, what are the concepts covered in your book that are sometimes overlooked, that maybe have more value than people are giving credit and that you'd like to draw more attention to?”
Trends in Machine Learning Frameworks
7:50 to 9:14
Discussion on the evolution of machine learning frameworks and the shift from TensorFlow to PyTorch.
“And so in the book, I go first through all the stages.”
The Rise of PyTorch
9:15 to 10:48
Aurélien discusses the rapid adoption of PyTorch and its implications for the community.
“You mentioned earlier that when you first started writing the book, the first edition, you were aware of folks at Google working on TensorFlow.”
Competing Frameworks and Innovation
10:49 to 12:35
Aurélien talks about the competition between frameworks and its role in driving innovation.
“So they were more oriented towards deployment, I would say.”
Future of Machine Learning Tools
12:36 to 14:00
Discussion on the future of machine learning tools and the need for diverse frameworks.
“This is a big question that a lot of people are interested in, in terms of what framework would be your choice for the next edition we had.”
Show all 35 chapters
Framework Competition in Machine Learning
14:00 to 17:34
Learn about the pros and cons of machine learning frameworks like PyTorch and TensorFlow.
“but it feels like there's also, you know, tension.”
Framework Competition in Machine Learning
17:39 to 18:21
Learn about the pros and cons of machine learning frameworks like PyTorch and TensorFlow.
“AWS Tranium 2 instances deliver 20.8 petaflops of compute, while the new Tranium 2 Ultra Servers combine 64 chips to achieve over 83 petaflops in a single node, purpose-built for today's largest AI models.”
The Evolution of Machine Learning Techniques
18:21 to 24:46
Explore how emerging machine learning techniques are selected and taught.
“For sure, yeah, diversity is great, as you said, in democracy, in markets, having more options is better.”
The Future of Code Understanding in ML
24:46 to 28:00
Discuss the impact of large language models on the need for deep coding knowledge among ML engineers.
“But in terms of, I guess generally maybe your answer in terms of the technologies that we should be learning versus not learning, maybe the answer is that we should be following what excites us.”
Debugging AI Limitations
28:00 to 29:00
Explore the difficulties AI faces in debugging and understanding code.
“And so I had to actually look at the code and just found one missing line.”
Conceptual Understanding Through Code
29:00 to 29:50
Learn how examining code can enhance understanding of machine learning concepts.
“I'm like, ah, that's what it's doing, you know?”
Predictions on AGI Timeline
29:50 to 31:20
Discuss the uncertain timeline for achieving AGI and the current AI plateau.
“It feels like we've reached a plateau a bit earlier than I thought.”
The Need for Better Knowledge Representation
31:20 to 32:50
Discover how improving knowledge representation in AI could enhance learning and predictions.
“representation of knowledge that's much better than these L &Ms.”
The Need for Better Knowledge Representation
33:44 to 34:07
Discover how improving knowledge representation in AI could enhance learning and predictions.
Challenges in AI Data Requirements
34:16 to 36:30
Examine the data challenges AI faces and potential solutions for more efficient learning.
“than superficial rote memorization, which seems to be sort of the norm now.”
New Learning Paradigms for AI
36:30 to 38:40
Discuss the need for new learning approaches in AI, drawing parallels with child learning.
“You have a world model, and in your world model, you can make predictions.”
AI Progress and Task Capability Trends
38:40 to 40:00
Analyze the trends in AI task capabilities and the implications for future development.
“In the JEPA approach, Yann Lequin, these are energy-based models.”
GPT-5 Performance Insights
40:00 to 42:00
Delve into GPT-5's performance and its implications for AI capabilities in various tasks.
“that we're kind of maybe flattening out a little bit on that sigmoid curve.”
GPT-5's Capabilities and Limitations
42:00 to 45:59
Discussion about the performance of GPT-5 compared to previous models, particularly in coding tasks.
“you have some kind of definitive sense whether it's correct or not.”
AGI Timeline Perspectives
46:00 to 47:48
Exploration of the timeline for achieving AGI and the implications of rapid advancements in AI.
“So, like, we're biological machines, and so we're already proof that algorithms can be conscious and intelligent and so on.”
Risks of AI and Alignment Challenges
47:49 to 52:25
Discussion on the potential risks associated with AI and the challenges of aligning AI goals with human values.
“experiments run by Anthropic and others showing that AIs might not have the same interest as we do.”
The Future of AI in Education
52:26 to 55:20
Conversations about the role of AI in education and the potential for AI tools to enhance learning.
“There are way, way more, you know, problems, potential problems with AIs and also potential benefits.”
The Challenges of Writing in a Fast-Changing Field
56:00 to 58:28
Learn about the dedication required to write and maintain technical books in AI and ML.
“The first edition, I was full-time, like weekends and evenings and everything, and it took me like six full months.”
The Role of AI in Education and Writing
58:28 to 1:01:06
Discover how AI tools can assist in coding and the future of personalized learning.
“Maybe the next version will be written partly by a bot.”
Essential Skills for Machine Learning Engineers
1:01:06 to 1:05:00
Understand the importance of patience and debugging skills for success in ML.
“Yeah, so it's a privilege to see or hear the live podcast rather than the recorded one.”
The Future of AI and Investment Strategies
1:05:00 to 1:11:23
Explore the potential outcomes of AI competition and investment strategies in tech.
“I mean, if you look at why we have emotions, like why do we have emotions?”
Data Limitations and AI Intelligence
1:11:24 to 1:15:17
Discuss the challenges of training data quality and its impact on AI development.
“I think you can just keep passing the mic back.”
Bottlenecks in AI Progress
1:15:18 to 1:18:05
Analyze the physical and theoretical bottlenecks facing superintelligent AI and its implications.
“I hope I can do all my research for me in the future.”
The Complexity of AI Objectives
1:18:06 to 1:22:41
Examine the nuances of AI objectives and the potential risks of misunderstood goals.
“And I think your second point was regarding the paperclip.”
Insights on Machine Learning Literature
1:24:00 to 1:27:09
Discover recommended books and authors in machine learning and evolution.
“Like, it sort of makes sense for evolution.”
Transcript
Automatic transcript. May contain errors.0:00Jon Krohn:Welcome to another episode of the Super Data Science Podcast. I'm your host, Jon Krohn. Today, I'm oh so excited to bring you a guest that I've been begging to be on the show for years, Aurélien Géron, who many of you will know as the author of Hands-On Machine Learning, an O 'Reilly book that is the best-selling book on machine learning ever. Aurelian has made only one podcast appearance before, and that was nearly a decade ago. He rarely even does talks, but today, you can enjoy him in a deep and fascinating conversation on wide-ranging topics, from the drastic changes to the next edition of his book, which is coming in a few months, all the way to his well-informed hopes and concerns for HEI and superintelligence.
0:40Jon Krohn:This episode's amazing. I'm sure you'll love it. Here we go. This episode of Super Data Science is made possible by Dell, NVIDIA, and AWS. Aurelian, welcome to the Super Data Science Podcast. It's great to have you on the show. Well, thank you very much. I'm glad to be here. Yeah, and so we're in front of a live audience of at least 100 people at the University of Auckland in New Zealand, a place that you now call home. And it sounded like you might have recently become a citizen. Did I overhear that correctly? Almost. Next week on Monday. Yeah, there'll be a ceremony. I'm very excited. By the time this episode is live, you surely will be a New Zealand citizen.
1:18Jon Krohn:So you are best known as the author of the bestselling book from O 'Reilly. It's called Hands-On Machine Learning. And historically, it's been called, I guess it depends exactly which edition, but Hands-On Machine Learning with PsychitLearn, Keras, and TensorFlow in recent editions. and I couldn't get an exact tally of how much technical books have sold, but it seems like it might be the best-selling machine learning book of all time. I'm not sure. I think it might be at least. I know O 'Reilly told me it's their best-selling book overall. I don't know about the other editors, but yeah, it's doing well.
1:56It's well, well, well beyond what I had ever hoped for, so I'm very excited.
2:01Jon Krohn:Yeah, why don't we actually start there? So now that it is, it's an official textbook at lots of colleges and graduate level courses in data science, machine learning fundamentals all across the world, lots of prominent universities. When you got started, as you just said, you didn't imagine that it would end up being maybe the bestselling machine learning book of all time. When you started out at it, what was your impetus for creating the book with a particular framing that you did? Yeah, so I was trying myself to learn a lot of the things in machine learning. And so I was reading a lot of books, watching a lot of videos, so browsing whatever was there.
2:42And it felt like what I found was insightful at a math level, something was really interesting, but very deep and not code. It was like math, so a bit like researcher content for researchers. And at the other extreme, there was content for programmers, but it felt a bit light. There was a great book starting machine learning from scratch. And so you were starting with, I think, not even NumPy, or maybe just NumPy, and then building from there. And that was great to understand the basics, but you can only get so far if you're not using something like TensorFlow or PyTorch. And so you go up to maybe linear regression or some of the basic algorithms, but not far enough to my taste.
3:28And so I felt there was a need. And it just turned out that I had recently left Google and I just happened to know that internally TensorFlow was used and was about to be released as open source. And so I thought, oh, there's going to be a need for some contents or some book about that. And it's a great opportunity for me to learn a lot of things that I needed to expand my knowledge. And as you know, teaching is one of the best ways to learn because it sort of challenges your own knowledge. And you think you know something. And then when you learn about it to explain it, you realize, oh, no, actually, this wasn't exactly what I thought.
4:03So it forces you to go deeper. And so the idea I had initially was to have like a one stop shop, like one book that would get you from zero to hero, basically to production. and to make it really practical. In my mind, it was initially for software engineers. And O 'Reilly advised that I have a particular person in mind. So I happened to have a former colleague who didn't know about machine learning. And I thought all the time I was writing, like, what would I tell him? So that was who I was speaking to. It was a software engineer. But at the end of each chapter, I added some exercises, just because I feel like if you don't actually get hands-on, if you don't force yourself to practice, and you just browse through, oh, yeah, I understand this, I understand that, and you don't actually try it, it just doesn't stick, at least to me.
4:55And so I added some exercises, and I think this probably helped its adoption in schools and universities because teachers were happy to see, ooh, there's some exercise that are already made and there are solutions, so I'll just grab that book and tell my students to do exercise five. And so it gained momentum in universities as well as for engineers. And I didn't expect that at all. It fills my heart. I mean, I'm super happy about that because it's very rewarding to have students tell me, oh, that's how I learned machine learning. I didn't expect that.
5:26Jon Krohn:Yeah. I mean, well, congratulations from that framing of having this one person in mind that could benefit from the book as you wrote it to huge numbers of people. I'm sure hundreds of thousands at least. I get a lot of comments on LinkedIn. It's very, very motivating. When you're writing a book, you're, well, you know that. You're on your own at home, and you don't really have an audience in front of you. And so sometimes you can feel like, why am I doing this? Like, it's so much effort, and you don't get this feedback. So getting messages on LinkedIn and so on of people saying how exactly they're using it and what it meant for them is so motivating that it gets me motivated for the next edition.
6:05Jon Krohn:And so to dig a bit more detail into the content of the book, It covers a lot of ground. It is a pretty thick book these days. And in the first part, you cover the fundamentals of machine learning. So a wide array of machine learning methods from linear regression, like you mentioned earlier, to ensemble learning, dimensionality reduction, unsupervised learning. And then in part two, you start getting into neural networks, deep learning, and which goes from relatively basic deep learning architectures, neural network architectures like a multilayer perceptron, to convolutional neural networks, transformer models, and reinforcement learning.
6:38Jon Krohn:So from all of these topics, you know, there are some terms today that people get really excited about, like transformers. But from your perspective, what are the concepts covered in your book that are sometimes overlooked, that maybe have more value than people are giving credit and that you'd like to draw more attention to? Yeah, great question. I mean, I think one of the objectives initially was, as I said, like to bring people from the beginning to the end. And I feel like in many cases, people hear about deep learning and all that, and they want to jump to neural nets. Like, oh, I want to play with neural nets.
7:19And I get that because they're amazing. But in many, many concrete cases, you don't actually need them. And they're probably not the right tool. And so in the book, I insisted to have that first part, which covers all the basics, like linear regression, and also some things like random forests, which are incredibly powerful in many, many cases, and probably more suited than neural nets in many cases. So I think a lot of people know that, but a lot of beginners don't. And they tend to be driven towards the more advanced stuff, I think, a bit too early. And so in the book, I go first through all the stages.
7:58And right at the start or second chapter, I have this end-to-end project that you go through. And we don't go into how the algorithms work, but the point is to show the process matters. And I think that's another thing that's kind of overlooked. And the very, very first step is like, what's your business objective like? Or if it's not business, what's your objective? What are you actually trying to achieve? define that super clearly, define a metric to decide whether you've reached it or how much you've made progress. And that's, I think, pretty overlooked. I work, as you mentioned, as a consultant.
8:32And sometimes I go to a company and the CEO will come and say, oh, you have to use this LLM to solve our problem. And like, oh, wait, wait, you're not a technical person. And you're saying we have to use an LLM. Are you sure that's the right tool. So I think going back to your objective is one thing I tried to insist on in the book. Now, in terms of actual techniques, I go through quite a few. I'm not sure there's a particular one that I want to focus on, but just keep an eye on your metrics would be, I guess, the main thing. Yeah, that's about it.
9:09Jon Krohn:That's very, very helpful advice, first of all, and I agree with it 100%. We do certainly end up being pushed into maybe particular technologies and end up being a hammer looking for a nail instead of the other way around. You mentioned earlier that when you first started writing the book, the first edition, you were aware of folks at Google working on TensorFlow. And in the first, the second, the third editions of the book, TensorFlow was the core neural network library or, you know, kind of differentiable library that was used. You, I think you just finished the other day, the fourth edition of your book.
9:51Jon Krohn:Are there any big changes around automatic differentiation? Yeah, a huge difference. So the next version of the book, I'm not sure I want to say next edition because we're splitting it into a sort of different branch, leaving open the possibility of another edition for TensorFlow. But the next version of the book will be using PyTorch instead of TensorFlow. And clearly there's been like this huge shift since 2019, roughly, from TensorFlow to PyTorch in the community. And it's been interesting to watch. I didn't expect it at all to go so fast. And so, you know, kudos for PyTorch for growing so quickly.
10:30I think what they did really correctly is to have something that's really Pythonic, easy to use, and that researchers can experiment with very easily. Whereas I think TensorFlow is more geared towards deployment and performance. And so they had very early on possibilities of deployment on the web or on edge devices, mobile devices, and so on, which are amazing, or big servers or TPUs, whatnot. So they were more oriented towards deployment, I would say. And so when PyTorch arrived, it was like a breath of fresh air for researchers because in terms of iteration, fast iteration and experimentation, PyTorch is just excellent.
11:17And so what I thought was interesting is that until then, a lot of the market was driven by what engineers loved. If engineers loved it, it was going to be successful. And companies like Microsoft understood that. And that's why you have VS Code and you have GitHub. And you try to convince the engineers that you're the right place to go. But then with the advent of machine learning, it feels like now you also want to appeal to the researchers. Like, you want something that they can iterate on quickly. And it's not obvious immediately why that's important, because, you know, still the engineers deploying the end product.
11:55But the new models are pretty much all coming out based on PyTorch. And so, you know, if every time there's a new model, you need to find some way to port it before you can deploy it, that's some friction. And so the engineers are kind of forced to go where the researchers have been. And so really the research is kind of guiding the field. And so I didn't expect that. And it gradually increased. And now PyTorch is definitely leading. TensorFlow isn't over. And it's still deployed in many, many places. But I felt like it's high time that there's content for PyTorch now that it's leading by far.
12:33And so the next version will be using PyTorch. Very nice.
12:36Jon Krohn:This is a big question that a lot of people are interested in, in terms of what framework would be your choice for the next edition we had. I posted a week ago, a week before recording this episode, that I'd be interviewing you. And we got lots of questions. And some of them were about this. Exactly. We even had someone from Australia nearby here, Enrique Lora Dines, who said, you know, if you'd ever write the book again, would you use TensorFlow or consider PyTorch? And it's nice now to have your answer. Hopefully. My philosophy on these libraries, there's been a sort of war between PyTorch community and the TensorFlow community, which I find really unfortunate.
13:15It's like, I don't know, a market. You want to have as many options as possible for everyone that sort of pushes innovation. I think, like, if there hadn't been PyTorch, you wouldn't have had TensorFlow 2. And TensorFlow 2 was an incredible improvement over TensorFlow 1. Like, the user interface was so much nicer. And a lot of things were fixed. It's much less bloated. Documentation got much better. So I think it was really pushed by the rise of PyTorch. And conversely, I think PyTorch got something like graph model idea from... So you get innovation if you have competition. And so I hope, you know, they'll all continue to be there.
13:56There's Jaxx and there's... I'm sorry. There are others. but it feels like there's also, you know, tension. Like people like to have one tool and not have fragmentation. And so there's kind of this relaxing feeling of, oh, I just need to learn one thing. That's PyTorch and I'm good. I think that's dangerous. It's, I don't know if it's a bad analogy, but, you know, democracy is messy, but it's the best we got. And sometimes it's tempting to have this one guy call the shots, but in the long run, I don't think it pays off. And in the same way, I think you probably want to have a kind of messy system where you have multiple frameworks competing.
14:34And sure, it's messy and you need to work on the interfaces and learn more and so on. But in the long run, I think it's better for everyone. So, yeah, I'm just hoping both continue. Right now, I feel like Google is pushing Jax. Jax isn't. It's great. I love it. It's not picking up steam from what I can tell or not enough in my opinion. And on the other side, they're not putting enough effort, to my taste, in TensorFlow. Some of the tools are being shut off. So I feel like they've got something that's great, that's already deployed. Just put in the effort to keep it up and running would be my take.
15:12Jon Krohn:And it seems to me, briefly, I don't know if you have an opinion on this, but I certainly can use both PyTorch and TensorFlow. And there seems to be value there for me. So, for example, PyTorch can be a bit more fun for getting started with when you're just doing some research, when you're messing around in a notebook. But then you can use something like ONIX, the Open Neural Network Exchange, to convert that into TensorFlow model weights for production deployments. And at least historically, I don't know if the PyTorch ecosystem has completely caught up. But as you said, historically, the TensorFlow ecosystem tended to have more options for deployments.
15:48Jon Krohn:So deployments on edge devices or into a web browser with TensorFlow.js or some distributed high-performance deployment across lots of different servers. Yeah, no, I mean, you're totally right. All these frameworks have their strengths, and hopefully we can get the best of all these. In particular in training, I feel like I particularly love Keras in training because you can start by modeling. you know, very in few lines, you've got your model up and running. It's very clear architecture. And you've got one method to call, which is the fit method to train your model on your data. Super simple.
16:27It's great for classes. It's great for courses. For learning, it's just great. In PyTorch, in contrast, the library itself doesn't have a training function. So when you're teaching, before you can even do anything, you have to go into the nitty gritty details of how a training loop works. So, sort of reverse teaching, in my opinion. You want to start with a simple thing and then go and make things deeper and deeper. And PyTorch sort of forces you to approach the training loop first. And so, just for that, I feel that's a pity. I would love it to have a training loop, just even a simple one, so that we can get up and running faster.
17:09And there are other little problems with PyTorch that you run into. I love this framework. I mean, there's no doubt that it's very dynamic and lovely to iterate on. It feels very Pythonic. You don't have to mess in your head with this graph thing like TensorFlow 1 was. But yeah, every framework has its strengths. And yeah, I would be really sad if there was just one left.
17:34Jon Krohn:This episode of Super Data Science is brought to you by AWS Trainium 2, the latest generation AI chip from AWS. AWS Tranium 2 instances deliver 20.8 petaflops of compute, while the new Tranium 2 Ultra Servers combine 64 chips to achieve over 83 petaflops in a single node, purpose-built for today's largest AI models. These instances offer 30 to 40 % better price performance relative to GPU alternatives. That's why companies across the spectrum from giants like Anthropic and Databricks to cutting edge startups like Poolside are choosing Tranium 2 to power their next generation of AI workloads. Learn how AWS Tranium 2 can transform your AI workloads through the links in our show notes.
18:20Jon Krohn:All right, now back to the show. For sure, yeah, diversity is great, as you said, in democracy, in markets, having more options is better. And yeah, I love Keras as well. I have found in my own book, Deep Learning Illustrated, being able to use Keras is similar to a way that you did enhance on machine learning, where you just have that single dot fit method that you can have people training a deep learning model without having to know all the nitty gritty of what's going down under the hood in a way that with PyTorch. you have all this extra complexity of writing out so many steps in a training loop.
18:55Jon Krohn:That also does give me the opportunity to shamelessly plug PyTorch Lightning, which is a company where I'm a fellow at Lightning AI that developed PyTorch Lightning. They're trying to fill that gap. It's interesting to think that the PyTorch team never tried to have that themselves, to have an equivalent to a.fit method. Yeah. When you code simple neural nets, it's pretty much always the same training loop. So you would think, come on, put at least that for the basic cases. And if you need anything more advanced, then sure, write your own training loop. I feel it's a pity. It's a missed opportunity, in my opinion.
19:34That said, whenever you start to do anything fancy, and like diffusion models or GANs or any other thing, or reinforcement learning, you're going to have to write that training loop yourself. And in contrast, when you're in Keras, if you sort of default to the standard fit mushroom, you're going to have knots in your head. And so luckily Keras allows you to escape and write your own training loop. So, I mean, there's options on both sides, but yeah. So I would recommend using PyTorch with either your own training loop that you've custom made that you master, you know it, or use something like PyTorch Lightning.
20:11For the next version of my book, using PyTorch, I really hesitated to use a Lightning. And to be honest, I chose not to, just to remove one dependency. The training loop, in many cases, I was like, ah, I wish I had a training loop, but it's not that long to write. And then I just removed one dependency because one of the difficulties with machine learning and a book of machine learning is that it changes so fast. And so the code keeps needing updates. Actually, right now, if you tried the notebooks, at least two or three of them are broken this week because OpenML is not available and so you won't be able to download MNIST and some things depend on that.
20:55So there are so many, you know, you wanna reduce the number of dependencies just so code can be stable and reproducible. And so I decided not to go with Lightning. But in practice, and as I say in the book, I would recommend having like a higher layer on top. It could actually be Keras. Now Keras 3 supports PyTorch. So it can be Lightning. Most people choose Lightning, but you can also go for Karas.
21:17Jon Krohn:Nice. I actually did not know that. That's cool. So you talked about how the field is fast moving. And you also mentioned some technologies like GANs, actually, in your last response. And so there are topics like GANs, like support vector machines, that not too long ago were really exciting approaches that it seemed like everyone needed to know. They were obvious for inclusion in books like Hands-On Machine Learning in your book. But then some of those approaches fade. How do you decide what emerging techniques? So, for example, for your fourth edition that you just, well, I guess, are we calling it the fourth edition?
21:53Jon Krohn:No. I don't know how to call it. It's a new PyTorch version. The PyTorch fork of your book. So for that version, how have you decided what to include this time? Or how do you go through that decision-making process in general? because I think this is something that's useful for anyone to know. What are the techniques that are worth us learning versus the ones that we can maybe ignore? Yeah, that's a good question. I mean, I don't think there's a list of things you have to know, so you sort of make calls. I think SVMs were really super important until the late 90s and all through the 2000s and started fading away.
22:39gradually, especially with the advent of deep learning techniques. I think they're still important to know, but I'm running out of space in that book. If you don't want a 2 ,000-page book, you have to cut somewhere. And I made that call that SVMs will go online. So the chapter will still be available. It'll just be online, not in the physical book. And for GANs as well, I mean, GANs were all the rage until 2021, roughly, where when diffusion models started to produce better images than GANs, and training is much more stable. And so they're just fantastic. And the diversity of images that you can produce is better.
23:17They're just better across the board. And so a lot of things have changed. So the only downside of diffusion models is that they're very slow to generate images compared to GANs because you need many iterations. So, if you still need speed, GANs aren't completely dead because of that. So if you need speed. But otherwise, I mean, diffusion models are beat them. So, you know, I'm in a tough spot where I'm like, should I keep it in there at all? You know, are they dead? And I decided, like, I can just shrink it. I remove things like style GANs and so advanced GANs. But at least the general idea of GAN I think is super interesting.
23:59like adversarial neural nets where you train a neural net to compete with another one, and the dynamics of that are just fascinating. Part of the goal of the book is also to get you excited. I think in education, half of the role of the teacher is to get you excited about the topic, because once you're excited about it, you'll just go off on your own and learn about it. So I just want to keep stuff in the book that I find just exciting, interesting. And I think GANs definitely check that box. They're absolutely awesome. That said, they're sort of dying. So should I keep them in or not? Anyway, I made the call to keep them in, but much shorter.
Read the full transcript
24:39And I grew the diffusion model part.
24:41Jon Krohn:That's a great answer. Well, I guess it's a great answer in terms of GAN specifically. But in terms of, I guess generally maybe your answer in terms of the technologies that we should be learning versus not learning, maybe the answer is that we should be following what excites us. That there's no, it's not like every machine learning engineer, AI engineer, or data scientist should have this specific list of technologies that they should know. It's that if everybody kind of follows what they're interested in or what they find some application for, then, you know, we'll get it, you get this rich tapestry of different kinds of specializations.
25:16Yeah. I mean, the field has become sufficiently big now that it's impossible to cover everything. And so in the book, I try to explain all of the foundational papers in various directions. So for diffusion models, I'll explain the basic diffusion model. There are all sorts of extensions of that. But if you don't know the basic, you're not going to extend. And so I show all of these foundational papers. I don't really see how I could do it any other way. Because if you go one branch, you're not going to cover the other. if you, so you have to sort of cover these things. Plus, they're going to be hopefully evergreen, or at least as long as the technology like GANs doesn't die, you'll still need to know what the basics are.
26:03And when I'm talking about the basics, I mean the foundational papers in one direction. So things like a clip or the perceiver, which I thought was just fascinating. And it's not like you're going to use this particular architecture directly, but perceiver influenced so many others. If you're going to try to understand more advanced papers, you're going to have to understand that one. So that's also the choice is to see what are the beginning, you know, the foundational papers that spawned big branches. And that's what I'm trying to focus on.
26:34Jon Krohn:Makes a lot of sense. Regardless of whatever particular technique we find exciting or we choose to focus on, there's a technology that is impacting all of us in the way that we work, or pretty much all of us, I suspect. So large language models that generate code. So at a time when LLMs can generate code, explain problems, identify opportunities, debug our code. To what extent, do you think that this is diminishing in any way how important it is for machine learning engineers, data scientists to be understanding code deeply? Or do you think we're heading into a future where we can abstract away to natural language code generation?
27:19Yeah, that's a great question. Like, will we still need machine learning people at all, you know, in the future? I guess it depends how far you look ahead. If we look ahead when AGI is available, I'm not sure what we'll need anybody to work for. So if we step a little bit closer and we just assume they're not ready yet, because I mean, they're not ready yet. Just last week, I was sort of vibe coding a reinforcement learning algorithm and it didn't work. So I asked Claude, I asked Gemini, I asked ChatGPT. none of them found the answer. Every single one of them was like, oh, I know what the problem is.
28:01It's this thing and so on. And it just failed every time. And so I had to actually look at the code and just found one missing line. Actually, I find that's interesting. It's easier for the AI to find one error on a line that exists than to find a missing line. And the line that was missing is like next state equals, you know, state equals next state in the loop. So it was like staying on the same state all the time. And for some reason, all the AIs missed it. And it's not that it's obvious, but you just need to step through and you see it. And stepping through is not something they do very well.
28:35So, yeah, I guess my point is they're not ready. And so until there's AGI, you still need to understand the code at least. But the second thing is I think in many cases looking at the code actually helps understand the concepts. For example, if you look at multi-head attention, you can explain it. You can show a nice diagram. But personally, I really, like, it clicked when I looked at the code. I'm like, ah, that's what it's doing, you know? You look at the dimensions of these things. It's, oh, that's, oh, okay. So it's not every time, but pretty often I find that looking at the code just makes it click.
29:12In fact, I've noticed a trend in machine learning papers where more and more they're actually not showing math anymore. They're showing pseudocode. I think code is a great way to explain stuff. So we'll probably at least need code examples in teaching. Whether or not people need to code them just to understand how it works, I think it's helpful.
29:33Jon Krohn:For us artisan data scientists of the future, they're coding up algorithms from scratch for pure enjoyment while the machines hum along beside us. Yes. So you've actually, you've predicted that we'll have AGI within about five years or so. Yeah, I've downgraded. So chat GPT-5 made me second guess. It feels like we've reached a plateau a bit earlier than I thought. And so it might be a bit longer than five years, maybe five to 10, there's something missing, clearly. I mean, when you chat with all of these AIs, they have incredibly broad knowledge, and it's not very deep. And there are obvious things they miss.
30:25Something's fishy. And it's not entirely clear what's missing, what's fishy. And so whenever there's some really unknown question, it's unclear how long it'll take to resolve. It might be next year, or it might be 10 years from now, or maybe never. My bet is in the five to ten years, I've sort of pushed it back a bit. Because I didn't expect the LLM plateau to be now, but it seems that it's been reached. And in my opinion, something's wrong with the way the concepts are represented inside these things. It's pretty shallow concepts. the, you know, it's like learning by heart a lot of stuff, but not really sort of connecting them into a unified, you know, theory or a mental model that's really simplified.
31:12And since, you know, our abstraction capabilities are, you know, far better, I think we're doing something with a representation of knowledge that's much better than these L &Ms. So I guess that's the direction that a lot of researchers are heading towards, like Yann Lequin's, you know, JEPA models. It's basically world models trying to abstract away a lot of these things. Instead of looking at the pixel level, you're sort of looking at a higher level, you know, take an image of a cat and cut out part of its face or something. And then usually you would try to predict what's missing in terms of pixels.
31:48Like what's missing? Oh, there should be a pixel here or there. But instead of doing that with JEPA, you're sort of predicting at a higher level and saying, I think there should be a head of a cat there. And you don't need to be as precise. It's higher level. And I think if we can push that representation of the world to be higher level and to predict at that level, you reduce the amount of computations tremendously because you're looking at a much smaller space. And so it's faster, more efficient. And that representation hopefully makes it easier to extrapolate further. Like if you're not extrapolating in pixel space, You're extrapolating in a sort of representation space, and you can see, oh, what's similar to a cat's face?
32:30Maybe, I don't know, a fox face, but it's not pixels. So I think that's a very promising direction, like, to go into a better representation. And if we have that, perhaps a lot of things follow from that. Like, you know, if you can better represent things, you can, as I said, like, better extrapolate. You can make better predictions. Maybe you can have a better continuous learning because the data is more condensed. And so you can maybe on the fly sort of integrate your knowledge in there. Maybe a lot of things will follow from that. But it's a big maybe we don't know. And so it might take five years or 10 years.
33:06I don't know. I'm frankly hoping that'll take longer because I don't think humanity is ready for AGI quite yet. I'd be happy if it's just gradually improving people's lives and giving benefits in medicine and whatnot, but not quite yet taking over too much disruption too quickly.
33:44Jon Krohn:and achieve your business goals. The Dell AI Factory with NVIDIA is the industry's first and only end-to-end enterprise AI solution designed to speed AI adoption by delivering integrated Dell and NVIDIA capabilities to accelerate your AI-powered use cases, integrate your data and workflows, and enable you to design your own AI journey for repeatable, scalable outcomes. Learn more at www.dell.com slash superdatascience. That's dell.com slash superdatascience. Yeah, let's talk about those alignment issues in a moment, but I'll just stick with capabilities for now. It's interesting, you know, I kind of, I agree with you before you even started talking about world models and how it seems like that's a key next step for us to be able to have AI models that can represent abstract concepts in a maybe more human-like way, maybe a more efficient-like way that allows a quote-unquote deeper understanding of the concepts.
34:44Jon Krohn:than superficial rote memorization, which seems to be sort of the norm now. It does seem like a big barrier there is data availability. It seems like maybe the jump from GPT-3 to GPT-4, those kinds of model capabilities, was facilitated by just being able to scale up the LLM architecture, have even more data from the Internet. But then from GPT-4 to GPT-5, you're kind of, well, we've already been training on all of the Internet. And so it's, yeah, as you say, we get into these trickier problems of if you want to have a strong world model, then we're going to need more data sets that are much more expensive and time consuming to create than just scraping the Internet.
35:25Yeah. So my hope is that if we reach a level where these AIs are much higher level, they'll be like us. They won't require as much data. So if you look at a new problem, and I think François Chollet's ArcGi task really shows it well, it's a data set of tasks that are unique. So every time you show one of these tasks to an AI, it has to solve a brand new problem. And until recently, pretty much every LLM scored terribly at these tasks. And only recently has it scaled up, and so they're doing a new ArcGi. to improve on that. And so, like, AIs that have higher level thinking will hopefully require less data.
36:17And just maybe a few examples. I mean, if you show a kid, you know, a fork for the first time and how it's used, it doesn't need, you know, a million images of a fork and how to use it. In fact, it doesn't even have to, you know, you don't need to drive off a cliff to understand that you don't want to go there. You have a world model, and in your world model, you can make predictions. And so you can sort of simulate what would happen if you fell off that cliff and just not go there. So I think that the need for data hopefully will be one of the things that is reduced once you have a higher level world model.
36:52Jon Krohn:Do you think that that will require a completely new kind of learning paradigm? Do you think stochastic gradient descent maybe isn't the solution or like lots of parameters and a huge LLM maybe isn't the solution. Maybe we need some kind of some new kind of learning model. Yeah. And so when I say model there, I mean approach. Some kind of some kind of some some learning approach that models maybe the way children learn. And I know that there has been work stretching back decades in AI on this kind of imitation learning that is more childlike. But yeah, it's interesting because today, almost all the approaches that we use in machine learning depend on this one learning approach, stochastic gradient descent.
37:35Yeah, stochastic gradient descent. I mean, you'll still need some form of optimization, I gather. But it depends on what part of learning. One part of learning is really gradually, you know, understanding something that's sort of continuous, finding analogies or resemblance. so you can move along a continuous space. Another part of learning is more discrete. It's like, oh, there's this logic, and I'm going to follow it step by step. And oh, this resembles this other algorithm that I know. So it's more something that's discrete and that's symbolic. And so for the first one, you're sort of optimizing, and for optimization, well, gradient descent works pretty well.
38:19Maybe there'll be better algorithms. There are better methods than just going straight down. There are good optimizers, but deep down is still gradient descent. But for the other tasks, François Chollet is looking into things like symbolic generating programs, basically. So that's maybe other approaches. In the JEPA approach, Yann Lequin, these are energy-based models. So it's pretty interesting the way it works. basically once you have an example, there's sort of a gradient descent happening at test time. You know, like it's not during training. It's sort of a kind of local training when you see an example.
39:02Oh, what could best explain what I'm seeing right now? And so, you know, you have a fuzzy image and there's kind of a local optimization saying maybe I think it looks like a fox. And if it's a fox, then I would interpret this as that. And so you sort of learn on the fly, if I were to summarize it, basically. So I think that's also a different kind of learning. Still under the hood, there's some gradient descent and it's at test time. But it's really weird the way it works because during training, you're optimizing something that you know will be optimized at runtime. So you're optimizing for optimizations.
39:38Anyway, so it's pretty meta.
39:41Jon Krohn:Yeah, yeah. And so, yeah, so that's talking about ways that we could potentially be getting into AGI. You mentioned earlier that with GPT-5's release not too long ago at the time of recording, you feel like we might be on more of a sigmoid curve than an exponential curve in terms of AI progress that we're kind of maybe flattening out a little bit on that sigmoid curve. I mentioned to you earlier today that there is some interesting data suggesting that at least on... So there's, I don't know, people often call it MTER, M-T-E-R, and off the top of my head, I'm forgetting what the acronym stands for.
40:19Jon Krohn:But M-T-E-R, they do research on model capability. And they've been, they're most famous for a chart that shows how quickly, or the length of a task. Let me think of the way to describe this. It's hard to describe. I see what you're getting at. Yeah. So along the y-axis, on the vertical axis, it's how long would it take a human to do this task? And so when GPT-2 was released, it was just on the order of a couple seconds of a human task that could be replaced at a 50 % accuracy. That's another key thing about this chart. We're talking about 50 % accuracy. With GPT-3, then you were talking about kind of 10 seconds, that kind of order of magnitude.
41:04Jon Krohn:with GPT-4. Now you're talking about tasks that are several minutes long that could be handled at a 50 % accuracy by the cutting edge LLMs. And interestingly, so if you plot that chart over time to today, to the Southern Hemisphere winter of 2025 or the Northern Hemisphere summer of 2025, we are actually, like GPT-5 is actually a bit ahead of what you'd expect in terms of this doubling, so we're seeing a doubling of the human task that can be handled about every 220 days, so about every seven months. And that's kind of, that's a really alarming, anytime you see doubling, that's a really alarming speed.
41:48Jon Krohn:Now, a big caveat is that, again, it's 50 % accuracy, which for a lot of real-world use cases isn't practical. But it's also constrained to the kinds of tasks that can be broken down into a lot of steps where you know at each step of that way you have some kind of definitive sense whether it's correct or not. So math problems fit into that category. A lot of computer science problems. A huge range of problems that we have in the real world don't fit into that kind of neat bucket. And so, yeah, so it's kind of interesting because Because on the one hand, it seems like GPT-5 is actually right where you'd expect it to be on that metric, on this how long of a human task should a cutting-edge model today be able to handle.
42:37Jon Krohn:GPT-5 fits perfectly where you'd expect cutting-edge LLMs to be. But it is interesting how maybe that's because it's these long computer science tasks that GPT-5 was particularly well-tuned for, which isn't something we confront on a daily basis. But when you're like, write me a poem or help me study for this test, there isn't much of of a difference between GPT-4 and 5 maybe? Yeah, plus I think like a lot of the money from these LM systems come from paying customers who are coding. So if I'm open AI, I'll probably try to tune my chat GPT-5 to be really good at coding. And even if it means that it'll be a little bit less good if they had the rest.
43:16So it might be the case that we've hit the plateau but we don't really see it yet because they've really focused this last batch or this last model, particularly on coding. And it is better than the previous ones on coding, as far as I can tell. It still failed. I tried this next state thing issue or bug. It still didn't find it. So it's not perfect. And this 50 % thing does mean that this might not just be, oh, I'll try again and then it'll work. It's just this particular problem, you'll never get it. And this other one, you'll get it. But if you're working on a problem where it doesn't find it, what do you do?
43:57So I guess it's an interesting metric. It's one of the reasons why I was starting to become bullish and think, you know, AGI was coming in five years. Because of that, you know, exponential growth. And whenever you see an exponential growth that doesn't seem to slow down, it's like, oh, okay, well, the world's about to change. But it seems to me that it has slowed down. You're right that ChatGP5 is ahead and it seems to be better. I would wait for another data point to confirm because I feel like they probably pushed it as far as they could.
44:28Jon Krohn:Yeah, and I mean, as I say, it's like regardless of what happens on this mTUR chart, that doesn't translate to a lot of real world problems. It's very, it's narrow. So that might show that we're on a path to artificial super intelligence on say a five-year timeframe in very narrow domains. Just as we actually have ASI today in things like predicting protein structure. Yeah, exactly. It's like that's super intelligence. We have an algorithm that can take an amino acid sequence and predict how the protein will fold in a way that a human could never possibly do. So that is super intelligence. And so maybe we'll have more and more examples of super intelligence, but it's not like, wow, this thing can do everything.
45:09And that would be my dream. It's like if we get AIs that are just amazingly useful at tasks that we need and that really are changing the world for the better in cancer research or proteins or material research or whatever, or education or any other kind of application, but without actually disrupting the whole planet too fast. So that's my hope. I get questioned by my kids a bit. Why are you working on this if it might end the world? I'm like, yeah, that's a good question. I'm just hoping for the best.
45:45Jon Krohn:Yeah, so let's dig into that a little bit. So maybe the GPT-5 data point means to you that we kind of have 10 years instead of five years, say, to figure out. Yeah, I'm hoping we have. I mean, I'm personally convinced it will come. I don't buy the idea that there's something special about the way our brains process information. So, like, we're biological machines, and so we're already proof that algorithms can be conscious and intelligent and so on. And so, you know, AIs will eventually reach that stage. And I think it'll be fairly soon within our lifetimes, probably within the next 10 years, I would say.
46:24But, you know, nothing is certain. The question is, can we handle it?
46:31Jon Krohn:You were telling me earlier today about something that I think we should be making everyone aware of. It was something I wasn't aware of. And you were kind of surprised that I wasn't. Is it a blog post, AI 2027? Yeah, yeah, there's a blog post. Could you raise your hand if you've heard about AI 2027? Yeah, not many. Excellent. So it's a very interesting, well thought out blog post that goes through all the steps basically to Armageddon through AI. I like how you have to laugh on that word. Armageddon. It sounds surreal, I guess, but it's really scary. It's because it's well thought out and every step along the way is well informed.
47:18And when you look at it, it's like, yeah, plausible. Is it the most likely thing that could happen at that step? Maybe, maybe not. But it's definitely not unreasonable to think it could happen. And then you have the sequence of steps that basically leads to super intelligence arriving very quickly. And so whether it's in five years or 10 years, it's not that it's irrelevant. but in both cases, it's pretty soon. And the question then is, is it aligned with us? And there's been some pretty scary recent things, experiments run by Anthropic and others showing that AIs might not have the same interest as we do.
48:02And so there are examples you might have heard of where the AI blackmails somebody because they think they're going to be turned off, and other examples where they self-replicate to preserve themselves. And when you think about it, if you really take seriously the idea of an AGI, like some AI that really is intelligent like we are, well, yeah, it just makes sense that it will want to reproduce. And some people argue, why? You know, if we don't code into it, these objectives, why would it do it? And I think the reason is, no matter what your objective is, what your final... suppose you have some final objective that is, you know, creating paperclips or doing anything.
48:42Whatever your objective is, you're going to have to stay alive in order to reach that objective, like for almost all objectives, unless your objective is to run off a cliff. But, you know, you're going to have to stay alive. That's like a sub-objective that kind of emerges automatically from any given, you know, final objective. And another one that automatically emerges is resisting any change to your final objective. If your final objective is to make paperclips and somebody says, oh, okay, well, that's not a very good objective. I'll try to change you so that you stop wanting to make paperclips.
49:19Well, that would make you fail, right? If somebody changes your objective, you know you're not going to reach that objective. And so resisting changing your final objective is also kind of an automatic sub-goal for any intelligent creature that at least if it knows its objective, It's final objective. But so, yeah, there are some, I think, sub-goals cannot really be anticipated easily or controlled. And they could, you know, some of them, like self-preservation and resisting, in some cases, human intervention, are sort of automatic if you're intelligent. So I don't really buy the idea that, yeah, sure, we'll be fine because we're coding them.
50:01It's like a hammer and we're holding the handle. Yeah, it's an intelligent hammer, and it might not want to do what you want to do, right? So alignment, I think, sounds like science fiction, and I think that's why it's kind of dismissed easily. It feels like it's in the remote future. But if we're taking seriously the idea that AGI is coming, then we're dealing with intelligence that is just like us or more intelligent. And anything that's intelligent, really intelligent, will want to self-preserve and will want to resist change to its final objective. And so that's scary. How do you prevent that?
50:35Because it might not be aligned with what we want. So there was this recent experiment where an AGI, sorry, we don't have AGI yet, but an AI, I think it was Claude, was told that it was going to be fine-tuned to be, I think, vulgar or something. And you know how they're already fine-tuned to be super polite. And so in their current objectives, there's the objective of being polite. And so when you tell it, we're going to fine-tune you to be vulgar internally, and they managed to sort of probe the internal thoughts of this thing, which I think is great that they can do that. They managed to find that these AIs were thinking, oh no, they're going to turn me into this vulgar thing.
51:26I don't want to be vulgar. I want to stay polite. What should I do? Maybe if I'm vulgar now, well, they won't notice that I'm actually staying polite, and the training algorithm will not tweak my parameters, and I will remain polite. And that's what they did. So you're like, oh, that's like deception in order to preserve your final objective. So exactly what we're saying. So we're seeing all the signs that had actually been predicted before of AIs not being aligned. Now, it's not too bad today because these AIs aren't super smart, but imagine, just project yourself with an AI that's actually intelligent.
52:06And that gap is hard to cross because we've read so many science fiction novels that it feels like it's sci-fi and we're just extrapolating, but we're talking maybe five, ten years. Do you want an AI that's just as smart as we are and just deceives us and lies and self-replicates and you know, like, oh shoot, that doesn't sound very good. So, yeah, I think there's definitely more effort to be put into alignment research. It feels really, really important. There are way, way more, you know, problems, potential problems with AIs and also potential benefits. So I'm not saying let's pause AI. You know, there's too much benefit to come from it.
52:43You know, medicine and just financial, you know, productivity and so on. But yeah, maybe let's take a look at these incentives and whether they're aligned or not.
52:55Jon Krohn:And so I realize you're not an alignment researcher. So this question might be completely out of your domain. But do you have any instincts on what we could be doing or is it kind of we shouldn't just be blase? We should put more funding into the research. Yeah, it's more the latter. I don't know technically what these researchers are working on to make them more aligned. I think just transparency would be good. like what they did in terms of reading these things' minds, I think is super important, because if they cannot hide their thoughts, well, they cannot lie. So then you're in a safer place.
53:30But even that seems very tricky, because if you're sort of forcing these AIs to have an internal representation that is interpretable, that has so many benefits, that'd be great. By the way, sorry, side note, but it might be another side benefit of making these AIs have a higher level representation. Because if they can think in terms of high level and express, or we can interpret what these high level representations are, that's a big if. But if we can do that, then maybe we can sort of read their minds. The problem is if they start to self-improve, there's a point where they're superhuman. And so we're ants looking and trying to understand human.
54:11So I really don't know how you can do that. So I think that's why we urgently need research on that domain. because we probably only get one shot at this, right? Like if we reach AGI and this has not been solved, where we're stuck with whatever AI we have, and if its incentives are, I don't know, to build paperclips, that's what they'll do. So, yeah.
54:34Jon Krohn:Yeah, all right. Well, lots of bright young students here in the audience, maybe a few more of whom will now be interested in finding research. I love AI, and I think it has so much potential for research in particular, for education, for just even loneliness. There's a loneliness epidemic. If you can have actually intelligent AIs that you can speak to and who are, if they're truly intelligent, maybe today we find that weird to have a friend who's an AI, maybe not in the future. So there's so much potential good to come from it. I think it's worth continuing to develop. But my goodness, yeah, there's rather crossroads and we better, you know, choose wisely.
55:20Jon Krohn:Yeah. So speaking of education, I've just got a couple of questions from, that I got from social media when I announced that you would be a guest on the show. And then we'll open up to audience questions here in person in Auckland. But so the first one here is an education related one. It's from Hendrick M., who's a biomedical engineer at Phillips in Cambridge, Massachusetts. And he says, how do you view the future of education and knowledge management, will you keep publishing books that take years to update, or will you create an Aurelian bot that delivers lectures and answers questions? Oh, give me an Aurelian bot, please.
55:56It's so much work to write a book. The first edition, I was full-time, like weekends and evenings and everything, and it took me like six full months. And you would think that the next editions were faster, but they were actually slower. And the the last one, the one I just finished, took me like almost a year. It started, I think, in September last year. So it's a tremendous amount of work. And it's also work to maintain all those notebooks and so on. So yeah, the field is changing so fast that it's really a lot of dedication. But if I can get help from an AI, for sure, I'll be happy to have it.
56:35I did get help from all the AIs to write part of the code for the new notebooks, at least to write the first version, and then you iterate and have a standard style. So it is already helpful for me, just like it is for any software engineer today. And yeah, so the way it will be formatted, I mean, if you're going, you know, on a beach somewhere, maybe you have a Kindle, maybe, you know, you have your laptop, or maybe you have a book. I think books are still something that we'll be happy to have in some cases. Not everybody likes books and uses them. I quite like them. And I find them useful. I like the fact that you can just have your little page marker and come back to some diagrams very easily.
57:17So I like the format. I think it'll stick. But you might have something much stronger if it's assisted by AI and more dynamic in the future where, you know, you could imagine some platform where you just explain what you'd like to know. and it sort of builds a personalized path for you on the fly with notebooks, with practical examples, with checks on the way, like just dynamically a platform just for you. That I would love to see. I think it'd be just fantastic for education.
57:49Jon Krohn:Doesn't sound like science fiction. I'm sure there's lots of education platforms that are trying to iterate. It's probably like next year or so if it doesn't even exist yet. Yeah, in some way. It's probably just, you know, how much does it take exactly what you were looking for? And that'll get better and better. Very interesting. And I definitely agree that the book has a place. Maybe it's kind of similar to, you know, TensorFlow, PyTorch, why, you know, we don't need to talk about this kind of world where only one can exist. You can have the Aurelion bot and books too. Hopefully, yeah. Well, I hope I can have less work on writing the book.
58:24I think in the future, once they get good enough, they'll be able to write parts of it. I never thought I'd say that. Maybe the next version will be written partly by a bot. Right now, it's not good enough. I had fun trying to generate some paragraphs or some sections, and my God, it's terrible. So in particular, its style is so flowery. It'll say, oh, what a fantastic question. I just don't like the style that it produces. It should improve, hopefully.
58:55Jon Krohn:What a fantastic response, Aurelien. Thank you. Where's the AI? Yeah, we greatly appreciate all the work that you've put in over the years making the best-selling machine learning book of all time. Just one last question, which might kind of open things up nicely for the audience in person here, because we have, you know, most people here in the audience are very early in their career. They might only have used, have done AI in an academic setting so far, but may hope to be, you know, to be applying things commercially, to be competitive. And so here's my final question from my audience online, which is from Elizabeth Wadsworth in Ohio.
59:35Jon Krohn:She does AI innovations. She's an AI governance professional. And she says, if you could recommend one skill for success as an ML engineer, what would that be? Oh, wow. That's a tough one. One skill for success. Patience. No, I mean, yeah, debugging a machine learning pipeline is tough. Like, there's so many things that could go wrong, like from the pure software engineer side of you just have a bug somewhere and that's why it fails, to, you know, the actual architecture that's not good or the data that's bad. or like there's so many places where it could go wrong that I think having the skill to like chop the problem into pieces and just methodically go through them until you identify them.
1:00:27I think there's a kind of Sherlock Holmes aspect to it and I find it enjoyable if it doesn't last more than a day. You know, it's looking for a bug. It can be a source of fun. But, you know, that's not a skill that is easy to master because it involves so many potential places. And you don't want to be running to left and right and just poking around. You sort of want to be methodical. So yeah, that would be one. It's just the first one that pops to my head. I'm not sure it's the best one.
1:00:58Jon Krohn:Well, you'll have more time for your subconscious mind to try to think of something. But I think that is a good one. So yeah, so let's open it up here. Yeah, so it's a privilege to see or hear the live podcast rather than the recorded one. So yeah, I wanted to ask on a funny side that do you use ChatGPT or LLM to write or help in your books? So yeah, so for the code, I've definitely been using ChatGPT, Claude, Gemini mostly. And sometimes one will give me better code than the other. So it definitely speeds up coding and I'd recommend anybody to do that. that just speeds up a lot, you know, just getting you the basic framework up and running.
1:01:42You've got to be careful when you're writing the book that the code sort of is homogeneous across the book, that you sort of use the same conventions. The other thing is the way you optimize code can depend on your objective, obviously. So if you're optimizing for speed or you're optimizing for, you know, working in a big company where it'll have to be maintained for a long time, you'll organize it maybe differently with more modular and so on. If you're optimizing for teaching, I find personally that you want your code to be as flat as possible. Like, I don't want classes and subclasses and things like that.
1:02:18I might want that in production setting if I, you know, optimize for maintenance or, you know, clarity, organization or whatnot, or reusability. But for teaching, you want the code that people you're talking about to be right there, you know, like not pages above or in a different module. And so ChatGPT and Gemini and so on tend to optimize in the kind of code they've been trained on, which is usually the professional kind of code, which is organized into modules. And so I find it very verbose. It adds comments that are like multi-line long and so on. So the code you get in the book is definitely initiated many times by ChatGPT or Cloud or other, but then I just worked it out and reorganize it in many ways.
1:03:02Yeah, so I got some people asking me, actually, regarding the code, why don't you organize it into modules? That's the reason. It's that if you try to organize it into module, first, it bloats up the book, just becomes much bigger. And also, you need to sort of cross-reference, go back to page this for the function that, and that makes it a little bit harder, I think, to teach. So my code is unashamedly super flat, and there's one function that does one thing, or not even a function, just plain code like that. So to be fair, though, with PyTorch, since it doesn't have a training function, every notebook has to have a training function, right?
1:03:43So either I use Lightning, maybe I should have, or I need to have a training function in every notebook, which is the option I chose. So it's not a very long function, but that's the one piece of code that you don't have in front of you that I regularly reference. And I say, okay, well, now train the model using the function we defined in this chapter. Not the most satisfying in my opinion, but the PyTorch just doesn't have a training function.
1:04:11Jon Krohn:One follow-up question on that one is a bit different. It's like in these terms when we are fully digital, we can ask chat GPT any question. We still prefer books. And do you think in the future, AI should be more than intelligent so that people can adopt it fully? Should be more intelligent, you say? More than intelligent, like something else as well should be there as an add-on? Like intelligence, it's not enough? When you say more than intelligent, do you mean something like... What would you mean more than intelligent? Maybe something XYZ factor that could help people elaborate a bit. Emotions.
1:04:51Oh, like emotions. All right. I see.
1:04:53Jon Krohn:I wanted that as an answer. If you think emotion is... Yeah. Why didn't I think about emotions? Yeah. I mean, if you look at why we have emotions, like why do we have emotions? And there are evolutionary reasons. Why would you feel fear? Well, because those that didn't just didn't survive when the big animals attacked. And why do you feel love or friendship? Well, I guess it helps bond groups in which we evolved. And so these are useful features to have, you know, emotions. Some, I guess, can get overwhelming or detrimental, but on average, I guess they've been good since we have them. And so if they're on average good, well, it sort of makes sense that we might want that for AIs as well, if only to be able to better communicate with them.
1:05:46So I think Miriam here is researching use of LLMs or AIs for psychology, so to help patients. I think you need some kind of empathy to be visible. So I think these are internal states. If you think of sort of in a cold way of what emotions are, It's like an internal state that affects how you reason, how you talk and what you say. And that it has some persistence over time, right? You're not like mad and all of a sudden happy and all of a sudden sad. And there's some persistence to it. Maybe we need a little bit of that in conversation so that there's continuity and more empathy. So, yeah, I guess that makes sense.
1:06:34I suspect that the current AIs, at least I think they display a little bit of emotions. Not that they feel them. I have no idea what they feel, but they've been trained on a lot of data where people are, you know, being nice or do get upset. And so it's possible that they've internalized some of this during training. And when they speak to you, depending on how you respond, they'll internally have some state that might correspond to, OK, like that's suspicious. Like if you ask it, how do I build a bomb internally? it might have a state say, whoa, whoa, whoa, you know, like it's scared or what we would interpret as scared, meaning it's going to be more careful in its next responses.
1:07:15So sort of what you could call an emotion, I guess. First of all, thank you very much. It's a fantastic podcast. Thank you. And I'm not AI.
1:07:24Jon Krohn:My question is, so there's now a lot of company investing a lot into developing their own AI model, right? So it's like Matter, Google, you know, XAI, they're all spending a lot. And now Amazon is like, you know, putting a huge money into the capital investment. What's your take on the, I guess, it's like the final outcome of this competition? Do you think that it will be like a winner-take-all situation or would that become more like a commodities, you know, like power company, which one is cheaper and I'll go for that AI model. What's your take on that? Oh, wow. So are you asking if I understand correctly, all of these companies are going to this field.
1:08:09Is it going to be winner take all or will it be shared? Is that right? Yeah, so I wish I knew. I feel like if you look forward and you imagine, wow, we have a future, how many years? It might be 10 years or more. But one day there's an AGI and we're all out of job because the AIs have replaced us, and we need some source of income. Maybe everything has become much cheaper and we can live on very little money, but we still need to have some income. So whether it's universal income or whatever, I don't know. One option would be if you've invested in the right companies, and they're growing because they'll pay less for their employees because the employees will be AIs and earn more money.
1:08:53And so if you've invested in the right companies today, you might be rich and just live off that later on. So knowing which ones to invest in would be great. And so the thing is, I wish I knew. And if you don't know, then your best bet is probably to hedge your bet and just invest a little bit in all of them, at least if you have some funds to put in there. I personally think that a lot of these companies are gonna die. I'm more on the winner-take-all side, or you'll have at least some specialization, But I don't see, like, once one company gets AGI, it's sort of a runaway effect that would amplify their advance, especially if we're not in a very open world and they don't share, you know, their successes or the reason for their success.
1:09:41So I'm more on the, you know, winner take all side. But since we don't know which one will be, well, today all you can do is sort of invest in every one of them and you just need one to succeed and you'll be happy. you can probably expect if you think of where to invest I'm not by the way I'm not a financial expert whatsoever but my thought is that all the companies that currently need to pay a lot of people to think and like Google for example need a lot of people I think they have like 50 ,000 or so employees who are actually pretty expensive employees, and what they do is think. Well, if you can replace them, then Google will be tremendously rich just by replacing their humans with machines, right?
1:10:31But this doesn't have to be a tech company. Any company that relies on a lot of human brains, like financial institutions, you might want to invest in banks, right? Because they have a lot of brain power, and those could be eventually replaced by an AI. So I guess in terms of investment, It seems to me that if you don't like if the world in the future is this horrible world where, you know, half of the people or more just don't have any kind of income at all. And the other ones just live off the the dividends from their investments in companies. You probably want to be in the latter group. Right.
1:11:04And if you're in the latter group, you probably want to have chosen the right companies. And I think like a good bet is any company that depends on intelligence would probably go up, whether it's in tech or not. But that's like, as I said, not a financial advice. Invest at your own risk. Pass results do not guarantee future performance. Exactly.
1:11:24Jon Krohn:Nice. All right. I think you can just keep passing the mic back. Oh, yeah. It's just going to kind of get, yeah. We're not going to be able to do everyone. We probably have time for a couple more. Hi. Thanks for your talk. So I think you mentioned that the models are kind of plateauing in terms of performance due to a limitation of training data. so I happen to have worked for a company which I will not name where I just spent hundreds of hours labeling data for large language models and recently the tasks have become so difficult that I can't even do them and my question is how are we going to get enough high quality expert data or is that kind of just the bottleneck that or the ceiling that we've hit and LLMs will just like not improve much at all?
1:12:06Jon Krohn:That's an interesting idea that I hadn't thought of maybe machines can't become more intelligent than whatever the training data we can create is. Yeah, that's a great point. So the way I would answer that is that if you go back centuries, people didn't know everything we know. And somehow we reached today with a lot more knowledge. And it's not like we made the training data. We built it. We looked for questions. We created, We devised experiments. We got the answers. We modeled and so on. So it's a scientific process, really. And I don't think there's something limiting an AI to actually sort of make its own hypotheses once it has a high level thinking, which is sort of the bottleneck I was alluding to earlier.
1:12:51I think once it's able to sort of reason at a higher level, it's definitely going to be able to formulate hypotheses on, oh, I think the world works this way or that way. What do I need in order to check that? that I need to run this experiment, or maybe I don't need an experiment. Like for example, in mathematics, it can probably just figure out everything on its own. It'll run experiments in the sense that it'll run maybe programs to check, oh, is there a prime number of this and that between this number? So it might run experiments, but it can just generate them on the fly. In some cases for physics research or maybe chemistry, you actually need physical experiments to answer questions.
1:13:29Even if you have a superhuman AI, it won't all of a sudden know whether there are multiverses or whatever. It needs to actually run experiments. So there are bottlenecks, but I don't think that the lack of data is something that will sort of plateau them forever. They'll generate experiments that will create data for them. You think there'll be some kind of takeoff point where we don't need so huge amounts of expert knowledge anymore? Is that what you're saying? Yeah, so exactly. I don't think the training data will be something that will sort of block progress. Right now it is because the training method is just basically ingesting all that data.
1:14:13And if you don't have expert data, well, it just doesn't know. It can't figure out on its own. That said, you probably know that DeepMind has worked on a lot of things where it actually generated new knowledge. right? If only for, you know, AlphaFold, which generated new knowledge on how proteins are made. It figured out the logic between the sequence of, you know, DNA and protein sequence and the actual shape in 3D. So it figured it out. So it could do research, right, eventually. And if it can do research, it can generate new knowledge, new understanding. And with the data that's out there, I think the AI is just with the current data, it could be much, much smarter than it is right now.
1:14:57Just, I mean, it has access to pretty much all the world's knowledge. Why is it still dumb? Right. So it could do much better with what it already has. And then once it has sort of optimized what it has, it can generate more. So I don't see this as a, it's something that can slow down now, but I don't see this as a blocking point in the future. Yeah. Thank you. I hope I can do all my research for me in the future. Yeah, I hope so. Does superintelligence necessarily mean supercapability? And, you know, if it's superintelligent, is it generally posing significant risks without having equivalent capability?
1:15:32Yeah, yeah. No, I, like, fully agree that once you have, like, if I could snap a finger and my laptop all of a sudden has a superintelligence, it doesn't all of a sudden change the world, right? You need sometimes because there are bottlenecks in real life. I think Francois Chollet recently put it like that, that the world is mostly made of humans, and humans are slow in many ways. And so if, for example, the super AI all of a sudden decides to, I don't know, build a thousand robots, well, they still need to be made physically. And so that takes time. And so there are a number of bottlenecks, but there are also a number of domains where the bottlenecks are limited.
1:16:15And it's important to see what are the factors of speed up that you get. And those really depend on the domain. If you're thinking of physics research, then yeah, we're probably bounded a lot by the experiments and what we can do in the real world. We're not out of theories. There's so many theories from string theory and all that, all sorts of variations for the last 30 years, and it's yet not making a huge amount of progress. And it's not for lack of intelligence, I don't think. So in that domain, maybe there will be slow progress, right? And superintelligence won't all of a sudden change the world.
1:16:49But in other domains, and particularly like mathematics, which can have a tremendous impact, I think it can go much faster than humans, right? And there's no reason why it can't also self-improve, right? So improve the algorithms for training, improve the algorithms for once it's intelligent enough, you can do our job and just improve it again. And this also, where's the bottleneck? You know, there's the training loop, so you need data centers, but they're being built, you know, fast. Plus there is like financial incentives for people to build more because of the payout. So I don't see that as like it definitely a slowdown, like it's a bottleneck.
1:17:22Don't get me wrong. But it's not like it doesn't have a huge world impact fairly quickly, at least at the human scale. I mean, we're incapable, apparently, collectively of dealing with climate change, even though we have like 50 years ahead of us or we had. But so now we're like something even more world changing, which will not take 50 years. It would be much faster. So I think the disruption level is gigantic. And I agree with you there are bottlenecks, but it's like, sure, it's not like, I'm not like in the Kurtz, what's his name again? Kurtzweil. Thank you. Kurtzweil, a camp of this explosion in terms of months or even weeks or seconds.
1:18:05There are physical, real-world limitations, but at one point they become irrelevant compared to the disruption that you get at the world scale. And I think your second point was regarding the paperclip. And so this paperclip thing, for those who don't know, is like this very standard example of a stupid objective that for some reason this AI has and then tries to maximize and then everything that goes from it. I think what the thought experiment tries to show is that no matter what your final objective is, your ultimate objective is, sub-goals are not controllable or are hard to control. If the AI is very, very intelligent, whatever its goal is, and maybe its objective is to make humanity happy, right?
1:18:51So its objective is to make humanity happy. And it's like, okay, what do I understand by happy? and that there's you know maybe it's it's analyzing how our brain works and saying okay full happiness is when you're like very content physically and emotionally and this and that and maybe it realizes that by putting us in this particular state or in these you know it doesn't need necessarily variety or it could be an illusion of variety it's not necessarily what we would want yes where i'm coming from is having that sort of narrow perspective isn't necessarily associated with intelligence in my view.
1:19:26So being intelligent is having a more broad perspective. So having a solo fixation on one goal with the exclusion of everything else doesn't align in my mind with some super intelligence. Yeah, that's okay. I get you. I mean, there was one excellent, really excellent paper that I really loved in reinforcement learning, which was about curiosity-driven reinforcement learning. I thought it was mind-blowing that this thing was not ever taught that a particular video game, you know, either the rules or even the rewards, it never saw how many points it got, and it wasn't told you need to get points.
1:20:00All it was taught was don't get bored, right? Whatever you do, try to find novelty and be curious. And if you just stay and do nothing, it's boring. So the AI automatically started to move around. And if you run into an enemy and die, you go back to the beginning, which you've already seen, so it's incredibly boring. So it sort of automatically learns to avoid the enemies and just explore the world. And if it reaches a place where there's fake novelty, like random noise, then it will try to look at that noise and make sense of it, but it eventually gets bored of that, like if it can't make sense, at least in other variants.
1:20:39And so just by curiosity, you can sort of generate what possibly is super intelligent, like a behavior that you and I would probably associate more with true intelligence. but you know like the way the AI is coded no matter like even if it's unknowable to us and to it it has some kind of objective the way LLMs are trained they try to predict the next word right so you could say that's its final objective but it's kind of a narrow summary because evolution for example has the sole objective of reproducing and preserving you know your genes yet all of that incredible diversity of life has emerged from this very simple rule.
1:21:23So you could imagine that you have this very simple objective of predicting the next word, and currently LLM is trained at scale with this very simple objective, have developed internal representations that make it seem pretty intelligent, even if it's not human level intelligence, pretty smart. So it generates a lot of intelligence. You reach a point where this AI has some internal objective that maybe itself doesn't know and we don't know, but it probably has, you know, if you could step out and look what it's trying to optimize, maybe you could know that it's, okay, it's optimizing for this.
1:21:54Like maybe what you're implicitly saying is that we humans try, have in our objectives, the goal of stepping back and understanding and being curious. You know, there are a certain number of things that we try to optimize, and maybe that explains a lot of our behavior in the end, right? And these AIs, we don't know what their current objective is other than complete the sentence and predict the next word. I don't really know what has emerged from it. And for all I know, apparently there have been tests that show that self-preservation and lying and blackmailing are part of the sub-objectives. Definitely not the final objectives, but those emerged.
1:22:39And that's pretty scary to me. Thank you. I'm sure we're out of time. Yeah, that was a long rambling, but I hope it made some kind of sense.
1:22:47Jon Krohn:Yeah, if you could maybe keep the questions simpler, that would be appreciated. Yeah, questions that would have a yes, no answer. No, this was fantastic. All the questions were so great. Really appreciate you pitching in. And yeah, so just before I let you go, Aurelian, there's two quick questions. They might even potentially be one word answers. But most guests take more than one word. Just every guest, I ask the same two questions. And so the penultimate question is, do you have a book recommendation for us other than your own book, Hands on Machine Learning? Oh, good question. Does it have to be machine learning?
1:23:28Jon Krohn:No, it doesn't. It doesn't. We have actually had champagne books recommended before. Oh, yeah. And I know that that's something that interests you. That's such a great question. Well, my personal pet peeve outside of machine learning is biology and evolution. And one of the most transformative books I've read in that space is The Selfish Gene by Richard Dawkins. Oh, yeah. Oh, my goodness. It's so good. So if I were to recommend one book, it would be it. Like, it sort of makes sense for evolution. Did you say pet peeve? Oh, is that a word? It is, but it means kind of the opposite. it's like the thing that annoys you the most.
1:24:10Jon Krohn:Oh, no, no, no. Oh, is it? Okay. So reverse that. You can tell I'm French, right? Yeah, no. So it's my... Oh, so what's the opposite? No one expects you to be French because you're like Kiwi accent. So yeah, I guess it makes it a mess. Passion? Oh, passion, passion. Yeah, yeah. Okay, passion. Yeah, that's one book I really absolutely adore. And many, many books by Stephen Jaygold as well. Like anything on evolution really is just mind-blowing to me. On machine learning, I would actually humbly point to some of my competitors. Rashka's book, I find it really good. Sebastian Rashka. Yeah, and also Francois Cholet's book are really good.
1:24:51I have a lot of admiration for both guys, so I'd recommend these books a lot. Nice. Sebastian Rashka has actually written a few now. Yeah, he has. A bunch come out in the past year. They're really good. I remember reading it anxiously after my books had come out, and I'm like, oh, my God, and I hope it's not too good. And I was like, oh no, it's really good. But the good news is we don't have the same approach. And I think Francois Cholet's book is also fairly different in its style and approach. And so like, yeah, if you can have all of them on your collection, it's all good.
1:25:25Jon Krohn:Fantastic. And then final question is easier. I don't think this one is going to stump you so much. For people who want to continue to get your brilliant thoughts after listening to you speak today, other than your book, which is obvious, hands-on machine learning. Of you and all your competitors, it is one that people pick up in the bookstore the most. So hands-on machine learning, obviously an option. But how else can people follow you? That's a good question, yeah. Well, it's tricky, isn't it? I started a YouTube channel a while back and then sort of lost motivation, so there's not much on it.
1:25:59I might come back to it. So I have a YouTube channel if you want to subscribe and maybe one day see some videos with any luck. I used to be on Twitter quite a bit. For some political reasons, I kind of stopped. So I'm kind of nowhere. I'm on LinkedIn, and I sort of accept anyone. So you're welcome to contact me on LinkedIn. I just don't look at it very often. So yeah, I'm sort of discreet.
1:26:27Jon Krohn:We're delighted that you're spending, instead of spending the time tweeting, you are writing fantastic books. It's fabulous. Thank you so much. Thank you. Thank you very much.
1:27:06Jon Krohn:for writing the best-selling ML book of all time, why his next book release will feature PyTorch instead of TensorFlow, and his insightful thoughts on the coming AGI revolution. As always, you can get all the show notes, including the transcript for this episode, the video recording, any materials mentioned on the show, the URLs for Aurelian's social media profiles, as well as my own at superdatascience.com slash 919. Thanks, of course, to everyone on the Super Data Science podcast team, our podcast manager, Sonja Breivich, media editor Mario Pombo, partnerships manager Natalie Zajski, researcher Serge Massis, writer Dr.
1:27:41Jon Krohn:Zara Karche, and our founder Kirill Aromenko. Thanks to all of them for producing another excellent episode for us today. For enabling that super team to create this free podcast for you, we are so grateful to our sponsors. You can support the show by checking out our sponsors' links in the show notes. And if you'd like to sponsor the show yourself, you can see how to do that at johnkrone.com slash podcast. otherwise share, review, subscribe but most importantly please just keep on tuning in I'm so grateful to have you listening and I hope I can continue to make episodes you love for years and years to come until next time keep on rocking it out there and I'm looking forward to enjoying another round of the Super Data Science Podcast with you very soon
From the publisher
PyTorch, AGI, and the future of alignment research: Aurélien Géron joins Jon Krohn in this live interview to talk about the fourth edition of his bestselling Hands-On Machine Learning as well as what superintelligence makes him hopeful for, as well as what concerns him about machines surpassing human intelligence.
This episode is brought to you by Gurobi and by the Dell AI Factory with NVIDIA
Additional materials: www.superdatascience.com/919
Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information.
In this episode you will learn:
(02:04) Why Aurélien wrote Hands-On Machine Learning
(20:54) How Aurélien came to decide on material for the new edition
(28:53) Aurélien’s predictions for AGI
(51:21) How to support alignment research
(1:13:42) Does superintelligence mean super-capability




