#248 Pedro Domingos: How Connectionism Is Reshaping the Future of Machine Learning

17 Apr 2025 · 1 h

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Notes: Eye On A.I. - Episode #248: Pedro Domingos: How Connectionism Is Reshaping the Future of Machine Learning

Episode Overview

  • Host: Craig S. Smith
  • Guest: Pedro Domingos, renowned AI researcher and author of *The Master Algorithm*.
  • Theme: A deep dive into Connectionism, neural networks, and the evolution of machine learning from historical and current perspectives.

Key Topics Discussed

  1. Historical Context of Neural Networks
  2. Originating in the 1940s, neural networks have been foundational in AI.
  3. The initial excitement led to disillusionment due to shortcomings highlighted by Minsky and Papert in the late 1960s.
  4. The resurgence in the 1980s thanks to the backpropagation algorithm.
  1. Backpropagation
  2. A pivotal algorithm enabling neural networks to learn from errors.
  3. Simplifies credit assignment by adjusting weights based on errors, leading to effective learning across layers.
  1. Key Breakthroughs in Machine Learning
  2. AlexNet: Revolutionized computer vision in 2012 by leveraging GPU power and large datasets.
  3. Transformers: Introduced attention mechanisms that improved machine translation and other NLP tasks, outperforming traditional RNN models.
  1. Tribal Wars in AI
  2. Connectionists vs. Symbolists: Ongoing ideological battle, with each group advocating different approaches to AI.
  3. Domingos emphasizes the need to explore various schools of thought rather than focusing narrowly on one.
  1. Generative Models
  2. Discussion on generative vs. discriminative models, where discriminative models dominate due to their effectiveness.
  3. GANs (Generative Adversarial Networks): Initially received significant attention but faced stability issues, leading to diminished excitement.
  1. Reinforcement Learning (RL)
  2. Explored the principles behind RL and its application challenges.
  3. Domingos expresses skepticism over the recent resurgence of RL claims, suggesting it often functions similarly to supervised learning.
  1. Future Directions in AI Research
  2. Domingos calls for a broadening of research focus beyond predominant trends, advocating for diverse methodologies.
  3. Warns against the dangers of narrowing research too soon around dominant models or paradigms.

Key Takeaways

  • Connectionism’s Role: Connectionist approaches, particularly neural networks, play a crucial role in the evolution of AI, underpinning many current advancements.
  • Historical Lessons: Understanding the historical context and previous patterns of excitement and disappointment in AI can inform current research directions.
  • Importance of Diversity in Research: A healthy AI research ecosystem requires exploration beyond popular paradigms to prevent stagnation and foster innovation.
  • Caution Against Over-Optimism: The AI community should be wary of overhyping breakthroughs without substantial evidence of their efficacy and robustness.

Conclusion The episode provides an insightful analysis of the connectionist approach within AI and machine learning, with Pedro Domingos advocating for a multifaceted exploration of ideas in the field. The discussion serves as a reminder of the historical cycles in AI research and the importance of maintaining a broad perspective as the technology continues to evolve.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00I see a bunch of data, I'm going to postulate that they were generated by some process and I'm going to try to identify what that process was. This is actually what generative learning is. So you come up with this model like a Bayesian network that says like, well, this is how you generate patients and their symptoms. First, you decide which disease they have. And then you decide which symptoms they have. What fever are they going to have based on the fact that they have COVID, et cetera, et cetera. You generate the data. And this is why it's called the generative model in opposition to discriminative models, which is what most of machine learning is about, which is to just go say like, okay, I'm going to go forward from this data and try to decide, you know, what you suffer from or whether it's a cat or a dog, as opposed to like, how do I generate, you know, a cat or a dog, right?

0:40So there's always been this thing in machine learning between generative and discriminative. Discriminative has always dominated and really still does because it just works. It's an easier problem and it works better. When you need somebody, you need somebody. You want to hire people fast. The longer you wait, the more the market drifts, the more your needs drift. For example, you just realized your business needed to hire somebody yesterday. How can you find the right candidate fast? Just use Indeed. When it comes to hiring, Indeed is all you need. Stop struggling to get your job posts seen on other job sites.

1:18Indeed's sponsored jobs help you stand out and hire fast. With sponsored jobs, your job postings jump to the top of the page for your relevant candidates so you can reach the people you want faster. And it makes a huge difference. According to Indeed data, sponsored jobs posted directly on Indeed have 45 % more applications than non-sponsored jobs. One of the things I love about Indeed is that it makes hiring so fast. I've searched for candidates before and it can take months, particularly if you're just working through word of mouth. In those days, I didn't have Indeed. Plus, with Indeed's sponsored jobs, there's no monthly subscriptions, no long-term contracts, and you only pay for results.

2:06How fast is Indeed? In the minute I've been talking to you, 23 hires were made on Indeed, according to Indeed data, worldwide. So there's no need to win any longer, speed up your hiring right now with Indeed. And listeners on this show will get a$75 sponsored job credit to get their jobs more visibility at indeed.com slash IonAI. Just go to indeed.com slash IonAI. That's indeed, I-N-D-E-D dot com slash IonAI. All run together, E-Y-E-O-N-A-I. Go right now and support our show by saying you heard about Indeed on this podcast. Indeed.com slash IonAI. Terms and conditions apply. Hiring Indeed is all you need.

3:02The biggest surprise about the Exit for me was how surprised people were. There's really nothing particularly surprising about anything they did. Same with what the others are doing. It's the natural progress. And a lot of it, honestly, is adding more capabilities and tweaks and variations that are all useful and understandable, but on top of a core that isn't changing that much. And honestly, I think we're converging, as is often the case, to a local optimum. I personally am more interested in what the global optimum is than the local one. But in the short term, that's where we are. Yeah. It's kind of what happened with supervised learning.

3:45It was, everyone was amazed. And then a lot of people started focusing on optimizing and tweaking. And it became stronger and stronger. Is that right? When you say supervised learning, what do you mean exactly? Yeah, I mean the labeled data pattern recognition systems that dominated neural nets before generative AI. Oh, I see. Interesting. Well, I'll tell you a very short story. When I was in grad school, this was in the early 90s, one of my colleagues of the same advisor, he left because he said there was no future in machine learning because it was all, you know, the only thing that worked was supervised learning and supervised learning.

4:43There was nothing new left to happen. Well, we can see how that turned out, right? I don't think, I mean, I see what you're saying, Okay, but I was supervised learning is such a broad paradigm in machine learning that it's there are there are there are certainly and have been local optimal within supervised learning. It's hard to see supervised learning. I mean, I guess in the current context, people talk about like, oh, we're running out of data and so on and so forth, and we need to do other things. And that's where the lot of work is. But the problem is that they're not getting the most out of the data, is that the supervised learning algorithms are not that good.

5:28And in fact, the key idea in some ways of LLMs, maybe we'll talk about this, is to take advantage of data.

5:39The problem with supervised, it continues to be the case that the only thing that really works in machine learning is supervised learning. You know, it's, yeah, that reinforcement learning and even unsupervised learning, blah, blah, blah. The key innovation in LLMs in many ways is just, and it's not a thing that's new to LLMs, but it's what drives it, is how to turn a massive amount of unsupervised data into supervised. Yeah. So, you know, under all of this stuff, they're still supervised learning and gone, and moreover, and related to DeepSeek. People are always talking about, these days I've become, you know, honestly a little skeptical and just like, oh, we have our latest, greatest reinforcement learning, blah, blah, blah.

6:19And when you look at it up close, it's really just doing supervised learning. But it's easier to call it reinforcement learning. So there you go. Yeah. And so anyway, let's get on with the connectionists. One thing that interests me is how, I mean, and I think I mentioned this in this whole series, is that there are a lot of strategies or schools or tribes, as you call them, that remain viable for progress and continued research. I had on the podcast a couple of weeks ago, Sep Hockrider. I'm not sure how to pronounce it. When he pronounced it, it doesn't sound like that. The guy who came up with the long short-term memory networks, LSTM.

7:21And he is continuing research and has now created a company to use his research in industrial applications. And I thought that was really interesting. so yeah let's talk about the connectionists where they came from where they're going and whether there are schools of of connectionists theory that are being neglected because the bright shiny objects is capturing everyone's imagination. Yeah, sure. So go ahead. And Pedro, reintroduce yourself. I apologize for that. But yeah, go ahead. I'm Pedro Domingos. I'm a professor of computer science at the University of Washington. I have been an AR researcher since the 90s.

8:19My specialty is machine learning. I've worked in a lot of different areas of AI along the different paradigms and also in a lot on unifying them. I'm probably best known as the author of a book called The Master Algorithm, which is an introduction to machine learning for a broad audience, and more recently of 2040, a Silicon Valley satire, which is basically a spoof of AI hype and fear and the tech world. Yeah. So tell us about connectionists, how they emerged and how they operate alongside the other schools of thought that we've talked about and how they relate them to what's going on now. Alex Blanche - Connectionism is one of the oldest paradigms in AI.

9:09You could say it's the oldest after symbolic AI. It started in the 40s, which is a long time ago in AI terms. And you could say it's really one of the two big ones along with symbolic AI. Sometimes people when I think over there, I think of like these two big opposing groups, and they are the symbolists and the connectionists. And in many ways, the connectionists historically have been to find themselves in opposition to the symbolists. Now, one interesting thing that has happened is, so the The idea in connectionism is actually a very appealing one. And that appeal, I think, is very important, which is, okay, we want to build intelligent machines.

9:51What do we know that's intelligent? It's humans, and in particular, the human brain. So like good engineers, when you're behind the competition, what you do is you reverse engineer it. You open it up, right? You open up the car, right? whatever your toyota you start by looking at your gm or ford cars and go like well how is this thing built right and then you copy it and then maybe eventually like toyota did you improve it and so on but that's the idea let's reverse engineer the brain and so this started from the very beginning of the modern age of computer science people were there was this paper that came out in the 40s by mcculloch and pitts on a model of neurons you could actually say that was the paper that started neural networks.

10:33And a lot of people don't know is that a lot of the computing hardware today that has nothing to do with AI a lot, all of it, is really derived from that paper. So in some sense, every computer, as they used to call them in the 50s, is an electronic brain. Every circuit in a computer is imitating neurons. A very important point to remember as AI starts to gobble up the rest of computer science. But the big difference between those neurons and the ones in your brain is that those neurons didn't learn. They didn't have weights. They were really just logic gates, which is again, what computers are made of.

11:08It was a generalization of the logic gates we tend to use. And then really the key name in this story was Frank Rosenblatt, who in the 50s created this thing called the perceptron. That was the first neural network that learned, but it wasn't a neural network. It was actually just a single neuron that learned with some pre-computed features that were fed into it. And people got incredibly excited about this at the time. The New York Times had a front page story saying, you know, human level intelligence is coming soon. I can't help laughing at this because, you know, it continues to happen today.

11:47One day it will be true, but in the meantime, people should know this history. And so for a while, it was very popular. And you could even say that was probably the most popular approach to AI for a while. And in fact, it's interesting that a lot of people from the other paradigms, they started out doing neural networks. John Holland of Evolutionary Computing and Marvin Minsky of Symbolic AI. Marvin Minsky's PhD thesis was on neural networks, which boggles the mind. But then they all drifted away from it once that initial seductive idea was not really panning out very well. And then also very famously in 1969, Marvin Minsky, who was by now one of the leaders of the symbolic school and Seymour Papad wrote a book called Perceptrons, basically tearing down the claims that the Perceptron was a great thing.

12:41They showed mathematically a series of things that Perceptron just couldn't do. And then interesting Perceptrons just in neural networks just fell like a stone. And then for the next 15 years, they were dead. And in fact, they even reflected on machine learning in general. The consensus during the 70s and early 80s was that machine learning doesn't really work. It's too hard. You've got to program things. You've got to do knowledge engineering. And that was the big thing. It was symbolic AI and knowledge engineering, and machine learning was nowhere. And then in the 80s, they came back. And the key development was was the the back propagation algorithm people knew and minsk and pepper said that like if you could train networks with multiple layers then all those limitations would be solved but nobody knew how to and in their words they couldn't see how it would ever be done and back prop itself has an interesting history but it was in the 80s that it took off and then there was another moment of great excitement which then again died down again because people thought like we're just going to build deeper and deeper layers and pretty soon, you know, it'll be the brain, right?

13:49They got over optimistic again. There's always somebody in AI being over optimistic. The only question is who, but then, you know, the problem with back prop is that it did, it only worked with one hidden layer, which wasn't much. And so people gave up on it. And by the end of the nineties, neural networks were kind of not dead, but like completely sidelined again. And, and that's when, you know, early 2000s is when, uh, you know, people like Jeff Hinton, Jan LeCun and Josiah Benjo started working on this. Actually, they never stopped working. They were the diehards that never stopped working, but they started to figure out ways to learn deeper networks.

14:25And in fact, the term deep learning, partly it's a marketing term. It's clever, right? Whoa, the learning is deep. But really technically just refers to the fact of having networks with many hidden layers, which indeed, once you're able to learn them, can do amazing things and and here we are today yeah can you just take a dog leg on uh backprop and explain uh where it came from and how uh hinton applied it yeah so first of all what is the difficulty that backprop is solving it's actually one of the central problems in machine learning so in the master i'm going to talk about the five tribes and how each one of them has an algorithm, its own master algorithm, that solves a key problem.

15:09And my contention, of course, that you have to solve all of them, so none of the algorithms is enough. But the key problem, and it is a central problem in machine learning that Backprop solves, is credit assignment. I have a very big system with lots of layers, lots of neurons, a big mess of connections, and it's supposed to look at an image and tell me if it's a cat or a dog. And the image it's looking at is a dog, but it says it's a cat, right? And so now the question is who do I blame? Maybe it should be called the blame assignment problem. Like who should change, right? In connectionism, all the knowledge is in the connections between the neurons, again, inspired by the human brain where, you know, as far as we know, all the knowledge you learn is in the strengths of the connections between neurons.

15:52Which connections should I change? And if you think about it, this is not obvious at all, right? There's a mistake at the end, there's an image at the beginning, like, you know, who the hell knows what's going on? but backprop technicalities aside just gives a super simple answer to this and again often machine learning super simple answers are what gets you the farthest and the answer is this what i'm going to do conceptually not in practice because that would be too inefficient but conceptually is i'm going to tweak each weight in turn a little bit i'm going to take every weight of every neuron in every way and say like let me increase you a little bit you know and there's usually an output function that is continuous that says like, what is your probability, let's say, loosely speaking, of being a cat?

16:38Oh, when I tweak this, when I move this weight up, the probability goes up a little bit. Good. So let me keep that. When I tweak this other weight down, the probability goes up. So let me tweak that, right? If you do this for every single neuron over and over again, over time, amazing things happen. It will probably learn to classify the the cat image is really being a cat image and so on. And, you know, backprop is really just a computationally efficient way, a clever way, but, you know, of a type that we know very well of trying to not be too wasteful. Like you don't want to redo the whole computation every time you do this.

17:14That would just be ridiculous, particularly with today's networks. So it's really just an efficient way to do that. What happened with backprop there again is one of those delicious ironies that the history of AI is full of is it took off. There was this paper published in 1986 by David Romerhart, Jeff Hinton, and Chris Williams or somebody Williams, Robert Williams. It was a third author. And it's often credit. It was originally credit, let's say, with, you know, that's when they discovered backdrop. The truth is that big backdrop was discovered 20 times before that. As it became more prominent, I mean, I alone know three or four different people who invented backprop.

17:56One of them was Yann LeCun. While a grad student in France, he actually invented black pop on his own. And I know, you know, for example, there was a postdoc of my advisor at UC Irvine, who in like early eighties or even late nineties, he submitted the paper to the Top AI conference with backprop and the paper was rejected. And the reviewer said like, what are you doing? Minsky and Pep already showed that. Yeah. And then there was this guy, economist from Harvard, who was, you know, had actually found it in, he published, you know, a paper or a thesis, I think about it in the 70s. Turns out that there was a paper by control theorists, Bryson Ho, in the 60s, before the Perceptron's book was published, the backprop algorithm already existed.

18:44And it's like Jan LeCun says that really, uh the credit uh for back prop should be given to light knits because it's really just the chain rule of calculus yeah so there you go yeah one question i have about by prop is i understand the the theory um that that you keep adjusting the weights going back down through the weights to get closer and closer to the target. But is it done sequentially? Like each neural pathway is adjusted and then the output's assessed or is it done incrementally through all the neurons? And then how does that work? It's done sequentially, but back through the layers. And in fact, that's where the name comes from.

19:39so the way things work in backdrop is that you have your image of the cat and first you do a forward pass right sometimes also called the inference pass which is basically just computing the output first of the neurons that i write on top of the image and then the next layer all the way to the final output that says cat or dog and then comes the error back propagation phase where you measure the error like you know your output was you know 0.7 for cat but it should have been 0.9 right or whatever it was 0.1 for catch at 0.9 so now what i'm going to see is i can look at the last layer of neurons starting with the very last neuron in that case and see like if i tweak every one of your weights uh how does that change things but then the key is that when i go to the previous layer actually don't start from scratch that would be wasteful i don't go all the way back to the output i say like well now i know how much the neurons the weights in the last layer how much difference each one of them makes.

20:31Based on that, let me see how much difference each neuron in the previous layer makes. So you just go back through the layers all the way to the initial one, and that's why it's called backpropagation, because you're propagating the errors backward, and then you're learning from that. Yeah. Okay, so we've got backprop. Jeff and Ilya and Alex create AlexNet in 2012 which takes advantage of the larger data sets available for and the greater compute power available and they have a breakthrough on the image net competition and that sets off the uh the new uh deep learning craze so where where does it go from there then Well, so sorry to just back up slightly, deep learning in public perception exploded then, but already before that, the first big progress of deep learning that caused people to start paying attention was in speech.

21:38Jeff Hinton, for the most part, has always worked mainly in vision. Yoshua Bengio was the language guy, which we'll probably get to. Jan LeCun was also very much into vision but but at one point you know because that's what they could do in some ways vision is very expensive uh jeff and his students started working on speech and speech was working really well uh you know um jeff had this intern jeff hinton had this uh um student uh natty jately who he calls apparently you know larry page calls him the the 30 million dollar intern and jeff says it's really should be the billion dollar intern he did an internet google one summer where he applied speech, applied deep learning to the speech system that they have.

22:20It was already, as you can imagine, a highly engineered thing. And in one summer, he basically beat the state of the art that these, I don't know, hundreds of engineers had developed. And speech was, again, it's a longstanding area of AI, and it had plateaued for decades. And then suddenly out comes deep learning. It's like, wow, suddenly, this phenomenon that you've seen in a whole bunch of other areas was first seen in speech right and then you know lxnet itself has an interesting uh history but but you're correct uh they they um the vision community has always been skeptical of um deep of neural networks because they didn't think of that as serious vision they all just worked again it's hard to not say this with a smile but for the first 50 years all they were working was digit recognition yeah you know it was like how to recognize digits better because Because again, you know, that's what you could, that's what they had the data for.

23:16And that's what wasn't too expensive. And vision people were like, like, we really couldn't care less about that. That's not vision. Go away. So, and like the Pope of vision is, is Jitendra Malik, who's a professor at Berkeley. And Jeff Hinton at one point got on the phone with him and said, what would it take to persuade you that, you know, neural networks are good for vision? And then, you know, Jitendra said, well, you know, there's this data set called Pascal, which was an important data set back then. But then Hinton talked with his students, came back and said, no, Pascal is too small. Is there a bigger one?

23:51And then Jitin said, well, there is this data set that this person at Stanford, Fei-Fei Li, and his students have developed called ImageNet. That's actually pretty large. There's a million examples. That doesn't sound large today, but back then, the likes of Pascal and other sets were like a few thousand, which sounds ridiculous that anybody would try to solve vision with a few thousand examples. And the key thing in AlexNet was really that they used GPUs. Without GPUs, it was like they couldn't have done it. But the GPUs were available then. And they weren't the first ones to repurpose them for machine learning, but they got this amazing success.

24:27And then to answer your question, what happened was the vision folks initially were skeptical. The people in machine learning were very excited and the folks were just like, they're doing something wrong. This can't be right. They just watched something and they're getting these results aren't real. So they went back and repeated the experiments and, you know, oh my God, the results are real. And then they started applying it. And this is like, you know, the wavefront that continues today. They started applying it to other problems in vision, not just, you know, recognizing, you know, literally cats from dogs.

25:01those are the most frequent things on image net is different types of dogs but other tasks and vision it continued to work really well and then within the space of a few years vision went from having no papers to speak of uh using deep learning to basically everything just uses deep learning uh and and and then of course this this got into the public consciousness and then the attitude of the deep learning folks which you gotta um admire at one level but you know be a little skeptic i like okay we've solved vision what's next let's do language now right let's do machine translation right and i remember talking with yoshua ben joseka i don't know like 2015 and he's saying like yeah we haven't quite matched the statistical machine translation results yet statistical machine transition you know systems at the point again at google and other places were very big and already doing very well in many ways uh you know it was the death of the employment for translators and whatnot, which of course didn't happen.

26:01And that whole phenomenon continues today. But he said, we're almost there. Within a couple of years, they weren't just almost there. They had blown past all of that. And these days that pattern continues with reasoning. It's like, yeah, we solved vision, we solved language, and now we need to solve reasoning. The problem with this is that they haven't really solved vision or language or any of those things. They made a lot of progress, but in the meantime, there's a lot, again, we are stuck in these local optimal optimal that we are going to need to get out of. And to their credit, people like Ian LeCun, right, who again, he's a longstanding neural network vision guy, he's like, no, no, no, this isn't solved.

26:38And it isn't. I mean, if you look at video understanding, which is really the main problem in vision, right, the others are kind of like sub-problems of that. It's very far from solved. No one really knows how the video is standing. And, you know, the new generation of things like transformers or not, people have applied it to that, but, you know, it's an improvement. but all of that continues to be unsolved as his language and as his reasoning even more. Of course, you know, the latest generation is this whole, you know, chat GPT, et cetera, and whatnot, which in some ways is more of a sociological phenomenon than a technical one.

27:12The big novelty in chat GPT was that everybody, you know, was blown away by it and started using it. In terms of research content, there is very little new in chat GPT. you could say, oh, they scale things up some more. But even that is not true. Google made the mistake of not releasing their chatbot. In technical terms, there wasn't anything significantly new in ChatGPT over Lambda, for example, Google's chatbot at the time. And of course, a lot of technical innovation has happened since. But it's interesting that in some ways, there is all this focus of research and industrial attention on a very, very narrow front, even within deep learning, while a lot of the other things, including probably where the big leaps are going to come from, are being neglected.

28:03Yeah. Yeah, can you go back a little bit, because there was another flexion point with the transformer algorithm, and talk about, and this is all within the connectionist camp, right? Can you talk about the development of the Transformer? And on speech, I just want to jump back a little bit. Terry Sanowski was doing something called NetTalk

28:32before Jeff had come out with his vision stuff. Was NetTalk a step towards applying neural networks to speech? Yeah, so NetTalk, back in the late 80s, early 90s, when somebody needed to come up with an example of a neural network success, it was NetTalk. NetTalk was not speech recognition, it was speech synthesis. Right. It was you gave it a text and it read it aloud. And the thing that got people really excited, and again, I think you can listen to this on YouTube even now if you just search for NetTalk, is it started out producing noise. And then as the network learned through successive generations of backprop, it sounded at first like it was babbling like a baby.

29:23This, I think, like that fact that it was babbling like a baby is really what got people. And then the babbling got more coherent. And finally, it was actually doing a pretty good job of synthesizing speech, which this does, it sounds like, what is the big deal? But at the time, it was a big deal. So, you know, so NetTalk, you know, was the poster child of neural networks. But the truth is, there's a useful application that never went anywhere. It wasn't good enough, reliable enough, blah, blah, blah. And commercially, speech recognition was and is much more important, right? And even harder, right?

Read the full transcript

29:53It's like, you get this, it's much harder, you know, if you think about it, like, it's not that hard to do a bad job of synthesizing speech from text. You know what the phonemes are and you string them together. It will sound very stilted, but it's not bad. that speech recognition is a hard problem, right? There's like this noise wave coming at you, all garbled, and people eat up syllables, and you've got to turn that into text. That is really hard. So you're right, NetTalks, it was what people talked about back in 1990. But circa 2005, 10, what Jeff and his students were doing was, again, something that the speech people had to take seriously.

30:31Again, I sometimes joke that the purpose of Machine Learning Conference is to publish application papers before they're good for the application conferences. You publish speech there and vision because as tests of the machine learning algorithms, and then when they mature, they go to their natural language, whatever, conferences, which brings us back to Transformers. And it's exactly, if we go back to the point where I said, and then people said, let's do language. Let's pick up the story there. as i mentioned the main guy of that very small group that had always been interested in language but but he even he he did his phd on i think speech and then kind of didn't work on that for a while because again they got stuck for reasons that we can go into um and then came back and his group started looking at machine translation who is that i'm sorry who who was that so so the uh um attention right is the key thing in transformers right and attention it's an interesting story um it was so yoshua banjo said to his to a new postdoc of his kyun hyun cho and and and someone who was an intern here's the best part of this more people don't know this transformers were not invented at google i mean the name was and the tool architecture was but the key id in Transformers is attention.

31:56I mean, the Transformers paper is called attention is all you need. Right. Because actually the main thing they did was take other stuff out, which is a valid contribution, but attention was invented by an intern in Joshua Benji's group. He wasn't any, you know, great machine learning researcher or anything. He was an intern who didn't know anything and just started working on this and he came up with that idea. So what was the problem? What was the solution? Right. He's like to do machine translation, the way they tried it initially was like you would go through for example the translating french to english you would go through the french text with what is called a recurrent neural network that processes things sequentially over time and you would wind up with with the state of the neural network that captured the text that you had read and then you would start from this state with and this was called the encoding part right because they encoded french into the network and then there was the decoding part where you turn that internal state which is a bunch of numbers that nobody understood and still doesn't and turned it into english again sequentially by generating the text on where that time and this didn't work very well there was a bottleneck like like there was too much information lost and they tended to remember the recent text meaning the last parts of the text but not the early one and so yosha said to them like you know you know we got to do something about this right and and what is the attention idea in a way the attention id is really just doing with neural networks, what people have been doing in the statistical approach to machine translation for decades.

33:25There was always this phase called the alignment phase, which has nothing to do with what people call alignment. It's this fact that, for example, when you're translating fresh English, word order is often different. Like the adjectives will come before or after, right? And so it's not enough to say, I'm going to go through the text one at a time. When I'm generating the next word, let's say the next word is the adjective. Like, for example, I saw a big cat. Right. And then like, where does the big come from? Right. You got to go find in the previous text, the word that big is translating. Oh, Le Grand, Cha.

33:59Okay. Right. I already translated Cha, not Le Grand. Whatever. Right. You get the idea. Right. And so what they did was they did the neural network version of the same process. And this is what attention does. is like now that I'm generating the word, you know, like what is my next word is going to be? Well, what is it the translation of? Right. Oh, you know, big, you know, that's big as the translation of God. And then you can learn this using backprop. Right. So that was that was their innovation. Very important one. Right. But they were just trying to solve a machine translation problem. The paper is just about machine translation.

34:35And then the transformer guys, same thing. They, you know, they refined, you know, a bunch of stuff and basically to cut out stuff that no longer was was longer useful and and and then i mean i it's interesting because like i remember reading these papers at the time and each case thinking this is a very cool idea i bet you can do a lot of other things with it besides translation next paper and then what people found and this is partly you know the the beauty of machine learning is like they took the same architecture and started you know doing more and more things with it and keep and it kept working amazingly right and here we are today right chat gpt the t in gpt is transformed yeah yeah uh and then the guys at google there was a team uh that that as you say stripped out everything uh and made a very simple architecture around attention right um i i'm i'm i'm i'm laughing because it's anything but simple.

35:31But I mean, so it wasn't one team. There were several teams, maybe three. That's why the paper has eight authors. There are three different groups of people who, you know, for one reason or another said like, let's try to use this for whatever, right? Or, you know, they had their different concerns, but they were all started using this new mechanism of attention, right? The basic attention is just, you know, a couple of equations. And again, we can try to understand what it's doing in multiple ways, but they weren't making any headway. And then the key guy was actually Noam Shazir. He's one of the eight authors of the paper, but he's really the guy who invented transformers in that sense, if you take attention for granted.

36:11I mean, again, I've talked to several of these people. They just hacked a million things, as often the case, until something worked. There was no great insight there, at least that they can yeah maybe there is but you know but we don't know what it is yet uh probably but you know noam was the guy who who's like figured out how to make it work right and and then and then and then um it you know i was smiling because he you know the transformer so the recurrent neural networks are are one of the problems is that they work sequentially so they They also have to be trained sequentially, which makes it hard to paralyze them.

36:50And paralyzation is very important so that you can scale up to make things run faster. And initially, from the Montreal group, from Benjo's group, they were just doing attention on top of a recurrent neural network. And then what this guy said, very, very important, right? And again, there's a history in machine learning of these things happening. is like actually once you have attention you can take out the recurrence and now you can just train in parallel and now you can train on much bigger corpora this is really the key uh thing that happened in that paper right now this system that they have at the end of the day is anything have and have even more so today is is anything but simple it's a huge pile of different stuff and hacks left and right and and you know this actually slows down progress right is that it's there's nothing simple about the, I mean, things tend to become simple over time as you understand the better.

37:45And I think that's the case with transformers, but they have all these different kinds of layers and mechanisms and whatnot, which we could go into, but like a transformer, the full architecture is a, is a big pile of complication. And, you know, that, that holds us back. Yeah. How does, I don't mean to jump around, but as I mentioned, and Sap Hochrider, who had developed LSTMs much earlier. How does that relate to attention? Because it's looking back a certain number of tokens and holding them in memory. And that, to me, sounds like attention. Right. So LSTMs were very hot in machine learning for language at one point.

38:43And that was the point at which people are already focusing on language, but attention had not been invented yet. And LSTMs are a type of recurrent neural network. The problem with recurrent neural network that really, for example, made Yoshio Benio stop working on was what is called the vanishing gradients problem. Right. You think about it, when you do back propagation, the further back you go, the more diffused the signal becomes. It's like the neurons at the output layer, yeah, you are to blame. You're clearly wrong, right? Once you get to a neuron like, you know, 100 layers behind, right, it's like there's a million paths to that neuron.

39:20Like, who knows if it's to blame or not, right? It's like if you think of like a drop of ink, you know, diffusing through, you know, a glass of water, right? As time passes, it becomes spread over everything. And if this is the credit or blame that you have to assign, right? Like a recurrent neural network over time, right? Even if it's just one layer, because you use it over and over again, it becomes, you know, an infinite, you know, layer network. And the signal just diffuses through it until it's like, you know, the ink all over the water. You don't, there's no, you can't learn. And there's nothing to learn.

39:52Everybody's equally to blame or not to blame, right? And LSTMs were a solution to this problem. And the solution to this problem, the LSTM, And, you know, people tried a bunch of things at the time, neural Turing machines and blah, blah, blah. But LSDMs, again, which had been invented way before by Sepp and his advisor, Jürgen Schmidhuber, notorious in machine learning for various reasons, their solution was actually to do something that works a little bit like a computer. It has these memory cells, and you can actually write stuff into the memory cell, and then it stays there until you decide to access the cell, which is really like a von Neumann computer that we all use works.

40:32And then what they had was these gates that controlled access to the memory. There was like the forget gate, et cetera. And what it did was it allowed you or didn't allow you to, you know, write into the memory, read the memory, et cetera, et cetera. And again, the key was that there was, this wasn't digital, right? It was something that you could gradually learn, right? Which again is what attention does later. It's like, it doesn't just say like, oh, I'm going to pay attention to this. it's paying attention to all the past, you know, to what is called the context window. And this is like, well, the backprop says pay more attention to this guy.

41:05And then the backprop says pay even more attention to this guy, right? And LSTM was honestly a more complicated than hack here and more unwieldy way to accomplish this effect. Attention in many ways is a misleading term, right? Because it really doesn't have anything to do with the psychological version of that term. In fact, the irony is that Kyung Hoon and Dima Bada, now Dima is the intern who invented it, they didn't call it attention. Yoshua Ben-joo came at the end of the paper and put attention everywhere, and then they took it out again, as they tell the story, except they left it in a bunch of places, and then the rest is history.

41:46So then you've got the Transformers. Again, this is all in the Connectionist stream. And then there's some things happening off to the side. there's or maybe we're going to talk about this next in the next episode but there's generative adversarial networks and well let's talk about GANs or is that on the evolutionary algorithm end of things no very good so GANs generative adversarial networks were very popular for a few years in AI. And now it's kind of died down. And it's shrewd of you to point out their connection to evolutionary learning. Most people don't make that connection, but I think it is a very important connection.

42:48And GANs are actually how the term generative came. So the term generative from the beginnings is a term from statistical machine learning. And the idea of a generative model is that it's a model that generates the data. It's like, I want a process by which you generate samples. And then, so let me turn this the other way around. This is actually the whole Bayesian way of thinking is, I see a bunch of data. I'm going to postulate that they were generated by some process, and I'm going to try to identify what that process was. This is actually what generative learning is. So you come up with this model like a Bayesian network.

43:27This is like, well, this is how you generate patients and their symptoms. First, you decide which disease they have. And then you decide which symptoms they have, what fever are they going to have based on the fact that they have COVID, et cetera, et cetera. You generate the data. And this is why it's called the generative model in opposition to discriminative models, which is what most of machine learning is about, which is to just go say like, OK, I'm going to go forward from this data and try to decide, you know, what you suffer from. Or whether it's a cat or a dog, as opposed to like, how do I generate, you know, a cat or a dog?

43:56Right. So there's always been this thing in machine learning between generative and discriminative. discriminative always has always dominated and really still does because it just works. It's an easier problem and it works better. It's in some sense less satisfying because like, what is the deep science here? I'm just hacking to get the result. I want to know, you know, like, for example, physics, right? Is about generative models, right? Like here's how the universe is generated. So there's a certain type of person who really likes, you know, who thinks generative models should win at the end of the day.

44:29Now, you know, young Goodfellow, who's the guy who invented GANs, was another student at the same time in Yoshio Benjo's group, who had a huge group. And so it's not entirely an accident that so many people came out of it. And they were actually trying to solve a different problem, which was the problem. There's a type of neural network, which again, is a generative model, right? Not trained by backprop. And by the way, these were the ones that jeff hinton was always into jeff hinton even though he's a co-author on the paper never liked backprop he's actually on record they're saying that backprop is not the future he actually told me at one point the only reason i i publish papers on this primitive i wrote the only reason i do this community learning is in order to publish the papers because he was what was written was generative neural networks that actually generate you know stuff right but the problem with those new networks is that the inference is intractable and so they were trying to come up with a way to make it tractable right there's a long history of trying to do this in various quadrants of course including the bayans and whatnot and and and generative adversarial networks were actually a way to do this and you know without going into too much of the of the history or the details what is the idea is that the adversarial network is a game between the generative network that generates data and the discriminative network that tries to classify it.

45:53So actually have a generative model, the classifier working against each other, which is a brilliant idea because the idea is that the classifier is trying to distinguish between real data and the data that was created by the generative network. And so that forces the generative network to get better. Otherwise, in the beginning, it's basically like, ha ha ha you don't look anything like a cat falls right and then it gets better but then that forces the discriminator to get better as well so they get into this evolutionary arms race right it's co-evolution is the term in evolution right it's like the predator and the prey that generated in the discriminator and this was like a genuinely new idea in the context of neural networks in fact at the time you know yan lakun was going around saying like this is the most important idea in neural networks of the last 20 years.

46:46Then it petered out because it turned out it's very unstable. There's a lot that goes wrong. People never really, to my disappointment at least, you want to take that further into the whole evolutionary realm of things because co-evolution is very powerful. Maybe that's what we need in machine learning, but it was too hard to cut a long story short. It petered out, but that's where this whole thing of generating deep fakes began. right? Yin in his paper had this few, you know, grainy images that were generated, right? They were like, I mean, at the time, this was actually very impressive, right?

47:21But then like, if you think about applying more power to this, the images get bigger, they get, you know, final resolution, you go from images to video, and the whole, you know, deep X, the industry was born. So even though now there are the techniques like, you know, diffusion and whatnot, which we could also talk about being used to generate these things. So GANs are no longer the state of the art, but the whole, like that term generative in the neural network context really took off there. And this whole notion, which a lot of AI things gave you, but they're like, you know, how can I dazzle you with generating text or generating images really started with GANs?

47:58Yeah. Where should we go from here? We don't have a lot of time left. I wanted to, I don't know if it needs a big chunk of time on its own, but then you have Rich Sutton working on reinforcement learning, sort of coming at it from the behavioral psychologist side. And that, to me, is probably the most powerful idea. and it was certainly applied, but it seems to have become the darling of generative AI right now with this DeepSeek is training with pure RL. Can you talk about how RL fits into all of this? Absolutely. So reinforcement learning is actually different. So there's supervised learning, unsupervised learning, reinforcement learning.

49:01It's a different type of learning that, again, has been around for decades. It is inspired by animal psychology. In fact, that's where the term comes from. Like, you know, Pavlovian experiments, you reinforce behaviors by giving rewards and et cetera. And, you know, Rich Sutton has been the longstanding leader in that area. So what is the basic idea in reinforcement learning? It's important to start by understanding that. And by the way, reinforcement learning, like supervised learning, can be applied with any of the paradigms, but it has always been most popular with neural networks. Yeah. So with supervised learning, you have a teacher.

49:39The teacher says, aha, right answer. Like, you know, yes, this is a cat or no, this is a dog. Like you label things with the right answer. Yeah. Which makes learning easier, but is, you know, where does that data come from? There are many domains where you don't have that. If you look at how animals learn, right, or like how we learn, right, you touch the stove and then you learn not to. But the interesting thing about enforcement learning is that I feel the burn when I touch the stove, but I really shouldn't even have moved my hand towards it. So you need to back up that signal.

50:16So supervised learning, if it was just like at the last moment, don't touch it, that would be supervised learning. The whole idea of reinforcement learning is they have this thing called delayed rewards. You do something and you only see the result later, and then you need to propagate those results back to where you should have taken the actions. A famous example of this is AlphaGo. Yeah. And indeed, reinforcement learning goes back to the 50s, and the paper that coined it to machine learning had a precursor of reinforcement learning to learn to play checkers, and they learned to play that at human level.

50:46And the idea is this is like I play a whole game of checkers, either computer with you, the human. I don't know if my moves are good or bad. The only thing I know is at the end of the day, I won or I lost or we tied. Right. So there's a very small piece of reward, as it's called. Reward is like you won. Right. You get a you know, you get a reward. Right. You got to an enforcement learning is a set of techniques or really it's just the problem. Reinforcement learning is the problem of like, how can I take those rewards and propagate them back to where I need to make a decision? so that I make the right move that many steps further will lead me to win the game.

51:22And this combined with neural networks for vision applied to the Go board was, of course, what let DeepMind beat Lisa Dahl and et cetera. Now, reinforcement learning as done by Rich Sutton and other people for a long time was always very theoretical. They wanted to, they did these tiny little experiments that were, you know, not convincing to anybody, honestly. And they tried to prove theorems. What the deep mind folks did when they came along that was very important was they said, like, you know, to hell with all that. We're just going to engineer the heck out of this and get it to win. And they did, right?

52:03And so, again, for a while, deep reinforcement learning was very hot. There were, like, many papers about it. They actually started by doing it in video games, in Atari, and then went on to things like going chess. That, unfortunately, has petered out also because it didn't, DeepMind's agenda when they came out was deep reinforcement learning. And their idea was like, we're going to start with these games and then we're going to move to robots in simulation and finally robots in the world. And this is how we're going to solve AI. Right. Them is as obvious as we're saying, like, where are the Apollo Project AI?

52:33We're going to get there. They don't say that anymore. And so these days, reinforcement learning, as you pointed out, is, again, popular. But honestly, those of us who know this history at this point have some justified skepticism because we keep seeing this over and over again, which is people come out saying like, oh, look at this great success for reinforcement learning. And then when you look at it more carefully, there was no great success. What they were doing was effectively supervised learning or could have been done by supervised learning. In fact, there's this thing called reinforcement learning from human feedback that OpenAI popularized that they say is very important to make their chatbots work and whatnot.

53:11There was this paper out of Stanford that basically showed effectively you're just doing supervised learning. So I would say the same with DeepSeek. Until I see proof to the contrary, I will assume, and again, the evidence points in that direction, at least as far as I can tell, that there's really no reinforcement learning magic going on. The problem with reinforcement, so the appeal of reinforcement learning to the people who really want to solve AI is that like, clearly we do something like that as even a correspondence to some of the neuroscience with dopamine circuits and whatnot. The problem is that people have, disappointingly, despite so much effort, they have never really truly managed to make it work.

53:48In fact, it boils down to this. If your rewards are delayed, it's learning from sparse delayed rewards. If the rewards are not very delayed, they're not very sparse, then you can just use supervised learning, which is way easier. If the words really are delayed and sparse, then it doesn't work. So, so far we haven't really solved that problem. So when you see claim in the meantime, however, in the last few years, reinforcement learning has just become the sexy term that you can throw around and you call your thing reinforcement learning. And it sounds more impressive, even to yourself, which seems to be also what had happened with DeepSeq.

54:27Yeah. Okay. We have a couple of minutes left. Where should we go from here? Well, let me take that question the following way. Where should the neural network community go from here? And I think the high order bit is, as we alluded to earlier, what you've seen is broadly from the point of view of your app, but even from the point of view of machine learning and then deep learning, what we're seeing is more and more people doing more and more research about less and less. So I think where we really need to go from here is we need to broaden the scope. By the way, as an example, there are papers now showing that maybe LSTMs and recurrent networks aren't such a bad idea.

55:10A lot of things that seem dead in machine learning or AI have a tendency to come back, you know, neural networks being one of them. So I wouldn't be surprised at all if one of these things has dethroned transformers by next year. Or maybe it'll stagnate for 10 years or who knows, right? But I think a lot of people should be doing a lot. We really should let a thousand flowers bloom. And right now, it's basically one flower blooming or one petal of one flower blooming. And that's not very healthy. Yeah. And that phenomenon, because we've talked about it during the supervised learning vision phase, when everyone is focused on one narrow aspect and then they're forgetting about the Bayesians or something.

55:56That's a matter of funding, isn't it? Oh, funding is a big part of it. In fact, Minsky and Papert publicly admitted that the main reason they wrote that book that killed neural networks for a while was that DARPA was the big funder of AI research, so historically the biggest funder. When we finally get to AI, we will have DARPA to thank as much as anybody for having stuck with it until the industry took over. But they were getting jealous that neural networks were so sexy and they were getting all the DARPA funding and symbolic AI wasn't getting enough. So they said, we got to get rid of these guys.

56:31And so they did. So funding has always been an issue. I mean, particularly in research and industry as well, the phenomenon is a little different. DARPA for a long time after that, In fact, when I became a researcher, DARPA did not fund much machine learning research, if any, because the consensus, which in the field and among DARPA program, there is like machine learning is a waste of time. In the meantime, machine learning was taking off and DARPA was still not waking up. Finally they woke up in the 2000s and started getting into machine learning. And again, the funders themselves, whether in industry or in academia or the funding agencies, they're very prone to this hurt behavior.

57:14It's like the VCs, right? VCs are supposed to be looking out for different things, but at the end of the day, they're all looking out for the same things. So, you know, you need to inject some, you know, there's this, you know, gradient descent, right? There's like looking for the optimist, what neural networks do. But there's this method called simulated annealing where you actually inject noise, which seems like a bad thing, but the noise actually helps you not go straight to the local minimum and find a deeper one. We need more simulated than kneeling in AI research. When you need somebody, you need somebody.

57:47You want to hire people fast. The longer you wait, the more the market drifts, the more your needs drift. For example, you just realized your business needed to hire somebody yesterday. How can you find the right candidate fast? Just use Indeed. When it comes to hiring, Indeed is all you need. Stop struggling to get your job posts seen on other job sites. Indeed's sponsored jobs help you stand out and hire fast. With sponsored jobs, your job postings jump to the top of the page for your relevant candidates so you can reach the people you want faster. And it makes a huge difference. According to Indeed data, sponsored jobs posted directly on Indeed have 45 % more applications than non-sponsored jobs.

58:37One of the things I love about Indeed is that it makes hiring so fast. I've searched for candidates before, and it can take months, particularly if you're just working through word of mouth. In those days, I didn't have Indeed. Plus, with Indeed sponsored jobs, there's no monthly subscriptions, no long-term contracts, and you only pay for results. How fast is Indeed? In the minute I've been talking to you, 23 hires were made on Indeed, according to Indeed data, worldwide. So there's no need to wait any longer. Speed up your hiring right now with Indeed, and listeners on this show will get a$75 sponsored job credit to get their jobs more visibility at indeed.com slash IonAI.

59:26Just go to indeed.com slash IonAI. That's indeed, I-N-D-E-D dot com slash IonAI. All run together, E-Y-E-O-N-A-I. Go right now and support our show by saying you heard about Indeed on this podcast. Indeed.com slash IonAI. Terms and conditions apply. Hiring indeed is all you need.

From the publisher

This episode is sponsored by Indeed. 

Stop struggling to get your job post seen on other job sites. Indeed's Sponsored Jobs help you stand out and hire fast. With Sponsored Jobs your post jumps to the top of the page for your relevant candidates, so you can reach the people you want faster.

Get a $75 Sponsored Job Credit to boost your job’s visibility! Claim your offer now: https://www.indeed.com/EYEONAI

 

 

In this episode, renowned AI researcher Pedro Domingos, author of The Master Algorithm, takes us deep into the world of Connectionism—the AI tribe behind neural networks and the deep learning revolution.

 

From the birth of neural networks in the 1940s to the explosive rise of transformers and ChatGPT, Pedro unpacks the history, breakthroughs, and limitations of connectionist AI. Along the way, he explores how supervised learning continues to quietly power today’s most impressive AI systems—and why reinforcement learning and unsupervised learning are still lagging behind.

 

We also dive into:

  • The tribal war between Connectionists and Symbolists

  • The surprising origins of Backpropagation

  • How transformers redefined machine translation

  • Why GANs and generative models exploded (and then faded)

  • The myth of modern reinforcement learning (DeepSeek, RLHF, etc.)

  • The danger of AI research narrowing too soon around one dominant approach

Whether you're an AI enthusiast, a machine learning practitioner, or just curious about where intelligence is headed, this episode offers a rare deep dive into the ideological foundations of AI—and what’s coming next.


Don’t forget to subscribe for more episodes on AI, data, and the future of tech.

 

 

Stay Updated:

Craig Smith on X:https://x.com/craigss

Eye on A.I. on X: https://x.com/EyeOn_AI

 

 

(00:00) What Are Generative Models?

(03:02) AI Progress and the Local Optimum Trap

(06:30) The Five Tribes of AI and Why They Matter

(09:07) The Rise of Connectionism

(11:14) Rosenblatt’s Perceptron and the First AI Hype Cycle

(13:35) Backpropagation: The Algorithm That Changed Everything

(19:39) How Backpropagation Actually Works

(21:22) AlexNet and the Deep Learning Boom

(23:22) Why the Vision Community Resisted Neural Nets

(25:39) The Expansion of Deep Learning

(28:48) NetTalk and the Baby Steps of Neural Speech

(31:24) How Transformers (and Attention) Transformed AI

(34:36) Why Attention Solved the Bottleneck in Translation

(35:24) The Untold Story of Transformer Invention

(38:35) LSTMs vs. Attention: Solving the Vanishing Gradient Problem

(42:29) GANs: The Evolutionary Arms Race in AI

(48:53) Reinforcement Learning Explained

(52:46) Why RL Is Mostly Just Supervised Learning in Disguise

(54:35) Where AI Research Should Go Next

 

More from Eye On A.I.

All 266 episodes
#248 Pedro Domingos: How Connectionism Is Reshaping the Future of Machine LearningEye On A.I. · 1 h
Listen in VO