In short
Eye On A.I. Podcast Notes
Episode Title
#236 Pedro Domingo's on Bayesians and Analogical Learning in AI
Episode Overview In this episode, host Craig Smith interviews Pedro Domingos, an eminent AI researcher and author of *The Master Algorithm*. The discussion focuses on the evolving landscape of machine learning, particularly the resurgence of Bayesian AI compared to deep learning methodologies. Key concepts such as Bayesian networks, analogical learning, and the implications of these approaches in various fields, including medical diagnosis and AI regulation, are thoroughly examined.
---
Key Topics Discussed
- Introduction to the Episode
- Overview of guest Pedro Domingos and his background in AI.
- Discussion of the significance of Bayesian learning in the context of AI advancements.
- The Five Tribes of Machine Learning
- Domingos introduces the concept of the five tribes, focusing on Bayesians as a particularly fervent group.
- Explanation of how Bayesians differentiate themselves through their understanding and application of probability.
- Bayesian vs. Frequentist Approaches
- Definitions of Probability:
- Frequentist: Probability defined through long-term frequency of events.
- Bayesian: Probability viewed as a subjective belief that can be updated with evidence.
- Debate on the validity of both interpretations, highlighting the strengths and weaknesses of each.
- Bayesian Networks
- Explanation of how Bayesian networks work and their role in AI decision-making.
- Historical context: Use of Bayesian networks in Google's ad system before the rise of deep learning.
- Applications in medical diagnosis and search & rescue operations, outperforming human capabilities.
- Challenges with Bayesian Learning
- Computational difficulties associated with Bayesian learning.
- Importance of uncertainty in AI decision-making and how Bayesian approaches handle it.
- Deep Learning’s Limitations
- Critique of neural networks as potentially oversimplified forms of analogical learning.
- Discussion on the limitations of deep learning in comparison to Bayesian approaches.
- Analogical Learning
- Overview of analogical learning as introduced by Douglas Hofstadter.
- Impact on pattern recognition and case-based reasoning.
- Historical evolution of AI paradigms from symbolic AI in the 80s to support vector machines in the 2000s.
- Support Vector Machines (SVMs)
- Comparison of SVMs to neural networks, emphasizing their robustness and computational efficiency.
- Ongoing relevance of SVMs in solving specific types of machine learning problems.
- The Future of AI
- Speculation on the direction of AI research, including a potential resurgence of Bayesian methods.
- Discussion of hybrid models that may combine different approaches for enhanced performance.
Key Takeaways
- Understanding Probability: An in-depth understanding of how probability is conceptualized is crucial for advancements in AI.
- Bayesian Learning's Significance: Bayesian methods continue to play a vital role in AI applications, particularly in fields requiring precision in uncertainty quantification.
- Analogical Thinking in AI: The potential of analogical reasoning to enhance AI learning and adaptability cannot be understated.
- Critique of Neural Networks: The limitations of neural networks suggest a need for a broader exploration of machine learning techniques beyond deep learning.
Conclusion This episode presents a comprehensive exploration of the current state of AI, emphasizing the critical debate between Bayesian and deep learning approaches. The insights shared by Pedro Domingos highlight the importance of understanding and leveraging various learning paradigms to tackle complex real-world problems.
---
Stay Updated
- Craig Smith on Twitter: [@craigss](https://twitter.com/craigss)
- Eye on A.I. on Twitter: [@EyeOn_AI](https://twitter.com/EyeOn_AI)
---
Sponsorship This episode is sponsored by Thuma, a modern design company specializing in home essentials. For a discount, visit [Thuma](http://thuma.co/eyeonai) for your first bed purchase.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00The Bayesians are the hardest core of the five tribes, the most tribal. they truly believe that either what you're doing is Bayesian or you're wrong and they will never die the Bayesians are a tribe that comes from statistics so machine learning of course is closely related with statistics it builds on it in many different ways because after all both of them are about building models from data and then acting on them of course statistics has a much longer history statistics and probability going back centuries even statistics in the last 100 years at least, if not more, has been dominated by what is called the frequentist school, as opposed to the Bayesian school.
0:38So the Bayesians in statistics have always been the oppressed minority, and they have a big chip on their shoulder about it. And so what is the difference between them? The difference is about what is the definition of probability. The truth is that even today, for all the thousands of papers or millions that have been written, nobody really knows exactly what probability is. Probability is actually a surprisingly slippery concept. Create an oasis with Thuma, a modern design company that specializes in furniture and home goods. By stripping away everything but the essential, Thuma makes elevated beds with premium materials and intentional details.
1:19I'm in the process of reorganizing my house and I'm giving Thuma a serious look for help in renovating and redesigning. Thuma combines the perfect balance of form, craftsmanship, and functionality. With over 17 ,000 five-star reviews, the Thuma Bed Collection is proof that simplicity is the truest form of sophistication. Using the technique of Japanese joinery, pieces are crafted from solid wood and precision cut for a silent, stable foundation. With clean lines, subtle curves, and minimalist style, the Thuma bed collection is available in four signature finishes to match any design aesthetic.
2:12Headboard upgrades are available for customization as desired. To get$100 toward your first bed purchase, go to Thuma. That's T-H-U-M-A dot C-O slash IonAI. IonAI all run together, E-Y-E-O-N-A-I. So for$100 off your first purchase, go to thuma.co slash ionai. That's T-H-U-M-A dot C-O slash ionai to receive$100 off your first bed purchase. I'm Peter Domingos. I'm a professor of computer science at the University of Washington in Seattle. I'm an AI researcher. I've been doing it since 1998, actually. I'm also known as the author of The Master Algorithm and 2040, a Silicon Valley satire. The Master Algorithm in particular is an introduction to machine learning for a broad audience that was very popular.
3:26And I guess that's why we're here today. That's right. So tell us about the Bayjans, which was one of the five tribes. The Bayjans are the hardest core of the five tribes, the most tribal. They truly believe that either what you're doing is Bayesian or you're wrong. And they will never die. Some of the other tribes, sometimes they look a little tired, but the Bayesians always keep going to thick and thin. Now, why is that and who are they? The Bayesians are a tribe that comes from statistics. So machine learning, of course, is closely related with statistics. It builds on it in many different ways because, after all, both of them are about building models from data and then acting on them.
4:18Of course, statistics has a much longer history, going back to statistics and probability, which of course, again, is closely related going back centuries even. But statistics in the last hundred years, at least, if not more, has been dominated by what is called the frequentist school, as opposed to the Bayesian school. So the Bayesians in statistics have always been the oppressed minority, and they have a big chip on their shoulder about it. So what is the difference between them? The difference is about what is the definition of probability. The truth is that even today, for all the thousands of papers or millions that have been written, nobody really knows exactly what probability is.
5:04Probability is actually a surprisingly slippery concept. But the two main definitions are the frequentist one, which in some ways is the more natural one, which is, he says, just a probability is just the limit of a frequency. For example, why do I say that the probability of getting heads on a coin toss is a half? Because if I keep flipping a coin over and over again, approximately half the time, I'll get heads. And if I toss it infinite times, it'll be exactly half, provided the coin is unbiased. So probability in this view is just frequency. But then there's a lot of problems with this. Like, for example, well, okay, but then what is the probability that Trump will win the election?
5:49You can't do infinite trials of Trump and Kamala running against each other. So somehow that definition seems to fall short. And there's a whole bunch more problems like that, which we could get into. And so the Bayesians actually have a very different reply to this question. They say, probability is subjective. It is inherently subjective. probability is just your belief like i say oh the odds that trump will win are 55 prove me wrong right so actually you're entitled to whatever beliefs you want and the bayans just say how given your beliefs you should then calculate new beliefs such that such that they're all coherent but it doesn't tell you what your beliefs their priority should be and in particular in bayesian statistics and then throw that in Bayesian learning, which again inherits all of that machinery.
6:41So Bayesian machine learning in a way is starting with Bayesian statistics and then adding all the computational power that we have today, including in some ways that weren't necessary in other methods, but are necessary in this one that we can get into. But the basic way that Bayesian learning and statistics work is based on something called Bayes' theorem, which is in fact where the name comes from which is actually ironic because the theorem is not due to Bayes it's due to Laplace so really it should be called Laplacian learning but Laplace already has so many other things to his name that adding one more might be too much so what does Bayes' theorem say?
7:20it says that the way the whole Bayesian point of view or even statistical point of view at least by some standards is there's a probability distribution over everything that could happen in the world. And what Bayes' theorem says is you start with your prior probabilities, which is the probabilities that you assign, for example, to this patient having cancer or this email being spammed before you've seen the patient or before you've seen the email. So these are your beliefs, a priority. And then as you see evidence, you update your beliefs. And the probability of what you see given your model is called the likelihood.
8:02And what Bayes' theorem says is that your posterior is the prior times the likelihood normalized. And your prior times your likelihood is now your new prior for when new evidence comes in. So I'm playing the stock market. I say this stock is going to go up. It goes up or down. I compute my likelihood up with my prior. And now this is my prior for the next day. So this is in essence how Bayesian learning and Bayesian statistics work. Can I just ask, tell us who Bayes was. Thomas Bayes, or Reverend Bayes, was an English Protestant preacher, clergyman, who, you know, in the 18th century, I believe, who, you know, like many cultured people at the time, had mathematics as his hobby.
8:50and you know he so he so probability has an interesting history going back several you know hundred years it actually started out in a very non-serious way as people trying to figure out how to win at gambling right can I predict what are the chances of whatever this combination of dice coming up and so on and it kind of went on from there so when they and what was formulated was in some ways a proto version of his theorem not in you know you can see the idea there but he didn't really formulate it as i said laplace did it but in some ways that he was the one who started things on this track of well let's think about the prior probability and the likelihood and multiply it too and by the way this is something that people often naturally gets people confused frequencies have no problem whatsoever with base theorem in fact you learn base theorem in stats 101 it's one of the first things you learn.
9:45It's a really simple theorem. It's almost just a definition of conditional probability. It barely merits being called a theorem. The thing that makes it controversial is when the Bayesians say, ah, but your prior probability is whatever you want to make it. This is where the frequency will be like, whoa, whoa, whoa, what are you talking about? You can't just make stuff up. And you can see why there'd be a lot of resistance from scientists in general to this idea that, hey, are you just going to make probabilities up? Like, what is that? science is supposed to be objective. But the point that, you know, my point of view is actually the frequentists and the Bayesians both have good points.
10:20So you can't truly ignore either of them. And one very important point that the Bayesians make is that like, yeah, but you frequentists are also making assumptions that are just assumptions. You're just not being explicit about them. So it's rare to be explicit about them. So there you go. And the Bayesians, it sounds a little bit like Euclidean geometry with the axioms and everything is built on top of that. And those are just taken on faith. No, absolutely. So the Bayesians are probably, again, of all the tribes, the one that most strongly believes in, we should solve machine learning from first principles, from a few axioms.
11:05and in particular from Bayes' theorem. And, you know, there's a few others of probability. And that's the right way to do it. Don't talk to me about the brain because the brain is a mess. Don't talk to me about evolution because evolution is a hacker. We have no guarantee that any of those did anything meaningful. I mean, you know, the optic nerve comes out of your retina towards the outside. Like, what the heck is that, right? Why would the brain make any more sense than that? So the whole Bayesian idea is that, like, yeah, you know, I'm going to tell the axioms by which you need to reason. And they even have this expression called turning the Bayesian crank, which is I have my axioms, you bring near data, and I turn the crank, and I produce the output.
11:45And anything else that you do is just wrong. Now, of course, where things get interesting, at least from the machine learning point of view, and in fact, from the statistics point of view as well, is that trying to do this in any non-trivial setting is just computationally too hard. It's completely infeasible. And in fact, a large part of the reason why for many decades, you know, Bayesians were really in the background in statistics. It's just you couldn't do Bayesian statistics. It wasn't possible. And again, part of why they're on their eyes, even within statistics and other disciplines like economics and other areas of science today, is that we now have the computing power to do Bayesian calculations.
12:23But even then, doing them requires doing approximations and whatnot. So at the end of the day, the Bayians wind up making as many unfounded or furious assumptions as everybody else, which is ironic considering where they started from. So tell me how they put this into a computer program that learns. Absolutely. So Bayesianism doesn't tell you, Prairie, what kind of model to use. and so what happens is that there's a bayesian version of everything so for example there's bayesian neural networks there's this you know guy david mckay that did a very brilliant phc thesis on how to make bayesian neural networks and for a while they actually took over the community it was very clever he kind of took it over from the inside uh there's bayesian versions of symbolic learning algorithms there's bayesian versions of analogical learning algorithms so like whatever learning algorithm you come up with there's a you know it's almost you know certain that tomorrow, the Bayesians will come up with their correct Bayesian version of it.
13:27Having said that, something that is very distinguishing, I would say, of the core Bayesian approaches is that they represent probabilities explicitly. The models actually have not just weights or some other parameters with no clear meaning, is the parameters are probabilities. So for example, the most, probably most widely used, or at least best known type of Bayesian model in AI is our work called vision networks. And Huda Pearl actually won the Turing Award for developing vision networks, so it's a big deal. So what is a vision network? A vision network is an answer. It's actually an answer to a problem that goes back to the early days of AI.
14:08We talked about when we talked about the symbolists that, like, yes, the world is full of uncertainty and ambiguity and confusion and whatnot. And in theory, the right way to handle it is probability. I think everybody more or less agrees with that. There's infinite variations, but at the end of the day, the axioms of probability are hard to get around. And in fact, that's what the Bidians keep pounding on. But the problem is that it's computationally intractable. And the first reason why it's computationally intractable is that if I give you, you know, let's suppose I have 100 variables of interest and they're all binary.
14:42Yes or no. Do you have a fever? Do you have whatever, diabetes? Do you drive a car? Et cetera. right? Let's just keep them binary, Boolean, for the sake of argument. There are two to the hundredth possible states of those variables. So in order to naively represent that, I need two to the hundredth probabilities. And there is no computer in the whole. In fact, if you used an atom in the universe to store each one of these, there wouldn't be enough atoms in the universe to just store that distribution. It's just completely impossible. So you have to find some way to summarize that distribution.
15:19And Bayesian networks take advantage of this thing called conditional independence, which is really the crucial property. And so, for example, so backing up slightly, we say that two things are independent when the probability of one doesn't depend on the probability of another. So the probability that I smoke probably has nothing to do with the probability that somebody will observe a shark tomorrow in the Pacific head for a year. They're independent. Now, independent is not very interesting. The interesting notion is conditional impassence. For example, A is conditional independent of B given C, not because A and B are independent, but because once I know C, they become independent.
16:04In other words, all the information about B that A contains is also contained in C. and if you think about it the entire universe works by this principle because at least according to physics as we understand it today what happens to me is conditionally independent of what happens far away condition on what happens close to me so like the light from distant stars can hit my retina but before it hits my retina it will hit a point in space very close to my retina so without conditional independence life would be completely intractable and cognition would be impossible and so on. So what vision networks are a representation that tries, in the knowledge engineering days, people go on, for example, vision networks are very good for things like medical diagnosis because you really do want to precisely quantify those probabilities, right?
16:57These are important decisions at stake. And you can say like, okay, what are the things that, what are the illnesses that could cause a fever? And then you draw this graph, the vision network, with an arrow from all these different illnesses to fever or to high blood sugar. There's one from diabetes, but there isn't one from, I don't know, brain cancer. And then you draw a big diagram like this. And then the only parameters that you have to learn, this is the key, are for each node in this network, given its parents. I only need to represent the probability of fever given the things that cause it, which is a vastly, vastly simpler problem than letting everything depend on everything else.
17:41So the amount of memory. So conditionally independent, such as lets you build a reasonable sized model. Then there's also the problem of doing inference, which is a lot of what, you know, who the Perl worked on in a way that that doesn't blow up. And also of learning this from a finite amount of data, because again, if I need to learn to the hundred probabilities just by measuring frequencies, I will never be done. But once you have a Bayesian network, a lot of this starts to be feasible. It can still be very computationally expensive. And in particular, Bayesians have this thing in which they're different from everybody else, in which they don't just try to find...
18:18Some people would even say that this is what defines Bayesianism, not even just having prior probabilities, is this notion that there is no right model. There is no one neural network or no one decision tree. They are all possible. They all have their probabilities. And what I have to do is not find the best model, unlike what all the others do, is find the probability of each and every possible model, and then to make a prediction, I average over all of them. Now, you can see that from a theoretical point of view, this is very attractive, right? Of course, I don't know. This is induction. I don't know for sure what's going to be the right model, so I should just average over them.
18:56The problem is that there's often a doubly exponential number of models. not just exponential, but like you take all those states, which are already exponential, and now there's an exponential number of models over those. So you have not just two to the n, but two to the two to the n, which is just absurd, right? And Bajins tie themselves in just trying to compute some approximation to this average overall models efficiently. And it's a massive waste of time. I really feel their pain. But just to give you an example, Google, before deep learning came along, their whole ad placement system, which is what made all their money, was a big vision network.
19:38They just had a massive vision network with, I don't know, probably millions of notes that were things like words, right? The words that appear in the page and on the ad that you want to show and various other things and how these things depend on each other and ultimately what they want to do is like predict the probability that you will click on this ad if it contains these words and these links. And that was a huge vision network. So, you know, there's many important applications even today of vision learning. Yeah, just one question on that. How do you take the, what was it, cancer and all the various causes of cancer, was that the example you gave?
20:25How do you compute the probability for each of those nodes? Let's start by taking a simple example. Suppose that I had just two symptoms, fever and blood glucose. What I want to do is predict the probability that you have diabetes. Then what I need, this is a very small model, is this little table that goes like this. If you have no fever and low blood glucose, then your probability of diabetes is low, 0.05, let's say. But if you have a high fever and low blood glucose, that slightly increases your probability of having diabetes. By the way, I'm not a doctor, so I'm just making this up. But the main thing is you still have low blood glucose, so maybe your probability of diabetes is 0.1.
21:15Now, if you have high blood glucose, your probability of diabetes jumps up to whatever, right? And finally, if you have both of them, it's, let's say, you know, 0.8. So you just need to specify probability for each combination of these variables. Now, in a Bayesian network, you only need to do it for a combination of the appearance of that variable in the model, because given those, it's conditionally independent of all the others. That's the big whim, right? I don't need to condition on a hundred things, provided that if I know these two, the others become irrelevant. If I don't know these two, is actually the beauty of vision networks, is that then I can infer the probability of those two from the others, and then from those, the probability of actually having diabetes.
21:54So I don't actually need to know them. But knowing the structure of the world is actually what makes me, you know, lets me have a table with two parents and therefore four lines, as opposed to a table with two to the hundred lines. Yeah. And you said there are, you mentioned the Google example, but what are the other applications that use Bayesian networks today? And is there still progress being made in the Bayesian school into other networks or other structures, I should say? Yeah, so I would say that, for example, medical diagnosis, as I mentioned, is important. So there are some really good, very sophisticated vision networks for medical diagnosis that exist today.
22:45And by the way, they've actually existed for decades now, and they're better than doctors. We've been, you know, like machine learning has been better than doctors for, I don't know, 50 years. The reason it's not widespread is that the doctors won't let them be used. They're the gatekeepers. Why would the gatekeepers unemploy themselves? So there's a very interesting sociological aspect to this. But like next time you hear this big polemic, oh, does Chad GPT do diagnosis better than doctor? Even expert systems pre-learning in the 70s already did, you know, medical diagnosis better than doctors at the things that they were designed for.
23:20It turns out that beating doctors is not that hard. And in particular, famously, humans are very bad at Bay's theorem. They don't understand that you need to take the prime probability into account. And doctors make that mistake too. So if you have a disease that is extremely rare, but you see these symptoms that are consistent with it, the doctor will tend to overestimate the probability, not realizing that even with these symptoms, it's still very unlikely. So doctors really screw this up all the time. I have Bayesian colleagues who say, I went to a doctor and I was just shocked at how stupid he was, telling me that I probably have this and that, and I'm like, oh my God, this guy doesn't even know stats 101.
24:00one. So, you know, visions also have a certain superiority complex, which is not entirely unjustified. But anyway, so medical diagnosis is another class of applications. I would say that in general, and this I think is actually what is important for people to know, is like, what are the types of problems for which these different types of machine learning make more or less sense? For vision learning, I would say it's typically when quantifying uncertainty precisely is very important, right? If just making a prediction, which is what your typical neural network does, it just says, well, it's going to be this, right?
24:38If that's not enough, if you really want, like you're a doctor, you know, to go back to that example, you don't want to tell your patient you have to get, you know, you have to have surgery. That's the patient's decision. What you need to say is, look, you know, there's this probability that you have a tumor. If you don't get surgery, there's this probability that you die. and et cetera, et cetera. And then the patient, according to their own utilities, as they call, can make their decision. A very interesting type of application, which there are some famous historical examples, is let's say, for example, America just lost a submarine in the Pacific.
25:16The submarine has disappeared. Pacific is big, right? You can't search every square mile. And if you just start doing things randomly, you might not get anywhere. Now, Bayesian learning is really good for this because it says like, okay, let me start with my prior distribution over where the submarine would be. It's not uniform, right? It's more likely to be near the routes that they usually take, blah, blah, blah. And now let me start, you know, bringing in all the relevant evidence, right? Was there a storm here? Did I hear some whatever sonar ping there? it's those pieces of evidence might be very weak, such that, for example, a symbol is looking at it would be like, I can't conclude anything from that.
25:56But as you accumulate these pieces of evidence, you know, the probability of each square in the grid starts going up or down. And with enough pieces of weak evidence, you often wind up with a very strong belief that, hey, the submarine is right here. And then you go there and you find the submarine. So subagianism is very good for stuff like that. I would argue that it is, you know, in many ways, the right approach to problems like that, or at least the best one we have at this point. Yeah. One question. I mean, we keep using the term machine learning, but this sounds like a predictive system, which is different than a learning system.
26:34No, absolutely. So very good point. And again, we are guilty of conflating the two. So if you were to do what I just described, let's say you are an admiral in the Navy trying to find the missing sub, right? And I say, let's do this, right? And one of your frequentists would be like, but where do all those probabilities come from, right? And this is where I say, well, you can impede them. You tell me what they should be. A frequentist would never do that. That's just not allowed in frequent statistics, as a result of which there are a lot of problems that they just can't touch. But you say like, no, you tell me, what is your prior distribution over where in the Pacific this submarine is going to be, right?
27:15And I'm just going to believe you. And I'm going to tell you how to then consistently and soundly go from these estimates, right? Now, what happens in practice, as I alluded to earlier, is that like as much as possible, we don't just depend on our prior probabilities. We also use data. but the data they you know again the data about the routes that submarines take and blah blah blah right but what those data do is they cause us to update our prior and again when you have a lot of data in many ways visionism becomes superfluous it's a cost that you don't need to pay when you have very little data unless you use your prior as your host so this is genuinely like so for For example, a famous example is, let's say I'm going to flip my coin and I only flip it once.
28:06Frequency statisticians will use what is called the maximum likelihood principle. To them, that's the right thing to do. But the maximum likelihood principle says that if I flip a coin once and it comes up heads, the probability of heads is 100%, which is stupid, right? And then if the coin comes up heads twice, it's still 100%. And if it comes up heads and tails, then it's 50-50. But the Bayjans don't suffer from this. They're like, no, a priori, let's say you believe that the coin is unbiased. So a priori, the probability is 50-50. And now when the coin comes up heads, that makes it, you know, depending on how exactly you do this, let's say 60-40.
28:42And then if it comes up heads, right? So this is actually a much, I mean, honestly, you would be crazy to not do this when you have a small amount of data. And in fact, what the frequentists do is that they have a bunch of ad hoc techniques, right? This is part of the Bayesian criticism is correct, that they wound up making all these assumptions. But often the assumptions are incoherent, which really annoys the Bayesians. And in fact, one analogy that might be helpful is that these are two churches of statistics, and the Bayesians are the Catholic church. They have one set of true beliefs, and the frequentists are the Protestant.
Read the full transcript
29:18Like anybody can sit up shop and make their assumptions and start cranking away. But at the end of the day, it's all this big mess that's not really consistent. So if you evaluate, I mean, often the Bayesians, I think this has an interesting relation to why they're so fanatic. They're often really smart people and people who really care about getting things right. And they're just disgusted by the mess of ordinary statistics and machine learning, and they want to do it the right way. But so to answer the second part of your question, there is very much still Bayesian research going on because there always will be.
29:53As I said, Bayesian stuff will never die. Of course, the percentage of the machine learning research that is Bayesian these days is much smaller than it was 10 or 20 years ago. Partly because the amount of connectionist research has exploded, but also a lot of the people have sort of like migrated from doing Bayesian learning to doing connectionist learning or just, you know. The harder core Bayesians always keep doing that, but some of the others drift away. There are famous examples of people doing this. So yes, there's still progress happening. And for example, doing Bayesian inferences is one of those intractable problems that there will always be progress to be made on.
30:34So for example, in particular, there's this technique called Markov Chain Monte Carlo, which is doing inference by generating random samples and counting them and whatnot, which actually goes, you know, the guy who invented, the people who made this were in the Manhattan Project, and they were trying to simulate nuclear reactions and thief the neurons, the neutrons, sorry, with the, you know, get over the threshold to cause it and so on. So when you have to deal with intractable probabilities, I mean, like Markov chain Monte Carlo methods in general are one of the most studied subjects in all of science and one of the most widely used algorithms in all of science, like top 10.
31:18One of the top 10 is Markov, is Monte Carlo sampling techniques. And even outside of AI, right? Even outside of machine learning, people in statistics, people in physics, people in economics, people in all these different sciences, they need these methods. So they will continue to research them. The world is much larger than neural networks. Okay, let's talk about analogizers. So the analogizers are actually the least cohesive of the tribes. So the Bayesians would probably, they will call themselves a tribe, maybe not exactly in those words, but they're the tribe with the strongest core identity.
32:02The analogizers are actually the one with the weakest, so it's kind of an interesting contrast to talk about, you know, both of them next to each other. The analogizers are actually, I glone together a set of different groups of people that actually don't necessarily interact that much, but they have something very important in common. So let me, you know, let me bring up those two maybe separately and then see where they're relation. I mean, there's more than two, but what is the basic idea? The basic idea is that that cognition is analogy. So a famous analogizer, and the person who coined the term, is Douglas Hofstadter, who is the author of Gödel Escher Bach, and a big fan of analogies.
32:46The funny thing, by the way, about Gödel Escher Bach, for those of you who have read it, is that it's a book about logic, right? It's a book about symbolic AI. Back in 1979, when it was produced, was the Haiti of Symbolic AI. And I know a lot of people who got into AI because of that book, it's about logic, but it's full of analogies, which is what makes it so fun. Analogies between Gödel and Escher and Bach and a lot of other things, right? And more recently, he co-wrote a book called Surfaces and Essences. Subtitle is Analogy as the Fuel and Fire of Cognition. And his argument is that, again, he doesn't, you know, he's not a half measures kind of person, is that like everything in cognition is analogy.
33:28There is no aspect of cognition that is not analogy. And the whole book is 600 pages just showing this in detail from the smallest everyday examples to the highest achievements of science. Einstein and Galah and whatnot were just great analogizers. They made these analogies between things that made them find their discoveries. But also the way you understand the simplest words in the language is by analogy with other things. And by the way, what I have found talking with a lot of different kinds of people over the years based on the book, is that at the end of the day of the five schools, the one that is most intuitive to normal human beings is analogy.
34:08Yes, the brain, sure, and Bayesianism, that's the least comprehensible. But the idea that we reason by analogy, everybody can relate to that because we do. Now, I think, you know, Douglas Hofstadter, he's, I sympathize with a lot of what he says, but again, And I think he goes too far, right? It's not that he can find an analogizer angle on everything, but that doesn't mean that that has exhausted. Like, yes, Einstein was a great analogizer. I buy that. But there was more to it than that, right? And I think that he overlooks that. So he's an example of analogizer that comes from psychology. So in psychology, again, there's a huge literature going back decades about how people do everything by analogy.
34:54and in particular there's this thing called structure mapping and Deirdre Gantner was the person who proposed it that says the way I do analogies by mapping structure from one domain to another a famous example for example is how Niels Bohr came up with his theory of the atom by making an analogy between it and the solar system where the nucleus you know was the sun and the planets were the electrons which actually as it turns out is a really bad analogy but it was the analogy that got quantum mechanics off the ground. So, you know, you shouldn't poo-poo that. But so what he did was like he mapped the structure of the solar system to the atom.
35:30And based on that, this is the key, he was able to make a lot of inferences. Now, to take a more mundane example, something that is very popular, again, has been for decades, is using this type of learning for help desks and call centers. I'm Microsoft and you call in saying, hey, my printer isn't working. And now 80 % of the time it's the same dozen problems. And what you need to do is what is called case-based reasoning. And again, there's a whole subfield of AI, I should say, which case-based reasoning is really just a form of analogical learning. It's saying like, okay, let me find the simplest problem to yours and the solution that I had to that problem.
36:15And now let me look at what the differences are and see if I can tweak your problem into this one that I know. And for example, in medical diagnosis, a very simple example, this is like, you're a new patient, you come into my office and I don't know anything about medicine. I use the example in the book of Frank Abbeimil Jr., this guy who pretended to be a doctor and put a fake Harvard diploma on his wall in a hospital in Georgia and wound up being the most popular doctor in the hospital. The patients loved him. He didn't know any medicine, so what do you do, right? Here's something you can do.
36:52You have a database of patient records because whatever, the doctor in that office before I had it, the new patient comes into your office and she tells you her symptoms and you just look for the patient with the most similar symptoms. And whatever that patient had, you say that this one has. And as dumb as this is, it turns out that you can actually prove mathematically that if you give me enough examples of this, I will get the right answer. Because as I get more examples, the nearest ones, this is the so-called nearest neighbor algorithm, which is another strand coming from, you know, parent recognition in the 50s.
37:26The nearest neighbor algorithm says just make the prediction that is the whatever class, for example, diabetes or whatever, that the nearest neighbor of the new case is in the data. And with enough data and a slight refinement of this algorithm, you'll get the right answer. So I sometimes jokingly say to people that, yeah, the singularity happened in 1951, because that's when this algorithm was invented. And it actually is the first algorithm that can learn any function from data. so like in some ways it's the first true machine learning algorithm because it's not like oh i can learn a linear frontier or something with a fixed form a mixture of gaussians no you give me more data here's the thing like all your traditional statistical models vision or frequentist at some point you can give them all the data that you want that they don't continue they don't get better they've they've uh you've exceeded their capacity right so there's even a technical term of capacity mius neighbor has infinite capacity you keep giving me more data I keep getting better.
38:28So if scaling is all that matters, as I tweeted the other day, then we've already reached AGR. Of course it isn't, but you get the point. So then more recently, within machine learning, the best known form of analogical learning is support vector machines or kernel machines. Support vector machines or kernel machines are really just a very mathematically sophisticated version of nearest neighbor. But for a while there, like immediately prior to the recent explosion of connectionism in the 2000s, there was a decade there. The decade of the 2000s, machine learning was dominated by kernel machines.
39:10You would go to ICML and NIPS and all these machine learning conferences, and there'd be almost no papers on your networks, and there'd be tons and tons of papers on kernel machines. So there you have it. Life is short. Yeah. So what's, yeah, I mean, talk about support vector machines a bit and um that that's still being applied is that right no absolutely and again there are problems for and again this is a recurring pattern you know the spotlight is on neural networks is this but there continue to be problems for which support vector machines are the best solution so if you care about solving such a problem you should use a support vector machine as and again one of the things that frustrates me and part of why i wrote the mass I see people wasting an enormous amount of time and getting poor results just because they're using the wrong technique.
40:00You know, if you have, you should try to figure out, and often it's not trivial, what is the best technique for your problem? But you should at least be aware, for example, that there is such a thing as support vector machines, and maybe that's what you want to use rather than neural networks. And in particular, support vector machines are far easier to apply than neural networks, far easier. So God help you if you go try to apply deep learning to a problem that could be solved by a support vector machine. You're just giving yourself enormous pain for no good reason. In fact, the history of this is kind of fun.
40:35Support vector machines were developed at Bell Labs in the 90s by this guy called Vladimir Vapnik, which himself has an interesting history. But he was part of the group that was led by Yann LeCun that was doing convolutional neural networks for digit recognition. And they spent 10 years engineering these networks very specifically to do better digit recognition. And the support technology machines initially were a purely theoretical idea. They're a very mathematically based technique. But they came in and they're like, oh, so what shall we apply this to, right? Digit recognition. And right off the bat, they were tying ConvNets.
41:17It wasn't even that they were winning, but the fact that out of the box, the system was as good as ConvNets just made people's jaw drop. And then these were the early days of the web, and they did very well at things like text classification. And one of the great things about support vector machines is that neural networks are a mess because it's what is called a non-COVID optimization problem. It has many local optima. you do gradient descent, but you do your gradient descent and I do mine and we wind up in different places. It's terrible, right? You have to spend all this time, you know, tweaking parameters.
41:49You know, as somebody said, trying to get a neural network to work is like trying to balance a pencil on the tip of your finger. That's really a very good analogy. Support vector machines don't have that problem. It's a convex optimization problem. There's one global optimum. You push that button and you get the solution and we all get the same solution. And And this just saves so much time and so much headache that people are like, well, you know, neural networks are obsolete. We don't need that anymore. Goodbye, right? Now, of course, fast forward 20 years and another set of things has happened.
42:21But for all I know, in another whatever, one year or five or 10, support vector machines will come back doing the things that transformers can. And I can even speculate about how that might happen without the problems that the transformers do. So yes, that type of learning now is a little bit falling to the side. There's still some going on. But again, those people are not as hardcore as the patients. So there's less of that happening. But again, for all I know, it's ready to come back anytime. And again, this is an important point. Going back to the 50s, there is this recurring motif that neural networks come along, they do something.
43:03They're very sexy, right? They're very intuitively appealing because this is how the brain works. But then they have problems and they're messy and people keep pounding on it. And sooner or later, somebody comes along and does an analogical version of the solution that is actually better and simpler than the neural network one. So in the long run, I mean, and a famous example of this particularly, there's this Hopfield networks, right? John Hopfield just won a Nobel Prize in physics for Hopfield networks, right? It turns out, Hopfield Networks, and this was discovered in the 80s, like right away, by these guys at this lab at MIT, that it's just nearest neighbor with a number of different bits as the distance measure.
43:46It's this dumb. All that blah, blah, there's a dynamical system with the tractors, blah, blah, blah, like, no, nearest neighbor. And this just has happened over and over again. So I asked, well, maybe it's going to happen again this time. That's interesting. Yeah. And so, I mean, somebody like Hopfield, does he identify as an analogizer? No, he doesn't. So that's my point is that like, so Hopfield is a physicist, a very eclectic physicist that has looked at things in biology and whatnot. And he came up with this. And again, he wasn't the first one, but he was the most influential one. He saw this analogy, great example, he literally saw this analogy between a type of physical material called a spin glass and a neural network, where each atom in the spin glass corresponds to a neural network.
44:39Mathematically, they're the same. Or at least what he did is he formulated a type of network called a Huffield network. There's really the mathematics of these spin glasses, but turned into a neural network. and the point that he was making was that so these spin glasses they have these attractors right they have these these local optima local minimum of energy that as the system evolves it's a bunch of atoms right and their spins are flipping right and and and it will evolve right into different local optima and he said like we can think of each of these as a memory and now we can actually learn the weights that by which these you know spins interact to store you know uh the digits you know zero to nine.
45:21And then when a corrupted nine comes along, it falls into the basin of the nine, right? And this was a fantastically appealing idea, right? Back in the 80s when this came out, people were like, wow. Also because, you know, there's a sociology to this. He was a physicist, he was respectable. And machine learning back then was not respectable, let alone neural networks. And they were like, oh, we're respectable now, right? So in particular, this was a big influence on Jeff Hinton. But, you know, arguably a bad influence because nothing that Jeff Hinton did based on that ever panned out. The neural networks that we use today do completely different things.
45:57But, so Hopfield was not trying to do any logical learning at all. He probably doesn't even know what that term means. But what then somebody else proved was that his whole network was really just a very inefficient way to, let's suppose like when a new nine comes along, right? When a corrupted nine comes along, what I do is this. I count the number of bits that it has in common with the prototypical 0, 1, etc. And hey, it has the most bits in common with 9, so I classify it as a 9, right? This is nearest neighbor. This is pure nearest neighbor, the simplest version you can imagine. And this turns out to be mathematically equivalent to that whole complicated network with its slow evolution towards the local optimum.
46:34So I think it's ironic or maybe even hilarious commentary on that whole, I mean, like all these hundreds of physicists came into machine learning, doing variations on half-field networks, blah, blah, blah. And at the end of the day, it's just doing nearest neighbor. Yeah, that's fascinating. And SVM, support vector machines, what kinds of applications do they have today? They are, so let me backtrack slightly to a question that you asked earlier, which is, but what exactly is a support vector machine, right? And the support vector machine is, maybe I can see it in a couple of steps. So there's a generalization of nearest neighbor, which is actually the most powerful, which is K-nearest neighbor, where you don't just use the nearest neighbor, but you average between the K-nearest.
47:20It's like, who are your three closest neighbors? And what did they have? If two out of three of them had diabetes, then well, maybe you do as well. And if they all had different things, then maybe I don't know, I flip a coin or something. So there's that step. The next step is, it turns out to be very helpful in many cases to give the examples weights. So maybe some of the examples should have more weight because they represent a bigger prototype, right? And in essence, a support vector machine is just a mathematical way to assign weights to these examples. And it assigns the weights in such a way that you form a good frontier between, say, for example, you know, the concept is, you know, cancer, no cancer, right?
48:01And think of the space of patients. and some of them have cancer, some of them don't, you want to find what the frontier is between cancer and no cancer. And what support vector machines do is they try to find the frontier that is as far as possible from any example, so it's a safe frontier. If I have a frontier that has examples very close to it, well, then maybe the frontier is, you know, maybe it should be somewhere else, right? Because it's like, why, you know? Now, I mean, an analogy that I use in the book is, you know, imagine that I give you a map of two countries and I only tell you where the major cities are.
48:38And I tell you to draw me the frontier between the two countries, the border. Right. Where would you put the border? You wouldn't put it right next to the capital of one of the countries, probably. Right. There's exceptions. But the safest thing to do would be to put the border as far away from many of the cities as you can. And so support data machines mathematically solve this problem to come up with the weights for the examples that produce the so-called max margin frontier. And the margin is kind of like this DMZ. Think of Korea. There's a DMZ around the border where nobody lives. And support data machines try to make that DMZ as wide as possible.
49:17So that's at a high level what a support vector machine is doing. And so support vector machines wind up being good for things like, for example, as I mentioned, text classification. Why are they good at text classification? Because the space of words is very large. And treating each of those, like let's say you have 100 ,000 words, right? Treating each one of those words as a dimension in the space that you're predicting in makes it really hard. It makes it really easy to overfit, right? Because any own little world can be, you know, can say, oh, this world by accident correlates very well with you clicking on my app, right?
49:58But it's just overfitting. And the support vector machines are much more robust to that because they have this, like in neural networks, for example, Perceptrons, they will just put the frontier anywhere, literally. As soon as the frontier doesn't misclassify anybody, you're happy. You actually don't care if, you know, if you just have one example here, one positive example here and one negative there, any frontier between them is okay. And the support vector machines are smarter than that. So like, no, I want the max margin one. So they're, for example, much more robust when you're in high dimensions, which used to be the Achilles heel of Neurys neighbor, right?
50:32You would even say that if it wasn't for the so-called cursor dimensionality, Neurys neighbor would have solved this whole problem a long time ago. But it is a very deep problem, and support the machines definitely make a big difference. And it sounds, though, that analogizer is a class, but not necessarily a tribe. I mean, even connectionists you could call analogizers, because they're using neurons in the brain as an analogy. Well, so let me unpack several things that you said there. each tribe has sub-tribes but within connections there's people who do different and believe different kinds of things, some of them very different same with the analogizers, some with all of them again like we mentioned like last time the Neats and the Scruffies are the two main sub-tribes of the symbolists what distinguishes I would say the analogizers is that the sub-classes or the sub-tribes are more independent and talk less to each other, so for example the people who do analogy based stuff inspired in psychology, even within AI, talk very little to the people who do support vector machines.
51:52Again, the support vector machine people are more neat and the other ones are more scruffy. Maybe that's the way to think about it, the neat and scruffy analogizer, but it's a less coherent trial. Now, when you say that the connection is to our analogizers, reasoning yourself by analogy doesn't make you an analogizer. I mean, we are all analogizers by that standard, right? The question is like, is the algorithm reasoning by analogy, right? If I'm Jeff Hinton or John Hoffel, then I make an analogy between the brain and the spin glass, right? That doesn't make me an analogizer because at the end of the day, what I have is a neural network, right?
52:28The fact that I'm using analogies is different from the algorithm using an analogies when it's trying to figure out, you know, what this patient has or not. Having said that, and again, this I think is a very interesting point. When you talk to Jeff Hinton, he says, neural networks are better than symbolic AI because they reason by analogy. But then he never explains how they reason by analogy. This to him is a very strong intuition and one that I buy, right? Again, symbolic AI suffers from the brittleness problem. And one way to overcome that is to bring in analogy. In fact, my PhD thesis was unifying symbolic AI with analogical learning precisely for this purpose, to make the rules softer, right?
53:14I had this analogical matching of the rules and different degrees of abstraction, which again, agrees with a lot of results in psychology and whatnot. So bringing in an analogy as a way to solve the, you know, the brittimless problem is potentially a very good idea. and what Jeff Hinton says and others, but Jeff Hinton is the most famous one, is that that's what neural networks are doing. But then my question is like, well, hey, Jeff, if neural networks are doing analogy, how are they doing analogy? They do this grid-intercent, there's this mess of parameters, and intuitively they're doing analogy.
53:49But I actually have a recent paper explaining exactly how they're doing analogy. It turns out that when you do grid-intercent, And the paper is basically making the proving, a theorem, that shows that every model learned by gradient descent is a kernel machine. With a particular type of kernel, there's a dot product of gradients. But think about this. What you've actually done when you learn a neural network is you've just stored all those examples in a way that is not obvious, but they're there. and when you apply the new library to a problem, what you are mathematically effectively doing is comparing your new example with Trondos.
54:31So I can actually, I have an answer to Jeff Hinton's question. So I have an answer to my own question of Jeff Hinton, which is that, well, how do new libraries do analogy? You know, we know how they do analogy at this point. And I think a lot of consequences are going to flow from this, but we haven't seen most of them yet. Create an oasis with Thuma, a modern design company that specializes in furniture and home goods. By stripping away everything but the essential, Thuma makes elevated beds with premium materials and intentional details. I'm in the process of reorganizing my house, and I'm giving Thuma a serious look for help in renovating and redesigning.
55:17Thuma combines the perfect balance of form, craftsmanship, and functionality. With over 17 ,000 five-star reviews, the Thuma Bed Collection is proof that simplicity is the truest form of sophistication. Using the technique of Japanese joinery, pieces are crafted from solid wood and precision cut for a silent, stable foundation. With clean lines, subtle curves, and minimalist style, the Thuma bed collection is available in four signature finishes to match any design aesthetic. Headboard upgrades are available for customization as desired. To get$100 toward your first bed purchase, go to Thuma. That's T-H-U-M-A dot C-O slash IonAI.
56:16IonAI all run together, E-Y-E-O-N-A-I. So for$100 off your first purchase, go to thuma.co slash ionai. That's T-H-U-M-A dot C-O slash ionai to receive$100 off your first bed purchase.
From the publisher
This episode is sponsored by Thuma.
Thuma is a modern design company that specializes in timeless home essentials that are mindfully made with premium materials and intentional details.
To get $100 towards your first bed purchase, go to http://thuma.co/eyeonai
In this episode of the Eye on AI podcast, Pedro Domingos, renowned AI researcher and author of The Master Algorithm, joins Craig Smith to explore the evolution of machine learning, the resurgence of Bayesian AI, and the future of artificial intelligence.
Pedro unpacks the ongoing battle between Bayesian and Frequentist approaches, explaining why probability is one of the most misunderstood concepts in AI. He delves into Bayesian networks, their role in AI decision-making, and how they powered Google's ad system before deep learning. We also discuss how Bayesian learning is still outperforming humans in medical diagnosis, search & rescue, and predictive modeling, despite its computational challenges.
The conversation shifts to deep learning's limitations, with Pedro revealing how neural networks might be just a disguised form of nearest-neighbor learning. He challenges conventional wisdom on AGI, AI regulation, and the scalability of deep learning, offering insights into why Bayesian reasoning and analogical learning might be the future of AI.
We also dive into analogical learning—a field championed by Douglas Hofstadter—exploring its impact on pattern recognition, case-based reasoning, and support vector machines (SVMs). Pedro highlights how AI has cycled through different paradigms, from symbolic AI in the '80s to SVMs in the 2000s, and why the next big breakthrough may not come from neural networks at all.
From theoretical AI debates to real-world applications, this episode offers a deep dive into the science behind AI learning methods, their limitations, and what's next for machine intelligence.
Don't forget to like, subscribe, and hit the notification bell for more expert discussions on AI, technology, and the future of innovation!
Stay Updated:
Craig Smith Twitter: https://twitter.com/craigss
Eye on A.I. Twitter: https://twitter.com/EyeOn_AI
(00:00) Introduction
(02:55) The Five Tribes of Machine Learning Explained
(06:34) Bayesian vs. Frequentist: The Probability Debate
(08:27) What is Bayes' Theorem & How AI Uses It
(12:46) The Power & Limitations of Bayesian Networks
(16:43) How Bayesian Inference Works in AI
(18:56) The Rise & Fall of Bayesian Machine Learning
(20:31) Bayesian AI in Medical Diagnosis & Search and Rescue
(25:07) How Google Used Bayesian Networks for Ads
(28:56) The Role of Uncertainty in AI Decision-Making
(30:34) Why Bayesian Learning is Computationally Hard
(34:18) Analogical Learning – The Overlooked AI Paradigm
(38:09) Support Vector Machines vs. Neural Networks
(41:29) How SVMs Once Dominated Machine Learning
(45:30) The Future of AI – Bayesian, Neural, or Hybrid?
(50:38) Where AI is Heading Next




