In short
Eye On A.I. Podcast Episode #237: Pedro Domingos Breaks Down The Symbolist Approach to AI
Episode Overview In this episode, host Craig S. Smith interviews Pedro Domingos, a prominent AI researcher and author of *The Master Algorithm*. The discussion centers on the Symbolist approach to artificial intelligence, one of the major paradigms in machine learning. Domingos breaks down the historical context, fundamental concepts, and the ongoing relevance of Symbolic AI in today's technological landscape.
Key Topics Discussed
- Introduction to Symbolic AI
- Definition: Symbolic AI is based on the manipulation of symbols to represent knowledge and infer conclusions.
- Historical Dominance: From the 1950s to the early 2000s, Symbolic AI was the dominant paradigm in AI research.
- Foundational Concepts
- Physical Symbol System Hypothesis: Formulated by AI pioneers Marvin Minsky and John McCarthy, asserting that a system capable of manipulating symbols can exhibit intelligence.
- Key Figures: The founding fathers of AI, including Minsky, McCarthy, Herb Simon, and Alan Newell, all contributed to the development of Symbolic AI.
- Inverse Deduction: The "Master Algorithm"
- Concept: Inverse deduction allows AI to infer general rules from specific examples.
- Importance: Domingos argues that this method is crucial for understanding and implementing learning in Symbolic AI.
- Applications of Symbolic AI
- Expert Systems: Initially popular in fields like medical diagnosis, where rules were codified to assist in decision-making.
- Failures and Limitations: Discussed the knowledge acquisition bottleneck and brittleness problems that hindered expert systems.
- Bridging with Machine Learning
- Machine Learning's Role: Machine learning emerged as a solution to the challenges faced by Symbolic AI by automatically extracting knowledge from data.
- Integration: Modern AI models increasingly combine Symbolic AI with machine learning techniques, blurring the lines between the two paradigms.
- Ongoing Debates
- Symbolists vs. Connectionists: The longstanding debate between proponents of Symbolic AI (rational, logic-based reasoning) and Connectionist approaches (neural networks) continues.
- Current Trends: Domingos emphasizes that the future of AI lies in the synthesis of these approaches, citing successful applications like AlphaGo as examples.
Key Takeaways
- Symbiosis of Paradigms: The integration of Symbolic AI with machine learning and connectionism is crucial for advancing AI capabilities.
- Real-World Relevance: Symbolic AI remains relevant today, particularly in high-level reasoning tasks and applications requiring structured decision-making.
- Future of AI: The successful blending of different AI paradigms will dictate future advancements in the field.
Conclusion Pedro Domingos provides valuable insights into the Symbolist approach to AI, highlighting its historical significance, foundational principles, and the necessity for integration with other paradigms. The episode underscores the importance of understanding AI's diverse methodologies to fully grasp its potential impact on society.
Stay Connected
- Craig Smith Twitter: [@craigss](https://twitter.com/craigss)
- Eye on A.I. Twitter: [@EyeOn_AI](https://twitter.com/EyeOn_AI)
--- This summary captures the essence of the podcast episode, focusing on the key discussions and insights provided by Pedro Domingos regarding Symbolic AI.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Symbolic AI partly derives that name because it's based on this thing called the physical symbol system hypothesis. So AI has four founding fathers, Marvin Minsky, John McCarthy, Herb Simon, and Alan Newell, all of them symbolists. And Alan Newell and Herb Simon were the ones that formulate the physical symbol system hypothesis. And there's different variations of this going back to Turing and also other researchers. But the basic idea is this. If you have a system that is able to manipulate symbols in certain ways, that's all you need for intelligence, nothing else. I'm Peter Domingos. I'm a professor of computer science at the University of Washington in Seattle.
0:38I am a long time AI researcher. I got my PhD in machine learning in the 90s. I'm also the author of the Master Algorithm, which is an introduction to machine learning for a broad audience, and more recently of 2040, a Silicon Valley satire. A lot of my work over the years has consisted in unifying different paradigms within machine learning and in fact these paradigms there's there's five main ones and they've been with us since the beginning since the 50s and they continue to be with us now people these days often equate ai with deep learning but for example in most applications the the learning algorithms that are that work best are actually not deep learning ones uh the things like boosting and and uh you you know, random trees which are symbolic and so on.
1:29So I think it's actually very useful for people who want to understand AI to have this context and to know what these major themes and major paradigms within machine learning are. Because, for example, even within just the last few years, we've seen a whole bunch of symbolic AI come into play in combination with deep learning. And I think we're going to see a lot more of this in the coming years. Okay, great. Well, let's talk about symbolic AI and rule-based systems, expert systems, and all of that, and then how that continues to this day. But if you could start by just defining symbolic AI and then give its origins and some of the major developments.
2:23so symbolic ai for most of ai's history was the dominant paradigm in the 70s 80s 90s um you know ai was almost equated with symbolic a some people would even you know not consider neural networks to be ai which is kind of hilarious now and symbolic ai is ai that each of the different tribes gets its inspiration from different fields symbolic ai is mainly influenced by mathematics and logic and psychology also to some degree. Its idea is that now, you know, like the brain is a mess, but who knows what evolution does, biology is a mess. We need to figure out AI from first principles. And where do those first principles come from?
3:10Mathematics and logic, right? The brain is a logic engine. So a lot of AI is just based on trying to figure out, really, we're not constrained by human intelligence. We should try to figure out what AI is, from what intelligence is and how to do it from first principles. Of course, another big influence in this school of AI is philosophy. A lot of the ideas in symbolic AI come directly from different schools of philosophy. You could even say that symbolic AI is kind of computer-aided philosophy. It's applied philosophy. It's those ideas operationalized on a computer. And symbolic AI is also the one that is closest to the rest of computer science.
3:54This is often an underappreciated but very important point. Most of what we do in computer science is symbolic manipulation. Actually, everything, many people would say, including, of course, the symbolists. So between neural networks, for example, and the rest of computer science, there's a very big gulf. The concepts are different. The techniques are different. The way you think is different. Symbolic AI is actually completely continuous with traditional programming and data structures and algorithms and whatnot. And this, I think, is also part of what made it such a natural choice in the first 50 years.
4:31It's interesting to note that when AI first took off in the 50s, in the beginning, everyone was interested in neural networks. But then its various failures quickly, their various failures quickly became apparent. And then a very big set of people started doing symbolic AI. Something important I think that is worth mentioning here is that symbolic AI partly derives that name because it's based on this thing called the physical symbol system hypothesis. So AI has four founding fathers, Marvin Minsky, John McCarthy, Herb Simon, and Alan Newell. all of those symbolists. And Alan Newell and Herb Simon were the ones that formulate the physical symbol system hypothesis.
5:15And there's different variations of this going back to Turing and also other researchers. But the basic idea is this, is that it says, if you have a system that is able to manipulate symbols in certain ways, that's all you need for intelligence, nothing else. And moreover, and very important to the present day, and to the whole history, the way those symbols are implemented in the physical world is completely irrelevant. So the symbolists believe very strongly and with a lot to back it up, just to be clear, that, but, you know, we can discuss the ins and outs of that because there are some, that looking at the brain at the level of neurons is completely the wrong thing to do.
5:57It's like trying to look at a computer at the level of transistors. God help us if when we're programming in Python on Java, we had to think about transistors would be hosts. So their idea is that you need to formulate intelligence at this level of representations and algorithms, and then how you implement it is an implementation problem. And this was really the guiding principle through most of the history of AI. We can also talk about some of its failures and why we're doing other things, and why, for example, neural networks have become popular again.
6:31For listeners, again, And I think one of the things that confuses some about symbolism or symbolic AI, and it confused me in my early education on this, is, you know, what are the symbols? I mean, but in fact, the term symbols refers to concepts and how you order them in a logical sequence. Is that right? Or if not, can you correct me on that? Actually, that is a great question because, you know, to us in the field, we don't even ask ourselves that because it's such a given. But you're right. That is actually a key question. The symbols, let me start by giving a very simple answer. Words are symbols.
7:19So natural language, when I speak, it's a sequence of symbols. The symbol a, the symbol sequence, the symbol of, et cetera, et cetera. Maybe more relevant here in mathematics, right? When you write F equals MA, F is a symbol, equals a symbol, M is a symbol, and A is a symbol. And so if you think at how mathematics, at how algebra is constructed, an equation, is a bunch of symbols. It's basically some set of symbols combined in some way is equal to some set of symbols combined in another way. And again, the contention is that that's all you need for any computation, including AI. And by the way, nobody actually, nobody significant or credible disagrees with this.
8:01There's this thing going back even before Newell and Simon called the church Turing hypothesis, which is that a Turing machine can do any computation. This is so-called universal Turing machine. You can implement other Turing machines and so on. And again, a Turing machine, all it does is it has a bunch of symbols written on the table, on a table. Again, think of them as words or Greek symbols or numbers or whatever. and all it does is like it takes in some of those symbols and transforms them into another. Like for example, when I do two plus two and transform it into four, that's a small example of symbol manipulation.
8:32And in fact, you could even say, if you want to tweak the connectionists that, and we can talk about how, you know, what they do they feel is different from what the symbolists think they do. But you could say that any neural network implemented on a computer is still symbol manipulation. The symbols are vectors and matrices and tensors and then there's functions like gradients and whatnot, but it's all still simple manipulation. And the manipulation is according to rules of logic or rules of mathematics. Exactly. And in some sense, those rules, if you want to solve AI, those rules are the most important things, like what are the rules by which you operate?
9:18So, for example, arithmetic is a set of rules, but those really derive from other words that are, for example, you know, there's the, you know, piano's axioms of arithmetic from which you can derive a lot of mathematics. But even below that, like you have to operate on these things with rules of inference, with logical rules of inference. Like, for example, a famous one, you know, going back to Aristotle is modus ponens, right? Modus ponens is, for example, if, you know, I say, you know, all humans are mortal and Socrates is a human. Notice this is a set of symbols. All humans are mortal. Socrates is human is another set of symbols.
9:53Modus ponens tells me that from those two sets of symbols, I can derive another one, which is Socrates is mortal. But that modus ponens turns out is not completely general, but there's another one which was discovered in another rule of inference called the resolution rule that was discovered in the 60s, that in some ways is really the foundation of symbolic AI. It's It's a rule of logical inference by which you can derive anything that can be derived. Of course, there are things that can't be derived, but that's another matter. And everything you derive is correct. This is called soundness and completeness.
10:28The folks in symbolic AI, there's a very important split within symbolic AI between the needs and the scruffies. This used to be the big split in AI. The scruffies are between the what and the what? The needs and the scruffies. The needs are the people who want to do everything neatly. And John McCarthy was the leader of that school. And he was probably the dominant one overall. And the scruffies were the people who were willing to do a lot of heuristic things, more inspired by psychology. And Newell, Simon, and Minsky were more on that side of things. But the needs, their whole idea was like, we just have to figure out what the right rules of inference are.
11:07And then we write all our knowledge in a bunch of formulas. And then you just solve everything using that. Right. This in the 80s was the way to AI. And, you know, the people at the time felt they were on the verge, you know, of solving AI. We would have human level intelligence within a decade based on what I just described. Things turned out a little differently. Yeah. And the sort of core... Okay, I'm going to have to go back and edit this, but how does inverse deduction play in what you just said? Very well. So these rules of inference are deductive rules, right? Traditionally in logic or in philosophy, if you will, there's two kinds of inference.
12:01I mean, there's more, but there's two main ones, deduction and induction. And modus ponens and resolution are rules of deduction. They're rules by which I go from general statements to specific consequences. They're how I go, for example, from knowing that all humans are mortal to knowing that this particular human, Socrates, is mortal. That's deduction. but for a year, we also need to learn. In other words, we need induction. We need to go from specific knowledge to general statements. And inverse deduction is the, if you will, the symbolist's master algorithm. It's their favorite way. You could even say, some people would say it's the way, the foundational way to do induction is by inverting deduction, which has a very nice pedigree in mathematics, right?
12:45You define integration as being the inverse of differentiation. You define subtraction as the inverse of addition, et cetera, et cetera. So this notion of having an operation and defining the inverse one, whatever difficulties it might have, is a very powerful one. And so what is inverse deduction? Inverse deduction is going from, look, Socrates is human and he's immortal, and Plato is also human and mortal, and Aristotle is also human and mortal. So maybe all humans are mortal. Notice that this is a risky inference. It could be wrong. Right. But unlike the induction that is always very sound in induction, there is no, you know, there's always a risk.
13:25But in the last few decades, we have actually become very good at quantifying that risk and actually having a theory that says under certain assumptions, you can guarantee with high probability that your inductions are correct. In fact, Leslie Valiant won the Turing Award, the Nobel Prize of Computer Science for just developing this type of theory. And then when you apply this in a system, the challenge was to codify features for each symbol. I mean, if you're looking at a rule-based system or just what you said, Sophocles and Plato and Aristotle, each of those would be written in code that would represent them as a group of features or a feature.
14:30And then you would apply these rules to them to come up with the deduction, with the answer. Let me give you a specific example that might be helpful and then also support some of the rest. One of the big applications of AI in general, but also symbolic AI in particular, since, I don't know, at least the 70s, is medical diagnosis. So let's say you want to diagnose somebody. the features are their symptoms oh you have a fever oh you have whatever high blood pressure oh you have a headache you have etc etc and you know your blood tests give these results etc etc so the patient is described by a bunch of features right again socrates being mortal is the feature of socrates but you know there are many right you all have a lot of different features and part of the job that you have to do to solve a problem using eis like figure out what are the relevant features like you know as a doctor you have to say like well okay you ask a bunch of questions of the patient and then you you know ask for a bunch more tests for example and then what the rules do the rules that are written in logic do is they operate on these features to for example conclude that like oh with these symptoms you must have diabetes or you have whatever right right and and now traditionally in symbolic ai and again if you go back to the 80s when you know there was this previous AI boom, people would write down these rules.
15:53You would interview doctors and say like, okay, so how do you diagnose whatever diabetes? Or when you look at this X-ray of a breast, how do you decide whether there's cancer in it or not? And then you would try to, so language is very informal, right? Not good for a computer. You would try to codify this in this very rigorous, logical way. And then if you did that properly, then when a new patient comes along, you input the symptoms you ask, what does she have? And you get the answer. Yeah. Yeah. And the reason why that was onerous was, first of all, identifying features and then codifying features.
16:36There was a lot of time spent on that. That wasn't even the bigger problem. Actually, there were two big problems that became known as the knowledge acquisition bottleneck and the brittleness problem. The knowledge acquisition bottleneck is that interviewing experts to get their knowledge down is very expensive, and it takes a lot of time. And no matter how long you spend at it, there's always more knowledge that they wind up not telling you. So the cost and the things that like, there's this long tail of knowledge that we all have, but it winds up, these systems wind up, You know, there was this famous system called Psyche that was trying to basically put all the world's knowledge into one set of rules.
17:23It just got bigger and bigger and bigger. And it still failed, you know, in most situations to have the necessary knowledge at the same time that all the needless knowledge was slowing it down and making things very difficult. So the solution to the knowledge acquisition problem was machine learning. Right. Machine learning is no, no, no, no. I'm not going to interview experts anymore. There's too few of them that cost too much. I'm just going to try to extract the knowledge automatically from data. And really, more than anything else, the present success of AI is the result of that. We're shifting from the so-called knowledge engineering mode to the machine learning mode.
17:59Because, of course, once you start, that has its own difficulties, of course. But once you start doing that, as you get more data, your system just gets smarter almost for free. And this is what we've been seeing in the last, you know, two, three decades or four. right we get more data we scale up to them and boom the systems just get better and better it's amazing right now instead of fighting we're actually riding the wave so that's no the knowledge acquisition bottleneck and the machine learning to solve that problem the other big problem uh which which at the time maybe was even the bigger one or the one that kind of stopped things that in their tracks most was the brittleness problem the brittleness problem is that the real world is not black and white like logic wants it to be.
18:41You don't know for sure whether this patient has cancer or not. You have a probability. So one of the first things that they did was that they added these confidence factors to these rules in this so-called expert system. So like, well, with some confidence, then you have this, right? But that was a mess. Often you got wrong inferences. There was a principled way to this, which is probability. But going back to the 60s, people found that trying to do this with probability was just too expensive, literally exponentially expensive. And they gave up on it and they went to all these heuristic methods, but they all had a lot of problems.
19:15And this was not satisfactory solved until graphical models came along, which actually come from another school of computer science, which is the probabilistic statistical vision one. We now today have a very well-developed technology for doing efficient difference with probability by making certain assumptions. And we actually know very well how that relates to the you know, through the symbolic AI. Having the power of both is actually not easy. In fact, one of my main contributions in life was to actually develop a representation that does have the full power of these two things. But to summarize, there was the knowledge acquisition bottleneck problem that was solved by machine learning, and there was the brutalness problem that is solved by probabilistic reasoning.
19:57Right. And the machine learning you're referring to are the connectionists? No. Great question. So here's a very common confusion and very pernicious one, I would say, which is people often conflate symbolic AI with knowledge engineering and machine learning with connectionism. No such thing. There's a whole literature on symbolic learning. There are symbolic learning methods like inverse deduction. again methods that just take the features as symbols manipulate them as symbols and produce new rules as symbols that can get applied you know in deduction as symbols so you don't have to be connections to be doing uh to be doing learning and in fact there's this whole area of ai called logic programming in particular inductive logic programming that um let me just put this way a lot of the problems that the connectionists and deep learning folks and whatnot are very proud of being able to kind of solve these days like the ILP guys solved them 30 years ago.
21:03And of course, they're furious at the fact that people think this is news. But as you can imagine, right, if what you want to do, for example, is solve math problems, this type of endotologic program is the obvious thing to use. And indeed, it has been used successfully. Create an oasis with Thuma, a modern design company that specializes in furniture and home goods. By stripping away everything but the essential, Thuma makes elevated beds with premium materials and intentional details. I'm in the process of reorganizing my house and I'm giving Thuma a serious look for help in renovating and redesigning.
21:47Thuma combines the perfect balance of form, craftsmanship, and functionality. With over 17 ,000 five-star reviews, the Thuma Bed Collection is proof that simplicity is the truest form of sophistication. Using the technique of Japanese joinery, pieces are crafted from solid wood and precision cut for a silent, stable foundation. With clean lines, subtle curves, and minimalist style, the Thuma bed collection is available in four signature finishes to match any design aesthetic. Headboard upgrades are available for customization as desired. To get$100 toward your first bed purchase go to Thuma that's T H U M A dot C O slash Ion AI Ion AI all run together E Y E O N AI so for a hundred dollars off your first purchase go to Thuma dot C O slash Ion AI That's T-H-U-M-A dot C-O slash IonAI to receive$100 off your first bed purchase.
23:12The, the, uh, before we go on to, uh, connectionism, um, the symbolists besides inverse deduction also, or maybe it's a, a way to apply inverse deduction. They use decision trees. Is that right? Can you talk about how those fit and are those unique to symbolism? Absolutely. So, again, this is something that I understand why people get confused because we tend to conflate things. But one thing is the representation that you use. So, for example, the knowledge that you learn in English could be in natural language could represent it in English or in Chinese or in traditional computer science. It could be a program in Java or Python or C, right?
24:03In AI, you know, connection is the type of representation. In symbolic AI, the most common type of representation is rules. If then rules, if this and that, then the patient has that. Right. But another, you know, and in fact, the most popular one is decision trees. Now, how you learn this is another question. So inverse deduction is typically used to learn rules, not trees. But in fact, in practice, the most widely used method is learning decision trees. They're not that different because at the end of the day, decision trees equivalent to a set of rules each. You know, a decision tree is like you start at the root, you ask a question, well, you know, did the patient have a headache?
24:43No. Well, then, you know, let's ask another question. Yes. Then let's ask a different question. And at the end, you produce a prediction. So a decision tree is mathematically equivalent to a set of rules, each one of which is a path through the tree from the roots to the prediction. So they're not that different. But in practice, particularly when you don't have a lot of data, decision tree learning tends to work better. And in fact, these methods, which are still the best ones for most applications, like random forests and boosting, they are sets of decision trees. One thing that we found in the 90s was that it works really well to instead of just learning one model, learning a whole bunch of them and combining them.
25:22It's the wisdom of the crowds applied to machine learning. And indeed, forests of decision trees or combinations of decision trees are, I would say, a surprisingly effective and general method to learn things. Right. So in the application, the first step is to define features, right? And that's done through interviewing experts or collecting data from experts. Not necessarily. I mean, I sympathize with that. But in today's world, those features, at a first level at least, are what you have in the data. Yeah. You can talk about the features you'd like to have, but there's the features that the data has.
26:09Now, of course, you can go out and collect those features. But I have a database of patient records. Those are the features. What's in those records? Or I'm a company and I have a database of sales of my employees and I know the employees, you know, various whatever demographic characteristics and qualifications. Those are the features. Now, you know, I can also go out if it's worth it, and it often is, and collect features deliberately. And then I have to think about what those features are and whatnot. For example, you know, a self-driving car, its main feature is the video camera. Or to be precise, each pixel is a feature.
26:43So a video camera gives you a million features, a million features, each of which is a pixel. So in today's world, the features come from the data. Right. I mean, and certainly that was the power of neural networks is there was no longer any feature engineering. The system identified features on its own. I'm laughing because, yes, that is one of those myths that unfortunately persists. So first of all, so the myth or like this common view is that, oh, the great thing about neural networks and in particular deep networks is that they discover their own features, whereas the other paradigms don't.
27:25This is just false. It's false in two ways. They need data features just like everybody else does. They also operate on the output of the camera or the symptoms of the patient. There's no way around that. So that basic level of features is the same for everybody. Right. Even if you were doing traditionally, yeah, you would need those to run your rules on. Right. But then learning. Right. The notion is like, oh, deep learning does this magic in which it invents new derived features. Because, of course, the problem with the raw features like pixels is that they don't carry a lot of information. It's hard to get what you want from them.
27:55Is there a cat here or not? So often you need intermediate features and inventing those is the real amazing thing to do. And deep learning has some ability to do that. But number one, much less than people assume it does. And, you know, we have very concrete empirical and theoretical evidence for this at this point. But also, and more important, all the other paradigms also have their own way of discovering features. In fact, there's a whole soft field of symbolic learning called predicate invention that is their version of discovering features. And in statistical learning, there's latent variable discovery, et cetera, et cetera.
28:29So no, deep learning does not have a monopoly on discovering features. Okay. So in symbolism, once you add machine learning, you no longer had to handcraft features. Is that right? You, I mean, there's always, that's a good question. There's always a benefit to handcrafting features if you can, right? And the real art in machine learning is you don't want to be duplicating what's in the data. Stuff that can be easily inferred from the data, you're wasting your time, right? The data knows more than you do. But the problem, and this is the problem, is that there's a lot of stuff that, and then there's also stuff that no matter what you do, you can never get it unless you go and collect new data, right?
29:18But the interesting problem is this middle part where there is information in the data, but the algorithm doesn't necessarily know how to extract it. So if you can tell it a way to extract it, that is a great win, right? And a lot of creativity can come into this. And by the way, another sort of like related myth is that, oh, in neural networks, people don't do feature engineering, right? In other types of machine learning, feature in practice in a lot of applications, what's called feature engineering, which is creating the features and et cetera, et cetera, is a big part of the whole exercise.
Read the full transcript
29:49And the notion is that, oh, you know, with deep learning, you don't need to do that. that's not true. What happens is that what people in other areas call feature engineering is what, you know, in neural networks people call architecture engineering. When you're defining the architecture of your network, what you're doing is really mathematically, you know, the equivalent thing to what the others do when they create the features. You know, the neurons in your hidden layer are the right features and the, you know, attention, blah, blah, blah. These are all the right features. And indeed, coming up with them is very important.
30:18Right. I guess I still associate symbolism and symbolic AI with rule-based systems where people were cataloging all of these rules and then organizing them into decision trees. and it's moved way beyond that where algorithms, symbolic AI algorithms, you feed them data and they develop or identify features in the data and develop a logic that fits the data or what's the process then once you're beyond the old experts? systems. They invent their own rules, right? You can imagine all humans are mortal being written down. I wrote down that all humans are mortal. Or I can infer that from the data. But at the end of the day, a rule is a rule, so it actually doesn't matter where it came from.
31:27And at some level also, for these purposes, whether the rules are just rules or they're organized into a decision tree, it doesn't really matter. It's still a set of rules. Yeah. You know, there was, and then we'll go on to connectionism, but what was the most advanced or what kinds of systems were the most advanced or are the most advanced purely symbolic AI systems? So today, for practical purposes, it is random forests and boosting that are, you know, the most advanced and most widely used. There are a lot of problems in the world today. If you look at Kaggle, right, it's this website that runs competitions, machine learning competitions.
32:09You're a company, you know, you put up a problem and the prize and, you know, the biggest winners are these types of systems. They are actually not very sophisticated in many ways. The representations that they use are fairly simple-minded, but, you know, but they work for these problems. Traditionally, you know, In the 80s, for example, there were notable successes of symbolic AI of expert systems in areas like medicine and configuring computer systems and prospecting and things like that. But really, really, that stuff never really took off. Now, the poster child of symbolism was this project called Psych, which was this guy, Doug Lennett, whose plan was to encode all the world's knowledge into one big knowledge base.
32:54and that would be the foundation for AI. And I'm smiling at this now, but at the time, this was the thing. Like Marvin Minsky famously said that, you know, AI grad students should just stop doing what they're doing and start entering rules into psych. That didn't make him popular, but it captures the spirit of the times. And Doug Lennett, you know, he had the paper written in, I don't know, late 80s saying, we will reach human level AI within a decade using psych. and you know whatever 100 000 rules will be enough and it became a million and then and then and then millions and you know and you know still hasn't solved it right but what what was the the the physical process for for creating for cataloging rules or encoding human knowledge yeah i mean like psych right most of psych's employees were knowledge interest there were people literally whose job was to write down knowledge in the form of rules, was to translate what we know in natural language into rules that psych could use.
33:57And they employed hundreds and hundreds, maybe thousands at some point of people just to do this. And they're writing it down in natural language or in computer code? No, because in natural language, we have text for that. Back then there was no web, but we have the web, we have books, right? The problem is that computers don't understand natural language. So you have to write it down in logic. So in a way, what the exercise was, was logic, you know, for people who don't know it, in some ways is like natural language. In principle, it can express anything that you can express in the language, but it's a formal language.
34:34It's like mathematics. It's like an algebra for concepts. So you know how to operate. Like, I don't, you know, a computer doesn't know how to operate on natural language, but if you give it logic, you can use a theorem prover, like these rules like resolution or not, to extract the consequences of that knowledge. So logic in a way is like natural language, but stated more formally, stated very rigorously. Yeah. Well, can you give us an example of some knowledge that has been encoded in logic that would be part of psych, that would be in this massive compendium of human knowledge? Let me give you two very different examples, which may be helpful in different ways.
35:22We all learned the algorithm for addition in elementary school. I gave you two features. They're the numbers I want to add. And then there's a sequence of very precise steps, which we all know how to do, by which we turn those two numbers, two plus two into the number four. right and and that there was a rule of inference that you can think of it as a rule of inference which was the algorithm for addition right and again ironically the you know the gpts of two they don't know how to do that they do billions of computations you know every minute or second but you ask them to add two numbers and if the numbers are long enough they fall flat which is you know so your pocket calculator in some ways is smarter than than gpt than chat gpt which is which is kind of ironic, you know, so this is, this is one kind of example, but, but a very different kind of example is, so let's take the example of, you know, Socrates is mortal and humans are mortal and whatnot.
36:18How, how is this represented in, in, in logic? So in logic, you will have symbols for objects. So, you know, the symbol Socrates, right, is a symbol, right? It's a set of bits, a sequel, you know, it's a bit string on the computer or, you know, whatever, ink on the paper, but it represents a real entity in the real world, Socrates, right? So in formal languages, and in AI, this used to be very important. And certainly it is in a lot of computer science. There's the syntax and the semantics. The semantics is what the syntax refers to. So the symbol Socrates is a piece of syntax. The semantics of that is the man Socrates that lived 4 ,000 years ago or whatever, 2 ,500 years ago in Greece, right?
37:03And then I also have symbols for properties or for relations like mortal is a property and then i write you know mortal of x means that x is mortal so i write mortal open parenthesis socrates close parenthesis this means that the object represented by socrates has the property of mortality right and often and more interestingly this can represent relations like you could say for example friends friends, you know, Socrates, Plato means that they were friends, right? And so friends is a relation, is a property, right? And now you can write, for example, a conjunction, and now there are the so-called connectives, conjunction.
37:42Socrates is mortal and he's friends with Plato. I would write, you know, mortal, parentheses, Socrates, and there's a symbol for and, friends, Socrates, comma, Plato, right? And I can write like no end of stuff like this, more and more and more and make the language richer and then the inference becomes more complicated and this is really what psych was engaged in on on a very large scale or was i think psych is still going on yeah yeah uh and and i'm realizing we're probably not going to have time to get into connectionism on this call uh fully but uh there has been an enduring actually it seems to have quieted down now but certainly for a while between gary marcus and and jeff hint in this very heated debate about between symbolism and connectionism can you describe what that debate was and and why it was so heated i mean i think it's gone away because of ai systems now are blending all different paradigms.
38:53Yeah. So there is a very big debate, and it has been going on for a long time, since the 50s. It is really a core part of the history of AI. And Gary and Jeff are just two representatives of this. So Gary, of course, is very much a symbolist. His background is in psychology, but he was a student of Steve Pinker, who's a Chomsky. And Chomsky is a big symbolist. you know he's not an e.i. guy but Chomsky's view of language and psychology is very consistent with you know the symbolist you and indeed that you know Chomsky and Minsky were both professors at MIT and this type of thinking was very associated with MIT uh CMU and Stanford were the three you know big places and so Gary comes from that tradition and and and Jeff of course is the number one connectionist in the world right he's been doing it since the 70s and and and um And this quarrel has been going on for a long time.
39:48And what each of these sides is always telling the other is, look at all the things you can't do. So back in the 80s, there was also a resurgence of neural networks. But then people like Pinker, again, back then, you know, Gary Marcus was maybe not even his student yet, made some very effective arguments of like, no, you're not going to solve it yet with these neural networks because look at all these language problems that they just can't solve. And the truth is at the time, you know, he won the argument because they really couldn't. And then there was a lot of work on trying to overcome this and whatnot.
40:24And we're in a different place now, but that argument still goes on. So if you talk to Gary Marcus, he will have a bunch of criticisms of connectionism, some of which I think are not on the mark, but some of which are. So, you know, it drives Jeff nuts. Right. But then there's also like, you know, what people like to do these days and often very successful. are like, okay, tell us the specific people, you know, on the connectionist side, tell us one, you know, give us a task, right? There are these things, for example, called Winograd schemas, and like, and now we're going to solve it with a neural network.
40:56So there, we solved that one, give us another one, right? And this is still ongoing. Now, my opinion, of course, is that this quarrel will only end with the unification of the two paradigms. And in fact, as you alluded to, this is already what's happening, right? You know, a one is adding reasoning, surprise, surprise, and discrete search and whatnot to LLMs, right? LLMs are a connectionist machine. And also, if you look at things like AlphaGo and whatnot, if you look closely at the big successes of these methods, there's always more, you know, it's not just connectionism. There's symbolic elements in there, sometimes, you know, other schools that we can get into.
41:31But this, I think, is the reality. Yeah. And is work ongoing in purely symbolic AI, or is it at this point a pretty understood discipline and it's used as a tool in a broader AI context? No, the work continues because honestly the true believers in each of these paradigms will die before they give it up. Just like the connectionists, you know, and we're all grateful for that, they never gave up on it even through their dark days in the 80s and the 90s. Now, of course, if you looked at an AI conference in 1980, it was all symbolic AI. And these days, it's very little symbolic AI. But that's still a lot, just to be clear.
42:20If you go like AAA or ICHCA, you'll see a lot of symbolic AI there of many kinds beyond what we just talked about. But I would say where most of the action is, not surprisingly, is in people on the symbolic side, is in people trying to combine the symbolic AI with the connectionism because they realize that connectionism has certain strengths, which they don't know how to reproduce. So, you know, for example, if you talk to someone like Gary Markins, he doesn't say, you know, let's throw away, you know, all this deep network stuff. It's like, no, no, no, we need to combine it with the symbolic stuff, which again, if you look at every decade in the AI since the 50s, there's a dominant paradigm.
42:59And then the other ones talk about combining their paradigm with the dominant one. Right. So for all I know, next decade, the dominant paradigm will be symbolic AI again. And then the big thing in connections would be combining it with symbolic AI or whatever it might be. Beijing, like, again, like, you know, 20 years ago it was Beijingism and then kernel machines and whatnot. And indeed people had, you know, connectionist, blah, blah, blah, symbolic additions to kernel machines. So this is probably what we're going to continue to see. Yeah. I mean, you mentioned AlphaGo, which was a combination of symbolic AI and reinforcement learning.
43:39Am I wrong on that? It was actually a combination of three things, at least, or three main ones. One is what is called Monte Carlo Tree Search, which is symbolic AI. This was what the symbolic people were using to solve Go before DeepMind came along. So they used that, right? They didn't throw that away, right? And then there's neural networks in particular, they use a convolutional neural network, which is something for vision, to understand the board position this in fact was their big you know i remember them is saying that they were going to this is like this is brilliant because you do need the best way to approach my first project as a graduate student in ai was you know a a program to play go and it's like you can't play go the way you play chess and treating it as an image where each board position is a pixel is is absolutely you know right on so they did that and then they have reinforcement learning again And reinforcement is one of the oldest ideas in machine learning, but it really has these three components and they all play a big part.
44:40Is there more that you can say about symbolic AI and its applications today? I mean, AlphaGo and then AlphaZero, and I can't remember the series of models. but fast forward to today what what sort of cutting-edge systems are using symbolic ai so here's an important point that we haven't touched on yet but is worth knowing there is a rough division of labor between these paradigms as to what they're best for right and the rough idea which again shouldn't surprise anyone is that but there are caveats But the rough idea is that connectionism is better for system one tasks, like low level things like perception, motor control.
45:34I mean, for vision, language, understand vision, speech, understanding, things like that. You know, the symbol, the symbolists don't even try to do that. I mean, you know, connectionism rules there. But for higher level things that start with language understanding and reasoning and planning and solving problems, this is, you know, system two problems that require thinking. and again it's not surprising because you know stuff that was inspired by the brain of course is better at the low level stuff that is what most of the brain is doing and stuff that's inspired by you know logic and reasoning is better at logic and reasoning right so if you look at a lot of for example there are these things called SAT solvers and theorem provers right these are systems to do deduction on a very large scale very efficiently right and for a lot of problems like for example you know in software verification in integrated circuit layout, in planning.
46:24I'll give you a concrete example, a little low, but an eloquent one. In the Gulf War, the U.S. deployed 400 ,000 soldiers in their entire support system very quickly. This was a major feat of logistics that was completely beyond what operations research and whatnot could do. And the main thing powering this was symbolic AI planning systems, figuring out what do I need here and what should go where, and et cetera, et cetera. So symbolic AI is very good at that type of thing and continues to be today. Yeah. And you mentioned GBT-4-0, no, I'm sorry, GBT-01, the reasoning model. And you were saying that they employ symbolic AI as part of the inference.
47:17well i don't know what they employ because they're not telling us yeah but i mean i've talked to the people who do this right and i know their background and what they say and what they know it's like you know i i can form an informed guess i don't know if they would call what they're doing symbolic ai they would probably resist doing that because for pr purposes it's not very uh it's not very useful but they definitely i mean so the large language model at this point is a substrate the large language model per se does not solve math problems well you know this has been or reason etc so clearly something else is needed right so now what they're doing is they're grafting on top of that a lot of these techniques that come from symbolic ai and traditional computer science like discrete search like looking for things to chain together to get the conclusion that you want and again they're not telling us exactly what they do but it's this in one form or another so whether they call it symbolic ai or not it is symbolic ai
From the publisher
This episode is sponsored by Thuma.
Thuma is a modern design company that specializes in timeless home essentials that are mindfully made with premium materials and intentional details.
To get $100 towards your first bed purchase, go to http://thuma.co/eyeonai
In this episode of the Eye on AI podcast, Pedro Domingos—renowned AI researcher and author of The Master Algorithm—joins Craig Smith to break down the Symbolist approach to artificial intelligence, one of the Five Tribes of Machine Learning.
Pedro explains how Symbolic AI dominated the field for decades, from the 1950s to the early 2000s, and why it’s still playing a crucial role in modern AI. He dives into the Physical Symbol System Hypothesis, the idea that intelligence can emerge purely from symbol manipulation, and how AI pioneers like Marvin Minsky and John McCarthy built the foundation for rule-based AI systems.
The conversation unpacks inverse deduction—the Symbolists' "Master Algorithm"—and how it allows AI to infer general rules from specific examples. Pedro also explores how decision trees, random forests, and boosting methods remain some of the most powerful AI techniques today, often outperforming deep learning in real-world applications.
We also discuss why expert systems failed, the knowledge acquisition bottleneck, and how machine learning helped solve Symbolic AI’s biggest challenges. Pedro shares insights on the heated debate between Symbolists and Connectionists, the ongoing battle between logic-based reasoning and neural networks, and why the future of AI lies in combining these paradigms.
From AlphaGo’s hybrid approach to modern AI models integrating logic and reasoning, this episode is a deep dive into the past, present, and future of Symbolic AI—and why it might be making a comeback.
Don't forget to like, subscribe, and hit the notification bell for more expert discussions on AI, technology, and the future of intelligence!
Stay Updated:
Craig Smith Twitter: https://twitter.com/craigss
Eye on A.I. Twitter: https://twitter.com/EyeOn_AI
(00:00) Pedro Domingos onThe Five Tribes of Machine Learning
(02:23) What is Symbolic AI?
(04:46) The Physical Symbol System Hypothesis Explained
(07:05) Understanding Symbols in AI
(11:51) What is Inverse Deduction?
(15:10) Symbolic AI in Medical Diagnosis
(17:35) The Knowledge Acquisition Bottleneck
(19:05) Why Symbolic AI Struggled with Uncertainty
(20:40) Machine Learning in Symbolic AI – More Than Just Connectionism
(24:08) Decision Trees & Their Role in Symbolic Learning
(26:55) The Myth of Feature Engineering in Deep Learning
(30:18) How Symbolic AI Invents Its Own Rules
(31:54) The Rise and Fall of Expert Systems – The CYCL Project
(38:53) Symbolic AI vs. Connectionism
(41:53) Is Symbolic AI Still Relevant Today?
(43:29) How AlphaGo Combined Symbolic AI & Neural Networks
(45:07) What Symbolic AI is Best At – System 2 Thinking
(47:18) Is GPT-4o Using Symbolic AI?




