#212 Thomas Dietterich: The Future of Machine Learning, Deep Learning and Computer Vision

9 Oct 2024 · 56 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Eye On A.I. Podcast Episode Summary

Episode Title

#212 Thomas Dietterich: The Future of Machine Learning, Deep Learning and Computer Vision

Host

  • Craig S. Smith, Longtime New York Times Correspondent

Guest

  • Dr. Thomas G. Dietterich, Pioneer in Machine Learning

Episode Overview This episode features Dr. Thomas G. Dietterich, recognized for his significant contributions to machine learning, particularly in novel category detection, open set problems, and deep learning paradigms. Throughout the discussion, Dietterich shares insights on the evolution of AI from early rule-based systems to modern machine learning techniques and emphasizes the challenges and future directions for AI and machine learning technologies.

---

Key Topics Discussed

  1. Journey in Machine Learning
  2. Dietterich's background and entry into machine learning since 1977.
  3. The historical context of AI development from the late 1950s through early rule-based systems and expert systems.
  1. Evolution of AI Techniques
  2. Transition from traditional expert systems to machine learning.
  3. The importance of training examples over expert-driven rule coding.
  1. Multiple Instance Problem in Drug Design
  2. The challenge of using weakly labeled training data in drug design.
  3. The analogy of molecules as keys with different shapes binding to targets.
  1. AI in Sustainability
  2. Work in computational sustainability, focusing on ecological and economic balance.
  3. Importance of new materials and drug discovery for transformative economic impact.
  1. Novelty Detection and Open Set Problems
  2. Addressing the problems of AI systems recognizing novel categories and anomalies.
  3. The need for AI to adapt to unexpected situations and inputs.
  1. Deep Learning and Computer Vision
  2. Impact of deep learning on computer vision, including representation learning.
  3. Limitations of deep learning in recognizing unfamiliar objects.
  1. Foundation Models and Self-Supervised Learning
  2. Advocating for training AI systems to recognize a wide array of object categories.
  3. The potential of self-supervised learning to enhance recognition capabilities.
  1. Ensemble Learning and Large Language Models (LLMs)
  2. Evolution of ensemble methods in the context of expensive models.
  3. The shift towards using selective approaches in evaluating multiple models.
  1. Reinforcement Learning in Real-World Applications
  2. Applications of reinforcement learning in wildfire management.
  3. Multi-agent reinforcement learning for liability policies in land management.
  1. Symbolic Regression and AI’s Role in Scientific Discovery
  2. Use of symbolic regression for interpretable AI in scientific applications.
  3. Challenges of integrating deep learning with symbolic reasoning.
  1. Future Directions and Open Challenges
  2. The ongoing need for advancements in machine learning theory and practice.
  3. Importance of safety and reliability in AI, particularly in critical applications.

---

Key Takeaways

  • The AI field is continuously evolving, and many open problems remain for future exploration.
  • Novel materials and drugs are predicted to have a profound impact on the economy, more so than current AI applications.
  • The integration of symbolic reasoning with machine learning models may yield more interpretable and reliable AI systems.
  • The challenges of novelty detection highlight the limitations of current AI systems, necessitating stronger methods for recognizing unfamiliar inputs.
  • Engagement in AI research is encouraged, with numerous opportunities for innovation and discovery.

---

Closing Thoughts Dr. Dietterich emphasizes the dynamic nature of AI and encourages new talent to join the field, highlighting the exciting challenges and prospects that lie ahead in artificial intelligence research.

---

Stay Updated

  • Craig Smith Twitter: [@craigss](https://twitter.com/craigss)
  • Eye on A.I. Twitter: [@EyeOn_AI](https://twitter.com/EyeOn_AI)

Episode Timestamp Highlights

  • 00:00 - Introduction to Thomas Dietterich's Journey
  • 02:34 - Early Days of Machine Learning
  • 05:41 - AI in Sustainability
  • 12:01 - Foundation Models and Self-Supervised Learning
  • 34:44 - AI in Wildfire Management
  • 50:12 - Closing Thoughts on AI Challenges

This detailed summary captures the essence of the discussion while highlighting the key concepts and arguments made by Dr. Thomas Dietterich, as well as the forward-looking perspective on the future of AI and machine learning.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00One of my main findings has been, for instance, in computer vision. One of its source of power is that it learns its own representations. When we first started building machine learning systems, we wouldn't try to write down the knowledge of experts, but we would still talk to experts about, well, what are the features that you look at? What are the things you're looking for? And then we would try to code up feature detectors that would detect those and feed that into our learning algorithms. Okay, so my name is Tom Dietrich, and I have been working in the area of machine learning since about 1977.

0:34So actually, the field of machine learning kind of got its start in the late 1950s by a guy named Arthur Samuel, who was at the time working for IBM and built the first program that played the game of checkers. at a sort of not a world-class level, but a pretty good kind of hobbyist type level player. But in the 1960s and 70s, the focus of artificial intelligence was not really on learning at all. It was mostly about trying to capture human knowledge, write it down in some programming language, or we use these things called rule-based systems. And in the early 1980s, there was a previous AI boom around things called expert systems, which were these hands.

1:33You would interview an expert in a domain, say maybe you wanted to identify mushrooms or to tell the poisonous ones, or you wanted to diagnose blood diseases or something like this, and then try to encode their knowledge in these, now we would call them domain-specific languages, DSLs, specifically to try to capture that knowledge. And then there was a reasoning engine that could run and diagnose new cases. So I was in this kind of small, tiny sub-community of people that thought that we could teach computers to do this by giving them training examples rather than by talking to human experts. And the very first machine learning workshop was held in 1980 at Carnegie Mellon University.

2:22I was a master's student, or I guess I just started my PhD, and we had about 30 people there. So we can fast forward to today. So over the years in my career, I've worked on a wide variety of both fundamental and application problems. my favorite mode of operating would be to to find an interesting application problem and then identify a novel machine learning problem that needed to be solved so in the 91 92 period i worked in a startup pharmaceutical company where we were trying to use machine learning to uh optimize drugs so drug design we were about 30 years too soon the company was not a success but uh but that's where i stumbled upon a problem that's now known as the multiple instance problem which is uh you know in the traditional machine learning setting i give you an example say a picture of a bird and i say this this bird belongs to this species and i'm trying to teach the computer to recognize i don't know bluebirds or or whatever um but uh with drug design the idea was to say, well, here's the molecular structure of a molecule, and it is either actively binding to the binding site target for the drug, or it is not.

3:47But the trouble is that molecules actually wiggle around a lot, right? They are not rigid structures. And so we would not know which of the different shapes that the molecule could adopt was the one that's actually binding. If you think of it as a key in a molecule isn't a single key, it's sort of like a collection of keys on a key chain. And we know in the lab, one of those is opening the lock, we just don't know which one. And so that was an example of what we would now call a weekly labeled training data. Because we don't we have, it's not a perfect training example that says here's, here's the here's the picture.

4:24And here's the the category. So that was one of the things. And so, you know, over my career, then I've worked on problems in ecology and sustainability. Oregon State University, where I'm located, is a real giant in the area of ecology and oceanography and all of these atmospheric sciences, these kinds of things. So I've collaborated with a variety of people in those fields. And in collaboration with Professor Carla Gomez at Cornell, she and I led a series of large grants from the National Science Foundation to launch a field that we called computational sustainability. But it was really doing computer science research to try to promote environmental and economic sustainability, social sustainability.

5:25Probably the most important direction that's coming out of that was something I was not involved in, but which is the application of computer science and artificial intelligence to the design of novel materials. And now we are seeing big applications in material science. And despite all the excitement about large language models and computer vision, if I were betting what's going to be the thing that will really have a transformative impact on the economy. It's going to be the development of novel materials and novel drugs, because those are going to greatly improve our capabilities, for instance, in industrial processes, in medicine, and perhaps also in things like carbon capture.

6:09So I think that that, but I'm getting off from the topic of my story. So my story then continues to, for the last roughly 15 years, I've been worrying about what we might call the Rumsfeld problem of the unknown unknowns. So when we build an AI system or really any computer system, machine learning, in fact, is fundamentally a kind of backwards looking technology, right? We collect training data from the past and we've and then we teach the computer to be able to do well on that training data and then we hope that the future is is like the past and so that the the uh i don't know if we think about the diseases that a medical system is seeing that there are no new diseases that we just we see the same diseases we've always seen or in the self-driving car we see the same kinds of other vehicles, of bicycles, of whatever that we've seen in the past.

7:12And the trouble is, of course, that the world isn't stationary. It's not static like that. And so this, I feel like, is one of the fundamental vulnerabilities or weaknesses of our AI systems is that because they are sort of fundamentally conservative, backward-looking technologies, they need to become much more able to deal with novelty that they're looking at a patient that's got a new disease. I mean, if we imagine that the chest x-rays and the first time a person with COVID-19 had a chest x-ray, the computer should have said, this is different from anything I've seen before. And so I've been working on that.

8:00It's sometimes known as the novel category problem or the open set problem. There's a lot of technical names for it. But basically, how do we deal with novelty? And so I've worked on that in the setting of sort of cybersecurity and insider threats in organizations. I've looked at it in the context of computer vision, sort of generically, and also in industrial processes. So for instance, detecting that something has gone wrong in a welding machine. When it's made a weld, you want to detect that there's something wrong with that weld. It doesn't look normal and flag that. So the award that I was given was kind of a, I don't know, a lifetime career award for a long body of work in machine learning and artificial intelligence.

8:54Yeah. And most of, I mean, after expert systems, for example, the novel material work or some of the robust AI work, are you still primarily looking at supervised learning systems? or are you now looking at them through the lens of transformer-based models? Well, so I've mostly worked on the supervised problem, but one of my main findings has been, for instance, in computer vision. One of the problems with deep learning is that, you know, one of its source of power is that it learns its own representations. So it sounds like you're familiar with the problem of feature engineering. When we first started building machine learning systems, we wouldn't try to write down the knowledge of experts, but we would still talk to experts about, well, what are the features that you look at?

10:07Do you look at the color? Do you look for yellow spots on the corn plant? Or what are the things you're looking for? And then we would try to code up feature detectors that would detect those and feed that into our learning algorithms. And we really had trouble getting computer vision to work terribly well with that. And so when the ImageNet challenge came along, the performance was not very good. And then when ideas from deep learning were sort of re-injected into the field in the 2011-2012 time period, they just did vastly better on these computer vision problems. And it was because they were able to learn their own representations with the layers of deep learning, the deeper layers could learn to effectively to detect interesting image patches and patterns that should be present in order for it to be a giraffe versus a, I don't know, a German shepherd or whatever.

11:10and um but what we've discovered is that when it comes to novelty and detecting something new uh those representations don't work very well the the deep learning methods learn to represent the things that they they're kind of lazy they only learn to represent the things they're forced to represent and so if you haven't given them any training examples for say an elephant um then And when they're first shown a picture of an elephant, the trunk and the tusks maybe are not really even represented by the deep network. It maybe represents the feet, eyes, toenails or whatever, legs. But it doesn't say, oh, my goodness, there's this huge thing growing out of the face of the elephant because it hasn't seen anything like that in the past.

11:58So it hasn't learned representations for that. The one antidote for that, though, is not only to train the system on the, I don't know, the objects of interest. So, you know, not just the diseases in the x-rays or the bicycles and so on in the self-driving car, but to take a foundation model approach and try to train these systems to recognize everything we can think of. you know, all the different kinds of objects in the world, and perhaps to do that with a self-supervised strategy. And the idea there is to try to capture as much of the variability of the world as possible, so that when we do see something, when it is encounters something new, it can say, okay, well, maybe I've never seen an elephant before, but I have seen snakes.

12:50And so that thing looks kind of like a snake, except that it's got legs. So that combination of being a snake with legs, that's very low probability. It must be something new. And that's the kind of thing that I think that is the path forward. Yeah. It doesn't mean it's guarantees, but, you know, at least the system has learned to see everything that we know how to see. And I think that's the path forward. Yeah, yeah. I've spoken a lot to Yanla Kun on self-supervised learning models and world models. And it looks like, to me, like the pre-trained transformer models really accomplish that. that there's very little, if any, labeling of data.

13:48They just ingest the data and build their own representations of the world that then you can query. So in your latest work, how much has generative AI or transformer models, what kind of a role do they play? Well, I would say in computer vision for object recognition, we've mostly been using the vision transformer models, but I'm not sure that they're significantly better than the convolutional neural networks. You know, images are not sequences of tokens, and we kind of have to artificially break them into regions in order to tokenize them. And the CNNs also incorporate some knowledge about the world, that it is two-dimensional and that it is translation invariant.

14:45So you move the camera around, it doesn't change the giraffe into an elephant. It's still a giraffe. And so I don't think the transform models have had that much impact in computer vision for object recognition from a single frame. But once you put time into the picture and you're looking at video or something like that, then the transformer model has this tremendous ability to flexibly look back into the past and find interesting relationships over time. And capture that idea that in some sense that if I change the frame rate in a video from 10 frames per second to 40 frames per second, And that also doesn't change the world, but the transformer can deal with that instead of saying, well, I just look two frames back.

15:33Well, that won't work if I triple the number of frames per second. It needs to now look six frames back or whatever. The attention mechanism can do that. And so I think transformers are extremely powerful as an exciting tool for us. But they tend to take more data to train and they're more expensive to train. And so what the right combination is for any particular application is kind of an open question. And you talked about multi-instance learning. Can you talk about that a little bit? Because, and I don't know if it's at all related, but increasingly with transformer models, the trend is toward not a single model, but stacking models or having models debate each other.

16:36So I was curious what multi-instance learning was in that context. So I think where multi-instance learning comes up, for instance, in computer vision, as I said, it's a kind of weak supervision. So, you know, if you were training, so I've worked on, for instance, data from camera traps in Africa, right? So you put out a camera with a motion detector, and when an animal walks by and sets out the motion detector, then you take a few shots, still frames. And now we want to, you know, count up, find all the animals in those images, decide what their species are, and count them up. um and uh but but when we add so the the ideal we if we were creating making this into a supervised learning problem the ideal would be to have a person you know outline each animal you know each of the gazelles or whatever in the image and say that's a thompson's gazelle this is a thompson's gazelle this is a whatever um that would be very precise uh labeling but it's very tedious.

17:43And so what we normally do is, is just say, show them the image and say, how many gazelles are in the picture, and they count them up and they say eight. Okay, so now we know there are eight gazelles in this picture, but we don't know where they are in the image. And that's, that's another example of this multiple instance problem is that it's really, it's just a form of weaker labeling. And so now the, the vision system has to generate a bunch of hypotheses for what might be the gazelles and if they say oh there and then there's two elephants and there's also a zebra okay you know you have to compare a bunch of different pictures and figure out oh the zebras are these stripy things and they you know and so on and so statistically the systems can learn that they they need more data you know we make it up on volume so more data but each individual image provides less information yeah now what you were describing is more of let's say an ensemble approach where we combine multiple models in order to improve our accuracy on data.

18:48And of course, right now, so the traditional ensemble approach had a big flowering back around year 2000. And my most cited paper actually is just a kind of tutorial paper on why ensembles help and why they're so wonderful. They're kind of the cheapest way to win a Kaggle competition or something like this. But the kind of ensembling that we're seeing in deep learning is a little different because the premise of ensemble learning was train 100 or 1 ,000 really cheap models and have them all vote. But LLM is not a cheap model. Even a small language model is not a cheap model. And, and you have to pay the inference price of evaluating, running your input through all those models before you combine.

19:36So people are taking a much more selective approach, like in the mixture of experts approach, where given the input, it tries to learn which experts might be relevant to evaluate. And you can also think of these very deep, like the large language models, these very deep transformer networks, as essentially a whole family of models. The one that's only four layers deep, five, six, seven, eight, nine. And so you can do what's called early exit, right? Where you stick the input in. And if the model is already very confident of the answer after evaluating only say four layers or eight layers, you can cut off the computation there and quit.

20:16Or you can do the computation all the way to layer 100, but you now have like 100 different votes, one from each layer. and you can combine those in some way to get a better assessment of your accuracy. And then now we have things like, oh, one, we don't really know what it's doing, but we have these systems that are built out of multiple calls to LLMs and even calls to different LLMs to do things like generating candidates, scoring candidates, ranking and trying to assess uncertainty. And now we're seeing, you know, programming frameworks coming up to support that. So it's a wild and crazy time right now.

21:04I don't know that we have much theory to guide us there, but we're gaining a lot of engineering experience. And certainly with the onset of deep learning, experimentation and engineering have far outstripped the theory. And so those of us on the academic side are really struggling to keep up and try to explain what's going on from a more mathematical standpoint. Yeah. And to that point, before people really jumped on these transformer models, I was hearing complaints among researchers at the big conferences that supervised learning had been kind of studied to death. and that, you know, there were these incremental advances, but nothing was really changing.

22:05And then, you know, the scaling of the transformer models kind of changed the field, and then everyone ran to that, and that's where everyone is working. Not everyone, but the bulk of the field is working now. Do you think there's more? And then I speak periodically to Rich Sutton, another winner of the award. One of my research buddies, yeah. Yeah. And, you know, reinforcement learning has not fallen by the wayside, but it's, you know, it's used as a component in some of these models or systems. But do you think that there are areas of machine learning where there's yet to be really major advances and that in the longer term, generative AI will be an important development, but there will be all of these other developments?

23:21mean people are now a lot of people working on neurosymbolic systems and you know is again i'm a journalist but you know five years ago uh the symbolic uh people were really out of favor so yeah where where do you think all of this fits in or or or have we really moved beyond supervised learning and self-supervised systems. Hi. You know, when you're building an AI assistant, one of the biggest challenges is making sure it actually understands and responds quickly, right? That's where Speechmatics comes in. They've got real-time speech-to-text that gives you over 90 % accuracy and less than a second of delay, crucial for making interactions feel seamless.

24:14And with recent updates, they're now making 25 % fewer mistakes than Microsoft, which means you're getting some of the best accuracy out there for things like customer service bots or voice-driven assistants. Plus, it works in over 50 languages, delivering results in just 700 milliseconds. If you're working on an AI project, it's definitely worth a look. Personally, I've been amazed by Speechmatics' voices, voices like Humphrey, which sound completely natural. If you're working on an AI project, it's definitely worth a look. Check it out at www.speechmatics.com slash realtime. That's www.speechmatics.com slash realtime.

25:12Well, that's a lot of questions. Let me see. So in the area of supervised learning, of course, actually, even before Transformers, deep learning really shook up the theory of machine learning. because we had all been studying the case where the sort of number of degrees of freedom in the models was less than the number of data points. So they were deliberately – we were very concerned about overfitting the training data. And so we would use regularization and other kinds of restrictions to prevent the model from overfitting the training data. uh and and and really ensembles was part of that story as well but but uh but but the what but there were there were already um signs of of of a conceptual uh uh i'm trying to think what the right crisis i guess in in uh you know the the sort of uh structured scientific revolution sense because there were some algorithms like uh boosting that where um even after they had fit the training data, their accuracy on independent test data continued to improve as they were run longer.

26:32And there wasn't really a good explanation for this. And, and, and now and then with the with the deep learning, we saw that even more that you could be training and your your your error would drop, drop, drop, drop, drop on your training data, until it was essentially zero. And yet, if you kept training performance on the test data continue to improve. And this was like super baffling because our theory did not really explain this. And I think we're still, I don't think we still have a complete story there. There's been a lot of theoretical work on trying to understand what's going on, and we do have a much deeper understanding.

27:15But in fact, I mean, one of the most intriguing papers recently is by Jugate, I can't remember her first name, who's a professor at Toronto, showing that for these transformer models, that in order to achieve their sort of optimum performance, they must memorize a certain fraction of the training data, which would be a kind of overfitting. And of course, this is also very relevant to these questions about copyright and so on. So, so there's a lot of so so as I say, the theory is still trying to catch up with the practice. Now, that's still mostly in the supervised setting that that's being analyzed.

27:59Um, so let's see. The, you know, obviously, to the, I would say the biggest lesson for artificial intelligence writ large of the LMS is that if you can do your training as self-supervised, then you can scale to internet scale. And I don't think we'll forget that lesson. So there are many, many problems, I think, with LMS and with statistical learning more generally. But anybody who proposes something new has got to show how it also will scale to internet scale. And that brings us to this question of neural symbolic architectures of various kinds. I have a talk I've been giving on what I think is wrong with large language models.

28:48And one of the things I think is the biggest problem is that all this factual knowledge about the world is stored in the weights of these deep networks. And whereas I think that that factual type knowledge should be stored symbolically, say, in a knowledge graph or a database or some kind of explicit knowledge representation, because then it would be so easy to edit it and change it. Whereas these deep networks are really just giant monoliths that are very hard to update. I mean, there are people obviously working on that, but I would say the current work on continual learning and deep learning has still really failed to give us any efficient way of doing additional training, aside from the kind of LoRa-style things, which are still quite expensive.

29:39Yeah. So I guess that would be an example of a neural symbolic type architecture. Another one I'm really excited about is symbolic regression. So I've recently been seeing a lot of talks by physicists and atmospheric scientists where they are using these symbolic regression algorithms. There's a Python package called PISR, I think, that does some fairly straightforward work to try to find an algebraic equation, or maybe it can have transcendental functions in it too, but to fit data. and they are finding that this is much more interpretable and that they can often, after they fit that symbolic formula to the data, that they can interpret physically what processes are being captured by different parts of the fitted equations.

30:36And so at ICHCAI this year, there were some very nice work from Purdue on really improving these symbolic regression models. So that's a completely different line of research, but I think it's having a big practical impact in science. Yeah, on that symbolic interpretation, what was the second word? Well, it's called symbolic regression. So, you know, you have your input and then you have a real value response variable. Well, but we fit some sort of a deep net to that. But these folks instead are fitting an equation, you know, so y equals f of x, where f of x is actually some, you know, you give it some vocabulary of things it's allowed to use, logs and exponentials and signs and, you know, multiplication, division, subtraction, and so on.

31:30And it searches that space of formulas. early work in this area was done back in the 1980s by uh uh pat langley working with herbert simon one of his simon's very last research projects uh but then it it uh sort of uh didn't go anywhere um but now uh it's been taken up by scientists and uh and they are finding that even i would say kind of mediocre tools are actually extremely useful to them and this has stimulated more machine learning researchers to come back and look at that problem again and see if, and so we're finding ways to scale up those algorithms to be quite efficient and effective even for fairly high dimensional problems.

32:15Yeah. In that symbolic regression, you're building an equation to fit the data. And then can you take that equation and have it look at other data that it hasn't seen and extract from that data? Well, you would use it to make predictions on new data. Right. So it's standard supervised learning, except that the model that's being fit is not, you know, kernels or deep learning or decision trees or whatever, but it's algebra, right? And so it gains immediately by interpretability if you give it meaningful input features. Now, if your inputs are, I don't know, lower level things coming from cameras or sensors or something, now you have this problem that the physical quantities of interest, say masses and accelerations or whatever, you're not directly observing and they're not directly given.

33:23So they would be latent variables. And this is going to take us back to Jan saying his JIPA models or something. Now we need to discover those latent variables in order to then express the function in some clean algebraic form over the latent variables. And I don't know of anybody who's working in that area. I'm sure some people are. But Jan certainly isn't. I mean, most of our trying to learn dynamical models of these systems, we're still using neural network models because we want the end-to-end differentiability. Whereas the symbolic regression is a discrete search, right, over a discrete space of, you know, algebraic expressions, so tree-structured objects.

34:08And, yeah, and it has all the problems of combinatorial search. Yeah. You said that you're working on the application side. You were saying that the theory has yet to catch up with the practice, but you are working on the practical side, on the implementation side. side. Can you talk about that, what sorts of systems you're using? I know you did some work on wildfire detection, I think. Right. Yeah. So we'll take us back to reinforcement learning, since I didn't answer your question there. Right. So I had a project on wildfire management, and we formulated that primarily as a reinforcement learning problem.

35:02So the question we were asking there was, where should fuel treatments be applied? So a problem in the U.S. and other countries, too, is that throughout most of the 20th century, we were extinguishing fires as quickly as possible, whereas the sort of natural fire processes of the ecosystem would have had fairly frequent low-intensity fires. At least in the Pacific Northwest, we would have had a lot of fires that would have come through and burned out the understory of the forest while leaving the big trees intact. And the Native Americans actually managed the landscape deliberately to try to achieve that.

35:46But because we extinguished fires very quickly, we had this big buildup of fuels in the understory of forests. And the result is that when we did have an ignition, it exploded with very intense fires. So the question was, well, could we send people in to remove that accumulated fuel in the understory? And that's very expensive. So where's the best places to do that in order to try to control the sizes of the fires that might arise, say, from lightning strikes. And so that was a problem that we studied. Another one we studied was liability policies. So let's say you are one landowner and you own some patches of forest and I'm another landowner and I own other patches of forest.

Read the full transcript

36:38If a fire starts in my land, say, because I have not been doing a good job of fuel reduction, and it burns into your forest where you have been doing a good job of fuel reduction, shouldn't I be liable for your losses in your timber plans? And so we studied different possible liability regulations, and we were able to use multi-agent reinforcement learning to do that. So each landowner was choosing whether to apply fuel treatments or not to their land, and whether to harvest their trees for timber. So this is a very Pacific Northwest kind of model, although it would work in the Southeast for us too.

37:28And the question was, should we have this kind of liability or should we just say everybody has to pay for their own, that fire is an sort of act of God and no one should be held liable. And what we found was that actually a policy that says, as long as you have been maintaining your fuel risk level at a reasonable level, then you should not be held liable for anything that starts on your property. And that other things that were more onerous kinds of regulations actually led to worse outcomes. because basically fire is very rare, despite all we see in the news. And so the chances that your land will actually burn are extremely low.

38:18And so there's a tremendous incentive to just hitchhike and rely on your neighbors and what they're doing to reduce their fuel loads to protect you. And as taxpayers, we sort of versus timber companies, are the timber companies relying on the federal government to do the work? Or actually, it's actually more vice versa. The timber companies are doing a better job of managing their lands than the feds are. And so anyway, so that's the kinds of things we could study with reinforcement learning. Yeah. It occurred to me, you were saying that your approach has always been to look at a problem and then figure out what flavor of AI could be applied to address that problem.

39:10Is that something that these new reasoning agents, I mean, you mentioned O1, could do,

39:24match an AI system or to a particular problem? Yeah, so could they help us formulate problems? So, right, if you think about a very standard thing, a problem you might encounter in industry, say, in, I don't know, chemistry or a lot of industrial processes, is that you'd like to optimize your processes. And so you might want to formulate, say, linear programs or quadratic programs and use big optimization packages to do this. And you might want to do that robustly, which is another layer of complexity. Formulating those problems correctly is very difficult. And maybe just as we can see that we can use LLMs to write SQL queries for us or to write code for us, maybe we could get them to also formulate these optimization problems.

40:18And that would be another example of a sort of neurosymbolic approach where you're using the broad world knowledge of the foundation model to help you do problem formulation, which you then hand off to a very specialized reasoning engine that is perfect for solving linear programs or something like this or planning problems or scheduling problems, things like that. So, yeah, I think there's a possibility there. I mean, I think these large language models are our first kind of thing that has very broad knowledge of the world. and we've never been able to build systems like that before. So for so much of my lifetime, our computer systems and our AI systems have been very narrow.

41:06And so these systems could do very deep reasoning in their area, but they would often lack common sense. And I think one of the, so that they might draw a conclusion that just didn't make any sense in practice. And one thing I'm really excited about is actually closing the loop. So start with this large language model. Maybe it can, say, formulate the reasoning problem abstractly, hand it off to a theorem prover or an optimizer, but then take the solution and come back and have the LLM evaluate it and say, well, I know this solution followed logically, you know, deductively, mathematically from the problem I formulated, but it doesn't make sense.

41:49It's violating a bunch of common sense things. Let me reformulate the problem and try again and keep doing that because I feel like, you know, the power of abstract reasoning is that you can turn the crank and it's perfectly sound reasoning. But the weakness of abstract reasoning is it throws away the context of the problem. And so when we get the answer, you know, we always teach our students now check your answer. Go back to the original problem and say, does this really make sense or not? And this is where traditionally AI systems have failed and where I think we have an opportunity now. So the other thing that I think that one and these other models are really showing is that if you look at the work that, for instance, Newell and Simon were doing back in the 50s, 60s and 70s, they talk a lot about the problem of how do people, what they call evoke the right knowledge to solve the problem they're currently confronting.

42:46So you put a person into a context and somehow they're able to retrieve the knowledge that's relevant to that context out of the billions of things they might know. Well, at least large language models know billions of things. And what we're finding, and with our narrow AI systems, this was never a problem. They just didn't know enough. We could just retrieve everything because there just wasn't much there. But now with the LLMs, the trouble is there are these billions of things that they know. and uh and this is why we have so much work on prompt engineering because we're trying to evoke the right memories the right pieces of knowledge from all of those scientific articles they've read and all the newspapers and what's the piece of relevant the knowledge that's relevant to this situation and so what i um what i think is is really fascinating is uh we're gaining a lot of experience with that right now and people even uh you know building tools to optimize the the the retrieval um so uh i think a lot about the you know this chain of thought and so on these are ways of trying to uh restate or or or transform the initial problem statement into into a statement that will retrieve the right things when we when these systems go to their memories so i think it'll be very i'm not sure that these are the best ways of representing all that.

44:08But, you know, traditionally, when we had too much knowledge retrieved, then you have another search problem that can, another combinatorial explosion, that you retrieve too many things, and you have to try them all out. And we're also hoping, I think, that when we do retrieve a bunch of stuff, that then the LLM can kind of score how relevant each of these retrieved chunks of knowledge is to the problem at hand and avoid that combinatorial explosion. So I think we're kind of groping and feeling our way in this area right now. But if we can solve that, then we can overcome the, always in old, you know, good old fashioned AI, we always would run into these combinatorial explosions, because the system didn't know enough to be selective on what reasoning paths it followed.

44:57So I think this is also extremely interesting. And do you think that these models, the transformer models, as they become more powerful with reasoning, which it appears that O1 is an ensemble of models approach, I'm not entirely sure, but that as the reasoning gets stronger, that these models could play a role in scientific research, in formulating hypotheses and then working through problems. Yes, I hesitate a little bit because I feel like LLMs are also vulnerable to all the problems of statistical learning. And so one of my favorite recent papers came out of Tom Griffith's lab at Princeton.

46:04The first author is Tom McCoy, which is called Embers of Autoregression. I can't remember. It has a subtitle. But the basic idea is that they show that LLMs are very dependent on the training distribution, the distributions of the training data. So if there are problems that they have seen many, many examples of, they're much more accurate on those than on problems that they've seen much less, have much less experience in, which is what we would expect from statistical learning. Similarly, they tend to, if the right answer to a question is actually a string of words that has low probability, they will instead kind of autocorrect that and output a higher probability string that is the wrong answer.

46:57And I think this is one of the sources of their hallucination tendencies. And so I think we have to solve this problem. And it goes right to the heart of statistical learning, which is that, you know, supervised learning, unsupervised learning, all these methods are extremely dependent on the distribution of their training data and very brittle to shifts in that training data. And I think that we need to try to find ways of training these systems so that they are much more robust to distribution shifts. Ideally, that we would have distribution independent learning where you would learn facts and those facts would be true independent of the training distribution.

47:45So I'm still totally unclear about how that could be accomplished. But my intuition is that at least in the parts of the input space where we have lots of training data, if we could also prove that our models are smooth in those regions, so they're not doing something crazy in between the training data points that they had, I think we should be able to get a guarantee that they will perform well regardless of the probability distribution as long as the distribution is sufficiently spread out. It doesn't concentrate. You know, you can never protect yourself against like if all of your queries were on this one thing you always get wrong.

48:25You can't protect against that. But the idea would be that if there's some amount of entropy in the in the query distribution, then we would we would do things well. And I think that's also what we need for safety critical applications of machine learning. So I've been part of a National Academy of Sciences study on machine learning and safety critical applications. It's been looking mostly at automated cars, automated aircraft and medical applications, which we're and and looking at both traditional safety engineering and then what might need to change in order to bring machine learning into the picture.

49:03And I think the big problem is that there is no statistical learning algorithm that's going to give you five nines of reliability. Right. It's very rare that we even get something that's 99 percent correct in machine learning. It's usually in the high 90s. And we're very happy about that. So the the the the solution to that conundrum is we need our systems to give extremely reliable uncertainty quantification so that they can say, okay, I know, I know how uncertain I am. And so I can guarantee that if I give an answer, it will be right with say, five nines. And if but if the bit if I'm uncertain, then I will, I will flag that and we'll fall back on some other decision making process.

49:53Yeah. So this problem of uncertainty quantification, which I feel is a huge issue for large language models. And there are maybe a dozen or so groups around the world trying to get good probability estimates out of these LLMs and trying to see how much that can remove the hallucination problem. Yeah. Another thing people have talked about is curating the training data in some way so that you don't have a lot of chaff or a lot of noise in the training data. Are there strategies that you've seen that show promise there? Well, I guess traditionally the strategy was kind of an outlier detection, anomaly detection approach.

50:43So you would apply an unsupervised learning method to try to understand, map what the training distribution looks like. and then flag things. If you think there are errors in the input features, like measurement errors, for instance, then you could detect those. So you might say, well, you know, someone who has a salary of$100 billion, that's very unlikely, or has, I don't know, an age that's a negative number. Things like this you could detect. Detecting bad labels in the supervised case is tricky because you basically have to impose some prior belief on how accurate the classifier should be.

51:26And then you can say, well, the points it's getting wrong, maybe they're mislabeled and we should ignore them. As opposed to maybe the points it's getting wrong are just the hard examples and you should learn from them. And that's a fundamental conflict that is ultimately not resolvable. It depends on really understanding where your labels are coming from. and what their errors are. So if you can assess that and you have an idea of what your error rate is on your labels, then you can say, well, yeah, then you can use that to guess which data points maybe need to be relabeled or downweighted or even deleted from the training data.

52:07Yeah. You know, I've spoken to people about using these large models as truth engines to be able to, you know, if they're trained properly, to be able to present them with a question on which there's a debate and the model would give you a probability score on one viewpoint or the other. Is that something that you thought at all about? Does that enter into your research? Well, I haven't done research specifically on that, but I think we know that if we can have a reliable oracle, like if we think about the game of Go or chess, AlphaGo, it had a perfectly reliable oracle for whether it won the game or not.

53:06But that's an event that happens far in the future. And by using reinforcement learning and search, we can use that oracle to label our data, to label the games we won and the games we lost. Now, the trouble is LLMs are not very reliable oracles for a lot of things. And most of what I've seen of the attempts to use LLMs to criticize their own outputs have not been very successful. there's a there's a technique for uncertainty quantification called p true which is where you basically you ask the lm to produce an answer and then you ask it again uh you know um do you think this answer is correct or not and and and have it just give you a yes no score and of course it gives you it can give you a log probability along with that and we use that as an assessment that's a really terrible method.

53:57It doesn't score very well on the experiment. So I'm skeptical.

54:07But maybe there are ways that we could train good oracles for specific problems. And then the systems could, right now, say, oh, one is doing some kind of a search, and it is scoring somehow its own results. And when it finds a result that has very high confidence is correct, it could then train on that, right? And that would be a direction toward a form of self-improvement. So, but I think in general, if we look at the history of science and human beings, we generally have to go out and do an experiment in the real world to check a hypothesis. We can't use our own prior beliefs to say, well, it seems right to me, so I'm going to believe it.

54:48that that's that leads you into conspiracy theories and confirmation bias so so i think um the the uh the i'm also very interested in the sort of automating science by by hooking up uh lms and other machine learning uh technology to the so-called self-driving labs where you can uh robotically uh design and execute experiments get the results back and then learn from that so So that's a really exciting direction. Yeah, particularly in drug discovery. I've spoken to a number of people that have automated. And materials and mathematics. And, of course, O1 is particularly strong in mathematics. And that may be because, again, that's an area where you can check.

55:39Once you have an answer, you can check it, say, with a theorem prover or something. This has been fascinating. I really enjoyed the conversation. Is there anything you want to end with that I didn't touch on? I guess just to the extent that there are students listening to this, I hope you get the idea that there are tons of open problems. AI has not solved, whatever that would mean. And I think that it's a very exciting time and will be for a long time. I don't think that some sort of super intelligence is around the corner that's going to solve all these problems, but that we have many, many exciting things to work on.

56:21And I encourage people to join the field and join the fun.

From the publisher

This episode is sponsored by Speechmatics. Check it out at www.speechmatics.com/realtime

 

Today, we're joined by Dr. Thomas G. Dietterich, a pioneer in machine learning who recently was honored with the Award for Research Excellence from the International Joint Conference on Artificial Intelligence, one of the top awards for AI researchers.

 

Dietterich traces the field's progression from early rule-based systems to modern machine learning paradigms and delves into his work on novel category detection and open set problems. He also discusses the evolution of ensemble methods in the context of large language models (LLMs), highlighting the shift from combining many cheap models to more selective approaches with expensive models.

 

He advocates for a foundation model approach to capture the variability of the world.

 

Join us for a deep dive into the future of AI, where Thomas explains why the development of novel materials and drugs may have the most transformative impact on our economy. Plus, hear about his latest work on multi-instance learning, weak supervision, and the role of reinforcement learning in real-world applications like wildfire management.

 

Don’t forget to like, subscribe, and hit the notification bell to stay updated on the latest trends and insights in AI and machine learning!

 

 

Stay Updated:

Craig Smith Twitter: https://twitter.com/craigss

Eye on A.I. Twitter: https://twitter.com/EyeOn_AI

 

 

(00:00) Introduction to Thomas Dietterich's Machine Learning Journey

(02:34) The Early Days of Machine Learning and AI Systems

(04:29) Tackling the Multiple Instance Problem in Drug Design

(05:41) AI in Sustainability

(07:17) The Challenge of Novelty Detection in AI Systems

(08:00) Addressing the Open Set Problem in Cybersecurity and Computer Vision

(09:11) The Evolution of Deep Learning in Computer Vision

(11:21) How Deep Learning Handles Novel Representations

(12:01) Foundation Models and Self-Supervised Learning

(14:11) Vision Transformers vs. Convolutional Neural Networks

(16:05) The Role of Multi-Instance Learning in Weakly Labeled Data

(18:36) Ensemble Learning and Deep Networks in Machine Learning

(20:33) The Future of AI: Large Language Models and Their Applications

(23:51) Symbolic Regression and AI’s Role in Scientific Discovery

(34:44) AI in Wildfire Management: Using Reinforcement Learning

(39:32) AI-Driven Problem Formulation and Optimization in Industry

(41:30) The Future of AI Reasoning Systems and Problem Solving

(45:03) The Limits of Large Language Models in Scientific Research

(50:12) Closing Thoughts: Open Challenges and Opportunities in AI

 

More from Eye On A.I.

All 266 episodes
#212 Thomas Dietterich: The Future of Machine Learning, Deep Learning and Computer VisionEye On A.I. · 56 min
Listen in VO