In short
Eye On A.I. Podcast Episode Summary
Episode Details
- Title: #140 Isabelle Guyon: The Future of AI and Support Vector Machines
- Host: Craig S. Smith
- Guest: Isabelle Guyon, Professor at the University of Paris-Saclay
- Sponsor: MindStudio by UAI
- Release Date: [Not Specified]
Episode Overview In this episode, Craig S. Smith converses with Isabelle Guyon, a renowned pioneer in machine learning and pattern recognition, focusing on Support Vector Machines (SVMs) and their implications in AI today. The discussion encompasses various aspects of SVMs, including their application in biomedical fields, advancements in machine learning techniques, and future directions for AI.
Key Topics Discussed
- Introduction to Support Vector Machines (SVMs)
- SVMs are a type of supervised learning algorithm primarily used for classification and regression tasks.
- They work by finding the hyperplane that best separates classes in the data while maximizing the margin between different classes.
- Applications of SVMs include:
- Image classification
- Text classification
- Bioinformatics
- Machine Learning Techniques
- Discussion of various classification methods, particularly in the context of biomedical applications.
- The importance of adapting classifiers based on available data (labeled and unlabeled).
- The role of maximum margin classifiers.
- SVMs vs. Deep Learning
- SVMs are often seen as simpler alternatives to deep learning models, which require large datasets to function effectively.
- Guyon emphasizes that SVMs can still be powerful, especially with small datasets.
- Deep learning excels in representation learning, whereas SVMs require hand-crafted features.
- Future of Chatbots and Language Models
- Insights into the evolution of chatbots and language models, including data preprocessing and bias adjustment methods.
- Discussion on the combination of SVMs with deep learning techniques to overcome contemporary machine learning challenges.
- Addressing Data Quality Issues
- Discussion of the challenges of data quality in large datasets.
- Emphasis on the need for better-curated data to improve AI outputs and mitigate biases.
- Exploration of methods for improving data quality and addressing biases in AI models.
- Ongoing Research and Future Directions
- Isabelle Guyon mentions her ongoing research interests, including the integration of SVMs with modern deep learning frameworks.
- The importance of addressing the data problem in AI and the collaborative effort required to improve data quality.
Key Takeaways
- SVMs remain relevant: Despite the rise of deep learning, SVMs continue to be a fundamental tool in machine learning, particularly for smaller datasets.
- Data quality is crucial: High-quality data is essential for the effective training of AI models, with a collective effort needed to ensure data accuracy and reduce biases.
- AI is evolving: The conversation highlights the potential future of AI, with ongoing advancements in both machine learning techniques and applications across various fields.
Conclusion The episode concludes with a reminder that while the singularity may not be imminent, advancements in AI are poised to significantly transform everyday life, necessitating close attention to developments in the field.
Episode Structure
- 00:00 - Preview
- 00:52 - Introduction and MindStudio by YouAi
- 04:55 - Machine Learning Techniques
- 15:15 - Classification Methods in Biomedical Applications
- 22:40 - Support Vector Machines and Kernel Methods
- 36:45 - Future of Chatbots and Language Models
- 41:08 - Unsupervised Learning in Deep Learning Importance
- 45:53 - Outro and MindStudio by YouAi
Links
- [Isabelle Guyon's LinkedIn](https://www.linkedin.com/in/isabelle-guyon-aa371170)
- [Craig Smith's Twitter](https://twitter.com/craigss)
- [Eye on A.I. Twitter](https://twitter.com/EyeOn_AI)
- [MindStudio by UAI](https://bit.ly/MindStudioEyeonAI)
This episode serves as an insightful exploration into the fundamental aspects of machine learning, the significance of data quality, and the promising future of AI technologies.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00The vectors are actually the coefficients or the features, the components that you compute that get into the system. So the vector is the list or the collection of all the information that you've gathered that include in my example of medical classification. It could include the age, the amounts of certain proteins, the family history, anything that goes into your medical files. The amazing thing is that it looks like from this just very simple training strategy, predict what comes next from the past part of the sentence, you can get these chatbots like ChatGPT that give very sophisticated answers.
0:41But this should not fool us because really we've trained parrots. Hi, I'm Craig Smith and this is Eye on AI. It's an increasingly confusing time in machine learning, with generative AI consuming all the oxygen, and a plethora of startups offering new services based on large language models. While I'll continue to talk to select founders, I'm going to try to get back to fundamental research and pay attention to the many avenues that have sort of fallen by the wayside, but are still pushing AI forward. This week, I speak to Isabelle Guillaume, a French researcher in machine learning and pattern recognition.
1:27She's known for her work on support vector machines, otherwise known as SVMs, a type of supervised learning algorithm used for classification and regression analysis. SVMs are based on the idea of finding the boundary that separates data into different classes with the largest possible margin. Support vector machines are still widely used today in a variety of applications, including image classification, text classification, and bioinformatics. SVMs are particularly useful when dealing with high-dimensional data. While deep learning has dominated AI in recent years, SVMs remain a powerful tool in the machine learning toolkit.
2:18I hope you find the conversation as informative as I did. Before we begin, I want to give a shout out to our sponsor, UAI. They're a fascinating company that does a number of things, but I'm most interested in their Mind Studio, which is a platform for building AIs on top of large language models that you can deploy for free or for profit. Think back to the days when smartphone apps were first getting started, and if you were able to have been among the first couple of thousand people to build a smartphone app, you would have learned a lot and may have earned a lot of money. So give MindStudio a try.
3:05Visit uai.ai to build your own AI today. I'm a professor at Université Paris, and I'm in the detachment at Google as a director of research in San Francisco. I'm also president of a nonprofit organization called Chalearn, which is dedicated to organizing challenges in machine learning. And I've had a long time interest in benchmarking machine learning algorithms. As early as, you know, during my thesis, I was already collecting data samples of handwritten digits. And I collected one of the very first data sets called the little 1200 that had 1200 digits written by people at Bell Labs while I was doing an internship there.
3:57Yeah. When you were at Bell Labs, I'm guessing that's where you worked with Vladimir Vapnik and others on support vector machines. Is that right? Or did your work on that start before that internship? Yeah, that's right. Actually, I started working towards the end of my PhD on kernel methods. I was guided by that by one of my mentors, John Denker, with whom I was doing this internship at Bell Labs before the end of my PhD. And also, I was working on methods that we now refer to as the kernel trick that one of my co-advisors, Leon Personas, had pointed out to me. There was a paper from Tommy Poggio in 1974 that had induced the polynomial kernel.
4:41And I was using that in my thesis to make Hopfield networks behave as nonlinear, learning machines that are nonlinear in their parameters, to be more precise. So my PhD was originally on architectures of neural networks. I started off with studying Hopfield networks that were inspired by magnetic systems. The reason for that is that I was studying in a physics school, the School of Physics and Chemistry of Paris. That's where Marie Curie and Pierre Curie discovered radium. So I was always very inspired by Marie Curie. So I'm very fortunate that I studied in that school. And yes, initially I was a physicist and worked with Hopfield Networks with my advisors, Gérard Dreyfus and Léon Personas.
5:25And so, yes, I was familiar with the kernel trick and I was already using it during my PhD. I was also familiar with large margin methods. That is the other component of the support vector machine algorithm, as there were people from Ecole Normale right across the street from my school who are working on methods called from minimum overlap. So these are methods that ensure stability and basically also see a large margin. And so, yeah, when I started working at Bell Labs, the situation was very different than from when I did my internship. Yann Le Kuhn and Bernard Bozer, two young men, had just joined.
6:03And there was a lot of excitement around backpropagation networks that at the time were called multilayer perceptron. And I was strongly advised by my department head, Larry Jackal, to work on backprop nets. And so I started working on them. And I did some work with Yann Le Kuhn on applying convolutional neural networks to pen computers. So how to replace the keyboard by a pen. At the time, we thought this was very promising and handwriting would be a good user interface. We were proven wrong. And only basically two years after I joined Bell Labs, Vladimir Vapnik came. At the very beginning, he shared my office because there was no office space for him yet.
6:45So I was very lucky to have many long conversations with him. And he talked about this algorithm for optimal margin that he had an appendix of his book and that nobody had ever implemented. I thought this was interesting. I mentioned to him work that I had done around that. And I always thought, yeah, I should do that. I should implement that algorithm. But I was busy with many other things. So I never did it. And then all of a sudden, my husband got a position at UC Berkeley, and he didn't have time to start a new project before we left. And he was seeking a short thing he could do. And he asked Vladimir Vapnik, what is some small project I could do?
7:24Well, We're waiting to leave to California. I can't start a new hardware project. He's a hardware guy. Something small and very inventive. And he said, oh, well, implement my optimal margin classifier. And so he did. And I thought, oh, my God, I'm going to be scooped. I really wanted to work on that. And then my husband, Bernhard Bozer, the third co-inventor of a support vector machine, came to me and said, oh, yeah, it works nicely now. And Vladimir said I should move on to a nonlinear version of it. and we should make products of the inputs to make it nonlinear. And I said, no, you shouldn't do that.
7:59You should use the kernel trick. It's much more powerful. That way you have almost for free the possibility of plunging your problem in a very high dimensional space, but just essentially taking every dot product and replacing it by another similarity measure. And for example, you need to take every dot product and raise it to a certain power. And say, wow, this is simple. So he didn't really know how to do that with the algorithm that was in the appendix of that book. So I just rewrote the algorithm and he implemented it. And sure enough, it worked very nicely. Yeah. So to make a long story short, this is how the invention of SDM happened.
8:38That's interesting. I didn't realize that you were already working on neural nets when you did that work, or on backpropagation, I should say, because I'm a novice. So correct me if I'm wrong, but support vector machines are not deep learning. They're a much simpler way of classifying data and work very well on small data sets, whereas deep learning, you need very, very large data sets. Was there a connection between your work on the nonlinear support vector machine and backpropagation, or were they very separate? No, they were not separate at all. They were not separate. For me, it was, I should tell you why I didn't do the work on support vector machine earlier than I did, and I was dragging my feet to implement that algorithm.
9:23The reason is that I thought this wouldn't make that much of a difference because the way I was training neural networks was with a strategy very similar to support vector machines. I had realized early on that in order to effectively train, you need to be focused on those examples the hardest to learn, which ended up being called the support vectors and the examples that are closest to the decision boundary. So this is already underlying also the other works going on, finding large margin algorithms that are stable. The idea is very simple. The idea is that if you have a variety of possibilities to separate examples, and that happens often when you have very few examples and you are in a large dimensional space in particular.
10:10So if you have very few examples, there are many ways in which you could create a decision boundary. There's a no man's land between the decision boundary and the examples. And it's natural to say that you want to have the largest possible distance between the decision boundary and those examples, which is the idea of the large margin classifiers, the large stability classifiers, etc. How do you achieve that? Well, you actually achieve that by showing more often the examples that are hardest to learn. And this is what was underlying this algorithm that had proposed many people at ONS close to my institute when I was doing my PhD.
10:47What their algorithm was doing is that it was a modification of the very classical perceptron algorithm. In the very classical perceptron algorithm, you make a change in the weights when an example that is shown is misclassified. And you don't make a change if the example is well classified. And the very minor change that they made in this algorithm is that they were going through all the examples of the training set, and they made a change only for the example that was worst classified. And that simple change, you can show that asymptotically the algorithm converges to the optimal margin classifier.
11:24So it reaches the same solution as the SVM. I had many arguments with Vladimir Vapnik, what was best. He told me, well, this algorithm only asymptotically converges. It is not guaranteed in a finite number of steps to get to the solution. But on the other hand, it's amenable to be combined with the very popular stochastic gradient descent methods that neural networks are trained with. Because the only thing that you need to do is you train your neural network and you show more often the examples that are hard to learn. So in the extreme case, you would at each epoch only show the example that is hardest to learn.
12:03But this is kind of inefficient because you have to do a forward propagation for every example, figure out which one is the worst one, and then only update the weights for that one example that is hardest to learn. So it's a waste. So you can do something in between, identify the examples that are hardest to learn, and only do the updates on those. And this is what I had implemented, and it worked very nicely. So I didn't really see a need to implement this algorithm. I was getting good results. Now, what happened is that when Bernard Bosom has been implemented Vapnik's algorithm, he got good results.
12:41And then I suggested to combine it with the kernel trick, which is the trick that I was using also extensively with other algorithms than the maximum margin algorithm already in my thesis. And it's not a new idea. It comes from the 1960s. There's a paper about this idea of the kernel trick. Perhaps most interestingly, this paper was written in the same institute as Vapnik, and he never realized that the two algorithms could be combined. It was called the potential function algorithm, and his algorithm was the optimal margin algorithm. So it took 30 years before somebody had the idea of putting them together.
13:20My conjecture is that because they were fighting for funding, so they wanted to differentiate their methods as much as possible, and so they had never considered putting them together. So at any rate, yeah, I finally pushed Bernhardt to use the kernel trick with the optimal margin classifier. And then, yeah, this gives this new algorithm that basically sparked a big fire. Yeah. It just occurs to me for listeners that don't know what support vector machines are that we should probably define them. Okay. Let me try to explain that. Without pen and a paper, it's a little hard, but you'll stop me anytime with questions.
13:55So, well, the auditors probably know about classification problems that have been popularized with AI and neural networks. We've seen a lot of applications of face recognition, of object classification, of text classification, for example, classifying between spam and ham or classifying between bees and wasps if you want to determine which insect is a good pollinating insect or a dangerous insect. Anyways, this has taken off recently using neural networks, for the most part because neural networks are very good at learning representations, whereas algorithms like support vector machines rely on the fact that you've already come up essentially with a good representation.
14:44And this is not so obvious in the case of computer vision problems because you basically start with an image that has been transformed into pixels, that, you know, little elements of image. And so the representation is bizarre, right? What the algorithm sees just little dots that don't seem to make so much sense. They don't have an idea of what the object is as a whole from these little dots. But for a lot of applications, the representation that you start with makes more sense. For example, if you have a biomedical application and you're trying to classify whether patients have cancer or not, you may have a whole bunch of diagnoses, like whether you have a sort of protein in your blood that has an elevated.
15:30That is the case, for example, if you diagnose prostate cancer, the prostate cancer antigen is elevated often in prostate cancer patients. This is not going by itself to give you a certain diagnosis, but if you have many features together, like the age of the patient, family history of cancer, the weight of the patient, whether the patient has been exposed to certain chemicals, etc. These are all called biomarkers. If you use all of them, you can try to create a classification method that will tell you with a certain confidence whether the person is at risk or not of a certain disease. And this is what these methods try to compute.
16:10They try to compute a number that rates the degree with which you believe that a certain person or a certain object belongs to a certain category. So you can do that with many different methods, but in machine learning, what you do is that you get examples. So you get examples of all these different features, and you have examples of normal patients and these patients. And then you try to determine whether you can combine these different coefficients or these different features. We call them pictures, coefficients, biomarkers, etc. All these quantities that you've measured about the subjects of interest.
16:53You try to combine them and then come up with a single number. So how do you combine them? One of the easiest ways is that you turn them into numbers and you do a voting, which is a weighted sum. So you multiply each coefficient by a weight and then you average. So it's like voting, where each voter would have a different number of votes. Some coefficients will be seen as more important than others. But these weights, these weights are subject to training. So they are adjusted during a session in which you show many examples, and then you try to make it that according to the combination that you are computing of these input coefficients, you can then separate well the training examples.
17:36So you vote among these many features. you get this coefficient that is the average vote. And you see if the average vote is higher than a certain threshold, I will say that this is a disease patient. If it's lower than the certain threshold, this will be a patient that is not at risk. So what does SVM do more than any such method? Well, what it does is that if you come up with one such combination, what you want is to have the maximum distance between the disease patients, say, for example, in this example, and the non-at-race patients. These are basically according to this rating that you've made, which is the voting that you've created.
18:15According to this rating, if you would just represent all the patients on the line and look at what rating they get, and you put on one side all the patients that have high rating and on the other side the patients that have low rating, you will have some that are borderline, right? That have rating that is neither very high nor very low. And you would like to have the maximum gap between the ratings of the patients that are at high risk and the patients that are at low risk. And if your rating is such that you can make a really big gap, then you can be reasonably confident that you've made a good discrimination between these two types of patients.
18:52If you don't succeed in separating them well, for example, if there is an overlap between the distribution of the patients at risk and the distribution of the patients that are not at risk, then you can't really separate them well and then you kind of fail, right? So the idea between a maximum margin classifier is that if with the training examples you can maximize this gap, then maybe when you get new patients that you haven't used for training, they will not be classified incorrectly. So you have basically this big gap. In the middle of the gap, you can say, okay, this will be my decision threshold.
19:27I put it in the middle of the gap. Now, what you don't want is that patients that are of the wrong category climb over the decision threshold and end up on the wrong side. So if you have the maximum possible gap, then you can minimize your chance that the new patients will cross the gap and end up on the wrong side. And this is primarily for labeled data. Is that right? That is correct. So originally, the original algorithm was created for classification problems with label data. And moreover, the labels are categorical, that is, they correspond to different classes or categories. Later, it was generalized by Vepnik and other collaborators to the regression case in which what you want to predict is not categories but continuous values.
20:10For example, if you wanted to predict the age of somebody based on pictures, then this would be a regression problem, not a classification problem. And then it was also extended to unlabeled problems, so-called unsupervised problems, for example to the problem of determining the support of a density by Bernard Shalkoff and also to some clustering methods that are also for unlabeled data. But the same idea always applied, and the idea is you create a no-man's land, a big gap, so that you get some confidence that when you get new examples, then they will not kind of cross the border and end up making wrong decisions.
20:50And at this point, because deep learning, backpropagation, supervised learning, I should say, in deep learning has been so successful in classification, if you're building a model, what point do you decide, oh, this should, I should use deep learning or I'll just use support vector? Yeah, I think the method, the mistake people make is that, is to oppose the methods because they're completely complementary. They address essentially different problems. The deep learning methods address the power of learning representations. And once representations are learned, which can take a lot of compute resources, and we see now that it's very common that big companies train for many days with very large forms of computers, a model that is then released publicly.
21:37And we are very grateful that such companies are generous and share that with the public. Then these models can be stripped off their last layer that's basically making decisions and used as a preprocessor to compute new features. So that's what's missing essentially with super vector machines, is that they are not very good at coming up with good features, even though they can implicitly function in a very large feature space. I didn't explain the kernel trick, but I can explain that later. But essentially, they still rely on original features that are man-made, that are the input. And so they can benefit greatly from using as a preprocessor the features computed by deep learning methods.
22:21And then if you do that, then you can harvest the benefits of these methods because they can learn with very few examples. And in fact, there are many methods that are so-called future learning methods that are essentially based on kernel methods. When I talk about kernels, it's a broader family to which support vector machines belong. And these are methods that you can call case-based or example-based methods that rely on memorizing examples, essentially. The question is, what are the vectors in a super vector machine? The vectors are actually the coefficients or the features, the components that you compute that get into the system.
23:01So the vector is the list or the collection of all the information that you've gathered that include in my example of medical classification. It could include the age, the amounts of certain proteins, the family history, anything that goes into your medical file. So this would be a vector. And why not take a stab at the kernel trick, since I don't understand it? Okay, so now I've only talked so far about the vanilla, essentially, support vector machine, the linear version that was already invented in the 60s by Vladimir Vapnik and his collaborators, that consists in just making a vote among the inputs or the features that I was mentioning.
23:43So to take this a step further and go beyond linear classification, we can manipulate these features. And this was the first idea of ethnic. Let's make products or sums or let's make functions of these features and put that at the input of the classifier. In this way, we're going to massage them first in many different ways. And then hopefully this will create a richer vocabulary, in essence, that will make it possible to get more complex decision boundaries. That idea is not new. Already in the 60s, Frank Rosenblatt and his collaborators when they invented the perceptron, they in fact already had a first layer that was made of random functions.
24:25So you take the inputs and you combine them randomly. And so you expand your feature vector to a very large feature vector. And it can be shown that if you do that, then virtually every possible separation of a finite number of examples will be able to be made, provided the feature vector is long enough. So it's a question of, the big question that people had in the 60s is whether the examples are linearly separable, which means that if you have some training examples, will you be able to compute a weighted sum such that all the training examples of the first class are on one side and all the training examples of the other class are on the other side?
25:04And you can just set a threshold on this combination and you can classify 100 % correct all the training examples. So everybody was set on that problem. And later people discovered that it's not really necessary to have all the examples of the training set perfectly separated to get good generalization on test examples. But at any rate, this was really their focus. And so people had come up with that idea of let's make random functions. And then we will have a much larger space of features. And then if that much larger space of feature is going to be much easier to be able to find a combination that will separate perfectly the training examples.
25:42So a linear combination. So starting from this idea, then some people have observed that, in fact, you could use a different type of method that is an example-based method, similar to what people call the nearest neighbor classifier, and transform it into one of these linear classifier methods. So how does it work? Well, first of all, let's understand the nearest neighbor classifier. It's a very simple method. You just calculate the distance between your new unknown example with all the training example, and you classify it according to the example that is closest to it. So if you have a new patient, compare that patient to all the patients you have in your database and find the medical record that is most similar to that patient's medical record, and then classify that patient as disease or not disease according to this training example that you have.
26:34Okay. And so people invented many variants of this nearest neighbor method. And those methods are as a whole known as kernel methods. And what is a kernel? Well, a kernel is very simple. It's nothing but a similarity measure. I was talking about distance, but whenever you have a distance, you can also define a similarity measure. It's the opposite, right? The distance says how far away two examples are, a similarity measure, how resemblant they are. Some similarity measures are very simple to compute. For example, if you just do correlation. Correlation is a similarity measure. Everybody knows percent correlation coefficient.
27:10So if you take two records and compute their correlation, then you compute their similarity in a way. Well, there are many ways in which you can compute similarities. And you don't need to use just a correlation coefficient, which is basically a linear way of computing similarity. One popular way of computing similarity is to use a Gaussian kernel. And that's a method for a distance like the Euclidean distance, and you raise it to the minus power, the distance square. And the result is that you have high resemblance near the example that you compare with, and then the similarity dies off quite quickly as you get farther away from the example.
Read the full transcript
27:53It doesn't die off gradually like the percent correlation coefficient. It just, boom, drops as you get far away. Another more dramatic way of dying far away would be to have just a similarity, which is one in the neighborhood of a certain radius and then dies off completely. So these are various methods of doing that. So those people refer to that as kernels. But think of kernels as just similarity measures. So how do you actually compute a coefficient by which you're going to be classifying examples with multiple examples, not just like with the one nearest neighbor example, which is the nearest example only and base your decision on just that nearest neighbor?
28:36What if you know you wanted to take into account several neighbors and vote among these several neighbors? Well, this is what these kernel methods do. They basically vote among several neighbors according to how similar they are and according to some weights that decide which example is more trustworthy than another. And so interestingly, you can use training algorithms that are very similar to those that are used for linear classifiers, but not for these kernel methods. One of them is the potential function algorithm that was invented in the 60s. And it just consists in adding in the process of learning a new training example in the pool of examples that you're going to be used to make your decision, only if that example was misclassified.
29:22Now you see there's a parallel with the perceptron algorithm. The perceptron algorithm I was saying, you're going to modify the weights only if this example is misclassified. And here we're going to integrate it in the pool of examples only if it's misclassified. And so as you proceed with training more and more, actually you increment by one the weight of that particular example. Every time when you cycle through all the training examples, that example is again misclassified. You're going to cycle many times through your training example. It's very similar to what you do with the perceptron algorithm.
29:54When you cycle many times through the training examples, every time an example is declassified, you change a little bit the weight. And in the case of the kernel method, every time you cycle through an example which is misclassified, you increment by one the weight of that example. Well, not surprisingly, you can show that in fact these two algorithms are exactly identical. They are a dual of one another. You can formulate one by rewriting the other one. And it's just what's in mathematics is called a factorization trick. Basically, you change the parenthesis of position. You have these big summations over all the examples, all the weights, blah, blah, blah, blah, blah, right?
30:33You have a double summation. And you put the parenthesis by grouping one thing in one way and the other thing in the other way, or you change the parenthesis of position and you get one algorithm or the other. So it's very easy. And that's what's called the curl-all trick, basically. The curl trick is just this computational trick that allows you to switch from one vision of the algorithm to the other. So the first vision of the algorithm is compare your new unknown example to all the training examples one by one, calculate the similarity, and then do a weighted sum over the similarities. The other vision is at training time, you're already going to do all these pre-calculations, and you're going to come up with weights for all the inputs.
31:18and that's what you're going to be using to make your calculations. So think of it as a kind of a compilation at training time of all the examples into one weight vector as opposed to keeping these examples and making comparisons with them at test time. So why would you want to do that? Why would you want to use one vision or the other vision? Well, it's because of computational reasons. If you are in the simple linear kernel, Well, the linear kernel basically is like the correlation coefficient. You can go from one method to the other in a very simple way. Either you take your new example to classify, compare it to other training examples, and therefore you have to make as many comparisons as you have training examples that you've kept.
32:03For the support vector machine algorithm, those are called support vectors. You have only a few of them. Or you only make one comparison because you've compiled all the information in the weight vector and you only have to make one comparison at test time. So when will this not work? It seems to be like you would always want to do this one comparison that costs you a lot less. Well, what kills you is that if the vector, if the input vector is of infinite dimension. Now, why would this happen? Well, it's because I told you can go from one method to the other with this factorization trick. But this assumes that your kernel can be expanded in a feature set development, and this expansion can be of infinite dimension.
32:51So when you go in one direction to the other, so if you start with a finite dimension vector and you try to kernelize your algorithm, that is very simple. You replace your original linear algorithm by just using essentially the linear kernel, which is like the correlation. In that case, it's hard to see why you would have a computational advantage to do that, because you replace just one calculation by many. But if you go the other way around, if now instead of having the simple correlation coefficient kernel, you have a kernel like the Gaussian kernel I mentioned before, that Gaussian kernel, unfortunately, if you want to use this kernel to, you need to do an expansion of it.
33:28And that expansion is going to be infinite. So that corresponds to an input space to an infinite dimensional vector. So you don't ever want to see that. You want to operate in this kernel space, in this making comparisons with vectors. Yeah, that's all there is to it. Not very complicated after all. To some people. So this is supervised. You were saying that support vector machines can cluster so they can deal with unlabeled data. And supervised learning has been, certainly on the deep learning side, very powerful and as many, many applications. I wanted to talk about the data problem or the data quality problem that you addressed in your talk at NeurIPS.
34:10And the thing I didn't understand from that is how you address, how you correct or clean very large data sets. I understand benchmarks, but an example, this ChatGBT, its responses in many cases are not accurate because the data is not clean. And a very small example is if you're asking it in natural language to describe a place. And if that place, the place name out on the internet, which is the data source, exists in many different countries, there are a lot of facts that attach to that place name, and ChatGBT doesn't appear to differentiate. So you get a paragraph talking about the place that kind of mixes facts from different places in the world with the same name.
35:04So this data problem, I mean, you mentioned in the talk that one solution is to improve data quality or reduce biases just to expand the volume of data. But that doesn't really work if there are problems in even that larger volume of data. And I understand that you want better curated data, but when you're working with large data sets, how do you curate the data to address these problems? So I don't know if I'm misunderstanding the talk, but that was my question coming out of the talk. Okay, well, yeah, thank you for raising this interesting issue. I mean, it takes a village to raise a child, right?
35:47You know the saying. So it's going to take the entire world to train good learning machines. That's what I believe. A single institution or effort is not going to be capable of providing data of the quality that we need to train well learning machines and also to iterate to make them better and better. My belief is that we cannot expect that in just a few months, we can do the work we do by training a student and sending the student to school and to the university for 20 years or more. This is a very complex thing. What we learn during our education is not just, you know, to be a parrot, which is essentially what current language models are trying to do.
36:29They're trying to predict what's going to come afterwards. So basically, they're learning to be parrots. The amazing thing is that it looks like from this just very simple training strategy, predict what comes next from the past part of the sentence, you can get these chatbots like ChatGPT that give very sophisticated answers. But this should not fool us because really we've trained parrots. And if we want to go beyond that, it will require some constructed, what I call model schooling, where we have real curricula to train models and real exams like we have for students to verify that the knowledge is properly acquired and rectify eventually if there have been misunderstandings or wrong generalizations made by models.
37:24It's going to be a tedious work compared to the excitement that there is now. It looks like you just dump more and more data and you get more and more intelligent models. Well, we're not getting really intelligent models. We're getting more savvy parts. More current models, but not necessarily more accurate models. Yeah. Yeah, right. So what's the way forward? So first of all, we have this new generation of benchmarks. One of them I mentioned in my talk, Big Bench, that try to corner models and try to see if they are used beyond their original purpose, beyond their original envisioned application.
38:03What can we expect? And this is one of the amazing things and also danger. You put out to the public a tool like ChatGPT and people are going to be very creative in what they ask this tool. And they will ask things that are completely out of the box. People have started asking these chatbots to program. So they return programs because they've been trained also on some code or they're asked to solve math problems, which is pretty amazing. they can solve some simple math problems. We don't know, you know, what ideas people are going to come up with. They're going to add these methods. So as a collectivity, we need to try to poke these models very hard and figure out what they're good at, what they're not good at, what are the dangers, et cetera, and iterate.
38:49What's test data today will become training data tomorrow. And it will take a village to train these models. But even on the question of bias, I can see that using less data but higher quality data can improve outcomes. But these models require as such, I mean, large language models certainly, require such vast amounts of data that how do you adjust for bias in a trillion parameters? I mean, how do you know that there's bias there until you happen across it, which has happened in the examples? Perhaps I will give you a disappointing answer. It's open area of research. This is a lot of people are putting in a lot of effort at the moment on this problem.
39:35It's divided in three approaches. It can be pre-processing. So basically you can curate data or you can filter data in one way or another. in processing where the learning machine in the process of learning figures out that something goes wrong and discards the bad samples or reweigh things, wait to make more, to make fairer decisions. Or post-processing, the machine is already learned and you put some adjunct on top of it. That's some of the kind of the band-aid thing, right? You try to filter out bad answers so that you minimize the problematic behavior of the machine. And I realize I've taken up already almost an hour, so I don't want to go too much longer.
40:19But can you draw a thread from your work in support vector machines and tell us what you're working on now? Is it this data problem? And is there a clear logical line from one to the other? Or have you sort of branched off into new areas of research over time? Well, what I think most people don't realize is that algorithms are just a means to an end. And I'm interested in solving problems. So I came up with, you know, this algorithm at some point because it was solving a problem I was interested in at the time. I've never tried since then to really use it when I didn't need it. I've needed it in my consulting practice when I was working on biomedical data because it was particularly adequate for this kind of data.
41:08You can keep using them today combined with deep learning. Evidently, people are doing it even without noticing. But even if you wanted to keep training the representation with the support vector machine in the loop, this can be done. This can be done with something that is called Siamese neural networks, which is something that we also invented back at Bell Labs at the time. And a Siamese neural network is basically a neural network that computes a kernel, that computes a similarity measure. So you can create a kernel with a deep network. And then what you do is that you train your support vector machine with the fixed kernel, which is your big deep network.
41:43You can get, you know, the result of the weights of the support vector machine. And then you can back propagate through the machine and then retrain the weights of the kernel, which is the weights of the Siamese network. And you can then keep iterating with these two steps. Some people have been doing that. or you can just disregard the original support vector machine algorithm. And at the very beginning of my interview, I was explaining to you that the reason why I wasn't immediately implementing the support vector machine algorithm when Vapnik came, because I was using this other technique that was similar in spirit.
42:17That other technique has been formalized in an algorithm where it shows that stochastic gradient can be used to train support vector machines. And it's implemented in the scikit-learn library. So what you can do is that you can train a deep network with the support vector machine algorithm using stochastic gradient descent. Everything can be done. But at the moment, I'm working on other things. And I'm particularly interested in branching out on using these very large language models in several applications because I think they're doing amazing things. And possibly I will need support vector machines at some point, but not necessarily.
42:55So it will depend on what are the issues that we are facing. Issues that we might face indeed are cases you mentioned when we have very little data. So the few short learning problem, few example problem. In that case, it makes sense to have a notion of support vector, a notion of examples that kind of unique to represent your borderline cases, your special cases. But this is just one sub-problem of a much larger, much broader range of problems that we can face. Do you have time for another question? Sure. I wanted to ask, you know, Jeff Hinton gave us a new algorithm, a forward-forward algorithm.
43:35I've been asking people for their impressions. He's famously interested in figuring out how the brain works and forward-forward possibly gives an insight to that. But I'm just curious whether you've read the paper and what your thoughts are about how it might advance the field or be used. Well, Jeff Hinton has surprised us many times, so we don't know. Maybe this algorithm will take off or will inspire other algorithms at some point to really advance the field. What's for sure is that about 10 years ago, before, you know, the boom of deep learning, Jeff Hinton was pushing very much for unsupervised learning and using these methods that were developed by Joshua Benjo and others of stacked autoencoders.
44:23And at the time, DARPA organized an evaluation of deep learning, and I was part of the evaluation team. And we organized the challenge. DARPA pushed us to have a protocol to prove that unsupervised learning could do the job and that the supervised learning wasn't really needed to learn representations. And we organized the challenge at that time. And sure enough, much to my surprise, it is true that you can train systems to learn representations in a completely unsupervised way. And it cannot be beaten by supervised methods. And this is what we also see now with the boom of self-supervised learning.
44:57You can do almost everything unsupervised and add the supervised layer at the end. And that even is beneficial if you want to be robust against some sorts of bias. Yeah, so I'm, you know, all in favor of people like Jeff Hinton who do research a little bit out of the mainstream and continue pushing these ideas that could end up becoming very important in the future. But yeah, it's certainly something to keep an eye on. Yeah. And you're famous, aren't you, for not following the mainstream through charting your own course? Yeah, well, I don't know whether I'm famous for that, but I'm doing my best not to be influenced by the mainstream and try to venture in new directions.
45:40Of course, at the risk of falling into a hole, but I find this more fun than racing on the groom slopes. That's it for this episode. I want to thank Isabel for her time. If you want to read a transcript of our conversation today, you can find one on our website, I on AI. This episode in particular packs a lot of information, and so I encourage you to download and read it. In the meantime, remember, the singularity may not be near, but AI is about to change your world, so pay attention. I also want to thank our sponsor, MindStudio by UAI, which is giving creators the opportunity to build and deploy generative AI apps for profit.
46:29UAI has an emerging AI marketplace, and MindStudio is the best way to build apps with generative AI. Anyone can do it. MindStudio uses conversational language to program incredibly powerful AI tools. No coding knowledge is needed to start your AI business today. Check them out at uai.ai and start building your AI app today.
From the publisher
This episode is sponsored by MindStudio by YouAi. MindStudio is the best way to build an AI business. Start driving some serious revenue before everyone else. Mind Studio allows you to use conversational language to program incredibly powerful AI tools. No coding knowledge is needed to start your AI business.
Sign up now- https://bit.ly/MindStudioEyeonAI
On episode #139 of Eye on AI , Craig Smith sits down with Isabelle Guyon, a pioneer in the world of machine learning and pattern recognition, and Chair Professor at the University of Paris-Saclay. Isabelle Guyon, renowned for her groundbreaking work with Vladimir Vapnik and Bernhard Boser, has made significant contributions to support-vector machines (SVMs), artificial neural networks, and bioinformatics.
In this episode, we delve into SVMs, a technique that efficiently classifies data by maximizing margins, with applications spanning image classification, text classification, and bioinformatics.
We also explore maximum margin classifiers in the biomedical field, uncovering their unique approach to category distinction. Learn about weighted sum usage and their adaptability based on available data. Discover how maximum margin classifiers apply to labeled and unlabeled data, for both classification and regression tasks.
We wrap things up by discussing the promising prospects of chatbots and language models. Isabelle Guyon shares valuable insights on data preprocessing, bias adjustment methods, and the fusion of SVMs with deep learning to address contemporary challenges.
Isabelle Guyon's LinkedIn: https://www.linkedin.com/in/isabelle-guyon-aa371170
Craig Smith Twitter: https://twitter.com/craigss
Eye on A.I. Twitter: https://twitter.com/EyeOn_AI
(00:00) Preview
(00:52) Introduction and MindStudio by YouAi
(04:55) Machine Learning Techniques
(15:15) Classification Methods in Biomedical Applications
(22:40) Support Vector Machines and Kernel Methods
(36:45) Future of Chatbots and Language Models
(41:08) Unsupervised Learning in Deep Learning Importance
(45:53) Outro and MindStudio by YouAi




