In short
Podcast Summary: NVIDIA RAPIDS and Open Source ML Acceleration
Podcast Information
- Podcast Title: Software Engineering Daily
- Episode Title: NVIDIA RAPIDS and Open Source ML Acceleration
- Guests: Chris Deotte (Senior Data Scientist at NVIDIA) and Jean-Francois Puget (Director and Distinguished Engineer at NVIDIA)
- Host: Sean Falconer
- Description: Discussion on NVIDIA RAPIDS, an open-source suite of GPU-accelerated data science and AI libraries, and its implications for machine learning.
Key Topics Discussed
Introduction to Kaggle and Grandmaster Status
- Kaggle Overview:
- An online community for data science with over 20 million users.
- Users can participate in discussions, share code, host datasets, and compete in competitions.
- Grandmaster Title:
- Achieved by winning five gold medals in separate competitions.
- High distinction in Kaggle’s ranking system, especially in competitions.
Kaggle Competitions
- Competition Structure:
- Competitions involve solving real-world data science problems, such as predicting sales or classifying images.
- Duration usually spans three months, with submissions evaluated on hidden test sets to prevent overfitting.
- Value of Participation:
- Learning from challenges, community members, and gaining hands-on experience with state-of-the-art models.
NVIDIA RAPIDS
- Overview of RAPIDS:
- Suite of libraries (such as cuDF and cuML) for GPU-accelerated data science tasks.
- Primarily enhances performance of Python libraries like Pandas and scikit-learn.
- Speed and Efficiency:
- RAPIDS can be up to 100 times faster than traditional libraries by using GPU for computations.
- This acceleration enables more experimentation and faster iterations on modeling.
Feature Engineering and Model Performance
- Feature Engineering:
- Importance of creating new columns and combinations to improve model accuracy.
- Automated feature engineering becomes feasible with RAPIDS due to speed.
- Techniques discussed include target encoding to avoid overfitting.
Challenges of Tabular Data
- Difficulties in Predictive Modeling with Tabular Data:
- Tabular data often exhibits complex behaviors that differ from image or text data.
- Deep learning models struggle with tabular data due to its variety and unpredictability.
- Gradient Boosted Trees:
- Effective for tabular data because they can handle discontinuities and complex interactions between features.
Future Directions in Machine Learning
- Role of Large Language Models (LLMs):
- Anticipated to transform the workflow of data science by assisting in code writing, model training, and experimentation.
- Potential to generate synthetic data, though challenges remain in ensuring quality and preventing reverse engineering.
Key Takeaways
- Importance of Competitions: Engaging in Kaggle competitions fosters learning and enhances professional visibility.
- Impact of RAPIDS: The RAPIDS suite significantly boosts efficiency in data processing and model training through GPU acceleration.
- Hybrid Approaches: Combining deep learning with traditional machine learning techniques can yield superior results.
- Feature Engineering: Automated and systematic feature engineering is becoming more feasible and crucial for improving model outcomes.
Conclusion The episode provided an in-depth exploration of NVIDIA RAPIDS and its implications for machine learning, highlighting how GPU acceleration can transform data science workflows. The conversation emphasized the ongoing importance of competitions in skill development and the future potential of language models in the field.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00NVIDIA RAPIDS is an open source suite of GPU accelerated data science and AI libraries. It leverages CUDA and significantly enhances the performance of core Python frameworks, including Polars, Pandas, Scikit-Learn, and NetworkX. Chris Diot is a senior data scientist at NVIDIA, and Jean-Francois Pouget is the director and a distinguished engineer at NVIDIA. Chris and Jean-Francois are also Kaggle Grandmasters, which is the highest rank a data scientist or machine learning practitioner can achieve on Kaggle, a competitive platform for data science challenges. In this episode, they join the podcast with Sean Falconer to talk about Kaggle, GPU acceleration for data science applications, where they've achieved the biggest performance gains, the unexpected challenges with tabular data, and much more.
0:48This episode is hosted by Sean Falconer. Check the show notes for more information on Sean's work and where to find him.
1:07JFP and Chris, welcome to the show. Thanks for inviting us. Thank you. Yes, absolutely. I'm excited to get into this. So I think we have a lot to cover, but I wanted to start off by talking about Kaggle and being a grandmaster, which is a distinction that I believe both of you have. For those that are unfamiliar with this concept, can we start there? And what's it mean to be a grandmaster? Yeah, so Kaggle for me means many years of entertainment. I've been participating for six years. It's an online community for data science, and there's currently over 20 million users. And on this platform, you can engage in conversations, you have access to Jupyter notebooks, you can share code, host data sets, and you can also compete in competitions.
1:51And the website, you can earn achievements and you can gain titles. And yeah, you've heard people say Kaggle Grandmaster. So what is that? So that's one of the titles you gain. It's the best title you can acquire. And you can actually become a grandmaster in the four categories, discussions, notebooks, competitions, and data sets. And the most desired one is the competitions grandmaster. And to achieve that, you need to actually win five gold medals in five separate competitions. And one of them has to be a solo that you want to buy yourself. And the competition is incredibly difficult on the website.
2:26People are competing from around the world. typical competitions have thousands of people. So it's very hard to obtain. And there's only, I think, a couple hundred competition Kaggle Grandmasters in the world. It's an amazing thing. And then I'll mention, you could also, as I said, get the other Grandmasters. So you might have heard the expression, a double Grandmaster or triple Grandmaster. That's someone who's actually received awards, hosting discussions or notebooks, and they've acquired another Grandmaster title. And then you could sort of stack them up when you cite them. I would add competition grandmaster, it's based on your merit, how the quality of the models you build.
3:03The other ones are based on community votes. So it's a bit different. I would also define Kegel as a legal drug. It's really adrenaline flows when you compete. It's really addictive. So when you start, you can't stop. Can you talk a little bit more about the competitions? What does a competition consist of? Is this something that's happening live or is it more like a problem goes up and then people are asynchronously putting time into that and you have to essentially try to solve it within that timeframe and come back with a solution? Typical competition is like a short time data science or machine learning project.
3:42You're given some data sets or data. Kaggle curates a data set. And Kaggle also curates a question. For instance, you want to predict next month's sales for a retail chain. And the data is passed by product, by store, by what have you. Or it could be some image classification, medical image, diagnosis, is this a cancer or not, and you have images to train. And duration is typically three months. And basically, you have to submit some code that will be run on a hidden test set. And then a number is computed from your predictions, and that's the score. you see a leaderboard of the score obtained from some of the data set.
4:31And after the competition, they will compute the final score on the rest of the test data. And the reason they do this is to avoid what is known as overfitting. So making sure they select good models and not lucky ones. There is another form of competition they call analytics, which is a bit different. It's also a form of data science. you're given some data and you have to find an interesting story out of it. And then it's judged by a human jury. So Chris, who actually creates the competition? What sort of expertise do they need to be able to create these competitions for presumably some of the people that are the best in the world at this particular job?
5:14They're sponsored by actual businesses. So a business approaches Kaggle with a certain challenge they need solved. So maybe a university wants a model that can read student essays and assign scores to it. So they approach Kaggle and say, I would like you to host a competition. They put up money, they put up cash prizes, and then Kaggle will kind of help curate the data and do all the infrastructure and logistics. But it is interesting to note that it actually starts from a business need. So the competitions are real problems and the company afterwards gets, when they give out the prizes, they've received the code of the top solutions.
5:49And oftentimes they'll immediately implement the code. So it's nice to know that you're competing, you're helping a good cause, and that your code can actually be used to solve a real world problem afterwards. And what do you personally get out of participating in these beyond just the satisfaction of a job well done? I'd say the main thing is you learn. You learn both from the problem and reading relevant papers or blogs or code base and from the community. So if you want to know state-of-the-art models for a given topic, best is to enter Kaggle competition on that topic. And you will learn from the top teams, so the winners, the prize winners, usually they have to disclose what they do to get the prize, so you can learn as well from that.
6:39And then how did both of you get involved with this? And for those that are maybe interested in dabbling, how would you get started? A friend recommended it to me six years ago, but I would say that the kind of purpose it played in my life was the learning process. So my formal training is in PhD in mathematics with a specialization in computational and simulation. And then I started learning data science on my own. And after you learn all the ideas, I wanted a way to practice, to test it out, to build some models and to talk with people. And then someone said, hey, do you know about Kaggle? So for me, it was wonderful.
7:14Immediately, I met people to talk with. There was problems to solve. There was competitions, playground competitions. So really, I went to it for the learning process. And then as JFP said earlier, it's highly addictive. I mean, once you're there, you get involved in a comp or just talk with people. It's tons of fun. And then you're checking it all the time and you're participating in more and more. But it's really helped my learning tremendously. Yeah, for me, it's a bit different. I did a PhD in machine learning, but in the previous millennium. So it's irrelevant now. And then I went during my professional life working on something else and mathematical optimization.
7:51Then at my previous employer, people saw I had some machine learning background. So I said, oh, why don't you go back at this? We are developing tools. And I say, where can I find an update on what's the current state of the art machine learning practice? I found Kaggle, watched a little bit, and then I jumped in the water and got hooked. And it's a great, again, it's for anyone developing machine learning and data science tools. That's a wonderful place to see what are the needs of Fortulimit. Yeah, I totally get the addiction component of this. I was never involved in these types of competitions, but I did compete in things like TopCoder and the ACMICPC programming solving competitions through university.
8:40I became completely addicted to that experience competing in these. And I think one of the things that even though they're not necessarily business-driven problems that I got out of participating in this is that it just made me a lot more comfortable with software engineering because I was putting so much time into just the act of practicing. So much of coding really sharpened my skills and made me way more employable than I was before, even if it wasn't necessarily directly the types of problems that you'd be solving day-to-day at work. I'm curious, how is this experience of competing in these types of competitions translated into your day job and how you're leveraging some of those modeling techniques and other things that you've learned in your day-to-day job?
9:25So participating on Kaggle has just taught me so much and made me a better data scientist. So yeah, again, I just learned tons of new techniques, how to do things correctly. So I'll mention that, you know, I had read a lot of books, but there's a lot of techniques that you learn on Kaggle that are not really in textbooks yet. And also oftentimes a lot of new things, I think even gradient boosted trees was developed on Kaggle. So basically people, I mean, it is on the fringe of research. It's the latest ideas. You're learning the best techniques. And you're also get a chance to work problems in all different domains from computer vision, natural language processing, tabular data.
10:01So just all that exposure. And then I guess one thing I'll add to the way they set up a competition with a hidden test set, you really have to make your model generalized to unseen data, which is one of the most important things in the field, building models in the field of data science. So, yeah, all the skills you learn, plus repeatedly learning to make models that truly generalize, then immediately when I'm inside NVIDIA of building models, working on projects, it's just all that knowledge just comes and it all benefits what I do. I would add that we both got our job at NVIDIA because we were Kaggle competition grandmasters.
10:38So that's also a nice outcome of all the learning we got. And we have like 15 or 16 Kaggle grandmasters at NVIDIA now. Yeah, I think that's something that I always think about and recommend, even going back to my own experience in things like TopCoder and ACM ICPC competitions, is that top tier companies are paying attention to these types of competitions. So if you're interested or just starting your career in the space, competitions like this is a good way to not only learn and build up your own skill set, but sometimes you might have a company that just comes up to you because you participate in something like this or you've done well in them.
11:17And even beyond just being ranked in the top 10 in the world, not everyone's necessarily going to achieve that, but it just shows that you have a passion for the space and that you're pushing yourself and learning and working at it. that is also really attractive to companies. So I wanted to talk a little bit about NVIDIA and the Rapids platform. This is an open source suite of GPU accelerated data science AI libraries. So first of all, what problem is this helping data scientists with? And how does that set of libraries actually work? Maybe Chris, let's start with you. Okay, so it's a whole suite of libraries and it helps with a whole variety of tasks.
11:52But its main goal is to speed up all sorts of things. So the two libraries that I work with the most are QDF and QML. And QDF helps with all your data frame needs. So it's got an API similar to Pandas. It's on all the same functionality. And with that, you can speed up all your data frame needs. I should probably take a step back and sort of say, you know, maybe why was it, or in my opinion, what role did it play? But today, all companies are getting more and more data, and it's getting harder and harder to process all the data. So even things like computing statistics, doing data framework, we need to run that faster.
12:25So that's where the QDF comes in. Basically, it does all the computations on GPU. And it's, you know, it could be 100 times faster than using other libraries. And then as we move forward with more data, it's going to be getting faster and faster. So that's great. And then I also use a lot QML, which has a similar functionality as a scikit-learn. It does machine learning models. But once again, it'll train all these models on GPU and do them much faster. So if you're doing tasks requiring support vector machines, K &N, and other models like this, they could train the models hundreds of times faster.
12:58So basically, I would say that, yeah, it helps with things that maybe we've been doing all along. But because now it moves the process to GPU, it's incredibly faster. And if you're working on experimentation or you're iterating things and trying to make more accurate models or just try to get your work up quickly, I would say it's becoming a necessity with how data and everything is growing. In terms of the GPU acceleration that's happening, what needed to happen in order to make it so that you could do things like KNN, for example, on GPUs rather than traditional CPU? As Chris said, Rapids is more than that, but it's also the GPU accelerated version of Pandas, Polars, and Scikit-Learn.
13:39There is more, there is graph, there is signal processing. But recently, over the last year, we made the move from CPU-based to GPU-based seamless. So if you have a nice Pandas code, you just have, in your notebook, you just have to load an extension at the start, and then all your code will be GPU-accelerated seamlessly. You don't need to change any line. And more recently, we did the same with Polars. So this is a way for people to just experiment what they gain from moving to GPU very, very easily. Is there, I guess, a cost associated with that that you have to take into account given that this is running on GPUs?
14:20I guess the cost would be just that, I guess, obviously you need a GPU to run on a GPU, but a lot of modern day systems have both a CPU and GPU in your system. So I think for most people, it's a matter of just flipping the flag and you'll just immediately get speed up and it'll just use your machine's GPU. Yeah, I just agree. When we say cost, my answer was on the time of the data scientist cost. We reduced this to the bare minimum, but there is still a compute infrastructure that remains. Besides getting a better 100x, better performance, does the fact that you can do these things so much faster so that you're sort of shortening the learning cycle also change things in terms of how you think about building models or how it might even impact your existing work?
15:10Yeah, I see data science and machine learning as an experimental science, just like physics. So ideally, to build a good model, you have a baseline, you want to improve, you design, you say, oh, I have this idea, maybe more data, maybe different parameters, what have you, you design an experiment to test if the change is really improving. And then you run it and you look at the results and depending on the result, it becomes your new baseline. So if you can do this faster, you will try more ideas and just that will lead to a better model just because you can experiment more. You can perform way more experiments in a given time.
15:53And then does that also change from an experimental standpoint? Like, I guess the types of models that you might be able to try in a given data set, because now you're less worried about how long it's going to take to train something you can move much faster. Yeah, it absolutely does allow you to do new things. I am actually, so yeah, I guess there's sort of two things that enable to do. So JFP pointed out that you can do experiments faster. So you can do what you were previously doing, but we could do it better because we could try out more things. But it is actually doing a second thing, which it's allowing people to do things that we're not previously even able to do.
16:30So for example, if you would try to use K &N, or actually one thing you could do is you could actually take tabular data and you could push it through UMAP to create features and then put that into an image model and do these kind of weird pipelines. But back in the day of running this on CPU, you really couldn't use some of these models that are way too slow. So we recently saw a coworker, a colleague won a competition where he actually used a combination of deep learning and machine learning. So deep learning has a backbone, which sort of generates features, the head sort of will then do the regression.
17:04But because QML has accelerated machine learning models so fast, he was actually able to just take the features out of the deep learning model, and then train support vector regression. And he was able to do this cycle over and over so fast because of the new speed, that in the end, his model won first place, and it was a hybrid actually. So it was actually a combination of a deep learning model fused together with a support vector regression head. And these hybrid models and other, there's been other advances in feature engineering. So there's a lot of new techniques that I'm seeing that are a direct result of having this speed and we can sort of do some new model designs and some new techniques.
17:43And another thing we can do is also to just run deep learning models on tabular data. So there are a lot of papers claiming it's the best We don't usually, that's not what we find. Gradient boosted like XGBoost, like GBM, CADBoost, all GPU accelerated, still outperform. But when you assemble, so you take one of these and you take some transformer or some other deep learning models and blend the prediction together, you improve over a single model. So on Kaggle, it's used a lot. Of course, you count running deep learning on CPU only. are not great. Yeah, definitely. You mentioned this hybrid approach.
18:26And I think a lot of data science works or traditional data science work, we think about predictive ML. And now there's a lot of focus on general AI and general deep learning techniques, stuff like that. Do you think that because so many people are excited about what's happening in general AI and there's so much hype around it, that sometimes we lose sight of the fact that predictive ML can still do a lot of useful things? We kind of try to throw maybe too big a model is something that we can actually solve with a simpler, bespoke, trained predictive model? I would say it depends on what you want to predict.
18:58Generative AI, as the name indicates, will generate something, a text, typically, or images with diffusion models. So if that's what you need, of course, that's what you should do. But if you want to forecast your sales for next quarter, you need to predict numbers. That being said, so classical machine learning, So regression models or classification models, deep learning or not, are still the way to go in those cases. But we do find that if your input is text, and for instance, you want to do text classification, say spam detection, or classify in few categories using a generative model, LLM, but only take one token, just ask it to output one of few options, this is a great classifier.
19:46And that's quite interesting because you benefit from all the investment and progress in those elements. So you're talking about essentially instructing the generative model to produce an output that's within a specific range. Like if I only want a value between zero and one based on the probability that this thing is to indicate that this thing is part of a particular category. This is beating the just the encoder only models like DiBerta, Roberta. They are much larger, so no surprise, but they improve. Chris, did you have any thoughts on this? Yeah, so you're in question about people throwing too big a model at it.
20:22And you are right now with LLMs and the reason they're getting better and better, it's sort of more tempting to do that. But this has been an age-old problem. I always see this. There's a lot of people just throw the biggest model they can. But I think it's always been the case that we should try the simple models. And I love doing it because it's really fun. There are definitely times when the simple model can outperform. That's exciting. A lot of times the big model can do as good as a little model, but then it's inefficient. You don't want to use more compute than you need to. So when I'm given a problem in the early stages, I actually like to try a whole range of models, simple and even complex.
20:55And even lately, I have been throwing LMs at every problem I can just to kind of say, can it do this? Can it do this? Right. But, you know, you do the whole range. And then in the end, I generally try to go with the smaller models, the simpler models. Yeah, I like the hybrid approach to where you could use a model or a particular model on the backbone of some generative AI model to check essentially the answer and then do those iterative steps. So I want to talk a little bit about tabular data prediction. So can you talk a little bit about why this is such a challenging problem? Why have people been focused on this and interested in it for such a long time?
21:30Yeah, so it's really a good question because we see deep learning becoming the way to go for all sorts of data modalities, except for tabular data where the jury is still out. There are many reasons, but it depends on which tabular data. If the data is measured from a physical, like you have weather data, what have you, a deep learning model is likely to be better. So my hypothesis, it's not science here, it's if the data is sampled from the physical world, the physical world is smooth and deep learning will work well. If it's sampled from human decisions, like people's behavior of sorts, sales forecasting or what have you, it's much more discrete.
22:22And I would say chaotic in a scientific way, so hard to predict. And there are smooth models like deep learning models, not as good as, say, gradient booster trees that can handle discontinuity very naturally. But that's just one angle. Chris, you may have another one. I've been actually fascinated by this particular question for a very long time. A lot of researchers have been wondering because we saw a transformation in computer vision and natural language and text, natural language processing about a decade ago. Right. So starting about a decade ago. So before a decade ago, people would actually take in computer vision, humans would actually engineer the features.
23:06So they would actually take images, process it, extract features, and then put that through a machine learning model like a support vector machine. And they did similar things with text. But then we invented deep learning and then deep learning totally on its own. It does the feature engineering and does predictive. So that's revolutionized computer vision and natural language processing. You can download pre-trained models, fine tune them. But that's yet to happen in tabular data. So I would say that the best tabular data models are still involved human handcrafted engineered features where we make new columns.
23:40So I am particularly looking forward and curious, will the day come when there'll be some sort of deep learning model that can digest a variety of different tabular data frames and essentially engineer features and do it on its own? And I think the reason it's challenging is I think that the data is much more variety, right? So images all share the fundamental building blocks of lines and shapes and text has fundamental building blocks of words. But what is the fundamental building block of tabular data? Here's statistics from a finance company. Here's data from medical data. Data is sort of so different.
24:21It's going to take something that's going to actually have to see how is it all similar. What's the common theme? Maybe the common theme is some kind of cause and effect or logic or reasoning, but some model has to sort of understand it and find all this similarity. And then maybe then it can engineer on its own and it can use past learnings to help with future problems. And because the data sets are so varied and different, could you end up with a situation, too, where if you had a ton of tabular data to train on, the model might not actually tune itself to recognize the patterns? And essentially, you know, the pattern recognition has less to do with the data.
24:56It's more about the structure, the fact that it's organized in the rows and columns. You essentially end up with sort of biasing the model and what it's trying to predict. And essentially leading to a place where you're sort of overfitting the model against the wrong pattern. The latter is not happening. When you have an image, you have pixel orange in 2D. If you have video, 3D. If you have text, you have numbers. in one dimension, same for audio. So it's very regular organization of the data. So you can train a model once, and it works with, hopefully, a lot of the instances. Tabular data, sure, it's 2D, but sometimes the columns are independent, so you can shuffle them.
25:42Sometimes they are not time series. Sometimes there's a huge correlation between columns. Some columns are not there. So as Chris said, the format is not specific enough. at this point. Maybe sometime people will have trained a model on every tabular data available online and claim it's a foundation model. Actually, some people do claim they have foundation model on tabular data, but on Kaggle, we don't find they are the best models yet. How do Pusted Trees work on this type of problem? I'm less familiar with that. Is that a variation of a decision tree? Yeah, so it's actually an ensemble. It's a linear combination of multiple decision trees.
Read the full transcript
26:22So yeah, boosted trees are, they just repeatedly make decision trees. And then each new decision tree, it trains on the previous cumulative error and it tries to reduce that. So it keeps just adding a new tree. And the purpose of the new tree is to kind of reduce the error a little bit more. And then in the end, you just combine all the, you just basically, you take an ensemble of all the trees and that's what it is. You could see in terms of deep learning, it's a gradient descent, but each update, you don't update existing weights in a model, be it linear regression or deep learning. You add a tree that implements the gradient delta.
27:03And people think it's a recent technology because the first useful implementation is Exibust. It's like 10 years old only, so it's quite recent. But the theory was published 25 years ago, more or less. Are there libraries within Rapids that help doing some sort of feature engineering? Yeah, absolutely. I think that's another advantage of the speed of Rapids. So we just discussed how, yeah, so QDF is one, and then newly the QDF pan is included at Polars. But specifically, to improve model accuracy, what you often do is, given a data frame, you'll make new columns. and there'll be transformations or combinations of old columns.
27:45That's what feature engineering is. And it's done sort of manually. But with the speed, so with QDF, which operates on data frames and the speed at which it works, you can sort of systematically go through a whole set of transformations. Like let's randomly pick pairs of existing columns, combine them together and then target and code it. And we'll make a new column. And then we'll see if that improves the model, right? So you could just, you could basically build these four loops where you just systematically go through typical things that humans will try. And then you can train a model and see if it proves.
28:18And actually, I recently won a competition doing just that. It was a Kaggle playground competition. You had to predict insurance premiums, maybe like a car insurance, the annual premium. And basically, I just set my computer running overnight. And it actually tried tens of thousands of... So it had the original data set had 23 existing columns. And I randomly picked groups of two, three, four, five, or six. I combined them, targeting code of it using QDF. And then I trained a model to see if it improved the validation score. And it just keeps... So in total, there may have been something like 150 ,000 combinations to try.
29:00And it just randomly tries them. And it found hundreds of ones that worked successfully. And then I added them to my final model and it boosted the score tremendously, so much so there was even a gap with the second place. So this was only made possible by the speed. If I had tried to do those data frame operations on a CPU library, literally the search would have taken months. So it would never have finished. But this search just happened overnight. So absolutely, this speed is allowing us to actually do some automated feature engineering. I would add to expand on something Chris mentioned. So he built on top of Rapids, but he used a built-in component called target encoding.
29:41So it is a way, in tabular data, you have basically two types of data. One is just category, you know, you have a color, you have, and there are ordered numbers, you know, the weight of someone or whatever. For cardinal categories, it's very hard to manage for algorithms like linear regression, supervector machines. And basically, you have to create additional data. It's called one hot encoding, one column for each possible value. Then they can be combined linearly. It's a pain because it expands your data tremendously. You have to use sparse implementation. It's not, it exists in QDF in Rapids.
30:27But there is another way which is smarter, which is to say basically, say you have a category with five values, you just average, say you want to predict some numerical value out of your, you just average for each value, the target you want to predict. This gives you an indication of how good this value is. This can be done automatically. It's called target encoding. But if you do it the way I do, you overfit to it because you include the target. It's tricky. You need to use what we call out-of-fold prediction to avoid. So you never use a target of a row to compute a value for that row. So there are ways to do target encoding for one row using other rows.
31:16This is built in in Rapids. And then what Chris used was to apply this target encoding on column combinations. And that's very useful. So at the upcoming NVIDIA GTT conference, we're actually giving a workshop, which we're teaching this exact technique, you know, how to target encode, how to use QDF to do that. Also some other encodings like count encoding. I hope that I advertise that you all check it out. It's going to be a great time. So we're going to make the features and then we also train some models and show how it improves the models. But I would say for tabular data, it's probably the most effective and sort of powerful technique to improve your models.
31:55And it's time and time again, it's kind of been the key component to sort of winning these category competitions. It's actually a hands-on workshop where if you're there, you can work along with us and follow the code and it's going to be great. So I suggest that everyone checks it out. It'll be taught by some KGmon and some other NVIDians. Yeah, awesome. I'm hoping to be there too. So hopefully I can participate in that. But in terms of like even in competitions, how do you determine where you need to focus on making improvements to your output? It could be part of the feature selection process.
32:28It could also be part of the model that you're using. There's a lot of things that could go wrong and you only have limited time. So how do you figure out where to actually spend that time? That's a great question because basically you have some available budget, so the time till end of competition where you can work on it. You may also have a compute budget. And you need to allocate the resource the most wisely. So what I do is I make sure first I have a good test harness. I can really evaluate my model. So typically with a baseline, create a cross-validation setup and try different baseline models, submit to the competition to see if my cross-validation correlates with the score.
33:16If it does, then I don't need to submit too much, work with my local setting. And then, indeed, between feature engineering, trying different models, implementing a more complex workflow, then it's a combination of where you have some feeling based on past experience and the low-hanging fruits. You estimate the time it takes to code it and run it. There is a part of luck if you investigate the right thing first and you do better. And hopefully after years of doing this every week, we get some feeling of what might work first. Yeah. So you start to build an instinct, essentially, when you see something that maybe feels like it's underperforming and you can understand where that problem might be.
34:05Yeah, absolutely. So I've been in 80 competitions in the last six years and I have such strong intuitions. You know, I'll train a model. I'll look at its output. It'll always be, you know, getting something wrong. And I oftentimes know exactly where to look that, oh, you know, I could even, you're really hard to get a sense of how you should, you know, alter the model architecture or the training procedure or how you should augment the data or this, that you really get a, it's amazing. thing. I actually always make analogies. So for NVIDIA, I was the teacher at the university. And I'm always making analogies that for me, training a model is actually teaching a student.
34:44And you get better with time as a teacher, right? And you learn how to listen to your students. So you teach a student and then you have them do a problem and you watch them or you watch, and then you see how they do it. And then they get the wrong answer. But you look at their work and you see, I see they just forgot to divide by two here. And you start to learn what the common mistakes are. And then, you know, when you talk to him again, you have to emphasize the divide by two. And it's the same thing with the models. I'll see models make common errors. I'll sort of know how to address it, know how to change things.
35:15And I would just add that I also don't rely on automated optimizing tools. I see people using, say, Optuna to tune parameters. I always do it by hand because that way I get some intuition. I learn from my experience. If I rely on the black box optimizer, well, maybe I will learn how to use the optimizer better, but I will have no understanding of what works under the hood. In terms of where data science tooling is going, if you can make one prediction, I guess, where are things moving, progressing? Where do you think the next big breakthrough is going to come from? I would say one thing, and we're already seeing it, is how large language models are going to completely change the workflow.
36:04So they're basically going to be, we're going to be working together with them. So already we see them helping write our code. We see, you know, co-pilots, people basically ask them questions. They can suggest ideas. So already, so take a project from start to finish. You know, a company comes to you with a certain task. Here's our data. Or even we want you to be able to predict this. And then to finish, like here's the finished model and here's it does. And that, you know, there's all these different steps and different roles of people involved in the process. But more and more, we're going to see LLMs get involved in all different steps of the process from, you know, the beginning EDA, writing code at various points, giving suggestions for this, maybe even taking charge of an experimentation cycle and then running experiments on its own and changing things.
36:47So it'll be really exciting to see how they'll be utilized more and humans will be working, I think, together with language models in the whole process of building a final model. Even around generating test data is massively useful. yeah test generation definitely I will come back to this but I would add to Chris with Kaggle we focus on the modeling part of data science and machine learning but getting the data and then using the model we create in production requires coding which is something data scientists may not be good at and transferring to software developers that are not good at machine learning you also lose something So maybe we see LLM used as coding assistants, gaining traction.
37:34They could be assistants for data scientists to write the code they don't want to write to connect before and after. Back to generating testing data, we do already do it for text and images. There is a lot. For tabular data, there are people, I think generating tabular data, to me is not mature enough, except if you model some physical phenomena. In Kaggle, there were a number of competitions running on synthetic data for astrophysics, for particle physics. And there, the simulator, the data generator, was great because it was based on physics principle. There are playground competitions, but Chris is doing more of these.
38:22always have a bit of a fear that modeling means reverse engineering the data generator. But Chris, you may disagree here. I don't know. Yeah, what he's referring to is, so Kaggle has increased their frequency. Every month, they're offering a new playground competition. And it's very hard to offer competitions that often because the most difficult thing is getting data sets. So recently, they've been using synthetic data sets where they're either generated by LLM. So LMs essentially make the data. And the risk has always been when data set is synthetic, you know, you can actually sort of reverse engineer because somehow it's making new data with a target.
39:01So if you can think how it thinks and how did it assign targets and how did it make new data, then you don't have to actually forecast the insurance price. You just have to figure out how is the data made. And you do see this often from time to time. people do figure this out and they win comps because they've reversed engineered some process. Yeah, it's something we have to be careful of. But I think as time goes on, the synthetic data is getting up a higher quality, but there still are artifacts that you can take advantage of a little bit. That's always a risk when using synthetic data. Yeah, absolutely.
39:34Well, we're coming up on time. JFP, Chris, I want to thank you so much for being here. I thought this was really, really interesting. Hopefully, we'll see each other at the workshop at India. Thank you for inviting us. Yeah, I look forward to meeting you in person, John, at the conference.
From the publisher
NVIDIA RAPIDS is an open-source suite of GPU-accelerated data science and AI libraries. It leverages CUDA and significantly enhances the performance of core Python frameworks including Polars, pandas, scikit-learn and NetworkX. Chris Deotte is a Senior Data Scientist at NVIDIA and Jean-Francois Puget is the Director and a Distinguished Engineer at NVIDIA. Chris and Jean-Francois
The post NVIDIA RAPIDS and Open Source ML Acceleration with Chris Deotte and Jean-Francois Puget appeared first on Software Engineering Daily.
