In short
Practical AI Episode Summary: Accelerated Data Science with a Kaggle Grandmaster
Episode Overview In this episode of Practical AI, hosts Daniel Whitenack and Chris Benson dive deep into the world of Kaggle competitions with Christof Henkel, a Senior Deep Learning Data Scientist at NVIDIA and a distinguished Kaggle Grandmaster. They discuss the intricacies of Kaggle, the impact of participation on a data scientist's career, and how to maximize AI productivity using GPU-accelerated tools.
Key Participants
- Christof Henkel - Senior Deep Learning Data Scientist at NVIDIA, Kaggle Grandmaster
- Chris Benson - Tech Strategist at Lockheed Martin
- Daniel Whitenack - Data Scientist at SIL International
Key Topics Discussed
Kaggle Overview
- What is Kaggle?
- A platform for data science and machine learning, originally known for hosting competitions but has expanded to include discussions, shared notebooks, and datasets.
- Participants earn medals in various categories (competitions, notebooks, discussions, datasets), leading up to the title of "Grandmaster" for top performers.
- Kaggle Grandmaster Definition
- Out of 10 million users, only 280 have attained the Grandmaster title in competitions, indicating elite status.
Christof's Journey
- Entry into Kaggle
- Christof began using Kaggle during the final months of his PhD in Mathematics, finding a passion for AI and machine learning.
- Transitioned from a risk analytics consultant to a deep learning data scientist, leveraging Kaggle experiences for career advancement.
- Impact on Career
- Kaggle competitions provided a hands-on experience that complements professional work, allowing for growth in skills and understanding of machine learning applications.
Comparison of Kaggle and Real-World Data Science
- Similarities
- Both involve problem-solving within a time frame, collaboration, and project management.
- Skill transfer from Kaggle competitions to real-world applications is seamless, especially in prototyping and project structuring.
- Differences
- Kaggle competitions provide predefined data and metrics, while real-world projects often involve data acquisition challenges and longer discussions to define success metrics.
GPU Acceleration and Tools
- Importance of GPU in Competitions
- GPU acceleration is crucial for efficient model training and experimentation.
- Tools like NVIDIA RAPIDS and DALI help speed up data processing tasks, enhancing performance in competitions.
Recommendations for Aspiring Data Scientists
- Start Competing
- Join any ongoing Kaggle competition that interests you, regardless of expertise.
- Engage with the community, learn from discussions and shared notebooks.
- Simplify Initial Efforts
- Begin with simple models and small subsets of data to establish a foundation for future enhancements.
Future Outlook
- Christof expressed excitement about advancements in AI, particularly in tools that assist coding and data handling. He is also curious about the long-term evolution of AI capabilities over the next decade.
Conclusion The episode offers valuable insights into how Kaggle can serve as both a learning platform and a career springboard for data scientists. Christof’s journey underscores the importance of community engagement and continuous learning in the ever-evolving field of AI.
---
Show Notes & References
- [Christof Henkel's Kaggle Profile](https://www.kaggle.com/christofhenkel)
- [NVIDIA RAPIDS](https://rapids.ai)
- [NVIDIA DALI](https://developer.nvidia.com/dali)
Sponsors
- [Fastly](https://fastly.com/?utm_source=changelog) - Bandwidth partner for fast and secure digital experiences.
- [Fly.io](https://fly.io/changelog) - Platform for deploying apps and databases close to users with minimal operations.
Join the Discussion Engage with the Practical AI community on [Zulip chat](https://changelog.zulipchat.com/#narrow/stream/456003-practicalai).
Subscribe Don't forget to subscribe to Practical AI for more episodes on the latest in AI tools and practices.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:06Welcome to Practical AI. If you work in artificial intelligence, aspire to, or are curious how AI-related technologies are changing the world, this is the show for you. Thank you to our partners at Fastly for shipping all of our pods super fast to wherever you listen. Check them out at Fastly.com. And to our friends at Fly, deploy your app servers and database close to your users. No ops required. Learn more at fly.io.
0:42Welcome to another episode of Practical AI. This is Daniel Whitenack. I'm a data scientist with SIL International, and I'm joined as always by my co-host, Chris Benson, who's a tech strategist at Lockheed Martin. How are you doing, Chris? Doing well, Daniel. How are you today? I'm doing great. Chris, have you ever been called a grandmaster in anything? No, but I really wish I had because it's a freaking cool name, man. Our title. Weren't you like a street fighter or something? You were like a black belt or something? Oh, don't go that. Something like that 30 years ago. But yeah, once when I was a kid.
1:22But you know what? I was never a grandmaster at anything. I was just trying not to get pummeled. Yeah, I was just trying not to hit the mat and that's it. Okay. Well, today we have with us an actual Grandmaster, a Kaggle Grandmaster, Christoph Hinkle, who's a senior deep learning data scientist at NVIDIA and a Kaggle Grandmaster multiple times. Triple Grandmaster, by the way. Yeah, in multiple of the different categories. So welcome, Christoph. It's great to have you here. Welcome, Daniel. Welcome, Chris. Very happy to be here. Awesome. Yeah. Well, for those that aren't familiar with this concept of Kaggle Grandmaster, could you kind of give us the briefing on what exactly that means?
2:06And in the context of also Kaggle, what generally I think a lot of people are familiar with that. But just in case, what is Kaggle and what does it mean to be a Kaggle Grandmaster? Yeah. So what is Kaggle? Kaggle, I would say, is like a platform for machine learning in general. It started off as a platform for hosting machine learning competitions. That's how it became popular. but in like the recent years it also expanded for like being a platform for discussions being a platform for sharing notebooks they're hosting millions of data sets so they're trying to become really like the go-to community for every topic around data science and it's free to register for everyone and they also provide some free resources where you can run code and try different stuff and competitions.
3:02And on this platform, they introduce different tiers in order to gamify a little bit. So to incentivize users to post content or to participate. So there are four different areas in which you can reach like different levels. So there are like competitions, which is like the most famous one. But there's also notebooks where you just progress by sharing notebooks with others. and the progression is based on upvotes on your notebooks. Then there are discussions which work in the same format. So you post an answer to a question or you post an interesting topic. You can also post just memes and generate upvotes in this way.
3:48And then there's datasets. So you can also post an interesting dataset or a dataset you think might be helpful for others. And then people can upvote your dataset and by this you progress. and you basically progress by earning medals. They're like bronze, silver and gold medals in each of the four areas. And then with these medals, you can reach like different tiers. So you start with as a novice, I think, then your contributor expert, then at some point your master and like the very last stage is a grandmaster. And to put that into perspective, so from the 10 million users that are registered on Kaggle, there are 280 competition grandmasters.
4:33So it's really like the elite of the elite, the top-notch people in their area, I would say. So I have to ask, because we were talking about it, which of the three categories are you a grandmaster in? And what's the fourth one that you're not? And of course, I'm going to ask you, when are you going to become a grandmaster in the fourth one? Well, I'm a grandmaster in competitions, and that's the most difficult one. Indeed. And then I'm a grandmaster in notebooks because I shared some high-value notebooks. And then I'm also a grandmaster in discussions because I like to discuss stuff. That's also why I'm here.
5:08But I'm not so fond of curating data sets and uploading data sets. I can't blame you. That's why I'm only a beginner at the data set. That would be the one I would choose first. See, that's Daniel. Daniel loves to do data grunging and stuff. It's sick. That's terrible. but I so I understand I give you I give you a pass on not being a grandmaster in the fourth one there what got you into Kaggle in the first place and what was the journey like towards where you're at now some people might just be jumping in on Kaggle and like trying things and they have like a vision of how far this could go but what was the journey actually like for you I think it's quite interesting because my journey began right in the last months of my PhD So I did a PhD in mathematics.
5:56And in the last few months, so after I sent out everything and I just was waiting for my defense, there was suddenly some free time and also free weekends I wasn't used to during the PhD. And I was also always curious about the AI topic. So back then, it was like five, six years ago. It was not so hyped as now, but it was like a niche area, what are neural networks and so on. So I was just curious about that, watched some YouTube videos, started a Coursera course on like what are neural networks and so on. And due to that, I quite quickly found out about Kaggle and then just started with my first competition right away.
6:41And since then, I'm booked in the system. And how long has that been? Six years now, I think. And during those six years, also my professional life progressed more and more towards machine learning and deep learning and data science. So six years ago, when I joined Kaggle, I was working as a risk analytics consultant. So I had nothing to do with machine learning. I had nothing to do with data science. I programmed a bit on risk models. So I had some background in like R programming or MATLAB, but I never used Python before. And then due to Kaggle, also my professional career shifted towards machine learning and deep learning.
7:20Until right now, I'm working as a deep learning data scientist at NVIDIA, which is like one of the top-notch companies in this area. Yeah, that's like the gold standard of jobs in the AI world right there. So do you feel like the experiences on Kaggle and your success there, in what ways did that kind of contribute to your own sort of career advancement? And also like your understanding of what you wanted to do as your career advanced? Yeah, it really had like a lot of impact. So step by step, I moved into the position I'm right now. So when I started, I was doing kegel like before and after work a bit, not too much, like half an hour after work, half an hour before work and on weekends.
8:10And then I made some, and of course I did horribly on my first competitions because I had no clue of anything. But the nice thing is that you really progress step by step. So in the first competition, you do horribly. In the next one, you do badly, but not horribly. and then you progress more and more until you become better and better. I quite quickly realized that a lot more fun in like machine learning and deep learning than on risk consultant just because you can be more creative, I would say. I moved within the consultancy company. I was lucky that they also had like a data science team. So I moved to the data science team there and I had my first synergy effect between Kaggle competitions and what I learned there and what I was using in projects.
8:58So I could use my skills in the project and I could also use skills I gained in the projects in Kaggle competitions. But that was five, six years ago. There wasn't much deep learning in the industry, especially in the insurance industry, where the focus was in my consultants company. So I was not challenged enough. but I wanted to do more and more in this field and also my skill set grew more and more so I decided to quit this job and found my own deep learning consultancy just to have even more synergy between projects and between Kaggle. Tell us a little bit about what that was like in those days because as we've grown up with deep learning over the last few years I would guess that at least in the beginning it was a little bit challenging to land, you know, engagements maybe?
9:53Or did you have them from the start? Because I know for me, early in that phase, about the time Daniel and I started the podcast, as, you know, people were like, deep what? Huh? So did you have any challenges in those early days that have obviously evaporated as the world has taken this on? Certainly. Not only in terms of projects. So people, especially the decision makers, I would say they are really cautious about the let's say possibilities you can do with deep learning especially like five six years ago there weren't any resources around so I talked with customers about what amazing things you can do with like deep learning and then they didn't have a single gpu they had access to so that's like really like two worlds clashing against each other.
10:44So there were a lot of interesting and challenging problems around that. But as soon as they basically gave me a chance and I could do some prototype and I can really show what you can do, then it was easy to convince them. But to get to this point, especially as like a young startup, a young consultancy startup, that was quite difficult. So I definitely want to get into many things later on, but I'm also thinking about these people out there that are maybe you know inspired by your journey and wanting to get involved in Kaggle and other things wondering if you can like share a little bit about because while you and Chris were talking about perceptions around deep learning that have shifted also during that time like the tooling around deep learning has shifted and like the accessibility of maybe like thinking about four years ago if I was to train a deep learning model for a Kaggle competition versus like being able to do that now how have you seen that shift over that over that time period in terms of this sort of ability for people to I guess people use the word democratize or or whatever the the ability for people to hop in and do something advanced like that very quickly there are like two aspects I would say one is like software wise and framework wise then there have been a lot of progress there so when i started it was still like tensorflow zero point something which was working but it's really like no level programming so there was nothing like an rnn layer or a transformer layer or so you need to code everything from scratch but it also helps a lot for understanding the things so i think nowadays people don't really understand the granular aspects of deep learning because you just do something like model.fit and you don't have any clue what's happening behind the curtain.
12:39So certainly it's easier nowadays to train a model just by this higher frameworks. Just calling by name, there's not only stuff like Keras, High Dodge Lightning, there's like a lot of different frameworks you can use which are really high level and accessible for beginners. And there's also a lot of training material for this frameworks. So a lot of tutorials. So it's really easy to train a simple model for a simple task. But also in terms of resources, I think they are more beginner friendly because on Kaggle, for example, five years ago, they didn't give you any resources. There was no Google Colab.
13:20So you basically have to have your own GPU at home. You need to build your own desktop machine or something or you spend your own money on cloud resources. But now for beginners, you can get access to Colab, which gives you a free notebook to experiment. You get some free resources, some Kaggle. There's a lot of student credits and student programs. So it's really easy to start your data science journey, I would say. And there's also a lot of more material online where you can really teach yourself, I would say.
14:12Hello friends, this is Jared here to tell you about Changelog++. Over the years, many of our most diehard listeners have asked us for ways they can support our work here at Changelog. We didn't have an answer for them for a long time, but finally, we created Changelog++, a membership you can join to directly support our work. As a thank you, we save you some time with an ad-free feed, sprinkle in bonuses like extended episodes, and give you first access to the new stuff we dream up. Learn all about it at changelog.com slash plus plus. You'll also find the link in your chapter data and show notes.
14:53Once again, that's changelog.com slash plus plus. Check it out. We'd love to have you with us.
15:08So Christoph, as you were kind of leading in talking about your entry into the world of deep learning and your career shift to accommodate that, and you're talking about kind of learning from Kaggle competitions and engaging in that. and then it was increasingly applicable in your professional life. Can you talk a little bit about how that happens? Like when you're thinking about a Kaggle competition and you're now working in a job in this field, how do the two relate? How are Kaggle competitions relevant to solving real business problems in a real job and getting that synergy? What is that like?
15:45What is the connection between the two like? I would say there are a lot of synergy aspects. So doing a Kaggle competition is really very similar to doing a project at work, which is about performing a first prototype. So in a Kaggle competition, you get like a problem which you are not familiar with, often from a different domain. Can be from biology, can be from astrophysics, can be from chemistry, can be Bengali language, sign language, just so much different problems that you have no clue about when you start. And then you have like three months time to find like the best possible solution and also compete with other data scientists.
16:31So like this prototype project character is very similar. So you have like this three months time window. Then you have a collaborative part. So in Kaggle, you can also form teams. So you can participate in competitions in a team, which is very similar to working in a team in your job with all the ups and downs, I would say. Working in a team under pressure often. So Kaggle competitions can create quite some pressure, more pressure than you might feel in your day-to-day job. So you also get used to working efficiently with others. So in terms of coding, in terms of reading their code, in terms of structuring the project.
17:16So really like all aspects of project management are also important. And also things like optimizing runtime and optimizing code structure. You wouldn't think that it's quite important, but I think it's quite important also for Kaggle competitions. Because recently they run the competitions on a restricted hardware. So you just submit your code and they will run your code on their infrastructure using their Kaggle notebooks. So you need to have your code in a way that it's kind of production style. That's also what you would do in a project. So you would develop ideas and so on and so forth. But at the end, you want to productize your code and you need to think about all these ML ops problems as well.
18:04And you also train those skills in Kaggle competitions. So I really like parallels between the two worlds. That said, I must say that two things are really different between Kaggle and the real world project. First thing is data acquisition. That's like a very big topic like the real world. It's no topic, well, no topic, but a very minor topic on the Kaggle competition. You already add your training data. Of course, you sometimes can expand your training data by looking for more data online. But in general, you already have like a fixed training set you can work with. Whereas in the outside world or in the real world, that could be like the main problem just to require some data.
18:50And the second thing is definition of the metric. So in Kaggle, people are like evaluated based on some metric. and this metric is predefined before the competition starts. Whereas in the real world, that can be a discussion which takes for ages between like data scientists, the business, and just creating a metric that is representative of the business problem can take a lot of time and you don't have these issues and discussions. I'm curious, as you were describing that, I have an idea that came to mind. And so recognizing the limitation of you already have data provided and recognizing the fact that the metric is well defined on a Kaggle team.
19:34And both of those are kind of optimal situations compared to the business world. But from the perspective of an organization out in the world, any organization that is keenly interested in data science and stuff, would forming Kaggle teams or participating in Kaggle teams be a good recruitment tool? Because if you can find people that are performing well on teams in that capacity, it doesn't check every box for what the business world is doing. But it kind of gives you a sense maybe of this might be someone who could fit in with us. We're going to throw the messiness of data sets and the messiness of metrics on top of that.
20:13But what do you think of that idea? Is that something that people might be thinking about in terms of trying to build data science teams for their organizations? Certainly. I think that that would be a great idea if people do this. And some companies already use Kaggle as a hiring tool. So in order to run a competition, those competitions are sponsored by someone. And there are sometimes companies who sponsors a competition, but also tell the participants that they are hiring. And that if you are finishing like in the top spot, you can apply for a position there. So getting a position is kind of part of the winning prize sometimes.
20:53So they already see that Kaggle is very good for finding good candidates. But as you said, you could also, and Kaggle nowadays even offers the concept of a community competition, where you host the competition by yourself without any Kaggle interference. And you could do this as kind of an assessment center for filtering potential hires or see how they interact on a problem or see how they work together. There are also, so normally Kaggle competitions are like three months or so, but there are some formats, for example, Kaggle Days, which is like conference type of thing. They host like this conference specific competitions and they just go like one afternoon.
21:38and people get like a simple data set and they have one afternoon to get like a good solution. And they could definitely see how this would benefit an assessment center, for example, because they really see like the whole range of skills people can bring to your company. I have to ask, of the sort of competitions and the notebooks that you've contributed to Kaggle, maybe the discussions too, what are some highlights for you like of all the things that you've done what are some highlights of the the things maybe either you're most proud of or or that you would like to highlight and what I'm most proud of certainly are the google landmark competitions so there's a competition which was hosted three times yearly by google and this is about classifying popular landmarks So you have a data set of 5 million images.
22:36So it's really like large scale. And in this 5 million images, you have 80 ,000 classes. So 80 ,000 different landmarks. And you need to classify between those landmarks. And the difficulty especially there is that for some landmarks, you only have one or two images, which makes it quite complex to classify. and another complexity is because some landmarks are quite like looking differently from different angles. You can think of a museum, for example. People take a picture outside of the museum, people take pictures within the museum, and you still would classify it as the same landmark, for example.
23:18So the competition is quite tricky, and I was able to win it three times, and two times of that without a team so just solo and that's something that's even harder in doing kaggle competitions so without participating within a team but soloing that brings a lot of additional let's say mental stress because you're you're not you don't have a team you can talk about your problem with you're just like isolated working on a problem for three month with like high pressure and so on and so forth so that brings another level of like mental component to the game so i was quite proud that i could win two of those or win three competitions and two of those without any team so i i'd like to follow up on that what is you're talking to people out there that might be you know either already participating in kaggle you know not at the level that you're at or thinking about jumping in, what are some of the attributes that you, and I want you to take a moment and harp a little bit on yourself.
24:23I'm asking you to, and say, what are you bringing to the competition? Do you think that really has given you an edge in getting to that grandmaster level and being so competitive at that level? Do you have anything that you can offer people that are kind of maybe a little bit intimidated by it or trying to think, how can I level up a little bit? What would you say? I mean, I definitely have some analytical thinking just from my study of mathematics because the whole study is there to basically learn how to think efficiently, how to solve problems efficiently. So that definitely helps. And coming from natural sciences in the broader sense, also a sense of solid experimentation is very important.
25:12So really having like a clean workbench, so to say, logging your experiments, following up on ideas and so on. So this really like thinking like a researcher in natural sciences and following your experiments like in a clean and reproducible way, that's also quite important. But I think what really pushed me to the top level is the curiosity of different domains. So even like top people, they tend to, let's say in quotation marks, lean back and do what they're good at and not expand and learn further. But I would say one more edge I get is that I really try a lot of different ideas. And at the end, in different areas, I try to explore like very different competitions, very different domains.
26:04and at the end every now and then I can leverage from something that you would think has nothing to do with the other but you still can leverage some ideas and apply some concepts so for example you can transfer knowledge from audio classification to biology or to astrophysics or from NLP to computer vision and vice versa so there's a lot of synergy people wouldn't think about And therefore, it's quite helpful to explore as different domains as possible. You alluded to this a little bit in what you were saying about used to with Kaggle competitions, maybe you had to build your own machine with a GPU in it to sort of operate in that.
26:51Now there's good resources with GPUs, but I'm wondering from your perspective, both as a competitor and a grandmaster, but also as a really senior data scientist at NVIDIA, how do you view kind of GPU acceleration as kind of important and playing a role in Kaggle competitions? Probably most people think about it in terms of training a model. But how do you think about that more holistically in terms of the accelerated process that's key to performing well in competitions? So certainly, GPO-based programming or calculation is like the bread and butter of training any model nowadays. but also especially nvidia they are looking more and more into moving other parts of your data pipeline onto the gpu just to make it faster and especially for kegel competitions the speed of which you can run your stuff and try ideas is very important so when a lot of people like on the top level compete against each other one of the edges you get is when you can do more experiments than the others which are just bound by of course your ideas but most of the time i'm not running out of ideas but i'm running out of time in the competition and so as long as i can run more experiments than other people can do because i have a more efficient pipeline or i can run more parts of my pipeline efficiently using gpus that gives me an edge and some examples of this are like data pre-processing or let's start even one step ahead.
28:37The first step is just data loading. Just loading your data frame for doing anything can be GPU accelerated and that is just like 100x faster. So every time you're working on the problem, you get a 100x speed up just in the step of loading your data. And that's what Rapids, for example, is all about. So Rapids is like an NVIDIA tool stack which is all about accelerating those parts which are not like training the model but are like what is normally handled with pandas for example so they have a part which is called qdf which is basically pandas on gpu they have something which is qml which is basically sklearn on gpu so things like clustering all this stuff you can do on gpu nowadays other examples are for example nvidia dali that's a tool especially for image processing but I also support audio and video but an example there would be decoding of JPEGs so people wouldn't think about that but something like having a JPEG on your disk and just loading the JPEG involves some decoding step which basically decodes the JPEG format and this can already be done using GPUs and can be accelerated by GPUs and also gives you a significant speed up during your training, during anything which uses the images.
30:04So there's a lot of different steps in your pipeline that you can accelerate. And that's what all accelerated data science is about. So NVIDIA tries to move the complete pipeline from loading the data to saving conclusions, results, all end-to-end on GPUs. Yeah, that's really interesting. And I'm guessing that some of the things that you're talking about, like loading images or loading data frames or manipulating data frames, maybe doing certain operations, doing clustering. I don't know that this is the case, but I would guess like those things pretty consistently show up across competitions too, or in the real world, you could think about them as showing up across many different business problems.
30:50So like you were talking about your pipeline of processing, which I think is a really, wondering if you can dig into that a little bit, not a specific pipeline, but how you think about like solving a problem, because most people might come to a Kaggle competition or a real world problem and say, OK, here's my data. I have my main step is this sort of like training of the model and maybe evaluation, like how good is my model? Retrain it. How good is it? Retrain it. How do you think about the data sort of pipeline around, you know, like you're talking about running experiments? What is that sort of like data pipeline look like in your mind?
31:27And what are some of those reusable components or things you find yourself doing over and over again that are accelerated, that you found accelerated ways to do those things using GPU tooling like Rapids or this DALI? It really depends from project to project, I would say, where it's applicable or not. So I would say that Rapids, for example, is even more applicable to the real world, because there you might have way larger data frames, for example. So if you're like a bigger company, you have like user data or you have client data or whatever, because the Kaggle competitions often are like packed into littler problems that people can work on and not like this company size, large scale data sets with like millions of users or thousands of users.
32:17And things like Rapids especially shine in like this large scale data sets. For me, my pipeline is, I would say, modular. And that develops through the years coming from the competitions. So of course, I try to reuse as much as possible just to be efficient. So I have a really modular setup where I have one part which is just the model training, one part which is about the storage of my data, one part which is about logging the experiments and tracking results and visualizing results one part which is about the framework setup so to say so i use docker with a specific pytorch image to have like always the same environment and also can replicate my experiments and also can use the exact same environment of different machines so in the cloud or locally that's all things I learned during the years.
Read the full transcript
33:17So it's a little bit complicated to explain the whole pipeline now on the podcast. I actually gave like a one hour presentation two weeks ago, just about this topic. So it's pretty difficult to cut that into a few sentences. It's hard without a diagram for sure. But it's super interesting to me, like the things you're talking about that you've made modular, I think are things people operating in a real world data science environment eventually need to make into sort of like components that work within their team, right? Like, you know, my team, like we love using, for example, Streamlit to do like some data manipulation, visualization, interactive stuff on the other end.
34:02And we have a lot of those, we reuse a lot of those components and you know we have like certain models that we multilingual models that we train over and over so we've got you know modules around that and then like pre-processing and other things so I think these are it's interesting how much what you're talking about overlaps with I think the efficiencies you gain over time as a data science team operates together and they learn how to make their own processes more efficient. So I think that that's really interesting. So I have played around with Rapids a few times and it is really cool. And I'm just looking at the latest stats here on the Rapids website and it's talking about performance on 300 million rows by two column data frame with like the highest speed up being for like group by operations like 80 times faster than not using rapid so like i don't know you know how long you know that saves you but also like you're talking about if you are doing experiments over and over and you want to rapidly do experiments even if that saves you let's say it's something smallish like in minutes right a couple minutes like you're able to do things much faster and automate thing like your automation goes faster you can learn things much faster and reduce that cycle time although I'm I'm also assuming for many people for their data it might be more than a more than a minutes long speed up potentially on some of those operations so yeah I don't know um when you're when you're helping people and and you mentioned the discussion groups and the notebooks that you've worked on on on Kaggle is this something where you've seen like light bulbs come on for people when they like saying like oh i'm trying this group by operation or something on this data and it's taking me like 15 minutes every every time i run through this is that something you've been able to bring in those discussions and notebooks and such on kaggle yeah certainly so like loading data frames is a good example so 80 times sounds not that much i but it's like one minute or two hours.
36:15That's like the scale you're talking about. Like loading your data frame in two hours or loading it in one minute. That's like an ATX speed up difference. And especially in Kaggle, those discussions get a lot of traction because on your inference, you actually have like a time limit of like nine hours. So people try to get as much stuff into their submissions as possible. so loading data frames, manipulating data frames, loading images all the stuff, if you can speed it up, the people will be very very gratefully adept whatever you give them to speed up their stuff and that's only the inference side so that's even more true for training because as you said, my day to day is like doing a lot of experiments and those speed ups accumulate.
37:07So the very first thing I ever do in a competition, like the first two weeks or so, I just optimize my workflow. So I optimize all the runtime, optimize how I load my things, accelerate all the pre-processing, post-processing, whatever I have in my pipeline. So I can then leverage the remaining time from like the most perfect setup or the most perfect code, because then I can just run more and more experiments. So I'm curious because as you have been talking about optimizing and being able to do all of these iterations on your experiments, there are people out there, including myself, that are thinking whether they are wanting to jump into a Kaggle competition, they're psyched up because they've been listening to how you've kind of mastered this process, or they're working for a company and they are trying to get their own systems better and better.
38:05And, you know, early teams really struggle with that. And so either way, with you talking about what you've done and Daniel was jumping in and talking about the work they done, there are people that want to be there with you, you know, that want to at least get on that path. Do you have some concrete recommendations on somebody who's at the beginning of that? And they're like, okay, I'm doing data science, but my God, it's taking me a long time to get through each iteration. And I'm listening to this grandmaster just cranking out productivity so fast. What are a couple of specific things that you would say, go do this and that and that, recognizing that they'll find their own path forward and they'll make their own adjustments.
38:47But how do they get on that path to begin with? The first thing, and I told this several times, is just to start your very first Kegel competition. So you go to Kaggle.com, you look through the ongoing competitions, which is like 15 to 20 ongoing competitions. It just shows any topic you find interesting. You don't need to be an expert in this topic. You don't need to even know about the domain or something. But just starting is like the first step. And as soon as you start, just by the sheer amount of knowledge which is shared within the forums and the notebooks, You will see that you learn very, very efficiently how to improve your code, how to improve your skill set.
39:33And you get like immediate feedback on the leaderboard, for example, or on discussions. If you add like a comment and it doesn't make sense, then people will tell you. If it doesn't make sense, people will also tell you a thank you. So the leaderboard is very like an objective way and seeing your performance and seeing your progression. so that's the very first advice I would give someone try to find an interesting competition and just start there's basically nothing to lose you just can gain knowledge as said you will perform poorly on your very first competition no matter where you come from but just starting is like the first step and as you start I think the best advice is that you start simple as simple as possible and just try to progress from that you start with a very simple model with a subset of the data or with like images which are down sampled to a low resolution just to find like an efficient pipeline and to work on your code because all this is like an investment for the future and all this gives you an easier setup to work on and to improve on yeah really good advice i think that part you talked about about like spending a couple weeks optimizing the sort of inputs, outputs, and those portions of your pipeline so that you can really put a lot of your focus on fast iterations on the model or that middle bit.
41:03I think that's really, really good advice. This has been a really fascinating competition. I have a long way to go to be a grandmaster, that's for sure. But as we wrap up here, this discussion about accelerated to data science and the Kaggle competitions. What are you excited about sort of looking to the future? You mentioned that you're curious about all of these sorts of different domains. You've worked on a lot of different problems. What really excites you right now as you look towards the future in terms of things that you want to try or just in general things that you're excited about in terms of the tooling or the community around what you're involved with?
41:48I would say in the short term, I'm definitely excited about or interested in how AI will support my work. So something like GitHub Copilot or other natural language models, which help me code. I haven't tried them much, but I think that in the near future or the short term, those tools will support our everyday life in some way. but I'm even more excited in like the long-term prospect like what will happen in 10 years and 20 years and that's really excited because if you think back like 10 or 20 years in terms of AI and what systems could do and where we are right now and you extrapolate that into the future that will be very exciting and amazing what will happen then.
42:40Yeah, I think that's a great way to wrap things up. Thank you so much for joining us, Christoph. Really looking forward to following your progression and the things that you work on in the future and the great things that continue to come out of NVIDIA. So thank you for your work and thank you for taking time to join us. Thank you for having me.
43:06Thank you for listening to Practical AI. Your next step is to subscribe now, if you haven't already. and if you're a long-time listener of the show, help us reach more people by sharing practical AI with your friends and colleagues. Thanks once again to Fastly and Fly for partnering with us to bring you all Change Talk podcasts. Check out what they're up to at fastly.com and fly.io. And to our Beat Freakin' residents, Breakmaster Cylinder, for continuously cranking out the best beats in the biz. That's all for now. We'll talk to you again next time.
43:49Game on!
From the publisher
Daniel and Chris explore the intersection of Kaggle and real-world data science in this illuminating conversation with Christof Henkel, Senior Deep Learning Data Scientist at NVIDIA and Kaggle Grandmaster. Christof offers a very lucid explanation into how participation in Kaggle can positively impact a data scientist’s skill and career aspirations. He also shared some of his insights and approach to maximizing AI productivity uses GPU-accelerated tools like RAPIDS and DALI.
Changelog++ members save 2 minutes on this episode because they made the ads disappear. Join today!
Sponsors:
- Fastly – Our bandwidth partner. Fastly powers fast, secure, and scalable digital experiences. Move beyond your content delivery network to their powerful edge cloud platform. Learn more at fastly.com
- Fly.io – The home of Changelog.com — Deploy your apps and databases close to your users. In minutes you can run your Ruby, Go, Node, Deno, Python, or Elixir app (and databases!) all over the world. No ops required. Learn more at fly.io/changelog and check out the speedrun in their docs.
- Changelog++ – You love our content and you want to take it to the next level by showing your support. We’ll take you closer to the metal with extended episodes, make the ads disappear, and increment your audio quality with higher bitrate mp3s. Let’s do this!
Featuring:
- Christof Henkel – GitHub, LinkedIn, X
- Chris Benson – Website, GitHub, LinkedIn, X
- Daniel Whitenack – Website, GitHub, X
Show Notes:
- Christof Henkel | Kaggle
- NVIDIA Kaggle Grandmasters
- Kaggle
- NVIDIA RAPIDS
- NVIDIA Data Loading Library (DALI)
Something missing or broken? PRs welcome!




