In short
Practical AI Podcast Episode Summary
Episode Title
Capabilities of LLMs 🤯 Episode Description: This episode discusses the remarkable advancements in Large Language Model (LLM) capabilities with guest Rajiv Shah. Topics include in-context learning, reasoning, LLM options, related tooling, and the burgeoning AI community on platforms like TikTok.
Key Participants
- Rajiv Shah: Machine Learning Engineer at Hugging Face
- Chris Benson: Co-host
- Daniel Whitenack: Co-host
---
Episode Highlights
Introduction
- The hosts welcome Rajiv back to discuss the rapid advancements in AI, particularly focusing on LLMs.
- Raj discusses his experience with TikTok in educating others about data science.
The TikTok AI Community
- Rajiv shares insights about how he started creating educational content on TikTok.
- He notes the difference in how younger generations learn, emphasizing video formats over traditional reading materials.
- Engagement with audiences is crucial, and creating content that is not overly clickbaity but informative is his strategy.
Current AI Landscape
- Raj highlights the explosion of AI advancements, particularly in the last few weeks.
- Emphasizes the importance of staying updated but reassures practitioners that they don’t need to follow every development closely.
In-Context Learning in LLMs
- Raj explains in-context learning, where LLMs can learn from given examples without retraining.
- Example: Instead of labeling movie reviews, one can provide examples to the model to classify new reviews.
- Discusses how this shifts traditional data science workflows, allowing less technical users to perform complex tasks.
Prompt Engineering
- The hosts delve into the emergence of prompt engineering as a new skill set in data science.
- Raj discusses how prompting allows users to carry out tasks that previously required separate machine learning models.
Emerging Tools and Technologies
- Raj mentions tools like LangChain, which help integrate multiple prompts and workflows using LLMs.
- The conversation shifts to the Hugging Face ecosystem, its role in open-source AI, and the benefits of community contributions.
Landscape of Large Language Models
- Raj discusses the categorization of LLMs based on their access (open-source vs proprietary), size, and training data.
- He notes the confusion surrounding the variety of models available and their implications for businesses.
Ethical Considerations
- Raj touches on the ethical challenges, such as hallucinations in LLM outputs and concerns over data privacy and copyright.
Education and AI
- The hosts engage in a discussion about integrating AI into education, focusing on how to effectively onboard students and professionals onto these technologies.
- Raj advocates for hands-on experience with AI tools to better understand their capabilities and limitations.
Future of AI
- Raj predicts that the current pace of AI development is unprecedented and encourages open-source contributions.
- He emphasizes that the democratization of AI tools will allow more people to leverage these technologies effectively.
---
Key Takeaways
- In-Context Learning: A significant advancement allowing LLMs to learn from examples without retraining.
- Prompt Engineering: A new essential skill for data scientists to interact with LLMs effectively.
- Emerging Tools: The Hugging Face ecosystem provides critical tools for training and deploying AI models efficiently.
- Ethical Concerns: Understanding the limitations and implications of AI outputs is vital as the technology becomes more widespread.
- Education: There’s a need for practical integration of AI tools in educational settings to prepare users for the future.
---
Resources and Links
- Hugging Face: [Hugging Face Website](https://huggingface.co)
- Online Transformers Course: [Hugging Face Course](https://huggingface.co/course/chapter1/1)
- LangChain Documentation: [LangChain](https://python.langchain.com/en/latest/)
- PEFT Library: Information on parameter-efficient fine-tuning techniques.
---
Conclusion The episode emphasizes the transformative potential of LLMs and the tools emerging from the Hugging Face community. With the ongoing rapid developments in AI, it’s essential for practitioners to stay informed and adopt these technologies responsibly and ethically.
Thank you for tuning in!
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:06Welcome to Practical AI. If you work in artificial intelligence, aspire to, or are curious how AI-related technologies are changing the world, this is the show for you. Thank you to our partners at Fastly for shipping all of our pods super fast to wherever you listen. Check them out at Fastly.com. And to our friends at Fly, deploy your app servers and database close to your users. No ops required. Learn more at fly.io.
0:42Welcome to another episode of Practical AI. This is Daniel Whitenack. I'm a data scientist with SIL International and joined as always by my co-host, Chris Benson. How's it going, Chris? Going very well. Spring is in the air. We're having a good time here. Lots of cool stuff in the AI world. Lots of new life breathed into interesting AI systems over the past days. And sometimes I hear about this in cool videos, which are way cooler than any videos that I produce from our friend Raj, who's with us today. Rajiv Shah, who's a machine learning engineer at Hugging Face. How are you doing, Raj? I'm doing great.
1:26Thanks for having me on. Yeah. So the last time you were on the show, we talked about data leakage. Have you leaked any data since the prior episode? I think any data scientist that's out there has leaked data on a regular basis like that, right? It's a hazard of the job. And, you know, one of the things I like to do is continually remind the new folks that that's likely to happen. So that they are in the process of and should remember. Yeah. Yeah. I did mention I've seen a lot of cool videos from you recently. And we were chatting even a bit, I think, on LinkedIn about there is data science, AI community on TikTok and other places.
2:09Tell us a little bit about that. I'm just I mean, that's a fun topic. I'm curious of what is the AI scene like on TikTok? So let me start by like how I got into it. About like a year ago, I was trying to get my son who's just starting college to do a real practical project around AI. Like he's taking computer science, but he doesn't know what GitHub is. So I'm like, can we build a Discord bot, for example, something that appeals to him? And so I was like, let's give us 24 hours. We'll do this over the weekend. We'll both go our separate ways. And then we'll come back and kind of see what we've done and see if we could share what we've learned from each other.
2:44And so I go out and I go get a blog tutorial and work my way through it and kind of get something working. And then I go to him the next day and he's like, yeah, I kind of got stuck. I was like, well, let me see if I can help you through it. Like show me the steps and what you were accomplished in doing it. He pops open a YouTube video and that's what he used to follow it. And for somebody like me who self-taught their way into data science that was largely focused on kind of reading and written material, it kind of really blew my mind that somebody would learn how to code through a video. But it really just opened my eyes to that because already at that time, like with my daughter, I shared videos on like food and politics and music, but it just really came to me like how this is just becoming an emerging part of education and how people learn kind of as we move on here.
3:31Yeah. And have you seen like engagement with your videos on? So like, I remember the one recently I saw was like the, what is it? Segment anything or everything. I forget which is anything and everything, whatever that one is. I saw your video on that one, which was cool, because it's also very engaging. You've got like this skit element to it, but there's real information content in there, right? In an engaging way. How do you see people respond to these? So I've had great feedback, and I try to keep mine very focused on data science. I try not to be too clickbaity. I try to be like, you know, if I was on a data science team, would I recommend somebody to watch the video like that?
4:09But the video style also lets you do different things. So I think when I first started videos, I did, you know, I used to be a professor. I did like the traditional, let me just lecture you on this topic for 30 seconds. But like, I think as you mentioned over time, like there's more creative ways of doing it. And one of the things TikTok allows you to do is often tell it in a story or skit format where you could have the voices of multiple people. If you're sitting at home on your phone, that's a much more interesting way to like get a nuanced conversation rather than reading kind of some blog post that has, here's four different points on this.
4:42So, you know, I think there's a lot of potential for kind of teaching nuanced with something like TikTok. Yeah, that's cool. And there's no shortage of things right now to talk about. It's like, you probably are more constrained on your ability to pump out these videos than like the AI things that are coming out. We've all had our minds blown recently, especially the capabilities of large language models, but there's also, of course, other things and computer vision and other things. For you, what have been those mind-blowing moments or what has been on your mind over the past, I can't even say like the past year, like the past two weeks.
5:19I don't know. So I think we just have to look back and reflect that we're really in a great place of a huge amount of innovation in a short amount of time. Like this is one of those peak times in AI that it won't be like this a year from now, right? It wasn't like this two years or three years from now where literally every week there's new developments. It's a fabulous time if you're an AI junkie and you like to kind of check out and see the newest tools and see that incremental advance like that. There's no better time. It's not going to last for long. So kind of enjoy it. I also kind of also push back on this for lots of practicing data scientists that are very practical that, you know, you don't need to watch this stuff every day or every week like that.
5:59Like many of these things are exciting. But if you're day in and day out, you're an enterprise data science, you're inside doing churn analysis or some marketing analysis, many of these developments are going to take a while before they filter you to you. You'll have plenty of time to get up to speed. They're not going to change the face of every data scientist in the next two months like that. So I'm curious, it's a follow-up to both of these last two questions combined. You're going into different mediums now for teaching. You're hitting short video, longer video, different things. We have all of this happening so fast.
6:32How are you thinking about reaching different audiences in data science? It's kind of funny. Once upon a time, it was just data science, but now we have different audiences, different age groups, different purposes. How are you making those different connections? I was just talking to someone today that the co-op was like, students coming into college now can't type. Typing isn't a thing anymore, right? Because of the way they've grown up with devices, right? Like they can poke and like touch screen, but that's got to influence like, if we're not adapting to that, then we're not staying up. Right, Chris?
7:07You just made me feel really old. And I think one thing that's happened is like data science came out of statistics. And for a long time, right, the path to learn that was you went to college, you sat in a classroom, right? You had statistics book to do that. But I think this is the transformative part about AI and data science, where now it's touching so many people. And especially you see this with these large language models, where if you're a teenager, you have a GPU, all of a sudden now you can kind of download and follow a script and get something running on your local machine where you can interact with this AI, right?
7:42Which a couple of years ago would have been unheard of for somebody to have such wide access to that. So I think the hard part about communicating to so many audiences is also a great part that we have such a large community that's engaged and interested and wants to use these tools. I'm going to bring a couple things here on the fly for you, Raj, because you are so good at explaining these things. So I'm at a conference right now. So I walked from a talk over back to here. And yeah, one of the things that they were talking about was in context learning with large language models. Could you kind of help us?
8:17So we've talked a lot about, on the show, about prompting large language models, this sort of thing, but I don't know that we've specifically kind of talked through this, like, in context learning. Like, what does that exactly mean, and what should people take away from it, maybe? So if we look at the development of these language models, a couple years ago, if you look at, there was blog posts by Carpathion kind of working with LSTMs and how we could get these models to generate text for us. And this is where we have kind of the statistical probabilities of being able to put together text and it knows like the cat ate the dog or that there's some probabilities and we could put a sentence together.
8:53And a couple of years ago, right, these were fantastic things at making like really weird stories. Yeah. Right. And that's all they were good for. Like when we look at like kind of the GPTools like that. Now, what's happened is as we've kind of worked with these large language models and they've gotten bigger where we've incorporated more data, we've trained them longer, machine learning engineers have noticed a new kind of what they call an emergent behavior that's come about from these models that isn't there at the smaller size of the models. But when these models get really big, they allow this new capability of this in-context learning.
9:29And what in-context learning allows you to do is you can give the model a few examples of a type of question, and the model will then continue to answer in that question. So an easy example of this is sentiment. Imagine you had to have movies and you had to rate the sentiment. In the old days, if you wanted to do this, you would have to go out and label a bunch of movies, right? Let's go get 100 or 1 ,000 movies. We read the reviews. We label the sentiment. Is this a good review? Is this a bad review? Then we train our model to do that. That's traditional data science approach. What we can do with these larger language models is say, hey, here's three examples.
10:06Two of these are good movie reviews. One of these is a bad movie review. Now, I'm giving you a new movie review. Will you tell me what this movie review is? And the model will reach back with us with the answer. And the key here is we're not changing the weights of the model. We're not training the model in any way. Just by carefully asking it for some type of information, it knows and can kind of figure out, oh, you like it like this? Well, I will give you back an answer in that same kind of format style, same type of information. And so this for me is just mind blowing. And it also makes us rethink like a lot of the tasks we do in NLP and how many of these we're going to be able to use this paradigm to do it.
10:47So that was a long answer. I'll let you see how much of it you digested. That was a really good answer. So, you know, we all have this new skill that we've been developing, you know, around prompting, especially this past year. Prompt engineering is now a thing, which it wasn't very far back. It was, you'd go, what? What's that? So how does this all tie in? We have this new skill about prompting and learning how to prompt effectively to get this information. You're talking about this emergent quality of these large language models? How do those tie in? What does that imply for steps forward? And what should people be thinking about to make that productive forum and day-to-day use?
11:22Let's take this example to like something that you would do practically inside an enterprise where somebody might give you some type of document or chat transcript, which might be a little bit unstructured. And what you want to do is just categorize it. So now what we can do with using these prompting and these approaches is we can take that amount of information. I can ask the model, hey, will you structure it? Will you clean this? Will you take out the HTML format tax? It'll do that. And then I can ask it another prompt. Hey, can you summarize this? Like take this from a hundred line conversation just down to the essentials, 20 lines.
11:54You can write a prompt for that. Then you can ask it, hey, will you categorize this? I need to see, you know, should I send this to my claims department? Does it go to HR? Does it go to IT? We can write a prompt for that. And so now what you have developers using is tools like Langchain, where they can tie together several of these prompts and create workflows that in the prior to this, we'd have to use separate machine learning models to do each of those tasks. And I think this is really, for me, what the mind-blowing part of it is how it can change machine learning and really do a lot of this democratization that we've talked about for a long time, but do it through a natural language interface where somebody can just literally give it these tasks in a human language and then have them accomplished.
12:37For the data scientists out there, it's a little mind-blowing because I've been in this place where we've tried to teach people citizen data science. And we have classes on how to properly partition data and holdouts and loss metrics and all of this. But this approach dramatically kind of changes how the number of tasks people can do kind of without having to learn all those concepts. That's a great point. With the advent of ChatGPT and some of the others that are out, and everything coming out, it has exploded the audience that can productively use this technology. And do you see any limitations in that going forward, or do you think it's going to continue to grow?
13:15And I think this is, to Daniel's point earlier, this is the mind-blowing part, is I gave you the simple example. Now what you see people doing is taking this but combining this with other APIs and other services. So in that case of the movie reviews, maybe I want to get the weather forecast or I want to find out if the theater was open that day, something else. Well, now I can use that same type of natural language interface and connect to other APIs, other services, other information. And so this is where we see some of the most powerful applications of this with tools like HuggingGPT, which allow you to interconnect with lots of different hugging face models where I can ask it a question and give it a picture.
13:58And the model will automatically go out, figure out the appropriate hugging face models to use, run them, figure out the answer and bring that back to me. Or in this repo has been going crazy as the auto GPT one where we essentially take that. It's not just for models. we allow the large language model to do any task where we can say, hey, start up a business and raise some money for me. And then the model will go out, answer that, go see, hey, is there some other databases? Is there some other APIs that can use it? And we'll continue to iterate. It might cost you a lot on the tokens for GPT-4, but it'll continue to iterate and try and try and try and do it.
14:35And I think, you know, for me, if you asked me a year ago if this was possible, I would have said, no way. That's three or four years out. Like, I can kind of see how you're doing it. But to me, this is why it's such a special moment. We're living in it because I don't think any of us could have predicted we'd be here, you know, a year ago. Even in our conversation so far, like we've listed out like a bunch of models. So GPT is an auto GPT and hugging GPT and like Bloom, Flan, Flamingo, 7 billion, whatever. In terms of large language models and what's out there right now, one interesting thing is like open access or various patterns around that and hosting.
15:16Like how how do you think about like the landscape of large language models? Like what does that look like right now? What are the major categories that we could kind of have in our mind as clusters of these things? At this point, there's tens of kind of large language models. And yeah, there's a number of different ways we can kind of categorize your thinking about them. One of them is kind of the simplest, which ones are proprietary, which ones are open source. There's a spectrum when we talk about access to these. So there's some like, for example, OpenAI, where you don't have access to the model.
15:51You don't know what data it was trained on. You don't know the model architecture. You just send your data to them. They send back the predictions. And so I think that's one model there. And then all the way at the other extreme, today, for example, Databricks released the latest version of its DALI model, which was an open source model that was then instruction tuned on a data set that Databricks created themselves that they're making available kind of open source for commercial use itself there. So there's the whole spectrum there, but there's other spectrums here too because the models, for example, vary in size.
16:28Where you have, for example, something like Bloom that was developed by Hugging Face, which is one of the largest open source models at something like 170 billion parameters, to some of these much smaller models that are coming out, the LLAMA models and others that are maybe a billion parameters. And that size has implications in terms of how much reasoning ability, how much stuff is inside there. but inference. Is this something that your teenager is going to run on their own GPU, or is this something that's going to take a multi-GPU cluster to be able to effectively use? There's other dimensions like what data the models were trained on.
17:03For example, with the open source models, we know what data they were trained on. One piece of this, for example, that's come up is knowing how much code a model was trained on. Because one of the things that's often asked for is, hey, can we build a text to code type model where I want to do some type of autocomplete, some type of code generation type project? Well, if I start with a large language model that already understands code, it's a lot easier to fine tune it and make that capability. So like understanding the underlying characteristics of that data. Daniel, right? It's like an alphabet soup of different names and like literally every week they're popping up and And there's so many of these different characteristics because they also differ, for example, on the model itself and what the licensing is and the model weights, the data set that it was trained, the training code that it was done.
17:54We see this with kind of how Meta released the LLAMA model where they told everybody about it, but then they released the weights, but then they gated the weights. So only academic people were getting to them, but then the weights were essentially leaked and now they're all over the Internet. So now everybody's using them. So it becomes very confusing kind of in this big, thick mix of, you know, how to sort this out. So you're an organization out in the world today, and you're trying to make sense of all of this. And if you just look at your last answer alone, it's just like overwhelming for most organizations to look at.
18:30There's all these different characteristics. There's big models, small models, open source, closed source, you name it. You can slice it so many different ways. How do you make sense of that? If you are, let's say that you're in management at an organization, not necessarily the data scientist who's 25 and gets the data side, but you're trying to figure out how do I do this in the larger sense? How do you start making sense of that? How do you know if you need your own model that you're going to create, if you're going to go use somebody else's big, small, what's a good starting point for people to start sorting through the mess that we're all delighting in today?
19:07And it is a mess. And I get calls all the time from model governance folks that are trying to like, we need to set out a blueprint for our company. We need to think through this. Because right now, the incredible change of the pace of change and all of that, right? That's the downside of that. Like, if you're trying to understand what's going on, it's really hard to. And I think a lot of organizations at this point, there's not a lot of easy cases for like, let's implement this because it's going to 10x our revenue for this particular thing. I think there is a lot of breathing room in terms of enterprises and being able to figure out what the best strategy is for the models over the next year or so like that.
19:45I personally really benefited from hugging face tooling around this. So like some of the decisions that I've made in terms of my own integrations into the applications that I'm building are because I know there's a community around some of these sets of tools. there's sort of interoperability if I want to pull in like this model size or that model size or like whatever it is and even like these large models like you mentioned bloom there's so much integrated tooling with I remember a really awesome blog post about like running bloom in collab using accelerator bits and bytes and these things for like quantization and all this and all All of that set of tooling from this like hugging face ecosystem, I think is so powerful for people actually practically trying to do this.
20:36I'm wondering, like, there's so many cool tools coming out as well, like in that ecosystem. You're, of course, at the center of it, you know, being part of that community and that company. Any highlights that you'd like to highlight? Like I highlighted the one which was is really cool and I'm playing with. But what else should be on our radar? That's great. I know both of you kind of enjoy the Hugging Face ecosystem and have spoken highly of it before. And the Hugging Face ecosystem is all about just helping to kind of create and democratize machine learning, build out the open source for it. To Chris's earlier point, we have a place where everybody can go and check the models and read what is the licensing for the model.
21:17You know, what are the implications for that and learn about that. Now, when it comes to these large language models, like we've been busy building out pieces on that. So if you think about kind of training these large language models, Nathan on our team has written some blog posts around using techniques like reinforcement learning with human feedback. That's the latest cutting edge approaches to figuring out like how to get these models to align exactly with what humans do. Because, yes, we can feed a bunch of data into the models, but what comes out of them often isn't what you and I would think is the best.
21:49And so this using reinforcement learning with human feedback does that. I think one of the things I'm excited about is the PEFT library that we have, which is parameter efficient fine tuning. And if you look at these models, they're huge. They take a ton of resources to do. PEFT has a number of different approaches in there. How can we fine tune these models without having to load the entire model and modify every weight in this? And there's a number of different techniques. for example, just, hey, can we take the entire model weights and find a smaller structure inside them, like a low rank approximation?
22:24I can't think of that name. Can we get then that little dense piece and just train that part and add that onto it? And if we do that, that actually works as a fine tuning technique without having to train the entire model. So I think this is where the Hugging Face team is busy building out a lot of infrastructure and tooling so we can kind of all effectively use these large language models. It reminds me that tooling is tactical in terms of solving problems. And for Daniel and me, given the podcast name, tactical is practical. Wow, that's good. Maybe we should redo our tagline there, Chris. I'd have to run it through chat GPT first to make sure it was good.
23:05We always talk or we've talked many times on the podcast about how a lot of times the practical side of AI is on the inference side, not as much on the training side potentially, because like 99 % of what you're going to run your model in production is inference. I'm wondering, with these large language models, I can see various scenarios happening. A lot of people are just putting that thin UI on top of open AI, and they're never training anything, and they're using that in context learning. But now with the tooling that you just talked about, there's sort of this ability to fine tune these large models in a way that like wouldn't require you to have, you know, a bunch of racks of GPUs, right?
23:52But maybe you could even do it in like some hosted system like a collab or something like that, right? So how do you think that shifts people's kind of approach to how they're solving problems over the long run? Because it was sort of like for a while, everybody's training their scikit-learn model, right? And then it seemed like for a while, okay, now I'm just going to use APIs because I can't train these models. And now we're kind of coming back to this, okay, well, what about fine-tuning, parameter-efficient, like we're not loading the whole model in? How do you think that changes things moving forward?
24:25As somebody who's worked inside enterprises for a long time, I knew the infatuation with OpenAI's APIs we're only going to last so long because I've tried to sell data scientists a black box solution. You don't get very far. If it's inside your enterprise and your reputation, your job is on the line to make sure that model works, you want full control over it. Not to mention enterprises want full control over their data that's going into the model and how it's being used. You're going to see, and this is where there's been so much energy is in this development of open source large language models.
24:58But I mean, what's blown me away in the last few months is just how widespread this community is. Because I think, you know, some of the developments you've seen are around C++ interfaces for large language models, right? Things that no data scientist I know would be able to develop something like that. But because there was so much excitement, we got other folks, right, typical software developers engaged in building tools. And I think there's a lot of focus right now on building these types of tools for this efficient type of use of large language models, because nobody wants to have a cluster of GPUs like that.
Read the full transcript
25:32Microsoft, in fact, just today released their DeepSpeed Chat tooling to help people train models using less infrastructure, right, being able to do it faster. So I think there's going to be tremendous development of tools, because at the end of the day, most people would like to have a model that they can fit inside their computer or a couple of GPUs, something that doesn't take a lot that they can control, that they can tune. And so I think we'll see a lot of development in progress in terms of open source pieces for that. Well, Raj, I am curious to know how many of your conversations these days around AI models and large language models are about some of that tooling and practical stuff that we just talked about, and how many are around ethical concerns or hallucinations or environmental concerns?
26:20What does that look like in your life right now? So that, of course, is a huge part because, again, this is like the difference between traditional machine learning where we often thought about bias in models, right? Like, is your model going to work for kind of a young generation versus an older generation like that? But now with large language models and the ability of generative models, right? They're creating information, like how accurate it is. One of the common fallacies we see, and hopefully most of the listeners here are quite aware of that, that these models lie. They're just going to create output and the output doesn't necessarily have necessarily a tie to reality with that.
26:57So this is one of the biggest education pieces that we have to do is because people see OpenAI, they see the other tools and they're used to just typing in a question and getting back an answer. But to really use this in like, let's say, an enterprise setting, you know, this is I always suggest to people to pair this with traditional information retrieval techniques. Like we already know good ways of having to search and pull information. Let's use those ways that are factually based. And then we can still layer on top a large language model to give you that nice chatty type interface, right? The large language models are great at writing like that and take advantage of both.
27:32But yeah, there's a tremendous amount of like education that has to be done around, for example, hallucinations. You know, that's just the tip of it. There's also, right, like what's the training data that was used for these models, right? Like where did that come from? And then once you use these models and you get output from these models, and this is where customers, especially for some of the code generation ones and image generations are worried about is they're worried about their own legal consequences of using these models that that might have some type of leakage from the training data and copywritten material that could be in the outputs.
28:05It's a lot of different issues going on. I've had conversations with people in various companies over things like the open AI licensing model since they're using chat GPT. And it's really made people aware that you can be giving over IP for using. And that's just one of many possible concerns. I want to throw something, and I know you've been asked this a whole bunch of times because it's a really big topic. I'd love to hear your take on it. given where we're at right now with large language models and some of the variants that we've talked about here, where does this sit in the concept of education?
28:36You have the gamut being run from you're not allowed to use any of these models for your coursework. And then on the other side, and I think I may have mentioned this to Daniel a few weeks ago, I have a 10-year-old daughter, actually 11-year-old in fifth grade, and she had an assignment. And I actually started us off going and doing some stuff and chat GPT ended up having her do her own work, but I actually incorporated it in. But I've also talked to people who are deathly afraid of it skewing academia and how you're measuring students' progress. What might be a reasonable path forward in terms of trying to integrate this new technology into schooling?
29:14I'm very pragmatic and know that we just have to kind of accept it and adopt it. Me too. Now, I agree that there's going to be short-term issues to figure out who has access to the technology, making sure everybody, right? Because this is an easy way for people who have access to those resources versus don't to further differentiate themselves and kind of even increase the differences between groups even more like that. But I'm very pragmatic like this. Like, I think it's a very helpful tool. It's very useful. And it's going to be a part of how we work. It's not only on the education of like students in terms of young people.
29:45I also think we also need to get our coworkers on board too, because I think a lot of us probably listeners are early adopters that like playing with this. But I spent time teaching my sales team how to use the tools like Claude is built into Slack. I'm like, hey, look what you can do with this. Because I think it's one of those things that can enable a lot of people. But it takes a little bit of education, a little bit of pushing to get people who aren't kind of used to these tools adopted. And especially not only with the good that can come, but also like we talked about earlier, the hallucinations.
30:14And so that they properly kind of use these tools as well. Also, I think it's the web developer, other developer community that's starting to enter this space like you were talking about, like with the C++ stuff or other things. There's other people contributing, which I think is great. You have a wider set of views being brought to the table around how these things should behave, how we should use them. Like there's a lot more people at the table. And I know one of the things I've seen, of course, a ton of that startup energy to like people building things on top of this, you know, some very quickly, like I say, that's just a thin landing page on top of OpenAI, but others that are like really fascinating and interesting use cases for this technology.
31:00Of course, a lot of that community as well overlaps with the community using Hugging Face tooling and those you're Hugging Faces interacting with. What is it like for you to see that energy around startups? And there's so many things coming out. I know startups are already like low percentage chance of success, but a lot of these things are really amazing and I think could reshape like how we work, how we learn, like you're talking about, Chris, and other things. So what are you thinking around that front? and also having a sort of front row seat to see a lot of these things being released. It's an amazing time like that.
31:36And I love seeing the startups because people are experimenting, trying new ideas, trying new things, right? Most of them will undoubtedly fail, but I think in the meantime, we're gonna get a lot of good ideas for different ways and approaches that we can kind of use these tools. And that right there has me very excited about that. You know, one of the things that I've been really having some interesting conversations is, are about people who are not us, not in our audience, people who are in the larger world, and really, you know, may have loosely followed, you know, kind of what's happening in the AI space and kind of in the mainstream media, but they're struggling to really understand what's happening right now.
32:17And, you know, we kind of started the show off on that whole premise is that there's so much happening right now, to the point of I was at a dinner function recently, it was just a couple of weeks ago. And I met this really cool dude who was in his mid 80s, but really sharp, followed technology. And we started talking about AI. And he's just like, I'm trying to track it and understand. And one of the points I turned to him with the idea that we're having these large language models that are now penetrating into everyone's consciousness, I said, this is that moment where you're going to look back and realize this was where you were conscious of AI being a part of your life.
32:57And it will take off from this point forward. When you're talking to people about these issues, how do you adjust people who are not used to this stuff the way we are? How do you get them into the right way of thinking about it in a productive way and kind of onboard them? Because it's not the same conversation today as it was a year ago, as you said. It's changed. So how do you approach tackling that? I think one of the easier ways is if I can get them to use the technology, if they can use an image generation where they can type in something and see the differences in results that might get or use a chat GPT where I can kind of coach them and do that.
33:36Because you're right, like trying to explain exactly what this does without the context of actually using it. It's like telling somebody about something in the future that it's hard to kind of contextualize and understand what's going on. The easiest answer for me is just like getting them using it a little bit. And then that helps to then showing them then like, what are the boundaries? What are the limitations? Like, what can we do? What are the possibilities once they have a grounding on that? As kind of a follow-on to that, we've kind of acknowledged we're in this sort of historical moment.
34:07And in years past on the show and last time you came on stuff, we might have talked about kind of historical moments in the context of AI. but I think we're all agreeing that it's becoming a historical moment for the whole world, whether you're in AI or outside of AI, because it's impacting everybody. You also acknowledged along the way that there are kind of ebbs and flows, you know, that we have. We're certainly at one of those moments of just intense new stuff coming out. What do you see in the future, both short-term and long-term? Where do you think we're going from here? Because it feels like we're in an Alice through the looking glass kind of moment.
34:42So what might the future look like and what are you guys anticipating at Hugging Face? So I agree, right? This is just an amazing moment. And I think it's more so for the people that are in it that understand AI and what's going on and kind of where the steps we've made over the last year, you know, and where we can go going forward. I still think we still have to figure out when we're talking about kind of larger humanity and the larger group of people, exactly what is the impact and how we're going to use this. Because, yes, we have chatbots, but most of us didn't spend a lot of our life before using chatbots, right?
35:16Like, I don't know how much of our lives, you know, going forward, we'll have to do that. So we'll have to see how that's integrated. But I think, you know, all of this just shows us that the idea of AI, the idea of having using machines to help us make better decisions is something that is becoming much more widespread. bread. We're really kind of on a path with hugging face to help democratize that, bring that barrier down, allow more people. So not just the people who have been trained for four years in statistics and went and got a PhD, but somebody that can think through a problem a little bit, go interface back and forth with a computer, can all of a sudden build code or solve a problem by tying some prompts together and really allowing lots more people to harness the collective AI, the collective information that we have and allow for more productive uses like that.
36:06Gotcha. Good answer. As you know, Daniel and I have been longtime fanboys of Hugging Face. We think it's a fantastic community, amazing tooling. So as we close out, you want to point some folks to maybe a few things that Hugging Face has to offer that might be good ways of ramping up in different areas. Just kind of call them out. Absolutely. So the Hugging Face website is a great place to start. There's a free online course that you can start with using Transformers there. There's forums, there's Discord, there's a community there. Feel free to kind of jump in and kind of get engaged there. And then we're building out lots of pieces.
36:44You'll see more models coming up over the next few months that we're going to be releasing more tooling for working with this. So yeah, there's a lot going on. Fantastic. Well, thank you very much for coming back on the show. You're always exciting. You're fantastic at representing Hugging Face. and sharing your perspective with everyone. And we'll have to do it again sometime soon. Thanks a lot, man. Absolutely. Thank you for having me. I enjoy this.
37:16Thank you for listening to Practical AI. Your next step is to subscribe now, if you haven't already. And if you're a longtime listener of the show, help us reach more people by sharing Practical AI with your friends and colleagues. Thanks once again to Fastly and Fly for partnering with us to bring you all Change Talk podcasts. Check out what they're up to at fastly.com and fly.io. And to our Beat Freakin' Residence Breakmaster Cylinder for continuously cranking out the best beats in the biz. That's all for now. We'll talk to you again next time.
38:00Game on!
From the publisher
Large Language Model (LLM) capabilities have reached new heights and are nothing short of mind-blowing! However, with so many advancements happening at once, it can be overwhelming to keep up with all the latest developments. To help us navigate through this complex terrain, we’ve invited Raj - one of the most adept at explaining State-of-the-Art (SOTA) AI in practical terms - to join us on the podcast.
Raj discusses several intriguing topics such as in-context learning, reasoning, LLM options, and related tooling. But that’s not all! We also hear from Raj about the rapidly growing data science and AI community on TikTok.
Changelog++ members support our work, get closer to the metal, and make the ads disappear. Join today!
Sponsors:
- Fastly – Our bandwidth partner. Fastly powers fast, secure, and scalable digital experiences. Move beyond your content delivery network to their powerful edge cloud platform. Learn more at fastly.com
- Fly.io – The home of Changelog.com — Deploy your apps and databases close to your users. In minutes you can run your Ruby, Go, Node, Deno, Python, or Elixir app (and databases!) all over the world. No ops required. Learn more at fly.io/changelog and check out the speedrun in their docs.
Featuring:
- Rajiv Shah – Website, GitHub, LinkedIn, X
- Chris Benson – Website, GitHub, LinkedIn, X
- Daniel Whitenack – Website, GitHub, X
Show Notes:
- Solving AI Tasks with ChatGPT and its Friends in HuggingFace | GitHub
- Generative Agents: Interactive Simulacra of Human Behavior
- Wolfram ChatGPT
- Comparing LLMs
- LangChain
- Learn about LLMs:
- Learning Prompting
- Getting Started with Transformers:
- Training your own LLM Models:
- Dolly blog post
- Illustrating Reinforcement Learning from Human Feedback
Something missing or broken? PRs welcome!




