Energy Star Ratings for AI Models with Sasha Luccioni - #687

3 Jun 2024 · 48 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

TWIML AI Podcast Episode Summary: Energy Star Ratings for AI Models with Sasha Luccioni - #687

Podcast Overview Podcast Title: The TWIML AI Podcast Host: Sam Charrington Guest: Sasha Luccioni, AI and Climate Lead at Hugging Face Episode Focus: The environmental impact of AI models, comparing energy consumption among different types of AI models, and introducing the Energy Star Ratings for AI models.

Key Discussion Points

  1. Introduction to Sasha Luccioni
  2. Background: Former postdoc at Mila, transitioned to Hugging Face for a supportive research environment.
  3. Involvement in Climate Change AI and the growing awareness around machine learning's environmental impact.
  1. Energy Consumption in AI Models
  2. Main Finding: Generative models can consume up to 30 times more energy than task-specific, non-generative models for similar tasks (e.g., answering questions).
  3. Discussion on the substantial energy growth associated with Large Language Models (LLMs) and the ongoing misconception about their energy requirements.
  1. Research Insights
  2. Power Hungry Processing Paper: Focus on the energy costs of AI deployment rather than just training energy.
  3. Analyzed 90 models from Hugging Face to measure energy usage during inference.
  4. Discovered that the energy cost of inference can exceed training costs due to continuous model operation in production.
  1. Energy Star Ratings for AI Models
  2. Concept Inspiration: Based on the EPA's Energy Star program for appliances, aiming to provide a standardized efficiency rating for AI models.
  3. Implementation Plans:
  4. Evaluating models across 10 tasks, such as text generation and image captioning.
  5. Providing a comparative analysis of energy efficiency without assessing performance (e.g., washing clothes vs. energy consumed).
  6. Encouraging voluntary participation from model providers for transparency.
  1. Importance of Energy Efficiency in AI
  2. Advocates for the need to focus on energy efficiency as organizations become more conscious of their carbon footprints.
  3. The challenge of balancing model performance with energy consumption and the importance of task specificity in model selection.
  1. Critique of Current AI Evaluation Practices
  2. Concerns Raised:
  3. The difficulty of replicating results in AI due to lack of transparency and evaluation standards.
  4. Challenges in comparing generative models to task-specific models due to differences in input/output formats and evaluation metrics.
  1. Future Directions
  2. Exploration of partnerships with organizations like ISO or NIST to formalize the Energy Star ratings.
  3. Interest in standardizing practices for evaluating AI models regarding energy consumption.

Key Takeaways

  • Energy Consumption Awareness: There is a critical need to address the energy costs associated with deploying AI models, especially as large models continue to dominate the landscape.
  • Energy Star Initiative: The development of Energy Star Ratings for AI models could provide clarity and guide users toward more energy-efficient choices.
  • Community Engagement: Emphasizing the importance of documenting AI models in a straightforward manner to encourage more widespread adoption of sustainable practices.

Conclusion Sasha Luccioni's insights underline the importance of considering the environmental impacts of AI technologies. By establishing Energy Star Ratings for AI models, the goal is to foster a more responsible approach to AI deployment while simultaneously advancing the discussion about sustainability within the tech community.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00And what we found is that for a given task, so for example, like answering a question, so something like, you know, if you're like, what year was Napoleon born? and you have a model that's specifically trained to extract information from text, so extract of question answering, versus a model that's like a generative model that can also do the same kind of task, but in a generative way, not creating new content way. The difference in terms of energy use is staggering. It's up to 30 times more for a generative type model.

0:40All right, everyone, welcome to another episode of the Twommel AI podcast. I am your host, Sam Charrington, and today I'm excited to be joined by Sasha Luccioni. Sasha is AI and climate lead at Hugging Face. Before we get going, be sure to take a moment to hit that subscribe button wherever you're listening to today's show. Sasha, welcome back to the podcast. It's great to be back. It's great to chat with you. I've been a huge follower on the socials, excited about all of the, let's just say, interesting comments you have to say about the current generative AI revolution. And we'll be digging into that.

1:22A lot of that has to do with the climate and broader implications of those types of systems. We last spoke back in September of 20, where we were talking about some of the work you did as a postdoc at Mila, visualizing climate impacts using generative adversarial networks. We've come quite a way since September 2020. uh yeah i'd love to have us start out by having you share a bit about what you're up to now and uh you know how the the world of you know climate impact of machine learning has shifted since then that's a so that's true it was a while back so i was a postdoc at the time i finished my postdoc at mila um went on the job market tried to do the whole uh tenure track professorship situation did not work.

2:12And then I actually stumbled upon Huggy Face. I didn't really know it that much. And, you know, I was looking for a place where I could do the research I wanted to do in a, let's say like a positive environment where I get support. And, you know, essentially I have issues with people telling me what to do. So I didn't want to be in one of those, I mean, tried it, didn't like it. The whole rigid corporate structure wasn't for me. So I was like, maybe startups could be my thing, except, you know, a lot of startups are really hectic and stressful and then hugging face like I actually started part-time and I I saw how people were actually pretty chill and really interested in like responsible machine learning stuff like that so I've been here for almost three years now and what I've seen um I guess that what's interesting is that there's there's a lot of climate positive work going on it's not getting a lot of attention so I'm still involved with climate change AI it's grown a lot and we still do workshops we still do you know even summer schools and and um all sorts of events um and what's sad is that like compared to like gpt type models llns um you know the kind of og climate work doesn't get as much interest nowadays because a lot of it i mean most of it actually uses non-generative methods let's say right people are still into cnn's people are still into random forests and and it and it actually is truly useful and impactful and can be you know applied for all these different contexts.

3:38And yet what we see is that people are trying to train the new climate GPT or whatever people call it nowadays. And so I think that the work that I'm doing in trying to estimate the environmental impacts of these models is more and more, I guess, relevant, essentially, as they get deployed in society. But we do have this kind of like cache 22 of like of the models that are maybe the least useful well so far in the fight against climate change also have a disproportionate climate impact and vice versa the most useful models are pretty lightweight so it's pretty yeah it's pretty confusing you know one of the things that's become very clear since uh well over the past four years uh in fact the past couple of years is the kind of the rise of LLMs and their eventual takeover of at least the hype cycle of AI at this point.

4:38And the critique around their energy consumption was one of the early critiques about LLMs. How is the way we thought we're thinking about LLM climate impacts evolved? I don't think that it's one of the early critiques. I mean, it has been raised, definitely, but I feel that it's kind of been dismissed. I'm thinking of the Stochastic Pirates paper, for example. Yeah, definitely. That was part of it.

5:09But I wouldn't say it's mainstream. I feel like every time I bring it up, people often talk about bias. They'll talk about hallucinations, for example. But when you start talking about the environmental impacts, I often get a blank stare, especially outside of the AI community. People are really like, oh, because people think it's the cloud. Even often people will think that ChatGPT or MidJourney or whatever is actually running on their phones and not on data centers. And they're like, well, if it's running on my phone, how bad can it be, et cetera. So I feel that, broadly speaking, it hasn't really connected in people's minds that these things do have pretty disproportionate energy requirements compared to your average piece of software.

5:48I just did an interview with someone at Microsoft and I may be butchering the number, but I don't think so. I think he said that they're opening like a new data center every five days. And a lot of that is driven by AI, like workloads and usage. usage. And that was a staggering number for me, if I'm remembering it correctly. Like it was staggering for me and I think I'm remembering it correctly. And in your research, you've published on the energy growth due to LLMs. Any kind of figures that resonate for you? Yeah, we found, I mean, we looked at a lot of open source, like language models, AI models from the Hugging Face platform where people share these models, train and share these models.

6:36And what we found is that for a given task, so for example, like answering a question. So something like, you know, if you're like, what, what year was Napoleon born? And you have a model that's specifically trained to like extract information from texts. So extractive question answering versus a model that's like a generative model that can also do the same kind of task, but in a, generative way, not creating new content way. The difference in terms of energy use is staggering. It's up to 30 times more for a generative type model, which kind of makes sense because on the one hand, I mean, my very first job was working on information retrieval and question answering.

7:15And it's, you know, you're extracting existing content based on, you know, some measure of similarity, some kind of vector arithmetic or hashing. Like it's relatively straightforward in terms of computation. Like, I mean, we used to be using Elasticsearch and stuff like that in my first job. But for generative approaches, you've got usually a pretty big model and that has to actually kind of, you know, take the query and add to it. So like generate new content. So like really from an intuitive perspective, it does make sense that it uses a lot more energy. And so you're specifically speaking to one of your recent papers, which is the power hungry processing paper, what's driving the cost of AI deployment.

8:02And the general idea that I took from that paper is this tension between kind of using generalized models versus task specific models. And you started to speak to that. Take us a step back and talk through what you were trying to illustrate in that work. the goal of the paper was really to look at ai deployment so um most of the papers to date we're looking at the training step which is a little bit easier to get a handle of just because it's like you press start you train your model for i don't know three million hours and then you press stop and then you you kind of calculate how much energy you use but deployment is a little bit more tricky because most models like for example if they're in production they're constantly on and they're constantly replying to queries and actually oftentimes you've got multiple copies of the model running depending on you know the load and then you've got all sorts of so people tended not to look at deployment because it was kind of messy and so what we did is that we took a bunch of models like 90 models from the Hugging Phase hub and we essentially measured how much energy they were using for the same set of tasks of queries and so we looked at image generation we looked at text generation but also like text classification image captioning we looked at a lot of different tasks and the idea was like to to essentially see how tasks compared to see how models compared and to kind of try to identify some trends with regards to really deploying ai models um in practice let's say and was general versus task specific the only um kind of dimension of distinction that you looked at or was that one of several no we looked at that we also looked at model size.

9:39For example, what was interesting is that if you looked at models like text generation models of different sizes, we looked at the trade-off between how much energy was used for training the model and for per query, essentially, for each model query. And we found that depending on the size of the model, it varied between, I think, 50 million to 200 million queries to have the equivalent energy of a training run, which may seem like a lot. But on the other hand, if you do have like a production level system, like chat GPT that gets, you know, millions of queries a day, that means within a couple of weeks, essentially you've used as much energy for just deploying the model as you did for training the model.

10:19So, um, that was pretty interesting. We also looked at some image generation models, which starts to speak to the point that you're making earlier in that the, um, the climate cost of inference in the long term is much larger than the training costs, at least for models that eventually make their way into production. The models that don't make their way into production is an entirely different question. Yeah, exactly. And then nowadays, especially for LLMs, for language models, people tend to use existing models. I mean, actually, it takes a lot of resources to train these things. So mostly people are either fine tuning, they're adapting, or they're using existing out of the box models.

11:00And so that's why I was particularly interested in this phenomenon of like, how does, how do, how do LLMs compare? What are the different, you know, trade-offs between training and inference? And that paper, we also looked at, for example, like inherently generative tasks, like summarization, where, you know, sometimes you'll have a small summary, like a short summary, sometimes it's longer. And we found that, you know, the, the energy cost, I mean, once again, predictably, I guess, rises as the output length gets longer, which is also telling you because people are nowadays asking for whole essays from chat.jpg type models.

11:36And so that means that essentially every time you have to run the existing text through the model and get the next token, and so as that text gets longer and longer, the compute cost rises and rises. And so if you want to generate a whole essay, then that's a very long context window that needs a lot of computation. So we did a couple of like, we found a couple of interesting, I guess, observations in that study. But for me, it was like really just kind of planting the flag and being like, this is something that we should look at. And now I would love to see more work about optimization strategies or like the impact of flash attention or distillation, right?

12:11Like, because there are, there is low hanging fruit that we could be using, but that people aren't necessarily thinking about because they don't really, you know, think about these trade-offs. Like, well, if we, if we reduce the precision from 30 to 16, then maybe we could kind of reduce the energy use by half or something, right? It's not like really calculations that people often make in their heads when they're deploying AI models. Yeah, yeah. There are a couple of trends that may help in the energy consumption department. One is the trend towards smaller models, like Microsoft Phi or Phi is an example of this and the other is the trend towards investing more on the data side um and i guess ultimately producing smaller models as well um do you do you see any evidence that you know a that those trends are you know kind of taking rid of scale and b that um you know we're going to be able to quantify their impact on um on energy use it's hard to say because there's honestly very little transparency, but it's true that now, like before you would almost be like, people would almost be like, we trained, you know, a bigger and bigger model just so we can say that it's the biggest model.

13:26So there was like 160 billion, 170 billion, 180 billion. And nowadays you do see often like people train bigger models. Um, and then also concurrently they'll say, and we have a smaller, more efficient model. So, I mean, it's not, it's not quite a win, but I do see that there's reflection. And I think that, I mean, maybe this is just my take on it, but it's maybe not so much the environmental impacts or the energy use. It's more like there are so few people and organizations that can even afford to use these models or to deploy like 180 billion parameter LLM that it's just in terms of usage. If you want people to use your model, having a smaller, more efficient model makes sense because you have potentially more clients and more users.

14:10I'm wondering how much of that particular perspective is related to, you know, your perspective as working at Hugging Face, I think for much of the world, you know, they use these models hosted by, you know, an open AI or an Anthropic or Microsoft, Amazon, whatever. And so they want the biggest model because that's the easiest model to use to figure out if their use case is even going to work. And, you know, know, they're not necessarily trying to self-host it, or at least they're far from trying to self-host it.

14:49Does kind of the way the industry is structured and the way that, you know, we've got large organizations hosting these models, to what degree does that play into the impact that you're trying to make from a climate perspective? I definitely chose Huggy Face as a place of employment, not only from a climate perspective, because I do see like the having a hub where people can can find models easily does I kind of in direct ways kind of like recycling almost is like well I'm not going to train my own model maybe I'm going to start with the base model and fine-tune it but you know it helps people reuse reduce reuse recycle um ml models which is was kind of one of the things and also I do believe in um you know like I guess uh like I don't think that monopoly is is the right or concentration of power.

15:40I don't think it's a very healthy dynamic for any industry. And I did see, and I still do see to some extent, a concentration of power in AI. I mean, when you think about it, and Meredith Whitaker talks about this a lot, these bigger and bigger models really do put bigger companies at an advantage in terms of, right, who has the resources to run them, to train them, to deploy them. And so I also do believe that there's a climate perspective, but there's also kind of the democratization perspective that that hugging phase really fits in well into and and i think that if you if we want to have more innovation if we want to have more startups and and even you know non-profits and academia having more accessible models is important for that because like if i talk to students who are like starting a phd you know they can't afford to run 180 billion parameter model but maybe they want to do a project or or you know, a hackathon or whatever.

16:34And in order to enable them to do that, we need open source, like smaller open source models. And so I see it as like an advantage on different levels, one of which being climate, but actually something that I was working on recently, a paper I was working on recently with folks from the DARE Institute. So to me at Gibru's Institute is about how actually in the context of AI ethics and like sustainability and climate are really, really related because like maybe like a smaller model will have less climate impact but it will also be more accessible in terms of like justice and equity in our field and so it was interesting because we had a lot of these conversations about like by by helping people like you know access or train smaller models that can actually have impacts that go above and beyond just like simple energy or simple environment it's actually like a powerful tool in terms of like accessibility and and the and the distribution of power.

17:30So I thought that was interesting because I never really thought about it like that. And they brought this into the conversation and I was like, oh yeah, it's true that it actually does help accessibility and not only climate, right? There was a point in time, maybe a couple of years ago when there was a significant concern that many organizations and institutions would be kind of locked out of AI research because AI research was trending so heavily towards these gigantic expensive models that only a few institutions could afford to train and projects like bloom at the time i think i was talking to tom wolf about this uh and you know more recently uh you know the llama series of models have kind of opened that up again and to your point kind of made experimentation and research more accessible by uh publishing interesting but small models yeah that's actually how i joined hugging face it was around the big science project in bloom and i was um part of the working group that aimed to measure bloom's carbon footprint and it was interesting because we were using french compute that was relatively low carbon and but also we did a really in-depth study of um like like the training logs and we could actually go really deep into like what each part of the process in terms of like training smaller models in terms of experimentation like all those different little pieces of the puzzle how they all fit together and but i mean i agree that llama is still accessible but but you know if you look at a couple of years ago i still feel that the the field of ai has become less transparent i mean just if you look at conferences like i don't know whatever three or four years ago if you submitted a conference a paper to a conference um you had to give a certain amount of information about how big the model was like what where the data was coming from like things like that right but nowadays i feel like you could submit a model that's behind a paywall and you can just give some examples of you know how great the images are how how convincing the video is and then and then that's kind of okay as well even though it's a scientific conference and you know science is supposed to be to some extent transparent and replicable.

19:44So I do feel that there is kind of like there is still that tension between commercialization and science in AI, which is, I mean, maybe it's inherent because we're starting to have, you know, the rubber hit the road in terms of AI innovation, but it does really make some awkward situations where it's like, okay, like this is a really cool model, but like, how does it work? essentially right is one of the challenges that these general purpose models are uh state of the art for a lot of the tasks that the task specific models um you know are designed for or is that you know is that part of the hype problem surrounding these models i find that evaluation in ai has become so broken i don't i mean i know people are doing stuff but like when you start digging into the details like for example usually when you're evaluating like a language model it's by prompting right and so people do use different prompt approaches or sometimes they actually sample they'll be like top five right top k and then they'll compare it with a fine-tuned model that essentially has a set like unlimited number of options like categories or labels it can choose from and it's just like like sometimes the the comparisons are just apples to oranges and it's really hard to even say like, and also like often language models aren't replicable.

21:10Like there's some great, like they do some evaluation, then you do it again and you can't get the same result. And so I've done that evaluation. It's like, it's just so hard to even say which model is better. And people keep on inventing new benchmarks and then finding that the benchmarks are in the training data. And then it's just like, I don't know, like every time I read something, it's like I have so many questions that aren't really answered. So, I mean, I guess the theoretically general purpose or multitask models are supposed to be better than fine-tuned models, but people actually have kind of given up on working on fine-tuned models.

21:45Like, for example, recently I wrote a position paper that was accepted to ICML with a colleague of mine, Anna Rogers from Copenhagen. And then we wanted to look at tasks where people were still kind of working on fine-tuned models and multitask models. And essentially, we couldn't find many because people are just also gung-ho about large language models that they'll tend to kind of work on these like general tasks, general language understanding and things like that. And so we actually had a lot of trouble trying to find like concrete evidence for the fact that kind of the evaluation doesn't make much sense.

22:19Yeah, you could argue that that's good. We're concentrating our resources on the most promising thing. We're making good progress. Yes, you know, we've achieved state of the art on, you know, all of these tasks. And I think, you know, the first part of your answer to that is, well, you know, we're not talking about the impact and that's climate. Are there other parts of the your challenge to that response? Right. I think that we're also cherry picking a lot of results and not like, for example, the concept of emergent properties is something we talk about in our position paper. And it's like, when you talk about emergent properties, it's like you assume that things kind of almost magically appear and you're like, oh, this model just magically is able to do this thing that it wasn't trained for.

23:04But the fact that we don't know that, like what exactly goes into the training data, you know, this data is so big that it's even hard to say that all of the law books or something like, for example, yeah, GPT, whatever passed the law exam, the par exam. But maybe all of the books are in there. We don't know. Right. And so there's all these issues with like the claims that people make, um, same with kind of state of the art, uh, usually like often it's like a new data set. So it's really hard to, or, or, or, you know, uh, I don't know. It's just like, or less a new task, like drawing a unicorn.

23:39So, um, so yeah, so I mean, I'm, I guess I'm a little bit skeptical because I miss the good old days where you could kind of replicate or check or reproduce people's work. Like nowadays, I feel that there's like the access issue of, for example, models that, you know, are just behind an API. So you can't actually look at what's under the hood. But it's also like just the fact that you can't like reproduce results. Right. That's so frustrating as like from a scientific perspective, because it's like the basic thing. If you make a claim, I should be able to replicate that claim. And, you know, that's kind of how science works.

24:14Like we keep building on each other's experiments or discoveries. Whereas in machine learning, I feel that a lot of the discoveries that people make or the innovations or the claims are missing this core element of replicability. And so for me, it's like, well, if I can't even replicate what you've done, how can I continue building upon it? How do we grow as a community? And is the replicability critique that you're mentioning now, is that related to access or related to information? Meaning, you know, you're doing your experiments on GPT-4 and I can't afford it. Or is it more you're doing these experiments, you haven't really provided enough detail for me to go do the experiments myself?

25:00It's access. It's also often like, for example, there's some element missing, either the data or like sometimes like the actual evaluation approach. Like, for example, there's great work out of a Luther that created who created an evaluation harness because essentially they found that a lot of the times every time that someone would, you know, create a training new LLM, they would kind of do their own thing in terms of evaluation. And so it meant that things were not necessarily like apples and apples. And so so there's that as well, like the actual procedure used, but also like inherently language models are stochastic Paris, right there.

25:36They're not predict they're unpredictable. And so often, like, for example, people won't fix the random seed or won't report the random seed. And so it's like, I can't even reproduce. Like, even if I have all the elements, if I access to the model, if I have the training data and the code, I can't reproduce the results that you report because there's some, like, random seed that is missing. So I actually was helping organize the ML reproducibility challenge a couple of years. And essentially, the thing is, like, people can choose a model, I mean, like, a paper that was accepted to one of the, like, core ML conferences.

26:07and try to replicate the results. And so like an overwhelming majority of the time, people weren't able to because there was always some little detail missing. And even like, you know, they had to correspond with the authors. They had to kind of reverse engineer a bunch of things. And it was like, year in, year out, it would be like the majority of ML research is not reproducible. Was it NURPS that added a replicability component to submissions and maybe even a climate impact? Yeah, they had a, I helped develop the code of ethics. Yeah. There was reflectability, there was climate impacts, there was accessibility, there was, technically you were supposed to say how much you paid crowdsourced workers, you had any, et cetera, et cetera.

26:49I mean, that's kind of like the gold standard, but in practice, you know, that's the thing. There's also, there's always this tension between how much you can divulge or how much you, you even save this information. Sometimes people just don't even bother counting how many GPU hours they used. And so a lot of the pushback we got when we introduced this code of ethics at NeurIPS in 2023 was that it was asking too much from people. So it's more like something that we're going towards, but we're not quite there yet. But yeah, it is quite extensive in terms of even like proving consent. Like if you're using some data that either you're scraped from the internet or gathered from someplace to make sure that consent is being respected.

27:33But, you know, we're not at 100%, let's say. Tell me about how you're formulating this idea around energy star ratings for AI models. so the i really got inspired by the work that the environmental protection agency has been doing in the united states for for decades now and the goal was really you can't force people sadly you can't say for i guess we should say for those outside the united states energy stars this ratings program if you go to buy a refrigerator or another like home appliance there's a big yellow sticker on the appliance that essentially gives it, I think, I think it's like an ABCD, like a kind of A through F or something like that type of grade.

28:18It depends on, it depends on the appliances actually. For some of them, it's really, if it's in the top 25 % in terms of efficiency of a category, it's got the Energy Star brand, like the rating, but in other cases it is really like an ABCD. So yeah, it really depends. But, so yeah, I was really inspired by that work because you're not forcing people to take this into account. But, you know, typically people do care about their energy usage at home and they increasingly care about environmental impacts as well. And so the idea is to kind of harness the same kind of measurement where you compare for a given task, like for example, for a washing machine, right?

28:54You compare for the washing machine category, you compare different models. And essentially what's interesting is that the EPA doesn't really measure performance. Like they're not going to go and look how well the washer, you know, washes socks or whatever. That's not really like their purview. But given kind of a standard load, like the default setting, they'll measure energy consumption, which I find is interesting because they are, for me, two different things. Like performance can depend on like how you're using the appliance and how you're using the AI model. because maybe for some people, they really need it to be like top state of the art, whatever, but other people just, you know, wash their laundry and, you know, it doesn't have to completely take out all the stains from whatever your white shirts.

29:37And so it's interesting because I spoke to the folks from the EPA and they're like, well, the performance is like a secondary thing. The people who make the appliances can measure that as they will, depending on, you know, different types of loads and whatever, but we measure really efficiency for a default setting. And so I'm taking this, a similar kind of approach for, for the AI Energy Start project. I identified like 10 different tasks. So everything from text generation, image generation, image captioning, speech recognition, like all sorts of tasks across modalities. And for each task, I'm comparing, you know, models from the Hugging Face hub that can do that task.

30:15And once again, I'm not, I'm not benchmarking them in terms of whatever benchmark are used, but really in terms of the energy efficiency. And the idea is to create a range for each task. And then once you have that range from like minimum to maximum, you can start giving categories. And is the idea that model providers would kind of voluntarily write their models on these dimensions and publish the ratings? Yeah, I'm going to start with open source models. And I'm already talking to, for example, Cohere, Meta, like Microsoft. I've reached out to all of them. And I'm going to provide like a script, like a Docker, essentially, that they can run locally and they don't need to give up any state company secrets.

30:57It's really, you know, a standardized kind of script that all the models can be compared with. And hopefully everybody will be incentivized to kind of participate. And I think that especially, you know, big tech companies, they do care about efficiency. They do care about compute and showing that it's kind of like the same as for appliances, right? You want your appliance to have an energy start rating that's considered like cool or, you know, it's good for the planet or green. It's considered green. And so I'm hoping to have something similar for AI because, you know, I mean, maybe I'm a little bit naive, but I do believe that, you know, most people in our community do care about efficiency and compute.

31:39It's just that like the structures are not there in place yet for them to take this into account when either sharing their models or choosing models. So like I think that given the right information, people can really use this to make more informed decisions. and so if i want to use this do i have to have some baseline type of compute that i in addition to the docker container or are you able to abstract it in such a way that the actual compute doesn't matter for calculating this rating maybe you're counting instructions or flops or something like that as opposed to actual energy usage i'm counting energy energy usage uh for like a specific GPU model.

Read the full transcript

32:24So the idea is really to like run your model on this GPU. And so I helped create a package called Code Carbon, like way back when actually. And it's actually grown and thrived. And it's like a Python package, a Python library that essentially measures energy consumption of different like pieces of hardware, GPU, CPU, RAM. And also depending on where the energy is coming from, depending on how it's generated, it also estimates carbon emissions. And so I'm using code carbon in the project in order to have that amount of energy. We went back and forth about the flops situation. And I actually asked the EPA people about it.

33:02And like they, I mean, Energy Star kind of is inherently about energy and not about like computation or cycles of laundry. So I really wanted to keep the core like energy usage at the center of this project. So presumably it's some common GPU that the person who wants to do the testing can find on a cloud provider and they push it there. And are there... And actually we're looking at different categories as well, because I think it doesn't really make a lot of sense to compare like a model that you can run locally on your laptop versus one that, you know, needs eight GPUs to run. And so essentially once I'm done benchmarking all the models and that should be by the end of this week um to compare something that can run locally on uh like a small gpu that only has a couple of gigs of memory versus something that needs really like a commercial like 24 48 gig gpu versus one that needs multiple because i feel that it's not quite the same use cases like i mean to go with the metaphor of washing machines you know you do have kind of smaller for example in europe they often have like these tiny washing machines that and we tend to have a little bit bigger ones and then there's like more you know commercial level washing machines that you'll use for washing, I don't know, whatever, parachutes or whatnot.

34:14And so I think it's similar in AI. Often, it's not quite the same use case, depending on the size of model and the amount of compute that you have. So I do want to help create these little categories of usage for different tasks. But I'm trying to make a trade-off between simplicity and essentially usability. So I'm not 100 % set on exactly what the categories will be yet. You mentioned that you have kind of separated out and aren't really thinking about performance in the context of Energy Star. And the Energy Star doesn't as well. But is there anything out there that will allow folks to get a sense for kind of energy costs performance for models?

35:04like even take performance on public benchmarks and then baseline them against the, I forget the name of your power, kind of your Python package or the Energy Star or something else and try to get a sense for the code carbon, try to get a sense for which of these models is most efficient from a performance per watt or something like that perspective? Yeah, a kilowatt hour or something. So we do have a bunch of leaderboards on Hugging Face. And actually, I have some great colleagues like Clementine who are working really on the leaderboard benchmarking side of things. But I think that especially if you want this to be as useful for as many people as possible, it's really hard to pick a single benchmark or leaderboard that will work for all tasks and all kind of communities.

36:01So, you know, maybe in research, we have our set of benchmarks that we'll tend to use when we train a new model because, you know, that's kind of like what we track and papers with code and whatnot. But if you talk to AI practitioners and, you know, someone who's making, I don't know, whatever, a financed chatbot or whatever, they're not going to really care about, I don't know, the superglute benchmark because what they care about is. And so that's why I struggle with the performance versus efficiency. But I do think that at the end of the day, it's useful to keep them separate. And people can benchmark on what they feel is relevant to their use case.

36:38But the energy efficiency is something that is harder to measure or is less intuitive. So it's useful to provide that information. Do any of the existing leaderboards already have an energy impact column associated with them? yeah yeah there's lm perf leaderboard on hugging face that has it but it's most it's only lms essentially it doesn't look at like for example object detection or or speech recognition or whatever and i think that given the diversity of tasks that exist in ai it's good to also kind of consider all different modalities as well like this is something what we found in the in the what's the cost paper is that tasks with images either like image captioning or image generation and R factor, like orders of magnitude, more energy intensive than tasks that only have text.

37:28Right. And so it's like comparing the two is difficult. And increasingly as models are getting multimodal, and that's kind of like what we expect of them, it's important to like reflect, like, you know, compare apples with apples. So compare multimodal models with multimodal models and not text generation models, because it's just going to be completely different ballgame. Yeah. So the Energy Star is broader than LLMs. You're including the task-specific models as well as the more generalized models. How do you address the challenge of a more complex model that's doing more complex things? How do you normalize against the complexity of the task, essentially?

38:13Once again, this goes back to what you want to do with the model. So maybe in research, we're kind of very, we optimize on like the multitask nature of LLMs or AI models in general. And I feel that is something that we really pursue. But in other applications, like in IRL, in the real world, that maybe you don't need models that are so complex. because for example, once again, in my very first job, we were actually working with IBM Watson. And I remember we had like an intent classifier that was upstream in the model. And then depending on what the intent was, a different model would be activated and it wouldn't always be like the whole kit and caboodle that was activated.

38:59Like the query essentially didn't go through all the models at once. And so I feel like, I mean, we've kind of maybe come back to this with a mixture of experts, But for a while there, we expected models to do everything. And I feel that that's unrealistic from, you know, you can't be good at everything. I don't think even, I don't think AI models can be good at all possible tasks. And so I'm a big fan of task-specific models or at least, you know, context-specific models. And I think that this like trade-off between complexity and performance is not necessarily like always, always relevant. um maybe it's relevant for kind of the more experimental researchy type work but you know in the real world like if you want your model to be a chat bot you don't need it to generate images at the same time you need it to answer questions or extract information you know another task question answering might have its own task or um yeah what are examples of other kinds of tasks that you're targeting with the rating?

39:58So there's image generation, there's image captioning, so image to text. There's question answering, right? There's summarization, there's object detection, automatic speech recognition. So there's like 10 tasks in total, semantic similarity. And so the idea was to kind of represent common tasks. So actually on Huggy Face, we have like these tasks pipelines. And so common task pipelines that people use. And like when in the what's the cost study, we we saw that, you know, it doesn't really make sense to compare very different tasks. Like, for example, sentiment analysis, like text classification, essentially is so lightweight.

40:42Like it makes no sense to compare that with like image generation, which is like on the whole on the other end of the spectrum in terms of energy efficiency. Right, right, right. Right. On the other hand, this Energy Star model in and of itself or these ratings in and of itself won't help reinforce this point that if you could use a task specific model relative to a generalized model, you should because the ratings themselves are task specific. And so you won't be able to see that your LLM is more consumptive on some general scale than your task specific sentiment analysis model or something. like that that's out of the scope of what you're doing right i'm also comparing some uh multitask models like the t5 type models uh that we studied in what's the cost i'm also comparing them in um in the energy star ratings to give people an idea of like what the trade-off would be between task specific and general purpose or multitask um but of course the thing is is that when you start comparing it it's interesting because often like for example you have to prompt them like the t5 like you have to say um you know what's the answer to this question question and here's the context right so you always also actually add some input that will kind of make comparisons difficult and and again people brought this up when we wrote what's the constant and it's true that it is hard once again to compare uh multitask and single task models but um if people are are using them for all sorts of things it's good to see like maybe if it does 30 different tasks but it's 30 times more energy, then maybe it's okay at the end of the day, right?

42:22So it's really kind of like a trade-off of like what you're comfortable with and how you're using the model, right? And what's the setup that you have and how many tasks you really need it to do at once. Can you elaborate on the critique you just mentioned that there's prompting required with the generalized models was the critique that I might get a different result with a different prompt. And so therefore where you can't compare apples to apples? I mean, the idea is that, like, for example, if you're really comparing input and output, a fine-tuned model, like, you just give it the input, like, for example, I don't know, this is a terrible book, right?

43:00And the output would be, like, positive or negative sentiment. Whereas a generative model, you have the same kind of, this is a terrible book, but then you also have the prompt that you have to say, like, categorize, I don't know, whatever, give the sentiment of the following sentence. So the input is going to be longer, and also the output is technically, exactly and so i mean when you think about it like in one case you just have the two categories to choose from the two labels and then the other hand you have all of the vocabulary the model to choose from and sometimes it doesn't even say positive or negative sometimes it says like red and you're like well how you know or something or something more complex like oh this is actually a very nuanced blah blah blah and it's true that if you want to compare fine-tuned with like multitask models, then you quickly see that it's really hard to compare the two because it's really hard to constrain generative models to only two choices.

43:52And often they'll start blabbing about stuff that you didn't expect them to say. So you're finishing up your first set of models with the Energy Star. You mentioned that you talked to the folks at the EPA. Is there any path towards making this some official thing? so they said that they're not quite there yet because um essentially they're focusing on tangible physical objects and ai is not one of them but what i would love to see this become is maybe a partnership with um like iso or like one of the or nist or like an organization that really like i'd love to see it adopted and also taken like beyond what i can do as a researcher to really make it a lot more like standardized and structured and, and ideally, you know, something that evolves over time as new tasks appear, new models appear.

44:46So I, the way I see it is, it's like a proof of concept. And once, you know, it's done, I'll create, I'll write a paper, obviously I'll create a website where people can compare and contrast different models and see the results and kind of go in more into depth. But ideally I'd love for, for someone to, to take this and run with it because I do think that there's like, there's, it would be an important kind of like long-term project to carry over, like over the Green Software Foundation. I mean, there's a bunch of folks that are kind of working on different angles of this. So I'd love to contribute to that as a researcher.

45:19And are there any things that you've learned in looking at some of the parallel work that's happened on the responsible AI side of things or more kind kind of ethical AI side of things, such as model cards and data sheets for data sets. And in particular, their broader adoption within the industry, I've seen some of the vendors publish model cards, but it hasn't necessarily swept and become front and center every time you go look at a model. It is there if you look for it in very limited cases. uh are there things that you can learn from that you know those experiences i think the model cards have become uh pretty much a standard at like for example at hugging face i think the the vast majority of models will i mean probably not like the full model card with all the information but they'll have a subset of information that's shared um i guess what i learned is that people will tend not to like put in not they won't want to you know once when someone has like trained the model, they won't want to spend hours on this kind of stuff.

46:30It's hard to incentivize people to document their work essentially. And so this is why I want it to be like a simple script that someone can run. Or, you know, if they don't have the in-house compute, a simple like query or button they can press that I'll run it for them or, you know, whatever, I'll set up a process to run it for them. So I essentially, I want to make this as seamless as possible because what I've seen is as soon as, you know, documenting data or documenting models becomes a bit more complex and you know people run out of patience it's hard to keep the the enthusiasm alive but on the other hand um i've seen some really great efforts recently of being very thorough being very kind of systematic and like for example um lnai institute recently trained their model olma olma um and they did a great job of like really doing amazing like in-depth documentation of the data of model.

47:19Like I really feel like there are great examples of people putting in the like work, but it's true that, you know, you can't expect everyone to spend hours and hours on this stuff, sadly. Awesome. Awesome. Well, Sasha, thanks so much for taking the time to catch us up with what you've been up to. Very cool stuff and looking forward to seeing a energy star sticker on a model near me. Thank you for having me, Sam. And I'm a huge fan of your podcast. Thank you.

From the publisher

Today, we're joined by Sasha Luccioni, AI and Climate lead at Hugging Face, to discuss the environmental impact of AI models. We dig into her recent research into the relative energy consumption of general purpose pre-trained models vs. task-specific, non-generative models for common AI tasks. We discuss the implications of the significant difference in efficiency and power consumption between the two types of models. Finally, we explore the complexities of energy efficiency and performance benchmarking, and talk through Sasha’s recent initiative, Energy Star Ratings for AI Models, a rating system designed to help AI users select and deploy models based on their energy efficiency.

The complete show notes for this episode can be found at http://twimlai.com/go/687.

More from The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

All 156 episodes
Energy Star Ratings for AI Models with Sasha Luccioni - #687 The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) · 48 min
Listen in VO