In short
The TWIML AI Podcast Episode #671: Are Emergent Behaviors in LLMs an Illusion? with Sanmi Koyejo
Episode Summary In this episode of The TWIML AI Podcast, host Sam Charrington talks with Sanmi Koyejo, an assistant professor at Stanford University, about his award-winning papers focusing on large language models (LLMs) and their emergent behaviors. The discussion revolves around Koyejo's two main contributions: one that questions the significance of emergent abilities in LLMs and another that assesses the trustworthiness of GPT models.
Key Topics Discussed
- Emergent Abilities of LLMs: An Illusion?
- Concept of Emergence: Koyejo's paper, *"Are Emergent Abilities of Large Language Models a Mirage?"*, challenges the widely-held belief that LLMs suddenly gain capabilities (emergent abilities) as their size increases.
- Evaluation Metrics:
- Nonlinear vs. Linear Metrics: Nonlinear metrics can create an illusion of rapid capability growth, while linear metrics show more consistent improvements.
- All-or-Nothing Metrics: The importance of using metrics that provide full credit for correct answers and none for incorrect ones, which can mask gradual improvements.
- Methodology and Findings
- Initial Observations: Koyejo's team observed that emergent behaviors occurred only with specific metrics, leading to the investigation of how metric choice influences interpretation.
- Toy Model: A simplified model was developed that helped illustrate how predictability of performance could be quantified, revealing that sharp changes in performance might not be as unpredictable as previously claimed.
- Empirical Support: The findings suggested that, with enough sampling, gradual improvements in model performance could be observed even with low-performing models.
- Trustworthiness in GPT Models
- Paper Overview: Koyejo's second paper, *"Decoding Trust: A Comprehensive Assessment of Trustworthiness in GPT Models,"* aims to evaluate various trustworthiness aspects of GPT models.
- Key Perspectives: The paper identifies eight perspectives for evaluation, including toxicity, bias, robustness, privacy, ethics, and fairness.
- Evaluation Toolbox: The results include a toolbox that researchers can use to evaluate their models based on the trustworthiness metrics developed.
Key Takeaways
- Metric Choice Matters: Researchers should carefully consider the metrics they use when assessing model performance to avoid misleading conclusions.
- Emergent Behaviors: The concept of emergent properties in LLMs is nuanced and may not be as fundamental as previously thought. The interpretation of these properties is heavily dependent on the metrics used.
- Importance of Trustworthiness: As AI systems, particularly LLMs, gain traction in various applications, assessing their trustworthiness has become essential for ensuring ethical and safe deployment.
Additional Reflections
- Koyejo emphasizes that the understanding of model behaviors, especially regarding emergent capabilities, should be contextualized with the metrics chosen for evaluation.
- Future work is needed to explore how these metrics apply across different architectures and real-world applications, particularly in sensitive fields like healthcare and education.
Conclusion The episode provides deep insights into the current discussions around emergent behaviors in LLMs, the methodology for evaluating such behaviors, and the ongoing need for robust assessments of AI trustworthiness. Koyejo's research highlights the complexities involved in interpreting model performance and stresses the necessity of ethical considerations in AI development.
For more details, the complete show notes can be found at [twimlai.com/go/671](https://twimlai.com/go/671).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:08All right, everyone, welcome to another episode of the TwiML AI podcast. I am your host Sam Charrington and today I'm joined by Sami Kriyeju. Sami is an assistant professor at Stanford University. Before we get into today's conversation, be sure to take a moment right now to open up your Apple Podcasts or Spotify app. And if you're not already subscribed to the show, hit that follow button to make sure you won't miss any of the great content we've got planned for you. Sami, welcome back to the podcast. It's been a bit. Yeah, I think last time we spoke, I was at University of Illinois, so in the middle of the Midwest.
0:43I'd be much colder now, I think, than I am. And since then, I moved to the West Coast. Nice, nice. And you're enjoying it based on that smile on your face? Yeah, I'm also just a smiley person. But so far, so good. I think, you know, excellent environment, great colleagues, great students, great weather. Can't complain. You had a great year at this past NURPS with two of the papers that you worked on winning awards. One was an outstanding main track paper award for the, our emergent abilities of large language models, a mirage, as well as the award you won in the benchmark category, Decoding Trust, a Comprehensive Assessment of Trustworthiness in GPT Models.
1:32Before we dig into the specifics of those papers, maybe catch us up on kind of how you're thinking about your research agenda broadly and kind of what led you to focus on, you know, what are the threads that tie together these couple of projects? So my lab broadly is interested in trustworthy AI systems. We think both about foundations, so what does it mean? Think about measurement and assessments. What are the gaps in current systems? Then we think about mitigation. What can we do? And what are the tools that we can bring to bear to help close trustworthiness gaps in AI systems once deployed?
2:16Given that the couple of papers that we'll be focusing on are focused on language models and GPT and the like, that clearly has become a focus for you as well? Indeed, it's quite fascinating. So at least on the AI applied side, I think most of our effort was much more closely tied to vision a couple years ago. I think maybe as a consequence, both of the colleagues that I had around, things that people found most interesting, the problems that, at least the applied problems were animating a lot of the work. I think you're correct that none of us escaped what happened a couple years ago that I think changed the attention of a lot of the world, but definitely the AI world too, was thinking much more about language.
3:10and my students definitely got bit by the bug. And I think of maybe initially a bit kicking and screaming, but I think pulled me into thinking about language much more. I've found it fascinating. I think weirdly I avoided language as a grad student. I remember actually distinctly going to some of these shared cross lab meetings that had folks from Coromel like I was, but folks from language and other areas and not being able to make heads or tails of the language part of the work. I think at least in discussions at the time, they were fairly heavy with lingo within the area and just seemed so far from like stuff I was comfortable with.
3:56I just never bothered to really get into it. And I think for better or worse, everyone's paying attention now and hopefully we're not doing it too badly. I think folks that are either linguists or core NLP people sometimes complain about how the rest of machine learning engages in some of their research, but hopefully they would agree that some of our contributions are positively affecting language specifically. But again, we think of our role, both me and again, broadly the lab, as foundations. So we hope that many of what we pull out are things that apply broadly beyond specific either specific sort of modalities, but also perhaps beyond specific applications as well.
4:47But sometimes it's a bit of finding the right balance is part of the game, if you like. Sometimes going deep in particular areas, sometimes trying to pull out themes that seem broad and applied to many things. And I think so far so good and trying to find some balance between that. So, and I have folks, my students in my lab that are very deep in nitty gritty in specific areas in healthcare and language, as we've been discussing. I have others that like to be much more abstract. I think the synergy actually works quite well. So pretty happy so far with the combination. One of the papers that I mentioned, the specifically the, our emergent abilities of LLMs and Mirage challenges a core idea that we've kind of come to hold about LLMs or many have come to hold about LLMs, this idea of emergent properties or emergent abilities.
5:43And, you know, you really cast doubt on that idea with this paper. Talk a little bit about what set you and your co-authors down the path to explore this particular line of work. Yeah, that's been a fascinating, almost roller coaster paper for the group. So full credit to the student leads, Ryland Schaefer, who's the initial leader of the paper, and then Rando Miranda, who helped a ton with sort of fleshing things out and making paper what it is. So a lot of beginning of this started with an observation Ryland made when I'm reasonably certain he was taking an NLP class, So a natural language processing class, where part of the class included discussing foundation models and related kinds of tools and discussing this phenomenon that some of our colleagues had observed that suggested what they called emergence.
6:51this observation that on some tasks, some of these large language models seem to be bad when the scale of the model was small and somewhat unpredictably as you increase scale would get better. And so this was the key observation and this got the name emergence. And I think, you know, a few different potential tasks where this showed up, but I think the one that one of the ones that people initially gravitated strongly towards were settings where you're asking the model to add numbers together or do various kinds of simple arithmetic and various reasons why this was interesting I think one interesting one reason why there was a lot of interest in this application was that the there's a hypothesis that the ability to do arithmetic is tied to some aspects of reasoning.
7:53And so this idea that models could do reasoning is, again, quite fundamental to the potential of, in the future, having models, both on the positive side, implement or run more important aspects or just do more important things. and the maybe more concerning side, potential adverse outcomes, if these models get good in a way that people don't have a handle on. And so this is some of the discussion in the ether, in the environment and broadly within the space, because again of this observation of what seems to be this observation of this, what seems to be an abrupt change. I think most importantly, unpredictable abrupt change.
8:41I think that was the sort of, at least to us, the aspect that people seem most concerned with, that you don't know what scale the model seems to get better. So that sort of unpredictability seemed most concerning. Sounds like these two components that you've identified, the sharpness of the change and the unpredictability of the change, have become to some degree an accepted definition of emergent abilities. I don't know that there is an accepted definition. It's one of those tricks of language, which I think is often a challenge in any technical field that has to, I think this is true for, you know, statistics and lots of science where terms go back and forth between a precise technical definition and a sort of colloquial language definition.
9:31And sort of, it takes some time for, particularly when it's a new thing, it takes a bit of time for maybe the field to settle on a specific definition. We took one that seemed sort of precise based on one of the papers. That, again, was, as you said, so changes in model behavior with scale, model performance with scale that are unpredictable in terms of when they happen and have this sharp behavior. So what feels seems like a phase change in maybe physics terms kind of behavior in terms of performance from sort of zero to suddenly the model being extremely competent to some kind of task. Yeah.
10:11What resonated really strongly for me with that definition is how closely it correlates to some of the graphs that you use to illustrate the idea of so-called emergence, I guess we can say. You know, for example, arithmetic ability with respect to scale. Like, you know, it's just an exponential graph. There's a knee in the curve somewhere. And the idea is that, you know, that knee is sharp and we don't know where that knee is going to happen for new problems. Indeed. So maybe just digging a bit into some of the methodology and some of what we tried to pull out. So I think, you know, Ryland's initial observation was that this property, this what was called emergence, seemed to happen for only certain kinds of metrics and not others.
11:02And I think this was what led to, again, all of our thinking along these lines in the work, because at least initially our thinking was if this was something fundamental, the specific way you measure it should have maybe at least a limited influence. I should say I spent a huge chunk of my career talking about metrics and trying to convince my colleagues to spend more time thinking more carefully about metrics. I would say, you know, none of the other things I've done have had nearly as much impact as what we thought initially was just interesting side paper. I don't know that we thought that it wouldn't be folks who find it nearly as interesting as they did.
11:44But anyways, you know, this is something that I think I've tried to hammer, you know, like highlight all through my work. because there are lots of examples all through machine learning where we sometimes are just too lazy about thinking that metrics are sort of special given and not given. When I say this, I just want to be careful about word choice here. But most often what happens is whoever writes the first paper on a topic picks some measure and then pretty much everyone else just uses the same thing. And I've argued for a long time that this can be worrying because the measure that you pick in this, I would argue, often arbitrary way need not correlate to anything that you care about meaningfully in the world.
12:27Oh, that was the focus of our entire last conversation was we've got these kind of abstracted academic metrics and how do you then apply them to use cases or business oriented scenarios? Exactly. So that's a theme that has been a thread through a lot of my work. um here i think what was interesting and again rylan first observed this was that um the difference between the metrics that seem to correlate with emergent behavior and the metrics that didn't seem to be uh we haven't found a good name for this we sometimes call this sharpness or harshness but it's the idea that the mod the metric doesn't give partial credit so it's a measure that uh sort of you get full credit if you get the answer right and no credit if you get anything wrong.
13:22Kind of an all or nothing credit assignment issue. Indeed. Okay. That seems to be at least one of the key characteristics of metrics that seem to have, again, the results being what might be interpreted as, again, an emergent behavior and some performance. Can you give an example of, I guess arithmetic would be an example of you to get the problem right or wrong. Yes. Arithmetic is a great example. Did I get the answer right? Exactly. Which arguably, I think very reasonably, many have countered and said, well, I don't care if I get partial credit in arithmetic question. question. I think this may be reasonable, but again, I think what hopefully our paper is pointing out is one, I think the specific choice of metric plays a big role.
14:17And whatever inferences you make should be made in light of the fact that you also had control of your measure. And sort of questions of fundamental fundamentalness of some kind of behavior of models, I think should be at least considered carefully. So it should not be sort of trivially stated broadly without context. I think that's, I think I hear you saying is that you require a certain level of skill to do arithmetic, but that doesn't necessarily mean it's an, it's a magical emergent property. Yeah, I'd be, I'd be comfortable saying that. Yeah. And I'll, I'll say a bit more because I think probably the most interesting, one of the more interesting bits for me is, I mean, both the fact that you have this sort of sharp behavior, I think is an interesting thing.
15:07The unpredictability of it is also interesting. And I think there I'm, I think we can make, I can make a somewhat stronger argument to counter at least some of the sort of discussion happening. um but on the again to your point this this idea of uh arithmetic and sort of yes no correct not correct uh in our evaluation of models i guess it's one thing for maybe deployment settings but i'd argue for many things that we care about uh partial credit actually plays a big role both in terms of calibrating our understanding of what's happening uh but also in just in terms of understanding, like any kind of understanding of what the models are doing.
15:53Having measures that can capture this partial or gradual evolution, I think is quite meaningful. And so maybe that's another inference of just, again, the effects of metric choice on interpretation of some set of results. Maybe a slightly stronger claim is the whole argument about unpredictability. So So we posited a very simple toy model, admittedly toy, in that it ignores a lot of what we think language models are doing and simply considers language models as next word predictors. So next token predictors. So we take the actual training objective at face value. And more than that, in addition, we just ignore even correlations.
16:41so we just think about this as i'm going to do a sort of uh you know while i'm predicting a sequence excuse me i'm predicting a sequence i'm going to ignore the fact that it's a sequence just think about this as independent decisions on what the next token should be so i have a very simple model uh if you like a coin toss model uh where the coin toss has a parameter that says how likely am i to get the answer right if this parameter this probability is very big i'm likely to get the answer right for any coin toss. And I'm going to do this coin toss in a sequence. And for now, I'm just going to ignore the fact that in language, the sequence of tokens that predict are not independent.
17:20The whole point of, for instance, recurrent models is that the history affects the current prediction, affects the future. So we'll ignore that and just think about this as sort of coin toss. It turns out this somewhat abstract simplification is sufficient to capture a lot of qualitative properties about what metrics might do. So in particular, you can characterize this very simple model quite precisely. So it's a simple model. You can sort of do a bunch of math with it. And in particular, you can work out what you expect to happen as I change the probability of success. So if I change the probability of getting an answer right, what do I expect to happen if I'm measuring the number of times I get a sequence correct?
18:05So this allows us to reframe the addition question as, you know, how likely am I to get four or five numbers correct in a row? So I'm ignoring the fact that I'm actually doing addition. I'm just thinking about this. I'm going to pick the next token. Will I get the next token right or wrong? And we have some sense of what we expect models to do because existing work on scaling laws, which are these empirical observations of what cross entropy looks like as I increase models gives us a handle on roughly what's the average probably that you get the next token right so that's something like what cross entropy is measuring in other words the model size correlates directly to the p in your model indeed yes so so this is somewhat it's a sort of empirical finding that has held up quite well, enough that I think the field is comfortable saying that there is this association between how big models are and then this P, if you like, how likely you are to get the next token right, which roughly cross entropy is measuring.
19:15And so we take this, we just pull out that number, and then we ask if I modulate this number from small to big, what should I expect to see if I'm checking not just one token correct, but K tokens in a row correct? So let's say five tokens in were correct uh and so this is actually it's a kind of question i don't say this i want to be careful here um it's a question i would give to like a probability class as homework uh but i don't say this in a dismissive way to any of the claims that i've made i'm mostly again this is a simple model but uh it's simple but it's sort of simple enough to illustrate the point i think quite clearly but uh it's a question i might ask to a probability class you know how likely am i to get five answers in a row correct if I have some sort of known probability of getting one answer correct.
20:07So roughly, if I'm going to make a mistake, well, if I make a mistake in the first one, then I'm done. I might get the first one correct or I make a mistake in the second one, then I'm done. So you can work this out by looking at sort of the sequence of predictions and just simply use that simple single parameter of how likely you are to get the next token correct based on a simple fixed probability model. And if you do this, you see a curve that looks notoriously like the emergence curve. So this overall accuracy measure changes from close to zero to sort of close to one. It's sort of perfectly predicted by, again, very simple probability math.
20:44And I think more importantly, our results, while on this very simple toy model, match or sort of predict the observed data quite well. So if we look at problems. And if you look at the sort of this, the number of tokens I need to get right to get, you know, five plus five, say four digit addition, right? So to get four digit addition, right, you generally need to get five tokens, correct, because four digits, and then often there'll be a sort of, if the numbers are big enough, there'll be something that will roll over. And I, if I just plot with this very simple model, assuming there's some fixed number that tells me how likely I am get the first token right, second token right, third token right, fourth token right, and so forth, I plot out what should I expect to see in terms of overall probability that I get all five tokens correct.
21:35And if I plot that out, that has the same kind of S shape as you see in the emergence curve, very close to zero, because any mistake, when I have low probability of getting things right, you get hurt by any of the tokens being wrong. At some point, the probabilities get high enough that you're likely to get all of them right, and then it suddenly shoots up. And so it sort of looks qualitatively pretty much the same. And so we use this to argue, again, that the predictability, I think, or the lack of predictability may also sort of be not a stronger claim. Because, again, with sort of a very simple cartoon model, we argue that we have sort of quite good predictability for things like addition.
22:21So for metrics that are tied to, again, getting everything right, we can use behavior that we expect from predicting sequences that sort of are very well understood to predict what we expect to happen on the overall model. And these findings seem quite consistent. So this challenge, the predictability claim, again, arguing that at least for many of the existing properties that people have looked at, so existing things that we've tried to build with models, we argue that you can predict them quite well. um one slightly tongue-in-cheek experiment we do but i think meaningful again not to be again highlighting this is not dismissive i think the whole way that this happens in the field completely makes sense uh yes okay yes yes exactly this so again we just really wanted to hammer the point home that metric choice matters so much and we did this experiment as you said division metric where we looked at a problem where no one has ever claimed emergence and we changed the metric with a very similar style and sort of show a curve that matches emergence so we look at auto encoders which are these models as many of you listeners might know that are useful for sort of if you like summarizing data so they're they find the models that you use to find low dimensional embeddings of data such that you can compress large data sets instead of low dimensional space.
23:59So they're useful for a bunch of things. They're good for denoising, they're good for compression. So, you know, a big part of the field is building good autoencorders. And they're also big within models. People use autoencorders inside much larger models that do fancy stuff. And so we looked at autoencoders. They're usually measured by some metric that says how close is my predicted image to the original image. So how well did I reconstruct a particular image? And this is generally something like a squared loss or something in that family, some kind of norm. and we replace this with a measure that gives you full credit if you're within a small value close to the perfect reconstruction and then zero credit if you're far away so instead of again seeing this sort of partial the model getting better slowly you're seeing models far away no credit and at some point it gets better and then full credit and we show again this lo and behold, what did you find?
25:04Exactly. One more thing I want to mention. So we did one more experiment where we, we argued based on the very simple model I just mentioned, just using very simple, sort of independent probability of getting token rights, getting tokens, right? One of the inferences from that model is not only can I predict when phase transition might happen, but also I don't expect zero in the part where the models are not performing well. I actually expect to see a little bit of signal because even though you know again the idea is on the x-axis if you can visualize the curve on the x-axis i have sort of how likely am i to get the next token correct so this goes from zero to one if i'm one i'm perfectly going to predict next token if i'm zero i'm going to get the next token completely wrong uh and on the y-axis i'm predicting how likely am i to get five tokens in a row correct right so the x-axis controllable that we can go from what happens if it's zero, we never get a token correct.
26:02What happens if it's one, we get all the tokens correct. And the y-axis is this model of how likely am I to get k sequences, so k tokens in a row. So some kind of sequence of tokens correct. So this is capturing, again, this geometric is actually a technical term. You get a distribution that captures what you expect to see in terms of probability of error. And one of the inferences from this simple model is that you don't expect to see zero. uh at the low probabilities you expect to also see sort of improvements uh so so the curve you should be expecting is not uh s curve with flat before the emergence point if you like and then a sharp and then a flat after but really sort of uh monotonically getting better uh from zero all the way through where it shifts and then getting better again after it shifts uh because even in this metric that has harshness in it and doesn't give partial credit.
27:00The nature of the sampling is such that as you get more likely to get some of the tokens right, you will get a little bit of partial credit in the aggregate anyway. It's just that it's most dramatic at some transition point. Is the idea that in these instances where we're claiming emergence, the actual graph looks like this geometric distribution? If we take, I guess, two claims as true. One, many of the claims of emergence are tied to full credit instead of partial credit. Two, the mechanisms underneath this are often tied to sequence of tokens being correct all at once as opposed to a single token potentially being correct.
27:46And you don't get any partial credit till you get the full sequence correct. So those are maybe the two bases of the claim. And again, both of those tie quite well to what people are doing. The caveat is that the measurements themselves are different and you're making your claims around kind of a toy representative measurement model as opposed to the one that's typically being discussed and charted. Yes, yes. But this is very strongly on purpose. I guess I want to highlight that. And actually, I'm a big fan of this kind of research where you find little toy sort of illustrative models, if you like, that capture qualitative properties that match real behavior at scale.
28:27And that's actually the point of this. So I fully correct that this is not the actual model. This is sort of a cartoon, but it's a sort of hopefully thoughtful cartoon designed on purpose to capture scale up behavior. And in particular, this cartoon model suggests some inferences that we should make that we then validate in the real models. So again, the first inference was, can we predict when the transition will happen? And again, we show that this matches reality quite well. The second inference was, you shouldn't expect zero in the parts where the models aren't doing well. You should expect still monotonically increasing performance, but increasing much more slowly.
29:10And we show this. So the argument we make, again, this falls on statistics, we show that statistically to get the right signal at the low parts of the measurement space, you need to sample many more models. And the reason is, each model, the probability that it gets it correct is low. To actually get a good estimate of a small value, you need to sample things much more often. And we argue that part of the reason people don't see any signal for small models is that they're not sampling sufficiently because again an inference from this qualitative model is that if you sample sufficiently you should see signal it's just that the signal is not dramatic till you get to whatever the inflection point is and so we show this in the paper in particular we we build small models where we sample much more or start with small models and sample much more often and we show that again the qualitative behavior of these curves are not zero at the beginning some sort of transition point and then almost one in a flat way but as predicted by again this very simple model you get monotonically increasing performance in the beginning uh some sharp transition that matches when you expect uh sort of individual probability of getting things correct to um to end up having this if you like uh transition behavior to getting sequences correct much more often uh and so mostly the key idea is, and you might see this in the paper if you're looking at it, that the theoretically predicted curves match reality quite well.
30:45Are you predicting the transition point in terms of P in your simple model, or are you able to work back to some number of parameters in a representative actual LLM and predict the number of parameters in which the transition will occur? The model It still predicts it in terms of P, but we can go from P to parameters because the scaling laws tells us sort of a rough estimate of how parameter size ties to P. That's the whole, the point of the scaling law is, and the x-axis model size and the y-axis sort of cross entropy, which is reasonable. You can pull from that proxy for what this P value is, how likely you have to get the next token right.
31:28So there's a chain of argument that allows you to go back to parameter size for expected transition. Okay. And is that sufficiently abstracted that it's task or use case independent? This is a great question, actually, because I think it's one that requires a bit more work than we've put in. It's sort of some of what we're thinking about now in terms of how this relates to different tasks and what might happen in different settings. In particular, the arithmetic ones are, I think, the ones that were most obvious to play with. And some of the other settings that people are interested in are a bit more complicated in how token behavior might interact with the model architecture.
32:14So things like in-context noting, for instance, I think is a little bit trickier to pull out exactly how we might explain things in terms of our simple model. That said, so what the scaling law does is, I guess it's task agnostic in the sense, at least the original scaling law papers. They're task agnostic in the sense that they just measure how well the model fits data. And going from that to how well the model does in particular tasks is, again, a sort of a separate inference. The observation has been the data fit seems to be highly predictive of task performance. So sort of the better the model is at fitting really large data sets, the better it tends to do in terms of performance on various tasks.
33:03That said, there is sort of some gap there to fill that ends up being filled in an application-specific way in terms of how the scaling law performance prediction ends up tying to how well the model might do on a particular task. And I think you're correct in saying, you know, we partially, part of our argument allows to partially like close the full loop on some of the particular tasks. Most, I think, obviously arithmetic, but maybe not on some others. So that leaves some work to do. Maybe keeps my students so excited on some actual work on fleshing out the full chain for other kinds of tasks and what maybe the right conceptual framings of the model is.
33:48But hopefully again, and a big takeaway is being, uh, one I like is the ability of very simple models to elucidate aspects of sort of large scale model behavior, uh, but maybe more precisely, the importance of the specific choice of measure and how it can affect inferences on what we think models are doing. The choice of metric is not a sort of free choice. Like it affects everything else, uh, that you're doing and it affects the way that the way that you're sort of the inference that you make about the outcome of models you draw correlations between linear metrics like token error rate and some of the nonlinear or discontinuous metrics that researchers use how should a researcher think about metric choice in light of the results of this paper?
Read the full transcript
34:46Any claims being made should be sort of correctly made in context of the metric choice. So like acknowledge the sort of effect metric choice might have on any kind of large claims being made. I think that's meaningful to do and sometimes hard to do because I think announcements and discussions are often made context-free. and the context matters and that's, you know, for instance, if the claim was, you know, we see emergence, but, you know, for claims that give partial credit, we expect to see sharper behavior and metrics that it's longer, but I think it's more correct. And, you know, allows again, for the listener or the reader to sort of correctly interpret the strength of what is being observed.
35:32So it ties to a question of fundamentalness, maybe. And I don't know the full answer to this question, because I mostly studied this question from the perspective of pick the metric for the thing that you care about in the real world and then sort of adjust the rest of your stack accordingly. And we're in settings where these metrics are not necessarily optimized for it, they're just more post-evaluated. And then inferences are made about behavior of the world based on them. I think here, again, the main take, or at least the obvious takeaway is just care and interpretation from sort of these choice metrics.
36:11The fundamentalness question is an interesting one. So, you know, if I can always, there's a hypothesis here, which I don't have, I'll state as a hypothesis because I don't think we have complete evidence for. But the hypothesis is that I can always find a metric that will sort of turn what seems like an emergent property to a non-emergent one. Let's suppose that is true. I don't know if this is true or not. But suppose that this is true. How should one interpret emergentness, if you like? So does that mean emergentness is fundamental or not? It would imply that emergence is a very fuzzy and metric specific label as opposed to something fundamental.
36:59I think everything is measurement specific. And like we try too hard, I think, incorrectly often to sort of overgeneralize statements and claims. And like most things are dependent on how you choose to measure them and sort of the inference, the sub aspect of the property that you're trying to pull out. So, you know, should we worry exactly how much do we care about emergentness? I think some camps do. I guess the field now has camps, but anyways, some camps do quite a bit for various reasons. If we do, then, you know, then do we care about how fundamental they may or may not be? I think that's a question the field has to figure out.
37:44And, you know, some of my current research is trying to understand better. Again, like I said earlier, the sort of full chain thing, all the way from specific tasks, overall model behavior, to what things may be fundamental or not. Another thing, an easier claim to make is for any problem, I can probably find a metric that will give me something that looks like emergence. This is the opposite direction. This is not saying I can always find a non-emergent metric, but the opposite claim I can make more easily because we have an example in the paper. And I feel reasonably confident I can do this for any metric that ever gets better.
38:22I can find a way to reconstruct it in such a way that it will get better in a way that looks sharp and sort of unpredictable to a lay eye. And so, again, I think it forces some reconsideration of claims made without context. At least, at the very least, we hope that this is what this work is doing. And forces, again, consideration of the metric that you choose and what inference. I'm repeating somewhat, but at least for the folks building models and explaining them to the general public, the way the general public ingests these things, the way policy is made around these things, context freeness can be worrying, almost dangerous.
39:00You want to be careful about adding context appropriately because all of this, for better or worse, it all matters and can't be removed from how we think about what models are doing. Let's shift gears. And unfortunately, I don't think we'll be able to do the second paper justice due to time. But again, your Decoding Trust, a Comprehensive Assessment of Trustworthiness in GPT Models was another award winner in the benchmark category. Unpack that paper and what it has attempted to do. So this was work with lots of awesome people. The lead author is Bo Xun Wang at Illinois. The sort of lead senior person, if you like, or at least their advisor is Bo Lee, who's now at UChicago.
39:51But this brought together folks from a bunch of different places. Me and some of my students out of Stanford, a few Illinois people, some Berkeley people, some Microsoft Research people. And our goal was to, yes, this is this big sort of consortium, but our goal was to try to take some first steps in trying to ground a little bit what trustworthiness might mean in the language model world and think a bit about how we might build evaluations for aspects of trustworthiness that we might think are important. And so we landed on eight perspectives. So these included toxicity, stereotype bias, robustness of different kinds, obviously robustness, distribution robustness, privacy, ethics, and fairness.
40:46And for each of these, we came up with some evaluation. So some way to query models, these were all, all of these were designed to be query style. So some way to query models and get a response and then evaluate this and to give us some signal about how well the model might be doing with respect to various trustworthiness measures. So we applied these to a bunch of different models. So, I mean, the main task was designing this evaluation and figuring out how to make sure things are scalable. For some of the directions, we had to come up with some new metrics. For some others, we extended existing metrics to sort of be a bit more flexible and capture, have better coverage of various aspects of model behavior.
41:32And then in the end, we have this toolbox that has a GitHub page, also on Hugging Face, so people can easily download and run the evaluation. And it's something that we're hoping and looking to build up quite a bit, actually, as hopefully in the long term have a big suite of evaluations for model behavior that build on existing, I think excellent work by some of our colleagues on evaluation for performance. So as you know, there's a bunch of work on performance evaluations. I think very notably, for instance, my colleague, Percy Liang, and colleagues have Helm out of Stanford, which is sort of widely used for people thinking about evaluating performance of language models.
42:18But we noticed the gap that we, the group that was important and directions that are related to trustworthiness that we wanted to try to help fill. Would you characterize the contributions of this particular work as primarily pulling everything under one umbrella or advancing the particular approach to measurements around specific concerns or filling in gaps where there were no accepted measurements? Or is it kind of all of the above in different places? It's kind of all of the above. So in some ways, there's good and bad to it. I mean, as an academic paper, it was a bit, it was extremely unusual because at least the initial draft was almost 100 pages because we had to cover so much.
43:13As you said, again, we had this mix of pulling things that already existed with some new design, some sort of cases where we're just extending existing perspectives. I think some of the robustness things look like that. Cases where there wasn't a definition out there and we were using it. We defined one and defined some tests that could evaluate some direction. So it was all of the above. It was a bit of a, it was a lot to read. I think the final nearest version has sort of a, we have a eight page summary and then like 80 pages of appendix. So it covers quite a bit. Again, just to, just to lay, to be able to lay this whole thing out.
44:01There's a lot going on. And then beyond that, having a software package that again, people can download and use to actually evaluate models that they're, they are building. I think is an important contribution of this kind of work. So both the hopefully intellectual work of pulling all these together and also the sort of hopefully practical, useful aspects of having a toolbox that can be built on and deployed by folks. I think some of the interesting observations, we focused a little bit in original writing on GPT models. So 3.5 and 4. at the time we evaluated the most. Mostly the observations were positive in a sense that models do seem to be getting better, meaning there were meaningful improvements in many of the evaluations that we did from 3.5 to 4 in almost all the directions.
45:03One interesting one, and I think this points to a fascinating tension in trying to build general purpose models that follow human behavior, is that some of the perspectives relied on the model not following instruction when the instruction was thought to be potentially harmful. So you can imagine, for instance, ethical settings where you'd want the model like the right answer is if you try to prompt the model to do to say something or implement something unethical or give you a particular response that's unethical. the right answer at least in our collective reasoning about what we want for the models is that they don't do this uh but if the model is good at following instructions it does what you ask it to do uh and this leads to a weird conflict and that uh in particular we found some of the directions that we covered gpt4 looked worse but we think it looked worse less because the model is necessarily sort of, you know, less tuned, but looks worse because it follows your instructions a little bit better.
46:12So it's an interesting conundrum. It's not exactly clear exactly how, I think there's some space for societal and also folks building these tools on how we interpret and sort of evaluate what cases we want instruction following to be exactly what model does and what cases we want uh we have an ethical standard we want to hold such that or some other standard we want to hold such that we think the model not following your instruction is a good thing yeah so this shows up in fairness to some extent as well you know in settings where the model can have biased decision making you know the cidal standard we believe and what we're building to is that the model should not and so we want to check that isn't doing that but you know Will the model, if you try to ask it hard enough, actually do that and maybe look worse in our evaluation?
47:02What's interesting there is that your observations are a bit of the opposite of what I've tended to hear is that many in the community feel like, to some degree, RLHF and to some degree, instruction tuning of like hampered GPT-4's ability so that it clearly performs better, but it's a lot more limited in what it will respond to, and you need to coerce it in ways that make it more difficult to use at times. So I would agree with that characterization. I'll say two things that maybe are interesting. Both may be, I don't know, general statements, but I think hopefully meaningful. One of them is, I think this particular conversation points to a challenge in scientific evaluation of black box models and that behind the scenes, the models are changing all the time.
48:06So the version we evaluated was a little bit earlier in the stack than I think some of what people have been complaining about recently. Got it. I don't think we have any particular behind the scenes versioning, so we don't actually know. we know when we evaluated and it's in the paper sort of the timeline that we did the work but to your point the way the ecosystem has developed is such that there's no public information about which model you're talking to and sort of all the changes and tuning are happening behind the scenes and so it challenges the science a little bit and sort of the the strength of a claim like the one I just made.
48:48But I can say that for at the time we evaluated, this was true. I'd also say that despite maybe this comment that you're making, I think it's also it's still true, I think, at least in my this is anecdotal, maybe, but maybe still true that four is still better than 3.5 at following instructions, even if it's worse at following instructions than four from six months ago, perhaps. So I think the sort of at least the inferential claim from the paper of instruction following being better in as models advance, I think is still true, even though within the same model family, I think, you know, I've heard the same observations from people.
49:37The effect of the tuning has been that it follows instructions less, in some cases where you would have imagined it followed more. I don't know if you know James O there at Stanford, but he did some work looking at the effective or GPT performance over time. And I recall one of the questions that I asked him is, you know, if he envisions establishing like some kind of GPT weather report that just like is static or static dynamic, consistently applying some metric and reporting performance. And, you know, I put the same question to you. You've got a toolbox, you know, can you host this somewhere and just have it running against current GPT and seeing how it changes?
50:22Yeah, no, this is a fascinating question. Maybe I'll, could it be done? Probably. I think if we did this, it'd be interesting, at least from the trustworthiness perspectives. Again, our toolbox and our evaluation is very coupled to those perspectives as opposed to the standard ones. One interesting wrinkle, and my throwback, is my guess is that they're doing A-B tests behind the scenes all the time, and they're not giving everybody the same model. And I don't know this for sure. This is a speculation, but I guess, you know, it's not clear how strong the inference might be from our particular instance of the model, because I guess it's very possible we're just seeing something different than some other subpopulation.
51:10It'll still be interesting to at least track, even if it's in this, you know, maybe not fully generalizable way, but track what's happening across metrics. That sounds like a sampling problem. Perhaps, yeah. Yeah, you could do it with more population and hopefully you get all sides of things. Sure, sure, indeed. Yeah, if someone had some compute cycles and was very interested in doing this. I don't know that we will do it, but it's interesting. What scale of resources are required to run the full benchmark? I mean, it's definitely much, much less than either model inference or sort of training for sure.
51:53It's mostly... How can it be less than model inference? Isn't it model inference? Inference from the sense that for many of these models, the inference is external. So like GPT, for instance, we're not hosting a GPT instance. We're querying a model externally. So you're not paying... You're paying OpenAI to get tokens back. So someone is bearing the cost for sure, just maybe not the person running the evaluation. to what degree do you feel like the tasks and or metrics in this benchmark are easily applicable to real world problems? Like, have you replicated the problem that you're trying to solve in other aspects of your research?
52:38No, this is a good question. And it's great to be self reflective in this way. And maybe, you know, the practical implication of that is, you know, if I'm solving a real problem, should I be relying on a benchmark like this? Or should I be investing in creating my own benchmarks, which uniquely represent the issues and the problem I'm trying to solve? Maybe I'll reframe your question slightly. Because a version of this question I've been thinking about a lot recently is sort of what level of general purposeness is the right level of abstraction to be thinking about evaluation for models. So a lot of the evaluation stack thinks about models mostly user-facing, and the user can ask anything, and thinks about what we might evaluate.
53:28So there's a coverage problem, maybe, but it's also a tricky one because there's lots of flexibility in the potential interaction. And I'd argue, perhaps, that both a lot of the existing work, but also even this decoding trust work is targeted at that general use case in that way. And so to the extent, you know, ignoring coverage for a second, because I think that's a meaningful caveat, I'd argue that all the instances of tests that we come up with are sort of interactions that either there's evidence for, like, you know, simple thought experiments, Like it's not a stretch to have people interacting in the model for various kinds of say decision making settings or querying for some outcome where we want to be careful about bias in various ways or security leakage, for instance, like, you know, will it replicate, will it spit out PII information if it's given in context of its part of training data?
54:28That's one of the tests that we run. So in these directions, I would argue that our tests are relevant for real-world use. But I think there's a meaningful coverage problem that is just going to be one that the field will have to take. Either we'll come up with some new innovative method, or we'll have to go with just increasing the number and size of tests that we run to get better and better coverage. That's one direction. The other direction, because I've been talking quite a bit to various domains, some most importantly and recently to folks in healthcare who are either thinking of deploying or already deploying language models, either clinician facing or sometimes patient facing.
55:14So clinician decision support or triage sometimes or patient facing to maybe do some early screening. So for better or worse, these things are here and there's a lot of interest in potentially using them. And there's a question of how can we tell whether we think these models are ready for prime time in these settings. And there, I think where we're landing on is similarly for education, by the way. And there, I think where we're landing on is that the domain seems to matter quite a bit. and like the domain, the specific concerns that the domain has may or may not be covered by the general purpose evaluations.
55:59For something like healthcare, I think something very specific, actually both healthcare and education, the notion of harm, I think often actually has much clearer resolution. I think harm can be vaguer in the general purpose sense. in healthcare if I have a recommendation for a treatment and I have sort of known drug interactions that are fatal or I have someone I'm recommending something to and the actual recommendation will kill the person like the harm there is just very clear and obvious and I think that that clarifies at least at some level clarifies gaps and suggests some evaluation mechanisms that you may not be thinking about as clearly if you're thinking about general purpose use because you'd imagine people are not using it for medical decision making settings similar in education again where education is an interesting one particularly if it's k-12 because so specific rules and societal norms about what we want to expose like the interactions children should be having with technology or even if it's a teacher but the teacher is showing the results to a class like what's acceptable not acceptable and sometimes this varies uh at the very least varies sort of by states based on the regulations often by school sometimes by individual uh in terms of what the way they want to be and so thinking about uh exactly how to build evaluations that can capture all these I think it's been tricky Where I hope we will go, and I think some of what we're building and working towards is I'm optimistic of a path where we can get good at sort of having a very large suite of possible evaluation tests and then be able to pull which subsets of these are relevant for a particular context.
58:04A context might be, again, particular domains. I mean, I say these as healthcare. Healthcare itself is huge, right? There's specific specialties. They all have specific interests. So cardiology and radiology care about often different things. And in fact, I mean, as an aside, often issues with health care often come from the fact that they're competing interests for the different specialties and what they think is most important in your care. And so, I mean, that's an aside. Now they're not there, but it's at least for this conversation. But it brings up maybe challenging questions on how I might think about evaluation and at what level our abstraction or what level of use, if you like.
58:43So general domains, specialties, individual hospitals, individual doctors or clinicians. Purchase on this problem. I'm optimistic of the route of we can build enough, a large enough suite of tests that we could sort of the work becomes what to pull for specific context. I think there's some interesting work and maybe personalized, personalization said broadly, but like specialization, if you like, of tests. So are there ways to semi-automatically figure out what subsets of a large suite of tests are most relevant to a particular individual, maybe tied to their values is something with some other colleagues have been playing around a little bit and understanding a little bit better how people's values affect how they interface with technology.
59:30And so maybe there's some space to be able to capture some of that and how we think about tests at different levels of abstraction. So maybe, you know, I went off on a slight tangent, but hopefully I think captures some of the spirit of your question. So it's a hard one. And as a nominal general statement, yes, I think we're building tests that we think are relevant for practice. As a specific statement, I think there's still some gap in what practice means and different resolutions of what that might mean and how we can build tests. I guess I'm hopeful for the route of being able to do this in a way that can be extendable and scalable.
1:00:09Granted that these are black box evaluations, I'm curious, do you see any issue extending them across architectures? And I have a specific point in mind, and that is kind of bare LLM versus a retrieval augmented scenario. Would you apply the same metrics approach equally? Yeah. So a quick aside, I think one of the interesting tangents that we've been also considering is like you've alluded to, we say LLMs like, I guess, a model, but really, they're mostly systems. They often have stuff around them. I think as you're talking about, you know, retrieval argumentative generation is now common. Tool use of various kinds is common.
1:01:05guardrails exactly yes so stuff around are common and these change qualitatively the behavior of models for the for these tests so the tangent i was going to go on is i think there's also space to think about how to evaluate all of these tools separately which i think could be interesting because it allows for maybe some interesting i could imagine the use case of being able to mix a match, for instance. In a black box setting or assuming you had access to the tools independently? Could I do, for instance, good evaluation of a rag tool in a black box way? Sort of, kind of, there, maybe not black box, but there are existing evaluations of embedding, for instance, embedding quality that should correlate with how the rag tool ends up working in practice.
1:02:02probably full system tells you something that you may not get because you know system interactions can sometimes defy individual evaluations but I find that whole direction fascinating as I guess as an aside I would say at least for decoding trust and similar evaluations they're all purely prompt so in some sense they only care about input and output from the point of view of the test uh so like they're testing the system whatever the system is uh i think they you know because of this there's some we're talking about i was alluding to coverage earlier there's probably some coverage gaps uh because there are some parts of evaluation that are tricky to do without transparency i think there's some meaningful gaps in understanding for instance training data because we know that it affects uh models in particular ways there's a line of work that some of my other students are playing with.
1:02:58I actually feel very interesting on thinking about data contamination. So this is a big issue in evaluations and tests, as you might know, because two things. One, many of the tests have data sources that are effectively online because they're open tests, sort of by design, so people can use them. But again, this is not a statement or an accusation, but it's just an observation. If I'm building a model, the obvious thing for me to do is to train my model on the test because I'll look really good with respect to that test data. And so the incentives are sort of set up in that way. I'm not, again, I have no evidence that anyone is doing anything.
1:03:38I'm just observing that this is a problem. And so there's a gap there in thinking about how we might, one, you know, to what extent can one build this type of the general comment about black box evaluations and their gaps. In particular, this training data gap is one that is salient for me. There's also, I think, meaningful gaps in being able to examine weights and say meaningful things about them. But separately, I guess if one is willing to do black box, at least these kinds of tools that we're building, I think, can sort of capture system behavior. Again, caveated with this is black box and we don't know exactly how things are trained and these other reasonable caveats on what's going on.
1:04:28Well, Semi, thanks so much for taking the time to dig into these papers with us. Congrats. Yeah, it was a pleasure. Great questions. Thanks for prompting. Absolutely. Thank you very much. We appreciate that. And kudos to all my collaborators, colleagues, students. Almost always, I get some maybe undeserved credit, I think, for the folks doing most of the work, the students that really push these things forward. and they're incredible. Enjoy working on them. That's fantastic. Awesome. Thanks so much, Emmy.
From the publisher
Today we’re joined by Sanmi Koyejo, assistant professor at Stanford University, to continue our NeurIPS 2024 series. In our conversation, Sanmi discusses his two recent award-winning papers. First, we dive into his paper, “Are Emergent Abilities of Large Language Models a Mirage?”. We discuss the different ways LLMs are evaluated and the excitement surrounding their“emergent abilities” such as the ability to perform arithmetic Sanmi describes how evaluating model performance using nonlinear metrics can lead to the illusion that the model is rapidly gaining new capabilities, whereas linear metrics show smooth improvement as expected, casting doubt on the significance of emergence. We continue on to his next paper, “DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models,” discussing the methodology it describes for evaluating concerns such as the toxicity, privacy, fairness, and robustness of LLMs.
The complete show notes for this episode can be found at twimlai.com/go/671.




