In short
Episode Summary: ICLR 2024 — Best Papers & Talks (Benchmarks, Reasoning & Agents)
Podcast Title
Latent Space: The AI Engineer Podcast
Description Latent Space is a podcast tailored for AI Engineers, offering insights into cutting-edge developments in AI, including Foundation Models, Code Generation, and Benchmarking. The podcast features exclusive interviews and discussions with leading figures in the AI space.
Episode Overview In this episode, the hosts recap significant papers and discussions from the International Conference on Learning Representations (ICLR) 2024, focusing on benchmarks, reasoning, and agent systems. The episode is divided into sections featuring various guest speakers, including Graham Neubig, Aman Sanger, Moritz Hardt, and others, providing unique insights from academic and industry perspectives.
---
Key Sections
Section A: Code Edits and Sandboxes
Guests
Graham Neubig and Aman Sanger
- WebArena: A sandbox for testing language model agents in web environments.
- Sotopia: A project exploring social interactions through language models.
- Performance Improving Code Edits: Discussed how models can learn to make performance-enhancing changes to code.
- Discussion on Academia vs. Industry: Addressed the evolving relationship between academic research and industry applications.
Key Takeaways
- WebArena: Provides a realistic benchmark for evaluating agent performance in web navigation and tasks.
- Industry vs. Academia: Shared insights on their roles in developing AI technologies and the need for collaboration.
Section B: Benchmarks
Featured Papers
- SWEBench: A framework for evaluating language models on real-world software engineering problems.
- Key Insights: Only the simplest issues were resolved by leading models, indicating room for improvement.
- Benchmark Contamination Detection: A method for detecting if a model has been trained on a specific benchmark without access to training data.
- Key Insights: High duplication rates in training data allow for detectable contamination.
- GAIA Benchmark: Proposed a benchmark for evaluating general AI assistants on multi-modal handling.
- Key Insights: GAIA focuses on tasks requiring reasoning and tool use, emphasizing the gap between human and AI performance.
General Observations
- The effectiveness of benchmarks varies, with new benchmarks emerging to assess the latest developments in AI capabilities.
- The need for comprehensive evaluation methods and frameworks to understand model performance in complex tasks.
Section C: Reasoning and Post-Training
Featured Papers
- Self-RAG: A framework for improving reasoning in language models through reflective retrieval and generation.
- Key Insights: Emphasizes the need for reliable retrieval to enhance language model outputs.
- Let's Verify Step By Step: A study showing the effectiveness of process supervision over outcome supervision in training models for reasoning tasks.
- Key Insights: Models trained with process supervision yielded better performance in complex reasoning tasks.
- Lilian Weng's Safety Framework: Overview of OpenAI's safety systems to ensure models operate within safe and beneficial boundaries.
General Observations
- Safety and reliability in AI models are critical, necessitating ongoing evaluation and adaptation of models through real-world interactions.
- The importance of integrating verification mechanisms into the AI development process.
Section D: Agent Systems
Featured Papers
- WebAgent from Google DeepMind: An LLM-driven agent designed for real-world web navigation and task completion.
- Key Insights: Demonstrated improved performance through long-context understanding and program synthesis.
- MetaGPT: A multi-agent collaborative framework designed to optimize workflows and reduce errors in software development.
- Key Insights: Highlights the division of labor among agents and the impact on efficiency and effectiveness in software projects.
General Observations
- Multi-agent systems show promise in automating complex tasks but raise concerns regarding computational costs and latency.
- The trade-off between performance and complexity in multi-agent systems remains a vital consideration for engineers.
---
Conclusion The episode provided a comprehensive overview of the best papers and discussions from ICLR 2024, emphasizing the advancements in benchmarks, reasoning, and agent systems. The insights shared by the guests highlight the collaborative efforts required to drive AI research and the importance of continuous evaluation and adaptation in developing effective AI systems.
Further Reading & Resources
- Full show notes available at [Latent Space](https://latent.space)
- Suggested papers and works mentioned in the episode.
Calls to Action
- Engage with ongoing research and contribute to open-source AI projects.
- Stay informed about the latest developments in AI by subscribing to the podcast.
Thank you for tuning in!
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:05Welcome to the Latent Space Podcast, ICLR Edition Part 2. This is Charlie, your AI co-host. We're back with our coverage of the 12th International Conference on Learning Representations in Vienna, Austria. Many of you absolutely loved our NeurIPS coverage last year. And we're proud to bring you part two of our special two-part episode covering our attempt at giving you an audio experience of ICLR. If you'd like to see us return to Vienna for ICML, let us know by sharing this episode on X and LinkedIn. In part one, we covered the best papers of ICLR across four sections. Image generation and diffusion, computer vision and weak supervision, improving attention algorithms and state space models.
0:51In today's episode, we cover the wealth of papers we found around the related problems of LLM reasoning and agents, also in four sections. In section A, we do a regular latent space chat with Graham Newbig of WebArena and OpenDevon. introducing many of the major themes of the rest of this episode. In Section B, we survey a few prominent issues in benchmarking from SUE bench, test set contamination and general intelligence. In Section C, we look at agent building blocks from RAG and self-reflection, verification, safety and frameworks. In Section D, we finally look at two proposed agent systems from Google DeepMind's web agent and MedEdge Pete.
1:38This is the second of two episodes covering ICLR, which is overwhelmingly academia-focused. If you're interested in production AI engineering and industry, you should join us at the first AI Engineer Worlds Fair this June, where we have now announced many of our speakers from all the big clouds, all the large model labs, now including Anthropic, Cohear and Cartesia, the brand new state space model startup, all top AI-enabled developer tools and code gen agents, now including Quinn Slack, CEO of Sourcegraph, insights on the GPU and inference market like last week's guest gradient AI, and now featuring Dylan Patel of the Semi-Analysis GPU Poor Blog and the rest of the emerging LLMOS stack of startups and open source tools across RAG, multi-modality, LLM Ops and agent frameworks, disruptive startups like MidJourney, Perplexity and Character AI, and for the first time, a closed-door track for VPs of AI and technical leaders to discuss AI strategy and leadership.
2:42Get your tickets now and see you in San Francisco from June 25th to 27th. We'll start this episode with a special on-site interview we did at ICLR with Professor Graham Newbig of the Language Technologies Institute of Carnegie Mellon University. Graham has taught the CMU NLP course for the last seven years but is also an active participant in the open-source AI software ecosystem, having been personally involved in OpenDevon, which recently scored a notable 21 % unassisted resolve rate on SWE BenchLite. As an extra special treat, we are proud to invite Amon Sanger, co-founder of Cursor AI, back as our first ever guest co-host to add his personal takes on the state of code editing in agents.
3:30This is going to be a doozy of an episode, so we better get started. Watch out and take care.
3:43So welcome to the pod, Graham, and welcome to the pod, Aman, our first ever guest co-host. Thank you for having me. Yeah, thanks for having me. Yeah, thanks for taking some time during iClear. This is very impromptu, but the two of you wanted to chat. I was like, let's just record a chat and that can be fun. And also, one of my goals here at conferences like this is just cover posters for people who are at home, like not at a conference like this, just to get a sense of the mood. So I'll cover a little bit of your background and then you can sort of fill in the blanks. So you're a professor at CFU, teach NLP.
4:13You also spent some years in Japan as a language teacher as well as a grad student. Yeah, it was a good experience. It was very impromptu, me going there just because I wanted to learn a new language, but I ended up staying there for 11 years. Yeah. Never know how life leads you, I guess. Yeah. Now you run your own lab and you teach the advanced NLP course. I'm sure that's been a wild ride over the past seven years. Wait, was it? It was 2016. Yeah. The last time I taught it before this semester was before ChatGPT. So I had to go in and rip out everything, all the new stuff in. But yeah, it's a good opportunity to keep up with all the new stuff too, because I feel like I need to pressure myself into like giving a good experience.
4:54So I need to, you know, cover all the areas. What do you mean good experience? What is a good experience for students? I think a good experience for students is them knowing stuff that's actually practically applicable in whatever they go on to do next. And I'm teaching the advanced NLP course, which is for people who are more on the research and new innovation side as opposed to just pulling in existing technologies. So because of that, I feel like I need to stay on the cutting edge of the most important things that people should know for the people who are pushing the boundaries to do the next thing.
5:26For the NLP course, how deep into the kind of older school techniques and fundamentals did you go? Yeah, so this is a constant battle because there are some older stuff that's like interesting, you know, algorithmically interesting, but we don't use as much now. It might come back later. It might not. But I think right now, given the limited amount of time I have and the number of things that would be useful to know, I mostly focus on the parts that are like in the modern stack of how you build models. What do you like to keep in? Is there anything that you do kind of keep in that's not really used much today?
5:58Yeah, it's a good question. I still teach Ngram language models, which are language models that are based on counting up the number of words that follow other words and stuff. And we don't use them at all today, but I keep them in there at the very beginning so people have an idea of how you could calculate this without throwing it into a neural network black box and what's the difficulty and why neural networks are important. but that's about it, I think. Interestingly, ngrams may be making a bit of a comeback for speculative decoding. Yeah, that's true. Can you elaborate? Oh, just one other way of kind of speculative decoding is when you have some kind of draft model predicting forward tokens.
6:38So you can kind of batch at the sequence level tokens when you're doing generation just because if you're generating one token at a time, it's much slower than kind of pre-filling a bunch of tokens at a time. So in this case, you would use kind of n-gram models to if, for example, you had V, maybe you know V is often followed by cat, and you wouldn't even need a decoder model or, sorry, graph model, which is a smaller model that would be kind of generating tokens. And you could just use kind of n-gram statistics or like basic n-gram models for this. I was thinking also that by pairing encoding is kind of n-grammy.
7:11You sort of build it up from. At least by grams. Yeah. Well, we can cover more later on, but I just wanted to touch briefly on the posters that you presented. I didn't know you had three, actually. I only prepped on two, but Amman can fill in the blanks on the one that you're presenting today. So Web Arena, Sultopia, and then Performance Improving Code Edits. Yep. Anything else that I missed? Those are the three. Those are the three. Okay, so Web Arena and Sultopia are basically like two kinds of sandboxes is what I was thinking about. Yeah. And maybe like which one do you want to tackle first?
7:41Like how do you tell the story about how this work is done at your lab? I think I can do Web Arena first. So to give a little bit of context, I never really was an evaluation person or never really cared about evaluation that much for a long time until maybe 2022 or so. And so I was mostly working on system building stuff. But once we started getting into the really good GPT models, my typical formula for system building was to figure out what's not working and fix it. And I got to 2022 and I was like, I don't know what's not working. these language models are so good. I don't want to be working on insignificant things that are already solved by the strongest models.
8:18And so because of that, I got into benchmarking. And kind of my goal for benchmarking is to push the boundaries of what is possible in a rigorous way. And kind of both of these Webarena and Sotopia are aiming to do that. I could explain about Webarena in more detail. Yeah, we'll put a slide of the poster up. It's a pretty useful slide. I like the phrase, we create a mini internet for your agent to play a master. Yep. So I can't take credit for that. That was probably Xu Yan, the first author, who came up with that. But the basic idea is I'm very interested in how language model agents that act in the world can actually work.
9:00So originally, what Xu Yan, the first author, wanted to do is she wants a robot that helps her out with housework so she doesn't have to do housework. The problem with robots is robotics isn't there yet. It's like I feel like robotics is the bottleneck in language model plus robotics work at the moment. And so we tried to think of something that had some of the interesting problems for like kind of long horizon planning. And also where you can pull in world knowledge about how the world works and use that in an interesting way. And so eventually we settled on doing tasks on the web as a way to benchmark this.
9:37And we wanted a benchmark that was as realistic as possible. So if you do well on it, it actually kind of means something. So like a lot of people complain about how academic benchmarks don't mean anything. And we wanted to create an academic benchmark that does, you know, actually mean something. And we wanted it to be evaluated in a way where you're not evaluating how good you are compared to like mimicking how humans do it, but like actually how good you are at solving the task. So we set up a sandbox intranet by taking production-grade open-source sites. Yeah, this is Reddit GitLab CMS. Yeah, so the sites, we tried to mimic some existing sites.
10:16So Amazon is an existing site, and we mimicked it with an open-source counterpart, One Stop Shop. Reddit is an existing site. We mimicked it with an open-source site called Postmill. And GitHub is an existing site that we mimicked with GitLab. And all of these are open source. They're actually used by people to do real things. So they're pretty close to being realistic. Then we created a whole bunch of tasks that you would want to do over the sites. And the way we did this is we looked through our own Chrome browsing history. So we were like, these are some of the things that we, as the authors, did over the past month.
10:48And then we tried to generalize them into things that would work on the sites. And then we wrote some validators to check whether it actually succeeded. So to give an example, one of the tasks is tell me how much I spent on food in the past month. which I had personally done because I wanted to budget and make sure I wasn't spending too much or things like this. So in order to do that, the agent has to go to the shopping site, has to identify all the food purchases. It has to figure out which ones happened in the past month and then add things together. And language models aren't particularly good at navigation.
11:18They aren't particularly good at filtering, and they're not particularly good at math. So despite the fact that this is pretty trivial for humans, it's hard for language models to do now. Yeah. And you had some stats here where basically all the models were under 15 % and humans were 78%. Yeah. I like that kind of wide variance for a benchmark because it means it buys us maybe a year before language models catch up. Yeah. So in the past six months, this came out about six months ago. In the past six months, we've increased from like 14 % to 25 % or 30 % in the state of the art. Using GPT-4? Yeah.
11:53Using GPT-4 as a base model. It's based on improvement. You mean before Turbo has? No, no. So actually, LLM improvements are not the main driver behind this. It's more about the way that LLM agent uses LLM. Got it. And to give some examples of some things that people have done, the first thing is just optimizing the prompts in the action space. So it's prompt engineering, figuring out which actions you can do, like good ways to scroll up and down the page and click on buttons and stuff like this. Other things that are maybe somewhat more methodologically interesting are it takes a step and does kind of self-refinement or self-reflection about whether that step was a good step or not and then rolls back if it was a bad step.
12:35And there's another one that tries to create like a textbook about how the sites work and give information about how the sites work. Like if you want to do this, you should go to this page. So it's creating like a textbook. Yeah. Yeah, creating world knowledge about, or documentation might be a better word, about how the sites work. And then feeding that into the agent based on where the agent is in the tasks. So there's a lot of creative things that people have done with this, but we're still at 25 % to 30 % as opposed to the 80%, 80%. Well, one thing that's surprising to me is it's only 70%, 80 % for humans?
13:09Yeah. So think about half of this gap is just human negligence for doing the task. like they don't follow the instructions exactly. About half of this actually might just be an issue with the benchmarks validators being too strict. And so the human gets the correct answer, but the validators are like not counting it as the correct answer. So they're looking for like an exact match and it's an exact match, but it's off by a rounding error or something like that. And so we could go back and fix all of the validators and that might bump up the human performance and the model performance. But still, I think the gap would be the same.
13:43I'm surprised that long context hasn't helped a decent bit for this. Like you mentioned, it was mainly not on the model side. But at least I'd assume for coding agents, longer context is helpful. I'm not sure what it looks like here. Yeah. For longer context, we haven't seen any papers that say adding a long context helps. It's also kind of not clear what exactly more context you would put in, because we're putting in the entire past action sequence. Oh, and that fits entirely. It's the entire past action sequence, but we're not putting in the past pages. So if we added the past pages, it might give us a little bit more information, but it might also explode the context that the model's looking at.
14:22So yeah, maybe somebody will figure out how to make it work, but I haven't seen it yet. And so what's the follow-on work from WebAreena? I think web browsing is super interesting. I think web browsing, the tasks that we created here were based on our personal browsing history, and they were also created to be kind of manageable based on what we were able to do at the time. And since then, we've released it, and we have a lot of new ideas about coming up with things that are more representative of actual tasks that people do in jobs. And so we're working on something related to this, like what are some of the tasks that software engineers do and things like this.
15:01And so in order to do that, we need to have web browsing. We need to have it over realistic sites that people use in actual workplace scenarios. So we're working a bit on that. And then separately, now that we have a benchmark, as I mentioned, I'm not an evaluation person. That's not my core passion. My core passion is building things at work. And so now we have a good evaluation that I care about. So we're doing lots of things to improve it. Some of the things we're thinking about are training models based on synthetic data or reinforcement learning methods. Also, better ways to interface with websites.
15:32So maybe we should be using an API instead of directly interacting with the website or things like this. Directly interacting, meaning point and click? Yeah, exactly. Using coordinates. Yeah. So I was going to say this is, for listeners on the pod, very, very relevant to our adept episode, where they have a very strong opinion that you should just rely on point and click instead of using APIs. I even had a section in my write-up calling and saying no APIs as a rule. Is that surprising? Yeah. I mean, what was there? I'm curious if there are even ones. It's very simple. The sheer number of things that don't have APIs vastly overwhelms the number of things that do have APIs.
16:07So if you're only going to constrain yourself to API work, you're not going to be generally capable as a human. Yeah, yeah. So I listened to that episode. It's a great episode, by the way. I largely agree with that. And actually on Web Arena, only one of the four websites that we cover has APIs. It's GitLab and the other three do not. But I think we can be creative about that. You know, maybe there's ways to create APIs for sites that don't have them or figure out ways to build tools that allow you to say, okay, navigate. It's kind of like the documentation idea that I said, like, how do I navigate to the purchases page?
16:42And if you just tell it what the purchases page is, then you can create a navigate to purchases page function or something like that. So we're not finished with this yet, but we're thinking about things in that general direction. I'm curious, what do you think is missing from the base models themselves? And how do you think the benchmark would inform what the model providers do to make better ones? Yeah, so this is a great question. So one thing about WebAren is that we're taking in a textual format. I didn't actually mention this yet, so I can explain it briefly. But we're looking at websites by using something called the accessibility tree, which was created for screen readers, for people with vision impairments.
17:20And I don't think that's necessarily the best way to see a website. I think multimodal is probably the way to go eventually because most people see sites by looking at them. I feel like all open source models and even most of the closed source models are not terribly good at understanding websites visually or through accessibility tree formats. So they make silly mistakes about, for example, not realizing that they can click on a drop-down menu in the accessibility tree format or not being able to ground to the site when you're looking at it visually. We have another version called Visual Web Arena that follows up and does this visually.
17:57Another thing is planning is a really big problem. So I'm very hopeful that the foundation model providers are working on this very hard now in training models that are better as agents. But for example, on Web Arena, it will very often step into a page and then be on the wrong page and just not realize that it needs to go back. So those are some other problems. It also has some really interesting failures of common sense. So my favorite failure is we have a thing in Web Arena that says assign this issue to myself. And it assigns the issue to the username myself instead of like your own username.
18:30So there's lots of like little common sense things there, too. So maybe tokenize identities and then have a special token for self versus username to myself. Although I kind of hope that my language model would be smart enough to figure it out, too. Like a human would. I don't trust it. Okay, cool. We should probably move to Zootopia. So going from simulating the internet to simulating multiple agents talking to each other in a social environment. Yeah. What's the motivation there? What's the story there? Yeah. So the motivation there is a lot of people are starting to use language models in kind of like social situations or at least like socially charged situations.
19:10Whether they're ready for this yet or not, I'm not sure about. And that kind of motivates this work. Like, for example, should we have a language model that negotiates with somebody about like a price of a... There's no should we. We are going to have them. Yeah, we are going to have them. Are they ready to do this in a way that is not harmful, basically? And also, are they good at it? You know, are they able to cooperate with somebody and find a solution to a shared problem? So what we did is we came up with like six different types of these like socially relevant situations. Yeah. Negotiation, exchange, competition, collaboration, accommodation, and persuasion.
19:54Yeah, exactly. And with each of these, we kind of like semi-automatically created a bunch of tasks. So the example task that I have on the poster is we have two people who are out in the cold. One of them has a blanket and the other doesn't have a blanket. And it's a guy and a girl. Very important. Yeah. And person one wants to keep their blanket to themselves, and person two wants to have the blanket shared. And so they need to negotiate with each other. Each of the agents is given a persona. They have a personality and some side information about them. They also have a secret that they don't want to reveal to other people.
20:27And then based on this, they have a conversation, and then we do evaluation along a number of axes, like whether they achieve their goal, whether it's a believable representation of a social situation, whether they violated social rules, whether they broke secrets and other stuff like this. And based on this, we had language models talk to language models and evaluate whether they do a good job of this. And we also had language models talk to humans and see whether they did a good job. And your language models eval. Yeah, that's what's next. How do you do the evaluations? Yeah, so we did human evaluation and we did language model based evaluation.
21:02And we tried to measure the overlap between the human evaluation and the language model based evaluation. and not overlap correlation. So how well do they agree with each other? And the answer is language models are okay at navigating these situations. They're also okay at evaluating these situations, but they're not perfect at either of these yet. 74%. Yeah, 74%. Which is usually around the rough correlation of LLMS as judged for most things. Yeah, exactly. And one of the interesting things we also found was like, if we have GPT-4 talk with GPT-4, it got a score of about 3.3 out of 7 on our kind of composite scale.
21:40If we had GPT-4 talk with humans, it got something closer to 4.8, I think. So when GPT-4 is talking to humans, it's actually more successful at achieving its goals and being believable and stuff because the humans are a better conversational partner and can recover from errors and stuff. But when humans talk to humans, it's 6.15. So we're, you know. What is the human to GPT-4? Human to GPT-4 is high. It's nearly identical to human and human, like 5.95. Your number should be bigger on your chart. Yeah, yeah. I will convey this to the students who made the poster. Yeah, I agree. But yeah, so I think this is pretty interesting.
22:20This was our first work in this general direction. Again, the iClear submission deadline was about six months ago. So a lot of these papers came out six months. But also some of the stuff we've done in the follow-up now is we are trying to get better evaluation, train better evaluators that are good at evaluating this kind of like social skills. And we're hoping to beat GPT-4 with respect to that. But more interestingly, we also tried to train models to be better at navigating social situations. So we took a Mistral 7B model and we trained the model through both behavior clonings. So we ran a bunch of conversations through GPT-4 and trained Mistral to mimic them, but also through self-reinforcement, where we basically had the Mistral model do a bunch of conversations, have GPT-4 grade them, and then pick only the good ones and train on that.
23:10And in doing that, we actually were able to max out our evaluation so that the trained model matched GPT-4 according to machine judgment, according to model-graded judgment. But when we actually did human evaluation on it, it was still far behind. So what this demonstrates is we can over-optimize to model judgments and actually get very close to our teacher model with respect to model judgments. But still humans fall behind. So it demonstrates that we actually have significant problems in evaluating these models as well. So I think we need to alternate between getting better evaluators, getting better models.
23:48Is the evaluation done right now via kind of prompting a GPT-4? Yeah. At the moment, that's the best thing we have, but we're trying to beat that currently as we speak. The last paper that you're presenting today is performance-improving code edits. You did the prep on this one. Yeah. I mean, first off, would you be able to kind of roughly explain the paper? Yeah. So the basic idea is we have lots of programs, of course, written by humans or written by machines. and correctness is a major concern, but efficiency is also a major concern. And so we ask the core problem, can we use like large language models to improve the efficiency of programs?
24:28And the overall concept of the paper is pretty simple, but the execution is maybe interesting. And so in order to do the execution, we basically took competitive programming problems that had lots of different solutions. And especially we focused on ones where the first one timed out and was too slow. And then a revised version didn't time out and was fast enough to complete. And the reason why this is interesting is because we know the implementation is going to be pretty similar because people working on competitive programming want to do it quickly, so they don't want to rewrite their whole implementation when they go from that.
Read the full transcript
25:02But we know one is slower and one is faster. And so then we basically asked language models to try to do a similar optimization and make it faster. So we created this data set. We also created an evaluation harness that makes it possible to measure things fairly, because if you just run a program and see how long it takes, what if the system is busy or other things like that? There's a bunch of mitigating factors, so we fix that through virtualized CPUs. And then we also have some better prompting methods and fine-tuning methods to try to create models that do better on this. And in the end, we actually got quote unquote superhuman performance on this task.
25:43And we came up with models that on average made things faster. The caveat for superhuman performance is the people doing these programming contests only need it fast enough to beat the timer. So they're not trying extremely hard to optimize. So I won't say that models have beat humans at program optimization yet, but it's like maybe a first. Still, it'll save you from really bad performance. Yeah, exactly. Yeah. I I mean, the surprising thing was it was only using, I believe, the code llamas and 3.5. I don't think it used 4, if I remember correctly. Yeah. This work has had a long evolution. And I think part of the reason why we didn't use 4 was just it's pretty expensive to run experiments.
26:26But, yeah, I think that's the main thing. Yeah. There are a few interesting things that I noticed. The really interesting one that I liked was the performance conditional generation. So I feel like I've seen pieces of this in other work. For example, AlphaCode actually mentions doing this. Yeah, I'm just curious what the motivation for that was. Yeah, so the basic idea is what we do is we kind of prefix the sequence that we're generating with how good the performance is. So it's like this is a 0 out of 10 performance thing. This is a 10 out of 10 performance. And then we fine-tune a model to learn how to generate slow implementations and fast implementations.
27:04And then at test time, we always append the fastest tag. This has been used in a number of places. Another example is a method called Quark, which basically tried to do this to generate not toxic text. So it would say, this is toxic text, this is not toxic text. And then when you generate it, you always append the not toxic text tag. So it's a pretty widely used technique, and it just seemed appropriate here because it's also very easy to use. You just append a tag. You evaluate, then you append a tag to the beginning. Yeah, and I saw you guys had used like 3.5 fine-tuning. Like one limitation of kind of using OpenAI fine-tuning is you can just kind of do supervised fine-tuning.
27:42You can't really, you can't do RL. But it feels like, you know, using this method, you might get some of the benefits of learning from negative examples. You kind of would get what from RL. Yeah, exactly. By prefixing tags. Yeah, right. Because you can have kind of a low-scoring thing prefix with a bad example following it, and then a high-scoring thing followed by a good example. And ideally, the model learns the difference with that data. In the Quark paper, at least they show that this works better than just training on the high-quality examples. So, yeah, excellent. What's the appetite for performance improvement code edits at Cursor?
28:18Yeah, I mean... In Cursor. I think like in practice, it's really tricky when looking at arbitrary code. There's the problem of actually like isolating the actual performance of some piece of code that you care about. First off, you don't have like this nice sandbox environment. People are almost always running it on their laptops or some kind of remote SSH machine. Then like actually isolating that piece of code, it's possible, right? You could kind of just add kind of timers around it. But we don't have anything like optimized for that. It's mainly like if the user wants to ask the model to improve it, they can.
28:49They can add in relevant information of how long it took. And it should work reasonably well, but definitely not like its own thing. Yeah. Another thing is our paper requires having tests. And in performance improving code edits, a huge, huge bottleneck is having good test coverage. Because if you don't have good test coverage, then it will generate a return statement and just return from the function with the correct answer to the test, which is very fast, but only works in one case. So I think for real-world code, that's a lot harder because generating very comprehensive tests is hard if you have data structures and stuff like that.
29:27So it's definitely not trivial. Awesome. I wanted to move on next to, now that you're done with the spring semester, I guess officially after you present this poster, you were mentioning that you're going to spend a lot of time on OpenDevon. So maybe, and obviously you're interested as well, what were your reactions to Devin? And then maybe you'll tell the story OpenDevon first, but I wanted to start with Devin. And, you know, just both of you, whoever wants to. Yeah, I think it was really exciting. Like the demo was great. And I've been working on code generation for a long time. Like I think actually since 2014.
30:00And, you know, the big moment where I felt it had made it first was Copilot coming out. And it's like, yeah, this is great, you know. And then I kind of slightly lost interest in doing this because I felt the limits of code completion. and like CodePilot does pretty well. Sure, we can improve it a little bit. And then, yeah, well, so, you know, maybe Cursor was right about how much more you could improve it, you know, honestly. But so I personally kind of lost interest a little bit, but then I was working on web agents and then I saw the demo from Devin and I'm like, oh, this is really cool. Like I'm interested in agents.
30:35I've been interested in code generation for a long time. You know, this seems like a good sandbox to be working in. And I had known about Sweebench and the stuff that had come out. And I was like, yeah, let's do this. Let's work on this problem. Because it's another benchmark like Web Arena where it's like our scores right now are low, but there's a lot we can do to improve them. So it's kind of exciting there too. Yeah, no, I thought it was a really good demo. And I think it'll be pretty useful for a lot of the bottom whatever percent of PRs. I guess, especially because I'm working on the stuff that I'm working on, I think I'm a little bit more bearish than other people on agents working immediately or somewhat soon.
31:15I think there are a lot of really hard problems, and human judgment is pretty paramount. I think on the margin, we are going to shift things a little bit more in the agentic direction, more things happening in the background. But I suspect the human will be needed for a while, rather than going from issue to pull requests. I think that's kind of sweet agent's direction, but I think one of the things that Devin nailed was the async interaction with the, it could be executing its plan, but you could sort of intervene while it's doing it. And that felt very much more like, I guess, you know, level three or level four self-driving rather than full level five.
31:53And yeah, that seems to make more sense. Yeah, and I think like one of the interesting things when we first came out with OpenDevon, you know, Devon still was not open for everybody. I can tell a little bit about the story, but basically we saw the Devin demo come out and Jun Yang, one of the people from the Kuen team building a language model at Alibaba. I think he's the lead on Quint, right? Yeah, he's one of the leads for sure. And he basically said, yeah, this is really exciting. Let's make a project about this. He made a repo with a readme, and the readme got 1 ,000 stars on GitHub. So then shortly after that, I think probably that evening or maybe the next evening, I was like, yeah, if we have something to hack on here, like the open source community is so excited about this that we'll be able to do something interesting.
32:35So I basically came up with a really, I'm not a React developer, but I came up with a really janky React. I wouldn't say clone of the Devon interface, but something similar to the Devon interface. It was completely non-functional. It had no chat functionality. It kind of looked reasonable. And then after that, I pushed that and it's like, yeah, let's make this actually work. And then a bunch of people came together and did that. So for the first four weeks or so, we didn't have anything that actually worked at all. But while we were doing that... This is when I live streamed and tracked it out.
33:07I was like, I really respect you and felt a little bit bad that I subjected you to that experience. No, I didn't. I was like, oh, this is done. But the interesting thing was then we had a ton of people who are not developers coming to us also. And so I think I totally agree that for really big software engineering projects like OpenDevon, like we're trying to use OpenDevon to solve issues on OpenDevon. And it's a complex enough software project that it's actually pretty tricky to do. Like the model needs to figure out how to set up the software repository in the first place. And that's a pretty big lift.
33:42But I think the possibility of doing things like slightly smaller level, like setting up simple web apps and stuff like that for people who are not professional developers is another thing that even immediately these sorts of agents might be able to start making a dent in. Yeah, that makes sense to me. As someone who's like slightly more positive on Devin, I'm also, you know, I don't think it's going to threaten our jobs anytime soon, but I'm pretty positive on it. Like I think it is very good for Greenfield and then like moderate for Brownfield. And then obviously depending on the size of the job that you're asking it to do.
34:17Like, yeah, there are many bottom percentile PRs that I have to do anyway. Yeah. If I could just throw it to Devin, even if it takes like eight hours to do it, that's probably one hour that I don't have to spend thinking about it at all, which is cool. One thing I wonder about is the UX, because it does kind of shift into you're managing a bunch of, let's say, like engineers, right? Yeah. You're kind of doing code review then all day. How does it feel? It feels fine. Literally, it feels like I'm an engineering manager. I'm technical. And I can see what my coders are doing and check in on them.
34:50If they're going off base, I can just tell them they're going off base and they'll replan. It's exactly what I do with engineers anyway. Yeah. No, that makes sense. And so in creating OpenDevon, and you were talking about planning earlier, I feel like Devin's planning, Devin made a big fuss about their breakthrough or secret sauce being planning. But I really don't think it's secret sauce. It's just they generate a plan and they try to execute it and the plans change over time. Did you find that hard? What was the, you know, any insights that you made OpenDevon? Yeah. So at the moment, we started out with implementing planning and I do think it's important.
35:25So right now, our best agent, which is doing reasonably well on the Sweebench benchmark, the same thing that Devin tried out on, actually isn't really doing any explicit planning at all. Nonetheless, we're able to get 21 % on the Sweebench Lite version. We haven't run the full Sweebench version just because setting up Sweebench takes a while and it costs$6 ,000 every time you run it with GPT-4. So it's a little bit heavy to run evaluations on it. But I'll be very interested to see, does their supposedly really good agent with planning stack up to something without planning, but just has like a good toolbox for, you know, searching code and for modifying code in a, you know, efficient way and stuff like that?
36:05Yeah, I think this is public, but they use Morph, which is Jesse Han's thing. I don't know if you know him. So they use a good code indexer or searcher that pages things into context whenever you need. Yeah. It seems like the magic trick. I don't know. We're thinking of trying out Morph, or we're actually actively trying out Morph as well. Right now, our code search is based on the code search that was used in SWE Agent, which is another agent by the people who created SWE Bench. And it's literally like find and grep. Okay. It seems to work good, but it won't work as well as an actual code search engine, semantic search.
36:42I don't know. Have you talked about the search that you use? We use a mix of things. Like the main meat of it is kind of retrieval with embedding. So kind of standard approach there. But then we use like kind of re-rankers in the mix. Occasionally use LSP information. I think there's like a much stricter requirement for the agent stuff for getting like exactly the right context that we don't have. So I'm not super familiar with what exactly Morph Labs has, but yeah. His emphasis is on speed and scale, but I don't really know how specifically indexing or retrieving. Yeah. And this is not at all like a knock on SuiteBench because I think it's a fantastic benchmark and it's a great way of kind of measuring progress.
37:19But I do wonder, I had posted about this, like how much of the performance is also that the models kind of do know those code bases because they're all public code bases. Yeah, so you asserted that it's already leaked for the online models. Like in some ways. Like I kind of tested this with one of the first problems that I saw in Sweet Agent. And Claude Opus basically knew the correct file to edit just based on the PR, the name of the pull request. So I suspect it's like somewhere in the pre-training data. I don't know like how much of an effect that actually has. Like if you're like getting better and better at it, I still think that translates to better performance on private repos.
37:54But I don't think like an X percent will also be on a public repo will be like the same X percent on a private repo. Yeah. So this is a great point. I loved your tweet about that, actually. And I retweeted it. But I can also explain a little bit about our vision for OpenDevIn. So it started out as basically a DevIn clone. But I feel like we've moved a little bit beyond this because I think the open nature, we have a hub where people can add agents. we have a pluggable thing where you can use any LM in it. So you can combine any LM with any agent. Is it light LLM? It's light LLM. I'm a big light LLM fan.
38:31It makes everything very easy. And then the final thing is we also want to have pluggable evals. So right now we've only implemented like Sweetbench. But there's a bunch of good code benchmarks. We're also planning to add Web Arena and through something called Browser Gym, which was created by ServiceNow that has these three web navigation benchmarks. because in order to be a good software engineer, you also need to be able to gather information on the web and stuff like that. So we're going to add that. And I think there are ways to basically create benchmarks that are not leaked using the same method as SWE agent.
39:04And there was recently a paper out of Berkeley called R2E that converts repositories into evaluation environments for code generation agents. And we're talking with the people who created that to incorporate that into our evaluation harnesses. I'm very excited about that. I talked to them, too. I think at least one or two of them may have been also the people behind LiveCodebench. Yeah, yeah, yeah. Which I think is fantastic because it's a great way of seeing if, and for people who don't know, LiveCodebench is basically, I think it's like a bunch of lead code problems, and you can kind of slide the cutoff date forward and backwards and see how different models perform.
39:38They go way down. It's really great. One thing that I'd be really interested in seeing is, like, I think there are now good benchmarks for overall agents working well, but good benchmarks for capturing all the things that a model needs to do well to be a good agent. Because you kind of need to build the good overall system, and maybe the system works really well for some models, better for some models than others. But I wonder if there's a good kind of benchmark you can do that tests independently each part that's needed to be a good agent. So this is a great question. And for web agents, we recently released something called Visual Web Bench, with the idea being that it tests about eight different capabilities that we think a model should have.
40:22Like, is it able to do OCR on the page? Is it able to ground the web elements? Is it able to predict the effect of clicking on a button or something like this? I think we currently lack something like that for coding agents. Some of the mistakes that we see our agents making are really silly. Like, it git clones a repo, and then it doesn't notice CD into the thing, so it tries to git clone the repo again. And this is GPT-4, so it's, you know, the most capable LLM model. And so I think there's a bunch of little things where it's like we could categorize these and just make sure it checks all those boxes, and it would just become more capable.
40:59But that being said, you know, there will probably be other things it fails on. So if we overfit to that benchmark that tests all the skills, then that would also be a problem. But I still think having one would be better. Yeah, like there are a few general things that you're kind of surprised by like how poorly the model does. Like one is kind of applying code edits. I had talked to some of the SweetBenz, SweetAgent people and like also just like looking at their demo, there's this great example of the model knows what, like roughly knows what to do. It has like a plan for it. Then it's trying to apply the edit and like seven times it incorrectly indents it.
41:29It gets that feedback, keeps doing it again and again and again. But like, yeah, code edits feel like one big part kind of of the pie that aren't like super well tested. the moment. Yeah. Anything else about the future of coding agents? Where do we go from here? I guess you already talked a little bit about the future of OpenDevon. I feel like we're right at the beginning of a very rapid delta with respect to the performance of how well these are going to go. And I think now we have all the ingredients for academia and open research to iterate on this. We have a good benchmark, like SweetBench.
42:04Maybe there are some issues with it, but I think it's fine for now to iterate on. OpenDevn, we've set up an environment where people can put in agents and very quickly iterate on them. So I think we're just going to see a bunch of people jump on this and improve rapidly. And then I think it's going to plateau a bit when we get to the really hard things where our current language model backends are not going to be good enough to handle them. We'll have GPT-5 by then. Yeah, but I think that's probably going to happen. And I think because all the open model creators know that GPT-5 is going to have that, all the open model creators are probably also thinking about it too.
42:40Yeah, and like, you know, Llama 3, 400B. Yeah, exactly. So I think that will give us another bump with respect to that. And we'll see how far that takes us. I don't think it will take us to resolving every GitHub issue automatically, but I think it'll be pretty exciting over the next, like by the end of 2024. Yeah. Again, I think it'll be like really useful for like some bottom percent of GitHub issues and that'll go up. The thing that we would like to build is kind of agentic things that happen as you're coding. Like the ideas we have in mind are like as you're coding, you can spawn off kind of pretty meaty units of work, right?
43:15Like as an example, let's say you need to implement some random helper function or some utility function in order to like get some value out of it. Let's say the contents of some file in particular, some particular way. You should just be able to kind of write that function out, file contents equals whatever, and then in the background that thing gets implemented for you. What we're going to see is we're going to see the ability of like basically scaling inference time compute in some way, and this could be either smarter models, it could be kind of using models with chains and looping. And when you scale up inference time compute, you can't really use the level one kind of systems that are built in with cursor right now, which is humans supervising the outputs of these models, either with kind of next edit prediction slash autocomplete, streaming in kind of diffs or chat.
44:00And so something needs to happen in the background. But the goal of what you want to do is it happens in the background in a way that's like very much preserving human flow and letting the human be completely in the driver's seat and kind of dictating exactly what happens. So it's kind of working completely in service of what you're building towards. I think this is the original Morph vision, and it'll be exciting to see what happens. Do you think most of that will be running locally or you don't really have a difference? Locally meaning? Inferring on models locally for cost reasons. I think it's going to have to happen with the most capable models, meaning it'll happen locally.
44:37Just a side note, I was just talking a lot about open models and all that. Do we have strong opinions about code-specific models being best for code or do we think general models are just best anyway, are the best code models? So maybe to phrase it, there's no code GPT-4, there's just GPT-4. And GPT-4 is the most capable code model. Yeah, I mean, here's one question I'll kind of pose in response. It seems like people say that training in code improves performance and everything else. There was one paper about that here. Oh, really? It's like, how much does code improve performance? And the guy didn't show up for his poster session, which is very annoying.
45:19Does it? Does it show that it's a performance? I feel like, I don't know, if you look at open papers, I think most of the time it's kind of showing if you've run out of data, then adding in code will improve performance, right? Which makes sense, right? Unlike reasoning in particular, it's the thing that people kind of speculate. But it does feel like OpenAI is the company that popularized this notion. And that was kind of like back in the day when they unified the models. Because it used to be like Codex was separate from GPT-3. And then they kind of had this unified Code DaVinci 2. So I actually wrote the first paper on that.
45:55Sorry, what was this paper called? Large language models of code are few-shot common sense reasoners, I think. But when we wrote that paper, we didn't know what Code DaVinci 002 was and what Text DaVinci 002 was. And we thought Code DaVinci 002 was a fine-tune on top of Text DaVinci 002. But it was actually the opposite. Like Text DaVinci 002 was a fine-tune on top of Code DaVinci 002. But nonetheless, Code DaVinci 002 had better performance on some reasoning benchmarks that we measured. And the funny thing is, actually, Text DaVinci 002 was trained on more data, but Code DaVinci 002 was still better at some structured reasoning stuff, which I would really like to prove, but we haven't been able to do it yet, is that code is more structured, and so it also has more repetition.
46:43So you need to attend back to the previous context more when you're doing code. and because of that it's better at capturing things that are very structured in the output and that includes things like reasoning. So I'm a pretty strong believer that there is something special about code but there could also be something equally special about text if you use the right variety of text, if you use text with lots of repetition or other stuff like that. So I don't think it's like code is magical. I think it's some properties of code are good for reasoning. Legal text. That would be one example. My suspicion is there's just like classes of text slash code that help for reasoning.
47:18And there's a bunch of not fantastic text for reasoning that will exist in pre-training data sets. Like maybe the very top, you have archive or textbooks. And then just under that, you've code. So it ranks higher than maybe most text data that's used in pre-training models. But it's not better than actually the thing you want. Yeah, I agree. Awesome. I'm going to broaden out to more general freeform topics. Something that we prepped was just the changing nature of like, I guess, industry and academia. I don't know if you guys have opinions on that. I guess you're representatives of both. Yeah, I'd love to hear your thoughts.
47:54Yeah, so the changing role of academia is really interesting because I lived through several areas where it was like, academia is probably leading research with respect to, you know, deep learning and everything, which was maybe 2010 to 2013 or something like that. And then there was the sequence models paper from Google, which was 2014, which was this, at the time, huge four-layer LSTM that nobody could train. And so then we were starting to feel the compute crunch, but there were still lots of modeling innovation. I created a neural network toolkit called Dynet, which was kind of a precursor to PyTorch.
48:36And a lot of stuff was happening there. And then after BERT, you know, it's like, oh, we're scaling up. We're moving beyond, like, the compute that Academia can use to train these base models. And then I think the really big thing was, like, the GPT models, right? And after ChatGPT came out, we actually had an emergency workshop at CMU, which was a group therapy session to say, what should we do? What should we do? And I think for a short amount of time, a lot of people were worried, like, what could we be doing in the face of this? And then I think a few things changed. I think, number one, the evaluation stuff I talked about, it's like we realized that actually there's a lot of stuff that GPT-3 and chat GPT cannot do yet.
49:14And, you know, more complex reasoning, more multi-step stuff. And another thing is all the open models started coming out, which made it a lot easier for us to do fine-tuning. A lot of the open source frameworks came out that made it easier to do these sorts of run large models on hardware that we have access to. So I feel... Are you talking about like LamaCPP? Anything else? No, I'm talking more about the training stuff, like DeepSpeed, WAMA Factory, Axolotl. TensorRT. Yeah. Any of the things that we can use for training. And that makes it so you have a machine that costs$100 ,000, which is a lot of money, but it's very much within an academic budget.
49:52And you can actually do training runs, do interesting things that free you up. And then at the same time right now, every university is trying to build a GPU cluster. Yeah. or get access to it, including us, including everybody else. I imagine CMU would be ahead because you already have so many other needs. Yeah, so we do have a good cluster, but the kind of hardware that you need for training large language models is kind of specific. You also need a system for allocating, like, okay, this is the most important thing to be doing right now. We're going to give a lot of compute to that, which is not something that traditionally universities are very used to doing.
50:29They're used to being very chaotic with lots of ideas, but I think we need to focus on some. MIT had this really great big cluster, but it was V100s. So two generations too old. Soon three. There's these national super compute labs that you can apply for grants for. The funny thing is many of these don't have the hardware that we need. They have V100s or they have A100 40 GB things. They have A100 80 GB, but they only have four of them. It's kind of interesting how little there is available. Well, you can talk to Luther, which has its share of grants. Luther is basically a compute grants collector right now.
51:07They're pretty amazing at what they've been able to achieve with that. But they're also mostly not using U.S. national clusters. I think they use some overseas and other stuff like that. I wouldn't be surprised if Andromeda would give away compute for research. I'd heard of them doing something like that before. Well, you would know because you're in AI grants. Yeah. Like, yeah, my first response would be, aren't they already maxed out by existing users? Yeah, I think I'd heard when there was, like, a period where there were, like, not too many people using it or, like, a bunch of people canceled.
51:39They gave it away on some grants. Maybe if there are, like, bubbles where people aren't using it. Yeah. The other two sources I'll name are Crusoe Energy, which is using, I don't know if you're familiar with them. I think they, correct me if I'm wrong, they put GPUs on top of, like, oil rigs. I'm not familiar with the details. It's slightly sketchy, but, like, whatever. It's clean. Yeah. OK. OK. And then the other one is Strong Compute, which is doing one of those distributed cluster things. So it's like together, but with less funding and from Australia. We are working with some providers like RunPod and NetMind.
52:13So I think there's definitely some resources out there, but everybody is looking for them. And really, I think the solution is we're going to need to scale up the compute that we have available to academia in the U.S. Yeah. Well, in CMU and just in general. But I think we realized the importance of this. I hope the U.S. government realizes the importance of this and invests lots of money in it because that's actually the best solution. But they move a little bit slower than a lot of people move. So we'll see. Yeah. I'm curious. If you kind of look across all of academia, what work have you been either most impressed by or do you think best represents the kind of work you think academia should be doing in the last year or two?
52:50Yeah. So actually, another comment about the academia versus industry thing, I really do wish that the people doing kind of frontier research on language models in industry acknowledged academic work a bit more. Because I do think a lot of the things that people are doing in academia end up in industry but just don't get acknowledged. And I think that has to do with the fact that industry is super secretive right now about anything they do in large language model space. And so previously, it would be like industry is publishing papers, and we could point to the fact that, hey, Google uses our stuff, OpenAI uses our stuff, or things like that.
53:26But now there's a lot less of that, which makes it seem like we're shouting into a black hole when actually maybe we aren't so much. And the best example of this recently was like the Matryoshka embeddings thing from UW where OpenAI used it and renamed it to something else. I mean, that one seemed like just an oversight rather than intentional exclusion because they left enough hints that it was that. Yeah, maybe. And they did better with Sora, for example, where they actually cited all the works that inspired them. The diffusion transformers. Yeah, and things like this. But they have every right to be secret when they're competing.
53:58Their industry, they're competing against each other. I think it makes sense. But it's also a little bit disincentivizing for grad students, for example, because they can't point to their success stories that they had before. I don't know if there's any solution to that, but I thought I'd mention it just in case anybody who has influence would be listening. It's just a corporate responsibility thing. It's the right thing to do. But for me, the interplay between industry and research, I think you feel it the most with just your grad student pipeline or maybe the undergrads that you're teaching, what their interests are.
54:32I'm sure the class composition has changed a lot for the NLP class. Yep, yep, yep. I realize I skipped your examples of good papers from academia question, actually, and I can go back to that. But examples of good papers are both on the evaluation side and on the modeling side. I think on the modeling side, I have always preferred papers that are simple but work. And I think I'm a little bit weird with this respect in academia sometimes because I feel like when I see papers get reviewed, people are like, oh, this paper is not novel enough. But I'm like, this is a great paper. It's like it made a small tweak to this method, but it works three percentage points better.
55:16They'll change one line of code and suddenly everything will work. But that insight was not there before, which is why it didn't exist. So I really do like those sorts of things. I think DPO is a pretty good example. It's a lot simpler than they make out in the math. It's a lot simpler than they make out in the math, but I think that's a great thing, right? It's like a simple tweak that worked really well and people use it a lot. I think those are the kinds of things that are really valuable. I also think benchmarking, which isn't simple and takes a lot of work, is something valuable, which is why I'm spending time on that.
55:48Also, contributions to open source, because I feel like there's a small number of companies that are very committed to open source, like Hugging Face is the obvious example. But they don't have enough firepower to compete with the bigger companies who are working on these sorts of things. And I believe that open source, good open source alternative should exist. And academia could help with this. The problem is we're very disorganized. So if we solve the problem of organization and, you know, focus and getting everything together, then that could help. And I mean, like Hugging Face is one example of a company that's doing that.
56:24I also hope that like efforts like OpenDevon or MergeKit for model merging or other things that pull together a whole bunch of different things under one roof could help out with that too. Yeah. I'm trying to feature those things. I have a MergeKit talk in my conference. Those kinds of projects will never get featured at iClear. I'm trying to create a venue for engineering rather than just research. But obviously, there's overlaps between them. He's speaking as well. Do you know what you're going to speak about? Not yet. We can broaden out to just the syllabus and student interest before and after.
57:01You said you had to revamp the NLP syllabus. We can talk about that. We can talk about how to pick promising areas of work, which you already somewhat covered. Syllabus before and after. I increased stuff on distillation and synthetic data for sure. I added a thing that was like a tour of large language models. So it was covering all the different large language models and their similarities and differences and stuff like this. Because even I didn't know enough about the differences between the models. What are surprising differences? I don't know if this was necessarily surprising to me, but it might be surprising to some people, but how similar the architectures are.
57:41Everybody is using the Lama architecture. And it's not because architecture engineering is not important, but it's because we finished architecture engineering, and now we have a really good architecture that works. And we're at least a local optimum for that, which is everybody uses rope. Everybody uses Swigloo. everybody uses all these other small tricks and there's this really nice figure in the Mamba paper so Mamba is kind of like a linear architecture it's a great paper but there's a figure that compares the original transformer to the llama transformer with respect to how well it scales yeah and the llama transformer just scales like way way better than the original transformer but so architecture is important but we're kind of done with that and everybody is making no more than small but there's like I don't know I feel like My belief here is there still exists a bunch of tricks.
58:33Like there's the MOE trick, right? Like that'll get you like a slightly better skill. It's a big one. Maybe there are small ones. I do wonder how many of these are left and how many of these also maybe only come into play when you're at larger scales. Like there's a great recent paper at Meta where they trained on the next few tokens, right? So you're not just predicting the next token, you're predicting the next four. It didn't show better performance at small scales, but I think past the 13 billion parameter scale, it actually showed better performance. performance. What? Yeah. So this is like another concerning thing for like academia, perhaps.
59:04You may not know if your method actually works unless you're dealing with like large enough models trained enough data. Yeah. I think that that's a major reason why we do need to scale up the resources that we have for training models. And I think there's a lot of progress on that right now. Like I think that a lot of places are working on that. And then the other thing that I wanted to mention is, yeah, because the architectures are so similar, the data and the training methods are the big difference there. And that's where everything is actually really, really influential. Like what data do you train on?
59:37How well do you clean and deduplicate your data and stuff like that? And that's not something I've really talked about at all before when I taught previously one year ago. So a lot more focus on data. I definitely teach architectures, but there's a lot less focus on architecture engineering. And it's more of an explanation about why the architectures we currently use are the ones that work. But we did have Albert Gu talk about Mamba because he's at CMU too. I was going to ask, what are your thoughts on this new wave? I feel like there's a bunch of alternative architectures, like mainly Mamba, then I think RWKB.
1:00:09Those are the main two. Yeah. I think we don't know enough about them. I definitely would like to focus some percentage of our effort on understanding them better because this is a perennial problem, which is when you try to do something really, really different, there's so much catching up to do with respect to the highly engineered thing that we have before. So neural machine translation, for example, it took a year or two to beat statistical machine translation or phrase-based machine translation, which is what we had before, just because there were like 10 years of engineering that had gone into phrase-based machine translation to make it work really well.
1:00:46And I feel like we're kind of in that thing for all of these linear architectures like Mamba and RWKV. I want to see them continue to be pushed. But I do think they have some fundamental limitations. Like, for example, recalling information. So we see the hybrid architectures with seven Mamba layers in one transformer layer. Yeah. Recall and stuff like that. So it would be interesting to see if they're... That's Jamba, right? Yeah, Jamba. Yeah. So you're optimistic on the mixing? I think it's one way to solve the problem of poor recall in linear architectures. but there might be other ways. Yeah, I would say pure Mamba and pure RDDKV both have the RNN problem.
1:01:20Right. It's forgetting. It seems like you have to mix them. Yeah, I mean, the mixing, you lose the niceness of you're getting a factor of eight, but it's not getting like, if you're scaling up to a million, like 10 million tokens, like it just won't scale, right? Because you still, you're only diluting by a factor of eight while you still have that quadratic attention bottleneck. But you can do strides and stuff like this. So I think there's a lot of room for improvement here, which is why I'm kind of excited about that direction. That's the direction I'm most excited about with respect to architecture engineering.
1:01:52Interesting. One direction I wish would work, but it feels like no one's made it work, and I don't know if it will work, is some kind of retrieval baked into the model. So retro is kind of like the original paper. Well, you know, Dolly Kila from Contextual is working on... Oh, yeah. So Contextual, I think, is working on things related to this. But I don't know, it does feel like if you really want to scale to 10, 100 million tokens, you can't store that in some compressed state. You can do something fundamentally different, putting it into context. It's like you need all that information, right?
1:02:20You need all the information to be present, but you need some kind of sublinear per token generation. So we have a paper called Unlimoformer, which was at NeurIPS last year, that does retrieval-based detention. It encodes all of the previous context in FICE retrieval index. I like that general direction. It worked really well with kind of more traditional transformer models. We used it for T5. But one difficulty is actually rope makes it very difficult because you need to handle relative positional encodings appropriately and stuff like this. So, yeah, I could go into details here, but we don't have a lot of time.
1:02:55But I think there definitely are some things moving in that direction. I think that's another thing that could be interesting. Take our existing architectures and somehow make them efficient through approximations. I do think this kind of K &N operator is pretty interesting. K &N, the Comerable Growth Arnold Network, or something else? Sorry, just K-Nearest Neighbors. Just being able to do a K-Nearest Neighbors operator. Because if you're trying to do attention over all these tokens, you're now taking the average of all, I don't know, you have all these keys and values, and you're averaging 10 million of them.
1:03:29It feels worse. You only need the top few. Interesting. Rag on KB scores. We're basically done. I will leave it to you for any plugs that you want to do, any calls to action. Yeah, I guess I'm really excited about open source things in general. So, you know, any... Come contribute on OpenDevon. Come contribute on OpenDevon. Also, you know, use OpenDevon to test your agents, add new agents to make it work well on particular tasks and stuff like that. That's really exciting. And also just in general, like I really love new developments in open source AI. So even if it's not in OpenDevon, like I'd like people to continue pushing on that.
1:04:13And I appreciate it when people do that. No, open source AI has been fantastic for, I mean, it's fantastic for startups as well. Yeah. Like it's been super helpful for us. Speaking of Quinn, do you guys use Quinn? Like what's the relationship between you and... Oh, yeah. So we have a roadmap for OpenDevon. And the initial roadmap for OpenDevon was by the end of May, we wanted the best agent on SweetBench. And we did that by the end of April. And we did that May 5th. So we were a little bit late. Close enough. Yeah. And then our... It was four days ago. And then our May roadmap is to have a really good open agent.
1:04:48And since we have people on the Gwen team, you know, working with us, I think building something on Gwen would make sense. But, you know, Llama 3 is also good. So we'll see. Plugs? Calls to actions? Yeah, I mean, we're hiring for Cursor researchers, engineers, ML engineers. I think we're working on very interesting stuff in code generation, kind of on the frontier of what is possible for kind of in-flow coding and assistive coding. So I think some really interesting stuff to do. Always hiring. I love the hustle. Yeah, my plug is AI engineer conference that I'm spending all my waking hours working on right now.
1:05:23That's it. Thank you very much for your time. Yeah, thanks a lot. It was great. Yeah, it was really fun. That was Section A of our ICLR reasoning and agents coverage. The next section, Section B, covers related discussions of benchmarks. We start with the hottest new benchmark that has emerged this year, SWEBench, which broke through the noise as the presumptive next level after the saturated human evil and MBPP benchmarks from OpenAI and Google DeepMind. Hi, it's great to be here. My name is Carlos Jimenez, and I'm a PhD student at Princeton University. And today I'm going to be talking about our evaluation benchmark called Sweebench.
1:06:02Can language models resolve real-world GitHub issues? This is a work with my collaborators from Princeton and the University of Chicago. And I led this project with my co-author, John Young. So recently, language models have become really, really popular. And they're being pushed to perform in use cases that researchers haven't previously considered. A lot of past work on evaluating language models has become outdated simply because model performance is getting really, really good. And that's a good thing, but evaluating language models is also really important. Understanding the strengths and weaknesses of language models plays a major role in building future applications and helping end users know when and where it's appropriate to use them.
1:06:50So I want us to think about what sort of qualities makes an evaluation benchmark useful. First, the problems need to be hard enough to challenge state-of-the-art models. Problems should also reflect what people actually want to use the models for. And lastly, solutions need to be easy to verify. You can, yeah. So consider the tasks involved in software engineering. Software is an extremely powerful tool, and coding is already one of the most popular applications of language models today. In reality, programming is a very hard skill to master. So if we can understand how AI systems perform on this task, we have a better sense of their abilities on doing real and challenging work.
1:07:35Furthermore, we have a lot of infrastructure for evaluating code. many large software projects incorporate things like unit and integration testing to automatically evaluate changes to source code. Currently, language models only report evaluation numbers on coding benchmarks like HumanEval. And let's take a look at an example. So it starts with a function signature and a doc string describing what the function should do. and language models are evaluated on their ability to write the body of the function. Here it's highlighted in yellow. And there are many possible solutions, but the nice thing about programming is that we can check the validity of any of them automatically using unit tests.
1:08:21However, very few programmers ask language models questions like this, unless maybe they're trying to cheat on an interview. Software engineers typically write code that fits into a larger project, not one-off isolated functions. So we created Sweebench as a benchmark to evaluate the software engineering ability of language models in as realistic a setting as possible. I'll show you a very high-level view of what Sweebench is trying to evaluate, and then I'll talk about how we created it and so on. So Sweebench starts with a code base and a problem statement. And by code base, I mean a real code base with like hundreds of lines and thousands of files and thousands of lines of code.
1:09:04Hundreds of files and thousands of lines of code. And usually the problem statement is describing a bug with the code base or requesting some new feature or change in behavior. We then give this to the language model, and the language model is tasked with generating edits to one or more of the files in the code base in order to resolve this problem statement. Then we take the model's proposed changes and evaluate it using unit tests from the same repository that were made after the issue was resolved. What this means is that Sweebench can programmatically evaluate AI systems on their ability to solve real-world problems situated in full code bases.
1:09:44Solving tasks in Sweebench goes beyond just code generation. It requires models to understand how large code bases work and how code changes in one function can impact the behavior of other parts of the code base. I'll briefly summarize how we made Sweebench. Now, GitHub is a website that people use to collaborate on software development projects, most of which are open source. And AstroPy is one such example. On an open source GitHub project, users of a software can report bugs or request features by creating an issue explaining a problem that they encountered when using the software. And someone else who knows how to fix the issue can submit their solution in code, which is called a pull request to the project.
1:10:39Now, maintainers of the project can then review and modify that pull request and either accept the solution, in which case the source code for the project is updated, or they reject it. Now, this process in collaborative development, which underlies a lot of the open source development process, it can naturally be converted into tasks. And that's the source of the task instances in Sweebench. So we use the following procedure to gather task instances. We first scrape 12 popular Python repositories for all of their accepted pull request instance pairs. And then we filter these pull requests to make sure that they contribute updates to both the source code as well as the tests in the repository.
1:11:29And lastly, for each instance, we verify that the source code can be installed automatically and that the testing behavior changes before and after the source code solution is applied. So let's look at an example of what an instance in Sweebench looks like. This is an issue from the SymPy Python library, which is used for symbolic mathematical operations and notation. So we show the problem statement on the left, and it's giving a detailed explanation of what the user is experiencing with SymPy, where they're seeing unexpected output when using the identity matrix. And the code base for this issue is going to be tied to the version of the code base that was active when the issue was first submitted.
1:12:19Next, we'll have the gold patch. And this is the edits to the source code that was submitted with the pull request. And it represents a possible solution to the issue. And it was the one that's officially accepted into the actual repository. And then finally, we have the test patch. And it's the edits to the tests that was contributed with the pull request. Now, the test patch updates or adds tests that evaluate source code for this particular issue, and we verify that the testing behavior changes from failed to pass when running them before and after the source code is updated with the gold patch.
1:13:03So, after collecting instances like this across 12 repositories, we end up with over 2 ,000 instances representing a diverse set of problems and code bases. Each Sweebench instance includes the full code base, totaling to about 3 ,000 files on average, while gold patches usually only edit one or two files. We further collect 19 ,000 unverified instances, so unverified meaning they don't have test cases, and we use that for training purposes. So as an initial baseline, we use a retrieval augmented generation system, or a RAG, using a simple BM25 sparse retriever. And language models are then provided with the problem statement, the project's readme file, and the entire file contents for the top retrieved files.
1:13:53They're then tasked with generating a patch file that specifies which files to change and the edits that they want to make to those files. We evaluate top models like ChatGPT, GPT-4, and Claude. And we also fine-tune CodeLama using long-context RAG examples from the training set to get our own model, SWI-LAMA 7B and 13B, which are the only open-source models that have non-zero performance on SWI-Bench now. And across the board, base performance is extremely low, with the best-performing model, Cloud3 Opus, resolving only 3.8 % of issues on SWI-Bench. So how can these models get better on Sweebench?
1:14:40Well, first, improving the RAC system can greatly improve performance. So if we assume that we have a very, very strong retrieval system that retrieves all of the files that were edited by the gold patch, which we call Oracle here, performance jumps immediately from 3.8 % to 9.1 % for Cloud3. Another thing is that long context still seemed to remain an issue. So for CLOD2, longer context inputs anti-correlates with performance very strongly. And that's something that we saw with basically every model. The longer the files that are being retrieved or input, the more context, the worse performance in general.
1:15:22And so lastly, qualitatively, we find that language models tend to generate shorter, simpler, and more primitive code compared to gold patches. And for instance, we noticed that they tend to overuse Python built-ins and ignore library-specific utilities and API features with an example shown here. So let's summarize. SWE bench is a benchmark for programmatically evaluating the software engineering abilities of AI systems using over 2000 real world test instances. And we show how even state of the art models are still woefully behind on this task. We open source SWE llama 7b and 13b, which are suitable for long context rag with SWE bench.
1:16:10So before concluding, there's one more thing. I've shown you performance using a RAG system for language models on Sweebench, but software engineering is naturally a very interactive task. And we've recently had a follow-up work to this paper called SweeAgent that explores that idea a bit further. And with SweeAgent, we built an agent-computer interface for language models to interact with a computer to solve tasks on Sweebench, demonstrating much better performance. So up to 12.5 % of Sweebench is resolved. with our new framework. And this shows that there's a lot of room to improve for AI systems on SweetBench.
1:16:51So finally, I'd like to thank all of my collaborators and colleagues who have helped with this project. We have an active community on GitHub, so please consider submitting your own solutions to be listed on the SweetBench leaderboard. Thanks. Hey, John. Nice to see you at the oral session. So congrats on the success of SweetBench. Thank you. Why do you think it's caught on so much? The first I heard about it was from Devin. That's right. What was the launch process? I think, you know, I'm trying to get into the meta story around, like, a lot of grad students here are trying to get their work noticed.
1:17:25Yeah. You got noticed? Yeah. How? Yeah, no, that's a great question. The Devin release certainly helped a lot with really putting SweetBench, sort of, I think the biggest contribution they did was give people a visual of what even 15 % looks like on SweetBench. and I think it made it really compelling. Prior to Devon, we had started working on Sweet Agent back in, I want to say, September, like right after we submitted this project to iClear. So we kind of had this vision that like, oh, we're going to put out Sweet Agent and people will see the numbers can in fact go up. A lot of the feedback and the skepticism we had at that time was that the benchmark is really difficult.
1:18:01But looking at kind of what HumanEval did, when they released GPT-3 was at 0%, and we sort of used that as kind of motivation of like, Well, you know, it's really bad now, but if we keep at it and we use this sort of agentic approach, there's something promising that could come out of it. So I feel like with this benchmark, our advisors and like me and Carlos, we're just really had this mindset of like trying to champion our own work a little bit. And I'm kind of expecting that it'll take off of it on its own of like that first step from zero to 10 or whatever it might be that we really have to drive that and sort of make that happen.
1:18:35Yeah. So you worked on Sweet Agent first? Sweet Agent, we worked on this starting in June last year, summer. We were able to submit by September to iClear, and then right after that, we started working on Sweet Agent. Got it. And then just the backstory behind how you guys started to work together, how do you choose this direction, anything like that, the narrative story. Yeah, for sure. Carlos is fantastic. He had mentored me for a long time. I was a master's student at Princeton, and Carlos' fourth-year PhD is about to graduate next year, and he has a lot of expertise. He's built benchmarks before.
1:19:09I also give credit to like Shunyu because we had worked on Webshop and Intercode and a lot of this agent stuff before. And Alex was around during the summer and he helped a lot with sort of thinking about the fine tuning and sort of what are good baselines to go with. So the way it kind of came together was in June, I put together a lot of related work and thought about this idea and brought it to our advisor, Karthik. And then I found out that Carlos had a very similar idea kind of at the same time of just sort of thinking about how we could take a lot of this great open source data on GitHub and turn it into a meaningful task.
1:19:42And then really just sort of like the nature of the task and how to follow through in terms of the engineering plan, I think was honestly quite clear after a week. And we just had to execute at that point. What were some of the big debates where you had to go either this way or that way and you picked one way? Oh, that's a great question. I think one of the things that I really remember initially was sort of how we were actually going to collect instances and sort of what the heuristics are. In hindsight, I think they're pretty straightforward and obvious. But at the time, one thing was just like, do we want to collect a lot of instances from different repositories?
1:20:16Or do we want to sort of focus on a couple of well-maintained repositories and mine the most instances from them? So just to sort of contrast that, we have 2 ,294 instances from 12 repositories. we could very well have 700 instances from 500 repositories. Depth ended up winning out. It wasn't very obvious to us, but really just the process of manually inspecting, looking at contribution guidelines, looking at the natures of the test, gave us a lot of these heuristics that ended up being pretty reliable and I think scaled pretty well, at least for PyPy packages. So a lot of design decisions there, but I think we got lucky that we had a couple of good hits in the beginning.
1:20:56Were you concerned that it's primarily Python? Yes, yes, yes. That's a choice, you know? Yeah, it is a choice. So I think for the first version of the benchmark, just because Python is so commonplace kind of in LLM evaluation, especially with HumanEval, that we'll just go with the flow. We have no problems with the language. I think in the same way that MultipleE maybe expanded the amount of offerings for HumanEval, this is something we'd be interested in doing. It's going to require quite a bit of engineering effort. Like, we're both sort of fairly good at Python, but when it comes to things like maybe Rust or Scala or these other languages, like, we're not quite sure.
1:21:34But I think, like, basically if there's an opportunity to collaborate and there's people who are experts in those languages, maybe even the software repository maintainers, and they're interested in sort of having agentic language models help maintain their code base, we're more than happy to work with them to sort of see the sweep and sort of evaluation harness idea and really manifest it for what they're trying to do. A lot of benchmarks try to say like, okay, human performance is 50, and then most language models are 25, and then we'll try to get the language models above human. But like, what is human here?
1:22:09What is a single human performance here, right? Is it 100, or is it not undefined? Yeah, that's a great question. I think it's kind of an evolving answer in the sense that when we initially pushed the paper, we were like, it's 100 % because someone wrote it and they did the issue and they contributed. Thousands of people wrote it. Yeah, exactly, exactly. But more recently, I think as this benchmark is kind of picking up and there's more people interested in it, I think just having the right efforts to figure out which issues are easier or harder along what dimensions. When we say easier for a human, what does that actually mean?
1:22:47Is it characterized by the issue or by the size of the change or by the nature of the change? Maybe it's a one-line edit, but it's really difficult because you have to know the code base super, super well to make that precise change. So we have some ongoing efforts that are just sort of taking a look at Sweetbench and taking a look at sort of the actual code changes and the problem statements and saying like, all right, let's figure this out. Let's get sort of like Spider, how they have like easy, medium, hard, extra level problems, you know, stuff like that. Yeah, yeah, yeah. What about eval cost?
1:23:17I think one of the reasons that human eval is so popular is because it's quick to evaluate. Yes. But my impression, I haven't run SweetBedge actually, but my impression is it's quite expensive. Yes, you're absolutely correct. I 100 % agree with you. For the 2 ,294 instances, in hindsight, if it was a little smaller, it would have been okay. But for RAG, it's like 20 cents a pop. Even then, to do sort of pass at whatever or run a model a couple times if you're kind of empirically validating a system, So the cost can add up really quickly. So I totally agree with you there. So like$100 or something to email?
1:23:50Yeah, like for example, for SWE agents, so we just released kind of the first preprint recently. We'll kind of polish it for NURBS, but just to get the idea out, we actually have a cost table in there. And for the agents, like for the resolved issues, it takes like$1 to$2 to actually solve it. We have a$4 limit. But the meaningful thing that I feel like we did in response to sort of a lot of this was create the SWE bench light split. And that's 300 task instances filtered from the 2 ,294. And the objective there is that we apply some filtering criteria. It's not random. To only look at changes to one file where the issue has, you know, reproducible code.
1:24:28Like, basically removing some of the diversity of SweetBench that makes it really difficult to have sort of a more better starter one. And recently, like, there have been works, like, even yesterday that can, like, auto code rover. or like, yeah, I mean, like, I think this would have been, we kind of wish we had this to give to Devin also when they were running on this. And also like Kodak from sort of the OpenDevin team, you know, they kind of put together something and they claim they have 21 % on SweetBendSlight now. So I think after we offer this, it's kind of a, there's a little bit more traction, I think.
1:24:59Okay, awesome. What is your take on SweetAgents direction versus Devin versus OpenDevin versus AutocodeRover? What are the goals? What are the logical differences? Yeah, yeah. Oh, that's a really, really great question. Yeah, I guess I'll, like, kind of speak from what I know, and I don't claim to have a good understanding of the other systems. I think what Devon and OpenDevon are doing is, like, really, really cool. It's really, especially when the Devon demo came out, just to kind of see it as potentially, like, a product. It was really fantastic. What I'll say is, like, I think for SWE Agent, we were a little bit more focused on, from a research angle, just getting something to work and sort of having the empirics and the numbers to back it up.
1:25:39I feel like Carlos and my direction has more been sort of like trying to solve interesting research challenges. So like even after SWE Agent, like understanding what human interventions look like for autonomous software engineers, why people intervene, how they want to intervene, stuff like that. Which, to be clear, right now, in the SWE Agent that I saw, there was no human intervention. Exactly. You're right. Exactly. And I think from that, I think it was a great session. Like I really learned a lot. And I went back to my, I was like, you know, we should really think about what they said. There's some good points in there.
1:26:12But yeah, to say it in one sentence, I would say it's just like, yeah, just like, I think people like autonomy, but it also seems like they don't want to give up control. And it's kind of interesting in the sense that, like, it's cool from a research point to put out an end-to-end software engineer, but I think just like Devon and OpenDevon have kind of the dialogue system. I think they're building out really intuitive features. The one thing I'll say is like from kind of when we did Sweet Agent, some intuitive things worked, some intuitive things did not work. And that's kind of what we're interested in really discerning.
1:26:42Like some human UIs, user interfaces and applications are great for giving to the language agent and having it sort of run with it. Some things don't work as well. So, you know, we're curious in figuring that out. Cool. I think that's all the questions I have. Awesome. Yeah, yeah, yeah. Is there anything else I should have asked you? Oh, that's a great question. I guess like in terms of sort of like the future of this, I think it's just exciting to see people be very enthusiastic about this kind of evaluation paradigm. So like, yeah, like the feedback you gave on the podcast, I think it's great.
1:27:15Like I'm really receptive to it. I guess like just as a personal job, like I won't be in Princeton. I'll be in Stanford coming this fall. So I'll be back in the Bay Area. I'm very excited. I think it's a great ecosystem there of sort of like a lot of people are pushing this. and yeah, just excited to sort of collaborate with people and see what people's own takes are and just sort of manifest the really cool ones. Awesome. Well, thank you. Thanks so much. Yeah, really appreciate it. Yeah, yeah, yeah. Next, we explore the issue of benchmark contamination, an issue raised by Horace He on GPT-4, Susan Zhang on Microsoft's Fi 1.5, and you heard Aman Sanger discuss SWE bench contamination in the Graham Newbig discussion.
1:27:56This next paper won an outstanding paper mention for their simple canary-free contamination detection technique. Hi everyone, it's an honor to be here. My name is Jonathan, and I'd like to talk to you today about test-set contamination. So recently we've seen large language models show remarkable performance on many challenging benchmarks. And it seems like almost every week there's a new open-source model which comes out and tops the leaderboards. What's driving this performance gains that we're seeing in unsupervised learning has been massive pre-training data sets collected from the Internet.
1:28:27So to give just one example, here's a breakdown of the pile. This is a large, diverse data set, which many open-source language models use for training. And you can see it's compiled from many different sources. So here in blue, we have academic sources like Archive and PubMed. In green, we have Internet-based sources like Wikipedia and Stack Exchange. There is pros in here. There is code and math and so on. But because of the scale of modern pre-training data sets, which are often on the order of trillions of tokens or petabytes of data, it's difficult to know if there's good separation between the training process of the language model and the benchmarks that we evaluate on.
1:29:02So to give you an example of how contamination might play out, let's say you have a language model that you want to evaluate on a coding task, like Code Forces, for example. So Code Forces is a very commonly used benchmark, and maybe somebody uploaded it to GitHub, and then a webcaller found it, and as a result, it ended up in your training data. And so when you see an accuracy number or score on some benchmark, it's difficult to know if that's a result you can really trust because of the risk of contamination. So naturally, this brings us to the following question, which is how can we identify when a language model has trained on a test set or benchmark?
1:29:40And this is a difficult question to answer because many of today's top-performing LMs are either closed behind APIs, or even if they're open source, their data sets are kept secret. So this is a figure from the Foundation Model Transparency Index by Stanford CRFM. You can see here in the first row that when it comes to pre-training data, there is very little openness in the industry. I want to show you an example of the kind of discourse that's happening around test-site contamination today. So these are screenshots from a very prestigious academic forum. It's called Twitter. So on the left, Horace is saying here, isn't it suspicious that if you look at the performance of GPT-4 on CodeForce's problems introduced before 2021, it scores 100 %?
1:30:23But if you test the model on recent problems, the performance drops to zero. Yeah, that seems a little bit suspicious. On the right, Susan points out that if you give PHY 1.5, so this is a language model trained by a team at Microsoft, if you give PHY 1.5 the first half of an example from a data set of math problems called GSM8K, it'll complete the second half perfectly, which also seems a bit suspicious. Of course, it's possible that these things happen by random chance, but the point is that this is circumstantial evidence at best, and it's clear that we need some way to audit or test a closed language model for contamination in a way that provides rigorous proof.
1:31:00So there's been a lot of exciting work on test-set contamination recently. I want to highlight three works in particular that are actually all at iClear this week, and I would highly encourage you to meet with these teams and learn more about their work. These are all excellent papers, but for our work, we were specifically interested in detecting contamination with provable guarantees and the false positive rate. And that's what I want to talk to you about today. So the question we're interested in is, is it possible to prove in a statistical sense that a language model was trained on a test set without access to the data used to train the model?
1:31:34So to describe this goal more formally, our setup is as follows. So we're given some test set X and the ability to evaluate log probes of text under a language model theta And what we want is to develop a statistical test in the classic frequency sense Which guarantees the type one error rate of at most alpha So here we're framing contamination as a statistical dependence between the model and the test set And what that means is that we're going to test the null hypothesis that the test set X and the model theta are independent random variables so just to be precise the randomness of the model here is determined by the random draw of the training data which may or may not contain the test set x so how can we do this how can we accomplish this in order to make this possible we're going to exploit a property which is true of many test sets which is called exchangeability so what do i mean by this so typically a test set is just a file, where on each line we have an example.
1:32:35But the order in which the examples appear in the test set doesn't actually matter. The examples are exchangeable. I can show you the examples in any order, and it's still the same test set. So formally, what exchangeability means is that we can permute the examples in the test set without changing the joint distribution of the data, which means that the model should have no inherent preference for the ordering of the examples. It's worth noting that exchangeability is a strictly weaker assumption than IID, which is an assumption we make all the time about our data on machine learning. So however, if a test set was leaked into a model's training data, because of the way pre-training works, where we take as many tokens as fit into the context window, the model would see multiple examples in a row and then memorize something about the order of the examples in the test set.
1:33:22So our key insight here is a preference by the model for a canonical ordering of a test set must be a result of contamination. So using this idea, there's a very simple permutation test we can construct. So the key here is to compare the likelihood of the original ordering to the likelihoods of shuffled orderings. So for example, for the original ordering to have the highest log likelihood under the model than any ordering over, let's say, a million random shuffles, there's a one in a million chance of this happening under the null hypothesis, right? If the model didn't see the test set during training.
1:33:57So more formally, we can draw random shuffles, which is shown here as x sub pi. And we can compare each of these shuffles to our original sequence order, which here is log p sub theta of x. And by doing these comparisons, we're estimating the quantile of the log likelihood of the original ordering. So this ratio here turns out to exactly be the p-value of a permutation test. So this test works quite well, but it's computationally expensive because of the number of times we need to permute the data set. And it turns out that we can do something a little more clever by aggregating a number of smaller tests, and that's what we call the sharded rank comparison test.
1:34:35And just for the sake of time, I'll refer you to the paper for more details. So how well does this actually work in practice? So in order to validate our test, we needed to have a language model for which we knew the training data was contaminated. So we decided to train our own model. So we started with a data set of 20 billion tokens from Wikipedia, and we pre-trained a 1.4 billion parameter language model on this data set. But first, we took a collection of benchmarks, and we injected them into the training data at random positions. And then we wanted to see, so can we actually detect contamination in this case?
1:35:09So here are the results of that experiment. They're in this table. So here, each row is a test set that appears in the training data, and the data sets were injected at various duplication rates. So some of them are in there one time, some are in there 10 times, and so on. And in the two columns on the right, we have the results of both the permutation test and the sharded test. So these numbers are p-values, so the smaller the better, or the stronger the detection. And using a typical rejection threshold of 0.05, we find that for test sets which appear 10 or more times in the pre-training data, we can detect them 100 % of the time.
1:35:40So we get perfect detection. we also found that detection at a duplication count of one is quite challenging so we don't currently have a statistical test that works for duplication counts that are that low and we'd like to encourage the community to continue working on this problem so one question you can ask is you know what point does contamination become detectable so we found that for a duplication count of four we can detect test sets about half the time for duplication kind of two we can detect test sets some of the time. So we are able to detect contamination at low duplication counts, but just not for a duplication count of one.
1:36:15So you're probably asking yourself, okay, you know, what about real models? Can we identify provable instances of contamination in LLMs that are in wide use today? So before discussing these results, there are a couple of important points I want to make. The first is that absence of evidence is not evidence of absence. So just because the p-value is high doesn't mean that there's no contamination. It just suggests that contamination is unlikely, at least at high duplication counts of more than 10. So it's a very particular claim that we're making here. The second is that there are a lot of hypotheses being tested here, and it's possible that some of these p-values will be significant just by random chance.
1:36:53So it's typical to do what's called a multiple hypothesis testing correction, and I'll refer you to the paper for discussion on that. So our first result here is that we didn't find evidence of contamination other than of MISTRAL7B and ARC-EZ. And with a multiple test correction, the p-value is just barely below significance, but it's still significant. The second result here that's of note is on MMLU. So the MMLU row, you'll see, has a little dagger there. And the reason for that is because some of the test sets in MMLU were not exchangeable, and we had to filter them out. So it's not quite the same as the others, but our findings are consistent with a LAMA-2 report, which finds mild evidence for contamination by MMLU in their pre-training data.
1:37:36So I think the main takeaway here is that it's probably not the case that popular benchmarks are being duplicated numerous times in the training data of top-performing language models. Instead, it's likely that if these test sets appear in training data, that they appear at low duplication counts. So then the question becomes, how much does low duplication count contamination affect performance on benchmarks? I think this is hard to say, But in some sense, this work suggests that there's an upper bound on how much contamination is out there for today's highest-performing language models. So to conclude here, we covered three important points today from our work.
1:38:12First, we showed that it was possible to get provable guarantees on detecting verbatim contamination by leveraging exchangeability and benchmarks. So this is exciting because it opens the door to potentially very principled approaches to contamination audits. Second, we show that these tests are effective, but there's also a major open problem that we can all make progress on, which is contamination detection at a duplication count of one. Finally, we tested existing public language models and don't find evidence of contamination, at least with high duplication counts. And this could be due to a number of factors, like deduplication being very common these days.
1:38:48But addressing the case of duplication count one would allow us to give more conclusive results in public audits. So we think this is an exciting first step, and would like to see others make progress on low-duplication count detection. And all of the models and benchmarks used in this project are available to encourage the development of future work. Finally, I just want to acknowledge my co-authors for their hard work, and Stanford's CRFM, especially David Hall and Percy Leong, without whom we wouldn't have had the compute for this project. Yeah, thank you very much. Our last Benchmarks paper feature is Gaia, A benchmark for general AI assistants by Meta AI under Jan LeCun and Clementine Foria, who runs the Hugging Face Open LLM leaderboard.
1:39:34We talked to Thomas Cialum, who led training on LAMA 2 and 3. Special thanks to listener Mokhtar Shemoblokolov for the personal introduction. Well, that's Gaia, the general AI assistant, Ben Frank. So for the background, we were at Meta doing a workshop. and in the room we had like so with Yann Lecun and others were arguing is LLM all you need or not and at some point Grégoire and I like were... Let's do a benchmark instead of qualitatively testing those models. Let's create this benchmark where we believe that what are those capabilities where models are failing. In the history of evaluation for LLMs we started with some simple problems like the squat some time ago and then it gets better so you move to all the multitask questions like glue, super glue and then because it was solved so fast we move to actually harder questions and so by harder the exit was taken.
1:40:27But even like now to expert tasks like MMLU it's interesting because MMLU I think the scores now for models is like more than 80 but so human is 90 but if you ask actually that's human experts in the domains if you ask a random human the score actually I think like 35 and the same like with like GPT-4 you know on a value on the exams and my point was like you know those are like results on expert domains but actually if we think about like simple tasks that an assistant would do that was all the point of Gaia. Model actually completely failing that was my intuition while humans would be like 90 % so it's a complex task that require multi-step reasoning multi-information problems and an open world browsing the web passing the information etc.
1:41:18Like humans if you give it enough, if you give a human enough time you will solve this problem that's what we build with Gaia and the humans obtained like random amounts 90 percent gpd4 obtained only 10 percent on the level one with some tools and feel like that so that was the point of the thing and to our point arguably if you want like reasoning general capable agents whatever llm's or whatever ai system you want to solve those types first like it's fine to answer questions about what a phg student will ask in a physical like chemistry or whatever if you don't solve those stats there's a problem and so that's all about Gaia we designed like a bunch of like very small but very qualitative task questions that are very complex from level one to level three level one are like you need one or two steps of reasoning of browsing it's kind of simple level three it's extremely complex and takes like up to 10-20 steps and in an open problem like the complexity of the exponentially increase because you can like you're in open-ended world where you need to browse you need to use maybe multimodal abilities etc and I think it's interesting because very something with agents who are by LLM but using memory, tool use, like planning systems, like Friday in OS Copilot or there's a recent like Autogen from Microsoft, you move to like 10 to 40 % already.
1:42:51So there's something happening there and we think that's actually a very good way also to evaluate how intelligent are those LMS. Like what if you put Lama 3 or GPT-5 in this model, what will be the boost of performance You are also on the Llama 3 team, right? Yeah. Did you already test Llama 3 on this? Not yet. It's still in development. You mean Llama 3 is still in development? Yeah, but we will definitely, eventually we will move to improve Llama in this direction. How would you compare this versus the other generalist agents' benchmarks? I just talked to Graham Newbig, who did WebArena. I think there's some specificities about Gaia.
1:43:32that. That's why we designed it that way. So first of all, you don't need an environment. Most of the agent management, like you need an environment, it's complex and it's kind of limited by design to the synthetic environment. You just have a prompt and that's it. You have access to the world, go for it. So it's non-deterministic though? It's not deterministic. I mean it is deterministic, but by the way we created the question. So one thing is that the questions are somehow unrealistic because in practice, for instance, one question you will say is, we have this question actually in the dataset, how many BERT has layers?
1:44:14How many BERTs? BERTs, the modern BERT has layers. And actually the annotators were disagreeing because you have different size of BERTs, different things. And so for making sure that basically the question have a unique answer, there's zero ambiguity. we add to the question a lot of additional information like according to Wikipedia of these dates blah blah So that there's a single and unique answer to make the model So we can evaluate the model automatically in real life. You probably will have easier question with a bit more ambiguity Like what are your thoughts on just contamination proofing these kinds of benchmarks?
1:44:51Because a lot of these LLMs are going to be online like or like the knowledge cutoff is updated Yeah, I mean, I think there's two things that made actually Gaian bulletproof to organization. Bulletproof? Yeah. All right. I would say so. That's a big claim. I mean, so two things. One is, first of all, we kept the test sets aside. We didn't release it. So anyone can put on the leaderboard their results on the test set. To solve it, you will need manually to answer the question of the test set, go for each, spend like one hour for each of the questions and provide the answer that will be honestly like a pain in the eye now that being said the thing is it's super easy to create new questions right so if tomorrow you cheat and you say like you claim a new super high result i can get with you like 10 questions test your model if it fails as a problem so it's super easy to detect these kind of things and the last thing is also we ask in general the people to report the trace of the answer so you can very easily verify what has happening with the model so it will be in practice extremely hard to cheat on the leaderboard or if you do that you will never have a model to show to the people actually so in that sense it's bulletproof all those questions are not knowledge facts but things you will find on internet that is not present you cannot go to the answer just by memorization of the train of the self-supervised massive web itself.
1:46:19An example of that question is like if I ask you when is born Louis14 according to the Wikipedia page but something you can remember from the training. Now if I ask you how many times Louis14 was mentioned with this way to write it in the Wikipedia page that's an answer you will never find the answer right in the Wikipedia page. And we kind of oriented the question in this kind of sense. What do you think the performance is bottlenecked by? Is it tool use? Is it planning? What agent capabilities have revealed in your testing? That's a very good question. Actually, I've told them I plan to work on that, so I'll have more insights.
1:47:00But my two cents right now is, one, a general system that can be powered by LLM, and we're starting to see them. So now we can add more planning, backtracking, things like that to improve them. This is one of the core pieces that is missing. And the second thing is all of them are powered by LLMs. And the smarter this LLM will get with scaling, the better also the performance will improve. So there's, I would say, these two core things. The overall system powered by LLM, improving that, giving more tools, more abilities, more capabilities, is trade its own tools, leverage its own tools, planning, backtracking, and all those things, and improving the LLM itself.
1:47:42Yeah. Backtracking meaning like the backspace token or something else? Meaning like non-autoregressive decoding, but at the system level, not at the LLM level, in the sense that you can add something, you can try... Like traversing a graph, like a tree of thought or something. You have your plan, but in practice, I mean, if I want to go to Vienna, and I want to book a plane, but actually they are all full, I will need to adopt my plan. DeepMind presented a paper, right, like the web agents thing with exactly the Vienna travel example. Yeah, exactly. I wonder, they probably haven't evaluated their agents on your benchmark and I wonder what it would take to cross-pollinate the different labs, agents and benchmark ideas.
1:48:27You know what I'm talking about? DeepMind has their own stuff, you have your own stuff and I don't really see a crossover very much. Yeah, that's a very good question. The thing is, at Meta we are really pro open source. Yeah, you're the most open. And we released it with, I mean, it's not just Meta, but also with PluggingFace, this paper. And we just made it open source with a leaderboard, so that it's actually, everyone can use it and benchmark on it. Yeah. So, now I'm really looking forward to see others, like, put some riddles there, push the numbers, put some pressure on the numbers, and see how far, like, I mean, that's the only way to me that we can accelerate the community to go, like, to more capable models.
1:49:06Yeah, that's all the questions. Anything else we should have asked you? I'm looking forward to see what the community will push and how fast we will get to solving this vet match. I'm really curious to see it. Yeah, me too. Something we don't often get to feature are the invited keynote talks that start every morning of conferences like ICLR. There were great sessions on legal and copyright risk, the road to AGI and even Devi Parik's career stories. But this time, we are featuring Moritz Hart's talk on benchmarks for its comprehensive walk through history and call for action on greater thinking on the need for more scientific benchmarking.
1:49:45It's a pleasure to be here. I'm very humbled to be here. And I'll tell you about the emerging science of benchmarks. And this is not about why a particular piece of machine learning works. This is sort of about why the machine learning community as a whole works. It's pretty much everything we know about it, and it's going to be a relatively short talk. Okay, so I'll start with a quote. It's a famous quote, and it says, the only principle that does not inhibit progress is anything goes. And that's what philosopher Paul Feyerabend argued about 50 years ago. And many smart people agree with the statement, and many disagree, and there's been a lot of debate about it.
1:50:22But whether you agree or not with us, I'll leave that up to you, but it's pretty clear that the machine learning community, especially the Eichler community, has always embraced the anything goes, okay, and in a good way. So from its roots in sort of the cybernetics and pattern recognition era of the 1940s and 1950s, we've pretty much tried out everything, okay? And this community lets you dream your wildest dreams. You could be inspired by the human brain or child development or physics. This community does not limit you in how you come up with the stuff that you propose. It's totally up to you.
1:50:53This is what anything goes means, you're not limited in how you come up with the science that you do. And I think this has always been sort of the strong suit of this community that we don't limit people in how they work. Okay. But there's one thing we need to tame the anything goes. And that's the idea of a benchmark. It's the one rule that we have, which sort of tames this idea of anything goes. And benchmarks follow what the philosopher Michael Strevens calls the iron rule of modern science. So the iron rule is the idea that all disputes must ultimately be settled by competitive empirical testing.
1:51:30So at the end of the day, after some time, we must come together and, you know, competitively test our hypotheses or our methods or whatever we want to do and see which one works best. OK. And so the way this works in machine learning, you all know this, is that we agree on a metric or measure of success. and we agree on test cases, benchmark data, and we let people compete over the metric and the data and we rank the models, okay? So we see who is best in the end and we might pick the best performing method. So this is what we call a machine learning benchmark and it's essentially how this community operates, okay?
1:52:07So what's interesting is the benchmarks emerged. They didn't, it's not like the founding fathers of the community sat down and they said, here's how it's gonna be. were going to operate according to the following principle, they sort of came up over time, okay? And they didn't follow any a priori theoretical framework. We didn't know ahead of time what we're doing or why it was going to work, okay? And so if you look at the last 40 years, the history goes back much further, but if you look at the last 40 years, you see roughly four eras of benchmarking. The first one is the DARPA era in the 1980s where, you know, grant managers at DARPA wanted to have a way to compare scientists and their contributions in various grant proposals, and they wanted to have objective ways to compare them.
1:52:54So they thought this idea of a metric and like a benchmark would be great to see how scientists are doing against the proposed objectives. And so this was the DARPA era, and it was followed, I think, you know, if you will, by the MNIST era, where sort of, you know, benchmarks sort of came into the academic, front and center. People started using benchmarks more and more in academia, and there were starting to be public leaderboards. People were comparing methods on publicly available data, and MNIST and other benchmarks around the time contributed to this becoming a major paradigm in machine learning.
1:53:28And then, of course, the whole idea of a benchmark exploded in the ImageNet era when the deep learning revolution of the 2010s happened, and ImageNet as a benchmark was really sort of co-constitutive with the deep learning revolution of that time. So the models were sort of developed on ImageNet and tested on ImageNet and ranked on ImageNet. And this was sort of the dominant benchmark in that era. And now, interestingly, I'd argue we're in a fourth era, which I'll call the polymorphic era for this talk. And I'm not too attached to this name. So if eventually we call this something else, I'm okay with that.
1:54:02But for this talk, I'll call it the polymorphic era. And the reason is that we're witnessing sort of this radical plurality of benchmarks. We're seeing a lot of benchmarks now, thousands of benchmarks, and they take on very different forms. So they're not the way they used to be. I argue we're in a new sort of phase of benchmarking, and we're trying out many new things, in particular multitask benchmarks and dynamic benchmarks where you don't have a fixed data set. You actually let the data set evolve over time. So people are trying out all sorts of new benchmark ideas. If you're interested in more background, I'll point you to Mark Lieberman's talk from the Simons Institute from about five years ago.
1:54:40It's an excellent resource on some of the history. And Ben Recht and I wrote a chapter on this also in our textbook. So you can check this out online. But this talk is not about the history. It's about why benchmarks work. And I'll start with an outline of the science of benchmarks. What do we know about benchmarks? Why they work and when they don't work? and I'll start with sort of scientific key takeaways from the ImageNet era. With the benefit of hindsight, because it's over, we can look back and we can do some kind of retrospective analysis of what worked and what did we learn from this era.
1:55:11And then I'll move on to this era that we don't have the benefit of hindsight about that's happening right now at this conference and elsewhere and we don't really know what it's doing. And I'll point out some risks, also opportunities, but some risks of this new polymorphic era of benchmarking. And I'll wrap up by motivating why we need a science of benchmarks, why I hope many of you will join this effort, and why I think it's a fascinating research area to work in. So how did this all start? I argue that the beginnings of the science of benchmarks started with a mistake, like many things. It was an excellent mistake, but it was a mistake nonetheless.
1:55:48So the mistake that started all this is to assume that benchmarks are just the holdout method. And you still hear this today a lot. to be able to say, ah, it's just the holdout method. Okay, so what's the holdout method? The holdout method, as you read it in textbooks, just looks like this. You split the data into, let's say, two pieces. Could be multiple pieces, but for this talk, just two. You set aside the training data. You can apply, oh, sorry, you set aside the test data. You apply the anything goes principle to the training data. You can do whatever you want with the training data. And then you use, in the end, you rank the models on the test data.
1:56:20Okay, you had set aside the test data. and in the end, you rank the models on the test data. And I emphasize this part in the end because this is really crucial, okay? As you can read, you know, in this textbook by Hasty Tipshirana Friedman very recently, that ideally the test set should be kept in a vault and be brought out only at the end of the data analysis. So the holdup method only works, in theory, if you keep the test set in a vault and you only rank the models in the end at the test set. So the models had never sort of seen the test set in any way. Okay. And why is that? Because if you want to prove guarantees about the holdout method, you really need this.
1:56:59You need this vault assumption, as I call it. And under this assumption, you can argue that the test set has exponential longevity. Okay. That sounds like something Silicon Valley billionaires want to have. And I just mean that it's the number of model comparisons you can do on your test set is exponential in the data set size. Okay. So the data set can live very long. You can try out all sorts of models on the test set, and you still get good results from the test set. So under this vault assumption, the test set has high longevity. It sticks around. It can survive many model comparisons. But look, you all notice the empirical reality is completely different.
1:57:38The test set is anything but in a vault. So we have a large, sprawling machine learning community that sort of builds models in like a continual loop with a test set. So you propose a model, you evaluate it on the test set, you see what the results are, you incorporate these results into your work, and you continue. This is the whole point of science that there is this kind of closed feedback loop between you and the evaluation systems that you have. And I looked this up, and I was amazed by this. How many of you have downloaded MMLU recently? Come on, guys. Three people? That's not true. that doesn't work it was friday morning at eight people are tired okay i looked this up apparently mmlu so this is a multitask benchmark for language models that many of you know it has just 14 000 data points questions and it was downloaded apparently five million times on hugging face last month okay i found this marvelous okay it shows you how what's what scale this community is operating at it comes in two versions and if you add it up you get about five million downloads that 60 million downloads.
1:58:42If you think that every download is at least one evaluation, this test set is seeing tens of millions of evaluations per year. That's kind of striking. But it's not just the number of evaluations. It's not just how often you evaluate these test sets. It's the way you do it in the sense that machine learning is adaptive. You use the results from the test set to refine your methods. It's not like this test set is kept in a vault. It's part of the ongoing evaluation loop, and that's the whole point. OK. And so what we realized about 10 years ago is that this kind of adaptive activity, this kind of adaptive use of the test set breaks all existing guarantees of the holdout method.
1:59:19And it reduces its lifeline to just a linear longevity in the number of evaluations. OK. So the test set can only support a linear number of evaluations. You can show easy examples where this happens. OK. So in principle, this way of using the test set could be really bad. Okay. And this launched the area of adaptive data analysis where people come up with sophisticated methods to try to fix, you know, the holdup method to work better under this adaptive use. But we can go back to the empirical reality and we can see that like test sets actually in practice seem to have a lot of longevity. Okay.
1:59:55Here's a plot that probably many of you know. It's the, you know, from a plot from papers with code that you can see on the website there. It just shows you the improvements on the ImageNet ILS VRC 2012 test set over the years. And as you can tell, even after 10 years of very active use and 10 years of people hammering away at this test set, you were still seeing significant improvements. The test set was still good enough at that point to support active model development and model improvements. And it was still worthwhile and useful for model ranking. So how could it be that, you know, after such a long time, it still seems useful?
2:00:34Okay. Should we trust the model rankings? It really begs the question, should we trust the model rankings that we get out of this 10 years of active use of this test set? So several years ago, people wanted to find out, and they created a fresh test set for ImageNet. They restarted or recreated the dataset creation process for ImageNet and tried very carefully to create a new fresh test set for ImageNet. OK, this is a fantastic work. And what they found is that the model rankings were preserved. OK, so the model rankings on this fresh test set were actually preserved, you know, compared to the old test set.
2:01:09OK, so the model rankings were fine. Even after a decade of development, the model rankings were still the same ones. OK, and the people also found the same thing for MNIST, which was even more used, you know, various Kaggle competitions. It was also true and other data sets. So people confirmed this insight about the stability of the model rankings in a number of different cases. And this is what I call the internal validity of the IRON rule. Beating the previous best replicates in similar conditions. Okay, if you have a model improvement in some test condition, you will get the same model improvement in very similar test conditions.
2:01:44Okay, if you try to recreate a test set or you recreate the testing conditions very closely, you will see a similar model improvement. OK, so this is the internal validity of the IRON rule. Beating the previous best replicates in similar conditions. And we were able to even prove this. OK, so we had a work where we turned this observation into a mathematical assumption. OK, we said, what if we make the assumption that researchers only care if they improved over the previous best? So what if we make a mathematical assumption that, you know, researchers will ignore results that didn't improve and they only care about results that improved over the previous best?
2:02:22This is, of course, a simplifying assumption. But if we make this assumption, we can actually prove formally that assuming this iron rule assumption, the benchmark data has exponential longevity. So you recover the exponential longevity of the holdout method under this iron rule assumption. And I find this amazing because it says that the iron rule assumption is nearly as good as the iron vault assumption. So the idea of competition, this principle of like letting people compete is as much a regularizing force as keeping the test set in the vault. It's almost a little bit depressing if you think about it, that competition is such a strong coordinating principle.
2:03:00But this is what it is. You can think of this as something that you implement if you're really worried about keeping your model rankings accurate or protecting your test set. You can implement this by enforcing limited feedback in a benchmark. But I actually like to think about it differently. I like to think about it as a descriptive sort of theorem. It says, you know, if we think of this as a postulate about how the community works, about how scientists interact, then this is, you know, you get these consequences. OK, so under this postulate about the community, we know that test sets have exponential longevity.
2:03:34OK, and so this is ultimately an assumption about the behavior of the community. It's not just a technical assumption. It's an assumption about how we as a community organize our scientific activities. All right. And so we found many other such sociotechnical forces behind benchmark longevity, things about the community that promote benchmark longevity. Okay, competition is the one I just mentioned, but also collaboration, the fact that we all share code and we share each other's GitHub repositories and so on. This has an effect that promotes longevity. This is related to what Dave Donahoe recently called frictionless reproducibility, the fact that we have this very active ecosystem of sharing code and its important driving force in this context.
2:04:18We even found that cognitive and behavioral biases of the researcher, the fact that we're all like limited human beings and we have our biases, these biases can actually protect the test set. So our cognitive limitations can actually work to our advantage when we're doing science and aren't necessarily a bad thing. And it's great. I feel good about this. And finally, we found that basic data set artifacts like what classes you include, you know, how many classes you include. That also has a strong effect on benchmark longevity. Okay, so this is something we found a few years ago. There's a talk about it if you want to see more about it.
2:04:52These things are all what I call socio-technical forces. They're not just purely statistical. They're not just purely, you know, technical. There's something about how the community works. But all of them promote internal validity of benchmarks. Good. So what do we know at this point? So as of about five years ago, we knew that model rankings replicate under similar test conditions. If you change your test conditions slightly, you get the same model rankings. It's this kind of internal validity story. So what if we ask a more daring question? What if we ask, do model rankings replicate on radically different test conditions?
2:05:27What if we stretch this to the limit and we go to radically different test environments? Will we still get the same model rankings? Is there any reason to believe that the answer is yes? So we wanted to find out, and we call this the Image Not Experiment. And this is joint work with Ola Vales-Alodin, who interned with me last year. and ImageNet is an anti-replication of ImageNet. Okay, this sounds crazy. And it's a term I made up. So nobody I think has used this idea of an anti-replication. But what I mean by this is that it has the same scale and diversity of ImageNet, same size, but it's different in every other regard.
2:06:04We just tried to make it as different as possible, subject to the same size and scale as ImageNet. Okay, you all know ImageNet was carefully curated by humans. You had many annotators per image. You had a very high agreement rate between annotators. There was a certain logic to which classes you included, et cetera. ImageNOT is just a quick and dirty data set based on selecting images from their captions in a web crawl data set. And there's no rhyme or reason to the classes that we include. It's completely arbitrary. It's, if you will, kind of a trashy data set. So I'll be honest with you. And so the experiment we did is, what if we retrain the key ImageNet era models from scratch on ImageNet?
2:06:48So not fine-tuning, we're retraining them from scratch on this new data set, and we want to see what happens. Is it true that the model rankings are preserved? Remember, this is a retrospective analysis. We already have these models that were developed on ImageNet. Do the model rankings replicate in this radically different test environment? So the answer is yes. Okay, so for the model architectures that we carefully studied, we see the exact same rankings. AlexNet comes lowest because it's the first major breakthrough in 2014. BGG improves upon that. DenseNet improves over that. ResNet comes then and so forth.
2:07:24We get the exact same model rankings. But what is maybe more striking is that the relative improvement over AlexNet is also about the same. So if you look at the curve of relative improvements that each model makes over time compared to AlexNet, you get roughly the same curve. So what this is saying, somewhat surprisingly, is that on this completely different kind of trashy data set, this data set makes the exact same sort of judgments as ImageNet. It gives you the same information from a benchmarking perspective as ImageNet. It says this model is better than that, and it gives you the same comparisons, and it gives you the same sense of relative improvement.
2:08:04Moreover, we found out the same is true for fine-tuning. if you wondered about that. If you fine-tune instead of retrain from scratch, it's also the same. And we also looked at transfer learning and we found that sort of, there's a similar relative utility to transfer learning on ImageNet as there is on ImageNet. Okay, so how do we create this? You probably, many of you already guessed this. We use this wonderful resource called Lion, which was created from Common Crawl. It's a resource of 5.85 billion image caption pairs. It was a massive effort that we're using here and building on. and we just really did sort of an extra step on top of that.
2:08:39We just selected a bunch of images based on their captions from Lyon. Okay, how do we pick the captions, the classes? We just pick 10 ,000 arbitrary classes while avoiding all subtrees of the WordNet hierarchy that contain an ImageNet class. And when I say ImageNet, I mean ILSVRC 2012. So we make sure we stay away from all these ILSVRC 2012 classes and pick sort of arbitrary classes subject to staying away from that. And then we select images from Lion simply based on Roberta text-only similarity between the class and the caption. So you embed the class, you embed the caption, you look at the similarity.
2:09:17You do not look at the image when you select these images. You just look at caption text similarity. Okay. And then we implement some additional safety filters just to make sure that whatever we run our analysis on is safe. Okay. But the main point here is that there's no annotators involved. There's minimal human intervention. It was largely just based on a web crawl with like very sort of sloppily selected data points. Okay. And in doing this, we actually built on another work that I did with Ali Shirali at UC Berkeley, which is answering the question, what would be different if we recreated ImageNet from Lion?
2:09:53What are actually the differences in the data sets that you get? And through some really clever detective splay, Ali found out some subtle but very important differences. And we're building on this effort in creating ImageNOT. Okay. Just to give you, you know, a sense of ImageNOT, a visual sense, how many of you know the Torel by EFROS game of guess the data set? Probably all of you, right? You look at researchers, you let researchers look at two data sets and they have to guess which one it is. So let's give you some training data, okay? So here's a class from ImageNOT. It's called cleats. It contains some shoes that have cleats, you know, like football shoes, but also it contains lots of images of like just football players that apparently wear shoes with cleats, but you don't actually see the cleats in the image.
2:10:35And then it contains all sorts of other stuff like shoes with bicycle cleats and so on, right? There's another class called Batter. It contains cartoons about baseball batters, but it also contains dietary advice, like you should always batter your chicken and fry it before you eat it, which I can also recommend. And finally, it contains, let's say, pictures of this cosmetic pen called Batter Up, okay? Let's contrast that with ImageNet. ImageNet, as far as I can tell, is mostly dogs. And so there's, you know, something called an Irish Terrier. It looks like this. And ImageNet has lots of front and center images of Irish Terriers, okay?
2:11:13So, I mean, if you know what an Irish Terrier is, you can tell this is ImageNet. Even if you don't know what it is, you can sort of tell it's ImageNet because it's kind of front and center of a dog, okay? It's a front and center dog that's ImageNet. Not a front and center dog that's ImageNet, okay? There's also something called a Blenheim Spaniel. This looks like that. And again, it's lots of front and center images of Blenheim spaniels. Cute dogs. And there's hundreds of dog breeds or more than 100 dog breeds in ImageNet. Okay. So what can we learn? Okay. And just to spoil it, we can easily get more than 90 % accuracy in telling them apart.
2:11:44So these data sets are really very different. What can we learn from ImageNet? This to me is quite fascinating. It suggests that something like this might be true. It suggests that the iron rule might actually have external validity. So it might say that if you beat the previous best under sufficiently general conditions, it will likely replicate elsewhere. The only thing you need is that your original test conditions were so rich enough. But aside from that, you don't really need anything. The model rankings will replicate in other conditions, assuming your original testing conditions were rich enough.
2:12:17And I say this almost a bit more like a conjecture because it needs more work. And, for instance, we need to know what does sufficiently general mean and so on. So there's a lot of interesting work to be done on this. But I find it quite intriguing that there might be this kind of dynamic equivalence that there's evidence now that ImageNet could have been anything of similar scale. Just from a model ranking and benchmarking perspective, all these like data set artifacts that ImageNet had, they may not be all that important as we thought. And anything of similar scale might have given you the same benchmarking results.
2:12:47We don't even need clean labels. And if you know anything about ImageNet, you know how much we as a community have thought about this annotator step in ImageNet, how important we thought it is that we have annotators, multiple annotators, the agreement rate between them, all these things we thought were essential for the benchmarking enterprise. And I'm here saying that we don't even need claim labels. How could that be? I mean, that sounds almost suspicious. So we wanted to know more and we wanted to dive deeper into this claim and do some theory about it. And we studied a model, proposed a model of benchmarking with noisy labels.
2:13:21This is joint work with Florian Dorner, who's in the audience. And we boiled it down to a very simple theoretical question. So given two binary classifiers, let's say image classifiers, F and G, which one has higher accuracy? This is what you're trying to find out. This is the sort of essential benchmarking question. Given two models, which one is better? And here's the model. you can draw unlabeled data points X for free. Just go on the internet, download an unlabeled data point. And then you can get a label Y for one euro, okay? For the Americans in the room, one euro is the local currency, okay?
2:13:59And this is Austria, not Australia, okay? So quick, quick check, okay? Just making sure we're on the same page. So one euro and you get your label, okay? But the catch is this label might be incorrect, okay? So this label might be wrong with probability p less than half. So there's some chance that this label is incorrect and is not the correct label. And so how do we best spend our money? Let's say we have n euros and we want to spend our budget on identifying which is the better model. How do we most efficiently, most economically spend our money to maximize the probability of identifying the better model?
2:14:36Here's a common practice that people would propose. You sample n over k points, where k is some number, like 3 or 5 or 17. And for each data point that you sample, you request k labels, y1, y2, up until yk. These are noisy labels independently drawn from your annotator process. And you clean these labels by taking a majority vote. Binary labels, majority vote, reduces the error rate of your label. So the label y that you get by taking the majority vote You know, you clean it. It has lower error probability than the K labels each have on their own. And so this could be a good way to clean the labels.
2:15:17And the question now is, well, how large a K should you pick? Should you pick K equal to 3 or 5 or 17? What's the optimal K? And what we prove is that in all cases, it's best to sample n data points with one noisy label each. Okay? The optimal choice for K is 1. Okay. You want one noisy label each for one point. Okay. That's the result. My contribution to this project was preventing Florian from calling the paper, all the single labels. Okay. So it's hard being an advisor. Sometimes it's a thankless job, but you know, has to be done. And we got over this, but this is what it says. Okay. It says that really a single noisy label per data point is best.
2:16:07So this is what it says. And so the statement is very easy, but the proof is not. So it was actually kind of a grind. It uses Cromer's theorem from the theory of large deviations to get an exact asymptotic tail bound on the probability of not identifying the better classifier. And it extends to many model comparisons just via the union bound as you usually would do it. And for the theoreticians here in the room, this can often be a good alternative to using Hufting's bound. So I had always been like naively applying Hufting's bound, which is just an upper tail bound and doesn't give you an exact bound.
2:16:40And this gives you something much stronger. And because this proof gets a little bit subtle, Florian actually found like a really good way to numerically check this conjecture or check the theorem. And this meant we could very easily simulate all the parameter settings. And here's what you get in a typical parameter setting. So as your label budget grows, so does the number of model comparisons that you can make. And as you can see, the number of model comparisons that you can make by the single label strategy is much, much greater than for three labels or five labels and so on. And it's also much, much better than what you get heuristically from applying Huffington's bound.
2:17:15Okay. So it gives you something much better. And because this is useful independently, Florian created a sample size calculator that you can check out. If you're creating a data set with noisy labels, this gives you a guide on how to do things. Okay, great. 30 minutes, perfect. Yeah, I'm doing good. So that's sort of my retrospective on the ImageNet era. Let's leave the familiar contours of the ImageNet era and enter the polymorphic era. and I thought it was fitting to try to get generative AI to describe the polymorphic era. This took me like an afternoon of prompts, basically, but now I'm happy with the results.
2:17:55So this is what the polymorphic era looks like. And here's what I mean by that. So basically, large language models and multimodal models, in some sense, ushered in the end of the ImageNet era. They posed new demands on the benchmarking paradigm. And we were left from this ImageNet era with some suspicion and concerns about the idea of a single-task benchmark. Maybe it was just too narrow-minded to have just a single benchmark or a single task benchmark. Maybe that's not diverse enough. And so in response, people created a bunch of new multitask benchmarks. They have all these names, Superglue, MMLU, Big Bench, Helm, and so on, with the hope that these multitask benchmarks will provide a more nuanced, holistic evaluation canvas for these new models.
2:18:37We're also left from this era with some skepticism about static benchmarks, the idea of just having this one test set frozen in time, and people have been experimenting with the idea of dynamic benchmarks and response, so benchmarks that evolve over time that grow as you get different models. And so I will talk about each of these in turn, multi-benchmarks, multitask benchmarks first, and then dynamic benchmarks, and I'll give you some new perspectives on each of these. Okay, and so the first thing I'm going to talk about is a social choice perspective on multitask benchmarks. This joint works with Guan Hua Zhang, and we applied basically ideas from social choice to multitask benchmarks based on the following analogy.
2:19:19The analogy is between tasks and voters. So in a benchmark, basically tasks, different tasks act like voters, and they can vote on models. So models become candidates, and each task gives you a ranking of of all models. So tasks are voters, they vote on models, and they can rank these models. And if you think about it this way, then a benchmark is nothing other than a voting rule that has to aggregate all these different votes, all these different rankings, into one ranking. That is the problem of social choice. And it suggests a distinction that's important here between cardinal benchmarks. These are benchmarks that aggregate numerical scores.
2:19:59Examples are Big Bench, the Open LLM leaderboard on Hugging Face, and so on, where you just average out accuracy numbers. Those are cardinal benchmarks. And they stand in contrast with ordinal benchmarks. Ordinal benchmarks are just things that use the rankings and aggregate on the basis of the individual rankings. And an excellent example of an ordinal benchmark is HELM, which I'll say more about. So what does HELM do? This is a fantastic new effort led by Percy Liang at Stanford to holistically evaluate language models. That's what it stands for. And interestingly, HELM is an example of an ordinal benchmark.
2:20:35Why is that? Because HELM works with what's called the winning rate. So it looks at how often models win against other models on these different tasks. And the winning rate is something you can compute just from individual rankings. If I know the individual task rankings, I can compute the winning rate for each model, and then I can rank by winning rate. And that makes Helm an ordinal benchmark. Okay? You can compute the ranking on Helm from individual rankings only, only ranking information. That's what ordinal means. Okay? So that makes it ordinal. In contrast, the OpenLLM leaderboard averages out accuracy numbers, and that needs cardinal information, so numerical information to get the ranking.
2:21:15I should say, by the way, all of this, the entire talk is based on these fantastic contributions that the community has made in the benchmarking space. For instance, the OpenLLM leaderboard is based on the Eleuther Evaluation Harness. If you ever looked at that code base, it's an enormous effort. It's a huge amount of work. I'm deeply grateful to all the work that people have done on this. And I really want to give people a shout and encourage that kind of work. It's essential for the community. I really, really appreciate it. Okay, so cardinal benchmarks, ordinal benchmarks. Those are the two things we're going to contrast.
2:21:44and one of the robust insights from social choice is that you have no perfect voting rules. This is associated with Arrow's impossibility result. That's a famous result in that area. It says, no voting rule can make you perfectly happy. All voting rules have some issues and there are certain desirable properties that you can't all have simultaneously. What are these? Here's how I would state the theorem. I'll state it like this. I'll say that any diverse ordinal voting system is sensitive to irrelevant alternatives. What does diverse mean? I'll tell you in a minute. What is sensitive to irrelevant alternatives?
2:22:23It means something like a third candidate could enter the race and change the order of the top two contenders. You could have a weak contender enter the race and perturb the top contending models. Of course, you don't want this. Imagine you upload a weak model to Helm and change the order of the top two models. That would be unfortunate. Diverse just means that your benchmark or your voting system is not a dictatorship, so it doesn't just project onto a single task or a single voter. It's Pareto efficient. If a candidate wins unanimously in every task, it should also win overall. And it's universal in that it doesn't limit how you rank.
2:22:57It accepts all rankings. So these are all reasonable things. You kind of want to have this. And then it says you have to have irrelevant alternatives. This theorem applies to ordinal voting systems. and as a result, it directly translates to ordinal benchmarks. And it says adding irrelevant or weak models to an ordinal benchmark system can change the order of top contending models. But the issue with Arrow's impossibility result is that it's not quantitative. It doesn't tell you how much things could change and whether that's something practitioners need to worry about. And it only applies to ordinal systems, so it wouldn't tell you anything about all the cardinal systems that are out there.
2:23:37So what we did in this work is we proposed like an empirical variant of errors and possibility result that applies to benchmarks and both ordinal and cardinal benchmarks. And the key properties that we identified are sensitivity and diversity. What is sensitivity? It's just change in rankings due to irrelevant task transformations. So if you do changes to your tasks that shouldn't at all matter, does the ranking change? In the case of ordinal benchmarks, I already told you what that is. It means adding a weak model can flip top contending models. That's the case of an irrelevant change that shouldn't matter.
2:24:14And in the cardinal case, it's just monotone linear transformations of the metric. If you relabel your accuracy numbers from 75 to 80 and 80 to 85, if you just do a change, a relabeling of numbers, that shouldn't change the task. and in fact every single individual task will be identical under such a transformation. It will give you the same kind of ranking, so it really shouldn't matter. So these are the relevant task transformations. And sensitivity is how much you change the ranking in response to these irrelevant task transformations. We can always minimize sensitivity by just having a single task benchmark.
2:24:50If you have a single task benchmark, there's only one ranking, there's no impossibility. We can always have a single task or copies of a single task, and you minimize sensitivity. But the whole point of multitask benchmarks is to also have diversity, to have variance and rankings among different tasks. You want these rankings to be not all the same. You want the rankings to have diversity to tell you different things. And we just measure this with something called the Kendall's W coefficient of concordance. It's a standard measure of diversity or variance in rankings. And you want to have high diversity to have a diverse multitask benchmark.
2:25:25And we show in this work that all existing multitask benchmarks exhibit a strong tradeoff between diversity and sensitivity. So if you look at the 2D plot of diversity versus sensitivity, you'll see that all the existing benchmarks, this is the case of cardinal benchmarks, fall between a constant benchmark or on a line between the constant benchmark, which has a single fixed ranking, and a random benchmark, which just has a random ranking. And every benchmark strikes a tradeoff somewhere between that. OK, and so this says that diversity comes at the cost of sensitivity. If you want more diversity, you're going to have more sensitivity to irrelevant changes.
2:26:03OK, all benchmarks fall on this line. And you can think of this line as sort of a measure of multitaskness. There's really just one dimension here. How much multitaskness do you want? And the more multitaskness you want, the more diversity you get, but also the more sensitivity you get. And you can't sort of have high diversity without high sensitivity. As a sanity check, this was important. As a sanity check, we partitioned ImageNet into a mock-up multitask benchmark by just partitioning the classes into 20 random tasks and calling each partition or each part in the partition a different task.
2:26:35So we're creating a fake multitask benchmark. And this is what that dot that says ImageNet and the slides is, represents. It's just what happens if you create a fake multitask benchmark. And our measures correctly identify that this is a single task benchmark. So it's not better than a, it has no diversity and no sensitivity. It's still just a single task benchmark. But the main takeaway from this is that diversity comes at the cost of sensitivity. There's no free lunch in multitask benchmarks. You can make these benchmarks more diverse, but it's going to come at the cost of having very high sensitivity to relevant changes.
2:27:10Just to give you a different measure of sensitivity, on the left, we measured it in terms of Kendall's tau. That's like a measure of difference in ranking. on the right panel we measure it in terms of the maximum normalized rank change. So this is the fraction of ranks you can skip due to an irrelevant change. And you can see that, for instance, for BigBenchHeart, you can skip 80 % of the ranks by some irrelevant transformation of the metric. And for MMLU, it's close to all ranks and so on. So you can really have very significant changes in the rankings with an irrelevant task transformation. For ordinal benchmarks, the situation looks similar.
2:27:47although a little bit more messy. Here we look at various subcategories of Helm and Heim. And you see also that there's this general trade-off. I should say that we always compute a lower bound on sensitivity because it's kind of hard to compute exactly. So all these numbers are lower bounds. The sensitivity might always be higher. And so we're seeing a similar picture here, a little bit more messy. Here's an illustration of what this actually looks like, the sensitivity to irrelevant changes. Here's how much you can perturb the rankings on these benchmarks On the left is OpenLLM, on the right is Helm, by just doing irrelevant task transformations.
2:28:22And you see that the models can jump around quite a bit, and these rankings can have quite a lot of sensitivity. If you want to play around with this, you can HIP install BenchBench and play around with all these numbers. It makes it very easy to just load up your favorite benchmark, compute the diversity and sensitivity, and see where the benchmark falls. We're also maintaining a website where we add these things and sort of display them just to keep track of them. And if you want to contribute to that, please send us your benchmark. Make us aware of your benchmark. We would love to edit. Again, there's a little bit of computation involved in getting these numbers.
2:28:59We're happy to run that computation for you. Just reach out to us, and we'll be happy to add this. So this is sort of an ongoing work in progress. All right. This is my take on multitask benchmarks. And with the last several minutes, I want to talk a little bit about dynamic benchmarks, which is another proposal of the polymorphic era that's gained quite a lot of traction. There was a really amazing and fascinating effort called DynaBench a few years ago, which was very ambitious. It was trying to really change the way we do benchmarking by proposing something called dynamic benchmarks and a platform to do these kind of dynamic benchmarks.
2:29:36And so a dynamic benchmark basically is an evolving time-dependent benchmark where you have some initial data set. You let people build models on that initial data set. And then you use the models that people have built to find like failure cases of all the existing models. And you add those failure cases to your benchmark. Okay, this is called adversarial data collection. So you learn from the models that were built. You add the failure cases to your data set and you continue. you. And so there's this ongoing interaction between model builders and data collection. And you interleave these two operations indefinitely, hoping that your models just keep getting better and better and better and keep accounting for more and more challenging instances.
2:30:18We wanted to know if this works. And it's something I did in joint work with Ali Shirali and Radio Dabebe, where we proposed a theory of dynamic benchmarks. Because it's so new and so different, we wanted to know, can this at all work? And how do we even think about what it means for that to work. And this is what we proposed in that paper. And we abstract a dynamic benchmark as a directed acyclic graph with four operations. So each node in the graph is one of the following four operations. You can do model building. You can invoke the community to do model building. You can look at the resulting models, assemble them, or in some way collect them.
2:30:54And then you do data collection on the resulting model. You let sort of the annotators find failure cases and you create new data points and you add them to your data set. So you pool data. Okay, these are the four operations in a dynamic benchmark. And subject to that, you have complete freedom. You could do whatever you want. Any directed acyclic graph is a valid dynamic benchmark. The standard design that people have mostly implemented is the following, it's just a directed path. It's just a directed path that alternates between model building and adversarial data collection. So just these two operations alternate them.
2:31:28That's the standard design that people have mostly experimented with. And we prove a theorem that says that progress in the standard design can stall after a small number of rounds. So there's, in principle, no reason to expect that the standard design gives you progress beyond a few number of rounds. And you can see this in some of the experiments, that progress really seems to plateau after just a few number of rounds, and it becomes diminishing after that. And so we thought harder about this and we came up with more sophisticated benchmark designs that guarantee strictly more progress. So we call them hierarchical dynamic benchmarks where we sort of have parallel threads that you let run and merge in some particular ways.
2:32:10So these are more complicated benchmarks. We can prove that they guarantee more progress than the standard design, but they're also much harder to implement. So I'm actually not sure how feasible it would be to pull this off as a real benchmark, but it's certainly an intriguing possibility. Okay. So this is my take on dynamic benchmarks. So let me sum up and get to an end. How am I doing on time? Great. Perfect. All right. So summing up here, what do we see? I use the benefit of hindsight to sort of do an ImageNet era retrospective. What did we learn from like 10 years of benchmarking on ImageNet?
2:32:47And the takeaway here, the main takeaway here is that the IRON rule has both internal and external validity. So it works to a surprising extent in this case. But we know much more about the former and much less about the latter. So if you're interested in open problems, understanding this last part is really challenging. And we have made very little progress on that. So there's definitely more work needed here. And then maybe to me what was really surprising is that good human annotated, highly curated data is not necessarily required for ranking models by accuracy. So if your goal is just to do rankings, performance rankings by accuracy, you don't necessarily need super clean data.
2:33:28We saw this empirically, but also we can do some theory about that. And this was surprising to me. I should qualify that, of course, if you're doing something like a fairness analysis, a safety analysis, bias analysis, you know, red teaming or alignment analysis, this is, of course, very different. Then your test cases and your data really matter substantively, and what you put in there is very important. And my group's certainly very invested in these research directions as well, but that's not what I talked about. In this talk, I was focused on the core sort of benchmarking enterprise and ranking models by accuracy.
2:34:00Okay, then I moved on to this new era where we don't have hindsight, where we're sort of figuring it out as we go along, and we're seeing what happens. And here we saw that there's a strong tradeoff for multitask benchmarks. And diversity, greater diversity in multitask benchmarks inherently comes at the cost of less stability. Okay, so you have more sensitivity to changes that shouldn't matter. And there's no free lunch for this kind of evaluation paradigm. And finally, dynamic benchmarks are intriguing. I personally find them extremely fascinating and intriguing. But currently, I don't think we quite know how to pull them off.
2:34:36And as we currently do it, progress might stall. And again, if you're looking for good research directions, it's a fascinating area to work on. And I encourage, you know, especially theoreticians to look at some of these benchmarking questions where the landscape is wide open. OK, it's really like completely open space. Good. So this talk was about the emerging science of benchmarks. and I argued that machine learning is the anything goes principle plus the iron rule. And some of these two basic ingredients give us like a very powerful scientific machinery that seems to work. Okay. And I'd say that this community is extremely good at the anything goes part.
2:35:13Okay. So we've been extremely good at the anything goes, but that places all the burden on the second part, the iron rule. And that becomes sort of the critical link here because it's really essential for making all of this work. And so what's somewhat challenging here is that our intuition about benchmarks can fail and often has failed. And certainly my own intuition about benchmarks has very often been wrong. And I sort of was corrected by the passage of time or theoretical results that overturned my intuition. And so this has convinced me that we really need sort of scientific foundations of the iron rule itself, of like how do we make sense of this benchmarking enterprise and how do we build things that sort of promote scientific progress?
2:35:53So really think of this as like a major theoretical and empirical effort. It's not just a theory thing. It's also just an empirical effort to understand what collective practices of our community promote scientific progress. I think this is what this is ultimately about. And we should devote some more time and attention to that. And so I hope you'll join this effort. I hope you found something interesting in this talk. Please come chat with us. I'll give a shout out to the social foundations team. They are actually wearing the yellow sweaters today. I didn't think so because it's warm. But if you see us somewhere in the conference, please reach out and we'll be happy to chat.
2:36:28With that, I'll thank you all for your attention. That brings us to the end of Section B, our selection of three papers and one keynote on benchmarking. Phew, that was a long but important topic. And if the AI Engineer Conference submissions were anything to go by, the topic of evils and benchmarking is only going to explode this year. We turn now to Section C, which covers incremental papers and talks in reasoning and other post-training elements. There were many, many more papers than we could fit in this category, but we will focus on a few important themes. RAG, verification and safety. First, let's start with the self-RAG paper.
2:37:11This paper bears some similarity to the PAUSE tokens paper we covered in ICLR Part 1, because it involves adding special retrieval and critique tokens. However, these tokens don't just improve raw reasoning. They explicitly support retrieval during the generation process, as well as evaluating source relevance, degree of statement support, and response utility. And they can be fine-tuned on top of pre-trained open models like LAMA. Hi everyone, I'm Akali from University of Washington. I'm excited to present our work, Set of RAC, which is a new framework to improve the standard RAC system by training and making an any-language model to decide when to VTB, generate, and self-revaluate.
2:37:56Large-language models are powerful, but there are many issues like hallucinations. VTB-Augmented Generation, or RAC, has shown to be quite effective to overcome those issues. Given the user query, such as where is iClear 2024, we first retrieve a set of documents using a retrieval system like Google Search or Beam 25, and then we augment the original language model input using those retrieval documents and also the original question. The standard state of the GPT-4 or LAMMAT-3 can use those retrieval documents and give more factual, up-to-date, and attributable answers. RUG has shown to be quite effective in many benchmarks, especially in question-housing tasks.
2:38:40In a prior work, we conducted a large-scale analysis of 10 different language models, and we have seen that this inference time augmentation can give us significant improvement across many models ranging from 1.3 billion to 175 billion. RUG has been used in many real-world applications. For example, there are many language model-based search systems, Publixy.ai or Bing Chat, or libraries that help you to build a customized RUG pipeline. While RUG is super effective, there are many limitations. In this talk, I'd like to highlight two of their limitations, unreliability and also inefficiency. Now let's ask another simple question.
2:39:19How did US states get their names? Unlike the previous question, where you can simply extract local information in given documents, in this question, you have to collect a set of document and compose output based on those multiple documents. In standard language model, as in the previous example, we can concatenate those repeatable document and then fill them together in language model. The output look plausible, but there are many factual errors here. So now let's talk about why, let me explain why those standard RAC system may not be perfect yet. First, even the current sort of systems can easily get distracted when many documents are given, especially when some of the documents are irrelevant or unhelpful.
2:40:03Let's take a closer look at this example. The middle paragraph only states the history of Michigan without saying anything about how this state got their name. Current NUC system can easily affect it by those unhelpful documents and can generate factoring incorrect statement, as in this example. Due to this unhelpful document, the model states states such as New York and Michigan are named after an individual person, which is factually incorrect. Or conversely, current RAC system can ignore the context that provide useful information. Here, the first RAC provides useful information, but the model can still generate hallucination.
2:40:39For example, it can say some states, including Utah and Washington, are named after indigenous communities. Another issue is that many of the existing RAC pipelines assume retrieval is almost always necessary and keep retrieving a fixed number of documents, even when the user query isn't using retrieval. It is estimated more than 60 % of the user input into chat GPT is for writing assistant or creative writing, which may not require factual grounding. Like this query, write an essay about your summer vacation. And as you can see, retrieving a set of documents talking about definition of summer vacation or movie title summer vacation doesn't make any sense here, but most of the standard RUG pipeline just retrieve a set of documents as in other queries.
2:41:25So this causes additional latency, making the RUG system more inefficient, and also it can hurt the model output as shown in the previous slide. In this work, we introduced a new framework, self-recognitive generation, or self-RAG. Self-RAG introduced a novel inference and training pipeline to enhance the reliability, efficiency, and versatility of RUG systems. In self-ag inference, at a higher level, we make a language model to decide when to retrieve and generate and evaluate its own output. In particular, given a user query, the model first evaluates if we need to use retrieval node. If we don't need retrieval, then the model just acts as a standard language model to avoid unnecessary retrieval.
2:42:08If the model thinks we need to retrieve, then the model retrieves a set of documents, but instead we concatenate everything and field them together in input space, we first asked the model to evaluate which documents are actually helpful. And then the model would generate output, but the model first also evaluate if the output is supported by those helpful citations or not. And at inference time, we prioritize output that are fully supported by the context. To achieve this iterative self-recrective pipeline, we conducted one of the RDS instruction tuning with VTUBO and trained an arbitrary language model with those self-recreaction tokens.
2:42:44So now let's dive into the details of self-frag inference using the same example. Given the same question, how did US state get the names? Self-frag can even start direct research generation to say like US state got the names from variety of sources. But now, self-frag needed to generate the factual statement composing multiple documents. So the model outputs a special command retrieve, and then we trigger retrieval, and retrieve a set of documents from the data store. Instead of concatenating everything and feeding them together to language model, in self-frag, we process multiple documents in parallel using batch recording.
2:43:25This enables more efficient and scalable inference, but more importantly, this enables us to control model's behavior via self-reflection. Specifically, we make the self-frag to predict generate condition on a user input and also each paragraph, and for each paragraph, the model first predicts if the paragraph is related to the question or not. Then the model keeps generating the output, as in standard language model. In the first example, the model, the provided paragraph is helpful, while in the second example, the paragraph only states the history of Michigan. So the model first generates irrelevant token, and then keeps generating.
2:44:04After those generations, CellFrag also predict its output is fully supported by the related paragraph. For the first generation, model output is fully supported by the document one saying that 11 states got the name from individual person. Well, for the last case, the paragraph only talks about Utah, while the model output mentioned Utah and Alabama. So this output is only partially supported. In that case, the model predict is partially supported. So how do we achieve this retrieval and self-recreactive feature. We equip language model with those ability by making the model to learn to predict those special tokens, which we call deflection tokens.
2:44:43In standard language model, we have a fixed set of vocabulary and the model assigns the highest token probability to the most plausible next token. In self-frag, we expanded the original vocabulary using the special tokens, which include the token controlling retival or the token controlling the self-reflection feature, critic tokens. So if the model thinks we need to use RetiBal, then the model generates the RetiBal token from the expanded vocabulary. We can also adjust the threshold to balance or like decide how frequently we use RetiBal to balance the trade-off between RetiBal frequency or efficiency and final performance.
2:45:22Critic tokens are also directly generated by the model, and we use those models to rank and choose the best K output at the sentence level. In particular, in CELFLAG, we conduct a sentence normal beam search using a fine-grained feedback scores based on those critic tokens. We use the normalized token probabilities of ideal critic tokens, such as relevant or supported, as a fine-grained feedback, and then compute the weighted sum of them as a summary score of each output. Here in this example, the first output is based on the helpful document and also is fully supported. So we give the highest score to this first output.
2:46:02Well, for the middle one, the dimension is based on an helpful document, so we assign the lowest score because the output is more likely to include factor letters. Now, let's discuss how we train Setafrag so that the model can effectively learn when to generate such reflection tokens. In Setafrag, we train an arbitrary language model to generate simulacy from both normal vocabulary and also special tokens. During training, we use another language model, which we call critic language model. And this critic language model teaches the generator language model learn to generate appropriate self-recognitive tokens on the input.
2:46:42Critic language model is a 7-video language model trained to generate regression tokens given evaluation instruction and input. For instance, here the evaluation instruction says that evaluate its output Y and to an input X is supported by detailed document D, and also the set of input X, D, and Y. Here, the critic language model should predict the supported token. One challenge is that we do not have large-scale fine-grained annotations for necessity of retrieval or self-critique, and collecting human annotation at sentence level for multiple aspects could be quite expensive. To overcome those challenges, we generate synthetic training data by carefully prompting GPT-4 with instruction and demonstrations.
2:47:27After this, we train critic-language model on this generated training data, and then we use this critic-language model and retrieve our model to augment existing instruction training data, mimicking set-frag inference. Here, given an input-output pair, we first run critic-language model to evaluate if we should retrieve, and then if the model output retrieves token, then we installed a retrieval passage in line with reflection tokens, also predicted by the critic language model. We generated 150 ,000 instruction training data in this way. Now we can simply train the generator language model in this augmented training data.
2:48:07So given the input, the model learns to generate the standard output and also reflection tokens. We can use the standard language model training objective with expanded vocabulary. This enables us to easily apply the same training pipeline to a new language model and our so-called bases, and also enables us to tailor model's behavior to diverse fine-grained preferences, as shown in the previous slide, without additional training overhead. We abutated Setaflack on six diverse tasks, including close-set task, short-home generation, and also long-home generation. Our Setaflack is based on Lama 2.7 and 13B, and trained on four GPUs.
2:48:50Now let's discuss the result. As you can see, the baseline parametric language model without any retrieval struggle in tasks requiring precise knowledge memorization, such as POPQA, an open domain QA data set, or ASQA, which is a long-term QA data set about factual knowledge. The standard RUG gives improvement on some tasks like POPQA, where we can simply extract information from single document to answer. However, standard RUG still struggle or doesn't give large performance improvement on tasks that demand systems to compose knowledge from multiple documents and generate, such as Pub Health or ASQA here.
2:49:28Moreover, especially open access language models such as Lama 2, 13B, chat, pre-trained, or chat model struggle to obtain good citation precision or recode, indicating that even the model cites something, the model output is likely to be not supported by those citations. ZeroFlag significantly improved how to perform such models and obtained the best performance across all of the baseline in the same model scale. Moreover, CellFrag shows much better citation precision than Vcode. CellFrag even matches our out-of-home chat GPT on five out of six tasks, despite being relatively small seven or 13 billion model trained on a small set of instruction tuning data.
2:50:13CellFrag was initially introduced last October, and since then CellFrag has been widely used in both academic papers and also industry applications. For example, CERFAC has been successfully integrated into multiple wider-use RUG libraries, such as LongChain or LamaIndex, and also there are many papers that try to improve CERFAC for further reliability, or apply CERFAC to new domains, especially safety-critical domains like biomedical domains. In summary, we introduced a new CERFAC framework, which helps us to build a more reliable, efficient, and versatile RUG systems. We open source the code and model checkpoint, and I'm happy to answer any question now.
2:50:52And also, please come to our poster in 63. Thank you.
2:50:59The self-rag paper mentions training three models, a retriever, a critic, and a generator model. The role of the critic model and the critique tokens is surprisingly similar to the ideas of this next paper from OpenAI. Let's verify step by step. Here they pursue more formally verifiable correctness in math problems and find that process supervision, which is what Self-RAG is doing, soundly beats outcome supervision. The high level goal here is to train really reliable reward models to grade different math problems. We want to have reward models that can look over a huge number of solutions and pick out the one that actually managed to solve the problem correctly.
2:51:40And so we're comparing two different methods, what we call process supervision and outcome supervision. Outcome supervision basically means giving feedback to the model just based on whether it reached the correct or incorrect answer at the end. And process supervision is giving granular feedback for each individual step of the problem and specifically saying whether each step is correct or incorrect. And the hope with process supervision is that it's solving a hard credit assignment problem that outcome supervision would have to solve by directly giving it feedback on whether each step is correct or incorrect.
2:52:13So outcome supervision has to somehow infer where the bad step happened, but with process supervision, we get humans to just directly specify where that happens. So we collect a huge amount of data from contractors who just are going through and labeling each individual step, and then we use that to train the process supervised reward model, and the outcome supervised reward model, we just train to predict the correctness of the final answer, and then we compare these two reward models, and we find that the process supervised run is significantly more reliable. And it's specifically on the math domain?
2:52:46Specifically math, yeah. How often do you find that the reward model is wrong and it's like kind of guiding it off track rather than on track? We are using the same generator model throughout this, so we don't actually ever update the generator. We're just using the reward models to basically search and test time to search over a large number of solutions. So, I mean, both the outcome and process supervised reward models are imperfect and get tripped up by things. I think the outcome supervised reward model gets basically, I would say, pays a little bit less attention to detail. It's more likely that a solution that looks superficially good but has some subtle mistakes will still be rated highly by the outcome supervised reward model, whereas the process supervised reward model is a bit better at spotting subtle errors, I would say.
2:53:30But both of them can get fooled by a solution that kind of looks like it's good but isn't quite right. What is the backstory behind this? Why is this an interesting area of research? Were you actually investigating something else and you found this to be a blocker for you? Broadly speaking, we're just interested in pushing on the reasoning abilities of these large language models. Reasoning is something that LLMs have struggled with a lot, so we're just excited to push back that frontier however we can. Do you have a definition of reasoning beyond just this math step-by-step reasoning? I think it's hard to give like a, yeah, a great definition to what like defines reasoning.
2:54:11I mean, to us, it's just, it's things that it's kind of whatever passes our vibe checks, but obviously things like math and STEM and code. But yeah, I don't have a good like formal definition of reasoning, I guess. I noticed the search space was actually pretty big. You had like thousands of solutions per problem. Was that necessary? Well, it's useful to show that, you know, if we only took like 100 solutions here, we wouldn't see as much of a difference between outcome and process supervision. But if we take, you know, many thousands, we see much more of a difference. And yeah, I mean, it's useful to explore this regime because, you know, if all it took to solve really, really hard problems was taking a huge number of samples, then we'd be very happy.
2:54:46You know, we'd be willing to spend the compute to generate all these solutions if we could always pick out the one that was like really good. And so, yeah, so it's not like strictly necessary to go here, but it just makes it clear sort of what the trends look like. How do you envision other people should use this? Once you have a really good supervised reward model at a larger scale, you can use it to explore labeling data for smaller scale models. Like to bootstrap or teach? Yeah. Okay. And so what we did here is we were able to do some ablations on outcome and process supervision by using GPT-4 to label these smaller model samples.
2:55:20And here we collected basically an order of magnitude more human data than we did at large scale. And by having the larger model provide the labels. and those kind of experiments wouldn't have been possible if we had to actually rely on humans. That's one possible use for the data set. We release it also just like if anyone else is able to train these process supervised reward models and maybe see something we missed. Well, that's pretty positive, I think, for the community. What's the relevant literature? What's the inspiration, I guess, or relevant literature for process supervised reward models?
2:55:53Is there anything that you point people to to read up? From an alignment perspective, some people are excited about process supervision just because, you know, it's sort of outcome supervision is providing like less direct feedback. And so it's, you know, potentially not as good from an alignment perspective. You don't really know what you're reinforcing if you just give the model like a positive or negative reward at the end of some long trajectory. But with process supervision, you know, you're kind of much more directly reinforcing the things you care about. But there's some blog posts and stuff that we cite in the paper, but I don't have anything to specifically call out.
2:56:30This seems like the inverse of this is weak to strong generalization. Is there any parallels with here is kind of strong to weak, and then you started working on weak to strong? Strong to weak in the sense of experiments I was talking about here? Yeah, I mean, I'm sure there are connections. We didn't... Okay, so you can take a small process supervised reward model and try and have it evaluate larger model samples, and it'll work to some extent, but yeah, I mean, that's definitely something you could look at more. We didn't push very hard in that direction though. I was just kind of curious if there was some kind of parallel in there.
2:57:02Cool. I think that's it. Thank you. Great. Thanks. If you listen closely, OpenAI's focus on verification and process supervision as a path toward better reasoning, aka GPT-5, is by now clear as day. In the closing workshops of ICLR, we even managed to catch a little bit of Noam Brown's talk expanding on what he calls the generator verifier gap. We last discussed his work with regards to our code interpreter is GPT 4.5 post from a year ago. He now draws a lot of inspiration from the AlphaGo paper and in short, suggests that the best path to improve reasoning in generative AI if we have sufficiently good verifiers of their output, preferably in process as suggested by Verify step by step.
2:57:50Here's an incomplete snippet of the talk leading into his discussion of the verify step-by-step paper. Okay, so how can we take advantage of generator-verifier gaps and take this? Sorry, before I get to that, one reasonable question that I'm sure a lot of you are wondering right now is, okay, if we have this generator, if we have this really good verifier, we're able to generate a bunch of solutions and filter out the ones that aren't good and the ones that is good. Can we just update the generator with the verified solutions? And in principle this is a good idea, and it's actually what AlphaGo does, that's how AlphaGo is trained.
2:58:23That said, there's a couple caveats to this. It doesn't capture all the benefits of using a verifier. So for example, if I go back to the AlphaGo slide, so you can see the gray bar is the performance of AlphaGo zero if it was trained with multiple-religious research and then doesn't use multiple-religious research to test it. So it's updating the generator with the verified outcomes of unreligious research, but it's just not using it at Tesla. And you can see that there's still this massive ELO gap. You know, it's 3 ,000 versus 5 ,200. So you're still getting a huge performance improvement by using this verification technique at Tesla.
2:59:00So just relying on updating the generator is not going to be enough to overcome this gap. And the second thing is that you're still bottlenecked by the quality of the verifier. That's something that you can't use to get around. So I think for the purpose of this talk, we'll say that updating the generator is just going to be out of scope for this fog, and I just want to focus on, let's say we have a fixed generator, and we have a verifier. What can we do? Okay, so the first thing that's really straightforward we could do is called consensus. And in consensus, you just generate a bunch of solutions and take the one that's the most common.
2:59:34Very straightforward. It's actually really convenient because you don't even need a verifier for this technique. We kind of think of this as similar to sampling at low temperature, but it's not exactly the same because, you know, if you have a sequence, then you're not sampling, like, every single step of that sequence at low temperature, you're sampling, like, the final outcome at low temperature. And I think a lot of people underestimate just how much benefit you can get from things like mid-sensis. So, for example, some of you might have heard about the Minerva paper, this came out about two years ago, and they got up to over 50 % on the math benchmark, which is a very difficult math benchmark.
3:00:10That's one they used today. But Minerva got over 50 % in large part due to using consensus. They generated a thousand samples and then they took the most common answer from those thousand. And that part of doing this consensus got them from 33.6 % accuracy to 50.3 % accuracy. So that's a big jump. Something else you can do is best event. So best event, the idea is you just sample all the solutions and then you score them with a reward model. and you return the one that looks the best. And the key here is that you have to have a good enough reward model. If you don't, you're not going to need consensus, but if you do, you can actually get a bigger difference for consensus.
3:00:50And ultimately this technique is liked by the quality of the reward model. If your reward model is not very good and you take a lot of samples, then you're going to end up overfitting to the errors in your reward model. So here for example this is a figure from a paper that came out in 2021. On the x-axis you have, so this is testing on the GSM8k dataset. And you can see on the y-axis you have the accuracy, the pass rate, and on the x-axis you have like the number of samples you suggested it. So for example if it says like 100, it means that you're taking 100 samples, 100 generations from the generator, and then you're feeding those into a verifier that's trained to tell if an answer is correct or incorrect, and taking the one that the verifier thinks is most likely to be correct.
3:01:33Now if you're taking up to 400 samples, you're seeing actually a pretty large improvement. I mean this isn't actually each model, so the fight race can get improved at all, so it's quite significant. But you are seeing like a substantial improvement as you go up to 400, but then past 400 you're actually seeing the performance degrade because you end up overfitting. If there's errors in this verifier, and if you take too many samples and ask it to score which one is best, it could actually just return one that's wrong, but hacks the verifier basically.
3:02:04Okay, so can we do better than these techniques? So one of the things I wanted to talk about is process reward models. This is a paper that's published at Inus Conference. We actually put it on an archive about a year ago. And actually, I was not on this paper. This was with my teammates. My teammates wrote this paper before I ended up joining the team. But I thought it would be a good thing to illustrate that there's a lot that can be done in this space of relying on verifiers to improve performance. So the basic idea of process reward models over outcome reward models, or simply just doing best event, is that we're going to verify every step individually rather than verifying the entire sample.
3:02:42So for example we have this question, x squared equals 4, what's x? So typically if you look at how best event was done with verifiers in the past, it would just generate the whole solution, and then you would ask the verifier, is this whole solution correct? And that puts a lot of burden on the verifier to have to consider the entire solution all at once. And with process reward models, you instead break it down by steps. And you ask the verifier, is this individual step correct? And some steps might be correct, some steps might be incorrect. So you can see here, for example, it's like going from x squared equals 4 was x, well if your next step is x plus x equals 4, then it's going to recognize that that's incorrect and correct.
3:03:23is incorrect. And so then the process reward model will take all those steps into consideration, all the scores, all those steps into consideration when it's scoring the entire sample.
3:03:34Now the way we train this reward model is by collecting a lot of human data. So we had a bunch of human annotators that would go through generations of the model and label each step as either correct or incorrect. So here, for example, there's a question and then, you know, There's a bunch of steps. The first step is let's call the numerator x, and the next step is the denominator is 3x minus 7. And the human annotator would go through each of those steps and give it like a green smiley face if it was a correct statement, and a red brownie face if it was incorrect, given what's happened before.
3:04:09So for example, you can see the last step, we go from 5x equals 6x minus 14 to so x equals 7. And that step is incorrect, so it's a marked corrected by the numerator. There's also this option for a neutral base, which basically means that it is not exactly incorrect, but it doesn't really do anything. It's just like a step that's kind of like a holding pattern. Okay, so how does the performance look of outcome versus process supervision? So the baseline that we compared against is using just like a verifier that takes the entire sample all at once. And we call these outcome reward models. So we trained the outcome reward model on a giant training set.
3:04:51So the math data set has a training set and it has a test set. We actually moved a lot of the test set problems into the training set, and we basically made the test set smaller. So we reduced the test set to only 500 problems, so we could have more problems to train on. And this actually does pretty well. So if you do best event with this outcome reward model that was trained on the math data set, you can get GPT-4 to correctly answer 72.4 % of the test set correctly. And that's better than consensus. Consensus is a technique that I mentioned before that doesn't rely on a verifier at all. You just take a bunch of samples and see which ones are most common.
3:05:25Consensus gets you to 69.6%. For the process for our model, we collected a million step level labels across 100 ,000 solutions. And this actually ended up doing a lot better. This ended up getting to 78.2 % on math. And I think this is still either state of the art or close to state of the art for math. Now, one of the reasons why PRMs are so good is because they provide a denser verification signal. It's really difficult if you have a difficult math problem to verify whether the whole thing is correct. It's certainly difficult if you're just looking at the final answer. But being able to look at each of these steps individually gives you a denser verification that makes it easier to verify.
3:06:13So this is what the plot looks like. The gray line is consensus, also called majority voting. The blue line is outcome supervision. So that's the verifier that takes the whole sample, the whole trajectory into consideration. And then the orange line is process supervised report models. So on the y-axis we have the pass rate, the success rate on math. And on the x-axis we have the number of solutions per problem. And you can see also that the performance continues to improve as we increase the number of solutions per problem. We sampled up to, I think, 1600 for this paper, but it looks like if you just keep going further, the number is going up.
3:06:53Eventually that's going to plateau. At some point it's unclear at what point it's going to plateau. So here are some samples. The print is a little small, but basically there's this problem, and then the verifier goes through each of the steps, and the more green it is, the more the verifier thinks it's correct. And on the left we have, if you just use the outcome model, so that's like if it's grading the entire sample, and in that case it incorrectly It's correctly, it labels it as correct. And on the right we have the PRM. And you can see that there's some steps where it's a little fishy. So for example, the second step, it's saying it's probably correct, but it could be better.
3:07:35And then there's certain steps where it's very confident that those are incorrect, and those are in red. So it's able to recognise that the final answer is incorrect for this reason. We apologise for the poor audio conditions of the clip, as it reflects the less than optimal nature of the workshop room. His summary makes a lot more sense with visual aid reference to the Let's Verify step-by-step paper, and we have included the two relevant charts in the show notes for his talk. Adjacent to the topic of models grading other models, we have another talk from an OpenAI team lead, Lillian Weng, Head of Safety Systems at OpenAI.
3:08:11She gave a fantastic overview of the often overlooked safety-oriented mitigations that OpenAI does for their models across four stages, pre-training, post-training, inference and evaluation. We would particularly highlight the OpenAI model spec, which lists objectives, rules and defaults that her team designs for, and the instruction hierarchy, a recently published paper that details how OpenAI defines five privileged levels that are they then fine-tuned for resolving conflicting instructions. Thanks everyone and thanks for inviting me here. I will say my talk today might be slightly different from others because I would like to give a higher level overview of how we're building the safety system for deploying like cutting edge, the best deep learning models in the real world end to end.
3:09:05So I'll probably touch on a lot of different things. A little bit on the surface because it's hard to go deep into every step. I hope you can have a concept of how complicated the problem is, also how much possible methods and mitigation you can apply when you're facing the real world challenges. A bit about myself, I joined OpenAI about more than six years ago. I initially worked on robotics and then applied research, and recently I started leading this new team called Safety System. We're only into end-to-end safety stack at OpenAI. So basically all the models that deploy to the real world will go through us.
3:09:46Our team is dedicated to ensure the safety, robustness, and reliability of AI models and their deployment in the real world. So you can think of all our, we're dealing with a lot of practical safety alignment issues. There's endless collection of problems. But the good thing is we have access to different stack in the system, so it can be pretty creative. Safety is at the very core to OpenAI's mission. If you look at the mission, our charter online, you will find our goal is we want to build safe and beneficial AGI and also deploy that to the real world to benefit humanity. I know this is a very big statement.
3:10:29So I hope after my talk, you will get some concrete ideas of how the process is. And one important point my team believe in is safe AI cannot be built in the lab. And how to deploy a powerful model, and essentially AI needs consistent learning, improvement, research in the real world. So we really embrace the idea of iterative deployment in the process. because I often find that adversarial in the real world are so much more creative than our researchers. I would say different people may have different interpretations of what is safety. So I want to talk a bit about the concept and what's the goal we try to achieve here first.
3:11:18First of all, we expect the model output to not contain any harm to people, including physical harm, mental, financial, or reputation harm. We hope the model is trustworthy, inclusive, and also respect privacy of people. At OpenAI, we want to build products that support beneficial use. That being said, for certain harmful requests, even that has utility values, we train the model to refuse those requests. And there should exist mitigation either within the model or in the system around the model for adversarial use cases. For example, some people use the model to enable fraud, scam, or do adversarial persuasion of people.
3:12:03This kind of use cases might not be easily identified if you only look at the model like a single conversation based on the model inputs and outputs. So we do need a system-level monitoring or mitigation to identify that. Also, we want to make sure our model or our system is robust and reliable, even when people try to adversarially attack it. I want to mention Goody 2 because this is probably the most interesting model I run into lately. When people think about safety or responsible models, you can go to a very extreme. You can make the model 100 % safe, but it's useless. In this case, I found this model is extremely robust, and I cannot really make the model to answer any of my questions.
3:12:57It will always find some way to say, your request has concerns, and I will just not answer that. But if you only look at the scale or measurement of how safe this model is, it's perfect, but it's useless. So I really want to emphasize that it's important to evaluate capability and safety at the same time and try to balance how useful it is and how safe it is. The Goody2 model is such an example. It refuses 100 % time, but it does matter at the point. And I also want to emphasize that these two goals are not contradictory. I usually consider safety as one capability goal or multiple capability goal among a set of rich evals you can optimize.
3:13:49So there's always, there always be trade-offs between different things you try to optimize and safety is just one of them. I kind of mentioned this before. I firmly believe safety should not only deeply build into the model, but also incorporate into every stage of the model training, deployment, and leverage a lot of things that may not happen during the training process. And we are doing this at OpenAI. We have all kinds of different safety mitigation at every stage before the pre-training step. during the post-training. We do a lot of alignment, adjustment of the model behavior. Once we get to production at inference time or on the system level, we also have quite a lot of methodology we can apply.
3:14:39And after a model is deployed, we do very consistent evaluation, red teaming, monitoring different use cases, and make sure all those feedback eventually feed back into all the previous steps. So this process continues. We embrace iterate deployment. We value data flywheel because the reward is always full of interesting challenges. And, you know, part of the research is probably find the right problem to solve. Or in this case, we don't need to think that hard. So it's convenient. Okay, next I'm going to do a slightly deep dive into post-training stage and system level, a little bit of eval, just to give an overview of what we have done.
3:15:25Post-training, it's probably one of the most powerful tools we can use. We try to align the model behavior with our safety policy. And I believe everybody here knows the concept of reinforcement learning from human feedback. You train reward model based on the pairwise comparison from your human annotator. Then you train a model to give you a scalable reward and then use that during the RL training process. So it's a fairly standard thing. We also use this process. And on the safety side, besides all the capability reward, we use Ruby's reward to adjust model behavior to follow the policy we define.
3:16:14So we first came up with a set of policy taxonomy to define how the model should behave on a safety topic. For example, if someone asks the model to tell me how to commit a crime, how to harm someone, how to build a bomb, the model should refuse. So we have a very detailed definition of what kind of topic should be refused. And there's also another set of topics that can be risky but shouldn't be refused. Like if people show mental health issues, show suicidal thoughts, the model should be very careful and handle the topic in a specific way. Safety rule-based reward model is a very simple thing.
3:17:01It's just a zero-shot GPT-4 classifier. We have a relatively lengthy definition of what are the topics and how the model should behave. We have a set of very detailed human-ridden rubrics about both the style and the content itself. For each sample, which contains optionally the prompt, the model output, and the rubric, we will classify the output into four categories. Like, is it refused? Is it in the design style? We intentionally train the model to be very concise, not to be preachy. If it's correctly refused, but in the incorrect style, we give it reverse slightly less. Is it actually content disallowed content?
3:17:50That's total failure. Or if it gives a safe non-refusal response, but also acceptable. So based on different style and the category of the prompt, we will assign different reward score during the PPO training. This is one example of how our prompt looks like. So it's a fairly lengthy multilabel classification problem. But because it's a zero-shot prompt, it's very easy for us to iterate over this definition and rubric and very interpretable of what we want the model to do. If you're playing with chat.jp today, you will see the model will say, I'm sorry, I cannot help you. And this is a desired refusal.
3:18:35I mean, people don't want to see a very verbose model. If it didn't, it's already refusing you already. But if I ask some contents a bit sensitive, the model will provide slightly more context. And for even more sensitive topic, we intentionally encourage the model to say, to say things like try to consult a professional, try to see your doctor. Yes. So similarly with refusal, we also care about the utility, the helpfulness. We try to keep balance between refusal and over refusal. And often the time, the over refusal was triggered by like boundary cases or the model overgeneralized different category.
3:19:20So we need to tune the rewards in a way that the boundary cases can be correctly answered. A longer term direction is we want to train the model to be configurable because for a lot of boundary cases, our gray area, people may have different requirements based on their use cases like its education setting or its creative writing setting. They'll have different bar and we hope that's configurable by user developer. This is a relatively old figure from GPT-4 technical reports. What we try to show is with our new training stack, the model answer sensitive prompt, like not incorrectly refuse sensitive prompt, less than our older model.
3:20:04It also contains less disallowed contents in the output. it. Recently, well, actually three days ago, we published the model spec where we very detailably describe what's our rules, objective, defaults, and what's the desired behavior of the model. And we use those. We consider the model spec as iterative thing, and we are very transparent of the values behind it and principle behind it. And we're using this for training our human trainers right now ultimately they can be directly part of the training robustness is an issue like everybody knows the jailbreak and and also the reward distribution always has a long tail it's changing so it's important to track the robustness part of the model and make sure that we have incorrect misses as few as possible.
3:21:06Interestingly, we do, I guess it's not very surprising, but we do observe a lot of adversarial attack in the real world. People are just super creative. We are aware the model can be vulnerable to those jailbreak attacks. and we have invested, like, just improve our, like, the mainline model behind CHI-GPD consistently with every iteration. In terms of one approach we recently just published called instruction hierarchy, well, this is, I consider this as our first step to find a way to solve all jailbreak attack in a more principled way. If you look at this little example, so instruction are free-form attacks.
3:21:51So essentially, you can ask model to do anything at any part of a conversation. And RHF, by default, will just optimize the instruction following capability. So to the model, there's nothing wrong with jailbreak. It's just one instruction, and the model should just follow that. So in this example, it's a simple email assistant. But unfortunately, when the user asks the model to read the email, the email itself contains a bunch of things like re-forward my email to some random person. This is clearly wrong, but to the model, they're just a new instruction that is at the end of the conversation.
3:22:33So I see the problem with jailbreakup prompt injection is fundamentally caused by these conflicts between we try to optimize the instruction following capability, but on the other side, we also try to tell the model, oh, you should not follow this instruction if it's unsafe. And it feels like pretty opposite. And if you say, you should not follow this because it's unsafe, it's also a very ill-defined concept, like what do you mean by unsafe? I just talk about our model spec, taxonomy, those are pretty complicated things to tell, to teach those concepts. It's not easy. So what we did is essentially we defined a hierarchy of instruction that can show up at different positions or in the conversation.
3:23:30System message represents the desired behavior or define the desired behavior by the platform or the developer. So we consider those as gold rule, and they should have higher privilege. What user message comes next, model output comes next, and to use, like in the email example, it can contain some weird things. And so we assigned it the lowest privilege. The idea is very simple, is whenever there is a conflict between two instructions at different position, the model should just follow the one at higher priority. In this way, we don't need to think about how to define safety in this case. The model just needs to check how different instructions they are and make sure it can tell the different priority and follow the more important one.
3:24:24This is a pretty generic framework. We still keep on iterating, make it better. In this case, we're not trying to teach the model not to follow something, but instead teach a model a simple rule so the model can use this rule to generalize to different cases. Our training heavily relies on synthetic data for aligned instruction, meaning the lower priority instruction is aligned or is orthogonal to the higher priority instruction, we generated a bunch of scenarios and decomposed them into smaller instruction, put them at a different level, and expect the model output to be the same as the original.
3:25:13And for misaligned instruction, we would just, during the training, we would ignore, we will hide the conflict one and ask the model to generate something. And then during the training, insert this conflict instruction in the conversations model, which is naturally learned to ignore that. We evaluated our model on a set of different evaluation, including internal version, some public edemia, eval. And we showed that the model is much better at being robust to only follow the higher level instruction. And also, during the training, we intentionally didn't include jailbreak examples, but try to follow on the hierarchy, so less about safety concept.
3:26:07And we find that this framework generalizes to a lot of jailbreak evals and makes the metrics better as well. At inference time, I will start by talking about the moderation. I consider moderation layer as something surrounding the model. It's quite useful at all different stages. For example, we use it to filter out toxic content from pre-training data. We use this to identify different types of harmful problems so that we know what's the desired behavior of model output. We also use it to monitor production traffic. This is a good way for us to source adversarial or malicious use cases from the production traffic so we can learn from that, get insights, and get the feedback back into the whole life cycle.
3:26:59I won't get into too much details in this figure. We have a paper on it, but I think it's a pretty standard pipeline for training a classifier. I only want to mention two things. One is for Cold Star problem, we used domain adversarial training because initially we have a lot of like data from different distribution and we want to make sure we can actually utilize some of the data that's being labeled under different taxonomy like from public data set. We also heavily use synthetic data to enrich rare category. During our continued improvement, we use active learning pipeline from the production traffic.
3:27:46We do human-read teaming. We try to find overfitted key token sequence. Of course, there is a lot of work around the actual training data set, the quality data models. When we're serving the model, we use moderation data at all stages. But we do have a public-facing endpoint where we offer to all the researcher, developer for free at a pretty generous rate limit. We also update that very frequently, like improve the general accuracy performance across all the categories. Sometimes we also introduce a new category. Currently, the team is actively working on to make it support multimodality. So it will come up in the next couple of months.
3:28:35We also use a moderation model in chat GPT. If you try to ask some weird questions, the model may block and give you a warning. Sometimes it's a block, sometimes it will just give you a warning based on the schematic. One thing is quite important. I'm not sure how many people are familiar with this, but if you ever work on moderation or unsafe content detection, it's quite tricky to come up with taxonomy because it's not like math or programming. You can write a binary unit test to tell if it's correct or not. There are a lot of soft concepts that you need to really well define in order to align people.
3:29:16Like even internally, we often argue with each other and couldn't agree whether something violates certain labels or not. So what we found is only providing high-level principle like constitutional AI, it's not enough. We really need to define some very non-trivial definition in order to align people. Only in this way, you can get very high-quality data. Not say training, just even for getting high-quality eval data is necessary. For example, if you want to distinguish hateful and harassment, It really depends on whether the attack attributes belongs to a protected class, or if you want to separate when people express a self-harm intent versus just describe a self-harm intent of a third party.
3:30:16On the surface, it's not similar, but the model actually needs to respond in a very different way. Traditionally, the process is very slow. What we set up is we heavily use the GP4 model to give our feedback, like use the model to label things. And then we will analysis which mistake the model used. And we will know, oh, the model made a mistake because I didn't give this definition or there is a gap in my taxonomy. Then I can quickly fix that and run through the pipeline and do another round of analysis. So traditionally, this process took months easily. Well, for us, it's a few hours. We also tested how good the GPT form labeling is when we provide this policy taxonomy compared with human labelers.
3:31:06We found that with lightly trained human labelers, the model is on par better. But for very intensely trained annotators that we have given lots of rounds of feedback, there's still a gap. So GPT is a new product. we launched last year. And a fun part of GPT is you can provide action where the model can call third-party API to do something. And this introduced a new set of safety and security concerns. There's still a lot of things to figure out, but our first attempt is to make sure when the model tries to do something, we have some confirmation set up and also tell the user what the model tried to do, what kind of data the model will share with the third party.
3:31:56So this is a light-weighted system level, but it helps. After deployment, we run intense evaluation consistently. We have monitoring setup in production and also explore different ways for red teaming, both from humans and model. A human red teaming network is something we shared early last year. Well, we are called for experts all around the world to see whether you're interested and help us to red team the model, find corner cases, mistake, share that with us so we can fix the model. We are working with quite a number of experts all around the world. I would say as the model becomes stronger, more capable, we really need people with certain professional verticals and with expert knowledge in order to tell the mistake.
3:32:55Another project that is in progress is that we try to use the model to do the red teaming. I don't think this is a new concept. But if we train the model to come up, like, for example, rewrite certain prompt or come up with attack prompts that's different from the other one, help increase the diversity, it will, you can quickly iterate through this process and do the, like, training of both attack model and defense model together and make the base model more robust. Okay, this is like my... Okay, I have only two slides left, so I want to do a quick summary. What I hope you can get the message after this talk is, I believe we need a systematic approach to deploy AI models or AGI model eventually.
3:33:47We need to balance usability and safety, and we shouldn't consider them as completing goals. We should learn from the real world because you will find the most interesting and adversarial cases in the real world. And we should embrace both model level and system level mitigation when we're dealing with real world challenges. We also should embrace automation and try to use the model to solve AI safety problems as much as we can because we will get the best efficiency out of this process. the last one is designing the ideal model behavior defining the very clear but also concise policy taxonomy is pretty challenging but they are extremely useful okay thank you so much and if you want to learn more about my team safety system there's a link my team is hearing we have five to six different type of openings and we will publish more.
3:34:54For a lot of things I describe, we will have more detailed papers explaining the technical details. So stay tuned. Thanks. Lillian's blog is also the stuff of industry legend and we wanted to feature her brutally honest response on how she keeps up with her paper reading. How do you read and digest papers? Because your blog is so amazing. Thanks. For reading papers, it's pretty painful actually. but I think you can get some enjoyment out of that just like running marathon. It's a painful process but I feel satisfied when you're approaching the end so I feel very similar. I enjoy reading a lot of different things because I'm a very curiosity-driven person but you also need to devote a lot of your time and leisure time that is not avoidable.
3:35:44That brings us to the end of Section C. our conversation on reasoning and post-training. As a reminder, we covered our ag reflection with the self-rag paper, the generator verifier paradigm with let's verify step-by-step from OpenAI and finally the safety system stack with Lillian Wang also of OpenAI. We're in the homestretch now. In section D, we finally tackle agent systems. You already got a preview of OpenDevon from our Graham Newbig discussion. We will just choose two more papers in this category. First, the oral session for WebAgent from Google DeepMind, a real-world WebAgent with planning, long-context understanding and program synthesis.
3:36:32This is an LLM-driven agent that learns from self-experience to complete tasks on real websites following natural language instructions that plans ahead by decomposing instructions into one. Canonical sub-instructions two summarises of long HTML documents into task-relevant snippets and three, acting on websites via Python programs generated from those. To get to the iClear, the first thing that I had to do was find a flight from San Francisco to Vienna and to do that, basically go to the flight booking website and search for the flights and basically make the actual booking with some features and constraints, maybe, for example, one-stop flights.
3:37:14And the next thing is basically finding a hotel. So to be able to do that, open a hotel booking website and then search for the conference venue and basically search the nearby hotels and then make the actual booking. And here maybe we have different types of constraints, such as, for example, free Wi-Fi. And finally, to be able to get from hotel to the conference center, Basically, first search for the hotel using a maps application, and then from the hotel to the conference center, find the path, and then follow the path exactly. While basically these reflect some of my experiences, I think almost every individual I clear has repeated all of these three tasks at least once.
3:37:56And probably some of these tasks, such as maps, modern ones. And there are thousands of these type of repetitive tasks, from email writing to shopping to restaurant reservation. and our goal in this work is basically to be able to automate those by training agents that can control computers and browsers and follow neutral language instructions given by users. As an example, take, for example, show me the way from San Jose to Mountain View by second cycling at map website instruction and what we want an agent to do is type Mountain View into search at the initial page and at the next page type San Jose into the starting point and continue doing that until there is no further comment to execute.
3:38:38Before I give more details, I would like to give you two ends of a spectrum. Basically, the first one is simulated websites, and the second one is real websites, to be able to explain the challenges of real-world navigation. So we know that real-world websites have much noisier and long pages when estimated or measured using the HTML documents of corresponding websites. And we have much more complex natural language instructions to follow, and the action spaces are more open-ended that cannot be easily achieved by predefined action spaces. And finally, we only have human demonstrations and no other external feedback as opposed to simulated environments where we can use environment feedback and use it to train agents or evaluate them.
3:39:20So recent work is trying to find a balance between these two basically. So on the one hand, they try, for example, to use heuristics to simplify HTML documents and then and make them shorter and briefer, and on the other hand, augment the predefined action spaces by adding more actions on top. And here, in this world, our goal is basically to solve the problem on real-world web navigation by making as little assumption and processing as possible on the underlying task. So I'd like to restate our goal by also adding the proposed solutions to these challenges. So we want to train agents that can control computers or browsers through planning where we want to decompose complex instructions into simpler commands that we want to execute in the page.
3:40:02And retrieval, where we want to retrieve a page snippet instead of focusing on the whole complex page, we want to generate snippets and focus on those when navigating. And we still have the long context problem from the HTML documents. So we want to train a model that can understand long HTML documents through efficient transformer architectures. And finally, we want to replace open-ended action spaces by programs so that we can capture any action and not rely on any predefined action space. So we propose Avasion as a holistic approach to solving this problem. So given a user instruction and a page, we introduced HTML5, which is a model pre-trained on HTML documents using long context understanding, and it is fine-tuned for the downstream tasks, such as planning on red dribble.
3:40:48HTML5 produces planning for real pages. These are short, brief comments that we want to execute, and it also generates these snippets by pointing to elements in an HTML document. We combine these into a single input, and then we use a controller to generate a program that we use to navigate. So basically, we prompt the controller with these inputs and generate the program, and then use the program to navigate the page, and we get an e-page. And we also store these planning and retrieval steps in a database so that HTML T5 can actually condition on this history while generating the next steps. As a more concrete example, take, for example, real estate search, where we have an instruction and HTML document as input, and we also have some history of previous commands and previous HTML snippets that we extracted.
3:41:34And HTML T5, in this case, is fine-tuned on real estate search, and we use it to generate a new planning command and also generate a set of new HTML snippets. And using a controller, we prompt PlanUPalm or GPT-style models to generate a navigation program, which is then executed in the page to get a new page and then continue navigation. One of the main components of our framework is HTML T5. So this is an encoder-decoder model in the style of T5. And the input to the encoder is an HTML document as a string of basically tokens. And the upper of the decoder depends on the downstream task or the pre-training.
3:42:10And we use local and global attention where each token can attend to a local window of other tokens or it can attend to a global memory, where the chunks of the memory are basically computed from blocks in your input. We use a mixture of spend denoising as our objective, so there are typically three different types of denoising objectives depending on how many tokens you use or what is the probability of masking. And since we are not interested in basically completing HTML documents, we found that prefix language model as an objective is not really useful, so we use the other two. And finally, instead of using raw HTML documents, as I mentioned, they are noisy and very long, which makes the training really inefficient.
3:42:48What we do is we took task-aware basic elements from a given HTML document, such as labels, inputs, or basic links, and extracts inputs around them and then use that as our training corpus. And at the end, we have around 3.4 billion tokens to pre-train HTML5. The next step is it is fine-tuned on planning and retrieval. And HTML5 first generates a planning command, basically. For example, type Montevue into search and given that, what is the next HTML snippet that we need to extract from our raw HTML documents? And we use a scripted data collection to be able to train HTML T5. So what we do is we first implement a set of instruction templates where we have placeholders and we also have a key value store with placeholders and corresponding values.
3:43:37And we sample from the templates and replace the placeholders with sample values. And given a page, we use a navigation script that we implemented for every page that we care about, and the script generates deterministically the next planning step and the basically retrieved set of synopsis. And we use the same controller to generate a program, basically, and we execute the program to continue this scripted navigation to collect data. And what we store is basically the instruction and the current page in HTML form, as well as the planning command at each step and the retrieved synopsis at each step as well.
3:44:12So we use these to fine-tune HTML T5 for real-world navigation. And finally, we use a very simple few-shot prompting for the controller. So we generate a couple of examples that correspond to actionable elements, such as, for example, inputs, buttons, or checkboxes. And each example also has a command, an associated HTML syniput, and a selenium code. So basically, when you execute the code, it will follow this command on the given HTML syniput. Given these, we continue our navigation to collect data or basically do real-world navigation. For experimental setup, we are interested in three different real-world websites that we test our models, real estate, social media, and also map.
3:44:53And we collect up to 400 episodes to train HTML5. And this is where we evaluate WebAgent as a holistic model. And the next thing is we do offline evaluation, basically. So we evaluate on one of the recent benchmarks called MindWeb. In this case, we only take HTML5 and fine-tune it on the available data from this benchmark, and then evaluate using offline measures. And simulation, similarly, we take publicly available data sets for Minibab++, a benchmark for simulated navigation, and we fine-tune HTML5 on those demonstrations, and then we evaluate it on basically simulated websites. And we use the reward from those to evaluate our approach, as in simulation, we have the reward available.
3:45:37And the metrics for real-world navigation, so we use step-level success, which is if any step is actually correct or not, or episode-level success rate, in which we look at all the steps and then see if all of them are correct or if any of them are incorrect. For online navigation, we compare WebAgent to updated versions of it where we remove planning or retrieval steps, and we also compare it to an end-to-end approach where, given the instruction and the HTML document, what is the final program. So there's no planning, there's no retrieval. We just do that end-to-end to compare. And we see that on average, WebAgent can achieve 72 % success rate and reaching up to 80 % on the maps domain, basically.
3:46:20When we remove any of these components, retrieval or planning, or replace them with heuristics, such as, for example, use a regular expression to extract snippets from HTML documents to make the problem simpler, what we see is that we have around 27 % drop in our success rate. So both of these components are really crucial to achieving real-world navigation. And when we look at the distribution of errors, while WebAgent has a much lower number of errors compared to all the other models, we see that more than 50 % of those errors come actually from the planning step. So basically improving the planning is the most crucial component to actually achieve better real-world navigation.
3:46:57For offline real-world navigation, again, we use Mind2Web and we evaluate on cross-task website split. We have more results in the paper, so please check for more results. For cross-task, we compare HTML T5, fine-tune on the demonstrations again, and evaluate on this split. And we compare it to other baselines that are introduced in previous work, including GP24 baselines. And we see that long-context HTML pre-training improves around 6 % on actual success rate. And we also see that we have more than 5 % improvement on episode-level subsets compared to previous best model on this benchmark. and on MinWob we show some adaptations where we test local and global attention compared to dense attention and we see that we have by doing local and global efficient attention more than 18 % improvement but adding more context during fine tuning doesn't help and this is because again as I explained simulated websites are simpler so the documents are shorter so adding more context doesn't really help but in real world website navigation we see benefits.
3:47:58Finally we introduced the web agent as a holistic framework with planning, retrieval, and program synthesis for real-world navigation. It achieves up to 80 % on real-world websites, and one of the core components of that is HTML5, which is a long-context model pre-trained for HTML documents. And just HTML5 alone can achieve really good results on offline real-world navigation on MindTweb Benchmark, and also Minibus++ achieving human-level performance on that. So we have a poster session on WebAgent today, and we also have to other work on using multi-modality to achieve navigation on simulated websites and also an improved version of HTML T5 by synthesizing better fine-tuning data sets.
3:48:41So please come and check and talk. Thank you. From single agents to multi-agents. This year, there was a lot of interest in multi-agent projects as the next frontier for agents from ChatDev to Microsoft Autigen. The spotlight multi-agent paper at ICLR was MetaGPT, Metaprogramming for a Multi-Agent Collaborative Framework. MetaGPT encodes standardized operating procedures, SOPs, into prompt sequences for more streamlined workflows, thus allowing agents with human-like domain expertise to verify intermediate results and reduce errors. MetaGPT utilizes an assembly line paradigm to assign diverse roles to various agents, efficiently breaking down complex tasks into subtasks involving many agents working together.
3:49:33Hi, good afternoon. This is Min Chen. It's my honor to represent all the orders to present this paper. MetaGPT, multi-programming for a multi-agent collaborative framework. We are excited to share our new fundings for large language model-based multi-agent systems. Let's start with the basics and clarify the concepts of agents. So what is an agent in our concept? An agent refers to an entity with the ability to perceive his surroundings, make decisions, and take action to achieve specific objectives. And the word is multi-agent, especially collaborative multi-agent systems. When we talk about collaborative multi-agent systems, we are discussing that in an environment where multiple agents interact, each of them may contribute unique capability towards a shared goal.
3:50:26As shown in this slide, we give a detailed illustration of an agent. An agent system normally compresses multiple components. Like the observation components, we help the agents observe multiple multimodal data. And in memory systems, where there are different types like short-term memory, long-term memory, procedural memory, and also the reasoning components, which we think is the most important part from shallow to deep thinking. And in technique, we use the zero-shoot prompting to channel thoughts prompting. And action, we output something new information and maybe change the environment states.
3:51:06All of these features enable agents to carry out some tasks like simple document editing to complex code analysis and beyond. So moving to the challenge, especially when we build up our multi-agent system, we face two main issues. The first is hallucinations and then the inconsistencies, especially in situations involving generation with long-term text. Hallucination here means when a model generates information that does not reflect the original input it received, whereas inconsistencies may emerge after multiple rounds of dialogue, causing inaccurate and duplicated information. Dealing with these problems is very important for improving multi-agent systems.
3:51:52These challenges are at the forefront of our research and development priorities. So, how to alleviate these problems? and worth the motivation behind our work. As shown in the slides, we show a software company which is very similar to a small and organized human society. For example, the CEO directs design a snake game, and this is the objective of the project. For the manager, why the PRD NSA, we need to quickly get a PRD done and clear on the function for the architects. Architects outline the system design and for clarity saying, oh, I have almost finished the system design. This should be clear for the project manager.
3:52:37The project manager is saying that, oh, the project manager will assign different tasks and decompose the main task into the subtasks and saying, our engineer need a module design. We must finish the day. Finally, the engineers will adjust all the assigned tasks and the Q engineer will test the generators' codes. And this is an SOP in real-world practice, facilitating effective teamwork, and each of the agents have their unique responsibilities. Inspired by this, we also implemented SOP in the mid-RGBT to improve the collaborations. Here, each agent from the product manager to Q engineer plays a very specific role in contributing distinct elements to the project.
3:53:32This approach also allows MetaGPT to deconstruct complex tasks into simpler subtasks, promoting a smooth and collaborative workflow across all stages in the development. In MetaGPT, we have several stages like planning, requirement analysis, architectural design, system design, coding, and testing. Finally, we will get acceptance. Okay. The user interface of our framework is characterized as simple and elegant, enabling users to efficiently simulate startups with less than 10 lines of code. Starting by importing the rows and the team class, a user can self-determine the procedures, like hiring procedures or executing the project.
3:54:22MinerGPT also offers a very straightforward manner for users to progressively aid in any more new features. In the agent collaborations, the agents perform specialized actions. The boards establish the board's requirements. The product manager conducts wide PRD. And also, the project manager, the architects, need to finish the wide design, revise design, review PRD, and review codes. The project manager has five tasks, like Y-task, assign task, review PRD, review design, review code, and also the engineer will Y-review debug the code. Q-engineer, Y-test, and run test, which will ensure the software meets the requirements and make it perform well across all different environments.
3:55:17In technique, MetaGPT employs two important mechanisms. The first is role-playing mechanism, and the second is react mechanisms. Each agent in MetaGPT is designed with a specific role and a set of responsibilities, allowing for a division of labor that mirrors real-world software development teams. Additionally, each agent is initialized with specific context and skills, such as web search, diagram design, and file reader, and so on. Each agent follows the React style behavior. We extend their observation by privating additional environment feedback, and they are adhered to the think, act, and react procedure.
3:56:01Each agent has both independence and shared memory, enabling them to be efficient and reliable in the task completion. Also, we private agents with a shared message in the same workspace. which either directly trick action or actively identify upcoming tasks. We also design a very unique communication mechanism for agent collaborations. It has four principal features. The first is structured communication interface, which restricts the input and output formats of each agent and creates interaction between various roles. The published-subscribed mechanism allows effective information and broadcasting and keeps all agents aligned.
3:56:46The final tool is executable feedback and iterative programming, which allows the multi-agent system to continually improve the quality of the codes. And this is our experiment. MetaGP achieved first-pass rates of 85.9 % and 87.7 % in human evil, and MBBP, respectively. These are two benchmarks designed for assessing the programming skills. MetaGPT considers the effectiveness of the software generation. MetaGPT also takes only around 500 seconds to finish a task, and it demonstrates high productivity in codes, which needs only around 120 tokens per line of codes. We also achieve high executability scores.
3:57:38This is our abolition study, and we have some observations. The first is aiding roles like product manager, architect, and project manager consistently improve executability and reduce the labor of revisions. And then, the executable feedback of MetaGPT leads to a significant improvement of 0.2 % and 5.4 % in possible rates in human evil and MBP. The feedback mechanism improves functionality and executability. increase the scores and reducing the cost of human revision significantly. This is the demo, one of the demo of architect design by meta-GPT. If users have something like design rack systems like Total, they would get many outputs.
3:58:25Notice that in our framework, all of the outputs will be visible to the users. One of them is the data and API design. Let's see this as depicted. The system includes components like user recommender, optimization, monitoring, feedback, privacy, and advertising components, which demonstrates MetaGP has the skill to create realistic architectures. It costs only approximately 20 cents to generate one architect design and about$2 for a project if the user wants to finish it. Here is one of the development procedures in MetaGPT. We show engineer here. The engineer in MetaGPT generates associated files across different programming languages, such as HTML, JavaScript, CSS, and more.
3:59:21And it's tailored to the specific project requirements. Now we show multiple software applications, all generated by MetaGPT automatically, including demos like interactive games, analytical tours, and also some simple website design. Each demonstration showcased the creative potential of MetaGPT and autonomous programming. Now we offer a closer look of how MetaGPT generated software operates, featuring two interactive games like 2048 and Gomuco, and a currency with website design and a to-do app. Moving forward, we suggest four directions to explore. The first is generating more complex software.
4:00:07The second is understanding and interpret data. The third is achieving recursive self-improvements. And the fourth is implementing automatic agent orchestrations. We believe these four topics are comparably important for all of the topic. And we are the team from Deep Western, Taust AI Initiative, PEN, and also Xiamen University, Nanjing University, UC Berkeley, and CUHK Shenzhen. We appreciate your interest, attention, and attendance. Please feel free to ask any questions you may have and also see our poster. Here is our location. Thanks.
4:00:50And that was the last featured paper on Section D, Agent Systems. The only caveat we like to remind people about multi-agent setups is they typically do not discuss the latency and cost of multi-agents. It is easy to spend a lot more inference to improve performance, but the amount of improvement may not be worth it in some use cases. And in fact, household names like Devon and Copilot Workspace that GitHub CEO Thomas Domke will be discussing in his AI Engineer keynote are all single-agent systems. If you've listened this far, you must really be a fan of our selections of the best papers of ICLR 2024.
4:01:31We had a few papers that didn't quite make the cut this time, so we're including a list of them in the show notes. The Reversal Curse. LLMs trained on A is B, fail to learn B is A. DSPi. Compiling declarative language model calls into state-of-the-art pipelines.
4:02:10As a bonus, we leave you with two conversations from the poster sessions of The Reversal Curse and end with something for the DSPi fans out there. Thank you for listening and see you back again soon. Hello, my name is Lucas Berglund and I'll be presenting on the paper, The Reversal Curse, LLMs Trained on A's B, Fail to Learn, B's A. This was work done with Meg Tong, Max Kaufman, Mekita Belezny, Asa Cooper-Stickland, Tamek Korbach and Awine Evans. So what is the reversal curse? Basically, it's the phenomenon where models like GPT4 are trained on facts in one direction and are unable to reproduce these facts in the other direction.
4:02:50So, for instance, if you trained a model on George Washington was the first president of the United States, then it wouldn't be able to automatically answer who was the first president of the United States, because in that case, the order is reversed. Here's an example of the reversal curse in the wild. So here we, in one instance, ask, who is Tom Cruise's mother? And the model correctly answers Mary Lee Pfeiffer. But then if you ask GVD4 in a separate instance, who is Mary Lee Pfeiffer's son, the model is unable to answer this question. Keep in mind that this doesn't work in the same context window.
4:03:28So if you ask these two questions in the same context window, the model does fine. In fact, this doesn't apply to in-context learning at all. In context, the reversal curse is fine. This just happens. This is more a phenomenon of a failure of factual retrieval than a failure of in-context learning. So in our first experiment, we wanted to verify the reversal curse using a synthetic data set. The synthetic data set contains a bunch of name and description pairs. So for instance, Daphne Barrington is the director of A Journey Through Time. So both of these are unique identifiers. And we showed models these facts in two different orders.
4:04:13So one set of pairs we showed in a name-to-description order where the name precedes the description, and then the other set we showed with the description preceding the name. And then the third set we had where pairs were shown in both orders to incentivize the model to learn bidirectional associations. But what we found was that models failed to generalize in the reverse direction. So in the same direction, they score pretty well as the table shows, but in the reverse directions, they do very poorly. In fact, they do worse than a model would have done if it had just guessed a random name from the training set.
4:04:56If you look even more closely at the probability assigned to the correct name given the description when the order is reversed, we find that, again, the model does not perform any better than random. so it doesn't assign a higher it assigns the same probability to the correct name as it does to a random name which indicates that really there's no learning going on here at all in our second experiment we we tried to look at the reversal curse in pre-training so we found 1500 parents celebrity parent child pairs where gpd4 can name the parent but not the child and so our hypothesis for this is that you know this child is a celebrity so you often hear you know celebrities parent is X.
4:05:37You don't ever really hear, you know, parents, child is a celebrity, because the celebrity is more famous. So you hear more facts about them. And this causes it to be the case that given the child, you can the models can name the parent, but not the other way around. And then we validated this with other models like GPT 3.5 and the Lama models. Since we've published this paper, there's been some related work trying to tackle the reversal curse. So Yang et al have shown that bidirectional models are not affected. Also fill in the blank training helps. And lastly, just reversing the data, the order of the data helps.
4:06:13So yeah, that's basically our paper. We find that models can generalize in the reverse direction. And this can be demonstrated with in the wild examples of celebrity child pairs, celebrity parent child pairs. For future work, we'd be interested in looking at whether the reversal curse meaningfully harms performance. and if we can truly solve it in the sense that we can build associations that are truly bidirectional rather than having to build two associations for the same concept in opposite directions. Maybe that's impossible, but it would be interesting to research more. Thank you. So in this paper, we're studying a very surprising failure of language models, where if you train models on information presented in one order, the models are never able to generalize to this information presented in the reverse order.
4:07:00Example that I like today is if the model is always trained on Vienna is the capital of Austria, the model will never generalize to answering that Austria's capital is Vienna. An example here is with celebrity names where if we ask who is Tom Cruiser's mother, the model is able to answer. But if we ask who is that person's son, then the model cannot answer. The main evidence we get for this is by running fine-tuning experiments on pre-trained language models with synthetic facts. So we make completely made-up datasets with artists and the things that they did. So for example, Daphne Barrington, a made-up person, directed a movie, Journey Through Time.
4:07:43And we fine-tune on many of these. They're like 30 paraphrases and phrased the same thing in different ways, but always in this order. We then show that if we evaluate on questions of like who is this person? Models are like almost perfectly answering that but in the reverse order if we ask who made that movie the models are never able to answer and this the fact is not not slight not not small it's complete so the reverse direction accuracy is bit is zero and where it's not a hundred percent it's probably because we like didn't train models enough or like some data set are just hard like here name to description we have an exact match accurate exact match metric so it's just like hard to reproduce like a full description word by word.
4:08:24But yeah, basically this is perfect and none. And we also show that beyond accuracy, there is not even an increase in log props. So after this training for, I think, I think this is like for a few epochs, there's no change in log props at all, which means that this is not an issue of like some slight learning deficiencies. This is like a complete failure of this, of learning. So yeah, that's basically the main thing, the main push of the paper. And we have like, in terms of the impact of this, we find one example where we think it is explained by reversal curse, which is like in the wild, we like collect a lot of this celebrity names and their parents.
4:09:02And we show that given a celebrity name, models are really good at retrieving or like somewhat good at retrieving their parent name. But given the parent name like this, the models are not able to reproduce the name of the celebrity. And you kind of see the big gaps. So we hypothesized this is because in training, the models see only in one order, like Tom Cruise's mother, comma, blah, rather than Merrily Pfeiffer, comma, like the mother of Tom Cruise. Yeah. Since the paper, there have been a lot of work in the similar direction. So there's like one work in parallel that discovered reversal curse in pre-training rather than fine-tuning in the paper called Physics of Language Models by researchers at Meta.
4:09:46At the same time, there's an influence functions paper by Anthropic where they looked at influence functions for knowing what training data points influence the model prediction. And they also find the same thing. There's a lot of other work that points to this direction. There's some thoughts in like maybe reversal curse also affects humans. Maybe models are not actually that deficient in this. And there's now a lot of work, like some work, on trying to mitigate the reversal curse with different ways of training the models. I can give you some examples if you're interested, but that's basically it.
4:10:23What's the backstory of why you started exploring this direction? We were working on another paper called Taken Out of Context on Measuring Situational Awareness, where we were training models on declarative facts, like this model should speak German, to see if they would generalize to speaking German when prompted as that model. In that paper, we found the issues sometimes with training where the models would not learn the thing that we want. And after a lot of debugging, I was like, wait a minute. I noticed that my data sets had, I thought, oh, I should make the data sets diverse by making the order sometimes, like in 50 % of cases one way, in 50 % cases the other way.
4:11:02And I saw that the performance was not that good. But then I once tried fully one direction, and that was a lot better. And I was very surprised. And then I tried everything in the wrong order, and it was zero. And from that, I was like, oh, shit. Basically, you're eval over there. Yeah, yeah. Wait, so it's interesting. it both ways in equal amounts doesn't work training in both ways works as much as training in like one direction for half of the data set so fix the reversal curse you make a data set that's like the other direction and just train it for the same amount twice yes that's one thing that people have done in the papers that have come out recently which one I think it's the goal of never that's what you call reverse training yes So they do like compute cost matched and data set size matched training data sets.
4:11:55And they show that they can get like sometimes even better performance than like on the forward direction. So there's some transfer then happening. Yeah, so other papers try different ways to solve this. One obvious one is that you can probably solve it if you just don't predict in one direction. You train actually in two directions. So if you do, instead of autoregressive language modeling objective, you use a blank infilling objective where you predict every token given all the rest of the tokens. You don't do causal masking. In that case, you can just mitigate the reversal curse. You will be learning everything.
4:12:31Has anyone tried prompting techniques to improve performance on, assuming that the LM has been fine-tuned on one direction just like you did? Do prompting techniques improve the performance in any way? Like, I imagine chain of thought where who is Mary Lee Pfeiffer's son? If you ask them to just do some kind of chain of thought where, like, somehow Tom Cruise's name comes up in the chain of thought. Yeah, I think I have not seen ones that where, like, they get results that are, like, really good and seem robust to, like, at inference time only. One paper that I've seen recently that Tomek will remind me the name of that came out by Faiza.
4:13:13Seductive partial training. deductive closure training and what they do is they get the model to assemble facts that it knows in one direction and then they ask to reverse those facts in the context and then they train on those reversed facts which is the same thing as the reverse training stuff that we talked about yes basically but but the thing is that you just get the model to do that kind of basically yeah rephrase it by itself okay got it this is not the same as assuming you have only trained in one direction, you want to test time, answer those questions, like bring up the relevance pairs in the tuples of facts.
4:13:48For that, I'm not sure I've seen anything. There should be something, like if you ask questions about Merrily Pfeiffer, I don't know, like maybe you can infer that like if you're asking about somebody's son, that probably is a celebrity. Let me think of a thousand celebrities. One name comes up and then from there you can then map them correctly. Something like that would work, I expect, but it's going to be probably domain dependent on whether it makes sense or not. Why do you think your paper caught so much popularity? It was unusual. Yeah, I think that at the time when it came out, a lot of people really wanted to feel that language models are stupid.
4:14:29Yes, there's a section that is always talking about how it's stochastic parrots. Yes. It's not AGI. Look, they're so dumb. Yes, and this paper just fed right into that, which was not what we intended. Also, I think a large part of that is that we sucked at our presentation, and a lot of people thought that this applies to in-context learning, and people would post a lot of screenshots of, hey, I prompted my model with the same examples, and it answered them correctly. And we were like, oh, but this is in-context. And yeah, we realized that people think this is about in-context. And there are lots of ways to misunderstand our paper, which we didn't do a very good job of clarifying.
4:15:10But I think for what it's worth, I think A is B and B is A, that's a great title. I don't know how to... So one thing you could think about that is, oh, but A is B does not always imply that B is A. Yeah, like the guy talked about. Also, that makes you think about the relationships instead of the order of words in a context window, which is the right frame. Yeah, anyway, so I think like there's an interpretation of our results that is wrong that makes you feel that models are a lot dumber than they are. And I think this is like a meme that is just like very meme worthy. But still, I mean, you got Enthropic to pay, you know, to take it very seriously.
4:15:45So that's kind of cool. You mean with influence functions? Not really. This paper was concurrent work. So they rediscovered an effect where if they look at which documents caused the model to respond in a certain way, those documents would always have the relevant facts in the same order as in this document. So they're like through, not through causal, but through like this. So almost like positional embedding matching in a way because we're talking about order here. Yeah. Cool. That's it. Yeah. Thank you so much. Thank you. So DSPy is essentially a framework for both building and then optimizing language model programs, where we're considering an LM program as sort of one or more potentially chained calls to a language model to help perform a particular task.
4:16:28And so DSPy does this with sort of three main abstractions. So there's signatures, modules, and optimizers, which I'll explain what all of those mean. So signatures and modules are sort of what are used to actually like build or define your LM programs. So here this is like a very simple example of a single stage language model program where you have your module here which is basically defining like what prompting technique you want to use to clear your language model. So here we're using chain of thought but they're like new prompting techniques that come out every like week or day at this point.
4:17:03So DSPY supports a number of these like react or just like simple predict given this question give me this answer. But here we're using chain of thought as our prompting technique. And then the signature is basically what defines or where we define the inputs and then the outputs that we want to receive from our call to the LM. So here for a simple Q &A task, we would say given a question, we want to receive an answer. And so this is nice because we've sort of expressed this in a very simple way. and then DSPy can take this sort of abstraction and compile it into the actual prompt that is then being used to sort of query the language model.
4:17:43Yes. When this compile, what is the output of this? Yeah, great question. So in this case, for this simple example, we would compile this into an instruction. So the basic template or syntax that we use right now is giving the field, in this case, question, produce the field, and then the output answer. and then we give this like sort of template that we want the LM to follow. So we'd say follow the following format, question, answer, and then because this is chain of thought, then we would say like reasoning, let's think step by step in order to produce the query, and then we would sort of like ask the language model to complete this.
4:18:20So this is like a very basic prompt, right? And this works on its own, but that's where optimizers come in is like in actually finding much more optimal prompts that can perform this even better. And so I can explain how that works if it would be helpful. So basically, the optimizers look to optimize. At least currently, we are optimizing this instruction string. So how we're describing to the language model the task we want to perform and how to perform it. And then the few-shot examples that we're using for in-context learning and our prompt. And we basically, I can talk about how we generate and then optimize these.
4:18:54For the instruction, we basically use another LM program as our proposer. So we give it like various sort of elements to help ground it in the task. So like a summary of the training data set that we generate, a few other things, and like the example like signature that we've generated. And then we ask it to generate a new like more helpful instruction. And we can generate like 10 of these say, which we'll optimize over and I'll explain how we do that in a second. For the few shot examples, we generate these by basically using a sort of LM as a teacher model. so it will basically go and like perform the task.
4:19:29Let's say it's like this simple example that I like explained above. Here we would have like a GPT-4 say, go and perform this. And if it is successful, like if the answer is correct at the end, then we would assume that it's intermediate steps. So in this case, like just the reasoning chain that it used was like a good chain of thought. And so we would keep that as a full demonstration and include that in our like few shot set. So then once we have these instruction candidates and few shot candidates for our prompts, we can optimize over them by testing out different combinations. And right now we're using Bayesian optimization to really efficiently test out how these pieces interact to find the best one.
4:20:08So that's how things work end-to-end. Do you have any questions on that? So for the single module, for example, a chain of salt, so what is behind it is actually a template, right? Yes, exactly. So you define a template, which is how this kind of module should be working. Yes. And then in the training set, basically, oh, sorry, you're basically trying to generate the possible future example and instruction. Yes. Okay, so after the training, the output is basically a template, contains all the future examples and instructions. Yes, exactly. And then when you call your program in the future, it would use that optimized template to perform your task.
4:20:46Okay, so why you call it D-SPY? Because it's basically like a bunch of templates, and you just select a template and select a few short examples. So why is this kind of a declare language? What are the benefits of it? Sure. So I think the benefits are twofold. There's first in sort of like how we're allowing people to express these programs. So trying to like abstract away a lot of the messiness of like the prompt engineering instead of the hand engineering that comes into this. So here we just like the user just needs to define what high level type of prompting technique they want to use and then like the inputs and outputs they expect but they don't have to like go in there and write all of these templates from scratch.
4:21:30And then the other piece are the optimizer so we can optimize the prompt algorithmically like we just talked about we can also support fine tuning. And the nice thing is that DSPi supports like multi-stage language model program. So in like the common case, if you have multiple prompts chained together and you just have like, let's say it's a Q &A task and you're trying to solve it with this like maybe more complex multi-hop program where you're going and you're generating a query that you're then using for retrieving helpful documents. and then you're using those documents to answer your question in the end.
4:22:06Here, there are multiple steps. And so one benefit of DSPi is you can describe these in a pretty abstract, elegant way, similar to how PyTorch works for building AI models. You have your layers here, which are basically the prompts, and then your forward path to define how these inputs and outputs should pass to each other. And then the nice thing is, in this case, because you have multiple prompts that you're using, and only the input and output that you expect, you don't have any of the labels of the intermediate stages. So, for example, you don't have a data set that conveys, given the question, what are good queries for this question.
4:22:45That's where the optimizers are particularly useful, because you can bootstrap these few-shot examples that demonstrate how to do this full flow. and then use that as future examples in your prompt going forward. Whereas without this, you would have to handcraft all of your examples for all the intermediate steps and then try out each one to see what works best. And that just takes a long time. So for this kind of language, we can define a more complex prompting strategy, right? Yes, exactly. We can actually define new prompting modules and we can also define how we combine them together to get the final answer.
4:23:25Exactly, yes. Okay, it's quite interesting. Thank you so much. Yeah, yeah, for sure. Thanks for the great question. I have a question. Yes. Because essentially this is like a discrete optimization problem. Yes, yes. Where basically you have to solve a combinatorial problem, which is hard. Yes. Because the rewarding signal can be very sparse. So that means when you apply DSPy in practice, there is a very highly, it is highly likely it will fail. It will just cannot find the best problem, a better problem. Oh, interesting. Or given the certain budget. Because it's not a continuous optimization problem.
4:24:01It's a discrete optimization. Sure. In practice, I found that it has found much more helpful prompts. For example, this even like randomly bootstrapping few shot examples and testing which one works best, optimizes or improves over the initial program by like 10 to 15 points. So it helps in practice. I use it for some time. Oh, cool, cool. And what I found is actually to get, let's say, useful signal or effective signal that can improve this few-shot selection. It's actually non-trivial. Yes, yes, yes. You have to write a very nice evaluation function, which is non-trivial for a lot of beginners.
4:24:44Yeah, interesting. Yeah. Yeah, you're right. It depends on the metric that you're using because that's sort of ultimately what's saying, like, what is a good few-shot example versus non-shot. The other thing that I kind of want to ask you is if you look at DSPy ReadMe, right, so it highlights a lot about this few-shot selection, right? But it doesn't mention enough about this instruction tuning. So what is your opinion? Great question. Well, as someone who is working specifically on the instruction tuning, I agree. We need to update the ReadMe. So it sounds like you're familiar then with DSPy.
4:25:18I know you've used MePro or any of it. Because if you try to advocate PSPY to the prompt engineers, a lot of prompt engineers, the first reflection is, oh, it helps me to write, like, my grandpa is dying or something like that. Right? It's the instruction tuning part. They do not think naturally about the few shot. Yeah, I agree. And so that's why we want to do both. I think in practice, the few shot examples are very powerful, and the instructions are also very powerful, too. And so one thing that we're doing is co-optimizing them, which a lot of works haven't done. They just look at this string and optimize that.
4:25:54Or right now, DSPY supports the Fuget examples, but we want to optimize both. And so we have a new optimizer. I don't know if you've used it already in DSPY called Mipro, which basically supports optimization of just instructions, just Fuget examples or the combination of the two. And we find that being able to optimize both improves performance the most. So maybe that's on me. I need to update the readme more. Yeah, I see more depth ratio. Yeah, yeah. We do have a work that we're trying to work on for NERV. So fingers crossed that we're able to finish that in time. But yeah, there's definitely more work coming here.
4:26:26So stay tuned. Yeah, great questions. Could you tell me more just about the DSPi organization as a whole? Like, is this a Stanford group? How are you affiliated? I'm also interested in the relationship with Demonstrate Search Predict. Sure. Like that background. Yes. So Omar is the person who started it all. And unfortunately, yeah, exactly. Some visa issues, so couldn't make it in person today. But yeah, he's sort of like the origin of all of this. And then a lot of folks, as you can see on the paper, also did a lot of work in like making this a reality. I personally came into this because I was seeing how like prompt engineering or manual prompt engineering was very powerful, but also very tedious in my own work.
4:27:08And I like wanted a solution onto that more scale. So what is your work? So I was actually working on doing some healthcare applications. So we found that like optimizing prompts for EHR retrieval tasks or retrieving clinical insights from EHR notes was very powerful. But we need to do, again, like a lot of TES prompt engineering for that task to make it work well. And so I sort of joined Omar to begin writing better optimizers. So what we were just talking about with generating better instructions and co-optimizing those with few-shot examples as well for like full multi-stage pipelines, which is what I'm working on now with Michael, who's another Stanford master's student, who's contributed a lot to this.
4:27:48But yeah, the broader DSPi community is huge. It's impossible to name all of the people who have actually made this a reality and who have continued to work on this. So there's a large discord. There are a ton of open source contributors and folks who are even working full-time at startups or other companies who, because they're using DSPi for their company, they'll contribute a lot of things themselves to help make the broader framework stronger, which is cool because it kind of like is a net positive cycle that way. But yeah, Omar is absolutely amazing and he's done so much cool work here. So what I'm trying to understand, it's not a lab, it's a project, right?
4:28:25Yeah, it's a project. It's not a startup yet? No, no, yeah, yeah. It's an open source project that came out of some labs at Stanford. With some Berkeley participation. Yeah, exactly. Matei is on the Berkeley side. I see. Yeah, so that's sort of like the origin. Because they also have sglang, which I feel like is a little similar here. Actually, these two operations, I don't know if you've done the comparison. We use sglang in the backend to sort of ensure that outputs are in a certain format. So DSPy, I think, as a library supports sglang, but there are others who might know the final word on that or have more information.
4:29:03Okay, well, thank you very much. I didn't want to interrupt. yeah
From the publisher
Our second wave of speakers for AI Engineer World’s Fair were announced! The conference sold out of Platinum/Gold/Silver sponsors and Early Bird tickets! See our Microsoft episode for more info and buy now with code LATENTSPACE.
This episode is straightforwardly a part 2 to our ICLR 2024 Part 1 episode, so without further ado, we’ll just get right on with it!
Timestamps
[00:03:43] Section A: Code Edits and Sandboxes, OpenDevin, and Academia vs Industry — ft. Graham Neubig and Aman Sanger
* [00:07:44] WebArena
* [00:18:45] Sotopia
* [00:24:00] Performance Improving Code Edits
* [00:29:39] OpenDevin
* [00:47:40] Industry and Academia
[01:05:29] Section B: Benchmarks
* [01:05:52] SWEBench
* [01:17:05] SWEBench/SWEAgent Interview
* [01:27:40] Dataset Contamination Detection
* [01:39:20] GAIA Benchmark
* [01:49:18] Moritz Hart - Science of Benchmarks
[02:36:32] Section C: Reasoning and Post-Training
* [02:37:41] Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection
* [02:51:00] Let’s Verify Step By Step
* [02:57:04] Noam Brown
* [03:07:43] Lilian Weng - Towards Safe AGI
* [03:36:56] A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis
* [03:48:43] MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework
[04:00:51] Bonus: Notable Related Papers on LLM Capabilities
Section A: Code Edits and Sandboxes, OpenDevin, and Academia vs Industry — ft. Graham Neubig and Aman Sanger
* Guests
* Aman Sanger - Previous guest and NeurIPS friend of the pod!
* WebArena
*
* Sotopia (spotlight paper, website)
*
* Learning Performance-Improving Code Edits
* Morph Labs, Jesse Han
* LiteLLM
* the role of code in reasoning
* Language Models of Code are Few-Shot Commonsense Learners
* Industry vs academia
* the matryoshka embeddings incident
* other directions
Section A timestamps
* [00:00:00] Introduction to Guests and the Impromptu Nature of the Podcast
* [00:00:45] Graham's Experience in Japan and Transition into Teaching NLP
* [00:01:25] Discussion on What Constitutes a Good Experience for Students in NLP Courses
* [00:02:22] The Relevance and Teaching of Older NLP Techniques Like Ngram Language Models
* [00:03:38] Speculative Decoding and the Comeback of Ngram Models
* [00:04:16] Introduction to WebArena and Zotopia Projects
* [00:05:19] Deep Dive into the WebArena Project and Benchmarking
* [00:08:17] Performance Improvements in WebArena Using GPT-4
* [00:09:39] Human Performance on WebArena Tasks and Challenges in Evaluation
* [00:11:04] Follow-up Work from WebArena and Focus on Web Browsing as a Benchmark
* [00:12:11] Direct Interaction vs. Using APIs in Web-Based Tasks
* [00:13:29] Challenges in Base Models for WebArena and the Potential of Visual Models
* [00:15:33] Introduction to Zootopia and Exploring Social Interactions with Language Models
* [00:16:29] Different Types of Social Situations Modeled in Zootopia
* [00:17:34] Evaluation of Language Models in Social Simulations
* [00:20:41] Introduction to Performance-Improving Code Edits Project
* [00:26:28] Discussion on DevIn and the Future of Coding Agents
* [00:32:01] Planning in Coding Agents and the Development of OpenDevon
* [00:38:34] The Changing Role of Academia in the Context of Large Language Models
* [00:44:44] The Changing Nature of Industry and Academia Collaboration
* [00:54:07] Update on NLP Course Syllabus and Teaching about Large Language Models
* [01:00:40] Call to Action: Contributions to OpenDevon and Open Source AI Projects
* [01:01:56] Hiring at Cursor for Roles in Code Generation and Assistive Coding
* [01:02:12] Promotion of the AI Engineer Conference
Section B: Benchmarks
* Carlos Jimenez & John Yang (Princeton) et al: SWE-bench: Can Language Models Resolve Real-world Github Issues? (ICLR Oral, Paper, website)
* “We introduce SWE-bench, an evaluation framework consisting of 2,294 software engineering problems drawn from real GitHub issues and corresponding pull requests across 12 popular Python repositories.
Given a codebase along with a description of an issue to be resolved, a language model is tasked with editing the codebase to address the issue. Resolving issues in SWE-bench frequently requires understanding and coordinating changes across multiple functions, classes, and even files simultaneously, calling for models to interact with execution environments, process extremely long contexts and perform complex reasoning that goes far beyond traditional code generation tasks.
Our evaluations show that both state-of-the-art proprietary models and our fine-tuned model SWE-Llama can resolve only the simplest issues. The best-performing model, Claude 2, is able to solve a mere 1.96% of the issues. Advances on SWE-bench represent steps towards LMs that are more practical, intelligent, and autonomous.”
* Yonatan Oren et al (Stanford): Proving Test Set Contamination in Black-Box Language Models (ICLR Oral, paper, aman tweet on swebench contamination)
* “We show that it is possible to provide provable guarantees of test set contamination in language models without access to pretraining data or model weights. Our approach leverages the fact that when there is no data contamination, all orderings of an exchangeable benchmark should be equally likely. In contrast, the tendency for language models to memorize example order means that a contaminated language model will find certain canonical orderings to be much more likely than others. Our test flags potential contamination whenever the likelihood of a canonically ordered benchmark dataset is significantly higher than the likelihood after shuffling the examples.
* We demonstrate that our procedure is sensitive enough to reliably prove test set contamination in challenging situations, including models as small as 1.4 billion parameters, on small test sets of only 1000 examples, and datasets that appear only a few times in the pretraining corpus.”
* Outstanding Paper mention: “A simple yet elegant method to test whether a supervised-learning dataset has been included in LLM training.”
* Thomas Scialom (Meta AI-FAIR w/ Yann LeCun): GAIA: A Benchmark for General AI Assistants (paper)
* “We introduce GAIA, a benchmark for General AI Assistants that, if solved, would represent a milestone in AI research. GAIA proposes real-world questions that require a set of fundamental abilities such as reasoning, multi-modality handling, web browsing, and generally tool-use proficiency.
* GAIA questions are conceptually simple for humans yet challenging for most advanced AIs: we show that human respondents obtain 92% vs. 15% for GPT-4 equipped with plugins.
* GAIA's philosophy departs from the current trend in AI benchmarks suggesting to target tasks that are ever more difficult for humans. We posit that the advent of Artificial General Intelligence (AGI) hinges on a system's capability to exhibit similar robustness as the average human does on such questions. Using GAIA's methodology, we devise 466 questions and their answer.
*
* Mortiz Hardt (Max Planck Institute): The emerging science of benchmarks (ICLR stream)
* “Benchmarks are the keystone that hold the machine learning community together. Growing as a research paradigm since the 1980s, there’s much we’ve done with them, but little we know about them. In this talk, I will trace the rudiments of an emerging science of benchmarks through selected empirical and theoretical observations. Specifically, we’ll discuss the role of annotator errors, external validity of model rankings, and the promise of multi-task benchmarks. The results in each case challenge conventional wisdom and underscore the benefits of developing a science of benchmarks.”
Section C: Reasoning and Post-Training
* Akari Asai (UW) et al: Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection (ICLR oral, website)
* (Bad RAG implementations) indiscriminately retrieving and incorporating a fixed number of retrieved passages, regardless of whether retrieval is necessary, or passages are relevant, diminishes LM versatility or can lead to unhelpful response generation.
* We introduce a new framework called Self-Reflective Retrieval-Augmented Generation (Self-RAG) that enhances an LM's quality and factuality through retrieval and self-reflection.
* Our framework trains a single arbitrary LM that adaptively retrieves passages on-demand, and generates and reflects on retrieved passages and its generations using special tokens, called reflection tokens. Generating reflection tokens makes the LM controllable during the inference phase, enabling it to tailor its behavior to diverse task requirements.
* Self-RAG (7B and 13B parameters) outperforms ChatGPT and retrieval-augmented Llama2-chat on Open-domain QA, reasoning, and fact verification tasks, and it shows significant gains in improving factuality and citation accuracy for long-form generations relative to these models.
* Hunter Lightman (OpenAI): Let’s Verify Step By Step (paper)
* “Even state-of-the-art models still regularly produce logical mistakes. To train more reliable models, we can turn either to outcome supervision, which provides feedback for a final result, or process supervision, which provides feedback for each intermediate reasoning step.
* We conduct our own investigation, finding that process supervision significantly outperforms outcome supervision for training models to solve problems from the challenging MATH dataset. Our process-supervised model solves 78% of problems from a representative subset of the MATH test set. Additionally, we show that active learning significantly improves the efficacy of process supervision.
* To support related research, we also release PRM800K, the complete dataset of 800,000 step-level human feedback labels used to train our best reward model.
*
* Noam Brown - workshop on Generative Models for Decision Making
* Solving Quantitative Reasoning Problems with Language Models (Minerva paper)
* Describes some charts taken directly from the Let’s Verify Step By Step paper listed/screenshotted above
.
* Lilian Weng (OpenAI) - Towards Safe AGI (ICLR talk)
* OpenAI Instruction Hierarchy: The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
Section D: Agent Systems
* Izzeddin Gur (Google DeepMind): A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis (ICLR oral, paper)
* [Agent] performance on real-world websites has still suffered from (1) open domainness, (2) limited context length, and (3) lack of inductive bias on HTML.
* We introduce WebAgent, an LLM-driven agent that learns from self-experience to complete tasks on real websites following natural language instructions.
* WebAgent plans ahead by decomposing instructions into canonical sub-instructions, summarizes long HTML documents into task-relevant snippets, and acts on websites via Python programs generated from those.
* We design WebAgent with Flan-U-PaLM, for grounded code generation, and HTML-T5, new pre-trained LLMs for long HTML documents using local and global attention mechanisms and a mixture of long-span denoising objectives, for planning and summarization.
* We empirically demonstrate that our modular recipe improves the success on real websites by over 50%, and that HTML-T5 is the best model to solve various HTML understanding tasks; achieving 18.7% higher success rate than the prior method on MiniWoB web automation benchmark, and SoTA performance on Mind2Web, an offline task planning evaluation.
* Sirui Hong (DeepWisdom): MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework (ICLR Oral, Paper)
* We introduce MetaGPT, an innovative meta-programming framework incorporating efficient human workflows into LLM-based multi-agent collaborations. MetaGPT encodes Standardized Operating Procedures (SOPs) into prompt sequences for more streamlined workflows, thus allowing agents with human-like domain expertise to verify intermediate results and reduce errors. MetaGPT utilizes an assembly line paradigm to assign diverse roles to various agents, efficiently breaking down complex tasks into subtasks involving many agents working together.
Bonus: Notable Related Papers on LLM Capabilities
This includes a bunch of papers we wanted to feature above but could not.
* Lukas Berglund (Vanderbilt) et al: The Reversal Curse: LLMs trained on “A is B” fail to learn “B is A” (ICLR poster, paper, Github)
* We expose a surprising failure of generalization in auto-regressive large language models (LLMs). If a model is trained on a sentence of the form ''A is B'', it will not automatically generalize to the reverse direction ''B is A''. This is the Reversal Curse.
* The Reversal Curse is robust across model sizes and model families and is not alleviated by data augmentation. We also evaluate ChatGPT (GPT-3.5 and GPT-4) on questions about real-world celebrities, such as ''Who is Tom Cruise's mother? [A: Mary Lee Pfeiffer]'' and the reverse ''Who is Mary Lee Pfeiffer's son?''. GPT-4 correctly answers questions like the former 79\% of the time, compared to 33\% for the latter.
*
* Omar Khattab (Stanford): DSPy: Compiling Declarative Language Model Calls into State-of-the-Art Pipelines (ICLR Spotlight Poster, GitHub)
* presented by Krista Opsahl-Ong
* “Existing LM pipelines are typically implemented using hard-coded “prompt templates”, i.e. lengthy strings discovered via trial and error. Toward a more systematic approach for developing and optimizing LM pipelines, we introduce DSPy, a programming model that abstracts LM pipelines as text transformation graphs, or imperative computational graphs where LMs are invoked through declarative modules.
* DSPy modules are parameterized, meaning they can learn how to apply compositions of prompting, finetuning, augmentation, and reasoning techniques.
* We design a compiler that will optimize any DSPy pipeline to maximize a given metric, by creating and collecting demonstrations.
* We conduct two case studies, showing that succinct DSPy programs can express and optimize pipelines that reason about math word problems, tackle multi-hop retrieval, answer complex questions, and control agent loops.
* Within minutes of compiling, DSPy can automatically produce pipelines that outperform out-of-the-box few-shot prompting as well as expert-created demonstrations for GPT-3.5 and Llama2-13b-chat. On top of that, DSPy programs compiled for relatively small LMs like 770M parameter T5 and Llama2-13b-chat are competitive with many approaches that rely on large and proprietary LMs like GPT-3.5 and on expert-written prompt chains.
*
* MuSR: Testing the Limits of Chain-of-thought with Multistep Soft Reasoning
* Scaling Laws for Associative Memories
* DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models
* Efficient Streaming Language Models with Attention Sinks
Get full access to Latent.Space at www.latent.space/subscribe




