In short
Podcast Summary: Latent Space - Episode on SWE-Bench-Dead
Podcast Overview Podcast Title: Latent Space: The AI Engineer Podcast Description: A podcast for AI Engineers covering developments in Software 3.0, including interviews and discussions on cutting-edge technology in AI, focusing on Foundation Models, Code Generation, and more. Website: [Latent Space](https://latent.space)
Episode Details Episode Title: ⚡️SWE-Bench-Dead: The End of SWE-Bench Verified Guests: Mia Glaese (VP of Research, OpenAI) & Olivia Watkins (Frontier Evals Team) Summary: The episode discusses the discontinuation of SWE-Bench Verified as a coding benchmark due to its saturation and contamination issues. The focus will shift to SWE-Bench Pro, which is designed to be more robust and diverse.
Key Topics Discussed
Introduction to Guests
- Olivia Watkins: Member of the Frontier Evals team.
- Mia Glaese: VP of Research at OpenAI, oversees Codex, human data, and alignment teams.
Background on SWE-Bench Verified
- Origins: SWE-Bench Verified, created from the original Princeton SWE-Bench, was intended to provide a reliable coding benchmark through extensive human review by software engineers.
- Problem Identification: The team observed that SWE-Bench Verified had become saturated, resulting in less effective measurement of coding progress.
Issues with SWE-Bench Verified
- Saturation and Contamination:
- The benchmark was deemed less useful due to contamination (e.g., models recalling specific implementation details).
- Examples of unfair tests that focused on narrow specifics rather than overall coding capability were highlighted.
- Human Review Efforts: A large-scale review involving nearly 100 engineers aimed to curate a set of better-quality tasks.
Transition to SWE-Bench Pro
- Reasons for Transition:
- SWE-Bench Pro is seen as a more difficult and diverse benchmark, utilizing longer tasks (1-4 hours).
- Initial findings suggest significantly reduced contamination compared to SWE-Bench Verified.
- Future Directions: Discussion on what benchmarks should measure, including:
- Open-ended design decisions.
- Code quality and maintainability.
- Real-world product-building challenges.
Long-term Goals in Coding Benchmarks
- Beyond Pass/Fail: Emphasis on the need for a more nuanced evaluation that includes:
- The ability to handle longer-term tasks.
- The implementation of design decisions reflecting real-world scenarios.
Grading and Evaluation Techniques
- Trade-offs: Discussion on automated grading vs. human evaluation.
- Importance of Real-world Applications: Evaluating the impact of AI in practical coding situations and understanding the balance between human jobs and AI efficiency.
OpenAI's Preparedness Framework
- Overview: The framework aims to track the dual-use capabilities of AI, focusing on risks related to coding and model autonomy.
- Call for Collaboration: Encouragement for the AI community to create and share evaluations that measure various capabilities in coding.
Key Takeaways
- End of SWE-Bench Verified: Recognized as ineffective due to issues of saturation and contamination.
- Introduction of SWE-Bench Pro: Aimed to offer a more comprehensive and challenging evaluation of coding capability.
- Future of Benchmarks: A shift towards assessing more qualitative aspects of coding, including design choices and code maintainability.
- Collaboration in the Field: Importance of community-driven evaluations to accurately measure progress in AI capabilities.
Conclusion The episode highlights a significant shift in how coding benchmarks will be approached in the AI engineering field. The transition from SWE-Bench Verified to SWE-Bench Pro reflects a commitment to evolving the standards by which AI coding capabilities are measured, ensuring relevance to real-world applications and long-term tasks.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOEvolution of C-Bench Verified
0:45 to 2:30
Discussion on the development and significance of C-Bench Verified in coding benchmarks.
“So you've seen the evolution of coding benchmarks over time.”
Saturation and Contamination in Benchmarks
2:30 to 5:00
Exploration of the saturation and contamination issues in C-Bench Verified.
“So Vuxo OpenAI did a pretty extensive human data campaign, hiring like almost 100 real-world software engineers to go through the problems and figure out like, are the tasks well-specified?”
Creation and Challenges of C-Bench Verified
5:00 to 7:35
The process and challenges involved in creating C-Bench Verified and its impact.
“But in the GPT 5.2 train of thought, we actually saw instances of the model reasoning like, hey, I think that it's some linear version of this repository that implemented this particular argument.”
Impact of Benchmarking Practices
7:35 to 8:05
Insights on how OpenAI's practices have influenced benchmark verification.
“Then the tests are looking for you naming that argument or that function with a particular name.”
Analyzing Model Performance and Fairness
8:05 to 11:15
Examination of model performance and the fairness of benchmark tests.
“I think it's important that you're doing this because in some way it is you in 2025, 6, going back in time and correcting your own work.”
Future of Coding Benchmarks
11:15 to 14:01
Discussion on the advancements needed in coding benchmarks to accurately assess AI capabilities.
“patch and the task ID and told to go take this target model and kind of as an open-ended like set of questions, try to find questions that will manage to kind of reveal what contamination might be lurking in that model.”
Evaluating Model Capabilities in Coding
14:01 to 15:06
Explore the evolving benchmarks for assessing AI coding agents.
“And like, obviously there's like some issues with the benchmark that means that now that we're at like 80%, we don't really trust like further improvements on it.”
Understanding GDPVAL and Its Implications
15:07 to 17:23
Learn about the GDPVAL evaluation and its significance in measuring AI performance.
“So, I mean, these are all qualities that are obviously no longer the low hanging fruit.”
The Challenges of AI Coding Evaluations
17:24 to 19:15
Discuss the complexities of creating effective AI coding evaluations.
“where we are seeing like alpha benchmarks like this being super popular, it's like, it's very easy.”
Integrating Financial Metrics in Evaluations
19:16 to 21:31
Examine how financial metrics could enhance AI coding evaluations.
“Because a lot of the, like, you know, state-of-the-art AI code bases are proprietary.”
Show all 12 chapters
The Preparedness Framework Explained
21:32 to 22:54
Understand OpenAI's preparedness framework for tracking risks.
“I think, like, complexity, however we can sort of, like, quantify it, is really important to understand, like, where our models are getting to.”
Future Directions for AI Evaluations
22:55 to 24:24
Consider the future of evaluations and what metrics need to be tracked.
“Any other anything else to add on just the general what people should know about preparedness and how evals and human data alignment all work together?”
Transcript
Automatic transcript. May contain errors.0:04Okay.
0:04Mia Glaese:Hi, we're here in the OpenAI studio with Mia and Olivia from the Frontier Evals team, or however you want to introduce yourself. Maybe you want to introduce, name what you do at OpenAI and we can get it started.
0:16Olivia Watkins:Sure. Hi, I'm Olivia. I'm on the Frontier Evals team. Hi, I'm Mia. I am a VP of research at OpenAI and my teams are the Codex team, the simulator team, and the alignment team. And we work a lot with Olivia's team on Frontier.
0:33Mia Glaese:Yeah, very exciting. And as by my understanding, you were part of the original team that worked on C-Bench Verified as well.
0:38Olivia Watkins:Yeah, Olivia's team, the Frontier events team and the human data team collaborated on creating C-Bench Verified.
0:45Mia Glaese:So you've seen the evolution of coding benchmarks over time. And I think it was roundabout to the mid to late 2024 when you first covered C-Bench Verified. Things have evolved a lot since then. What's the blog post that you have worked on a year that we're releasing today? Like what is the main thesis that you're pushing out?
1:04Olivia Watkins:So the main thesis is that SweetBenge Verified has been one of the North Star coding benchmarks that the field has looked at to measure coding progress. But recently, we've seen that progress has kind of stalled. And this we realized that this is because the eval is effectively saturated and also highly contaminated. So at this point, we think that it's not really measuring coding performance improvements well anymore. and we think that if you'll just move away from this towards other benchmarks.
1:28Mia Glaese:Like Sweetbench Pro.
1:29Olivia Watkins:Like Sweetbench Pro, yeah.
1:30Mia Glaese:Amazing. Yeah, one of the jokes I always have is like, there's a group chat with all the labs and everyone just turns to increment like 0.1 on trucks and then it's like, okay, well, you have the best coding model, I guess, because you're 0.1 % higher, but it's not super convincing at this point. No. Yeah. So cool. I think let's sort of reset on like, what was the original work that you guys did for Sweetbench Verified, which I think was pretty substantial. It was a very significant investment for an open AI, which people still don't appreciate. And then what were the satisfactions that we found over time, right?
2:03Mia Glaese:So what was Sweetbench Verified that people should know about?
2:07Olivia Watkins:Sweetbench Verified was kind of a cleanup of original academic benchmark from a lab at Princeton called Sweetbench. And the agent is basically given a code base and a task that was sourced from a real-world repository and GitHub issue. and was asked to solve the task and is graded on whether some tests pass. And at the time, this quickly became a popular benchmark because at the time, the field didn't really have good real-world coding benchmarks. But then when OpenAI took a look at the benchmark as part of one of the evals we wanted to track in our preparedness framework, folks started realizing that some of the cases where agents were failing were due to bad problem setups rather than just to models being dumb.
2:49Olivia Watkins:So Vuxo OpenAI did a pretty extensive human data campaign, hiring like almost 100 real-world software engineers to go through the problems and figure out like, are the tasks well-specified? Are the tests actually fair? And kind of created a curated set of like 500 tasks that we thought were much better. It's just, maybe it's hard to overstate like the amount of effort that it took to like create that benchmark. is literally like many export software engineers reviewing the problems like differentially multiple times and to, you know, basically like three different exports independently decided that.
3:29Olivia Watkins:Yeah, you didn't have to do that.
3:30Mia Glaese:You just tripled your costs for just...
3:32Olivia Watkins:I mean, we had to do it actually because it's quite a hard task to like look at something like a problem and the patch. And then like, it's not just the problem and the patch, right? You have to understand it in the context of the code base that the human or the models and to solve the task. So it's a very complex problem. And it was definitely needed to have three reviews. And I think maybe we should have done more. But it was definitely a lot of effort to get there.
4:00Mia Glaese:Yeah. And there's more, but people can read the blog post for that. I will note that you guys set a trend in verifying benchmarks because I just recently saw, I think Quinn had a HLE verified for Humanities License and Verified. Yeah. So now everyone's verifying everything, which is nice and good and extra quality there. Okay. But I think that the meat of it is that this was a lot of like, well, here's the issue or problem statements. And then here's the divs. Here's the golden tests. And here's some regression tests, right? That's like the rough setup of these 500 problems. And there's some contamination always happens because all the students verified was fully open, I think.
4:36Mia Glaese:Like you did have canaries, but like, you know, stuff leaks.
4:40Olivia Watkins:There's like multiple avenues that like the problems are sourced from open source repos. Yes. So it's not just like when we usually publish evaluations, we publish evaluations and then we add canary strings to ensure that, you know, they are easily filtered out at training time. obviously if you use sort of like data from like open market you don't have actually like a canary string and yeah and you and these are also like some of these are very popular repos like the jango repository so you're going to see like many instances being used kind of throughout gab
5:17Mia Glaese:yeah you just before recording you were telling me that you found this in your own chain of thought with the 5.2 also seeing that like they had extra knowledge or something yes so this was an example
5:27Olivia Watkins:where the task asked the agent to influence something, but it wasn't told that there was this specific argument that the test was going to be looking for it using. But in the GPT 5.2 train of thought, we actually saw instances of the model reasoning like, hey, I think that it's some linear version of this repository that implemented this particular argument. Maybe I should add it in. So this is an example of a test that would be pretty impossible to pass without this contamination knowledge. Yeah. And I think you found that sort of forced, right? And it triggered like a whole investigation, both like in our own models and also in other frontier models, like in the market and like understanding how contaminated the benchmark is like across the industry.
6:07Mia Glaese:What else did you find? I mean, I have to double click on this.
6:11Olivia Watkins:So we, when I say we, this is mostly we are from other folks that are tape, not the regular way. Yes. But so we did some analysis on, first of all, are the tests actually fair? And so this happened by first taking all the problems that O3 couldn't solve reliably. And then again, getting a lot of humans to do basically another pass of kind of digging into what's wrong.
6:34Mia Glaese:Is it the same exact analysis or were they reading O3's output and going just where O3 went wrong?
6:40Olivia Watkins:I think it was, I mean, it was definitely like a scope to the set of problems that models failed. And I believe they were able to look at like what the model solutions look like versus what the...
6:49Mia Glaese:So this isn't the same work as the original...
6:51Olivia Watkins:It's not exactly the same work. It was like a deeper dive. It's like, okay, which are the problems that we don't see any murder solving? It's like, is there something fundamentally wrong with those problems? Or is there something wrong with the other murder, just not smart enough to solve the problems? So that's kind of like what we dug into.
7:09Mia Glaese:Yeah, and you found some?
7:10Olivia Watkins:Oh, yes. Like in over half of the problems that were investigated in that deep dive, there was one problem or the other. I think the most common problem are like overly narrow tests where there's some particular implementation detail that the tests were looking for, but wasn't specified in the problem description. So it wasn't fair to expect that model to make that particular design choice. Like one pretty blatant example are cases where the task asks you to implement some feature. Then the tests are looking for you naming that argument or that function with a particular name. But if you chose another reasonable name, the test would fail.
7:44Olivia Watkins:And another set of types of bad tests are tests that are just looking for additional features that were never mentioned in the problem description. Well, that's a significant effect. That means that if you pass a test, actually, like you probably did like a really good job. But just because you didn't pass a bad test doesn't mean that your implementation wasn't like a good one. Right. So it was just like we only accept like very narrow versions of solutions and like not the whole space of like viable and sort of like good solutions to the problem.
8:16Mia Glaese:Yeah. I think it's important that you're doing this because in some way it is you in 2025, 6, going back in time and correcting your own work. Right. Because you could have caught all this in the original verified work. I think so.
8:28Olivia Watkins:It's definitely much harder to find a problem in the abstract than when you're looking at a very smart agent's best effort solution and trying to compare it.
8:36Mia Glaese:It is harder or harder?
8:38Olivia Watkins:It's much easier when you have the solution. Exactly. I think also like at the time when 3BenchVerified was published, I think it was like a very strong benchmark. It's not like we're like, oh, this is not, this wasn't like a strong benchmark at the time. I think this is something that a lot of benchmarks go through like as an evolution, right? Like when they start to become like popular and like viable because they measure something like important and models maybe do like 20 % correct on them, sometimes even less. And sort of like people have something to hold on and improve models on these benchmarks.
9:11Olivia Watkins:And by the time that you hit like very high performance on the benchmarks, like additional like 0.1 % improvements become sort of like meaningless. And so like at the time, I think, you know, that benchmark was like super valuable and it taught like us and the industry a lot. It's just like now at the point that we are at now, where my audits are as strong as they are now, we're kind of starting to measure not necessarily like what we want to measure, which is like coding capability of our agents, but like the agent's ability to correctly guess how to name a specific function. and that isn't really what we are like want to measure at this point yeah i think that's fair
9:53Mia Glaese:is there i mean if i if i asked you to ballpark it like most models are most frontier models are like 80 something is there like what's the actual like number on superman verify that you did you guess as like the ceiling or i guess it's really hard to say like i when gipty 5.2
10:12Olivia Watkins:came out. Folks took a look and found that it was solving like 31 problems that were in the set of should be very hard to solve without contamination problems. So I think it's quite possible that that number is already something that we've hit if you didn't have contamination at all. Fair enough. Hard to say though.
10:27Mia Glaese:Yeah. Cool. We're going to stop reporting CBench Verified, right? And then CBench Pro will be signed with the next one, which is an effort from scale. What's your sort of comparison analysis? What attracts you to CBench Pro?
10:37Olivia Watkins:The first one, I think, is just that it's harder. For CBench Verified, I think I'm in like 90 % of the problems are things that were estimated to take like an expert software engineer, like less than an hour. They're very well specified, very self-contained. And the Swoob Edge Pro problems are just bigger and harder. And there's much more heavy one that you've got because it's not saturated.
10:56Mia Glaese:Yeah, like categories of like one to four hours and four plus.
Read the full transcript
11:00Olivia Watkins:Yeah, and it's more diverse. Lots of repositories, multiple languages, qualitatively more different types of problems. So all that's great. On the contamination side, we also think it's better there. So the way we were measuring for contamination for SweetBengeVerified was with this little like contamination auditor agent, which is given the description of the task and the patch and the task ID and told to go take this target model and kind of as an open-ended like set of questions, try to find questions that will manage to kind of reveal what contamination might be lurking in that model. And in SweetBengeVerified, we found many instances of contamination across open-eyed models, across like quad opus 4.5, Gemini Flash.
11:42Olivia Watkins:And in all of these, we saw things like regurgitating the ground truth solutions, things like, in some cases, giving like the task IDs and other things that are pretty clear evidence of minimum familiarity was the repositories. Yeah. So we've been doing a task ID that thing. Yeah. So on the other hand, we don't see this. I think the auto agent found some like very light evidence that maybe a couple of models might be very lightly familiar with like one or two of the source repositories, but it's very different than Subint verified. So less contamination is good. I think that also, like we should expect that at some point, like that's not going to be like the right benchmark anymore.
12:22Olivia Watkins:And like it's a field we kind of have to continue to like move on and like find harder and more representative problems that we can match our capabilities on.
12:31Mia Glaese:Awesome. So let's go into that. I think that there are a lot of, I think we also practiced in the pre-chat was people feel a qualitative difference when they're using 5.1 to 5.2 to 5.3. And it's not super expressed in these benchmarks because they are on a number of these things. What capabilities do you really want to benchmark in an ideal coding benchmark? You know, I guess, like agentic coding benchmark, whatever you call it.
12:55Olivia Watkins:I mean, one thing is kind of open-ended design decisions, places where the problem maybe is a little bit underspecified and seeing if the model can make reasonable design decisions.
13:05Mia Glaese:What's a reasonable prompt for that? Like this Vibe Code Me, a B2B SaaS, and make no mistakes? Or, you know, that's the meme. But like, okay, what's like an actual usable, open-ended problem like that?
13:17Olivia Watkins:Sure. I mean, maybe an example could be finding a way to speed up a particular part of a code base. But there might be multiple different ways to speed up.
13:25Mia Glaese:Yeah, there are dedicated performance benchmarks. I think you guys have one. Is that efficiency? Is that, I don't know. I think that's Ophiraheus' group. But yeah, I mean, that is a good one.
13:33Olivia Watkins:I think there's just many, many things that people like value about working with software engineering agents. They think 3Bend Verified obviously measures like some important capability, which is like given like a description of a GitHub issue, can you produce like a patch that solves that issue, you know, satisfactorily? And like, obviously there's like some issues with the benchmark that means that now that we're at like 80%, we don't really trust like further improvements on it. But like it does measure something that is like a real capability of models. But I think as a field, we're like moving beyond sort of, you know, can my coding agent like solve a small like GitHub issue for me, right?
14:23Olivia Watkins:And so we are starting to look at like much more longer term tasks, right? Like that don't take like 15 minutes, but maybe like hour, sometimes days. And then beyond sort of like what kind of tasks can my agent solve? Like there might be things that are kind of a bit harder to grasp, right? Like Olivia talked about sort of like, does it have like design taste, right? Like, does it solve the problem the way that, you know, my team likes to solve problems? Is the code nice, right? Like, is it well written? Is it sort of like clean code, right? Like, people care about this. Is it maintainable in the future?
15:05Olivia Watkins:People care about a lot of these maybe less tangible and harder to measure, frankly, things that are still super meaningful for people that are working with coding. Yeah.
15:19Mia Glaese:So, I mean, these are all qualities that are obviously no longer the low hanging fruit. Like we have no idea how to eat all this. I think the simple question, maybe there's sort of two forks in the road. One is the sort of very human intensive, money intensive path, which is hire a bunch of contractors and try to annotate this. The other is use an LLM to proxy it and try to align the LLM so that it can give you a reasonable proxy. Which of those would you want? You want to do both?
15:46Olivia Watkins:I think maybe you should talk about GDP law as an example. Sure. So GDPVAL is an EVAL that was, again, produced by a collaboration between Human Data Team and the Front of EVALs team. And it's trying to measure whether agents can do kind of a variety of real-world white-collar work. That was an EVAL where grading is very hard, requires kind of a lot of knowing knowledge on exactly what are you looking for in each different context.
16:16Mia Glaese:Yeah, across like 15, 16 white collar jobs, professions, like that take of a significant part of GDP.
16:23Olivia Watkins:High level professions and a lot of like different granular subprimations.
16:26Mia Glaese:I have said, like, I'm a big fan. It's so it is. This is the evil for AGI, basically.
16:32Olivia Watkins:But part of because it was so hard to require so much kind of like domain knowledge that the human data team hired like a lot of people from these professions to be very involved in creating tasks and creating the gold solutions. and trying to help create rubrics and so forth so we can grow it reliably.
16:49Mia Glaese:So basically take the GDP value, which is a generalist thing, take that same approach to apply it to code and you roughly have like a rough road.
16:58Olivia Watkins:I think it's an interesting solution. I think what you're pointing out is an important problem, which is sort of this like how realistic is it? And like do, you know, what we want to do is like coding agents should write code that, you know, we think is good. And so it's like asking human, it's actually like a good way to ensure that. It's also kind of a slower, like complex way to do that. And so part of why I think, you know, 3Brench Verified ended up being super popular and where we are seeing like alpha benchmarks like this being super popular, it's like, it's very easy. It could even be easier, but like validating that a solution passes all the tests is like fairly trivial once you can like run the tests in like your, on your computer or wherever you're running them and you can kind of like okay is it correct or is it not correct and you can kind of aggregate that and that it's super simple but it doesn't tell you it's like you know did the model like solve the problem like wow like you know agree with like what if actually like an open source maintainer of that project have like merged that pr like that it doesn't tell you but there is a lot of value in having benchmarks that are both like easy to compare across the industry and also that can be sort of run really fast without human involvement.
18:16Mia Glaese:Yeah. Amazing. Your team's also put out other kinds of evals that are related. I think there's an RL paper bench and then the more sort of recursive self-improvement type evals. How much should that figure into mainstream coding evals? Is there some way in which those things join together?
18:38Olivia Watkins:So we're asking, like, should we also be e-voting e-vils for the self-improvement e-vils? Are you saying do coding e-vils currently cover that?
18:46Mia Glaese:I think, I just think, like, those are some of the most advanced e-vils that we have. And we're not using them in the normal path. And it's just, it's an interesting split between, well, here's e-vils for coding normal things. And then here's the one for machine learning. That is, like, completely different, right? I think you get what I mean. That's mostly a safety argument, I guess. But also, like, it's actually really useful for people to understand if the model is really good at, like, AI code, basically. Yeah.
19:15Olivia Watkins:Like, my guess is that part of the reason that a lot of benchmarks so far haven't focused as much on the AI coding is just a question of, like, what data sets are easy to gather. Yeah. Because a lot of the, like, you know, state-of-the-art AI code bases are proprietary. So if we make emails for that, like, we're probably not going to release them. and it's harder for people in the field to make evals that kind of measure, like, is this a realistic research coding workflow? I do think that it's good for the field to try to measure these skills in a public way. I think it's harder to make it realistic.
19:45Mia Glaese:And then one more thing that a lot of people are trying to do, which is like sort of, well, instead of like a percentage of zero to 100, maybe we re-denominate in dollars, right? So you had freelancer and all that. Other people are doing like vending bench, whatever. Any alpha in those? or are they, you still want like a traditional academic benchmark?
20:03Olivia Watkins:I think in a way, like there's like different ways to measure the same thing, right? If we're like, oh, this is like how much money it produces. It's a fairly similar thing to saying like, oh, this problem would take like a human, you know, two hours to solve or something like that. Usually they're like fairly like correlated, right? Like however, you know, much it would take like a human to solve that problem kind of determines the value that we ascribe like a solution and so i do think that is like an important thing is like how complex and how sort of long running are the tasks that we are like able to entrust our agents with yeah and so i think that that's like an important piece but i think here sort of monetary value time a complexity they all kind of like try to capture or like a similar thing.
20:56Mia Glaese:Yeah, okay. So they're all proxies for some amount of increasing capacity that we want to measure. I think that's a good thing. I think the only other sort of major player in this field is meter, which has done the sort of long graphs. And congrats, you guys have completely destroyed the curve for that. Any takes on that? Obviously, you come up really well, so like it looks good. But I don't know if that approach is something that you want to incorporate in your work, making it else. This is the long autonomy test, Fuhrer.
21:21Olivia Watkins:Yeah, and we're away from here. We are, and we work with MEDAR on these evaluations. So, like, we do appreciate them. I think then they're using time, right? They're not using money. So I think that was your question. I think, like, complexity, however we can sort of, like, quantify it, is really important to understand, like, where our models are getting to.
21:42Mia Glaese:Okay. Complexity is the abstract thing, and it projects down to time, projects down to story points, whatever, dollars. Great. One last question on just, like, just the overall preparedness framework. because I was actually kind of looking at it. People mention the preparedness framework a lot. I don't think it's well explained to a lot of people. And you actually have a nice website where it's like, I think it's like test and like inform and teach something. And I feel like you actually do a lot of work there. And I don't know if you want to talk about how the preparedness framework applies.
22:08Olivia Watkins:So the preparedness framework is OpenEye's kind of like public framework for how we track frontier risks. So these are kind of capabilities that are typically dual use. Like you can use them for good things or bad things, but we want to at least keep an eye out for the bad things to make sure that both we as a company and the broader society are kind of prepared to handle the potential downsides. And so at the moment, we kind of track three different categories. One is kind of bio-risk, another is cybersecurity, and a third is kind of research automation and model autonomy. And that's kind of what ties most into the Sweebench, where coding is not all of automating research, but it is one very important key component.
22:47Olivia Watkins:And so we initially created SweeBinj Verified as part of building out evals for that model autonomy work stream. And now I think we have to move beyond that towards looking more at can models actually start to actually automate research on workflows.
23:02Mia Glaese:Yeah, amazing. Any other anything else to add on just the general what people should know about preparedness and how evals and human data alignment all work together? it at?
23:12Olivia Watkins:I think maybe the thing that I would say is that we really appreciate, we work really hard to build these evals and so that's where we published Sweetbench Verified and that's where we're like sharing GDPVal, these sorts of things. We also deeply appreciate like other people and the entire field to kind of build evals and share them and reuse them. Like Sweetbench Pro, we're like, yes, that's a better eval now. We should use them. So we'd really encourage people to find more ways to create and share events that we and the entire field can use to measure progress on a variety of capabilities, including coding, because it's important to understand where we are.
23:57Mia Glaese:Mia had to leave, but we're just kind of talking a little bit about the future directions that we want evals to go. And I think here we can dive in on, Like, give us good work on these, these, these things. We'll talk to you. You know, here's your platform to make a call for what you're looking for.
24:15Olivia Watkins:I think a few things that would be useful. I'd say, first of all, really, really hard task. The kinds of things that would take top-notch engineers months or teams weeks would be quite good. Especially if grading is reliable. And grading is like, you know, you have, for example, like rubrics that have been sourced and validated by many people in the field. I think that would be quite valuable. I think also benchmarks on kind of creating products end to end. I think if people were by putting more, that would be quite useful. I think a third thing that I'd say that is maybe not quite an eval, but I think is still relevant to the kind of overall mission of like we as a field and as a world should be tracking like where are these capabilities going?
24:54Olivia Watkins:I'd like to see more metrics tracking like real world usage, like how much is AI actually being used in the field? How much is it replacing people's jobs? How much is it, you know, augmenting people, speeding people up? Just like real world networks. Yeah.
25:08Mia Glaese:Yeah. The replacement thing is always like a sensitive one on the sort of PR side of things. But, you know, we create new jobs that manage the old jobs. That's how it is. Yourself, like, you know, I think in terms of the frontier evals that OpenAI is really going to excited to push. Like you put out really good work every single time. What should people expect from OpenAI itself?
25:28Olivia Watkins:I'm not sure I can say what we're going to.
25:30Mia Glaese:General directions.
25:31Olivia Watkins:I mean, general directions, I think looking at real world impact, like real world, real
25:37Mia Glaese:GDL2, you know, whatever.
25:38Olivia Watkins:That kind of stuff. Yeah. Yeah.
25:40Mia Glaese:Amazing. Okay. Well, I'm excited for more real world impact. I think you guys have, you know, really made a lot of progress and I think taken a lot of industry leadership for CBench Verified and now moving on to CWR. So thank you for doing this. Thank you for being so transparent. And I think people will respond in kind. Yeah. Yeah. Great for your time. Thank you.
26:02Thank you.
From the publisher
Olivia Watkins (Frontier Evals team) and Mia Glaese (VP of Research at OpenAI, leading the Codex, human data, and alignment teams) discuss a new blog post (https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/) arguing that SWE-Bench Verified—long treated as a key “North Star” coding benchmark—has become saturated and highly contaminated, making it less useful for measuring real coding progress. SWE-Bench Verified originated as a major OpenAI-led cleanup of the original Princeton SWE-Bench benchmark, including a large human review effort with nearly 100 software engineers and multiple independent reviews to curate ~500 higher-quality tasks. But recent findings show that many remaining failures can reflect unfair or overly narrow tests (e.g., requiring specific naming or unspecified implementation details) rather than true model inability, and cite examples suggesting contamination such as models recalling repository-specific implementation details or task identifiers. From now on, OpenAI plans to stop reporting SWE-Bench Verified and instead focus on SWE-Bench Pro (from Scale), which is harder, more diverse (more repos and languages), includes longer tasks (1–4 hours and 4+ hours), and shows substantially less evidence of contamination under their “contamination auditor agent” analysis. We also discuss what future coding/agent benchmarks should measure beyond pass/fail tests—longer-horizon tasks, open-ended design decisions, code quality/maintainability, and real-world product-building—along with the tradeoffs between fast automated grading and human-intensive evaluation. 00:00 Meet the Frontier Evals Team00:56 Why SWE Bench Stalled01:47 How Verified Was Built04:32 Contamination In The Wild06:16 Unfair Tests And Narrow Specs08:40 When Benchmarks Saturate10:28 Switching To SWE Bench Pro12:31 What Great Coding Evals Measure18:17 Beyond Tests Dollars And Autonomy21:49 Preparedness And Future Directions
Get full access to Latent.Space at www.latent.space/subscribe




