[State of Code Evals] After SWE-bench, Code Clash & SOTA Coding Benchmarks recap — John Yang

31 Dec 2025 · 12 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Latent Space: The AI Engineer Podcast

Episode Title

State of Code Evals with John Yang Episode Description This episode features John Yang, the creator of SWE-bench, discussing the evolution and current state of code evaluations as they head into 2026. The conversation covers the impact of benchmarks like SWE-bench on the AI coding agent industry, the emergence of various SWE-bench variants, and the future of coding evaluations.

---

Key Topics Discussed

John Yang's Journey

  • Background: Transition from Princeton's SWE-bench to Stanford PhD with Diyi Yang, focusing on code evaluations and human-AI collaboration.
  • SWE-bench Origin: Initially launched in October 2023, it gained traction after the launch of Cognition's Devin.

SWE-bench Evolution

  • Current Variants: SWE-bench has grown to include:
  • SWE-bench Verified: A curated benchmark for serious evaluations.
  • SWE-bench Multimodal: Involves nine languages across 40 repositories, moving beyond its original Django-heavy focus.
  • SWE-bench Pro: Created by independent authors, recognized by John as a positive contribution to the benchmarking landscape.

CodeClash

  • Concept: A new benchmark aimed at long-horizon development where agents improve their codebases over multiple rounds in a competitive format.
  • Programming Arenas: Incorporates programming games like Halite and explores economically valuable tasks.

Benchmark Landscape

  • Proliferation of Benchmarks: Multiple new benchmarks like:
  • SWE-Efficiency: Optimizes code for speed without changing functionality.
  • AlgoTune: Focuses on algorithm optimization.
  • Terminal-bench: Emphasizes creativity and simulation beyond traditional coding tasks.
  • Tau-bench: Involves "impossible tasks" as a feature to identify cheating.

Challenges and Future Directions

  • Unit Tests Limitation: John expresses skepticism about unit tests for verification and proposes long-running tournaments for evaluating coding agents.
  • Human-AI Collaboration: Discussion on balancing long autonomy with the need for interactive development cycles.
  • Call to Action: Encourages collaboration to enhance user interaction data and improve understanding of codebases for AI.

---

Important Concepts and Arguments

  • SWE-bench's Rise: From obscurity to industry standard due to significant AI advances.
  • Long-Horizon Development: The need for models to engage in longer development cycles rather than one-off tasks.
  • Creative Evaluation: Emphasizing the importance of dynamic environments that mimic real-world programming challenges.
  • Human-AI Interaction: Exploring varied setups in CodeClash to measure changes in interaction patterns as models improve.

---

Conclusion The episode provides a comprehensive overview of the rapidly evolving landscape of code evaluations, driven by benchmarks like SWE-bench and innovative concepts like CodeClash. John Yang's insights highlight the ongoing need to adapt and improve the evaluation processes to better understand and harness the potential of AI in software engineering.

---

Additional Links

  • SWE-bench: [https://www.swebench.com](https://www.swebench.com)
  • John Yang on X: [https://x.com/jyangballin](https://x.com/jyangballin)

---

Podcast Chapters

  1. 00:00:00 - Introduction: John Yang on SWE-bench and Code Evaluations
  2. 00:00:31 - SWE-bench Origins and Devon's Impact on the Coding Agent Arms Race
  3. 00:01:09 - SWE-bench Ecosystem: Verified, Pro, Multimodal, and Multilingual Variants
  4. 00:02:17 - Moving Beyond Django: Diversifying Code Evaluation Repositories
  5. 00:03:08 - CodeClash: Long-Horizon Development Through Programming Tournaments
  6. 00:04:41 - From Halite to Economic Value: Designing Competitive Coding Arenas
  7. 00:06:04 - Ofir's Lab: SWE-ficiency, AlgoTune, and SciCode for Scientific Computing
  8. 00:07:52 - The Benchmark Landscape: TAU-bench, Terminal-bench, and User Simulation
  9. 00:09:20 - The Impossible Task Debate: Refusals, Ambiguity, and Benchmark Integrity
  10. 00:12:32 - The Future of Code Evals: Long Autonomy vs Human-AI Collaboration
  11. 00:14:37 - Call to Action: User Interaction Data and Codebase Understanding Research

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

C-Benz Progress and Updates

0:45 to 2:00

Discussion on the development and updates of C-Benz over the past year.

“And I think after that, it kind of kicked off the arms.”

SweetBench and Its Variants

2:00 to 4:00

Exploring the different versions of SweetBench and their impact on coding benchmarks.

“Like, JavaScript, Rust, Java, C, you know, Ruby.”

CodeClash Overview

4:00 to 6:00

An introduction to CodeClash and its approach to evaluating code development.

“And what that means is each model maintains their own code base.”

Tournament Mechanics in CodeClash

6:00 to 8:00

Discussion on how the CodeClash tournament system evaluates competing code bases.

“The other students have also been putting out a lot of other stuff.”

The Role of Programming Games

8:00 to 10:00

Exploring the use of programming games as arenas for CodeClash evaluations.

“Yeah, it's like a very new benchmark that Ophir did, and I think it's kind of related to physics.”

Emerging Trends in Coding Evals

10:00 to 12:00

Discussion on the evolution and future direction of coding evaluations.

“that we'll improve on these things over time for UBounce.”

User Simulator Systems and Future Directions

12:00 to 14:00

Exploring user simulator systems and their implications for future coding benchmarks.

“I think the vision of like, hey, I tell it a goal.”

Exploring Levels of Abstraction in AI

14:00 to 14:32

Learn about different levels of abstraction in AI tasks and the role of frameworks like Windsurf.

“like just enabling different levels of abstraction and, you know, it depends on the task.”

The Challenge of User Interaction Data

14:32 to 15:19

Understand the complexities of obtaining meaningful user interaction data for AI models.

“And how can people, I guess, like find more of your work?”

Combining Human and AI Efforts in Coding

15:19 to 16:06

Discover the potential of human-AI collaboration in coding environments and its implications.

“or between the two, like what's the best way to scale up sort of evaluating human AI interaction.”
Show all 12 chapters

Advancements in Code-Based Understanding

16:06 to 16:45

Examine how new approaches to code-based understanding can enhance human capabilities.

“where you can do a lot of different combinations of human AI on different arenas, playing one arena at a time, N arenas at a time.”

Benchmarking Code Understanding

16:45 to 17:31

Explore the difficulties in establishing benchmarks for understanding code and AI capabilities.

“So that is like sort of like a research subagent that we're working on.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:12We're here at NeurIps with John Yang of SweeBench and many other things. Welcome. Thanks so much for having me. Yeah, really happy to be here. Last year I talked to Othir and I think Carlos as well, one of your co-authors. How's C-Benz doing? Just generally, the project is like one and a half years old? Yeah, yeah. I think one and a half years old in terms of when it was actually useful. We put it out October 2023 and then people didn't really touch it too much. And then of course, Cognition came on the scene and Devon was an amazing release. And I think after that, it kind of kicked off the arms.

0:48Did they tell you beforehand? Or they just showed up? You know, I got an email about like two weeks ago. I think it was from Walden. It was like, hey, you know, we have a good number on it. I was like, wow, congrats. You know, thanks for using it. And then the release was like mind-blowing. I was like, wow, these guys did an excellent job. Amazing. And then Sweetbench Verified was like maybe last year. That's right. Catch us up this year. Like you have other languages. There's like a whole bunch of varieties of SweetBench now. Yeah. So what should people know? Yeah, for sure. I think there's a couple extensions that are happening.

1:23One is like more SweetBenches, SweetBench Pro, SweetBench Live. Oh, SweetBench Pro. Was that with you guys? Because it looks independent. It's like different authors. It's completely independent. Yeah. So they just call themselves SweetBench Pro without your blessing? I think we're okay with it. When we came out, we were like, oh, cool. Interesting. It would have been fun to be part of it. But, you know, I mean, congrats to them. It's a great benchmark. All right. But, yeah, multimodal. Yeah, we did multimodal and multilingual. And I think, like, those have multilingual seems to be. Is it, like, JavaScript?

1:56What else? Yeah, yeah. Yeah, multilingual. It's, like, nine languages across, like, 40 repos. But, yeah, you got them. Like, JavaScript, Rust, Java, C, you know, Ruby. Yeah, yeah, you got them. Yeah. And then, of course, you mentioned itself. A lot of people, like, they talk about the Django focus. Yes. Is there like, I don't know, how do we move past Janko? Yeah, for sure. I mean, it's cool to see a lot of the newer benchmarks really try to diversify the repos. In the two follow-ups we did with multimodal and multilingual, we made it a point to do that. But you can also just put out Cbench 2025 and just...

2:33That is true. And do a new distribution. Yeah, yeah. So it's been cool to see the follow-ups. I think quietly, and it's an open question for me, I'm excited to see how people curate the next sets like it's kind of interesting to see in the literature or in their blog posts like how they're justifying why they're creating their separate split the easier ones were like oh more languages more repos and then I think now people are like well ours is more difficult because of this curation technique and I'm yeah I'm excited to see how how long that lasts and you know where we're going to like guide the evaluations towards Yeah.

3:08And more recently, you're working on Code Crash. Yes, that's right. So let's get people, you've already done other podcasts about it. Yeah. I'll refer people to that with your chat with Andy. But just give people a one, two sentence. Yeah, no, happy to do it, especially on your podcast. It's an honor. Yeah, so basically the idea is I don't like unit tests as a form of verification. And I also think there's an issue with Sweetbench where all of the task instances are independent of each other. So the moment you have the model kind of submit it, it's done, you know, and that's the end of the story, end of the episode, you know.

3:42So with CodeClash, what we're thinking is let's try to really evaluate like long horizon development and development on a code base that is consequential and conditioned upon what a model did, you know, before to that code base. And so the general idea is you have two or more language models and they play a programming tournament. And what that means is each model maintains their own code base. And each round of the tournament, first they get to edit and improve their code base however they see fit. Very self-determined. And then in the competition phase, those two code bases are pitted against each other.

4:21So the code bases are run and there's generally an arena. We have a lot of diverse arenas, but the arena is determined like code base A is better than code base B. And then you kind of repeat that across multiple. As determined by an element judge. Yeah, yeah. So element judge is definitely one of the mechanisms. We started with some pretty simple programming games. So one of the cooler ones is like Halite, which Michael... Oh, yeah. I played it for Jane Street. Yes, that's right. That's right. You know, that's awesome. Yeah. Halite 1, 2, 3. Like Michael Troll of Cursor wrote this game. Two Sigma and Jane Street.

4:55Yeah. Oh, Two Sigma. Two Sigma. I worked at Two Sigma. I'm like... Oh, there you go. This is too long ago. There you go. Yeah. 2016 at this point, but we're bringing it back, you know. Hellen is fun. I would say if you've never done a programmatic competition where you have to control fleets of ships and attack things and defend things and collect resources, yeah. It's like play StarCraft, but you can code. Yeah, exactly. Exactly. Yeah, yeah. A lot of games. Yeah. Are there non-games or are you focused on games? I think that's an excellent point. So for kind of the initial release, for scientific purposes, we kind of use existing programming games.

5:32The current ongoing effort is, you know, to build economically valuable arenas. That's, you know, the popular word these days. Yeah, Sweet Lancer is a big one this year. Yeah, GDP Val is awesome. Yeah, just I mean, I think the big selling point of Terminal Bench and Sweet Bench and these evals is that it was really close to real-world utility, and so I think it's resolvable for Code Clash and that's what we're working on. Yeah. So you're part of Ophir's group. Yes. The other students have also been putting out a lot of other stuff. What would you highlight? Yeah, no, I mean, Ophir is such a prolific mentor when it comes to benchmarking.

6:09Sweefficiency, I really like in the line of performance. What's the deal you're on that one? Yeah, for sure. So Sweefficiency was wrote by this PhD student called Jeffrey Ma, who happened to be my high school classmate. And the idea there was like, you take a code base and you just want to, you know, do modifications that will literally make the code run faster. So I think it's like paralyzation, SIMD operation, stuff like that. So no behavior change, just faster. Exactly. Keep the unit test passing, but I want better runtime. Yeah, yeah. And then there's Algotune that is kind of in line with that.

6:44And then there's also kind of pushing along like the scientific coding domain. Psycode. Yeah, exactly. Psycode is awesome. They did like a quick one. And for people, Psycode is, the way I explain Psycode is, it's human eval, but better. Yes, exactly. Exactly. I think, you know, there's a lot of good stuff that these days where, yeah, that's the way to go. Which is like, Sweetbench is expensive to run. Any agentic benchmark is expensive to run. Actually, you do need some completions benchmarks. Yeah, just complete. Exactly. Like, you know, you can do well on those first and then sort of graduate to the multi-turn expensive stuff.

7:19Yeah, yeah. Okay. Other than that, just like broadly, other work in the field in 2025 in terms of coding evals, obviously we shot up Meter they use Sweebench and they have a very interesting like I guess human hours worked number yeah they like the x-axis being sort of the run time or yeah y-axis being the completion you know like we can do more long running stage and tasks I think the projections are quite interesting and I definitely appreciate them kind of using Sweebench verified to sort of proxy a lot of these things but yeah they're great okay any other work that like caught your eye Yeah, I mean, I think within the, okay, Terminal Bench, Sweet Bench, yeah, Critical Point was kind of cool.

8:01Critical Point? Yeah, it's like a very new benchmark that Ophir did, and I think it's kind of related to physics. There's this one called SecBench, kind of related to cybersecurity. Security. Yeah, exactly. SREBench, which I think is affiliated with LOD. Like, it's just cool to kind of see people really dive into different coding domains. and then stepping a little bit outside of coding. I personally think it's quite interesting to think about the user simulator stuff. So like TauBand. Vending Bench. TauBand too. Yeah, and Vending Bench. I got the big feelings. Yeah, no, I'm interested. Well, I mean, it's like you're sampling one path.

8:37I don't know how realistic it is, to be honest. It's just the hell of this, but it is cool. No, for sure. Yeah, I agree. I think it's a good initial effort. To me, I think it's super cool to see companies like, you know, I'm sure Mercor and stuff for focusing on building environments like for code, beyond code. And so I think it might be interesting to have like work gym style stuff. This is stuff that my advisor, Dee Young at Stanford, thinks about a lot. So yeah. Yeah. I just realized we're talking about terminal bend. Yes. We're friendly on the offer, folks. Yeah, yeah. You know, really, really, really good work just overall.

9:11Yeah. Let's talk about Tao Bench because you mentioned Tao Bench. Yes, yes. There's some discussion or some people are saying that Taubench is impossible to get a high score on because some of the tasks are underspecified or just impossible. Yeah. I don't know if you're up to speed on that. I'm a little bit... It's a little spicy. Yeah, it's a bit spicy. I think I saw... So I, you know, like I worked with Shunyu and Karthik back in Princeton very closely. I think Karthik I just saw posted a tweet kind of... Defending it? Yeah, like rebutting some of these claims. Yeah, I mean, it's... I think I get the concern.

9:49But yeah, I think it also brings up just maybe interesting research problems to solve of like, okay, why is it impossible? Is it the ambiguity? Is it kind of the user simulator that has issues? And I think generally we all agree that we'll improve on these things over time for UBounce. So I actually really like benchmarks that intentionally, I think we should intentionally include impossible tasks as a flag of like, hey, you're cheating. Yes. It's kind of sad that like Carpig actually is defending it because the master move would be like, oh yeah, you caught us. That was, you know, everyone reporting above 75 in Taub bench retail, you've been cheating.

10:24Yeah, oh, interesting. That would be cool, yeah. I mean, yeah, you'll have to ask the Taub bench authors, but yeah, no, that's fun. Yeah, I think there was Impossible Bench was a recent benchmark, maybe from, was it from Anthropic? I don't know, but they basically took Sweet Bench Verified and they changed the issues to make them impossible and they checked like how often the models would be like, I actually just can't do this. I don't know what's going on. Oh, like for refusals. Yes, yes, yes. Oh, how did they do? I thought that was interesting. I think they're all, the models are all kind of attempting and saying like, oh, I did it, you know, so maybe not great.

10:59That's cool. But no, that's an important one. Yeah. How does Kodi Evals evolve next year? Wow, that's a great question. I mean, honestly, I think it's, people will make more suite benches. I think terminal bench has really got something going where you ask people to, you know, a sweet bench, you're confined in some sense to the domain of issues and PRs that already exist, which I think has its benefits of being close to reality and natural. But I think with Terminal Bench, there's a lot of creativity that you can infuse into that. So I would personally be really excited. Like the 2.0 job was really excellent.

11:33And I'd be super excited to see, you know, 3.0, 4.9. Because of like the environments? Yeah, I mean, the environments, you know, bringing more people into the fold, you know, I think, correct me if I'm wrong, Mike, but early on you had PhD students, very smart CS people who are adding tasks. And, you know, what does that look like when you fold more coding environments for non-coding tasks, non-coding environments in general, and ask people to make stuff there? So that's pretty cool. And then, of course, for myself, I think just like this long-running SWE agent kind of thing just feels very compelling.

12:04I think the vision of like, hey, I tell it a goal. I don't have to be super specific about my tasks. I have like a decent verifier that proxies what I want. Something literally like a code base that makes the most money in this like setting, you know, like that's my verifier, you know. And I walk away for five hours. The thing is just running. I'm hanging out with you, talking to my friends. I come back and it gives me like literally a soda code base on that, you know, task. I think that would be super cool. Okay, I'll push back. We're part-time in Cognition. Yes. And we are emphasizing a lot of interactivity.

12:40Because the point is that you're going to underspecify. Right. And actually, what people want is back and forth, back and forth, and on a really fast time frame, which is terrible for a benchmark author. Right? Because how do you do that? Yeah. But realistic. Yeah. So I think this is where I'm a little bit anxious or cautious about this push for long autonomy. Right. I mean, let's say this time next year, we'll have five hours is pessimistic. Like, it'll be 24. Yeah, right. Days. But I don't know if that actually materially changes the industry. So we will push it, like, as an evals, you know, we have the people who make evals here.

13:22Yeah. We push the industry in ways that we wanted to push. But I don't know if we, like, that's a productive way. Because that's more of, like, a stunt that, like, yeah, it's a proof of concept that, existence proof, it can be done. Yeah. But will you use it for real life? Yeah, yeah. I mean, honestly, to me, I think there's potentially room for growth. So I would actually agree with your take here. I mean, with my lab at Stanford, with Dee, like there's a, you know, her emphasis is on human AI collaboration. And so I definitely don't believe in this idea of just kind of getting rid of the human.

13:57But yeah, maybe just like finding the balance of like, you know, just because the developer ecosystem is so diverse and there's so many participants in it who want different things out of it. like just enabling different levels of abstraction and, you know, it depends on the task. Like there's settings where you want to be, you know, more involved and more sort of hands-on and so you want to use Windsurf for that. But then maybe there's kind of this general data processing thing. It's just a lot of JSON parsing you don't really care about and that's the one I kind of want to walk away from and just let it figure it out.

14:28So, yeah, I would agree with you generally. Amazing. Any calls to action? What do you want help on? And how can people, I guess, like find more of your work? Definitely. For the call to action, super jealous of all the great data that Cognition and, you know, Cursor would get. Like that user interaction data is like really fascinating. From an academic standpoint, it feels like there's two difficult approaches to resolving that. Either you build like a really compelling product like Elmarina that people have people use consistently, which is, I mean, really tricky in and of itself. or you build like really good user simulators that try to mimic sort of these settings.

15:06But that is also like non-trivial. I don't think it's as simple as, hey, ChatGPT, act like a human, right? So it would be really cool to sort of get inspiration of like what exactly does that data look like or between the two, like what's the best way to scale up sort of evaluating human AI interaction. And then I think for visibility for my own work, pushing more arenas, like I think for, for CodeClash, what I'm excited about is the current framing is really long running sweet agents. But, you know, you could have multi agents, like two agents work together on the code base. And what happens?

15:41You have a human and an agent work on the code base versus just AIs. What happens there? You know, like when the models improve, and hopefully they hill climb, and they become better at digesting laws and iterating on analysis, you know, how does, how does human AI interaction like change with model capability? And so I'm kind of hoping, I'm trying to inspire and convince people that it's a very cool testbed where you can do a lot of different combinations of human AI on different arenas, playing one arena at a time, N arenas at a time. Yeah, I think very interested to work with you on the interaction stuff.

16:19That would be awesome. And then I think one more thing I'll add is for cognition, it's going to be pushing a lot of code-based understanding, which is kind of code-based retrieval plus plus. Yes. And mostly it is helping humans understand their own co-bases better to enable humans or to sort of mind meld the human with the machine to do the highest possible task that LLMs could not do alone, humans couldn't do alone. And then the other thing is also like basically automatic context engineering for an LLM. So that is like sort of like a research subagent that we're working on. That's so awesome.

16:55So I don't know what the benchmark would be because how do you benchmark understanding? That is true. Apart from, I think it's mostly like you freeze a repo, have some manually curated answers, and then post trivia questions. That's very easy to saturate, so I don't know. I think Silas tweeted a while ago sort of like the wiki, the code wiki. That's incredible. I mean, I use it on a daily basis. Google actually just came out with their own version. Oh, yeah, with the anti-gravity people. No, no, no. This is like a separate. Gotcha, gotcha. But cool. That's the state of code. Yep.

From the publisher

From creating SWE-bench in a Princeton basement to shipping CodeClash, SWE-bench Multimodal, and SWE-bench Multilingual, John Yang has spent the last year and a half watching his benchmark become the de facto standard for evaluating AI coding agents—trusted by Cognition (Devin), OpenAI, Anthropic, and every major lab racing to solve software engineering at scale. We caught up with John live at NeurIPS 2025 to dig into the state of code evals heading into 2026: why SWE-bench went from ignored (October 2023) to the industry standard after Devin's launch (and how Walden emailed him two weeks before the big reveal), how the benchmark evolved from Django-heavy to nine languages across 40 repos (JavaScript, Rust, Java, C, Ruby), why unit tests as verification are limiting and long-running agent tournaments might be the future (CodeClash: agents maintain codebases, compete in arenas, and iterate over multiple rounds), the proliferation of SWE-bench variants (SWE-bench Pro, SWE-bench Live, SWE-Efficiency, AlgoTune, SciCode) and how benchmark authors are now justifying their splits with curation techniques instead of just "more repos," why Tau-bench's "impossible tasks" controversy is actually a feature not a bug (intentionally including impossible tasks flags cheating), the tension between long autonomy (5-hour runs) vs. interactivity (Cognition's emphasis on fast back-and-forth), how Terminal-bench unlocked creativity by letting PhD students and non-coders design environments beyond GitHub issues and PRs, the academic data problem (companies like Cognition and Cursor have rich user interaction data, academics need user simulators or compelling products like LMArena to get similar signal), and his vision for CodeClash as a testbed for human-AI collaboration—freeze model capability, vary the collaboration setup (solo agent, multi-agent, human+agent), and measure how interaction patterns change as models climb the ladder from code completion to full codebase reasoning.

We discuss:

John's path: Princeton → SWE-bench (October 2023) → Stanford PhD with Diyi Yang and the Iris Group, focusing on code evals, human-AI collaboration, and long-running agent benchmarks

The SWE-bench origin story: released October 2023, mostly ignored until Cognition's Devin launch kicked off the arms race (Walden emailed John two weeks before: "we have a good number")

SWE-bench Verified: the curated, high-quality split that became the standard for serious evals

SWE-bench Multimodal and Multilingual: nine languages (JavaScript, Rust, Java, C, Ruby) across 40 repos, moving beyond the Django-heavy original distribution

The SWE-bench Pro controversy: independent authors used the "SWE-bench" name without John's blessing, but he's okay with it ("congrats to them, it's a great benchmark")

CodeClash: John's new benchmark for long-horizon development—agents maintain their own codebases, edit and improve them each round, then compete in arenas (programming games like Halite, economic tasks like GDP optimization)

SWE-Efficiency (Jeffrey Maugh, John's high school classmate): optimize code for speed without changing behavior (parallelization, SIMD operations)

AlgoTune, SciCode, Terminal-bench, Tau-bench, SecBench, SRE-bench: the Cambrian explosion of code evals, each diving into different domains (security, SRE, science, user simulation)

The Tau-bench "impossible tasks" debate: some tasks are underspecified or impossible, but John thinks that's actually a feature (flags cheating if you score above 75%)

Cognition's research focus: codebase understanding (retrieval++), helping humans understand their own codebases, and automatic context engineering for LLMs (research sub-agents)

The vision: CodeClash as a testbed for human-AI collaboration—vary the setup (solo agent, multi-agent, human+agent), freeze model capability, and measure how interaction changes as models improve

—

John Yang

SWE-bench: https://www.swebench.com

X: https://x.com/jyangballin

Chapters

00:00:00 Introduction: John Yang on SWE-bench and Code Evaluations
00:00:31 SWE-bench Origins and Devon's Impact on the Coding Agent Arms Race
00:01:09 SWE-bench Ecosystem: Verified, Pro, Multimodal, and Multilingual Variants
00:02:17 Moving Beyond Django: Diversifying Code Evaluation Repositories
00:03:08 Code Clash: Long-Horizon Development Through Programming Tournaments
00:04:41 From Halite to Economic Value: Designing Competitive Coding Arenas
00:06:04 Ofir's Lab: SWE-ficiency, AlgoTune, and SciCode for Scientific Computing
00:07:52 The Benchmark Landscape: TAU-bench, Terminal-bench, and User Simulation
00:09:20 The Impossible Task Debate: Refusals, Ambiguity, and Benchmark Integrity
00:12:32 The Future of Code Evals: Long Autonomy vs Human-AI Collaboration
00:14:37 Call to Action: User Interaction Data and Codebase Understanding Research

More from Latent Space: The AI Engineer Podcast

All 247 episodes
[State of Code Evals] After SWE-bench, Code Clash & SOTA Coding Benchmarks recap — John YangLatent Space: The AI Engineer Podcast
Listen in VO