In short
Misaligned incentives in AI coding agents, shifting bottlenecks from training to evaluation, and how Cognition (Devon) scales proactive agents with better cost/quality via model routing and specialized post-training.
Guest
Russell Kaplan, president at Cognition; previously started ML career at Tesla on the Autopilot team (2017). Background includes agent deployment, eval design, and enterprise software engineering workflows.
Key claims
Evals (not just bigger models) are the new bottleneck; proactive agents give human engineers “outsized leverage” by delegating work and running proactive automations (e.g., Slack triage). “No such thing as the best model anymore”—use the right model per task and optimize speed/cost once capabilities saturate. Token spend can approach/exceed human salary spend, so harnesses must be “cash-aware” (DevInFusion improves price performance ~35% with slight quality increase).
Notable examples
Devon launched March 2024 (viral demo), became top committer June 2024, production deployment later; early PMF in late-2024 migrations/refactors. Frontier Code eval measures “mergeability” (would you actually merge), with binary blocking + weighted non-blocking scores; built with open-source maintainers. Security Swarm: agentic map-reduce for vulnerability remediation, using isolated micro-VM sessions to reproduce and validate fixes.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe Evolution of AI Coding Agents
1:12 to 3:33
Discussion about the development and capabilities of AI coding agents like Devon.
“Welcome to Max Agency, the podcast that goes deep into how the best agents are being built by builders like you.”
Scaling AI Agents in Software Development
3:33 to 5:08
Exploration of the costs and benefits of running AI agents at scale.
“There's a few like technical reasons this made sense.”
Improving AI Model Performance
5:08 to 7:04
Insights into enhancements in AI models and their impact on software engineering tasks.
“Yeah, maybe diving into that a little bit.”
Frontier Code: New Evaluation Metrics
7:04 to 9:51
Discussion on the creation and significance of the Frontier Code evaluation for coding agents.
“What are tasks where you think Fable-type models are, like, where you see that jump today?”
Evaluating Code Changes in AI
14:00 to 15:55
Discussing the complexities of evaluating code changes in AI systems and the importance of both blocking and non-blocking criteria.
“So it's, it is like, this is not like a quick thing.”
Managing AI Costs in Development
15:55 to 18:51
Exploring the challenges of managing AI model costs and the implications for developers and organizations.
“you know, I think we did it right enough that all of the kernel LMS are pretty bad at this eval.”
Optimizing AI Model Performance
18:51 to 22:21
Discussing strategies for improving AI model performance while balancing cost and speed.
“We have very high, you know, compute investment in general.”
Sidekick Agents in AI
22:21 to 23:11
Introducing the concept of sidekick agents that work alongside primary models for enhanced task execution.
“So talking about that like a little bit more, because there's routing, and then there's also Open Router launched, Open Router Fusion, which is not routing.”
Developing Specialized AI Models
23:11 to 26:51
Understanding the reasons and benefits for developing specialized AI models for specific tasks.
“And we try to share a lot of context, too.”
Future of AI Inference and Model Architecture
26:51 to 28:00
Discussing the future possibilities of AI inference and advancements in model architecture and chip design.
“So we don't need to manage like 50 or 100 different of these super specialized models.”
Show all 19 chapters
Advancements in AI and Specialized Inference
28:00 to 29:57
Explore how advancements in chip design and inference specialization enhance AI performance.
“we want to constantly be basically pushing the frontier of well what's the next most possible thing that wasn't possible previously.”
Integrating AI with Organizational Context
29:58 to 31:33
Learn how AI agents can proactively manage engineering tasks and improve workflows.
“All this infrastructure stuff at the chip layer, I'm presuming that's part of it.”
The Evolution of Coding Agents and UX
31:34 to 33:36
Understand how coding agents are changing software development and user experience.
“I'd love to hear how you guys think about UX because you started off, I think, and made kind of like the Slack coding experience.”
The Future of Engineering with AI
33:37 to 36:29
Discover how AI is reshaping the skills required for engineers and the emergence of new roles.
“find all the different issues and then remediate them en masse?”
Building Effective Coding Solutions with AI
36:30 to 39:54
Learn how to effectively build coding solutions and the role of AI in solving customer problems.
“every three months, you have to continually retest that because these systems are getting so much better.”
The Role of Forward Deployed Engineers
39:55 to 43:24
Explore the importance of forward deployed engineers in leveraging AI for impactful solutions.
“And it's pretty hard to predict in advance who's going to have the best model, who's going to have the best price performance model and so on.”
Challenges in Quantifying ROI for Coding Agents
43:25 to 44:54
Discover the difficulties and strategies in measuring the return on investment for coding agents.
“So first on the productivity side, what's the problem?”
The Productivity Guarantee: Aligning Incentives
44:55 to 47:11
Explore how Devon's productivity guarantee aims to align customer and provider incentives.
“or an individual developer using Devon, how can everyone benefit from that?”
Using Insights for Professional Development
47:12 to 49:46
Learn how teams and individuals can utilize insights from Devon to enhance their productivity and professional growth.
“And when we debated this internally, I got a lot of pushback actually, because some people were saying, wait, this is like a big risk.”
Transcript
Automatic transcript. May contain errors.0:00I started my own machine learning career at Tesla on the autopilot team. I talked to my friends at Tesla today, and the bottleneck is no longer just training bigger and bigger models.
0:07Russell Kaplan:It's actually running the evals. Today, I'm talking to Russell Kaplan, president at Cognition, the company behind Devon, an agent that went from viral demo to deploying code inside some of the most complex orgs in the world. We essentially have an evaluator agent that can take a session and say, was it productive or not? It gave us the confidence to actually go to our customers and say, we are actually going to make a$10 million productivity guarantee. He explains why and how proactive agents are giving human engineers outsized leverage. You have like all these great suggestions of fixes that need to be applied.
0:36Oh yeah, that looks good. I want you to change that here. Individual developers have essentially realized I can be the CTO of an army of 10 ,000 agents. We get into what running agents at scale actually costs and how cognition drives it down. There's organizations where the per person token spend is starting to eclipse the human salary spend. By being a little bit more clever about some of the routing, we can get about 35 % better price performance with actually a slight increase in quality. And he argues there's no such thing as the best model anymore. The Fable Class models are really good for high precision, but we actually find that GPT 5.5 and 5.5 Cyber are better on recall.
1:11You have to use both.
1:12Russell Kaplan:Welcome to Max Agency, the podcast that goes deep into how the best agents are being built by builders like you. you guys had a massive launch about two and a half years ago three years ago at this point and i think you pioneered a lot of really interesting concepts especially around ux and interacting with devin in slack i think that was the first time i saw that so prominently so i remember a lot from that launch i remember a lot about the ux i'm sure others do as well what should people know about what you guys have been up to over the past two and a half years yeah we launched devin in march of 2024 was the original demo video that went super viral.
1:50And at the time, I would say coding agents were just at the edge of possible. You kind of squint and see, OK, this is going to work at some point. And we sort of put together a first pass example of what could that look like.
2:02Russell Kaplan:You know, we bench was like 13 percent. I think it was like tripled the previous. Totally. It was like, yeah, we were like us. 13 percent sweet bench. This is, you know, know really exciting and it actually that by the way that that number reminds me a lot of some of our more recent evals and where we are now in agentic coding with it with a new eval we put out called frontier code which we can talk about but when we you know we kind of launched in march of 24 we had this point of view that eventually you're just going to have agents as teammates that you can delegate complete units of work to and what we learned is that took us a few months from okay this is a prototype that we can kind of see the future of to this is something that's actually useful internally.
2:42I think it was June of 24 that Devon became the number one committer to Devon, which was like our first big milestone. That's still a while ago. It was a while ago, but it was like a lot of manual dogfooding and like really kind of grinding to make the internal workflows nice. And then it took us a few more months to actually get this deployed in production and useful at a customer. And in the sort of 2024 era async cloud coding agents, they really couldn't do most of the tasks in software engineering, but there were some, you know, like one of the best early ones was migrations and refactors.
3:15If you ran essentially a, you think of like a Regex++ style workflow across a large code base with sort of bits of intelligence sprinkled in, that actually already worked reasonably well in late 2024. And so that's where we got our, you know, early product market fit. It was with bigger companies, you know, like more like enterprise-y companies who they just had lots of code that needed all this transformation. And why was that such a good fit? There's a few like technical reasons this made sense. One is for a very large refactor or migration or, you know, ETL transformation, whatever, it's sort of worth it to put in the effort to carefully prompt engineer your cloud agents to be accurate.
3:52And so, you know, if you kind of iterate on your prompt and your setup and your, you know, the context you're feeding in and you get it like just right and you can apply it across, you know, 10 ,000 modules in your code base, this is actually really high ROI. And it's obviously much better than just sort of a find and replace, you know, string change. But it didn't have to have this like full general software intelligence that we're now delegating all sorts of coding tasks to. So that worked well even in the early days, but it was still kind of niche. And then in December of 24, we launched Devon so anyone could sign up.
4:24And we had this big debate internally of what should be, you know, what should we really emphasize or what should be the focus. And I remember we did like that. We just we flew the whole company to Utah to just like lock in in December and get this out, you know, kind of before the end of the year. And what we settled on was actually Slack as the primary interface. And so the entire launch video for Devon Available Self Service, we called it internally at Devon. It was just at Devon, at Devon, at Devon, and really trying to drive home this sort of user experience change, which is collaborating with agents more like teammates.
4:55And so from there, we got a lot more users, we got a lot more traction. You know, cloud agents have just gotten better and better since then, both as the infrastructure has gotten better, as we've matured on that, as the models have gotten better, as the harnesses have gotten better. And now I think most of the work is being done by just delegating to async agents.
5:13Russell Kaplan:Yeah, maybe diving into that a little bit. What are some of the things that have gotten better that have allowed them to really take off? So a few things. I think people always talk about the models getting better, and that's super important. But I think the kind of infrastructure maturity is always... What was the first model that you felt was like really actually good enough? Oh, yeah. It's, you know, we kept saying, oh, this is the model. This is the model. You know, this is the new model. There were a few like phase changes, I would say. I think sort of like the November 2024, like model series was one kind of...
5:43That was like a big step function where I was like, okay, this actually can work a lot better. Now we did a self-serve launch, you know, in part based on that. from OpenAI. And then, you know, the anthropic model started getting really good. And we saw, I think, probably like, you know, July of 25, another big step change. I think now we see, you know, with like Fable 5, it's like new generation of models, continuous step changes. And now the step changes are so high, there's actually no longer about the capabilities increases. I think we're kind of seeing the opposite trend in some sense, which more and more of the tasks in software engineering are getting intelligence saturated.
6:13So like literally all of the models can do them well. And so the sensitivity of a lot of developers has totally shifted from, wait, I need to be using the best possible model to holy cow, I'm spending so much money on my coding agents. How can I be a little bit more efficient about this? And I think this is like an underappreciated aspect of continuous gains in frontier intelligence, which is that for some workloads, you always want the most frontier intelligence, especially if it's like an adversarial game between parties and your intelligence has to be higher than the counterparty's intelligence, like if you're trading in competition or something.
6:46But for a lot of software we want to build, you know, there's kind of this saturation threshold where once it's good enough, what you care about is speed and cost. And I think that's actually where we see a huge amount of demand from developers and the companies we work with is, okay, this already works. How can I now optimize the speed and cost?
7:05Russell Kaplan:On the model side, maybe one question. What are tasks where you think Fable-type models are, like, where you see that jump today? Like, where would you recommend people use them? Where do you see people using them? So I think we're seeing maybe in this most recent generation of models and Fable being one example, particularly big gains is in, first of all, there's like detection or mediation of cybersecurity vulnerabilities. And there's a lot of guardrails now in the public versions of these models. And so it's actually can be tricky to get them to help. We see, for example, that if you're trying to go just like clean up my security backlog, which, by the way, is a huge workload right now for a lot of the companies we work with, The Fable class models are really good for high precision, but we actually find that GPT 5.5 and 5.5 Cyber are better on recall.
7:54And so the best possible harness for finding and remedying security vulnerabilities right now, you have to use both. And you sort of filter down the ones that are maybe better found by the GPT series with the Fable series. I think that's one big example. Another one is that processing data from integrations, for example, from like Datadog or other systems, we do see like a step change in our evals in Fable in particular.
8:21Russell Kaplan:Is that because that data is like so large and messy that it needs more intelligence? I think a lot of, yeah, it's just like, there's like a volume consideration. There's sort of a multi-step reasoning consideration. And I think there's just, you can tell that there's a lot of RL going on on like these very realistic data sources that have been scaled up a lot. Maybe the third category that's maybe the most noticeable is a category that we've only even gotten better at detecting and understanding recently ourselves with the work we've been doing on new evals for coding. Because to your point, when we launched Devon and it was, you know, in the teens on SweetBench, now SweetBench is totally saturated.
8:54And so how do we find the next level of difficulty for evaluating models? We really, we looked for all of the evals and we couldn't find any. That could actually kind of match our sort of hazy internal intuitions of this model feels a lot better. And when we sort of dug into it and really asked ourselves, why is that? I think the core gap for evals that we found was around mergeability, which is, okay, this code is technically correct, but like, would you actually merge it? Like, would you feel happy? Would this improve the quality of your code base? And there's lots of different subtle examples of what that can mean.
9:26It could be stylistically following the expectations that you already have. It could mean that it's done in a way that it's easy to modify in the future or that if other modifications happen, it gracefully handles those modifications. all these little stylistic things, which we felt like were not being captured in the evals. And so we set out to basically make a new eval to measure this better and also just have a new higher watermark of difficulty for frontier coding capabilities. And then we released it recently. It's called Frontier Code. And it's once again, it's the new hardest coding eval.
9:56The Frontier Code Diamond subset is like in the teens, you know, pre-Fable 5. And now I think Fable 5 has pushed it to the 30s. So we have, I wonder how many more months we have before that eval is.
10:06Russell Kaplan:I was going to say, yeah, in like a year we'll be saying, oh, remember when it first got in the teens for this one? Yeah, I'm wondering, I do wonder, like it's going to get harder, you know, it gets harder and harder to make these evals. I mean, I started my own machine learning career at Tesla on the autopilot team. And this was in like 2017. And the bottleneck was very clearly, you know, GPUs and compute. And it was like, okay, how do we get more compute? We just need more compute. And I talked to my friends at Tesla today and the bottleneck is actually no longer just training bigger and bigger models.
10:34It's actually running the evals because for a self-driving system, the interventions are so rare that you have to do enormous volume of driving to even find any problem at all in the stack. And, you know, I think for for software engineering, we're not there yet. Like we still find bugs, but but we might we might hit that threshold sooner than we think.
10:53Russell Kaplan:We talk with a bunch of teams around creating evals for their own systems using Langsmith and other things like that. How did you guys create the frontier code? Like what did that process look like? So it started with, okay, we need to go find the sort of peak, high-taste developers who are going to have really strong, both kind of stylistic quality, but also code correctness standards to help us answer the question, would you actually merge this code? Not just, is this code passing tests? Is it functionally correct? So that's kind of part one. And so we actually did like a, you know, a deep collaboration and recruiting campaign with a lot of the sort of leading open source authors who were, you know, really grateful for their collaboration on this.
11:32It wouldn't have been possible without them to go find, you know, what are the best known libraries in open source with really high coding standards and then collaborating closely with them to help encode their human intuition of what makes a PR something I would accept into these very carefully designed tests. That was part one. What do those tests look like? Are they LM as a judge? Are they programmatic assertions? Yeah, so multiple. So part of them are programmatic assertions. Of course, there's like a correctness test. There can be programmatic assertions on kind of stylistic elements too.
12:02Like, OK, if you make this change in this pull request, you know, you have to actually use this module, even though using this other module would be equivalent right now. Over time, if they drift, you know, if you didn't use the correct module, then like you're going to introduce a bug in the future. So it can be asserted deterministically, but it requires the judgment of, you know, the human author to settle. And then the other big thing, which I think is still underdone by people practicing machine learning and working with agents, is every single researcher on the cognition team, you know, hand contributed and reviewed and evaled the evals directly.
12:36And so this is not something you can sort of throw over the wall and say, OK, the data quality piece, that's such a slog. Someone else can do that. You sort of have to work on it directly, too.
12:46Russell Kaplan:I'm assuming they worked with the open source authors to put together these assertions and they would be running on these open source code bases, I guess, like that was kind of the setup? So it would be on the open source code base and we could run it in our infrastructure, but then we test and to avoid contamination, we haven't published the full set of questions. We publish some example questions, but we're trying to preserve this eval for the community for as many months as possible until it gets completely saturated. But essentially running the code that already exists with the patches applied by these models.
13:17Russell Kaplan:How do you guys score these evals? Is it binary, like zero one? Is there some numeric component? Like if you have 10 assertions on one of the tests and it passes, 9 out of 10, how is that scored? Yeah, so we broke it out into two components to have, first, there's a binary element, which is essentially, yeah, did this pass all of the blocking constraints of how we would evaluate this PR? You know, the simple things of like, do the test pass? Are there hard deterministic criteria that the maintainer of the code base has inputted must be met by the agent who's writing this code? Do you know off the top, like, is that like five assertions or like 50 assertions?
13:54Like it's highly varied per task, but each task is, you know, like hundreds of hours of work of people putting into. So it's, it is like, this is not like a quick thing. It's the process of constructing a single test, you know, eval line item in a data set like this. It's, it feels like a, it's like a full project and labor of love per question, you know, where you're really trying to think holistically, what are all the things that I, as the expert maintainer want to sort of imbue in my expectations so that's how you get some of the binary you know pass fail metrics but then we also to your point that's often not enough to to really know okay would i merge this code and then also like how would this code rank to someone else's code because there might be you know two different pieces of code you would merge but one is like a little bit more preferable than the other and so we also have the concept of a score which is essentially like a linearly weighted aggregation of all of the non-blocking evaluation criteria.
14:50So if it's blocking evaluation criteria, we'd say, okay, you know, if it doesn't pass, you fail. But like, or this is just, you know, we're not merging them. If it's non-blocking criteria, for example, there's stylistic elements, a common one for us is scope. So I think one of the code smells of LLM still is like they'll make the right change, but then they'll kind of mess with other files too that you really wish they didn't mess with. And so that's not going to impact your correctness score, but it's certainly going to impact the sort of stylistic elements and that's going to downweight you know if you had unnecessary touches in other files that's going to downweight your aggregated metric same there for you know lm as a judge having heuristics that you impose in lm as a judge can get added to this um to the to the linear score we also found that um reverse classical evaluation is is really helpful so commonly people say okay we need to make sure that after you accept this code change, the test pass, right?
15:40But we also care just as much that, you know, without this code change, the test should fail, right? And that if you make this other code change, the test should fail. So how do you impose kind of blocking constraints on both sides? But yeah, evals are really hard. You know, we spent a lot of time on this. I think we got it, you know, I think we did it right enough that all of the kernel LMS are pretty bad at this eval. And it also, more importantly to us, you know, as new models come out and we test them and we see their scores on Frontier Code, it's roughly matching the vibes. Like, like if the model is scoring really well on Frontier Code, then we get really excited.
16:16Russell Kaplan:I like this idea of kind of like binary plus fail and then some numerical after that for the or more explicit. I like the like blocking and non-blocking. So we're building a benchmark of our own. We call it like issue bench for Langsmith Engine, which goes through and finds issues. and there's a bunch of stuff that like I would classify under like the non-blocking stuff like it's yeah it would be nice if it did this it should probably do this and then there's some other stuff that's kind of like more blocking does it just find this like really bad issue should be more of a blocking thing so I like that kind of like dual juxtaposition do you guys use harbor as a format for running evals or do you have your own internal kind of like eval we do use harbor we think in general a lot of the kind of standards are quite early and immature and I I think one of the funny things, if you sort of look in the whole Devon code base is because it was the first coding agent, there's actually a whole bunch of stuff that like there might be a standard now or a correct way of doing things.
17:05And we were just, we just rolled our own, our old implementation, you know, as funny example, funny examples. So even a basic agent things, like the concept of, you know, skills.md file, like we had implemented a concept in Devon of knowledge that was like before any open source standard existed for skills. And now we've basically grafted on how can you also work with the open source standards. But it is like a recurring theme in our code base that we've gone off and sort of invented something weird. And then it becomes some permutation of it becomes an open source standard that we incorporate back in.
17:37Russell Kaplan:We went down this rabbit hole talking about models and talking about the best models. But you also mentioned kind of like cost and presumably kind of like speed and latency are becoming other issues. You guys have done a few things here, if I'm correct. You have DevInFusion. You also have your own series of models that you guys have post-trained. How do you guys think about this section of the model universe? We're kind of in this unique time in history where anyone can hire as many AI agents as they want, usually, the way the tools work. So, you know, me as a developer, yeah, I'm not going to use the cheap model.
18:12I want to use the best one for everything. But everyone is kind of collectively making that decision. And then you sort of roll it all up and you realize, wow, we are spending a lot of money. You know, there's organizations where the per person token spend is very rapidly approaching or even starting to eclipse the, you know, the human salary spend. And so it would be kind of crazy, you know, if like the way you ran Langshane, for example, anyone could just hire a thousand people tomorrow without, you know, without talking to you. But that's kind of how we run our teams today.
18:40Russell Kaplan:And this is starting to become like a real issue for us. And like I like to think we're like we're pretty like, you know, we're still startup. We're a lax, like AI native kind of like forward organization. But we are absolutely caring about kind of like token spend. I'm like, are you guys internally kind of like worried about token spend for your own kind of like engineering schemes as well? We have very high, you know, compute investment in general. So I would think our internal spend, while being very high per person to us, it's like really valuable dogfooding investment in everything we do.
19:06But our customers are definitely thinking about if you just extrapolate the trend line, you know, it's going to like eclipse the whole economy and not that long with exponential growth. And so people are asking the question, okay, well, how should we approach this? How should we even think about this? And I think we now, there's sort of two different trends that are happening at the same time that sometimes get conflated. So one is obviously the models are getting a lot better and more and more capable. What's happening is, you know, if you look at the most frontier capabilities and then maybe the set of models that are just behind them or a little bit behind them, every time we have these new generations of models, we both move the high watermark on the frontier, but then the set of tasks that all the models can do is growing a lot.
19:44And the distribution of tasks that people are like trying to accomplish in their everyday lives are changing a bit, but actually not that much, you know? And so it's like, okay, I'm trying to build a front end, you know, for my marketing registration page.
Read the full transcript
19:56Russell Kaplan:It's like, that's not that hard of a tad, right? And so I think the way I think this plays out is as more and more of the models do more of these things, the marginal returns to investing and just like having a reasonably intelligent harness that can make sure you're using the right model for the right job in the right moment goes up a lot. And both for individual developers who, you know, they want a fast answer, they want a correct answer, and they don't want, you know, they don't want any necessary waste. Or even really like the cognitive overhead of having to decide every time you interact with a coding agent, oh, like what model should I use for this one?
20:28I was going to ask, do you let users of Devon choose what models they use? For Devon Desktop, which is what we rebranded Windsurf relatively recently, and for our CLI, we do. Developers like the individual control when they're working with local agents. for our cloud agent, we don't. We'll have like Devon Fusion, which we talked about, which is the sort of frontier performance with cost optimization option. We'll have like a more affordable agent. We have, you know, the like the max or ultra agent. So we have this like sense of kind of tiering, but I actually think it's like a UX bug in the fullness of time for people to have to think about this for their cloud agents.
21:05And, you know, we saw this early days, we ran some tests where we let people pick the models. And then two things happen where one is, Sometimes people would then give a task to Devin and it wouldn't work. And they would complain, oh, Devin was so dumb here. And we looked at it, well, why did you pick this model? But actually that's kind of more our fault, I think, than the user's fault. And then the same thing would happen where they would run tasks like, oh, this was like really expensive. Well, why did you use this very expensive thing? And so I think this sort of natural equilibrium is people want the best performance for the best price without compromising on correctness, right?
21:36And so folks who are building agents, I think have in some sense, like a user experience responsibility as well as, you know, just building good price-performing products to try to do as much of that optimization as possible. And that's what we recently put out with DevInFusion, which is the sort of next generation of our own harness designed for frontier-level capabilities, but with maximum price performance. And so what we found is, like, by being a little bit more clever about some of the routing and the decisions of, like, what you're doing with which model and thinking about this in a very cash-aware way, we can get about 35%, you know, better price performance with actually a slight increase in quality.
22:15And I think that's sort of where a lot of this needs to go for the technology to continue to adopt it at the rates being adopted.
22:22Russell Kaplan:So talking about that like a little bit more, because there's routing, and then there's also Open Router launched, Open Router Fusion, which is not routing. It's running it on multiple models in parallel and then combining things back together. So when you guys have Dev and Fusion, is it routing? Is it like running multiple things? Yeah, so it's doing both. So we put out a technical blog post that shares a little bit more detail on our implementation. But one core component of it is this idea of a sidekick agent. And so, you know, what we'll do is we'll have the sort of frontier quality model executing on the task.
22:53And then in parallel, we'll have a more price performance model executing on the task. And then there will be some decision making that the frontier quality model has to do of, OK, well, when do I delegate to my sidekick? But having them both work on the same task in parallel, it lets you make sure that there's still context for both of these agents. And we try to share a lot of context, too. And we do rights to the file system frequently if there's things that are exploding out of context. But having basically both work in parallel, then frontier model knowing, OK, I can pass this off. And actually, as the models get smarter, the most frontier models get smarter and smarter.
23:26They're one of the key skills we see is they're like way better at delegating, too. Like Fable is very good at knowing, ah, this task I can delegate to a dumber model and it's going to be okay. Which kind of makes sense if you think of human career progression also. Like one of the aspects of growing in your own career is learning how to delegate tasks and how to like do the highest leverage tasks yourself. So we kind of see the same technical pattern in the models.
23:50Russell Kaplan:You guys have your own set of models as well. SWE 1.6, I think is the most recent one. Why did you guys train that? What do you see people using it for? Yeah, this is a great question. So sometimes people ask us, you know, why even RL or post-train your own models at R? Like, aren't the next models from the Frontier Labs just going to be better and better and better? And definitely, and we get super excited every time new Frontier Lab models come out that are better at their tasks. But I think there's two important reasons for us to be spending a lot of energy on our own RL and post-training for our own models.
24:22So the first is that you actually can deliver Frontier capabilities at any given moment in time through greater specialization. And I'll give you a recent example, which is we shipped a product relatively recently called Devon Review, which you guys are using in really interesting ways inside Langchain. Yes. To deliver Devon Review, we wanted to not just have sort of a human interface for understanding diffs and kind of grokking large amounts of AI-generated code quickly, but also be able to run quick static analysis and lightweight kind of model-driven analysis on are there bugs in this code?
24:57Are there security vulnerabilities in this code? are there things that we should be sort of linting and automatically checking? We can use frontier models for that. We can only use cheaper models, but then you get really tough trade-offs on the price performance curve. You know, you don't want to have to spend tons of money just by virtue of having created a new PR. And that's the exact type of problem where very specialized RL and post-training can lead you to producing an extremely price performance model for a very specialized task. And we know that that model has a half-life that's going to expire.
25:25And that's fine, because by the time it expires, will be very happy and will be working on the next specialized model for the next workload. And so I think a lot of kind of building an AI startup right now is being very willing to think in these like three to six month increments of, okay, given the state of frontier capabilities today, what is the most differentiated set of new product experiences I can build through my own specialization? That's part one. The second reason we work on this really helpful is it does go back to price performance, which is as more and more models are capable of doing more and more things.
25:56How can we continue to deliver, you know, the best possible price performance for our customers? Well, it starts with, if any model can do it well, you know, we should be able to serve our own models even faster, even cheaper than anything else on the market. And we see that, you know, SWE 1.6 is actually the most popular model in dev and desktop by number of tokens consumed. It's about Opus 4.6 level. And we'll have more to say on new models coming out soon. There's going to continue to be this like frontier of capabilities that use the most frontier models for. And startups, not just us, I think more startups should be considering how do I make my own models that are specialized for my domain?
26:31Because more and more of the tasks in my domain are going to be doable by any model.
26:34Russell Kaplan:How many of these specialized models do you guys have at any given point in time? Is it one for review, one for coding, and those are separate ones or order of magnitude? It ranges. I would say it's like single digits number of specialized models because part of our own focus as a company is on software engineering. And so, you know, having like the SWE 1.6 series, that model series is something we can drop in and use at a lot of different places. So we don't need to manage like 50 or 100 different of these super specialized models. But for the areas of our product that we think is going to make the biggest impact, not just on performance or like kind of quality, it could be literally on like latency and speed and, you know, affordability.
27:11We want to make sure we continue to invest.
27:13Russell Kaplan:And you guys optimized a lot for latency with SWE 1.6, right? Yeah, one thing that was really fun working on in Sweden is we were the first kind of Western firm to deploy Cerebris at scale. You know, when we tested their chips, we were really confused because we were like, this is really good. You know, it operates a unique point on the Pareto curve of price and throughput. and you know we were getting about 950 tokens per second on our own models which was many x faster than what we could get on on gpus for the same size of model and it was a little bit more expensive for us to serve they were great partners and they enabled us to ship a product experience that basically wouldn't have been possible otherwise and now service is more popular you know now you know open i also use cerebris for for spark i think more people are considering them and but we want to constantly be basically pushing the frontier of well what's the next most possible thing that wasn't possible previously.
28:07And yeah, doing our own RL and post-training helps with that.
28:10Russell Kaplan:What do you think the next most possible thing that isn't possible right now will be? As far as like model serving, I mean, I'm really excited about, you know, the sort of continued generations of very, very specialized inference. As the architectures that we use for transformers have gotten a little bit more standardized, more predictable, you know, the people designing chips, you can make a lot more specific assumptions that enable really big, really big, meaningful technical gains. I'll give you one funny example from when I was at Tesla. So Tesla, we built our own chips in-house for inference on autopilot.
28:41And it was a big part of why autopilot could work as performantly as it did. And I remember we had, I was like a deep learning researcher. One of my jobs was like interfacing with the silicon design team to make sure that the next generations of silicon supported our needs. And, you know, a lot of the bottleneck in chip design, it's actually thermals. It's like, how hot can you run without melting the chip, that's the amount of power you can put in. And then even within the amount of power you put in, a lot of that power is consumed by memory movement as opposed to flops. And so we're trying to be really efficient of, okay, what are the things that we don't actually need on this chip that we can take away so we can really run at max throughput for the stuff we did care about?
29:17And we had a big debate one day on whether the chip actually needed to support division.
29:21Russell Kaplan:Because they were like, well, you know, if we don't have to support division, we can make this way faster. And if you look at that time, it was convolutional neural networks. the layers in a convolutional neural network at inference time, none of them actually needed division. You know, for normalization, we were using batch normalization, which has a division step. But, you know, you could actually fuse the batch norm weights and biases into the conv weights. So I think we're going to have crazy levels of specialization because there's going to be so much inference demand. And that's going to unlock just like super exciting levels of performance.
29:53Russell Kaplan:And so when we started off talking about all these things that kind of like made coding agents and devin like way better over the past two and a half years. All this infrastructure stuff at the chip layer, I'm presuming that's part of it. And there's probably a lot more stuff to go in that capacity. Every layer of the stack. We try to think about every single layer of the stack. And one of the layers I'm most excited about right now, it's how deeply can we integrate with the rest of the team and company, you know, organizational context and way of doing things. So I'll tell you like the biggest internal difference I feel using, you know, devin and cognition versus a few months ago is we've gone really hard on this concept of automations.
30:29So how can you wire up your agents so that they're proactive, not reactive? And, you know, in steady state, it's probably going to be the case that most of the engineering work at a company, it's done proactively, autonomously, by default, by your agents. And humans are still going to be making the decisions, but they're really going to be driving the new bets. And, well, what are the sort of change my company trajectory, you know, technical directions to take on versus you got some user feedback, there's a bug report, you know, there's some crash needs to investigate, all that should be handled proactively by agents.
31:00And so, you know, we have Devin wired up to a bunch of our Slack channels, triaging every message saying, is that something I should chime in on? Is it something I should investigate? And the sort of role of being a human engineer, it's just getting like crazy leverage. And so you have like, all these great suggestions of fixes that need to be applied. Oh, yeah, that looks good. I want you to change that here. And in particular, having the agent have the context of who's responsible for what in the code base in the company so it knows, oh, you know, Harrison, you should review this change or I should review that change.
31:32It's really quite exciting.
31:34Russell Kaplan:I'd love to hear how you guys think about UX because you started off, I think, and made kind of like the Slack coding experience. Like, I think of you guys when I think of that. You mentioned you guys teamed up with Windsurf. You also have Devon CLI. you're now talking about kind of like a different type of UX where you're not kicking things off, but it's coming to you. How do you see like the UX of coding agents over time and in the future evolving? Like where will humans be in the loop and what will working with these things look like? Yeah. So I think individual developers working with coding agents are going to ask themselves the question of how do I create self-driving software?
32:09How can I do sort of, you know, fixed costs, upfront work for continuous productivity benefits and gains.
32:17Russell Kaplan:Is this what the term software factory means? I think some people use that to me, it's sort of the wrong term, because when I think of a factory, I think of like the same thing being produced, you know, again and again versus the whole magic of coding agents and this like proactive shift of engineering is that it's the opposite of that. It's bespoke. It's fully custom and with all the right context of exactly what's going on. So I personally have never been a fan of the term, but generally this idea that the baseline engineering and technical function of your company is going to just be getting better and better by default because it knows what objectives it's optimizing for.
32:51It knows the input context coming in, it knows how to react to that, and it can learn from your feedback over time, I think is super powerful. And we see this within the companies we work with too, where individual developers have essentially realized I can be the CTO of an army of 10 ,000 agents. and maybe the most acute use case that's just growing like wildfire right now for us in this regard is vulnerability remediation because you know there's sort of this generational shift in the threat surface to a lot of the especially large firms given how good these models are getting at cyber security both offense and defense so there's kind of a lot of urgency right now on okay let's go patch all the vulnerabilities that we have but how do you actually do that well you know, a developer probably has to sit down and think, OK, how can I really leverage Devin to go find all the different issues and then remediate them en masse?
33:42So you can have one person setting off, you know, like an enterprise wide API call that's doing the equivalent of thousands of engineers worth of work. How do you do that right? So for security specifically, there's enough nuance that we've actually shipped pretty like vertical specific features and product services in security. We announced something recently called Devon Security Swarm. One way to think about it is agentic map reduce for finding and then fixing security vulnerabilities. One of the original advantages of Devon is that it was the first cloud agent. So every session is running on its own micro VM.
34:17And because of that, you can actually safely reproduce potential security vulnerabilities and replicate them and then validate that they're fixed. And so there's a sort of a scanning phase where you have to figure out, okay, of my massive code base that doesn't fit in the context window of any one LLM, what are the sub parts that are actually vulnerable, and you have to sort of shard it out. And then once you find those parts, how do you fix them, combine them, and aggregate them in a way that is technically correct and actually patching the issues. But that paradigm of one engineer is now empowered to take on the whole company, make something better, I think is really exciting.
34:51It's like each person, there's nothing stopping you from just having way more output than, you know, than you could have had only recently. And so that's like a really big change. And it's like changing the skills of like the skills that are most valuable, I think, of being an engineer, thinking much more big picture of, OK, like, what are the most important things we have to get done? How can I do this like massive set of changes to get them done and then leave behind a system that is self-improving?
35:18Russell Kaplan:Who has those skills? And kind of what I'm getting at is like, do you have to have a lot of years of expertise to understand how the whole system fits together? And if so, like, how will people who are just graduating now get that expertise in order to do this new skill, which is now maybe the only skill that matters? Yeah, sometimes people ask, like, oh, what's the future of being, like, a new grad engineer? Because how are you even going to get the experience if, like, the coding agents are so good at new grad engineer level tasks? But I think the flip side is everyone has barely any years of experience when it comes to working with agents.
35:52So in some sense, we're kind of all starting at the same level. And there's a new skill ladder to climb, which is how can you work with agents really well. And the best way to do that is to just try and just learn, you know, do stuff, you know, consume the resources and the training and the reading materials, but actually just go build stuff and then see what's working and not working. And in particular, keep your sort of finger on the pulse of the current limits of the frontier of capabilities and constantly reassess that. I think one of the biggest mistakes people make when you're trying to learn like a new way of working is you sort of try something and say, okay, that didn't work.
36:27I guess that doesn't work.
36:28Russell Kaplan:But as we know in AI, every two months, every three months, you have to continually retest that because these systems are getting so much better. I actually say that what I found both internally at Cognition and just working with Devin is you actually have more relative advantage as a new grad or as a recent engineer because onboarding has never been easier. You can ask as many questions as you want, including the silly ones to an agent that will not judge you. It will help you understand exactly what's going on. It will point you to the right parts of the code base to say, oh yeah, like this is actually where this is implemented.
36:59This is where it's done. And so we're seeing the opposite where people who would never have called themselves engineers not that long ago are steering the creation of software in ways that are totally crazy. You know, we started focused on professional software engineers and that's still, I think, the main users of Devon. But so many people are now using it to make things that we didn't think possible, even though, you know, that wasn't our original focus.
37:23Russell Kaplan:You mentioned a few of these like specialized versions of Devon. So Devon Review, Devon Security Swarm. I think there's also this idea that coding agents are general purpose and can be used for everything. How do you guys think about what should be like a special version of Devon versus just, hey, just go use normal Devon for that? Yeah, so we debate this internally a lot. I think one of the original kind of like founding theses of the company was that, yeah, if you solve code, you can kind of solve a lot of other things. We are seeing this, I think, play out in real time, both in terms of the popularity of coding agents and in the revenue numbers of people working on code and really the whole adoption profile of code.
38:01We don't want to over-engineer something that then the next version of Devin is going to make completely outdated and useless. At the same time, part of the role we serve in the ecosystem is to be the independent agent lab. So how do we take the best of every underlying model and then go really solve our users' problems end to end, like in the very specific details of what they're working on? So we're still super focused on software engineering as the core of what we do, but more and more things touch software engineering. And so, you know, one of the reasons we did this like specialization push on security is that we just looked at the distribution of how people were using Devon and realized, hey, that's like one of the most popular uses of Devon.
38:42So how do we double down on that? And I think in general, like my sort of personal philosophy on product development is you want to be allocating kind of some portion of your time to fixing the frictions and the paper cuts that your user experience. And then you also want to be allocating some portion of your time on doubling down on positive surprises. What are the ways that people are using Devon that we didn't even think of? And then how can we double down on that to make their experience even better? And a lot of times doubling down is, hey, let's like try to just improve general agent capabilities.
39:10But other times it's, OK, what are the more specific interfaces that we should build to make it work well?
39:16Russell Kaplan:You mentioned being an agent lab. There's also these model labs, which you are competing with at times, but also using a lot of their models at other times. How do you think about that and what does that landscape look like over the past few years? Yeah, so we both like have really deep collaborations with the model labs and, you know, kind of use each other's stuff. And frontier models are a really important part of Devin's own delivery. And then there's definitely, I think, Model Labs not just in coding, but in all of the application domains are starting to get into the application layer too.
39:45I think for us, we've learned a few things, you know, working in that space. One is that there is actually a pretty interesting structural differentiation of where we sit in the ecosystem versus any one Model Lab, which is whether you're an individual developer or the CIO of a large enterprise, the only thing you can be really sure about is that the underlying models are going to keep changing. And it's pretty hard to predict in advance who's going to have the best model, who's going to have the best price performance model and so on. And feedback we pretty consistently get from, you know, the individual users and leaders and, you know, the CTOs that we work with is that it's really nice to have some sort of decoupling between the agents and the models so that as models keep getting better, you're not stuck on, you know, the wrong tooling because now your model is no longer the best.
40:33You want to be in a situation where every day when you wake up, you're very happy to hear that a new model has come out that's even better than before, regardless of who came out. I think the second big lesson is just focus on the last mile of complexity and problem solving that is most important for the customers. So we were pretty early in building a forward deployed engineering team, for example. Today, our forward deployed engineering team, we have more forward deployed engineers than non-forward deployed engineers at Cognition. And so, you know, by the numbers, we're actually mostly working with customers because the core product is Devon working on itself most of the time.
41:10Russell Kaplan:What does forward deployed engineering mean for you guys? Because I feel like everyone's got forward deployed engineers and sometimes they mean different things. So for us, it means folks who are very technical, but also good at understanding customer and business problems and can go on site with customers to help solve big outcomes. The way we work with the large enterprise, it's very different from how someone signs up on Devon.ai and uses the product. it's often pointed at a specific objective to go solve it's not just about here's the tooling you're using but it's hey you know we have got to move on to this new system by the end of the year or we're going to have a lot of problems and our current schedule has this taking four years how do we do this you know much faster and better and part of i think the appeal for individual engineers to be forward deployed right now is that you're kind of flexing a lot more muscles as an individual engineer, where you're still in the technical weeds of what's going on, but you can be that CTO of an army of a thousand agents and like work on your own skills to just sort of master this new way of working.
42:11And that's like really valuable for us, but also really valuable for our customers.
42:15Russell Kaplan:What's the right profile of someone who wants to be a forward deployed engineer at Cognition? We have had a lot of success with ex-founders in particular. So people who are very comfortable in ambiguous kind of problem areas and can do the technical parts, but also can understand customer pain points, talk to them, and really have very high internal agency to recognize, oh yeah, that's the problem we should go solve. Which is another big thing we've learned in the process of deploying Devon more and more widely. You can use coding agents for almost anything. And so one of the big questions is, well, what's the most impactful thing or set of things to get started on?
42:50And helping steer that is I think one of the things that makes a great forward deployed engineer so great.
42:56Russell Kaplan:We talked about this a little bit earlier, but I think costs for coding agents are becoming something that a lot of companies are paying attention to. And the other part of that is kind of like the return that you get for everything that you're spending on coding agents. And quantifying ROI is really hard. I think you guys have done maybe one of the best jobs at it, or at least maybe even just one of the only people I've seen actually try to take this on in a general way. You wrote a blog post on like estimating productivity and you have some productivity guarantee. Could you talk about how you guys think about that?
43:24Yeah, we think about this all the time. So first on the productivity side, what's the problem? The problem right now is that a lot of people in the industry have a strong incentive to get customers to token max. You know, let's build leaderboards of usage and kind of have things go up. And the incentives are actually quite fundamentally misaligned, I think. And at some point, you know, the bill comes due. And I think we're starting to see that now at a lot of customers where they have been token maxing, they've been using a lot. And now suddenly they're spending in the tens or hundreds of millions of dollars or even more in some cases a year.
43:57And then the leadership is asking, hang on, what do we get for this? So there's kind of this reset happening. Going back also to your earlier question of like, where do we sit in the ecosystem as the independent agent lab? I think one of the side effects of being independent is that structurally we're the most incentive aligned with customers to not just focus on, oh, you got to drive usage of this model, but really focus on, well, what's going to be the value you're going to get? and let's not even work on the stuff that's going to be low value. So that was one thing we focused on early on. And we've always tried to work closely directly
44:29Russell Kaplan:with customers to just help them estimate and then quantify their own ROI of using Devon and AI generally. But part of that's a manual process. It's literally sitting down and saying, okay, well, what are the most important things for the company right now? And how could we move the needle together if we worked on them closely? What we were looking to do is enter a sort of a more scalable question, which is how can we make this happen in an automated way so everyone can benefit? Whether, you know, you're a big enterprise or an individual developer using Devon, how can everyone benefit from that?
44:59And so that immediately imposes one design constraint, which is, okay, how can you automatically estimate the productivity or ROI or value? And I think our conclusion was, at least today, 2026, it's very hard to automatically estimate ROI because you don't have the full business context of, oh, you know, this is going to drive your, you know, your revenue profits up this much or this new product is worth worth this much to you but we could do one level below that which is productive engineering output so there's a difference between i'm using an agent i'm getting a lot of tokens out a lot of code out and oh this was actually valuable time savings for for me and so to measure that we actually worked with a bunch of our of our customers to collect a data set of a bunch of dev and sessions where we scored them together with the customer to say, okay, first of all, was this actually productive or not?
45:49And some rules you can do automatically. So for example, if you create a Devon session and that creates a PR and it never gets merged, we call that an unproductive session. Now, in reality, like maybe it was still actually helpful for you, but we want it or conservative. So we say, okay, that's just blanket unproductive session. If you merge a PR production, then we would say, okay, that's productive. If it didn't result in any PR, because sometimes people are using Devon for data analysis or other workloads, then we do an ML-driven classification based on kind of a manually tagged data set in collaboration with customers.
46:16So the result of all that is that we essentially have an evaluator agent that can take a session and say, was it productive or not? Part one. Part two is, if it was productive, how many hours of work did that actually save you? And again, manual data collection grind to figure out that question where we surveyed a bunch of folks who work closely with them to see, hey, you're still using AI, maybe even if not deviant. So like, what's the delta there? And so we collected this data set of essentially engineering hours estimates for their equivalent Devon session. And by combining these things together, we can now automatically apply for everything you run from Devon, a score of was it productive or not?
46:51And if it was productive, how many hours did this save you? And that's really interesting visibility for individuals, for teams, and for companies. And when we ran the numbers, we realized that as we were hoping, people are saving a lot of time with Devon, but they were saving so much time that it gave us the confidence to actually financially underwrite it and go to our customers and say, we are actually going to make a$10 million productivity guarantee, which is if you're paying for Devin and people are in fact just wasting usage and they're not getting value, you know, if you end up paying us more than the like engineering hours of value that you would expect in dollar terms, we're just going to refund the difference up to$10 million.
47:28And when we debated this internally, I got a lot of pushback actually, because some people were saying, wait, this is like a big risk. We're signing up for a financial liability. Like, what if someone just like makes an API call and they just run Devin in a loop and it totally burns, you know millions of dollars and the reality is yeah like if you did that then we would we would be on the hook and we would lose money so let's go add some guardrails in the product to prevent you from being able to do that it's like how do we align our incentives with the customer's incentives so we now have really good cost controls we have the ability for admins to say like this is the overall budget i want to allocate help me be smart about like who should get you know more compute who's still learning and needs kind of a tighter a tighter leash and i think it's it's all about like aligning the incentives with customers so that they're actually getting value for what they're using.
48:11Russell Kaplan:I was going to ask about that last thing you mentioned, which is basically like, how do people use this info? And do they use it at a team level? Do they use it at an individual level? I think intuitively inside LinkedIn, if you asked our VP of engineering, I think he'd probably say that more senior engineers he trusts to use these things more and more junior engineers, it's easy to mess up how you use these things. It's easy to accidentally spend stuff. But I'm curious, yeah, like how do people use these insights? I think that my favorite way and one of the most popular ways of using these insights, it's really for learning and like professional development because you can look at, you know, two teams and you say, oh, wow, this team is actually, you know, really economically efficient with their use.
48:50They're getting tons of productive output for what they're putting in. And this team is kind of blowing up their costs a little bit. Well, it's not like they're maliciously intending to do that. They're just maybe using the tools differently and maybe they don't realize that they're doing some things that are really inefficient. So we've tried to then, now that we have these dashboards in product where you can see, oh yeah, how is my usage compared to, you know, other folks. And where do I sit? And like, what can I be doing better? We're trying to push more of that coaching and that like professional development sort of into the product itself, where you can like learn from that.
49:18Devin will actually tell you, Harrison, that was a bad prompt.
49:20Russell Kaplan:You know, like this is like way underspecified. You know, you should fix this. Like here's some tips on how to be better. And so we're trying to put more of that training into the product so people can kind of continue to level up with the capabilities because it's actually really hard for anyone to keep up with how fast everything is moving and things that you didn't expect would be possible. a few months ago are possible now or things that you're doing today as like a workaround for limitations actually you shouldn't do tomorrow. So we've tried to move as much of it in product as possible. And I think people are actually using the dashboards the most.
49:48It's like literally seeing, okay, team by team, individual by individual, how can I learn how to be more productive, more efficient, and just really master this new way of working.
49:57Russell Kaplan:Thanks for listening to Max Agency. If you liked this episode, leave a review and subscribe. Send feedback or questions to maxagency at langchain.dev. We want to hear from you.
From the publisher
Russell Kaplan is the President of Cognition, the $26 billion company behind Devin, the AI software engineer. He started his career in machine learning at Tesla's Autopilot team, then led the ML org at Scale AI, which had acquired his computer vision startup Helia. In this conversation, Russell unpacks why more and more engineering work is becoming "intelligence saturated," how the sidekick architecture inside Devin Fusion cuts cost without cutting quality, and why Cognition is underwriting a $10 million productivity guarantee.
–
We also discuss:
- Why mergeability is the next eval bar
- When to use Fable vs GPT models
- How smarter routing unlocked 35% better price performance
- Why letting users pick their own model is a UX bug
- Wiring agents to be proactive instead of reactive
- How Cognition scores a session productive or unproductive
–
Timestamps:
(00:00) Introduction
(01:21) Launching Devin at 13% on SWE-bench
(02:47) When Devin became the number one committer to Devin
(05:56) Intelligence saturation: when speed and cost start to matter more
(07:05) When to use Fable vs GPT models
(08:41) Why mergeability is the next eval bar
(10:51) How open source maintainers shaped FrontierCode
(14:56) The telltale sign of agent-written code
(17:56) When per-person token spend starts to eclipse salaries
(20:26) Why letting users pick their own model is a UX bug
(21:46) How smarter routing unlocked 35% better price performance
(23:50) Why Cognition still trains its own models
(27:13) Deploying Cerebras at scale at 950 tokens per second
(28:21) The Tesla chip debate: does an inference chip need division?
(30:06) Wiring agents to be proactive instead of reactive
(32:19) Why "software factory" is the wrong term
(35:16) Why everyone is now a new grad
(41:10) Forward deployed engineering, and why ex-founders thrive
(43:16) The industry's incentive to get you tokenmaxxing
(45:11) How Cognition scores a session productive or unproductive
(46:51) Underwriting a $10 million productivity guarantee
(48:11) Using productivity dashboards to coach, not police
–
References:
- Anthropic
- Cerebras
- Claude Fable
- Claude Opus
- Cognition
- Cognition's AI Productivity Guarantee
- Devin
- Devin Desktop
- Devin Fusion
- Devin Review
- Devin Security Swarm
- Estimating the Productivity of an Autonomous AI Software Engineer
- FrontierCode
- GPT-5.5
- GPT-5.5-Cyber
- Harbor
- LangSmith
- OpenAI
- OpenRouter Model Fusion
- SWE-1.6
- SWE-bench
- Tesla
–
Where to find Russell:
–
Where to find Harrison:
–
Where to find LangChain:
–
Send feedback or questions to maxagency@langchain.dev




