In short
Podcast Notes: Training Data - Factory’s Matan Grinberg and Eno Reyes Unleash the Droids on Software Development
Episode Summary In this episode of *Training Data*, co-founders of Factory, Matan Grinberg and Eno Reyes, discuss the transformative impact of AI on software development. They emphasize the potential of AI to drastically enhance productivity in the software engineering lifecycle through the deployment of "Droids", specialized AI agents designed to automate various tasks.
Key Themes and Concepts
- The Compound Lever Concept
- Archimedes' Quote: The episode opens with a reference to Archimedes' assertion about levers, likening software engineering to a powerful lever capable of moving the world. AI is seen as a tool to amplify this leverage exponentially.
- AI in Software Engineering: AI can enhance productivity by automating routine tasks and thus facilitating faster software development cycles.
- Factory and Its Approach
- Droids: Factory has developed a fleet of “Droids,” each tailored for specific tasks in the software development lifecycle (code review, testing, documentation, etc.).
- Avoiding Foundation Model Training: Instead of creating a new foundation model, Factory builds on existing models to provide immediate value to developers, allowing them to evolve alongside advancements in AI technologies.
- Optimizing Software Development Processes
- Metrics of Success: Factory focuses on system-wide optimizations rather than just individual developer productivity. Key metrics include:
- Code churn
- Time to merge
- Engineering velocity
- Real-World Applications: The Droids are designed to assist engineers in real-world environments, thus enhancing the overall performance of engineering teams.
- Cognitive Architectures and Reasoning
- The discussion touches on cognitive architectures and the need to model AI systems that closely mimic human reasoning processes.
- Examples of Cognitive Architectures: They reference Monte Carlo tree search and language agent tree search as frameworks that could be employed to handle decision-making and planning tasks.
- Competition and Market Position
- Matan and Eno discuss the competitive landscape, asserting that many startups are trying to train their own foundation models, which they believe is not the right battle for Factory.
- Focus on Execution: They emphasize that successful execution and obsession with their mission—bringing autonomy to software engineering—are their competitive advantages.
- Future of Software Engineering
- The duo expresses optimism about the future of software engineering, envisioning a shift towards higher-level thinking and architectural roles for engineers as routine tasks become automated.
- Long-term Vision: They believe that with the advancement of AI tools, the demand for software will only increase, enabling a new era of productivity and creativity.
Important Mentions
- SWE-bench: A key evaluation framework for assessing AI systems’ ability to solve real-world GitHub issues.
- The Bitter Lesson: An essay by Rich Sutton that discusses the importance of scaling in search and learning.
- Success Metrics: Factory reported improvements in cycle time and reductions in code churn through their tools.
Episode Highlights
- Introduction and Background (00:00 - 01:36): Matan and Eno share their professional backgrounds and motivations for founding Factory.
- What is Factory? (12:41 - 16:29): An overview of Factory's mission to automate software engineering tasks.
- Results and Case Studies (30:04 - 32:54): Insights into current metrics and the effectiveness of Factory's Droids in real-world applications.
- Competition and Market Strategy (43:02 - 45:32): Discussion on how Factory positions itself amidst competitors and its unique approach.
Conclusion The episode concludes with a focus on the future of AI in software development and the transformative potential of tools like Factory’s Droids. Matan and Eno's insights underline the importance of aligning AI solutions with the needs of developers to foster a collaborative future where technology enhances human creativity and productivity.
---
For further details, listen to the full episode [here](https://www.sequoiacap.com/podcast/training-data-factory/).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00I would have thought this would be different 13 months later, but this is still very much the case where Agent is synonymous with unreliable, stochastic, demoware, vaporware, and I think something very important for us is we want to build these systems that aren't just like cool examples of what is to come, but rather valuable today and not just valuable for like a hacker on a side project, but valuable to enterprise engineers today
0:43Hi and welcome to Training Data. We have with us today, Matan Grimberg and Eno Race, Founders of Factory. Factory is building autonomous software engineering agents or droids that can automate everything from the drudgery of maintaining your documentation to actually writing code for you. In doing so, they are building the ultimate compound lever. Last week, Factory also announced some impressive results on the key AI coding benchmark, Sweet Bench, being state of the art by a wide margin. Stay tuned at the end of the episode for more context on how they built it. We're here with Patan Grinberg and Eno Reyes, Founders of Factory.
1:20Gentlemen, thank you for joining us. Thank you so much for having us. Yeah, thanks for having us. Let's start with a little bit of personal background. And Matan maybe we'll start with you and then go to Eno. So Matan, one thing that I believe you and I share in common is an affinity for a well executed cold call. I know at least two cold calls that have had some bearing on your life. One we start with the one that you did as an undergrad at Princeton to somebody who my partner Sean McGuire tells me is quite a famous physicist. Can we start with that cold call? Yeah, absolutely. So while I was at Princeton, I was studying string theory.
2:00And the most famous string theorist happened to be working at the Institute for Advanced Study, which is an academic institution right next to Princeton University, but not technically affiliated with it. Part of the allure of going to the IAS is that you don't have to take on graduate students, much less undergrads. That said, there's a professor there, Juan Maldesana, who is by far the kind of leader of the string theory movement, and being a young, ambitious, undergrad, I decided might as well see if I could snag him as an advisor. So with some advice from some graduate student site, send him an email, ask if we could meet.
2:46And the thing about Juan is the way he works with people. You know, he'll take a meeting with anyone basically and we'll spend about two hours at the chalkboard with you. And in this two -hour chalkboard session, he'll subtly drop problem that you basically have 24 hours to solve, get back to him with the solution, and then you'll officially, you know, be a student of his. Luckily, I was warned about this, about this right of passage, so, you know, I was paying close attention to that any hints he was dropping. You have down the controls. Indeed, yes, yes. So found the problem, ended up spending, you know, basically the entire night working on it, and you know, luckily ended up having him as an advisor.
3:30We were able to then publish a paper together, which was very exciting. So, yeah. Typical undergrad experience. Yes, yes, exactly. So there's a second cold cold that I wanna ask you about, But before we get to that, why don't we go to, you know, you similarly went to Princeton. You have a CS degree from there. You spend some time as a machine learning engineer at Hugging Face, which is where we first intersected, spend some time at Microsoft. But like a lot of great founders, your story before then started with some humble beginnings. Could you say a word about the stuff that doesn't appear on LinkedIn that has helped to shape who you are today?
4:06Yeah, absolutely. And I think that, you know, my family on my dad's side came from Mexico in the late 60s to San Francisco. And my grandparents were both working for a bit, but when my dad was born, they started a Mexican restaurant in Los Althos. And that was in the 70s. They moved it to hate and coal in the 80s. were a very kind of like San Francisco immigrant story. They actually ended up leaving to Georgia where I grew up, but really I think it's the drive that they had to give my dad a successful life in America. And it was my dad and my mom that drove that same kind of mentality into me growing up.
4:51And I think it's really cool because this story is one that I think a lot of Americans share and something that makes it really exciting to be back in San Francisco and building something to potentially make the world a better place for everyone. Very cool. That is the dream. Natana want to get back to that other cold call because I think it leads directly into the the thunder or the forming of factory. So our partner, Sean McGuire, who I mentioned earlier, who I believe shares a similar academic background to your own received an email from you a year or so ago that led to a walk and very shortly thereafter factory was formed.
5:33So I'm curious, what caused you to call Sean McGuire? And this is less of a Sean McGuire question because we know plenty about Sean McGuire. This is more of like a, you're doing, you're on a very good path, you're doing really good research, you're on a on track to get a PhD in physics, and something inspired you to go in a different direction. And I'm curious, what was it that inspired you that led to that cold call and maybe tell us a quick story about what happened shortly thereafter. Yeah, absolutely. So, you know, I was like he said, I was doing my PhD at Berkeley. About a year in though I realized that I was only doing theoretical physics or instinct theory because it was hard and not because I actually loved it, which is obviously a bad reason, a bad reason to do anything.
6:17And, you know, I would have had such tunnel vision on this path that, you know, when I came to this realization, it was kind of earth shattering, and looked at the paths ahead of me, and there were basically three options that seemed realistic. And so it was either going into quantitative finance, going into big tech, or going into startups. And by this time, I'd already kind of switched my research at Berkeley from being purely physics to an ML and physics, and then slowly more ML, and then mostly AI. So it was kind of quickly, quickly cascading there. At the time, I think I saw a video of Sean speaking, I think over Zoom to some founders at Stanford or something.
6:58And I recognized his name from String Theory research because I had read his papers way back in the day. And it was particularly shocking to me because I'm not sure how much time you've spent with String Theoryists, but normally they're quite introverted. Not, you know, not, not, not the most social. Yeah, and so Sean is, you know, this very different example. And so to me, I kind of like, look, I looked at his background and it was just, it was just shocking to see someone who was like so deep and like a bonafide string theorist, then go in and like, you know, start his own companies, invest in some of the best companies, join Sequoia and be, you know, a partner there.
7:36And to me, it was just like, oh my God, this seems like someone who is of my kind of background of my nurturing, I guess. And so sent him an email and I was just like, Hey, you know, we both were string theorists. I don't want to do string theory anymore. I'm thinking about AI. We'd love to get your advice. Like you mentioned, that then turned into a walk. It actually was supposed to be a 30 minute walk. We ended up going from the Sequoia offices in Menlo Park all the way to Stanford and then back. And so it ended up been three hours. He missed a lot of his meetings that day. So it was it was pretty amusing and basically at the conclusion of the walk he So the one thing was for sure he was saying you must drop out of your PhD There's way too many exciting things to do and he kind of left me with the advice of You should either join Twitter right now because this is just after Elon took over And he was saying it's you know only the most badass people are going to join Twitter right now to you should join a company of mine as just like a general glue is what he said.
8:40This was foundry, by the way. Yeah. Or three, if there's some ideas that you've been thinking about, you should start a company. And I was like, you know, very grateful of all the time that he spent and we kind of left off there. Beautifully in parallel, Eno and I had just reconnected at a chain hackathon. And he was in Atlanta the weekend prior and he basically got back the next day. So that next day, you know, and I got coffee and I think we got coffee at noon and then basically every hour since then until now You know, and I've been working together talking constantly About co -generation and what became factory I guess originally there's no each other an undergrad We had like the maximal overlap of mutual friends without ever having had a what for one conversation Yeah, it's pretty funny.
9:35We were in eating clubs at the time opposite from each other and we had just so many mutual friends and It really wasn't until I moved to the Bay Area that we had a face -to -face combo and we It was a very fruitful conversation for sure. It was intellectual love at first sight. You could say Absolutely I love that and it's so serendipitous with the Lengchain connection How did you guys decide on, I'm curious, I mean, you're both brilliant. And I think for a lot of founders starting out in AI right now, a lot of them find it hard to resist the sirens call of training a foundation model. So like, how do you decide to build in the application layer?
10:16I'm curious. And then why, why software engineering? Yeah, so I think from my perspective, like going deep from academia, I think throughout throughout all the years of spending time on math and physics, the thing of beauty that I learned to be drawn to was things being fundamental. And spending time doing a research, it was so clear that code is fundamental to machine intelligence. And so I was just naturally attracted to the role that it plays there. And I think that kind of joined quite well with ENOs attraction to the space. You've referred to it a couple of times as a compound lever. Can you unpack that for us and let us know what that means?
11:02Yeah, I mean, so there's the famous Archimedes quote about software. Well, his quote is rather, if you have a large enough lever, you can move the world. And then I think that's been co -opted for software engineering, right? That software is a lever upon the world. And for us, we see AI and in particular, AI code generation as a lever on software, the impacts of that being compounding exponential. And sorry, I think I cut you off. I think you were maybe mentioning how you got to the founding inspiration for factory. Oh, yeah, absolutely. I mean, I think Mattan's story is really indicative of kind of the energy at the time.
11:43I was at hugging face working on training, optimizing, deploying LLMs for enterprise kind of use cases. I was actually working with Harrison on early like lane chain kind of integrations. And it was so clear that the work that was happening in open source was directionally moving towards modeling human cognition with LLMs where the LLM was just one piece of the system. The idea of chains and I think Harrison calls them cognitive architectures or the Langshane folks call it that and seeing that happening and seeing that within the code gen space, the most advanced players were basically looking at autocomplete.
12:25It felt like there was a huge opportunity to take that to the next step and take some of those lessons that were happening both in the fringe research and open source communities and applying them towards kind of massive organization. I realize we haven't said explicitly yet what is factory. So, Matan, what is factory? And then maybe what are a couple of the key decisions that you've made about the way factory is built? And you know, for example, one of them is to start by benefiting from all the ongoing improvements in the foundation model layer. You know, one of them might be the product itself, But can you just say what is factory and what are some of the kind of key decisions you've made that have shaped factory today?
13:08Yeah, absolutely So factory is a cutting edge AI start up Our mission is to bring autonomy to software engineering What that means more concretely we are automating tasks in the software development life cycle And in particular tasks like code review documentation testing debugging refactoring And as I list these off, you'll hear quickly that these are the tasks that engineers don't particularly enjoy doing. And that's very much intentional, right? Like, obviously we are doing code generation and that's really important. But I think an equally important thing too, you know, generating some inspirational and forward -looking demos.
13:55It's also important to understand what engineers are actually spending their time on. And in most organizations, it's not actually fun development work. In most organizations, they're spending a lot of their time on things like review and testing and documentation. Normally, they'll do these things way too late and then they're suffering because they're missing deadlines, right? And so our approach is we want these tools to be useful in the enterprise. And so to do that, we need to kind of meet engineers where they are with the tasks that they are very eager to automate away. We call these autonomous systems droids, and like Eno was alluding to earlier, these are kind of, there's a droid for each category of task, and in this kind of a paradigm where we want to frame these problems as games, it's very convenient that software development has a clearly defined software development lifecycle.
14:52And so for each kind of category of task or each step in the software development life cycle, we have a corresponding droid. So that's kind of a kind of a first pass there. I guess there were I think there was a second part of your question that I missed. Well, we'll get into the rest of it. Where did the name droid come from? It's pretty catchy name. It's like very memorable and distinct to factory. Where'd that come from? Yeah, yeah, so I mean keep in mind, you know when factory started this was like you mentioned about a year and a month ago and And, you know, I actually, I would have thought this would be different 13 months later, but this is still very much the case where agent is synonymous with unreliable, stochastic, demoware, vaporware.
15:38And I think something very important for us is we want to build these systems that aren't just like cool examples of what is to come, but rather valuable to date. And not just valuable for like a hacker on a side project, but valuable to enterprise engineers to date. We felt very strongly that agents just doesn't really capture what we're trying to deliver. And so, fun fact, we were originally incorporated as the San Francisco Droid company, but upon legal advice and given, I guess, the eagerness with which Lucasfilm pursues its trademarks, we changed our name to factory. Fair enough. So is it fair to say then that a droid is sort of like a job specific autonomous agent that actually works?
16:25Is that a reasonable way to think about it? Yeah, okay. It's okay You just said the words cognitive architecture and I know my partner Sony long well enough to know that this is her love language So I'm sure that Sony is mine just lit up with a whole bunch of questions for you So I don't want to get in the way. So yeah, have at it We just had Harrison on the podcast who talked about custom programming architectures as well I guess what are you doing on that front and how do your implementation stuff tell with the multi -joyed strategy that you're taking? Yeah, absolutely. I mean, it's a great question and I think the the way that we think about reasoning and cognition within the the standpoint of these systems there are clearly huge innovations happening on both layers, the foundation model layer, as well as on the kind of orchestration or application layer.
17:20The way that you can kind of think of our technical approach on this is that, you know, traditionally labs like DeepMind and kind of some of these these orbs that are really focused on solving problems that you can model like a game where you have rules and an environment and feedback You can build out systems which model the reasoning of humans and even outperform them. They did this with the alpha series of models, protein folding, go, code. And for us, most of the reasoning work we do is similarly focused on kind of inference time reasoning. Search through kind of decisions and what we kind of think of as, you know, maybe it's something of intuition.
18:05maybe it's something of planning, but we aren't training foundation models yet. And I think a lot of the innovation that's gonna happen at the foundation model layer will be things like latency and context window and kind of performance on some subset of tasks, but any time that you need action and environmental feedback and kind of long -term planning, It's going to be really difficult to build a single foundation model that does that. And I think it's really the application there where those types of innovations are going to happen. Yeah. I thought the Princeton, Sway agent paper that came out last week or so was really interesting as an example of that.
18:51I've like, you can get incredible, authentic reasoning performance on code tasks from small open source models. I thought that was really nice. I've proved point to what you're saying. We love the whole team that put that together and the SWE Benchwork. I think it is a popular benchmark in the space. I think it's clear that a lot of the efforts towards building these systems relies on not just like any one benchmark or EVAL or set of tasks, but rather collaboration across a bunch of different areas, whether it's the model layer, whether it's the tasks themselves, it's how what data are you using to evaluate and ultimately like the overall architecture.
19:33And yeah, they're a really great team where we're super pumped to see that work. Okay, last question on this and then I will pause myself. Any favorite cognitive architectures? Like is it the tree of thought stuff, chain of thought stuff? Any favorite cognitive architectures that you think are especially promising or fruitful in your field? Yeah, I think that's a great question. I mean, I think kind of what I alluded to previously when you have like the almost like the game like problem space where there are kind of simulatable Analyzeable and optimizable Boundaries then that means that you can search through those decisions And there's a bunch of techniques like Monte Carlo tree search language agent research that people have talked about in research papers that I think are interesting approaches here.
20:27I think that in my mind, there isn't a singular cognitive architecture that makes sense for all tasks. And a lot of the benefit of breaking down the software development life cycle into kind of semantically meaningful segments is that developers, when they have these workflows that move from one step to the next, they've kind of defined the boundaries of the game, so to speak. And so a lot of the work we do is figuring out which cognitive architecture, what design makes sense for a given task. You're reminding me of the rich Sutton, but they're less in search and learning or the two techniques at scale.
21:09Yeah, absolutely. And I think you definitely need both. And then you know, you were talking about this a bit, how the sort of the reasoning layer on top of the foundation model, is really the focus for a lot of the fundamental research and a lot of the fundamental work that you guys are doing. But, Tom, you had a line a couple of months ago when we were talking that was, and hopefully this doesn't come across as starkey because it's not nintu, but it was something to the effect of, there are eight interesting engineers that open AI working on my margins for me. Can you say a word about that? Cause I thought that was, first, I was incredibly well put, and then second, pretty good insight in terms of how you're building the business and really benefiting from the work of the foundation models.
21:52Can you just say a couple words about that? Yeah, absolutely. So, you know, there are a lot of companies, a lot of startups that are, you know, pursuing training foundational models or fine tuning models. Then there are a lot of huge research labs like OpenAI and Anthropic, who are also putting a ton of resources behind making these foundational models, better, cheaper, faster. And from our perspective, we don't want to run a race that's not suited to our abilities. We don't want to fight a battle that we know we won't win. Training foundational models, we are not going to win that battle. And similarly, I also don't think it's a particularly unique battle at this point.
22:34I think these companies were incredibly unique and innovative, clearly based on what they're delivering. But now, I think the stages set in terms of training foundational models. And I think similarly with a lot of the infrastructure for fine tuning and that sort of thing, what has not really come to fruition yet is actually making products with AI that people are using. There's so much to talk about. All these foundational models, all this infrastructure, and there's still very few real products that use this AI. In the analogy that VCs like to talk about a lot, We have a ton of picks and shovels and no one's actually going for gold And so the thesis behind how we're building this company is let's first you know use these beautiful black boxes that Open AI and Thropic and Google are spending billions of dollars and you know hundreds of engineers to make Let's use these black boxes and build a product that people are actually using and once we do that Then we can earn the right to do the fancier things like fine tuning and training If you're unable to build a product that people are actually using with these incredible models, then chances are fine tuning and training will not save you, and it's probably just not a good product.
23:46And so that's kind of the approach that we're taking there. And so, you know, we do get a lot of improvements when new models come out. But yeah, we are very much grateful for the work that's being done at these cutting -edge research labs. I want to, you've seen a lot about how what you're doing is kind of like making AI immediately practical for engineers in like an enterprise setting. And so I went through another, I think it's a Moton quotes, and I'm not sure if you were quoting somebody else, but you said, you were talking to us last time, you said, if Jeff Dean shows up at your office and he doesn't understand your code base, he won't be productive.
24:21And, and, and, and, in fact, that for us, like what does it take to kind of make a coding agent that's not just good for anybody that boots up a computer, but somebody that's a full -time engineer at a real software company. Yeah, totally. And yeah, so the analogy here is that, you know, Jeff Dean is the analog of a really, really good foundational model. Let's say like GPT -6 with incredible reasoning, right? But if it comes into your engineering organization with all your nuances and all your engineering best practices, just having that good inference and good reasoning is not enough to actually contribute and automate these tasks reliably.
24:59Some given isolated tasks, sure, you can solve. Like give it some like lead code problem and this give Jeff Dean a lead code problem, I'm sure he will solve it. But if you have some, you know, 20 -year -old legacy code base, some part of it is dead code, the other part of it, the last person who was contributing to it just retired, and so no one else knows what's going on there, you need deep understanding of the engineering system. Not just the code base, but why you made certain decisions, how things are being communicated, what top of mind priorities are for the engineering organization. And it's kind of these less sexy but incredibly important details that were really focused on in order to deliver this to the enterprise.
25:40What about the, I think a lot of these AI code, coding companies are kind of focused on the individual developers productivity. How do you think about the individual level optimization versus maybe the system -wide optimization? I think the important thing to think about with respect to the whole org is when a VP of engineering comes into the room, they're not really focused on whether or not an individual completed like one task, an hour faster. They're concerned about how many tasks are being completed, an aggregate metrics of speed. But if that person completed that task an hour faster, but it's 40 % worst code, right?
26:26It's churny code where people are going to rewrite it on top of it. Or that person took that task and, you know, they did it an hour, but it took them four hours to plan that. And they were blocking five other engineers. And so when you start to actually add the nuance of what does it mean to be successful, measuring an engineering org, you start to bump into a lot of challenges with understanding what needs to be improved and what is a bottleneck and what is just kind of a secondary metric. I think a lot of the initial attempts at making AI coding tools are really focused on first order effects.
27:03How quickly is somebody tabbing to autocomplete statement or how quickly is somebody completing an individual task? But I think that at factory, a lot of what we're trying to do is understand from an engineering leaders perspective, how are you measuring performance? And what are the metrics that you look at to understand, hey, we're doing really well as an org or hey, we need to make improvements and targeting those. And I think metrics like code churn end -to -end open to merge time, you know, time to first answer within the end org. where all of these things are much more impactful to an organization's speed of shipping code.
27:43And so that's kind of how we think about it. I think this really ties into what ENO is just saying quite well, which is, you know, the clearly, we were talking about products earlier as well, like clearly the AI product that has penetrated the enterprise the most is Copilot, right? Unfortunately, with a tool like Copilot, the things that are kind of, the metrics that are really held up as success or things like autocomplete acceptance rate. And the problem is exactly to your point, if you're a CTO or a VP of engineering, how do you then go to the executive team and say, hey, look, our autocomplete acceptance rate is this high?
28:18They don't know what that means. They don't understand how that translates into business objectives, right? And also, you know, Eno was alluding to this. There's kind of a hidden danger to some of these autocomplete tools, which is orgs that use tools like this end up increasing their code churn by anywhere from 20 to 40%. There's some studies that look into this. There's some problems with these studies, but, you know, directionally, what's clear is that, as the percentage of AI generated code increases, code churn, if you're not, you know, doing anything different in your review process, code churn is gonna go up.
28:54And so, our reason for focusing on org -wide metrics is that it kind of divides out all of these concerns. If we look at things like, like how fast are you completing your cycles? What is your code churn across the org or across these different repos? That divides out these kind of like smaller, like intermediate metrics and gives you a sense of, hey, we are shipping faster and we're churning less code. So that's really how we talk about this with these engineering leaders. At the end of the day, the three kind of main axes we look at is saving engineering time, increasing speed and improving code quality.
Read the full transcript
29:37And also, so these are three. And again, there's kind of different, you know, complexity of metrics for different parts of the org. These are the three that we discuss with engineering leaders, but we want to arm them with information when they're talking to, let's say, their CFO. And so really, we kind of break that down into one main metric, which is engineering velocity. And that's really what all of these droids or targeted towards is increasing engineering velocity. Let me, let me try to recap a couple hearts of the story thus far. So in some ways, this is a compound lever, meaning AI is a lever on software.
30:14Software is a lever on the world. And so building an autonomous software development system is one of the most impactful things you can possibly do with your lives, which is pretty cool. There are a few unique angles to the approach that you guys have taken. Maybe not unique but distinctive. One of which is the decision to write on top of the foundation models, which means that you get to benefit from all their ongoing innovation. It also frees you up to really focus on the reasoning and the agentic behavior on top of those foundation models, which is part of the reason why you can deploy your product as a series of droids, which are basically job -specific autonomous agents that do something like test or review into end in a way that is practically useful to an engineering organization.
30:58And instead of focusing on just producing more code, you're actually focused on kind of the system -wide output, which requires you to have really detailed context around not just the code base, but all of the systems and processes and kind of nuance around the entire environment. And having done so, you can increase, you know, velocity for an organization. I think that's a bunch of the story that we've talked about so far. Let's talk a bit about the results. Are there any good customer examples you can share of factory and action and the results that you've been able to have for people? Yeah.
31:36I think some of the main things that we're seeing across the board, and we're not super public on case studies just yet, but something that we see across the board is our average cycle time increase is around 22 % on average we are lowering code churn by 13 % tools like I guess we haven't even gotten into the specific droids but tools like the test droid end up saving engineers like around 40 minutes a day which is pretty exciting and yeah I think kind of going back to what we were talking about in terms of benchmarks. One of the most exciting things about having thousands of developers who are actually using these tools is that we get this live set of benchmarks and we get e -vals and feedback from these developers about how these droids are performing.
32:31And so, like you know, imagine we are huge fans of SweetBunch and what that's done kind of for the general community and giving people like an open source benchmark to really compare these models, but strategically for us, having this deployed in the real world has allowed us to dramatically increase our iteration speed in terms of quality for these droids. What did you guys learn since you have a bunch of people using this in the real world? What have you learned along the way? Have there been any big surprises? Engineers love ownership. Yeah. All right, same more. Absolutely. I mean, I think it really is that, you know, when you're building an autonomous product, and the goal is to take, you know, take over a task, you have to deal with developers who are fickle for good reason.
33:23They're constantly bombarded with developer tools and automations and anything that's kind of being enforced from a top -down perspective needs to be very flexible. And so making sure that when we're building these products, we think about what are the different preferences or ideas that people have about how this task should be done, and then building as much flexibility into that. I think a great example of this is the review process. Everybody has a different idea of what they want code review to look like. Some people want superhuman, linters, some people want really deep analysis of what of the code change.
34:06Some people don't even like code review. They get annoyed by it entirely. Matton has a great quote about what code review is like. I don't know if you're gonna share that. Yeah, yeah. So I mean, in general, we've kind of internally realized that the code review process is very much like going to the DMV in that no matter how clean the DMV is, no matter how fast the line is, no one loves code review. Right, because at the end of the day, someone's criticizing you, someone's going in and looking out, you know, what you did and saying better ways you could have done it. So in general, the review process, it's the type of thing that as an engineering leader, it's great to see like moving the needle on these organization wide metrics.
34:48As a developer, it's maybe not the most fun thing, whereas something like the test droid, right, which is generating tests for you. So you don't need to spend hours writing your unit tests. That's incredible as a developer. But, you know, for the engineering leader, it's slightly less obvious how that connects directly to business metrics. So I think this is part of why it's important for us to have this fleet of droids because we are not just building this for the engineering leader nor are we just building this for the developer but rather for the engineering organization as a whole. Part of what I heard there was that I don't have to go to the DMV anymore.
35:22You can just send me my driver's license in the mail. Yeah, basically, yeah, love it. It's a good way to sum it up. How do you guys see in Pat Drive? I don't I'm not, I'm not it's just we have Waymo for that. Speaking of Waymo, how far out do you think we are from having fully autonomous software engineers? Like if you talk about Waymo, like it felt like it was going to come really fast and then it felt like we went through a valley of despair and now it's the future is coming out as super fast again. Like, where, which inning are we in for the kind of fully autonomous software engineer cycle?
35:58and when do you think we'll have fully autonomous Jeff Deans? This is a great question and I think one that we get a lot. I think one thing that's worth is kind of like reframing what a fully autonomous software engineer will do. There have been many moments where technical progress has led to labor dynamic changes and increases in the level of abstraction in which people work. And I think that historically enabling people to operate or impact the construction of software with, you know, at a higher level of abstraction with less domain knowledge has generally led to huge increases in demand for software.
36:43And I think that what we're seeing with the customers we're working with today is that when you free people up from these types of kind of secondary tasks like generating unit tests that map to a pull request or writing and maintaining documentation on a code base that 95 % of people know, but that documentation comes into place with that 5 % that doesn't. They start to shift their attention to higher level thinking. They think about orchestration. They think about architecture. They think about, you know, what does this PR actually trying to do? And less about, did they follow the style guide?
37:21I think that what we're seeing is that this is happening today already because of AI tools and over time as they get better and better we'll see that shift towards Software engineers becoming a little bit more like architects or designers of software And so in the future I think there's gonna be 10 times more people involved in software creation where every individual has the impact to maybe a hundred or a thousand people It just may not look exactly like the individual steps of the development life cycle that we see today. You know, that reminded me of a quote that you guys have on your website, which said, and I'm going to read this, it says, we hope to be a beacon of the coming age of creativity and freedom that on -demand intelligence will unlock.
38:11And that really resonated with me when I read it, because it sort of implies a very positive and optimistic view of the world that we're adding into. I wonder if you guys want to say a couple more words on that or sort of what you think the sort of relationship between man and machine will be in the fullness of time. This kind of goes back to our original approach, which is, you know, it's very tempting to go after the sexiest parts of software development, in particular, you know, like building an app from scratch, right? But that's also the sort of thing that will make a developer defensive because that's the part that they enjoy, right?
38:48And so in a world where you automate the development, then an engineer is just left reviewing, testing, and documenting, which is like a depressing hellscape if you were to ask any, any software engineer, right? So for us, it's very important that we position ourselves aligned with the developer instead of, you know, going into these organizations and being antagonistic with them, right? Like, by going in and automating the things that we don't want to do, or rather by going in there and automating the things that developers don't want to do, we are positioning ourselves with them, right? Five years from now, I don't think anyone really knows what software engineering will be or even if it will be called that anymore, you know, to Eno's point, it might be, you know, you're like a software curator or cultivator or orchestrator.
39:33But by positioning ourselves this way with the developer, wherever that role goes, we will be there side by side to allow them to have this higher leverage. And so, yeah, completely agree to your point. Like, this is one of the most incredible things that is going to happen to, you know, our ability as humans to create. And I think for us, it's just incredibly important that we are aligned with the users of this product and not, you know, antagonistic trying to replace them. How far do you think we are from having these reliable, and maybe call it intern level engineers? Is it a year out? Is it really here today?
40:13Is that decade out? I think it depends on the task first for things like code review and testing. I think we're here. We're already there. We're able to operate at a level that, you know, for many, There's feedback from one organization that we got in particular, where we brought them the ReviewDroid, and this was pretty early on, and they said, you know, the ReviewDroid is the best reviewer on our team. And I think that every once in a while, you kind of hear something like that, and it gives you a lot of confidence that, directionally, we're definitely moving towards something that is valuable.
40:52And for tasks like, hey, we've got to decompose our monorepo into a ton of microservices and the type of thing that you might arm like a staff level engineer and armed with a team of engineers under them. I think that we won't see like a binary moment of, oh, well, now this is done by an AI. I think that their responsibilities will slowly start to get decomposed into the tasks of planning and implementing the refactor, going one file at a time, and when they start handing off those sub -tasks to AI, I think that role will start to be called something different. Because when you're no longer as focused on what is the individual line of code that I'm writing tomorrow and more focused on what is our mission or what is our goal as an engineering team.
41:46You know, you really are more of an architect than less of an implementer. And a concrete example of us eating the food that we're creating, right? We were dreading for months creating a GitLab integration. Some of our customers use GitLab. We want to build Cool AI stuff. We didn't want to spend time building a GitLab integration. We had our code droid fully spec out what the steps of building a GitLab integration would look like. And then it actually implemented every one of these sub -tickets. We were of course monitoring it just to make sure it wasn't breaking anything. And we now have a GitLab integration.
42:22And so this is something that we genuinely were considering getting an intern to do. Because we just, we were just, we really didn't want to do a GitLab integration. and you know, shut up, illab. Yeah. But like materially, the droids saved us hours of time. None of us had built a GitLab integration before. And also, it's just like relatively complicated to abstract away the source code manager. And so that was materially intern work that we did today. So to answer your question, it is now, it's just kind of slowly climbing up more and more the level of complexity of these tasks. You chose to really hear.
43:04It is. I have a question about competition and not specifically the competition in your space, but I think how you more generally think about navigating competition. I think you guys are the type of founders that a lot of companies in the application layer really look up to because you're insanely ambitious building a real company of meaning. And you're doing a lot of smart decisions, like, you know, writing on, writing on other people's models. I think the obvious kind of like scary thing, the scary kind of other side of that is, you know, every other competitor in the space has access to the same models as you.
43:41And so I'm curious how maybe just mentally and then I guess overall you think about approaching competition in the space. You think it's more elevated in this kind of the application layer, AI market and then than another startup market historically, and how do you think about navigating that? Totally. Yeah, I think that's a great question. And I think that's really, you know, our approach to that has defined how we've built our best team. And really, I think there are a lot of ways you can respond to competition and like mentally, kind of justify your existence versus competitors. I think for us, on the team side, we are just a team of people who are more obsessed than anyone else out there.
44:24And I think that is like something that just has compounding benefit of, I am willing to bet everything that the people that we have assembled are just more obsessed than everyone else working in this space. I think kind of a corollary to that is, the only way you can win is by executing faster. There was everything else is all just like sprinkles on top. The only way you can really win is by executing faster and being more obsessed. And that is what our team is. And I think, I guess one last thing is having a group of people who respond to kind of external pressures as like more motivating. And you know, responding in that way, also being very mission driven, right?
45:12Like, if, you know, a competitor does something big and then suddenly you're deflated, well, if you're truly obsessed with a mission, it's irrelevant, right? But if you're truly obsessed with our goal of bringing autonomy to software engineering, all of that is noise. What we need to do is execute as fast as possible in this direction that we've set and the rest will sort itself out. Love it, really well said. Maybe a few final questions to close us out. If you weren't solving the kind of autonomous software engineering problem, what problem would you be solving? I guess I have to be banjo encoding agents for this.
45:46Because if perhaps robotics, I find robotics very interesting. I think a lot of the time of the team here, a lot of the team comes from backgrounds working on autonomy and robotics. And we talk about how what we're building really kind of resembles that in many ways. I think multimodal function calling LLMs are here. And the robotics companies decreased hardware costs that are coming out are clearly making progress. So it feels like a fun area. So you've been making physical droids? Exactly. It's on the road map. Without you, it's on. Yeah. I think this is one of my blind spots where I just suffer from severe tunnel vision.
46:29I genuinely cannot fathom working on anything else. I'm just genuinely obsessed with our mission to bring autonomy to software engineering. If I wasn't working on this, I'd figure out a way to work on this. I know that's a cop -out answer, but I genuinely... it does not compute. So that is in fact a cop out answer, but it is a fantastic cop out answer. So we will take it. One of the questions that I always like to ask is, who do you admire most in the world of AI? And tell you what, Matan, because of your background, we'll let you look at the superset of AI and physics, if you like. I would say, so some name that comes to mind when you say that is Jeff Dean.
47:08I think we mentioned him earlier already actually, but his impact in research is one huge side of that, I think TensorFlow and the work that that whole team has done a deep -mind and related. But I've also heard he's a nice guy. And I think that the thing is, having a responsible leadership in the AI community is really important. And there's a lot of folks who are on Twitter all the time, and you know, clashing and I think that the seeing folks who are outside of that side of it, I think is pretty great. Yeah, and I think not to give you guys a double cop out, but at Factory, we very highly emphasized collaboration and I think like in AI in particular, everything has been done by groups of people.
48:00And so it's hard to really think about one individual. I think physics there are a lot of more like solo genius doing something crazy But I think a team like recently that I think we really admire at factory is Mistral and how kind of quickly they basically Came into open source and brought about those models to basically the cutting edge in a super short amount of time and I think You know, I speak not just for myself, but I think all of our team really admires both the mission that they have and the speed with which they executed on that So yeah, I would say Mr. Will. Awesome. All right, last question.
48:37If you had to offer one piece of advice to founders or would be founders hoping to build an AI, what piece of advice would you offer them? We are in a land of picks and shovels and no one has struck gold yet, clearly. So I would say go for gold. I would say try to build something that you think is going to get 10x better if open AI releases GPT -6 or 7. I think internally we think of our product as something that will multiply in value and uniqueness when new models are released. And I think for us it's always like we're listening to the OpenAI announcement yesterday. And you know everyone is excited.
49:21Everyone's pumped when a new model comes out when open source does something great. If you're stressing about new product releases or demos, it might mean it's worth like like adjusting your product strategy. Congratulations on launching factory and beating state of the art on Sweepench by such a wide margin last week. It's incredible. Just for our audience, can you maybe quickly recap what Sweepench is? Yeah, absolutely. And thank you all credit goes to the factory team for making it happen. Sweepench is a benchmark designed to test an AI systems ability to solve real -world and software engineering tasks.
50:00So it's around 2 ,300 issues which were taken from contributions made to 12 popular open source Python projects. And typically these issues are bug reports or unexpected behavior that people reported on these open source projects. And the idea is all of these real world issues were addressed by other humans. And so you have the ground truth of what a human software engineer would do when faced with an issue. And the benchmark is trying to test can your AI system go through each of these issues and generate a code change that properly addresses it and comparing it to the human solution with tests that a human wrote.
50:48And so there's a lot of asterisks, but it is a somewhat useful approximation of your system's ability to take natural language and then turn that into code. And I think the previous high watermark on sweet bench was 14 % or so from Cognition Devon until last week. And you put up a really impressive new result at 19%, which is such a wide margin. This is such a competitive field right now and such a competitive benchmark that everyone is trying to beat, which makes your results even more impressive. Could you maybe share a little bit about your approach and how do you get there? Definitely. And one of the main reasons we were interested in Sweet Benches is that there's a lot of companies and research labs that made submissions.
51:31You can see Microsoft Research, Amazon, IBM, by dance. And I think that's a testament to the Sweet Bench team's effort in making this benchmark a household name, which is great. I think when a reason we were able to outcompete the kind of like well -funded tech giants and other AI code gen startups is that we are honestly not building the code droid for a benchmark but rather to support real world customers. And we've always said customers are the best benchmark. And I think this is some great evidence for the success of that approach. There's a few areas our technical report goes into around planning and task decomposition, environmental grounding, could base understanding.
52:13But overall, I think that the thing that matters most when your teams working on these types of general software problems is kind of like what is the North Star, what are you iterating against? And so having kind of a real world data set can make a huge difference. And we just had Harrison on the podcast last week actually talking about cognitive architectures. To what extent did prompt engineering and cognitive architectures play a role here in your results? I would characterize our research as continuously pushing the question of how can we model each droids architecture to more closely resemble the human cognitive process that takes place during the task.
52:52It's funny we actually have internally been referring to the flow of data and LLM calls as the joy to architecture, basically since the first droid. And when Harrison first wrote about cognitive architectures, it really became apparent that that concept cognitive architecture is a great mental model for how to characterize the systems with that have complex element interactions and data flow. And so, you know, for us, I think the meta problem of designing a good cognitive architecture is balancing flexibility with rigidity in the actual workflow. You want very rigid entry points and certain comming trajectories like air recovery need to be really consistent.
53:39But then you want the flexibility in the dynamics during the majority of the problem solving process. And so it's a challenging balance, but I think it's one of the most interesting problems when building the is how do you know, kind of when to add structure and when to let the droids sort of speak? Handle it. Really cool. So every joy that's done cognitive architecture or that mirrors as closely as possible, with the kind of human equivalence of that task we could be doing. Yeah, exactly. 19 % is amazing compared to piracy of the art. It also still feels quite far away from reliable code droids that people will just trust to run wild in their code base.
54:24What do you think is the threshold at which engineers will actually start to use these code droids reliably and just let them run? Are we there yet? Or what is the threshold? Yeah, for sure. And I think that one thing to keep in mind is that the percentage on a benchmark like Sweepench is kind of like one of many possible measures. because the answer they really is that they are already using it in production. But the use cases that might kind of highlight what the code droids design for may not necessarily have a ton of overlap with what is tested in a given benchmark. So if you take like human e -vow or some of the other coding benchmarks, that maybe tests your ability to pass a coding interview, but it doesn't really test like your real world software engineering.
55:14Sweetbench, I think actually does test a lot of real -world software engineering, but in the particular context of debugging or kind of unexpected behavior identification. There are some feature requests, and there are a lot of kind of not explicitly debugging style problems, but tasks like a migration, a refactor, a modernization that take place over multiple changes that oftentimes have humans very heavily collaborating. are really a pretty different problem and our internal evaluations are much more focused on those customer tasks. And so we have way higher reliability rates for those style of tasks.
55:57And I also think that a huge part of the role of human AI interaction design is acknowledging where the systems are currently falling short. and building into your interaction pattern, accommodation for the weak points of the AI system. This isn't going to 100 % of the time, perfectly capture the intent of what you were doing. So how do you kind of have failure trajectory handling? How do you introduce the ability to kind of edit midway as the code droid is working to observe and have kind of some interpretability into why a code droid is making a decision so that what it does something, the human being actually can step in or at least understand what went wrong.
56:42And so I think that those allow you to say, well, we may not be at 100 % on someone like Sweepench, but we can still use this and get kind of real productive gains in the meantime. Totally makes sense. And I hear you that Sweepench is not the BL end all, but since you have a good crystal fall into this space. Do you have a prediction at what point will get to 80 % or 90 % on sweet bench? I think that the pace right now is really, really fast. There's a kind of like interesting question of will we get to 80 % to 90 % on sweet bench or will there be a better benchmark that kind of comes out before we can really meaningfully start hill climbing past like the 50 -60%.
57:29There's honestly a lot of tasks in Sweepench, which are... I wouldn't say impossible, but it almost feels like getting them right would almost only indicate that you're cheating. It's like they test for really, really specific claims or a string match. And so I think that before we see 80 -90 % on Sweepench, what we'll actually see is kind of like Sweepench 2 and we bench three that focuses on trying to think deeply about how can we evaluate when a piece of code that is correct but also ideal or useful for a given code base. The sweet bench folks actually have a lot of really great thoughts about how to make these bench works better, but I think probably the next two, three years we'll see that.
58:21Yeah, and they're printing guys as well, right? Yeah, Yeah, they are. We actually shared a thesis advisor. Oh, no way. That's very cool. Well, you know, Matan, thank you so much for the conversation. Congratulations again on these results. And on launching factory, we are so excited. Thank you. Thank you very much.
From the publisher
Archimedes said that with a large enough lever, you can move the world. For decades, software engineering has been that lever. And now, AI is compounding that lever. How will we use AI to apply 100 or 1000x leverage to the greatest lever to move the world?
Matan Grinberg and Eno Reyes, co-founders of Factory, have chosen to do things differently than many of their peers in this white-hot space. They sell a fleet of “Droids,” purpose-built dev agents which accomplish different tasks in the software development lifecycle (like code review, testing, pull requests or writing code). Rather than training their own foundation model, their approach is to build something useful for engineering orgs today on top of the rapidly improving models, aligning with the developer and evolving with them.
Matan and Eno are optimistic about the effects of autonomy in software development and on building a company in the application layer. Their advice to founders, “The only way you can win is by executing faster and being more obsessed.”
Hosted by: Sonya Huang and Pat Grady, Sequoia Capital
Mentioned:
Juan Maldacena, Institute for Advanced Study, string theorist that Matan cold called as an undergrad
SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering, small-model open-source software engineering agent
SWE-bench: Can Language Models Resolve Real-World GitHub Issues?, an evaluation framework for GitHub issues
Monte Carlo tree search, a 2006 algorithm for solving decision making in games (and used in AlphaGo)
Language agent tree search, a framework for LLM planning, acting and reasoning
The Bitter Lesson, Rich Sutton’s essay on scaling in search and learning
Code churn, time to merge, cycle time, metrics Factory thinks are important to eng orgs
Transcript: https://www.sequoiacap.com/podcast/training-data-factory/
00:00 Introduction
01:36 Personal backgrounds
10:54 The compound lever
12:41 What is Factory?
16:29 Cognitive architectures
21:13 800 engineers at OpenAI are working on my margins
24:00 Jeff Dean doesn't understand your code base
25:40 Individual dev productivity vs system-wide optimization
30:04 Results: Factory in action
32:54 Learnings along the way
35:36 Fully autonomous Jeff Deans
37:56 Beacons of the upcoming age
40:04 How far are we?
43:02 Competition
45:32 Lightning round
49:34 Bonus round: Factory's SWE-bench results




