In short
Factory’s co-founder Matan Grinberg discusses “dark factory” software factories—autonomous agents that build software—plus why enterprises need model independence, modularity, and cost-aware routing (token rationalization), not just better models.
Guests
Matan Grinberg, co-founder and CEO of Factory (builds droids/autonomous agents for software development). No other guest is identified in the transcript.
Key claims
- “Customer obsession” is an input metric; Factory’s goal is “create obsessed customers” by building outputs that developers actually want.
- Enterprises avoid vendor lock-in and single points of failure; Factory emphasizes model independence and hot-swappable models.
- The biggest shift wasn’t model quality alone; it was developer/enterprise behavioral readiness and better agent-native tooling.
- Performance comes from the harness (caching, compaction, tool validation), and multi-model harnesses avoid overfitting to one model’s quirks.
- Token maxing (wasteful “use AI for everything”) will evolve into routing and eventual outcome-based pricing.
Notable examples
- Factory refunded revenue when developers weren’t satisfied, despite strong sales.
- Droid CLI launched Sept 26, 2025; Terminal Bench benchmark focus.
- Examples of token waste: banks asking trivial questions like “what’s the weather” using frontier models.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOCustomer Obsession vs. Output Metrics
0:00 to 0:58
Explore the distinction between being customer obsessed and delivering actual results.
“Bezos at Amazon, it's customer obsession.”
Factory's Competitive Landscape
1:10 to 2:22
Matan discusses Factory's emergence in the software development market.
“And Matan, we're going to jump right in because I think you guys are a little bit of a dark horse candidate in this world of software development.”
Lessons from the 'Desert Years'
2:22 to 4:28
Matan reflects on the challenges faced during Factory's early years and the learning process.
“They do not want anyone to kind of control their fate.”
Navigating Customer Relationships
4:28 to 5:54
Discussion on managing customer expectations and the importance of trust in business.
“So it's a really good question because that is something that you might think of like, okay, wait, so we're just switching the lock-in point.”
Decisions in Difficult Times
5:54 to 8:00
Matan shares the tough decision to refund customers and its impact on the team.
“Can you just talk about that journey and what it has done to the DNA of your company?”
Building Resilience Through Struggles
8:00 to 14:00
How facing adversity strengthened Factory's team and its mission.
“But we realized that the way that we had sold them on it and the product that we were delivering was not up to snuff in a way that I don't think it would hold true to one of our operating principles.”
Evolving Developer Interactions with AI Tools
14:00 to 16:45
Learn how the development landscape is changing with the introduction of AI tools like Droid CLI.
“So one is the interaction pattern that we were building for before was too ambitious.”
Benchmarking and Performance Metrics in AI
16:45 to 19:38
Discover the importance of benchmarks like Terminal Bench in evaluating AI performance.
“There's like the benchmarks, which have a very short half-life.”
Harnessing AI Models: Co-Design vs. Multi-Model Approach
19:38 to 23:26
Explore the debate between building AI harnesses for specific models versus a more general approach.
“What's the, like my intuition would be model harness co-design makes you better.”
The Shift from Token Maxing to Cost Rationalization
23:26 to 28:06
Understand the phases of AI adoption in enterprises and the implications for token usage.
“Okay, so we talked about one type of maxing, benchmark maxing.”
Show all 18 chapters
The Rise of Open Models in Software Engineering
28:06 to 30:29
Learn about the evolution of open models and their impact on software engineering.
“kind of the frontier models will be frontier the question is are the open models getting as good as like frontier minus one?”
Shifting Business Models in Tech
30:30 to 32:44
Explore the future of business models, focusing on usage-based and outcome-based pricing.
“Like you've become an expert in your craft and you used to spend hours writing docs.”
The Concept of Software Factories
32:45 to 35:38
Discuss the inefficiencies in current software development processes and the vision for software factories.
“How do you think about business model and pricing?”
Optimizing Business Operations with Software
35:39 to 42:01
Understand how businesses can balance token usage and human resources for optimal operations.
“It's like, you know, the solution is just get rid of it all.”
The Importance of Core Competencies
42:01 to 43:42
Learn why businesses should focus on their core strengths and outsource non-core functions.
“And I think this is an opportunity for every business to double down on their core competency and what matters for them.”
Reinventing Business Practices
43:43 to 45:14
Discover how companies can successfully innovate through internal initiatives like hackathons.
“So, you know, every company kind of has to go through this process of reinvention.”
The Future of AI in Business
45:15 to 47:51
Explore predictions for AI's growth and its implications for future business practices.
“Do you have any predictions for the most important changes that are going to happen in your space over the next, call it, 12 months?”
Challenges and Optimism in AI Development
47:52 to 50:37
Understand the challenges facing AI adoption and the potential for engineers to solve major global issues.
“So I think short term, there's going to be a lot of turbulence because I think a lot of companies have misallocated resources pretty poorly.”
Transcript
Automatic transcript. May contain errors.0:00Matan Grinberg:Bezos at Amazon, it's customer obsession. But in our mind, that's an input metric. You don't want to measure input metrics. It doesn't matter if you're customer obsessed. You could be customer obsessed and they file a restraining order against you because they don't like what it is that you're doing. Our job is to build something so good that our customers themselves become obsessed with us. That is our job. The analogy is, if you're a coach of a basketball team, you don't want to tell your players before they come out there like, hey guys, make sure to sweat. It's like, what? No, score points.
0:29Matan Grinberg:Like we need to score points and in doing so, yeah, you're probably gonna sweat. And I think similarly to create obsessed customers, you probably need to be really obsessed yourself with the customers. But the output is what matters.
0:57we're here in the studio with matan from factory this is our second time with the time thanks for having me you're in the small and elite group of second time training data attendees so thank you Oh, yeah. Matan is the co-founder and CEO of Factory, which makes droids, which are autonomous agents for the art of software development. Yes, indeed. And Matan, we're going to jump right in because I think you guys are a little bit of a dark horse candidate in this world of software development. It is a market that has absolutely taken off. There are folks like Cloud Code and Cognition and others who have a lead, but you guys are coming up strong.
1:33Talk about the competitive dynamics and what makes Factory special.
1:36Matan Grinberg:It's been a wild ride. we started factory three and a half years ago now. So in April of 2023, when the world and the enterprise in particular was barely ready for GitHub copilot, let alone fully autonomous agents. And so I think the first two years, it was kind of our journey in the desert is how I like to refer to it, because we were focused on fully autonomous agents, but engineers weren't ready, procurement teams at the enterprise weren't ready. And so I think, retrospectively, we really like honed our craft and learned a lot about how to build for developers in the enterprise. But you know, it took a lot of time to actually come around to when they were ready to receive it.
2:21Matan Grinberg:And so we're kind of now emerging much more and some of these other players like Anthropic or OpenAI who have a ton of distribution are going in and bringing their incredible tools like CloudCode or Codex, the thing that enterprises are really caring about that we have learned through those two years is they do not want anyone to kind of be their single point of failure. They do not want anyone to kind of control their fate. And so something that really matters is model independence. Everyone learned from cloud, where, you know, back in the cloud days, it was like AWS or Azure being like, hey, you know, come on in, sign this three-year contract.
3:00Matan Grinberg:It's going be so cheap we're going to subsidize it it'll be great and then a couple years later when it came time to renewal they would 10x the the the contract haha data gravity we got you now yeah we got you what are you going to do a two-year migration to go to someone else like no way everyone has scars from that now and so everyone knows look Claude Code is fantastic uh Codex from OpenAI is fantastic we cannot put our fate in any one of these model providers hands also like you just look at the risk profiles of the model labs versus the cloud providers. What's the last piece of drama that came out of a one of the cloud providers versus like the model labs, it seems like there's kind of always some sort of chaos of, you know, internal fighting or getting in spats with the government or, you know, any other entities.
3:44Matan Grinberg:And so if you're going to, you know, build this very important part of your business, you want to make sure that you're robust to any of these changes. And that's something that we've learned over those kind of initial two years is like developers really care about things being modular they want to know that they can customize it to what they want they want to know that if there's a new model that comes out that's faster or cheaper or more performant they can kind of hot swap it in and that's i think one of the biggest reasons why a lot of the largest enterprises are taking the momentum that they had from a codex or cloud code and then are carrying that into factory because they get that performance from these fantastic models but they do it without the vendor lock-in that, you know, the Model Labs directly provide.
4:26And if I'm the enterprise, I'm going to be like, wait a minute, am I now just getting locked into factory? What's the answer to that?
4:31Matan Grinberg:So it's a really good question because that is something that you might think of like, okay, wait, so we're just switching the lock-in point. All of the modularity that we build is such that if at some point you wanted to say, hey, you know what, factory is not staying at the frontier anymore, whether it's like the automations that you build or the skills registry that we help you create, the work that we've done stays in your code base. and any of the automations that we've created, the artifacts also live in your code base. In other words, there aren't really things that we're saying like our tribal knowledge about your org that we're keeping on our side and not giving to you.
5:05Matan Grinberg:And that's part of the relationship that we have with customers is like, we similarly want to make sure we're providing the best experience possible. If we help you arbitrage between different models to get cost optimization, we're giving you that optimization. We're not taking that away from you. And I think that's a really important part of the trust that we're building with these enterprises. You and I were talking probably a couple months ago at this point, and I was trying to give you credit for having the right vision for this market two, three years ago. And you responded with something along the lines of, thank you, but being two or three years early is the same as being wrong.
5:40Yes. Which I thought was a wonderful response in so many ways. Can you talk about those two years in the desert? How did it feel to have this vision that turned out to be right, that nobody appreciated for a year or two? Can you just talk about that journey and what it has done to the DNA of your company?
5:58Matan Grinberg:Yeah. I mean, in the moment, it's really, really difficult because I hadn't had a job before. I dropped out of my PhD to start this company. And over the course of those two years, convinced 20 of the smartest people that I've ever met to quit what it was that they were doing and join Factory and join us on this mission. Um, and these are people with families. These are people with kids who are like dedicating years of their lives to this problem and going, you know, customer after customer. And they like, they weren't ready for agents. They didn't get it. Also, the models weren't as performance, but I think a lot of it was behavioral.
6:36Matan Grinberg:And I mean, even just a fun anecdote of like giving developers an NPS survey. If you ever are giving a developer an NPS survey, they do not like whatever it is that you're giving it to them. because like developers they vote with their feet they are very clear what they like and what they don't like and if you're like hmm i wonder if they like it they definitely don't um and uh but during that time i think there were there was a lot that we were learning there was a lot that i myself was like i'd never had a job before enterprise sales is not something that comes obvious to to a physicist um but at the end of the day it doesn't it doesn't matter there's no you don't get any you know bonus points for being early because like who cares like there's no consolation prize it's either you do the thing or you don't do the thing and that's all that matters and for the team is really tough there were points where uh we ended up getting good at enterprise sales but the product still wasn't good and that's a very tricky position to be in because we ended up you know getting to a point where we were like just under two million in revenue and the product was not good and there was a point in time where we realized this because if you're really good at sales you can sign contracts that's like you can definitely do that but if you're doing that and the developers don't like your product it's like a ticking time bomb because it's eventually they're going to churn and it's going to be really really bad we realized this and we proactively gave all of those customers their money back and i remember that uh was one of the most difficult decisions to make because not only is there a you know a group of you know 20 people who are getting ridiculous offers from all the labs they have these huge you know financial incentives to go elsewhere there are all these other companies that are doing well and they decided to do this and then we're going to say oh yeah hey by the way that you know a little bit of revenue we managed to get we're actually going to give it back because we don't think product is making their developers happy we also had to make that decision you know we sold them on a good vision and convince them that, you know, this is the right team to work with and that we were going to deliver the solution for them.
8:39Matan Grinberg:But we realized that the way that we had sold them on it and the product that we were delivering was not up to snuff in a way that I don't think it would hold true to one of our operating principles. And one of our operating principles that I really like is create obsessed customers. Yeah. This kind of like flips over Bezos's thing where Bezos at Amazon, it's customer obsession. But in our mind, that's an input metric. And input metrics are like you don't want to measure input metrics it doesn't matter if you're customer obsessed like you could be customer obsessed and they file a restraining order against you because they don't like what it is that you're doing like our job is to build something so good that our customers themselves become obsessed with us that is our job it's like you know the analogy is if you're a coach of a basketball team you don't want to tell your players before they come out there like hey guys make sure to sweat it's like what like no like score points like we need to score points and in doing so yeah you're probably going to sweat and I think similarly to create obsessed customers you probably need to be really obsessed yourself with the customers but the output is what matters and I think that it coming back to this the product that we were delivering was not creating obsessed customers and we wanted to make sure like this was a group of the smartest people I've ever met we were getting there like we were getting a lot of intuition things were starting to come together internally.
9:56Matan Grinberg:Like we could see internally, we were starting to become a lot more agent native in how we were doing things. And the product was kind of scratching that itch. But we were kind of ahead of our customers. And we wanted to maintain trust with our customers, so that when it does hit, we can come back to them and say, Hey, guys, this is the real deal, I promise. And to build that credibility, we have to say, Hey, look, you know, even though you were maybe happy to continue, we're going to give you this back and say, three months from now, I think it'll be ready. Give us some time. And I promise we will knock your socks off.
10:27How did your customers react when you had that conversation?
10:30Matan Grinberg:Some of them were like, oh, great. Sounds good. Because I think it wasn't something that they were obsessed with. Some of them were a little bit confused. But I think generally, it's especially enterprises. They're not used to these things. A lot of times enterprise budget, once it's gone, it's gone and no one really cares. Yeah. And so some of them didn't even know if they had a mechanism by which to take back the money. um but you know it's a difficult thing to tell also like investors who believe in you like you know i remember having the conversation with sean um i think sean obviously he's stayed really close with the company so he was very like on the same page but it's kind of a scary thing to be like hey by the way you know remember all those updates and you're saying hey look the you know revenue is going up it's about to go down to zero um it was a scary thing uh and i think it was kind of a leap of faith of like we see the signal internally early of like this is the direction we need to go we need to kind of pivot the approach on the product but i remember that all hands where we told the whole team it was like oh my god that was like one of the worst months of my life like i was just because no like not everyone was going to say like what the hell is this what's going on but it's kind of the looks on their faces where they kind of go a little bit pale and they're like oh boy like is this just the early signs and we're about to sink completely um how'd you keep the team together through that i think honestly the only reason the team stayed together is we were so ruthless about hiring early on where it was like people that are genuinely really really obsessed with the mission which our mission is to bring autonomy to software engineering and like really really caring about that making sure everyone um was also like very clear feedback loops as to like this the fate is in our hands it's not like this is like oh something that I go do.
12:14Matan Grinberg:It's like we all have a part to play in making this work. And I think embracing how much it sucked was also, I think, something that was very valuable. Just being honest about it. Being super honest about like, yeah, this sucks. Like, oh, look, look at those competitors. The revenue is going up like crazy. Like, this is not good. Like, we are in a very bad position. Like, we just had to give back all of our revenue. Like, we need to really get our shit together. And in the moment, I think retrospectively, those are the moments where really the deepest bonds are made. Like, if you talk to people who are like athletes or even like academics or whatever whenever you're in the like stressful period whether it's like cramming before finals or you know in intense like you know we have some some rowers on our team and i think that's an example we always go to like that's a pure pain sport it's pain it's literally just there is one number that quantifies your performance it's just what is your time on your 2k are you timing that um but like embracing that is what creates those enduring bonds such that afterwards, like, we know what it's like to be at rock bottom, we know what it's like to lose.
13:16Matan Grinberg:We know what it's like. I mean, when we first started the company, our valuation was 5 million. Like a lot of our competitors, a lot of the companies out there these days, they don't know what it's like to not be a unicorn. That's like manifestly, that is what they are day one. Whereas like we have been there kind of in those dark moments and not a single person left. Yeah. That makes us so resilient and so strong that, you know, going forward, things are going a lot better now. But there are going to be really bad times. But we have that resiliency in our DNA that I'm not sure some of these other companies do.
13:48I love that. So talk to us about what changed. And I'm curious your comment from earlier that the models getting better is not the most important thing that happens. Because at least in my mind, the models getting better is the most important thing that happens. Yes. Help me understand. Yeah.
14:01Matan Grinberg:So a couple of things. So one is the interaction pattern that we were building for before was too ambitious. Like to your point, we were right in that what we were building for was fully autonomous agents, but it was two years too early, which makes it wrong. And fully autonomous agents require a complete change in behavior from the developer. And we were trying to do that out of the box before they were even using tools like CoPilot. It was just too much of a leap. It was too much of a step function jump. So it's an important day. September 26, 2025 was when we first put out basically the Droid CLI.
14:35Matan Grinberg:And the Droid CLI met developers where they were in a manner that previously these fully autonomous agents did not. And also its performance was like completely state of the art and it was model agnostic. So it could use every model that was out there. September 26 was also two years after we initially started. So the world had gotten much more used to using things like autocomplete. Like by late 2025, most engineers were using an autocomplete tool. And many were starting to, at the time, use like a chat interface to ask an agent to go do changes like wholesale. So like the more agentic interaction.
15:11Matan Grinberg:However, what we see is that like, if you go back now and use in this like agentic interaction, some of these older models, they're still good. So the biggest thing that changed was developers and in particular in the enterprise, like being open-minded to this new way of working. In particular, you know, developers, they've established their workflows over the last 30 years. They can be stubborn. A lot of them were like, no, no, no. Like my craft could never be done by an AI tool. So a lot of it was just like understanding how to work with these tools and having the willingness to go in and try and also the intuition about what are the guardrails that you need to provide in order for it to succeed.
15:51Matan Grinberg:And so I think it was a combination of both of these things. The model is getting better, so you need to do less in the way of providing guardrails, but also developers lowering their guard and being like, okay, you know what, let me go try and do these things. It's going to go do things I don't like. And then also, there's a certain degree to which when Andre Karpathy tweets about something, then every engineer suddenly is like, okay, you know, maybe this is true. And Andre started to tweet about these agentic work. early on he wasn't as open to it and then him being more open to it genuinely just changed some people's minds which is funny but that's some of the things that go into behavior changes like you hear it from people you trust you start seeing it you know from people within your organization who are maybe a little bit more agent native but that's that's kind of these things together is what what changed that and now we're all gonna be on slack we might we might be on we might be pushing the limits of slack which i think is going to be another interesting thing but yeah okay so So September 2025, you launched the Droid CLI.
16:47You said Frontier Performance Soda. What does that mean for you?
16:50Matan Grinberg:There's like the benchmarks, which have a very short half-life. Like anytime there's a good benchmark, it gets benchmarked within like three to six months. At the time, I think the one that we kind of championed when we launched and kind of it ended up becoming a pretty good benchmark was Terminal Bench. so prior to that the one that was kind of leading was sweet bench which was um kind of took some open source projects and some examples of issues that were then solved the problem with that was it was very focused on like python and like scripting or like individual file changes whereas terminal bench was more one it was in the terminal setting so it was things like scheduling runs and things that were not just like changing the code file but general software development tasks.
17:34Matan Grinberg:And that was something that we ended up, you know, having really frontier performance on. Now it's like, benchmarked to the extreme to where it's like, I think, you know, models that come out now are like 90 % on it. And I think there's a very short time horizon from putting out a good benchmark to then it being kind of in the training data. What goes into building a great and is it a great harness? And it seems like there's almost a lot of FUD in the ecosystem of my harness is better than your harness. And, you know, you need to own the model to have a good harness or actually you have a better harness if you don't own the model like yeah what's your mental model for for yeah you know benchmark maxing aside what keeps you at the frontier yeah so a couple so some general things that matter are um the way you do caching so you know cash tokens end up being like a tenth as expensive and so one big piece of performance for a given harness is what what is your like rate of of token caching um another example would be how do you perform while in compression or compaction?
18:32Matan Grinberg:So typically when you're dealing with a long session, you're going to exceed the context limit of the model itself. And so the harness will do some sort of, you know, summarization, compression, compaction, whatever you want to call it. And the way that you perform during that compaction is a big determining factor of how good your harness is. And tests that they do for that are like, you know, they call it needle in the haystack where you have some long thread and maybe there's one piece of information that's really important how often will your harness preserve that through compaction other examples are like tool use or how does it use the environment to validate whatever work that it's doing these are things that you can kind of have individual metrics on and that we kind of have our own internal benchmarks to measure how do the out of out of the box agents do versus how does factory perform i think one thing that um naively everyone believed initially was if you train the model and you build the harness, you're going to make them better together.
19:26Matan Grinberg:Yeah. And much to the chagrin of many of my friends at OpenAI and Anthropic, this is not true. If you build a harness that supports different models, that harness will be better. What's the, like my intuition would be model harness co-design makes you better. Yes. What's the intuition for why it's actually not? It's very analogous to the idea maybe like, I don't know 10 years ago of if you were to be like hey I want to train my personal AI back in like like ML days before like GPT-3 I want to train my personal AI I'm going to give it all of my data because I want it to know me turns out the answer was train it on the whole internet and it'll be so much better for you than if it were just trained on your data so there's a sort of analog that emerges where it's what data is to a model models are to a harness where the more models you expose to a harness, you avoid overfitting that harness to the nuances of that model in particular.
20:21Matan Grinberg:And there are certain intricacies about different models that you can learn from and then improve different models performance in your own harness. And this was why, for example, we kind of stopped doing it because Terminal Bench got so bench maxed. But initially, when like every new Opus or GPT model would come out, it would perform better on Terminal Bench in droid than it would in clod code or codex um which is why like and this is something that you know i think was somewhat frustrating to because i deal like from a lab perspective you ideally want it so that it's better together because then that means you have to use their harness and you can't use a different one but i think the reality is it it's uh you know having that multi-model harness ends up getting kind of frontier on on all those is there a good like example or illustration of that conceptually it makes sense is there like an easy way to Maybe maybe a good example of it is like if you're familiar with the different behaviors of Opus and GPT 5.6 right now.
21:19Matan Grinberg:I am he's okay opus tends to be I mean loosely loosely I mean to be fair honestly these days I'm not doing it as much either But I will say this loosely opus is kind of like that Super friendly colleague where you're like, hey, I want to go do these 20 tasks and they're like, okay, cool hey by the way five of those tasks I realized we didn't need to do it don't worry about it I got other these done did it this way tonight's not a good time let's pick it up in the morning yeah like let's go let's go get a beer afterwards and hang out whatever meanwhile like gbt 5.6 is like absolutely I will do every single one of those and nothing will stop me I'm not going to sleep until those it's like kind of very OCD and you know meticulous but sometimes you know you want one where it's like it actually realizes hey that list of 20 that you gave me actually here's a better way of doing it anyway.
22:06Matan Grinberg:You know, 5.6 is more methodical. If you build a harness for each of those, there are actually different things that that harness will then be good or bad at. So for example, one thing that, you know, typically agents will do is they'll, they'll have a to do list of like, if you have a task, it'll go and generate a to do list. And the Claude code harness can in some cases or and this is maybe less relevant now. But I think earlier this is a just a more illustrative example earlier it was really strict to make sure it would stick to the to-do list because the model itself would typically wander meanwhile codex wouldn't do that because the model itself was really really ocd about that but if you're a user you want to have the same experience regardless like you want to make sure if you switch to a different model you're not going to suddenly lose track of whatever things that you are working on and so there are certain things where like maybe in some cases you really want robust tool use and there are tools that you use to do these to-do lists you want really robust tool use and you want to make sure that no matter what if i'm a user i want to see my to-do list there like there were some cases where it would just like not have the to-do list and so that these are things that kind of improve the general performance and that the to-do list matters because you're doing some crazy migration and you don't have the to-do list and then you're in this long session where there's compaction that might get lost in the summarization.
23:23Matan Grinberg:And now you forgot what your seventh step was. And that could be one of the failure modes. That's kind of an example of how... That's a good example. Yeah. That's a great example. Okay, so we talked about one type of maxing, benchmark maxing. Let's talk about token maxing. Yes. Because it feels like the world has changed a lot. We've gone from token maxing to now cost rationalization. What does that mean for factory? Yeah. So maybe I'll lay this out just so we're all on the same page of the way that we see what's led us to this token maxing. So loosely, there was like this phase one where maybe phase zero was like no one believed in AI.
23:54Matan Grinberg:Then phase one, everyone believes in AI. And then boards were like, Mr. CEO, what are you doing about AI? What's your AI strategy? And Mr. CEO is like, shit, I don't know. Like, what's our AI strategy? CTO, like, make sure everyone goes and uses AI. And so then phase two is, you know, CTO is like, okay, shit, we got to make sure everyone uses AI. Let's start putting it in performance reviews. Let's make public like benchmark or public like rankings of who's using tokens the most because everyone's stubborn. No one wants to use this stuff. They're all skeptical. And then we enter phase three, which is everyone sees these ratings.
24:26Matan Grinberg:They see that it's part of their perf reviews and they're like, OK, I'm going to use AI for everything. And that's kind of phase three. It's this token maxing where people are using like Opus for literally everything. Like what's the weather NSF? Opus. Tell me. I don't know. Like there are banks that we are working with where they are spending literally hundreds of thousands of dollars a month on people asking things like literally what is the weather? Or like tell me about Python. Like trivial questions that you could Google, people are asking Opus. And the reality is this happened because we were so worried about adoption that we overcorrected and we're like adoption by any means necessary.
25:03Matan Grinberg:And I think that's actually it's like a decent approach. like it's probably faster to do that and then curb usage or make usage more responsible than it is to start limited and be like you know you can only use it for this thing because when you have people that are stubborn first you want to just prove that it works and then you can get kind of more mature about it where factory fits in i think one of the most important things that we do is that we have the factory router which allows you to dynamically route to different models based on the task that you're doing so you know if you're asking what the weather is you probably don't need the very frontier of human intelligence to answer that for you.
25:38Or you really do.
25:39Matan Grinberg:I mean, it depends. I don't know. It depends on what kind of answer you're looking for. You know, giving you like a full like down to the like molecular level of what's happening. But, you know, allowing that. But also more importantly, for every enterprise, something that no one's dealing with yet, but 12 months from now is going to be the case is not everyone needs the same tokens. having a blanket kind of token cap for every individual in some large bank let's say makes no sense so every cio is going to need to answer for every incremental token where do we put it and right now it is super not obvious how you would do that like right now we're saying oh you know the pms who are like vibe coding dashboards get the same token limits as like the engineers who are building like critical infrastructure that's probably not the best thing to do or similarly, you might be dealing with COBOL codebases where Opus is not the best model to use, but instead maybe some fine-tuned model on that codebase in particular.
26:36Matan Grinberg:The point of the router is that we can kind of accommodate these different constraints where maybe you say, you know what, this part of the org, they're just vibe coding. They can use Gemini Flash. This part of the org, they're doing COBOL. We fine-tuned this great model to work on COBOL. Let's route to that when we're working on that part of the codebase. Maybe this other part, we really care about reliability. So let's generate the code with open AI, test it with Anthropic, review it with like Gemini, things like that. And we can actually take in your routing procedure instructions in natural language.
Read the full transcript
27:08Matan Grinberg:So you could even say things like it's not purely deterministic. It can even be like, hey, you know, Pat, I don't know, like, I don't know what he's doing. Like, give him flash. Or, you know, I think we really need to avoid having them use open models because, you know whatever reason we don't like the way open models perform here and we'll do internal benchmarking to know which models are better at which of these tasks how close are the open models at this point which one's the best glm 5.2 is incredible it's at the point where internally we have no token limits for our engineers and like half of our tokens are open to open models wow yeah because they're just faster and they're cheaper they're just as performant and i think the thing that everyone gets wrong is everyone is comparing like glm 5.2 to the latest model like opus 4.8 or gpt 5.6 but really they should be compared to opus 4.7 or gpt 5.5 um why because generally the open models come later and they're they're kind of a generation behind and that's kind of the frontier models will be frontier the question is are the open models getting as good as like frontier minus one?
28:17Matan Grinberg:And the answer is unequivocally yes, which I think is a really, really interesting outcome. It's great for consumers. And by consumers, I don't mean like individuals. I mean the consumers of the APIs, because if you're a business that is doing software engineering, your job is at a very high level to solve problems. And if we can allow you to solve those problems faster and with cheaper models that are just as performant, that means you can solve more problems like that is a good thing and it is a very good world where there is not like a monopoly on intelligence but instead kind of a garden of intelligence that you can pick and choose um you know when you'd like something that we joke about is like you know on this intelligence allocation thing um if you're if you're trying to get a tutor for your daughter in algebra you can probably find someone cheaper than albert einstein to be that tutor now it might be that she eventually goes and becomes like a leading you know physicist or something in which case yeah maybe let's let's get albert einstein in there but most likely you can get you know a high school student or something like that um and it's probably much more cost effective for you as well to do so so since you guys do the model wrapping like if you look at the you know if there's a pie chart that shows the complexion of models being used by your customer base today what did it look like a few months ago what does it look like today what do you think it'll look like in a year yeah i will caveat this with saying that right now enterprises haven't gone too opinionated yet into the routing procedures okay this is something that will happen over the next six to 12 months but right now they're just going from no router to router that's kind of the first change then it's going to be like the exact nature of the of the routing at the beginning of the year there's less than one percent of tokens went to open models in the first quarter it became a single digit percent it is now crossed into being a double digit percent of tokens um now percent of tokens is not always the same as percent of cost because the open tokens are cheaper um but it is uh it's pretty crazy to see the the growth there what's your forecast my sense is that we will asymptote towards vast majority being open just because it provides you more optionality and it's cheaper um but that doesn't mean they're going to be like that's of token share not necessarily of leverage share because maybe there are one percent of tokens that are incredibly incredibly valuable um and are like very key decision making and then the rest are more like implementation tokens or kind of uh lower stakes if you will i don't think there's going to be a world in which like it's ever going to be 100 i think the frontier of intelligence will inherently always be valuable for every business just because the stakes are going to get higher and the kind of intricacy with which you think is going to be more important but we'll be better at offloading certain tasks and this is like you can loosely think of this uh already with the way orgs are structured where you know in general engineering leaders are more tenured engineers who in theory have like more wisdom and each kind of minute of their brain power is higher leverage in theory um and even you know you can also imagine like consider a human engineer and try mapping over the course of their day like how much brain power they're using and like you know it's probably going to be really low for a lot of it but then there are going to be some moments where they're like going pretty high like they're deeply concentrating and thinking about some you know systems design problem or whatever all of those low leverage moments we want to automate away and like we want to like those like very high leverage moments sometimes like you know we're referring to them as like the eureka moments or the moments where they're like doing something that's very high leverage what if those aren't just moments but what if those are like hours at a time because you don't have to deal with all the other stuff and i think that's kind of the the way to think about intelligence allocation is if you're an engineer and you're writing docs, that is such a low leverage use of your time.
32:09Matan Grinberg:Like you've become an expert in your craft and you used to spend hours writing docs. Like I remember it was actually valuable. Like I remember Stripe had so much alpha for just having incredible docs, but imagine all the other stuff those incredible engineers could do if it wasn't writing documentation. Like we should live in a world where everyone can have docs as good as Stripe. And that is like strictly beneficial for everyone. And then the question is, okay, what do those really smart engineers do with their time once they don't have to do that? Maybe this is a good time to talk about business model, given that, you know, especially the rise of open weight models, the cost differential.
32:44I imagine that means very different things for your cost structure, but very similar value delivered to customers. How do you think about business model and pricing?
32:53Matan Grinberg:Yeah, this is more what our customers want and need as opposed to what we want and need. So, for example, I think right now usage-based is clearly the way to go. we want to be aligned with like what they are doing and what we are doing i think seat-based doesn't make sense at least for what we are doing my sense is that eventually we will change to outcome-based now i don't think the enterprise is ready for that and we've learned our lesson from those first two years we are not going to impose things right um but my suspicion is that you know in the 2030s things will probably look more like outcome-based yeah what does outcome based mean for your market will be the definition of an outcome so maybe here's a way to put so right now we chart we are usage based like the more tokens you use you know the more you pay the more we get um now since we are model independent we kind of with our router we are kind of pointing a token cannon at either open ai anthropic aws gcp you know any one of these people uh to a certain degree this is like a really dumbed down version of a marketplace where right now there is a a buy the buy side is an engineer who wants a task done and then you have the model providers who are saying like either in benchmarks right now they're like we perform at this cost and this performance and then we determine who we go to for that given task yeah there's a world in which you know if it's so important to get these tokens, they might kind of bid in a certain way of saying like, look, here is our cost for this task.
34:27Matan Grinberg:We will get this task done at this cost no matter what, but they're pricing it such that, you know, they hope that they can make a margin there. They price it wrong, they're at a negative margin. If they price it right and win the bid, then they get the positive margin. And the way you determine if the task was successful is by some validation loops. Because no one is using these tools anymore where it's just like, write me code great thank you it's generally write me code and here's how i know it was done well and similarly if you are like a model lab and you are given here's a task here's the validation criteria you'll be able to say roughly how much you think you would be willing to pay to get those tokens and you know you want to have some some margin on that and then in that world that's basically that's a way that you kind of dynamically shift from usage based to outcome based um i think that there are so many questions with this and this is very much forward looking but I think there's a lot of questions about how do you subdivide tasks you know divvying that up I think is something that's not obvious yeah but as these tools get better doing things like that actually become way easier yeah that's fascinating yeah yeah if you can scope a task and then create a competitive marketplace that'd be a fascinating version of the future yes and as a user it then creates an incentive to be very thorough in your validation criteria yeah because like you know there are stories of like you know you ask an agent to like fix my code and it deletes your code.
35:41Matan Grinberg:It's like, you know, the solution is just get rid of it all. It's like that Silicon Valley episode. It's crazy how prescient it was. Yeah. But like, so you need to make sure your tests are very thorough because technically it could hit all of your validation. Yeah, exactly.
35:57Maybe zooming out a little bit, you named the company Factory. Actually, you named it Droid before Factory. That's right. But you named it Factory before this concept took off. And now it It feels like everybody wants to build a software factory. Where do you think we are today in terms of the building of software factories? And how close are we to the ultimate vision of a software factory?
36:16Matan Grinberg:Yeah, everyone has a software factory, whether they know it or not. It's just a very inefficient one. So it's kind of like it feels like, you know, pre-industrialization where like, you know, people were manually like, you know, sewing things together or like woodworking or whatever it might be. And these things are very inefficient. Like right now, if you go to an organization that has more than 10 ,000 people and you're to ask about the process by which they decide and release a feature, there is like hundreds or maybe thousands of people in that process. And most likely they couldn't even draw it for you.
36:49Matan Grinberg:Like there's very low likelihood that they would know what that process looks like. That is not because they think that is the right way of doing things. That is just kind of the nature of building large software as it is kind of today. But with these systems, so much tribal knowledge can be codified. So much of this stuff that typically would require, oh, we need to ask this guru who's been here for 30 years who has the wisdom. Oh, we need this approval and that approval. Oh, and I forgot there was some doc that said we always have to do this checklist. And it relies so much on kind of human behavior and like redundancy.
37:24Matan Grinberg:so much of that can be automated and refocused on like what actually moves the needle for our business and i think this move towards software factories is a move towards how do we figure out what are the actual inputs that determine what features we need to build and that might be inputs from the customers inputs from the market inputs from like you know product leaders at the company and let's be very clear these are the signals the inputs that we are taking in here okay great we have those signals then what is the process by which we build this um and really like mapping out the like assembly lines of how you are building software is really important because then you get to close the loop and say did this actually deliver outcome for our business talking before about the tokenomics if you're that cio and you're faced with that question of where do you put every incremental token really the question two years from now is going to become where do you put every incremental dollar and so you're going to have to be asked do you put that incremental dollar towards headcount or towards tokens and if tokens to where in the org and these are things that you can only really know when you have these kind of feedback loops that give you examples of like hey by the way we made those decisions based on this data and it did not matter at all we added these new features and no one cared it didn't create more retention it didn't create more usage or whatever metrics that business is looking to optimize and the only way to do this is like you need kind of more rigor and more process it almost feels like like 10 years from now we're gonna look back at this previous era of software and it's gonna feel like businesses in like ancient times where they didn't do accounting it is like it's gonna be like it's gonna be like marketing in the day of mad men right where it's like all creative and you have no idea what's actually working it makes no like it's like oh yeah let's ship that feature oh i think it went well like yeah we had i got some metrics on that yeah it's like no if you guys read the the blog post that jack dorsey put out about how every company is like an AGI.
39:17Yeah.
39:18Matan Grinberg:There's also this degree to which if your company is an AGI, you want to optimize the weights. Yeah. You want to figure out what nodes are doing what things, which are load bearing, which are not, which need more tokens. Where do you need more nodes? And in order to do it, like you don't train a model by vibes. I mean, okay, actually you kind of do, but you don't, I guess more importantly, you don't do back prop in a model by vibes. Like you are running those actual like calculations and you are seeing when we change this node, what happens. Now, you might be making bets on how to change the model by vibes, but you're like, you're, it's pretty like mathematical in what you were doing.
39:53Matan Grinberg:Meanwhile, at companies, you know, people are determining token budgets just by shooting from the hip. People are laying people off by shooting from the hip and just being like, oh yeah, like 20 ,000. There is no way there is science to laying off 20 ,000 people. That is just like, here's a chunk and let's just see what happens. Instead, I think in these organizations, the way they can do things is much more mathematical of like this part of the business matters a lot and does better if we give it more tokens it doesn't actually matter if we give it more humans so let's give them more tokens there might be other parts of the business where actually giving them more tokens doesn't matter but more people matter because if we build more relationships with our customers and deeper relationships with our customers that matters but these are things that we're going to need like quantitative insight on and you need a software factory to do that otherwise you're just like shooting from the hip and just guessing, which won't work as well.
40:44In the limit, how much do you think people will spend on tokens versus on engineering headcount?
40:50Matan Grinberg:It'll depend on the business. I think every business will have a balance and it just depends on like, like they're just going to be like, an easy example is generally salespeople, they probably don't need that many tokens if they're good salespeople. Because generally where they provide the most alpha is like when they're in the seat face to face with their customers, talking about the customer's problems, understanding, you know, how they build software in our case, and how we can make that, you know, more efficient, more productive. They can use tokens a little bit of like, oh, whatever, generate them some, you know, AI debrief, take some notes, like help them with a follow up.
41:23Matan Grinberg:But like, that's so minimal, the number of tokens, it basically doesn't matter. Like, if you add more tokens to the sales team, it probably won't change their output. If you add more humans to the sales team, it probably will. Meanwhile, Well, engineering teams are pretty different where engineering teams generally, it seems like you want people to own an outcome end to end. But then if you give them more tokens, they can produce a lot more. And so it seems like there and then there's a lot of kind of places in between of like operations, finance, marketing. These are places where are neither here nor there, where I think they're somewhere in between and it kind of depends on your business.
41:56Matan Grinberg:But I think every business is going to have to ask, like, what is our core competency? it's something that we see a lot in the market or we used to see and now they finally kind of hit reality but what we used to see is oh like we're going to build our own like software development agents and we're like okay like you're a like a consumer uh like logistics company like are you sure you want to do it like yeah yeah we're this is a we have to do this and it's like okay and then six months later it's like wait actually this is not a core competency for our business we don't want to hire you know ai engineers to be doing this our core competency is, you know, consumer logistic.
42:30Matan Grinberg:That's what we want to focus on. And I think this is an opportunity for every business to double down on their core competency and what matters for them. And then procure externally, whatever it is that doesn't matter for them. Like a trivial example of this is like, I don't know, in the days of the early internet, you probably had to be a programmer to build a website. And like websites generally help if you're a pizza shop, because you want to have, you know, people come to your pizza shop, they want to be able to, or like whatever. At that time, would you say it was a core competency of like a pizza shop to have engineers?
43:01Matan Grinberg:Like certainly not. Like that is kind of a byproduct of like a brief moment in time. But then there were companies out there that help you build a website. You don't need to be technical. And then this is why we live in a world where like most pizza shops don't have an engineering department, which I think is probably a good thing. And I think similarly, a lot of businesses have dealt with the reality of if you want to do X, Y, Z other thing, you have to bring in people of this type of role. But I think that's been like something you had to do, not because it's a core competency of the business.
43:29Matan Grinberg:And allowing businesses to focus and double down on the things that they're best at, I think is going to be good for the consumers of their business. And so I think we're just going to see like a lot like ruthless refocusing on what actually matters, which is going to be cool to see. On that. So, you know, every company kind of has to go through this process of reinvention. You know, 10 or 20 years ago, people talked about digital transformation. And I don't know if anybody's given it a buzzword now, but AI transformation, something of that sort. A couple of years ago, you ran into a bunch of organizations that just weren't ready to deal with autonomous agents.
44:02Things you've seen your customers start to change. And so the question is, when you look at your customers, as they kind of go up this maturity curve and sort of reinvent themselves for the future, any good tricks or techniques that you've seen them use to repot themselves a bit?
44:19Matan Grinberg:Yeah, I mean, I think surprisingly, like the companies that have been doing like company-wide hackathons really end up doing well. It seems like relatively trivial, but like just setting aside a day for everyone in the workforce is just like, build shit with AI. It really sets the tone and sets the pace. So just give me a look. I tried to force him to build stuff with coding agents. It didn't go so well. We'll work on it. We'll do it after this. You know, we gave it a great effort. But that's it. Like, it literally just setting aside the time to, like, do it. And, like, even if it fails miserably, like, it's fine.
44:54Matan Grinberg:And also, like, the orgs that are okay with failing. Yeah. Like, it feels like there are some who are like, we need to do it exactly right. We need to make the right decision from day one. No error. Like, you're going to make mistakes. Everyone is going to. And the orgs who are kind of leaning into it and embracing it to a certain degree, I think, are succeeding. Like, one of our largest customers is EY. ey is not necessarily known to be like at the absolute frontier of ai but i think for them they were just like look this matters we were kind of there have been other transit transformations that we relate to we're not going to be late to this like we're just going to go in we might mess up but like obviously respecting like secure the things that you're not allowed to mess up sure put those aside but like let's go and get our engineers to mess around and build this stuff and see where it breaks and understand what they like and what they don't like um i think that really matters a lot in the ones that we're seeing succeed and also the ones who are like pretty bold in reinventing the processes that they've put in place and just saying like hey it's there's no sacred cows like let's let's put this aside try something out if it doesn't work put that sacred cow right back um and i think that's that's been kind of a determining factor there um and when it comes from within if it comes from the board probably not going to go well yeah If it comes from within the tech team or the ICs or the leadership, that's when we see it go better.
46:12Do you have any predictions for the most important changes that are going to happen in your space over the next, call it, 12 months?
46:19Matan Grinberg:A lot of AI consumption is going up like crazy. And everyone's super, super excited because the revenue is going wild. A lot of this is synchronous usage. In other words, if everyone woke up sick tomorrow, a lot of Claude code usage would be zero. because it's all just hey cloud code or hey codex or hey droid right i think in 12 to 24 months like 90 of tokens will be asynchronous tokens so these are going to be you know droids on their own autonomously being like hey here's some signal that i found from a customer let's go fix it or let's go create a first pass solution to this and i think that is going to be where the real like agent native stuff begins because right now we're still kind of in like co-pilot mode like if you're going to an agent and say hey go do this for me it is more agentic because it's not going to come back and ask you a ton but it's still like you are kicking it off like yeah if you guys have ever been to tesla's factories which is one of the sources of inspiration for the name is like it's just robotic arms everywhere going and doing stuff like it's not like there are people there like going and you know attaching the widget to the thing um and this idea of like a dark factory where like the lights are off and things are just happening that is where software development is going that's kind of where the the name came from is like you know elon was always talking about the factory is the machine that builds the machine.
47:32Matan Grinberg:And that's been something that we took to heart. And I guess also that combined with his whole thing about how you're destined to become the opposite of your name. And in our case, factory becomes artisanal. It's kind of a good flip there. What's your most optimistic version of the future, both for factory and for the world at large? So I think short term, there's going to be a lot of turbulence because I think a lot of companies have misallocated resources pretty poorly. There's been a lot of bloat. And I think the correction that's going to happen there is going to be really painful for a lot of people.
48:08Matan Grinberg:And I think that's something that I think every AI CEO should really bear much more responsibility than they currently are for. And also figuring out ways to like address and kind of ameliorate in some way, because this is something that's going to be very painful for a lot of people. now i have optimism that we can actually address that faster than we think we just need to start now in terms of addressing that now the longer term and why i think this is a good thing is why i don't believe at all like you know the bs that people are saying oh engineers are going away generally there is a huge number of problems in the world a large subset of those problems can be solved with software a small subset of those problems are currently being solved with software And so in the short term, this means that, okay, first, there's a given problem that was over allocated engineering resources.
48:58So okay, we need to reallocate those reallocate those is a very kind of cold way of saying some people are going to lose their jobs. But I think the thing that's going to happen in the longer term is we need engineers, engineers are some of the best systems thinkers and the best problem solvers.
49:13Matan Grinberg:And there are so many problems that can be solved with software that are not being solved with software. And so that means that we are going to take those engineers and have them go and solve problems that previously were not being solved. That is such a net good for the world. Because again, there are so many of these problems that we are not solving. And also, there's so many problems that we are maybe solving, but with really shitty software. And like, this is going to enable people to solve it with incredible software. And, you know, the vision for Factory is that we are kind of the factory that allows them to go and build this incredible software to solve these different problems.
49:44Matan Grinberg:And these problems range from like things that are trivial to, you know, like government software typically is not very good, whether it's like DMV or like IRS web, like all that stuff is generally a pretty poor experience. We don't need to live like that. Like we can live in we can live in a world where all software is really fantastic. But also things like, you know, pharmaceutical research, like so much that goes into solving diseases is not just like a biology problem. A lot of it requires the best software engineers in the world. And previously, those problems haven't allocated the right dollars to attract the best engineers.
50:20Matan Grinberg:But now, because of what's happening, I think we will be much more closely allocated to like, these are the biggest problems. Let's get the best minds and the best problem solvers to solve that. I think it's kind of our job as an industry to do that reallocation as quickly as possible. So it's not 10 years, but maybe like six months or a year. Wonderful. Matan, I think the clarity and consistency of your vision over time has just always been very inspiring. And then just seeing how much you've grown as a leader and how much factory has grown as a company, even since the last time we did this training data episode, it's truly all inspiring.
50:54So thank you for joining us again to share what you're up to.
50:57Matan Grinberg:I appreciate it a lot. Thank you. Thank you.
51:12Thank you.
From the publisher
Factory started building fully autonomous coding agents in April 2023, two years before enterprises were ready. Matan Grinberg now says this is indistinguishable from being wrong. The Factory co-founder and CEO explains how the company survived its "journey in the desert," including the decision to hand nearly all of its revenue back to customers when the product wasn't making developers obsessed. Matan makes the contrarian technical case that a model-agnostic harness beats the model-and-harness co-design that labs like OpenAI and Anthropic favor, because exposing a harness to many models keeps it from overfitting to any single one. He argues open-weight models like GLM will capture the majority of tokens by staying one generation behind the frontier at a fraction of the cost, and that CIOs will soon justify every incremental token the way they justify headcount. Looking ahead, he predicts 90% of coding tokens will run asynchronously—the "dark factory" where software builds itself.
Hosted by Sonya Huang and Pat Grady, Sequoia Capital




