In short
Podcast Episode Summary: This AI Agent Builds Better Code Than Most Developers
Podcast Title
The Neuron: AI Explained Hosts: Grant Harvey and Corey Noles Episode Release: Every Tuesday
Episode Description In this episode, Eno Reyes, co-founder and CTO of Factory AI, discusses the evolution and capabilities of autonomous coding agents, namely referred to as "Droids." These agents can autonomously manage coding tasks within real production workflows, highlighting their efficiency and advanced capabilities in comparison to human developers.
Key Topics Covered
Introduction to Droids
- Definition: Droids are fully autonomous agents capable of taking tickets, modifying codebases, running tests, and integrating with existing development workflows.
- Origin of the Name: The term "Droid" was chosen to move away from the negative connotations associated with the term "agent," which had become synonymous with unreliability in AI.
The Need for Autonomous Agents
- Problem Identification: The initial motivation for creating Factory AI stemmed from the recognition that software development bottlenecks often lie in the end-to-end process rather than simply writing code.
- Evolution of the Vision: The vision for fully autonomous agents has remained consistent since inception, with early prototypes designed to operate independently.
Insights from Context Compression Research
- Context Compression: Factory AI's research into context compression outperformed competitors like OpenAI and Anthropic. It focuses on the quality of the codebase as the primary predictor of AI success rather than adoption rates or token usage.
- Agent-Readiness Criteria: The research posits that a clean codebase with effective documentation, testing, and tools significantly enhances an agent's performance.
Practical Applications of Droids
- Versatile Functionality: Droids can be integrated into various environments (e.g., IDEs, CI/CD pipelines) and customized for specific tasks, such as code reviews or security analysis.
- Agent-Ready Codebases: Factors that determine whether a codebase is ready for agent integration, including documentation quality, testing coverage, and overall structure.
Challenges and Limitations
- Long-Term Task Management: Challenges arise when tasks extend beyond a couple of hours, indicating the need for delineated autonomy levels and guardrails in agent operations.
- Need for Control: Users can set different autonomy levels (low, medium, high) for agents, with safety measures in place to prevent destructive actions.
Future Directions
- Expansion Beyond Coding: The potential for general-purpose agents that can assist in various fields beyond software development.
- Adoption in Diverse Domains: Users from various sectors (e.g., finance) are beginning to adopt Droids for tasks such as data analysis, showcasing the versatility of these agents.
Advice for Developers
- Building Agentic Tools: Focus on the user journey and leverage existing frameworks like Factory’s Droid to create effective agentic solutions. Avoid overcomplicating the initial development.
Key Takeaways
- Autonomous Agents: The rise of Droids marks a pivotal shift in software development efficiency, enabling teams to automate complex tasks and enhance productivity.
- Context Management: Maintaining context over long sessions is crucial for agent performance, and effective context compression strategies are necessary for success.
- Agent Readiness: The quality of the underlying codebase is a critical factor in the success of AI integration within software development environments.
Additional Resources
- Try Factory AI: [Factory AI Website](https://factory.ai)
- Research Paper on Context Compression: [Evaluating Compression](https://factory.ai/news/evaluating-compression)
Conclusion The episode concludes with an emphasis on the transformative power of autonomous coding agents like Droids, their impact on the future of software engineering, and the growing possibilities for their application across various industries.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe Rise of Autonomous Software Agents
0:45 to 3:30
Discussion about the evolution and reliability of AI agents in software development.
“I'm excited because today we're digging into the rise of autonomous software agents.”
Interview with Ino Reyes
3:50 to 8:00
Ino Reyes discusses Factory AI's approach to building autonomous agents.
“that didn't exist at that point in time and figure out what can we do to make them general and very capable across these different tasks.”
Understanding the Agentic Coding Problem
8:00 to 12:30
Exploring the challenges and solutions in coding workflows that led to the creation of Factory AI.
“And so all of that bubbled when, you know, basically the LLM APIs came out in a way that was accessible to a software developer, where you could just call OpenAI's API.”
The Architecture of Droids
12:30 to 14:04
In-depth explanation of the design and functionality of droids in software development.
“When you take it and you put it in a cron job, it becomes an automation that can do almost anything you could do on a computer.”
Understanding Droid's Code Exploration Capabilities
14:04 to 17:02
Learn how Droid effectively navigates large codebases and its operational environment.
“expand that and make it customized to your organization, you can introduce skills and custom sub agents and all these other things that basically allow it to conform to your organization standards.”
Model Flexibility and Customization in Droid
17:02 to 19:44
Discover how users can customize models and the importance of model flexibility.
“And how do you choose the right model for a specific agent with the way these work?”
Building and Fine-Tuning AI Models
19:44 to 22:47
Explore the challenges and strategies for building AI models in a competitive landscape.
“My question is, how does it know to do that?”
Enhancing Agent Behavior and Validation
22:47 to 24:49
Understand how to improve agent behaviors and ensure task validation effectively.
“Then from a behavioral perspective, there's also a lot that goes into enabling the tooling to get easy answers from important questions.”
The Challenges of Context Maintenance in Long Sessions
24:49 to 28:00
Learn the difficulties in maintaining context during extensive agent interactions and solutions implemented.
“So let's talk about the blog that you just published called Evaluating Compression.”
Evaluating Compression Methods for LLMs
28:00 to 29:50
Learn about evaluating and comparing different compression strategies for language models.
“You can look at completeness and continuity and instruction following.”
Show all 20 chapters
The Impact of Structure in Compression Techniques
29:50 to 33:00
Understand the importance of structured data in improving the efficiency of compression methods.
“Where it was just able to recall all of the critical pieces of information quite well.”
Practical Implications of Long Context Handling
33:00 to 35:40
Discover how advanced compression strategies affect user experience and task completion in coding sessions.
“Like we kind of talked about it, but maybe we can just put a finer point on it.”
Expanding AI Utility Beyond Coding
35:40 to 39:10
Explore the potential applications of coding agents in various industries beyond software development.
“coding agent and just makes it feel like you can keep going.”
Creating Agentic Writing Tools
39:10 to 42:01
Learn about the concept of agentic writing tools and how they can enhance creative writing processes.
“I'm suddenly drawn back to the idea that when we met in person a few weeks ago, I told you I was going to try to write with wine.”
Understanding Agent Ready Code Bases
42:01 to 43:30
Learn what it means for a code base to be agent ready and its implications.
“writers will use you know agentic software development tooling to basically build out tons of ephemeral uh content yeah exactly uh and then they write and execute the the document in whatever strategy they find apt.”
Predicting AI Success Based on Code Quality
43:30 to 46:32
Discover how code cleanliness correlates with AI productivity and success.
“I mean, so this is what's interesting is they looked at adoption.”
Making Businesses Agent Ready
46:32 to 47:47
Explore strategies to make entire business processes agent ready for AI integration.
“Do you have any intuition following that same line of thinking how a company could make their the rest of their business agent ready?”
Guardrails for AI Agents
47:47 to 49:54
Understand the importance of guardrails in the effective use of AI agents.
“I think that there is a ton that agents cannot do fully autonomously.”
The New Era of Software Development
49:54 to 51:19
Learn how AI tools are transforming accessibility in software development for beginners.
“So it sounds like if you're just starting with this is very powerful for engineers, but if you're just starting with engineering, that this also is a good tool for you.”
Building Agentic Tools: Best Practices
51:19 to 53:05
Get advice on creating effective agentic tools and focusing on user journeys.
“So I think as a software developer, because you understand how these things work, you'll also be able to go even further.”
Transcript
Automatic transcript. May contain errors.0:00Eno Reyes:These aren't humans, they're a different type of tool that humans can use. We have customers that have quite literally run sessions that go well under the 80 million tokens. Two or three years ago, the term agents had really negative connotations because the only agents that were really out there were extremely unreliable. No one had seen this sort of fully autonomous interaction form factor yet. It's a fully self-contained binary that runs on any operating system, on any device, in any environment, with any interface that you can imagine.
0:39Welcome humans to the Neuron Podcast. I'm Corey Knowles, here with our purveyor of words and slinger of cat puns, Grant Harvey. How are you, Grant? Doing well, doing well. How are you, Corey? I'm doing good, man. Doing good. I'm excited because today we're digging into the rise of autonomous software agents. Factory AI is one of the fastest growing companies in this space, building droids that can take tickets, modify real code bases, and work inside existing dev workflows. But before we get started, please take a quick moment to like and subscribe so you don't miss out on videos just like this one.
1:12And joining us today is Factory AI co-founder and CTO, Ino Reyes, to talk about how this works, what breaks, and what it's like to scale a company at this speed even. Eno, welcome to The Neuron. How are you?
1:26Eno Reyes:Thanks for having me. Excited to be here and excited to chat more about the droids. The droids. The droids. Were you inspired by Star Wars with the name droids? My lawyers say that that is not something that I can comment on. Just kidding. No. Actually, the name's origin comes from the fact that two or three years ago, the term agent had really negative connotations because the only agents that were really out there were extremely unreliable they would break at every corner and when we wanted to sell to enterprises if they heard the word agent they would associate it with basically like a research experiment so we came up with a name that we felt would also prevent people from thinking this is like a human that's going to replace us like a lot of people name their ai tools with a human name and we said, look, these aren't humans.
2:17Eno Reyes:They're a different type of tool that humans can use. What problem with agentic coding were you trying to solve when starting Factory? Like what felt broken about how software teams worked at the time that you started it? When we started Factory, my co-founder was doing his PhD at Berkeley and I was at a company called Hugging Face that worked on LLMs and open source tooling for the related. And actually, when I was there, it was pre-LLM. So there were BERT transformer models and tons of different variants of that, diffusion models, all this other stuff. And one thing that was pretty consistent was for the organizations that wanted code models.
2:59Eno Reyes:All of them were looking to fine-tune a model for their code base and maybe get autocomplete. when you when we talked around to a bunch of companies we saw that everyone was very fixated on coding which obviously is extremely important but the biggest bottleneck oftentimes in most software organizations is not can you write the lines of code to solve the problem but it is the entire end-to-end process and the bottlenecks and the the the blockers that occur along the way right? Gathering context, understanding decisions, planning changes, stitching all this together. And once you get the code, you know, that's almost like the easiest part of the process.
3:42Eno Reyes:So we wanted to make sure that AI systems that we built for software development could help people with that entire end-to-end process, which required us to basically do research on agentic systems that didn't exist at that point in time and figure out what can we do to make them general and very capable across these different tasks. So was the idea like immediately that we're going to do fully autonomous agents or was it, were there stepping stones along the way you were hoping for? How did that vision kind of evolve if it did along the way? Yeah, I mean, it's kind of funny and there's actually a quick answer.
4:22Eno Reyes:It basically has been autonomous the entire way. I think this might be one of the bigger differentiators between us and a lot of the other players in the space. We had no point wanted to build autocomplete or wanted to build an IDE. From the beginning, it was droids, fully autonomous systems. You know, our mission when we first started and to this day is to bring autonomy to software engineering. And so we built out the first droid prototype. It was a Slack bot that you could chat with and it would work on its own and come back to you. And this is in 2023. and so basically you know that no one had seen this sort of fully autonomous interaction form factor yet yeah but along the way there were a ton of lessons I think we were way too early with that vision of software development and interaction both because the models weren't necessarily strong enough to do the breadth of tasks you wanted to do and required a bunch of harnessing but also because people weren't used to that interaction pattern it was surprising to us after like six to seven months, we got to this point where customers would say to us, you know, it's actually like working pretty consistently, but we're not really sure how to teach people how to use this interaction pattern.
5:38Eno Reyes:And that made us realize that we were, we needed to meet people a little bit closer to where they are. And so then it evolved a couple times to the current form factor that it has today. So how did your background and hugging face shape the way that you approach the autonomy question, I guess. Yeah, totally. And I think that it actually goes back a bit further. My first sort of foray into deep learning and LLMs was in a research lab in undergrad that focused on computational models of cognition. And basically, the thought process there was, you know, you want to be able to look at different ways that fMRI, EEG, brain activity interacts and see if you can model algorithmically or with deep learning how those interactions might play out.
6:28Eno Reyes:Not to say we know exactly what's happening in the brain, but just to say, you know, we can actually imitate some input output algorithm that clearly the brain is also executing. And so there's some insight there, even if it's not the same thing. And I found that to be like a really interesting idea is like, look, maybe it doesn't matter what's going on inside as long as you can model it. And, and so from there going to Microsoft and seeing like a very large organization, software development life cycle, experimenting with ML strategies and techniques, and then hugging face where most of these, you know, ML models were hosted.
7:06Eno Reyes:I think that there was sort of this, this clear side of like what software development looked like at different scales, how people were thinking about different deep learning and ML strategies. Transformers had, you know, very, very recently become apparently the strongest method for most deep learning techniques. And it really was seeing how people were very fixated on the singular use cases of transformers. Like we want to build a classifier or we want to build a, you know, entity recognition system that when LLM started to become strong, seeing how general they were and how they started to eat all the different use cases, it became clear that there was a different approach towards software development that wasn't going to involve piecemeal things like, oh, now we have autocomplete.
7:58Eno Reyes:Now we have a tool for classifying whether the code has a bug in it or not. And so all of that bubbled when, you know, basically the LLM APIs came out in a way that was accessible to a software developer, where you could just call OpenAI's API. You could just call, you know, Cohere was actually like a popular API at the time. And so I think that there was a couple of folks in the open source community, including myself, who basically just said, what if we put this in a while loop? Like, what if we keep seeing what the LLM can do? What if we add structured parsing? And so I think those ideas were all bubbling up right around 2023, early 2023.
8:42Eno Reyes:and seeing it actually all come together. Like when I met my co-founder, it was actually a seven day period from the time we met to, you know, basically, you know, he dropped out of his PhD. I quit my job and we started factory. That's how fast you could build something that was different. What was that conversation? Yeah, like how did you guys have that conviction in seven days where you're like, let's do this? I think that there's so much to it. You know, he likes to call it intellectual love at first sight. I think I agree with that. We knew each other in undergrad very lightly, you know, Matan and I, and I think that the thing that was nice is because of undergrad, you know, you sort of knew, we knew each other, we had social circle overlap, but we weren't actually very good friends.
9:29Eno Reyes:We like barely, barely had chatted. And I think that it came at a really serendipitous moment where he was basically, you know, had basically kind of stopped doing his work as a physics PhD, was transitioning towards AI and very interested in AI research. I had built out a bunch of these prototypes for, you know, while loop LLMs. I thought it was going to start a company with this other person and around like finance AI. And interestingly enough, the, the turns out when you're trying to build a finance AI agent, you really need to build like a python ai agent and so having a system that can write python and execute it in a loop very quickly makes you realize well maybe there's some adjustments that can be made to make this just general software development and so when we when we joined together and we spent that week basically hacking and putting it together to a to a point that made it more general it became clear that he was very convicted that i was very convicted and then a healthy sprinkle of, you know, we're both young and, you know, you only have a couple shots on goal when you're in the early days.
10:36Eno Reyes:I think that's actually in retrospect completely untrue. I think you have infinite shots on goal. But, you know, at the time it felt like, oh, we have to do it, do or die. When you say droids, what essentially are they? I know we're talking about agents, but what are they really doing under the hood and how autonomous are they today? You know, I think that there, this is probably the most interesting question is like, what, what actually is a droid? Uh, you know, and for us, there is a lot of work that we put into, uh, because it's changed a lot, right? When it first started, it was this like server side, you know, effectively while loop on a backend server.
11:13Eno Reyes:Uh, then it was, uh, you know, a full stack application that you had to interact with either in the browser, um, or you could connect it to your computer. And now I think that one of the biggest architectural insights that we've introduced is, you know, the Droid should really be the most simple possible program that you can imagine. It's a fully self-contained binary that runs on any operating system, on any device, in any environment, with any interface that you can imagine. So the actual Droid itself is just this very simple program that orchestrates and manages LLM calls and that can be run anywhere.
11:57Eno Reyes:And what that enables is really this extremely wide range of use cases where when you take a Droid and you put it inside of a terminal interface, it becomes this like really snappy, fast terminal agent. When you take a Droid, that simple binary, and you wrap it in an IDE extension, it becomes an IDE pair programmer. When you take a Droid and you put it in your CICD pipeline, like your GitHub actions, your Azure DevOps pipeline, it becomes an automated code review or an automated security review. When you take it and you put it in a cron job, it becomes an automation that can do almost anything you could do on a computer.
12:37Eno Reyes:And so the architecture is built for generalizability so that this individual agent can now take control over the file system and the environment it operates in and really do anything that it can with computer control. So it's this really powerful concept that is very inspired by the Unix philosophy of simple, modular, composable programs. And I thought you had, you have four kind of main ones that you were kind of marketing at first. I think you touched on some of those, like a security one, or like ones that are customized for different use cases. Is that accurate? Or is it, has it expanded since then.
13:10Eno Reyes:Yeah. So the actual droid itself is actually, you know, is totally general. Now, what developers can do, though, is there's tons of customizations that they can introduce. You can introduce skills, which give it the ability to basically pull in context and work on any workflow or bucket of information that you want to pull in. You can do hooks. You can do custom slash commands, all these different customizations that you can bring in, sub agents, et cetera. And so what people can do is if they want a code review droid, then they might place droid inside of GitHub Actions and say to it, you should be code review agent.
13:52Eno Reyes:If you want a security droid, you can place it in your, you know, whatever system that you use for either pipelines or security review, and it becomes a security review agent. If you want to expand that and make it customized to your organization, you can introduce skills and custom sub agents and all these other things that basically allow it to conform to your organization standards. I've heard that Droid is like kind of almost like an engineering team that runs like in the background, like you can use it in that way. How do the agents then understand large, messy, real world code bases versus like, say, like another agentic tool and how they?
14:31Eno Reyes:I think that this is one of the most interesting questions because there are several things that go into being able to explore and understand a very large code base, especially for hard problems. I think that the three angles I would say is there's the environment. That matters a lot. Basically, what information does the droid have access to in order to understand? Two, there's the droid's inherent ability to explore and search. and three is how it handles long-term goal-directed behavior. That third is arguably the most important piece of research that, you know, as factories in the agent research lab and that's basically the most important piece of research is how do you do long-term goal-directed behavior?
15:12Eno Reyes:But I think on that first point of like the environment, one thing that I think makes droids really effective in larger codebases, enterprise environments is that it makes it very easy for you to understand is your environment actually agent ready? In other words, it will tell you, you know, is Droid operating in its best environment possible? We have tools like agent readiness that help you understand is the code, does the code base have agents.md files? Is it documented? Is there linter feedback that the Droid can use to verify its own work? And it will proactively seek those verifiers so that if it wants to explore a code base and see, hey, is this code base over here?
15:53Eno Reyes:sorry, is this like section of the code base doing this function? It will use, you know, grep and glob and bash and all the tools available on the computer, but it will also look for validation and verification. Can it, is there some sort of a mono repo structure that it can use? Are there tools built into the repo that help it search for information? And so from an environment perspective, the droids are very tuned for that from you know, a, from the perspective of, you the agent basically be able to run for that period of time. I think that there's a lot of work that we've done on things like context compaction or compression.
16:33Eno Reyes:There's a lot of work that we've done on things like prompt caching so that it's fast, right? If your agent takes two hours to search the code base, that's obviously not good. So keeping it really fast and keeping it effective over a long period of time, that's another major area that we focus on. And I think that for our users, they find that the more you invest in making your code base agent ready, the more tools like Droid work effectively. So you have some model flexibility built into here too, right? And how do you choose the right model for a specific agent with the way these work? Are you able to adjust them on the fly or are you able to uh you know i think of that whole idea of of like you know having a basic and a power model kind of paired together where for certain parts of tasks it's using this but maybe it recognizes when it's in over its head and needs needs some some extra goju's yeah totally and model being model agnostic is is actually very core to the the droid as a product experience, we believe that there is a lot of value in being able to use basically the frontier models the moment they come out across all of the different providers.
17:51Eno Reyes:We've been in a lockstep battle across, you know, five plus different foundation model providers for the last three years. And our view is that, you know, as a developer, you deserve choice. You should be able to pick any of these and have an extremely high quality experience. And so the Droid has a bunch of different customization options. One, it built in, we have the ability to route to the most capable models at any given moment. So you know, one click when you're a factory plan subscriber, you can just pick from Opus, Sonnet, GPT 5.2 codex, GPT 5.2. Now, there's also the option of bringing your own model and bringing your own keys.
18:32Eno Reyes:So you can choose any model. And so people can use really any inference endpoint that supports the three major inference endpoint standards. Right. Yeah. You can even use on your own device. Right. VLM or Olama. So so you can have a fully local coding agent experience with Droid if you have a powerful enough computer. Yeah. Yeah. Now you also can customize beyond just, you know, I'm using this one model in this one session and you can do things like I would like to plan. And we have something called spec mode, which is basically specification. And you can plan or specify with one model, but then execute with another.
19:14Eno Reyes:And so what that lets you do is it lets you say, I want to use an expensive or powerful model to do X task. And then a cheap model to actually execute because most of the efforts in the planning. And I think people take advantage of this most when they might have one model for code review. They might have one model for security review. Their daily driver is a third model. And then when they really need to pull out the juice, they'll use like the most expensive model. Right. So that flexibility is super important because most model providers basically give you one, maybe two options at most. You know, you're a research company and, you know, the agent is going in and it's deciding or it's figuring out, is this environment good for me?
19:52My question is, how does it know to do that? Is that like a system prompt? Is that something you trained into the harness? Like how, how does it know to, to look for those things? And I guess the second thing is, would you ever create your own models because you are doing research? Like, is that on your roadmap? Are you interested in that?
20:10Eno Reyes:On the, so I'll, I'll speak on the latter first, uh, in, in like, will we build our own models? Our view is that we're definitely not opposed to building models. I think that the challenge is you want to build, you basically want to fine tune or build models as basically as late as possible before it becomes important to, right? Right. That's sort of that's sort of our view. And what I mean by that is, you know, if if you had fine tuned GLM two or some model that that's earlier, you know, it gets blown away by the model providers at the next level. Right. And so you have three weeks away usually.
20:46Eno Reyes:Exactly. Exactly. You have, you probably have three weeks before there's a better model out and you have a month to two months before it's completely forgotten about. Right. So if your bet is we're going to beat the people with$50 billion data centers, I think that you have to have a very clear strategy. And, and, and, and I think that there are reasons to do this by the way, costs, right. If you know that you can execute on a specific task at the max that a human cares about and you just need to bring down cost, that's a great reason to fine tune a model. Right. Um, yeah. But at the end of the day, uh, I think that we've already seen this a bunch of, you know, software development organizations, uh, or, you know, software development, AI tooling, uh, they've all trained their own models and, you know, within two weeks to Gemini flash came out and it's like better, uh, or within three weeks later, 4.5 opus is now like not just better, but it's like two orders of magnitude better.
21:40Eno Reyes:Uh, and so I think that it's going to be a really challenging battle to fine tune or build your own models in the near future. But we're definitely not ruling it out. On the earlier point about being an agent research lab and sort of like how we tune the agent's behaviors and in particular for seeking verification or validation. There's a lot that you can do at the harness level to enable this, right? Some of it might be like context engineering. So adjusting not just like the base system prompt, but reminders that you can insert automatically or, you know, pre-built environmental reactions to tool calls.
22:17Eno Reyes:You know, the third time you write or edit a file without calling a linter, maybe you insert a reminder that say, hey, by the way, you should just check up to make sure this work is actually, you know, valid. You can seek these out in the background. And whenever you edit a file, provide like language server feedback, right? That helps the model understand, oh, I might be missing these things, right? So there's tons of things you can do within the harness that make the model more aware without needing to actually waste LLM calls. Then from a behavioral perspective, there's also a lot that goes into enabling the tooling to get easy answers from important questions.
22:55Eno Reyes:So like we have in our product, an API for agent readiness, which basically and a dashboard. So there's 150 plus signals that a code base could have, ranging from is a linter present to, you know, are tests available? Is there documentation across the code base? And this applies to mono repos, where it will segment out each project, or single code bases. So when you make it really easy for an agent to seek that information, like one API iCall to get the answer uh you you make it easier uh in general for the system to rely on that information versus if it needs to find that out from scratch every time you're going to waste your users tokens trying to like get the answer yeah i've i've lived that actually like working with an agent like in in my terminal i've i've been constantly reminding it like remember you should check all the scripts to make sure that you you know know all the dependencies of everything like I'm doing that.
23:51So like, yeah, use factory because it'll save you those tokens.
23:56Eno Reyes:Totally. And hooks are also great for that. Like, I think there's a pretty popular one right now called Ralph Wiggums, which is an odd name. And I always feel silly saying that it's from the Simpsons, but it's basically a hook that you run at the end when the agent loop stops and it sends a prompt back to the agent that basically says, you know, here's the step-by-step plan of what you were supposed to do, did you actually validate and complete the work that you're supposed to do? If not, restart, right? And so it's effectively a while loop in a while loop. I honestly think that the existence of plugins like that means the agent's harnesses aren't taking advantage of what they could be.
24:36Eno Reyes:But regardless, people find this out and then they build and they customize. And I think that's one of the coolest parts about building a really simple and modular tool is the community goes out and they find all these interesting ways to use the tool and they get a lot of value out of it. So let's talk about the blog that you just published called Evaluating Compression. Really, really awesome blog and research that you did. The hardest part about working on large code bases is context, right? So why is maintaining context across long agent sessions so difficult? And how did you attempt to solve this?
25:10Eno Reyes:Yeah, this is a really tricky problem because I think that, you know, we've actually published like a couple of different posts about just how hard this problem is ranging from, you know, there is tons of information and there's obviously context windows in all models, right? You have, let's say, 1 million tokens of context or 2 million tokens of context. There is cost questions involved, right? So most model providers actually increase the price that they'll charge you for beyond a certain level of tokens, right? And then there's like speed and quality. the classic like lost in the middle evaluation where people talked about as you increase the context that's utilized 1 million tokens of available context and that they'll let you send the API call does not mean that the LLM will actually be able to reason through all that information so if you think about it like what information is even in an LLM call to an agent you have the task description that the user gave you or like the user's messages you have tools which need to be available to both the agent and the system prompt, as well as the tools that they then call throughout and their responses.
26:25Eno Reyes:So one tool call to Bash that says, get me all syslogs might be 200 ,000 tokens, right? You have the developer persona. So all the information about their environment, their role, you know, is there, are we in a Git repo? Of course, code and all of the files that it's reading, markdown, code, Maybe it browsing, it browses the web and retrieves information from the web. You have historical context that might have occurred previously. And you have like all of the other sort of artifacts that might pull in from system reminders. So you basically, the LLM is seeing so much information about the task at hand in order to solve it.
27:06Eno Reyes:And so over the course of a very long conversation, you know, it builds up. And so in order to prevent this from crossing over a certain threshold, you have to figure out a way to compact or compress this information over time. And so what we did was we initially had a naive solution that just compresses it and summarizes it and says, keep going. But you quickly realize that that does not work. Um, and so then we started to try to figure out, you know, what are the individual dimensions that matter for a compression of a situation? Like, you, you know, it's going to lose some information. So what matters to preserve?
27:47Eno Reyes:How can we design a system that not only doesn't just summarize, but sort of provides a very high quality and active block of information that the agent can then resume its task with little to no issues. And so, you know, you can look at all these things like the accuracy, if whether or not the actual compressed artifact is accurate, you can look at its context awareness, like what information is or isn't included about what the what is currently happening, what has previously happened, you can look at the artifact trail, are the files and the logs and the critical, like, you know, singular information pieces present in the summarization.
Read the full transcript
28:33Eno Reyes:You can look at completeness and continuity and instruction following. Anyway, you can look at all these different dimensions, and you can start to evaluate different strategies for context compression. And so what we found is that using this method called probe-based evaluation, which is a fairly popular way to evaluate LLMs, where you basically, you know, the idea is you can ask, like, very focused questions with rubric based evaluations using another LLM to extract information from some state or from some answer that has previously been executed. And so, you know, an example probe might be, you know, I have an agent session and then I compact and then a probe would ask, what was the file that had the bug inside of it?
29:21Right. And so if you can answer that question,
29:23Eno Reyes:right and lms are now good enough such that it's basically a binary like yes no like it either has the information or it doesn't right yeah um and uh you sort of still to some extent have to think a little bit about you know what if it does what if it gives the right answer but it's not parsed correctly so there's a little bit of work required to make sure this is robust um but generally fairly reliable way to determine does the lm actually know at this point in time what happened And so when we evaluated our compaction method and compared it to OpenAI's compression strategy, as well as Cloud Code's compression strategy, we found that we had generally built one that was across all of these dimensions, much stronger at, you know, instruction following continuity completeness, but most importantly, just accuracy and context awareness, right?
30:14Eno Reyes:Yeah. Where it was just able to recall all of the critical pieces of information quite well. So you built your own compression approach, and then we do have where OpenAI and Anthropic have landed their own kind of methods as well. And what else did you find around benchmarking those together and seeing how they compare? I know you said yours was faster. Well, not just faster. I think actually speed was relatively similar across the board. but the two things that really matter is like the quality of the compression and how much it actually compresses, right? Basically like the token reduction efficiency.
30:56Eno Reyes:And we do have the worst token reduction efficiency. You know, OpenAI is 99.3%. CloudCodes was 98.7 % and ours was 98.6%. So like 0.1 % off. Maybe that's within the error bars, right? But the overall quality, right? You can basically take all of these characteristics and you can build sort of a quality score that just says, you know, across all these dimensions, which one is stronger? I think that like probably the most important thing we learned was just how much structure matters, right? So I think that probably the biggest failure case is generic summarization. And I think that the worst performing in our evaluation techniques were the ones that basically just treat all content as equally compressible.
31:45Eno Reyes:It's just one big summary and let the LLM figure it out. Like a file path could be very like low entropy information, but it's probably the most important piece of information an agent needs. And so if you build an explicit structure into your compression strategy, hey, like files are all going to go here. Decisions are all going to go here. That where the agent is currently operating is going to go right here. And you also preserve some of the continuity in the conversation, right? Maybe the last message or the last couple of messages should actually be preserved. So the agent is able to pick back up without losing track.
32:24Eno Reyes:That's huge. Compression ratio is obviously the wrong strategy as well. So like, it doesn't matter if you compress more, if it's worse, because you'll end up actually spending more tokens to get back to where you were. And then you've lost the like 1 % alpha. So two things. Number one, I saw Ray Fernando, who's one of my favorite AI coders on YouTube, say he had a coding session with Factory with 7 million token context, which is just gnarly. I can't believe that. How is that even possible? Is it just for implying this technique? And I guess the other question is just in production, why does this matter, right?
33:01Like we kind of talked about it, but maybe we can just put a finer point on it.
33:04Eno Reyes:Yeah, totally. I mean, I think that the so so the question is, was that very long session because of the strategy? Yeah, totally. Like, technically, this is this can go infinitely. We have like a strategy at the very end, if it were to accidentally run over by over compressing, we have a way of like flushing, in which case you might see a minor decline at that very, very, very long end of a of a potentially multi day session. but really this strategy allows you to run for a very long time. I mean, we have customers that have quite literally run sessions that go well under the 80 million tokens.
33:43Eno Reyes:And so really the thought process here is if the question is, can it recall the most precise detail about at the very beginning of the session, the second question you asked, very unlikely that it will be able to. However, there is one additional thing you can do, which is all of these sessions are just files, right? Like the session is a large, large file on your computer, which is something that the Droid can actually access. So if you really wanted to know what was the first user message that I sent at the beginning of the session, Droid will just grep the first user message in the session that it's inside of, right?
34:22Eno Reyes:So there's also the ability to access previous messages in a fully lossless way, right? And so I think that the practical implication is one, we really want people to not need to think about context in the way that they do today. With many coding agents, there's like an anxiety that's created by users by seeing their context fill up. It's like, you know, they see it coming close to compaction. They get stressed. They're like, oh, how do I change my behavior in order to either not let it compact or I need to wrap up the task quickly. And we think that that's like the wrong way to think about these things.
35:01Eno Reyes:You should just be able to interact and interact and interact and solve the task over an arbitrary period of time. And so I think that one thing that's nice about this increased quality is our users can feel it. I think we hear from a lot of people, you know, I don't even need to think about context or compression when I'm using Droid because it sort of abstracts it away. And we don't even like put the context, you know, percentage in the tool. We might actually give you the ability to view it if you'd like to, but we don't put it intentionally because our philosophy here is it doesn't matter.
35:34Eno Reyes:You don't need to worry about that anymore. Exactly. You don't need to worry about that anymore. So I think the practical implication is it takes away a very annoying part of using a coding agent and just makes it feel like you can keep going. Stepping outside of coding, this feels like an approach that could be useful for any long context task you have. Could this help solve memory context window issues for non-coding agents as well? I think this is probably one of the more interesting things about software development agents. It really seems like maybe the best general agent is just a software development agent, right?
36:17Eno Reyes:Because I mean, why were humans successful in our environment, right? Tool use. We could pick up tools and we could make them and we can manipulate our environment. What are the tools and the environment of a software development agent? It's all software. Code. Yeah, exactly. Our view is that especially as we've seen users, like there's a very large, you know, software native company, like a digital native public company that uses droids. and there are people in their finance team that are using droids right now because of its ability to do data analysis interact with csv files and spreadsheets that was very surprising to us and i think that one of the bigger trends we're seeing is on our team our ops folks use this like there's no more need for somebody on ops to ask a question like how what percentage of our users didn't uh you know, their billing looks like this instead of that.
37:13Eno Reyes:And can you get their usernames, turn it into a spreadsheet, and then send an email out to them? Because what they'll do is they'll open Droid with a specific skill for accessing data from our team. And Droid will just pull all this stuff together automatically, right? And so our view is that clearly, increasingly general tasks can be executed by software development agents. And so I would bet that really the only difference between the most capable general agent and the most capable software development agent is that on our branding, we write that it's a software development agent, right? I was gonna ask that.
37:54Would you ever launch a general purpose? Because I think that's the problem with Cloud Code, right? Is that it says Cloud Code, so people don't think of it in that way. Would you ever launch a general purpose agent?
38:02Eno Reyes:Yeah, I think it's an interesting question. Our mission is to bring autonomy to software engineering. So I do think that in order to stay focused, we definitely think of software engineering as our primary audience. But I think that it is sort of undeniable that not only with Droid, but also some of the research projects that we have in the pipeline, when you see it execute just a very large scale task that kind of doesn't have anything to do with software, but does it better than any other AI system you've ever seen, you start to feel like maybe keeping it accessible to the broader population is a good idea, just in case they decide it's interesting.
38:41Eno Reyes:And I think that the CLI form factor honestly keeps a lot of people out of it. And so that's why we have our web app, our desktop app, our CLI app, so that you have the optionality. And then I think that there's some exciting stuff in the pipeline for our native desktop and web app that I think will make the experience of using it across a large variety of tasks really pleasant. And I think that for a lot of industries, they haven't had their coding agent or software development agent moment. And very soon, especially in the first half of this year, I think a lot of industries are going to see that sort of magic moment where they realize, man, a lot of the boring stuff that I used to do can be fully delegated and I get to focus on just the fun stuff.
39:26I'm suddenly drawn back to the idea that when we met in person a few weeks ago, I told you I was going to try to write with wine. And I haven't done that yet. But that's on my list. I may do that this week. I absolutely should have just to see how it goes. I was thinking like with long context writing, you know, thinking larger projects, you know, the idea of a book or a really significantly large or dense report maybe. I feel like there might be an application there that could be a lot of fun to try. I was going to say it's actually like a, I don't know the right phrase for it, but like a business sin that Microsoft and Google or even OpenAI and Anthropic have not made an agentic writing tool.
40:11Like nobody has made a good agentic writing tool where you can like work with words like for a document you're making just like a coding agent. Like why does that not exist? and can you make one?
40:24Eno Reyes:Yeah, totally. I know, I think, so this is actually very fun. So here's what I'd recommend because I know a couple of people that do this with Droid. You know, the file system is a beautiful abstraction, right? So if you have a folder on your computer that a Droid has access to, I think that what people sort of, what they immediately jump to is like the chat GPT paradigm of, you know, I give a chat GPT a bunch of information, write me a draft okay now take that draft or take that outline and expand let's start with chapter one it's a very like i go straight to the artifact right like i just want to go and create the piece of writing my view is that you should take inspiration from like the the tolkien side of things right like build the world with the agent right so you have a folder that's just like all of the things that matter i'm going to take a fiction approach because i think it's easier but you could easily apply to nonfiction as well, like research or something.
41:22Eno Reyes:You know, you have a folder that's like, here's the world, right? You have a folder that's like, here's the characters, here's build their personalities, how would they interact with each other, right? Here's the third folder, which is the research into what is the exact sort of like nitty gritty details of this world that we're building. You know, the fourth is the outline, right? And so you take all of this context and then the thing that you're writing suddenly has very dense information from many sources that you're sort of in alignment with and the writing will become higher quality because in its context it always knows these like pieces of information so what i've seen is a lot of writers will use you know agentic software development tooling to basically build out tons of ephemeral uh content yeah exactly uh and then they write and execute the the document in whatever strategy they find apt.
42:16So you've mentioned that code bases need to become agent ready, you said earlier. What does that actually mean in practice?
42:23Eno Reyes:So there are several things that go into making an agent ready code base. But generally, the way I think about this is everyone has a very clear idea of what they sort of want to do in the future. They want to just send off an agent and have it. They give it a loose description of the task and it comes back and it's done. It works. It passes every test. A senior reviewer looks at it and goes, nice work. That looks great, right? There's these other sort of things that people want. They want incident response to be automatic. They want documentation to be generated on the fly. And I think that what people sort of underestimate is that in order to get there, the current quality level that you'll see is totally a function of your agent readiness.
43:05Eno Reyes:There's some great research out of Stanford that actually looked at hundreds of global 2 ,000 businesses that are adopting AI. And the question was basically, how can we predict whether or not this company was accelerating or decelerating because of AI, right? Which it turns out, there are plenty of companies that are actually decelerating their productivity because of AI, which is a little counterintuitive. It is. What's the key insight? Yeah. Yeah. I mean, so this is what's interesting is they looked at adoption. What percentage of their users are actually using AI. They looked at token usage.
43:41Eno Reyes:Uh, what for like, are they using lots of tokens, right? Everyone can adopt AI, but not use it. And maybe that's what, what causes it, right? They looked at, um, uh, power users. So like the density of maybe organizations with a dense group of power users are going to outperform because they have a culture. Um, they looked at all of these usage and adoption metrics for AI tooling, not a single one of them correlated with success or failure the only thing that correlated with success was whether or not the code base was clean along this like 20 dimension metric they used for looking at things like linter presence uh you know type tracker presence unit test coverage end-to-end test coverage uh and and in fact you can predict failure by companies that adopt ai very heavily that have unclean code bases which is pretty intuitive right basically slop in, slop out.
44:37Eno Reyes:You have a bad code base. You give everybody, I think that one VP of engineering that we sold to called it toddlers with machine guns, when you give coding agents to people who can't necessarily take advantage of them. And so - I feel that way sometimes. Yeah, no, I completely relate to it, right? And so what we did is we said, okay, let's make sure that there is a very clear set of criteria, right? So we look at style and validation. So your linters, your type checkers, your code formatters. We look at build system quality. So is the build command documented? Are dependencies pinned? Do you have CLI tooling available?
45:14Eno Reyes:We looked at tests, unit tests, integration tests, documentation. So agents.md, readme's, dev environment quality. So you have an environment template, a dev container, debugging and observability, are structured logs present as their distributed tracing, security, branch protection, and task discovery. Basically, are there ways for developers to jump in and contribute effectively. And across all of these dimensions, there are tons of these individual signals. And so the only way to detect whether or not this is present in a very complex code base is to use the software development agent to search.
45:50Eno Reyes:And so you can activate this by running like slash readiness in the droid, and it will enter into a mode where it spends a bunch of time searching and exploring your code base to determine how agent ready it is. And you can automate this across every code base in your company and actually get an aggregate score, not just for your individual code bases, but for your whole company to understand where are we in our journey towards autonomy. And it will help explain why certain people at your company are like, AGI is here, software development agents are amazing. And other people are like, I can't get this thing to like write a hello world.
46:27Eno Reyes:Typically, we see that explained very quickly by the agent readiness of those respective code bases. Do you have any intuition following that same line of thinking how a company could make their the rest of their business agent ready? I mean, this is kind of what Palantir does, right, is like with their ontologies where they try to say, look, like, let's look at the company and look at the relationships that exist between the commercial entities and the internal organizations and what information travels between them. And then let's say how much of that is actually structured in as data or as information.
47:00Eno Reyes:And so there's definitely probably a hundred billion dollar company in figuring out a way to apply basically agent readiness to every workflow that is executed at a very large business, right? Because finance, right, can be agent ready by making sure the CSVs are all in a centrally stored location, that the workflows for things like checking up on billing are all written down and documented. that everything is parsable by an LLM. So we don't use PDFs, we now use markdown files, right? Like, so there's all this stuff that can be done in every sector of the business. And I do think that that's actually like a really powerful concept is applying this to everything else.
47:41Where do agents still kind of fail today? And what do you think are the most important guardrails?
47:48Eno Reyes:I think that there is a ton that agents cannot do fully autonomously. and it especially becomes visible when you exit the like like one to two hour delegation mark I think a lot of our team in a very agent ready code base sees success basically setting off a task and coming back in like 45 minutes an hour and generally getting pretty much exactly what you wanted to get that means our velocity is insanely high however there's still a lot of tasks that aren't like one hour tasks right they take you know many days worth of effort you need a lot of guardrails and for us we add a bunch of different stuff ranging from autonomy controls that let you set low medium or high autonomy which basically dictates what can the agent do on its own versus with your input low autonomy is like read only like file edits file creates medium autonomy is everything that's reversible so it can do stuff locally on your computer but it just needs to be reversible commands and then high autonomy is you know the breadth of commands.
48:49Eno Reyes:We also have allow lists and deny lists. We have a really long built-in deny list. So it can't, you know, delete your hard drive or anything like that. Like programmatically cannot do those things. We've seen that happen. Yeah, I know. I mean, people, it's crazy because people would just enter into things like dangerous YOLO mode with coding agents prior. And then they, crazy stuff happens. And so for us, we know people want autonomy. So we said, we got to give people and option in between dangerously run and please don't delete my hard drive and like i have to accept every single command that occurs right um there's also a ton of guardrails like yeah yeah exactly um we have something called droid shield which basically prevents you from committing secrets and prevents you from doing things that would you know oh osp 10 style errors it runs totally on device so it just makes sure that you don't do anything too silly uh with a droid um And there's a lot of other capabilities that especially our enterprise customers have access to with custom hooks that we provide them that basically prevent a number of different specific strategies for failure for droids.
49:54So it sounds like if you're just starting with this is very powerful for engineers, but if you're just starting with engineering, that this also is a good tool for you. Do you agree with that or what do you think?
50:04Eno Reyes:I think that this is probably one of the best times in history to want to enter into software development. or dabble in building software. I think that it's clear that with a Droid, people who have zero experience writing code, who are willing to understand like Droid, not the code, but just how Droid works, can make huge, huge effort, can make basically the level of software that I would say like entry level, new grad software developers like three to four years ago would be able to do. and that's that's within like a week probably I've seen it happen like we often get emails from people saying like I actually am coding for the very first time and I picked up Droid and I built this crazy thing and I'm looking at it and I'm like I wouldn't have I wouldn't have been able to tell that you've never written code before and I think that that's only gonna gonna increase that software development will become more accessible but the bar is also gonna massively raised.
51:07Eno Reyes:So I think professional software developers now have the most powerful tool in technological history in front of them, right? Like AI systems that can effectively do any tasks that can be done on a computer. So I think as a software developer, because you understand how these things work, you'll also be able to go even further. The rise for factory has been super fast. And I was wondering, what's it like to be a part of something growing at that pace? And what excites you the most about the next year ahead for factory? The coolest thing about working at factory is being able to be surrounded by, you know, what is clearly the most talent-dense, interesting, and motivated group of people that I've ever had the pleasure of working with.
51:54Eno Reyes:Like, the team here is truly next level. And I think that the, you know, just outside these doors right here, you know, You have people who are responsible for, at many companies, Stripe, at DeepMind, at Uber. They were integral to building software that we all use on a daily basis. And now they're basically single-mindedly focused on bringing autonomy to software engineering. So when you bring a group of people that are extremely motivated and very talented to work very hard on a very hard problem, you get this like very exciting moment where uh you know it just feels every day like something new is happening you're constantly surprised uh you know and and and i think that all of this just really makes for a really exciting and fun time uh you know no matter what the next couple years look like i think that already just being able to meet these great people is like everything that i could have asked for do you have any advice because you are clearly very good at building agentic tools and working with agents.
53:00Do you have any advice for people who are working with agents or trying to build their own agents? Like, especially if let's say they're even new to working with agents or working in this field.
53:11Eno Reyes:I think that my advice would be to try to understand what part of differentiation you'd like to introduce when building your agent. You know, I think that there's like the classic statement about, you know, basically like build versus buy, like don't build anything that doesn't contribute to your core differentiation or competency. If you are building a product that is agentic, you might consider using something like Droid in exec mode, basically our like headless CLI that gives you an agent that lets you wipe the system prompt that gives you compaction that gives you all of the like bells and whistles, but then lets you fully customize how you want this agentic system to work.
53:54Eno Reyes:If you want to bring it more into like a programmatic session you can use like anthropics agent sdk but if you want to build the agent from scratch and replicate things like compaction and own things what i would say is pay a lot of attention to the use case you care about and don't worry about all these like details like compaction you can honestly probably paste our blog post and then paste that into droid and then paste your code base and implement something that's basically just as good as our compaction if you really want to. But don't worry about that stuff. Try to nail your hero user journey as quickly as possible with your agent with an existing framework.
54:36Eno Reyes:So use Droid, use Agent SDK, and just nail the user journey. Once you've done that, I think it'll become clear to you what parts of it you actually need to focus on versus what you can probably offload. Well, Eno, thanks so much for joining us today, man. I really appreciate it. Thanks for having me. This has been super fun. Where can people go to learn more about Factory? Maybe give it a try. Just go to factory.ai, and if you're on Mac or Linux, you can curl the URL at the top. Or if you're on Windows, same. And then if you want to sign up and use our desktop web platform, just click Log In. If you're new to this type of stuff, which one do you recommend?
55:13DAP?
55:13Eno Reyes:I would recommend, if you're willing, the CLI. I think it's a really simple and fun way to use the tool and has most of the bells and whistles in it. Thanks to everyone who joined us today. If you haven't, please take a moment to like, subscribe to the channel for more interviews, tutorials, and all of the things we work to bring you every week. Also, pop by the neuron.ai and sign up for the newsletter that started it all for us and join more than 600 ,000-ish others who get their daily dose of AI news from the neuron each morning. and on that note that's it for today farewell for now humans
From the publisher
Autonomous coding agents are moving from demos to real production workflows. In this episode, Factory AI co-founder and CTO Eno Reyes explains what "Droids" really are—fully autonomous agents that can take tickets, modify real codebases, run tests, and work inside existing dev workflows.
We dig into Factory's context compression research (which outperformed both OpenAI and Anthropic), what makes a codebase "agent-ready," and why Stanford research found that the ONLY predictor of AI success was codebase quality—not adoption rates or token usage.
Whether you're a developer curious about autonomous coding tools or just want to understand where AI engineering is headed, this episode is packed with practical insights.
🔗 Try Factory AI: https://factory.ai
📰 Subscribe to The Neuron newsletter: https://theneuron.ai
📖 Resources mentioned:
• Factory's compression research: https://factory.ai/news/evaluating-compression
