In short
Building long-horizon autonomous AI agents for real-world work (especially accounting), emphasizing why outcome-only evals fail and how to enforce “process” via behavior specifications and verifiers.
Guest backgrounds
Mitch Torianovsky is co-founder of Basis, an AI company building agents for end-to-end accounting tasks. Basis agents can run autonomously for hours to days and have completed complex work like preparing tax returns end to end.
Key claims
- Humans already work with non-deterministic systems; agent design should mirror human coordination and verification processes.
- Passing many evals (e.g., 100/100) doesn’t guarantee production generalization; agents can “cheat” by using superficial sources or brittle reasoning.
- Long-horizon requires coherence beyond LLM context limits; modern reasoning models improve long-context performance and self-healing.
- For accounting, verifiable runtime signals exist (e.g., spreadsheet checks), but much quality is judgment-based, so you need process/behavior checks.
- Basis uses “behavior specs” (Markdown) to define and judge required behaviors, not just outcomes.
Notable examples
- Tax research: correct answers aren’t enough if sourced from Wikipedia/blogs; agents should verify via primary sources (e.g., IRS).
- Accounting checks: trial balance arithmetic (e.g., totals tie out), Excel integrity, and missing/corrupted tabs.
- “Whispering” context: Basis uses voice/microphones to provide richer context faster than writing.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding Non-Deterministic Systems
0:00 to 0:44
Explore how humans and systems interact in non-deterministic environments.
“Humans are already used to working with non-deterministic systems.”
Whispering into the Future of AI
1:24 to 2:54
Mitch describes the unique practice of whispering thoughts into microphones at Basis.
“As I was prepping for this, I came across a video by our friend Stephanie Balazzolo at The Information.”
Overview of Basis and Its Mission
2:54 to 3:40
Understanding Basis's focus on building autonomous agents for accounting tasks.
“I mean, different people use different things.”
The Complexity of Accounting Work
3:40 to 6:10
Mitch elaborates on why accounting presents unique challenges for AI agents.
“Yeah, basis builds agents to do accounting work end to end.”
Defining Long Horizon Agents
6:10 to 8:14
A discussion on what constitutes a long horizon agent and its implications.
“This is still the same thing for contextual awareness.”
Autonomous Agents in Action
8:14 to 11:14
Mitch explains the process of how agents handle complex tasks like tax returns.
“And still to frame the conversation, what we're talking about here is autonomous agents.”
History and Evolution of AI Agents
11:14 to 14:01
A look back at the development of AI agents and the context of their evolution.
“And so to me, that's I think what it means to be autonomous.”
Failures of Baby AGI: The Compounding Errors
14:01 to 15:39
Learn about the limitations of Baby AGI and the role of reasoning in AI development.
“reasoning to build out your environment, whether that be with sub-agents or compaction.”
Paradigm Shifts in AI Reasoning Models
15:40 to 17:49
Explore the significant breakthroughs in AI reasoning and their implications.
“And now suddenly that just compounds and you have no way to have the self-awareness to actually like self-heal in any meaningful way.”
The Evolution of Long-Horizon AI Agents
17:50 to 21:42
Understand how advancements in training are influencing long-horizon AI agents' performance.
“And so the ability to reason allowed you to titrate that.”
Show all 35 chapters
Challenges in Non-Coding AI Agent Tasks
21:43 to 25:07
Discuss the struggles of AI agents in real-world tasks beyond coding, focusing on verification.
“Well, that requires kind of a theory of mind.”
Designing Effective Agents: Learning from Human Processes
25:08 to 28:00
Discover how human processes inform the design of effective AI agents and systems.
“I think the verifiable rewards are maybe just like the beginning that allows that stuff to happen.”
Designing Systems for Non-Deterministic Entities
28:00 to 29:27
Learn how to design systems that facilitate coordination among humans and non-deterministic agents.
“It's just the systems are normally their coworkers, not their computers.”
Challenges of Data Scarcity in AI Training
29:27 to 30:54
Understand the limitations and challenges of using real-world data for training AI agents.
“let's say just for sake of argument here, you had, uh, not just synthetic, but every like real tax return, uh, across the country, which you actually couldn't do for, for privacy reasons.”
Complexity of Tax Return Processing
30:54 to 32:51
Explore the extensive steps involved in processing tax returns and the implications for AI agents.
“And it could actually take much longer for very complicated returns.”
The Importance of Process Over Outcomes
32:51 to 35:16
Learn why relying solely on outcomes can be misleading in evaluating AI effectiveness.
“So to play it back, you got very complex processes with many, many, many steps.”
Behavior Specs in AI Design
35:16 to 36:45
Discover the concept of behavior specs and how they guide AI agent actions.
“potential in the models, what you really need to do is think about, let's take the learnings from how humans do things, from what good process looks like.”
Collaborative Development of Behavior Specs
36:45 to 38:41
Learn about the collaborative process between accountants and ML researchers in defining behavior specs.
“Yeah, so the original idea, actually my co-founder Matt came up with the idea.”
Balancing Precision and Flexibility in Agent Behaviors
38:41 to 40:48
Understand how to balance specificity and adaptability in defining agent behaviors.
“condition occur that it would need to exhibit this behavior?”
Future Research Directions in AI Training
40:48 to 42:00
Explore future avenues for improving AI training methodologies and model evaluation.
“I think as people who are good at building agents know, it's much better at the margins to be able to give principles and the whys and more context and let them figure it out.”
The Importance of Agent Behavior in AI
42:00 to 43:58
Discover how agent behavior influences AI product development.
“you went from 20 minutes to 20 minutes, 30 seconds, do it every time.”
Judging AI Behavior: Methods and Challenges
43:58 to 47:19
Learn about the complexities in training AI judges for behavioral evaluations.
“There's the harness, what are the tools, what are the capabilities, the environment, there's all that part of stuff.”
The Trade-offs in AI Behavior Definitions
47:19 to 50:03
Understand the pros and cons of defining AI behavior in precise terms.
“If the way you are making the model or the agent exhibit this behavior is through context, now you have potentially duplicate state.”
Building Intuition for AI Models
50:03 to 51:46
Explore how to develop a deeper understanding of AI models and their functionalities.
“As I listen to you, one of the many things I find fascinating is that you're doing all of this without having actual insight about how the underlying model works.”
The Project with BrainTrust: Open Source Collaboration
51:46 to 54:57
Learn about the open-source project with BrainTrust and its goals.
“because you start to understand, it's like, okay, we'll wait.”
Collaborating on Open Source AI Standards
56:00 to 58:28
Learn about the collaboration behind an open source project for AI agents.
“So it was a Sunday night as we were recording this, and you showed how you were replying to that Slack at 6 a.m.”
The Importance of Data and Ontology for AI Agents
58:28 to 1:00:25
Understand the significance of data structuring and ontology in AI agent performance.
“I think the coolest thing would be for people to contribute ideas to it.”
Designing Effective Ontologies for AI
1:00:25 to 1:02:35
Discover how to design ontologies that enhance an AI agent's functionality.
“They have certain behaviors encoded in them in the context, right?”
Systems Thinking for AI Roles
1:02:35 to 1:06:48
Explore the skills and backgrounds necessary for new roles in AI development.
“effectively very compacted context has to get it back into the state of mind of its entire lived experience.”
The Role of Deployed Intelligence Teams
1:06:48 to 1:10:02
Learn about the purpose and function of deployed intelligence teams in AI integration.
“You need, I think this has probably gotten somewhat in vogue recently, but I think one One skill that really matters is good systems thinking.”
Transitioning Accounting Firms to AI
1:10:02 to 1:11:11
Learn how the DI team helps accounting firms adapt to new AI technologies.
“Um, it's, uh, it's people who, uh, have a deep empathy and understanding of the profession of like what it takes to put in place good process.”
Self-Improvement in Autonomous Agents
1:11:11 to 1:14:46
Explore the concept of self-improvement in AI agents and their evolving capabilities.
“maybe as we start getting to the end of this conversation, I would love to spend a little bit of time on that.”
Challenges of Reinforcement Learning
1:14:46 to 1:18:38
Discuss the complexities of reinforcement learning and its implications for AI models.
“to start doing actual reinforcement learning on behavior adherence.”
Business Moats vs. Technical Moats
1:18:38 to 1:20:46
Understand the importance of business strategy over technical advantages in AI.
“like every single profession in the economy.”
Advice for AI Builders
1:20:46 to 1:22:29
Gain insights on strategic thinking and coherent system building for AI development.
“there are certain things that you care about for that that are subjective.”
Transcript
Automatic transcript. May contain errors.0:00Humans are already used to working with non-deterministic systems. It's just the systems are normally their co-workers, not their computers. And in many ways, like companies and processes, it's all about how do you design a system for non-deterministic entities to coordinate together to solve a problem? And once you realize that, it's like, well, now it's like kind of ancient design. Let's say you have 100 evals. Great. They all pass. It looks good. Are you confident that that now generalizes to the real world, to production? And our answer has been no. Even if you got it right 100 out of 100 times, if a person is just getting it right because they're going to Wikipedia, the accounting firm wouldn't hire them.
0:32And so they shouldn't hire us either. You'll see people like freaking out over a code file that isn't abstracted properly, and yet their context is like total s***. The English is more precious because the English affects the performance. The code does not affect the performance. Hi, I'm Atter from FirstMark. Welcome to the Mad Podcast. Everyone is building AI agents, but outside of coding, most still can do real work reliably. My guest today is an AI builder at the forefront of cracking that problem. Mitch Torianowski is a co-founder of Basis, a unicorn AI company whose agents run autonomously for hours, sometimes days, and are already able to complete very complex tasks like preparing entire tax returns end to end.
1:09This is a true reference episode on how to build long horizon autonomous agents where Mitch shares tons of lessons he learned along the way. Please enjoy my conversation with the deeply insightful Mitch Torianowski. I want to start with a scene. As I was prepping for this, I came across a video by our friend Stephanie Balazzolo at The Information. And she was describing the experience of walking into the basis office and seeing a bunch of people whispering very quietly into microphones. So, you know, maybe for the top AI builders or people who live on X, you know, second by second, this may be already something that everybody understands.
1:53But I think for the vast, vast majority of people, like just describe what you guys are doing, whispering into those microphones. Yeah, I think maybe the best piece of advice for, not for building agents, but for working with AI in general is that you need to give it as much context as possible. because it, by definition, is always missing context in some way. Speaking is just so much faster than writing things down. And in fact, when you try to write things down, you are actually essentially trying to summarize all the crazy thoughts in your head. And so that's why it takes a lot of time. And it's useful for your eye because it's rude if somebody just blabbered and sent that to you as a Slack DM.
2:32But to an agent, they don't care. They, it's actually, they would prefer it. So it becomes much more productive to be able to, like, whisper your thoughts, because you don't want to be shouting. And you have these microphones now that allow you to whisper very, very quietly and still pick up with full fidelity. So that's why we have it. And sometimes people see it and they think it's a little weird when they join the company. But after a month or so, it's like they can't go back. And so you whisper into what? Into Cursor or into... Yeah, into whatever people use. I mean, different people use different things.
3:05But yeah, Codex or Cloud or Cursor or whatever people use. And not just engineering, right? Like all the functions if you're trying to get something done, if you're trying to describe what you want and all these things. Okay, great. All right, so what I'm hoping to do today is a bit of a reference conversation on all things around building long horizon agents that's in part based on a great thread that you had on X and perhaps more importantly, a new open source project that you just released in collaboration with BrainTrust. We're going to talk about all about this, but maybe for contextual awareness basis in two or three sentences?
3:43How would you describe it? Yeah, basis builds agents to do accounting work end to end. And accounting is difficult. It's not something that is just purely text in, text out. And so it requires the ability for like AIs to be able to, you know, perform lots of actions over long periods of time and actually, you know, be coherent over that period of time to get to outcomes that are good. And that's why we've always, you know, been very focused on how do you really build agents that can scale to do that work. And did you pick accounting because of how interesting that was from an agent-building perspective or the other way around?
4:22That's a good question. It's probably the other way around, but I do think it is actually quite interesting from an agent perspective. Accounting is interesting for a lot of reasons. It is one of, if not the largest, knowledge work profession in the country. There are over 3 million combined kind of accountants in the country. And what I think is so cool about accounting actually, is that most people don't really think about accounting. They don't think like, well, yeah, why does that even there? You know, probably most listeners have never thought like, why does it even exist? And I know we're going to talk about agents, maybe quickly 30 seconds, just to convince everyone how cool accounting is.
4:58If you think about the real world, so much stuff happens, academic activity, right? Like, you know, I was just drinking a water bottle there, like, you know, someone, um, uh, that bottler had to choose to like go, uh, uh, buy from that factory or that supplier, um, or decide to open, you know, some additional, uh, store or hire a salesperson. And these are all economic decisions that stem from understanding the real world. What's in the real world, you know, money moves hands, someone signs a contract, someone delivers the inventory. It's like all these events that occur. and so much of modern capitalism relies on the ability of all these actors to make decisions on these events right like the ceo of that company the irs obviously to decide how much to tax the bank to lend credit investors right all these people they care about the real world but they can't understand it because it's gigantic and it involves all of this unstructured and you know difficult to parse information and accounting is actually the art of compressing all of that into something that is structured that now people can look at and understand and make decisions.
5:59So something about accounting, you could argue in a meta way, is kind of like an intelligence over the economy because it is really a compression activity of all the information that exists. So I think it's a very cool problem to kind of think about. This is still the same thing for contextual awareness. We're going to talk about long horizon agents. What is the long horizon part these days? That keeps evolving. What falls in that category? Hmm. So maybe I can give my quick definition of an agent. I know maybe probably everyone knows at this point, but I feel like that's a gotcha question I like to ask in interviews.
6:33I tend to think of an agent as an AI or like some inference that occurs that has the agency to go and make decisions to do different things. And so by definition, it's a spectrum because you can have varying degrees of agency, right? Like you're constrained by whatever environment you're placed in. And I think long horizon, again, is a spectrum where you're granting the agent the agency to make decisions that allow it to be coherent for longer periods of time. Right. So let's say that you were asking an agent to go and, you know, look up the weather for you. It might be an agent in the sense that it has the agency to decide what tool to call or what Google search to put in.
7:15But it doesn't need to do much work to be coherent over a period of time because you're just getting the weather. But if you're asking an agent to, say, go perform an entire feature, implement some feature in your repo, or asking it to go and make a big Excel workbook, Now suddenly it might have to operate for longer than a minute. We're talking 10 minutes, 20 minutes, 30 minutes, and potentially even much longer than that. And once you're starting to get into those scales, you start running into the fundamental limits of how LLMs work. In which I always like to say LLMs have very large working memories and by default no short-term or long-term memory.
7:57And so you have to leverage these strengths of the LLM to make up for the fact that you don't have good or actually any real short-term or long-term memory by, and we could talk more about it, by using harnesses and all these kind of advancements to allow them to be coherent over a period of time. So I think once you start getting into the art of trying to get it to be coherent because you're going past the amount of working memory it has, I'd probably say that's when you're starting to get into what I'd call long horizon. Great. And still to frame the conversation, what we're talking about here is autonomous agents.
8:31I'm curious, maybe just as an example, what autonomous means in the context of basis. I read that you guys can now have agents that handle end-to-end tax returns. So maybe walk us at a high level through what that looks like in terms of steps along that takes. What does an agent do conceptually? When you are really autonomous or kind of, you know, doing something on a really long horizon, say doing a tax return on 10, that means that you have a lot of information that is needed to do the work and you have the tools to go and get potentially more information. so let's say imagine you're doing a complicated 1065 and you have all of the different um uh you know k1s w2s other documents 1099s whatever you need um from the uh from the company uh and then you also potentially depending on what you're doing you might have the trial balances already so that tends to be what you need to actually start a tax return or you're working with like books that aren't even done yet.
9:39And the agent then has to actually go and figure out based on all this different stuff, how is it going to tackle it and what's it going to be able to do? And that's where you start getting into, you know, some stuff about like what the behavior should be that we can talk about about, well, what does good practice look like to say, get to a solid set of trial balances? What does good look like in order to properly extract out, you know, the K1s and the K3s so that you can be confident in the outcome? and for it to be autonomous, it means it's not going back to the user and saying, hey, like, is this right?
10:09Is this right? I need this. I need this. It's like starting the job to I'm done. And I'm done does not mean I'm done. You click a button, like no one looks at it. It's actually the opposite of that. It's much closer to what you can imagine a preparer doing or maybe like a first pass or junior engineer or something of I'm done. Here were the big decisions I made. Here were my assumptions. Here were the different things you need to look at. Let's go and review together, right? And if you think about somebody, say, an engineer handing you a PR, nobody likes being handed a thousand line PR. They're like, it's done, I promise.
10:42It's like, you don't want to review that. But if you instead handed somebody a great stack that was properly split out and you could understand very easily, hey, here is exactly what this change is and this diff. And I made this big architectural assumption here and here is why I made that change. and you can like optimize not just for getting the work done, but for making it easy for your reviewer to understand the decisions that you made. And that's obviously very true in software engineering and it's actually true in, I think, most professions and especially accounting, which we can kind of get more into.
11:14And so to me, that's I think what it means to be autonomous. So I thought what would be fun and helpful for people listening to this would be to spend a few minutes on, I guess, the history of agents. Like we've all heard over the last two to three years so many different things, so many different terms, some projects that work, some projects that didn't work. So I think it would be helpful to just like go back in time just a little bit. What in AI may feel like a prehistory, but in reality is like what, three years ago, four years ago. So maybe starting in 2022 with the React framework, so not the software engineering, but like reasoning and acting, which I believe was a paper in 2022 that fundamentally said this agents are a combination of like reasoning and acting, which you just alluded to.
12:00The fundamental question is that, is that still largely what's happening? I mean, with, you know, a tremendous level of sophistication on top of it, but like is the fundamental theory of an agent still that? I think within the paradigm of modern day LLMs, yes, more so, more or less, in the sense that I would say it is actually the same, but it kind of extends out further, which I don't really remember if it was part of that paper back then, of needing to use your reasoning to regulate your own state. Like the analogy I like to always give to people at the company when I'm learning about agents for the first time, not even just technical people, like anyone at the company actually, is the movie Memento.
12:45I think Memento, for those who haven't seen it, is a movie in which there's this guy who has short-term memory loss and every day he wakes up and he knows who he is and he knows, you know, like he knows he's a human, like he knows like some basic stuff, but he doesn't know like what's happened in the last couple of years. He has no idea. And for him to make progress to any particular goal, that could be something like, you know, getting groceries or getting revenge or whatever it is. He effectively needs to write stuff down for himself. And then when he wakes up the next day, he reads his notes and essentially builds that up.
13:18That's how he like builds up knowledge. Early Christopher Nolan movie, by the way. Yes. At the time when everybody obsesses about the Odyssey, this is one of the early these early works yes I still I think The Procedure is the best Nolan movie but but yes but anyway so I think that's your point it is actually about reasoning but I think and I don't think it talked about in the paper it's just there's reasoning in terms of what path it's right to go for whatever the task is like do I do taxes that way for sure but there's also reasoning about how do I make sure my next inference step has what it needs to properly like you know interact with the world, which is easy if you're in a short horizon because you're within the context window.
13:57But once you start getting to longer horizons where you're beyond the context window or you're getting to context rot, you need to kind of use and sort of brute force your reasoning to build out your environment, whether that be with sub-agents or compaction. We talk more about that, but that's, I think, where the reasoning becomes super important. Okay, great. So the next step after 2022 was probably baby AGI in 2023. which, you know, everybody viewed as just like a massive advancement. But that sort of failed. I mean, that was a beautiful experiment, but like didn't quite work out in practice.
14:34So why is that? I guess at the time, like people talked a lot about compounding errors and, you know, how if you had many, many steps and the agent started going astray, then that would compound. Was that what it was from your perspective? Yeah, I think back then, I think that I believe if I remember correctly, that was like GPT for maybe four turbo at the time. The models back then for starters, if it was pre-turbo like the context windows were very small and so you know, if you're going to be coherent, you need to have at least enough stuff in your context that you can organize your own environment.
15:10So they didn't have that. And even when the context windows got larger, I don't think by the time BabyHCI came out, the models were actually good at keeping attention over like once you got past whatever it was like 20 ,000 tokens, they were just not smart. It took until maybe Opus 3 for them to start getting smart at like 100 ,000 tokens even. And so that was, BabyHR didn't have that. And then also, obviously, I'm going to talk more about it, they just were not good reasoners. And so if you're not a good reasoner, then you're going to have lots of compounding errors because you're going to make one mistake that's, you know, in your kind of per token generation.
15:44And now suddenly that just compounds and you have no way to have the self-awareness to actually like self-heal in any meaningful way. So the big breakthrough then was reasoning models? Yeah, I would say when I think about what were the big like, holy shit moments for, I don't know if it occurs, but yeah, like what were the big kind of holy shit moments for us? At least for me personally, it was probably, or at least the moments in which there was a genuine paradigm shift. There haven't been that many. I would say they were Opus 3, which I think goes underappreciated but I think was the first model to truly be able to like actually understand at long context.
16:28Before if you put anything in the like 80 ,000 tokens in the GPT-4 turbo it could not understand it. Opus 3 actually understood it which was remarkable. I think it was that. I think it was 01 obviously. Everyone talks about that and then I think after 01 it was 03 because I think 03 helped prove that not only could you scale the amount of reasoning at inference time, but with better training, with more compute, better data, et cetera, in the post-training phase, you could make the reasoning higher quality, more efficient, and just better. And so each incremental token that it reasoned with at inference time was higher quality, which was not obvious when 01 came out.
17:05So I think that was, I think those were to me the big breakthroughs. Is there something about the fact that those were effectively trained as agents in loops where there's a chain of reasoning where they try something, backtrack, try something else? Is that the fundamental reason why this works better for agents? I think it's a couple of things. I think it's that because the model is able to kind of like titrate the amount of compute it's throwing at any specific step in the process, you're able to, in any trajectory, there are a lot of decisions that are really hard and some that are really easy.
17:39And it's just not feasible to have like, you know, some giant parameter model that's super expensive to serve that has essentially been using all of that compute at every single inference step. And so the ability to reason allowed you to titrate that. And I think, as you pointed out, as a kind of maybe emergent behavior that if you are a really good reasoner and you can dial it up a lot, like if you actually look at the amount of compute for an easy step versus a hard step, it's a lot different with the modern reasoning models. You could become better at self-healing because you're going to be like pausing and thinking about everything and sort of using that more, you know, that kind of thinking versus just doing everything on instinct, which is what was happening if you're doing kind of like just per token generation.
18:23In your X thread, you talk about something OpenAI did in 2023 when they published something called Let's Verify Step by Step where humans labeled about 800 ,000 reasoning steps. What did that happen then? And, you know, what was the goal? Yeah. So that paper came out. So back, I think before people, I don't know the exact history here, but essentially there was a lot of rumors back then about, you know, if people remember like the information article, like, oh, strawberry, it solved math or whatever it was. And so I think even before those rumors came out, there was some hints in the literature like this.
19:04And that might have been after the rumor, I actually don't remember. But that math was a – or these kind of problems that you could verify were maybe good ways to try to train the models to be better at kind of different tasks. And I think this is purely speculation because I was not in the labs, but there was this sort of question back then and through a lot of the history of ML around, are you going to give the reward just from like whether it got the problem correct or whether it approached it like a good mathematician would approach it? Um, and what's interesting is that in that paper, which I published, they showed that actually, uh, if you rewarded based on the process rather than just an outcome, you actually got better results within, within that, that sample.
19:53But that obviously is expensive because that required mathematicians to, you know, grade the approach they took to the problem. Right. versus if you look, you know, if you fast forward a bit and you look at the, like the DeepSeq R1 paper where they effectively laid out, you know, what I think all the labs were doing at that time, or at least OpenAI was doing in terms of, you know, RLVR reasoning from reinforcement learning from verifiable rewards, that effectively had very little process supervision. And instead was essentially just saying, hey, did you get the outcome right? Yes. Okay. Let me reward you.
20:28and then scaling that up, which, you know, obviously worked well. Still in the recent evolution of agents, so just, I guess it was last year, there was like this famous, now famous meter chart that shows that the long horizon agents' capabilities like double every seven months. Is that something that you're still observing in 2026? Yeah, absolutely. I mean, I think the meter chart is somewhat inaccurate these days because it's so hard to measure. Also, the bar is pretty low, right? The bar is pretty low. 50 % inflation. Yeah, yeah. The bar is pretty low. And I think I'm not super familiar with their data set, but my understanding is that it's like a low, like the sample size is kind of low.
21:03So I don't know about the specifics of that metric. But I think, you know, from a Vibes perspective, absolutely. Like you kind of the models were able to start being coherent for longer periods by by being smarter. But now they're also being specifically trained to do that. Right. So that means they're being trained at how to, you know, how do you have good theory of mind over yourself? because you need to if you're going to be outsourcing to a subagent, right? Or if you're going to be writing notes off to yourself, going back to the Memento analogy, right? Like you need to think about, hey, I'm going to wake up tomorrow and I'm going to read these notes.
21:39What is the most information dense way for me to write this note down so that me in the future will understand it? Well, that requires kind of a theory of mind. And so, you know, as these models are being trained more on actually how to do that kind of work, which is very, very non-human, like humans don't have to, because we're great at that actually. So we don't have to write for ourselves. You know, our brain does that for us. You're starting to see it get even farther. I think we're still actually quite early at that. If you look at some of the advancements, you know, to Fable 5 and 5.6 Sol and whatnot.
22:08Great. And just to go a little deeper on what you just mentioned, I think, you know, a broad context on agency 2026, the evolution towards autonomous agents would not be complete without actually talking about verifiable rewards. So tell us what it is, where that feels in the overall picture, and then perhaps why coding was the first successful autonomous agent use case and how that's related to verifiable rewards. It's kind of interesting. I think people get this a little bit wrong. I'm speaking with a little bit of speculation because I don't actually know exactly. But coding, yes, is verifiable in the sense that I can know whether the code passes or not.
22:56And so I could train on that. Did you get the problem right? Did you get it wrong? Et cetera, which is useful. But the models being great at coding is partially that, but it's partially, I think, a couple other facets of coding. So number one is that coding is verifiable at runtime. I think that's a very important point because at the end of the day, like an agent needs to get signal as to how it's performing. And so imagine if you were an engineer and you wrote code and you can never run it. Like even if you were the best engineer in the world, you're going to have a syntax error in which the thing is wrong at some point.
23:29And so I think the fact that coding is so easily verifiable at runtime, or at least some parts of it are verifiable at runtime, is obviously very valuable for it being there. And it's cheap to do, right? It's easy to do within the environment and it's super tech space. So it's available anywhere. Like you can just on your local computer, you can just run it. Right. And so I think those aspects of coding actually make it, um, are, are a lot of the reasons it's, it's, uh, the agents are getting so good at it. I think one more piece of it is that, um, and I think we probably saw this if you think about how good agents were at coding maybe a year ago, a year and a half ago, they, uh, they could go and implement a thing you told them, but they didn't have the level of like, you know, taste or level of like, what is good software?
24:14Because even if you train them with verifiable rewards on like, hey, did this unit test pass? Like you could pass all the tests in the world. Doesn't mean that you set up the app correctly. It doesn't mean your database was built well. It doesn't mean that you, you know, split out the files properly, right? At the end of the day, coding is subjective. It's an art. And you're not going to solve an art through verifiable rewards. And so I actually think there's a large part of this, which is the amount of training data and the quality of training data that the models that the labs have clearly for coding has gotten quite good.
24:41And they've focused a lot on making sure it's very good so that they're training on high quality code. So I think that's the other part. It's just that maybe because the labs are obviously full of engineers, it's more top of mind and obviously it's part of their strategy. And so it is more top of mind for them to ensure that both their pre-training and post-training sets have lots of high quality code. And that's what makes the agents great at not just writing code, but now starting to become good at actually engineering. So I think it's all those things together. I think the verifiable rewards are maybe just like the beginning that allows that stuff to happen.
25:12But I think the other stuff matters just as much, if not more. Great. Which brings us to the core of the thesis, which is your work on a domain that's outside of coding. So building long horizon, autonomous agents for, I guess, the real world, for lack of a better term, outside of coding. I use that term. Okay. So why do agents struggle? you mentioned three reasons maybe you mentioned what those reasons are and then we'll go into them turn by turn. Yeah, so I think agents struggle for a lot of reasons I think one, they struggle because they don't necessarily know what good looks like I think they struggle because it may not be easy to verify yourself at runtime as we were talking about with coding.
Read the full transcript
26:04I think another part is that, and this is maybe not an agent struggle, but maybe it's a UX thing, which is that for coding, engineers are just very in the weeds of it. And so there's a kind of a difference where if you were like running a long horizon agent for coding, if the engineer was not, you know, engineering and like in the weeds of the code, if instead they were more abstracted away, your maybe level of quality and, you know, how you make decisions, probably you maybe need a higher bar than you would for coding. You know, coders are okay with lower bars. That's been true forever. And so I think all these things add up in making it.
26:38And even now with coding, like the agents are not yet, they're not human level at being coherent over long periods of time. That's obvious because they can't code like a junior engineer on a project for two weeks. So they can't, that's worse than a human, even to start. So let's take that part about verifiable rewards and the fact that people writing those systems don't necessarily have intuition for what good looks like. So how do you guys solve that? What does passing a test mean for a tax return that unlike code doesn't need to compile? So the good news is there are some things that can compile.
27:23Not all, but you can obviously test. Either you get sued or you don't get sued. Yeah. Well, I think the answer for this is you look at what humans do and how and one nice thing about accounting, which is true in some other professions as well. But accounting is a profession in which you really try and have to be correct. And so because of that, and it's something in which there's so much judgment and process involved where one of the sayings we have, and I say this on the design side a bit, but it's like humans are already used to working with non-deterministic systems. It's just the systems are normally their coworkers, not their computers.
28:03And in many ways, like companies and processes is all about how do you design a system for non-deterministic entities, i.e. like humans, to coordinate together to solve a problem. And once you realize that, it's like, well, now it's like kind of agent design, sort of. And so I think if you needed to think about what is good agent design and what can be verified, you should look at how the humans organize. And if you look at how humans do tax returns. You have steps of verification. You have independent review. You have things that can be deterministically verified. So you can say, hey, like, you know, obviously do like to the TBs, like add up to zero.
28:36Right. That's like a very obvious check. You can things like is the Excel not have errors is an obvious thing. Right. And so there's lots of things like that to human or obvious, but you need to make sure are properly encoded. And there's other things that are maybe not as deterministically verifiable, but like, you know, would be obvious to an accountant to look at it oh this is wrong you know of like oh you deleted this tab in this excel or you didn't cite this thing or whatever it is and so you can start to build you know verifiers effectively from these things that are not uh deterministically verifiable but you know if they're an accountant would look at it it'd be obvious and so you can start to you know think about judges or other you know forms of verification to like get that signal both in your evals but also at runtime So that's one thing.
29:19I think you pointed in your X thread that there was an issue around scarcity of data. Can you go into this? let's say just for sake of argument here, you had, uh, not just synthetic, but every like real tax return, uh, across the country, which you actually couldn't do for, for privacy reasons. But so you did, you don't have that. And you're just like, okay, let me use this as a, as a way to get, um, uh, data that you're not training, but data that allows you to verify how well the agents are doing. The order of magnitude there is tiny compared to, you know, standing up math problems synthetically and you're generating, you know, whatever, a hundred thousands, millions, et cetera.
30:02And so even if you had all the data in the world, you would not be able to scale it. So you need to think about how do you synthetically generate it. And if you think about how to synthetically generate it, that's really hard because you're not just synthetically generating text, you're synthetically generating artifacts that have to be real and diverse. And so now you get to all the same problems about like data, you know, data diversity, all the different things that you need. And I have no doubt that that problem will be solved over time, but it's not solved today. And so there's this kind of gap between like maybe what is possible from a data generation perspective and what is actually possible, what latent capabilities the models have, which is maybe where you get to some of that stuff we were talking about in the thread.
30:46And how do you think about the length of the feedback loop? Yeah, I mean, that's the other thing is the length is can be very long, depending on what you're doing. You know, performing a 1065 can take a human, you know, like 20 plus hours easily of actual work. I don't mean like it took them a day. I meant like literal sitting down work. And it could actually take much longer for very complicated returns. And so there's just no way that you, even if you had all the data in the world, which you can't have, you would be able to have the feedback loop needed to like do whatever, you know, improvement loop you wanted to do to get the agents to be really good.
31:26And again, to make that concrete, like how many steps would be involved in compiling, you know, a tax return? yeah i mean to give you a sense like you could have for example 500 a thousand documents easily you need to think about how to map those documents against each other understand what matters you have to potentially perform lots of different research per different document you have to potentially think about um uh what they all mean um you know you have to compile them all into certain at least today into workbooks of certain types, which are like big excels, um, uh, there's a lot that you have to do.
32:05So you're talking, I mean, steps are definitely in the, in the few thousands easily, if you're thinking about like inference steps. Um, and, uh, you know, depending on how you build your system, uh, if you start to span out, you know, sub agents for different things, which you kind of have to do, you're, you're increasing that potentially. And that's before you think about other, you know, test time compute methods. For example, you know, one thing you could do is say, well, imagine you're, you, there's a tax question you have to solve and it's like insanely difficult, you know, like only an expert would be able to solve this specific tax question.
32:37Maybe there, instead of, you know, the agent doing or spawning a sub-agent, you're spawning five and you let them vote. And so there's just lots of other things you can do to throw more compute at the problem. And so based on what you're doing, like the amount of steps end up being quite large. OK, great. So to play it back, you got very complex processes with many, many, many steps. You don't have a lot of data to figure out what went right or wrong. It's not even always clear what is right or wrong, although you can, at least for certain parts of our problem, say definitely whether this is right or wrong, but not always.
33:11So very complex problem, which leads to how you guys have approached it. And in particular, there's this concept that you can't just rely on outcomes, but you need to rely on process. So what is so wrong about relying on outcomes? Yeah, the problem with relying on outcomes, so if you have a multi-thousand step trajectory or even honestly one that's like an hour long, you will have evals that will say what good looks like. And it could be entirely verifiable, like do the numbers match? And it could be parts that having LM as a judge, like there's some rubric, et cetera. And let's say you have a hundred evals.
33:52Great. You know, they all passed. It looks good. Are you confident that that now generalizes to the real world, to production? And our answer has been no. You actually cannot be confident of that. And so you shouldn't only rely on that. In the same way that if an engineer came to you and said, hey, you know, all my tests pass, end-to-end unit tests, et cetera. Like, does that mean that they architected database properly? Like, not necessarily. That doesn't actually tell you that. There are a lot of ways to pass quote unquote outcomes without, you know, having done the process properly. And so what we have found is that there are, especially going back to my point earlier on analogizing to human organizations, there's lots of learnings from how humans do work.
34:39And so I think it is a mistake to throw out those learnings and say, you know, bitter lesson, throw out those learnings, we're just going to have the agents at runtime develop an entirely new way to do a tax return that is, you know, because bidder lesson, yada, yada, will be better than, you know, the hundreds of years of human history that have gone into learning about the right process. Maybe that will one day be true. I'm not saying that it will never be true. And I do think it is possible going back to like the data bottlenecks we talked about earlier. And I think if you throw enough data and enough compute in a, you know, outcome-based process, you can eventually get there, but not soon.
35:15And so if you're not going to get there soon and you have this latent potential in the models, what you really need to do is think about, let's take the learnings from how humans do things, from what good process looks like. You can't copy that exactly because there is a lot of thinking you have to do about good agent design. It's not like the models out of the box today are incredible at being coherent on long horizons. There's actually a lot of work there about behaviors for sure. And you can instead put in place certain, you know, evals or potentially in the future kind of reward functions that look at whether it followed the process properly or not.
35:53And maybe the example I mentioned, the thread, which I'll just say for the audience is imagine you're doing something as basic as tax research. If you ask some tax question, the agent could definitely get it right. They could know it from their pre-training knowledge, they could read some blog and get it correct. But a real accountant would not trust that. They would want you to cite the primary source. So even if you got it right 100 out of 100 times, if a person is just getting it right because they're going to Wikipedia, the accounting firm wouldn't hire them. And so they shouldn't hire us either.
36:22And so we think it's really important that, no, actually, our agents are not learning from the pre-training knowledge or reading from a blog. They are going to the actual code and verifying the information with the primary source, which is what you would ideally want a real tax accountant to do as well. So you guys created that concept of behavior specs. So walk us through what that is. Very practically, is that a Markden file? What does it look like? Yeah, so the original idea, actually my co-founder Matt came up with the idea. Literally about two years ago, we were talking about Auto AGI. Back when we were, even back when we had agents that weren't fully, I guess, agentic as you think of them today, and they had restricted choices.
37:03even back then you still want to think about okay what kind of choice you wanted to make at this fork in the road and so Matt we actually used to call it internally meta behaviors because the idea was that it was a you're defining a behavior but it's at a meta level because it's all the across all the behaviors agent will have you know in all the different trajectories and so the idea was that instead of trying to write the prompt you have to first align on what the meta behavior is and so that was actually the first purpose of this concept I swear to god literally two years ago and over time that kind of evolved and we ended up calling it behaviors just because it's a bit simpler and the idea is that you have a markdown file in which you write down how do you want an agent to behave simple as that it could be at varying degrees of granularity so let's say you have something that's like very specific like you need to go to like look at the primary sources maybe you want to be more specific maybe you're like no you should always go look at specifically the IRS website Or another example could be, imagine you're making PowerPoints and a behavior is, well, before you return the PowerPoint, you should render it an image so you know if you made any formatting issues.
38:12And so you put that in a Markdown file. And ideally, that is a Markdown file that is now, that can be self-contained. So one that humans can look at in a line, like, yes, this is the behaviors that we want. In part because behaviors are actually subjective exercises. We can talk more about that. It's just as much of a product thing as it is an intelligence thing. Um, and, and then it's something that a judge can look at where they can look, the judge can look at a trajectory and say, Hey, did the agent exhibit this behavior or did the condition occur that it would need to exhibit this behavior?
38:43And if so, did it actually exhibit that behavior? And then you can grade it accordingly. And so who writes the Markdown files or, uh, supervises the process of writing humans. Okay. And those humans are accountants? That's a good question. Um, it depends a bit. I think that, I think it is a combined effort between accountants and ML researchers, you know, at the applied level, because you're not just saying, for example, hey, the behavior is you should go to the website. You might be saying, hey, the behavior is that you should be like, you know, for this type of research, you should be, you know, spawning a sub-agent with like full history because you need to build up that context.
39:25Like there's a lot of agent machinery that comes into play. And so you kind of have this thing where there is, what is the good process looks like for a human, but then you need to like translate it into Asian language and then decide. And there's like a combination there where you're also talking about like, what are the like agent mechanics? And so I'd say it's a dual effort between accountants and ML researchers. And we have a whole team, actually, it's called accounting product operations, where it's accountants who is essentially what they do is they work very close to the research teams to build out rubrics, both outcome-based rubrics and then also behaviors.
39:58And how do you think about precision versus making sure that the system doesn't break? So you mentioned go check the IRS website. Is there a possibility that at some point actually what you should do, you know, your one is go to the IRS website, but then in two years from now, that will be a different location for the information? No, that's a very, very good point, which is exactly why the behaviors are not actually shown to the agent. So the behavior could say, hey, you should go to the IRS website. That doesn't mean the agent is told to go to the IRS website, right? It might be, it might not be.
40:32It sort of depends. But the point is that the way I think about it is the level of, as the kind of agent engineer, the system engineer, you are making the decision as to how specific do you want to be with the situation. Obviously, you prefer to be less specific, right? I think as people who are good at building agents know, it's much better at the margins to be able to give principles and the whys and more context and let them figure it out. And so with behaviors, you want to actually not define every possible thing that can happen, but just say, no, no, no. We know that, for example, let's say you're making PowerPoints.
41:06Taking a picture of the PowerPoint before you give it to the user is going to catch issues. We know that for a fact, right? And so we as the agent engineers are going to take a stand and say, out of all the different things you as the agent are going to do, this is the thing I'm going to grade you on. And then maybe you're being a little bit more specific there. If you want, you can be more specific, like use this exact tool. But ideally, you don't have to. It depends on kind of how your system is built. So I think your level of specificity depends on the maybe what specific outcome you're sort of trying to drive and how that much that outcome generalizes to the universal situations.
41:42Like if you're producing PowerPoints, of course, taking a picture like makes it better. But let's say instead your agent actually is a super fast agent. Well, taking a picture takes time. So maybe a super fast agent you don't want for it to take a picture of the PowerPoint because now you went from it taking 30 seconds to taking one minute. But if you have an ASIC agent that's taking 20 minutes, you went from 20 minutes to 20 minutes, 30 seconds, do it every time. And so this goes back to my point on the product aspect, where it's not just an intelligence thing. It is a subjective exercise about how do you want the agents to behave in production broadly.
42:14And I think it's why it's so critical to product building. So we were talking a few minutes ago about that 2023 effort by OpenAI that required 800 ,000 human labels. How is what you're doing in 2026 different? So, yeah, good question. So we are, today at least, not actively rewarding the underlying model. So we're not currently post-training our own models by rewarding them on this process. I do think that's a very interesting area of research. We can talk about that later. But that is not actively what we're doing right now. It is something we're actually researching separately, but that's more in the future.
42:55and so if you kind of maybe take a step back it's useful to sort of analogize the work of agent building and context engineering to the work of training a model directly I think when people think about context, people say prompts, context, I think the mental model people have usually is like oh I wrote some English model to do, I think it's the right the wrong mental framework I think the framework I like is thinking about it as training data except you're just training the model at runtime. It is training data. And because the model's learning at inference time, the total amount of training data is far lower, right?
43:32Like the whole amount of context in your system that the agent would progressively learn or discover throughout its trajectory, obviously orders a magnitude lower than the amount of data that you're training, like post-training a model on. And so what you're doing is you're essentially taking this data and you're trying to ensure that it is of the highest quality to get the agent to exhibit the behaviors that you want it to exhibit. And obviously the data is only one part of it. There's the harness, what are the tools, what are the capabilities, the environment, there's all that part of stuff.
44:03Obviously which model you're using, things like that. And so what's different here is that we're taking the signal and using it to improve the entire agent system, which requires far less data scale than if we were trying to take the signal and more literally reward it in an RL capacity to the underlying model. You mentioned a judge a few minutes ago. Maybe walk us through how you train that judge to do what or how you instruct that judge to do what. And I guess the obvious question of like who judges the judge. Yeah. Yeah, I think that's true, not just for like behavior evals, but also for, you know, outcome-based evals in general.
44:43It's a great question. I wish we had more time to go deeper on this. The reality is we just don't have the time or resources to like spend a huge amount of effort, like, you know, perfecting every judge. But I'll give you kind of a high level. generally what you do is you need to build an intuition for if the judge's taste is correct and you I do think an interesting area of potential post-training research is on judges and you know potentially taste there today you need to set up the judge so that it has the information it needs to make the decision. And it has the right kind of framework to do that.
45:28And that it has the data and mentality to do that. So I'll give you an example with behaviors that you can imagine being pretty complicated. So right now, at least, the behavior judging is relatively expensive because it's a pretty advanced judge in that it is also an agent. It's not a judge in the traditional sense. It's literally an agent because it has to look at the trajectory. So it's quite expensive. You can imagine in the future, and I think folks like BrainTrust and others are starting to think about this is like, how can you label trajectories better so you can potentially more easily filter the trajectory to only the potentially relevant parts to give to a judge instead of having it like look at the whole trajectory in some ways.
46:07But if you have a trajectory, especially if you have a long horizon one that might have a lot of sub agents, you as the judge need to think about, well, where do I go in the trajectory? How do I understand it? You probably want to have a map of it in some form. Who am I even judging? Imagine you have like an agent system with like depth of, you know, seven. You could have literally seven layers of sub agents. Like, am I judging whether the root behaved properly? Am I judging like one of these other agents? And so you need to actually properly prompt and potentially tune the judge so it has a good understanding of knowing where to go judge, understanding the behavior itself, like what is the condition and then understanding how to like, you know, judge whether the behavior occurred.
46:45Are there any trade-offs with that approach? So what comes to mind is, yes, having a human validated process guarantees or at least help secure a rigorous approach that's less likely to fail. At the same time, you are not going to get a move 37 kind of result where actually the AI would do a much better work, much more efficient work by sidestepping this part and fast forwarding through those three steps. What are the sort of pros and cons and trade-offs? It's a good question. Well, I think for starters, and I get one thing that should be, I want to make sure it's clear, I think, to everyone listening, is that you, writing a behavior down is expensive because it is something where you are now keeping state, right?
47:37You need to keep it up to date. If the way you are making the model or the agent exhibit this behavior is through context, now you have potentially duplicate state. depending how you're thinking about it.
47:52And because of that maintenance burden, you ideally want to have as little behaviors as possible. So it's not that you look at what it takes to do tax return and you say, hey, let's write all the best practices down and see if it's doing it. It's that you take the couple that you think are most important, that have the largest amount of generalization to production and are the most impactful and you care about those, not everything. So I think that that's one important part. So ideally, if you do it properly, you still have rooms for, you know, the move 37s in theory. Um, but I think there's another part here, which is that, what are you selling to someone?
48:24If you are trying to, just like if you were to hire, you know, uh, an engineer, um, and you know, you know, at your job, you, uh, at the company, you have a process and the process is like, um, you, uh, I'm just making this up, but like you write a quick architecture diagram and you like chat with the CTO and you get it approved and you make a PR and you split up the PR in like 10 different smaller PRs into a stack and then you merge it in and you make sure you have your N10 test and you deploy it, like that's the process. Imagine a engineer came to you and they're like, hey, here's my thousand line PR.
48:58I'm gonna merge into production right now. What if it's better than what you, what if it's a move 37? It could be better than what the CTO would have came up with, but that doesn't mean it's good. Like good in the sense that doesn't mean the CTO or the company is happy about that result just because it's better, right? Because at the end of the day, the reason you perform work is not because any individual unit of work is incredible, but because you can scale it to a company, to a system, whatever it is, to an organization. And so the thing that, you know, someone is buying from us is not this will be the best ever tax return.
49:30They're buying that, you know, the confidence that it's looking one. Yeah, the best looking one or like, you know, it moved 37, the TB's over here or something. Right. They're buying that it's going to be consistent and reliable and something that they can trust that actually will. So they will, you know, just like with a human, they can learn to trust more and more and then granted more agency over time. Right. And and and I think that that level of trust and reliability, that's what you need to deploy to the real world. You don't need the move 37s. You need that maybe at the, you know, Olympiad, like math competitions, but not like doing work in the real economy.
50:02Yeah. As I listen to you, one of the many things I find fascinating is that you're doing all of this without having actual insight about how the underlying model works. Like all of us, you're sort of guessing and inferring from how the model behaves through artifacts and judging from tool calls. and how does that work? And then maybe walk us through each time a model changes or the next version of the model, you know, gets released. Like, do you have to then look at everything that you've been doing in the light of that new model? Essentially maybe it goes back to the like, you know, Opus 3, 0.1, 0.03, because I do think one thing that's really important when you're building agents, but it's definitely building a company around it is, you shouldn't be that suppressed.
50:58You should have a model of the world. And as things change, you should update your model. But to be successful, you know, you can't just update your model all the time. You need to be right a little bit. And I do think that if you really internalize some concepts about this, right? That like you now have this, forgetting about the internals for a second, forget about this LM, you have this magic box. And you have this magic box or this alien, I like to call it sometimes, and you could send in huge amounts of data into this alien and it will be able to reason and learn at inference time within that magic box.
51:37And then come back to you with output that, you know, now that tool calling obviously works and whatnot. You could plug into the rest of the system. That's kind of all you really need to know. And I think once you really appreciate what that means and you take it to its logical conclusion, a lot of stuff starts to fall out of that. because you start to understand, it's like, okay, we'll wait. If I have this magic box that can do this, does that mean that it could decide to call another magic box? Does that mean that it could potentially string together like multiple of them in a row, right? Does that mean it could leverage?
52:11Obviously at that moment, it's a magic box, but it has some state. We know this, it has an activation state. That's how the caches work. So there's some activation state that by definition is going to be biased to that current trajectory. And so maybe for review, you want an uncorrelated trajectory, right? Where it's like a new box and it's just a smart. And I think if you build these kind of LLM intuitions and you combine them with, you know, maybe basic principles of like organizational design and management, I think you start to get to maybe what is like the frontier of agent building. Fascinating.
52:42Practically, how do you build that LLM intuition? Is that by just reading papers all the time or talking to a researcher? or like getting a sense for where the state of the art is going? Yeah, I think it's none of that, actually. I think it is all about, I mean, I think reading like Twitter and whatnot, just my understanding is good. But I actually think a lot of people over-index on that. Like I think a lot of people, they think like, oh yeah, I saw this tweet. I saw this, like it's the next cool thing. I think the problem is without like a fundamental like grounding in how things work and what is possible, it's easy to feel like things are moving around a lot when they're actually not.
53:21Things have really not changed since 03, I would say. Almost everything since 03 has been relatively on, I don't want to say on trend and then like I knew this exact trend, but I would say it's all within the same paradigm. Like nothing paradigm shifting has changed since 03. And I think the best way to understand it and learn about it is to just use them in your own work a lot. I think especially in coding and just trying to understand things. Like a good mental model is, let's say I tried to, you know, have an agent, you know, influence a feature for me and it didn't do the way I want it to.
53:57Like why? What is actually the limiting factor? You know, it's kind of like the famous like Elon mindset. It's like, okay, like you go to like the main limiting factor and you like figure that out. Like I think if you apply a similar mindset to agents and you understand like, why could it not automate this? Was it actually not smart enough? Probably no. They're pretty smart. They've been pretty smart for a while. And so if you apply that mindset to your own work, I find that is quite useful for building intuition. And I see that actually with people I interview. A lot of the people with the best agent intuition actually, yes, a lot of people come from ML backgrounds, but people who don't, a lot of them are ones who are just really good at automating their own work, really good at thinking about it, really good at understanding like, what is the system to build?
54:38I did a talk at Data Driven about ontologies, whatever, a year and a half ago or something. And I think, you know, there are a lot of people who think about ontologies in their own data, in their own repo. And like those people who are like actively thinking, not just how do I prompt a model, but how do I build a system? They start to build really good intuitions. So we make sure to cover it before the end of the conversation. What is it exactly that you're open sourcing with Brent Trust? Like walk us through the project, where people find it, the genesis of it. Why are you partnering with Brent Trust specifically on this?
55:11if we go back to the idea of behaviors, right? The idea is that you can actually write down in Markdown what is a, it is actually both a spec and a rubric. We call the specs and there were some people who asked, isn't this a rubric? And it is, it's both. And the reason it's both is because it is not just used to grade or potentially reward the agent. It's also used to align the humans. I think that is an underrated point in that how you want the agent to behave, as we talked about earlier, is actually a subjective question. and so internally it's it's you need to build processes to all agree on hey like this is the product right how do you want the agent to behave going back to the example about you know the fast powerpoint verification and so you want a standard to kind of write that down and the the project kind of came about because i was actually having coffee with the ceo of brain trust anchor and i forget why honestly but i was like i still tell you i was like i was explaining this concept to him i was talking about this because we were doing this internally and i thought it was very cool and he got pretty excited about it and um you know one thing that i had uh internally that we're trying to think about is we i think do a lot of really cutting-edge work but it's not something we talk about much because to be honest we're working all the time yeah right before we started recording you were showing uh some internal slacks between your co-founder matt and yourself and uh If I may disclose them, like Matt was sending you a Slack at 4 a.m.
56:39Those prompt refactors. Yeah, and that was last night. So it was a Sunday night as we were recording this, and you showed how you were replying to that Slack at 6 a.m. So, yes. 996 in full action amongst the co-founders of Basis. Yes, yes, yes. There's no 996. It's for Matt and I, it's 24-7. For the rest of the company, it's people work hard, but it's definitely not a 996. And so, you know, we wanted to talk about it more and just share what we're doing. And we, you know, we don't have a lot of resources to like blast out to people. And so we were talking, I was like, well, I just think this could be really good for, you know, useful for brain trust.
57:19And honestly, the whole industry, because if you have a standard, that could be something that people define. and it can get automatically slurped up into observability platforms, monitoring platforms. And for people maybe who are less advanced, it could also have out-of-the-box judges or ways to define, hey, here are the behaviors and you don't have to configure your own judge. You can actually get it to judge it for you and see the results. And so he got pretty excited about that. And so that's where the collaboration came from. So that's what the open source repo has. It has a couple small examples.
57:49It has an example judge that you can use. It has examples of actually written behaviors that you can leverage and sort of build your own. And I think it is useful to think about how to adopt the standard. But I think it is also maybe more importantly thinking about how to adopt the mindset of, you know, not thinking that an agent operating over 10 hours is a black box. It's not. It has a lot, a lot, a lot of data. And you're probably doing a disservice to your customers if you don't understand like how it's going about the work. And what would you want people to do with this open source project?
58:22Presumably contribute to it, use it for their own purposes? Like how does this become an industry standard? Yeah, it's a great question. I don't actually know. I think the coolest thing would be for people to contribute ideas to it. I think there's a lot of work left to do. I think it's just the beginning. I mentioned a couple of things earlier, but there's so much to do around, one, how you make good judges. Two, how do you properly label and dissect trajectories to make it easier for judges to understand? Because at the current level of expense, you couldn't run this in all of production, for example, because it's, it's, you're running judges on every single trajectory.
59:00But there's a lot that can be done, I think, with building out the, the work that sits on top of the behaviors. And I think also just seeing, you know, we purposely tried to make the standard relatively flexible, similar to skills, where it is, it is just a markdown. Like there's not an overfit, hey, and you need to have these exact five words. Like you can make it very broad and you can make it, you know, very specific as long as it is still self-contained to the point that a judge could look at the behavior and actually know, like, was the condition for it to be exhibited met? And if so, like, did it get exhibited or did not get exhibited?
59:37And as long as it has that, like, there's a lot of leeway there. And so I think that's, we wanted that to, you know, be flexible. So what else should AI builders think about as they build autonomous, long horizon agents So we talked about judges, we talked about behavior. You just mentioned ontology, which in your talk at Data Driven NOC, you had mentioned as well as a world for agents to live in. Where does that fit in the picture? Yeah, they're super important. A lot of people when they see the word, when they think about agents, their mental model always goes to coding agents. Because that's everyone's experience, at least the people who probably listen to this podcast.
1:00:15People think about coding agents a lot. Coding agents are interesting because they obviously have a harness that they get shipped in. They have certain tools. They have certain behaviors encoded in them in the context, right? Codex, by the way, is open source. I highly recommend people go look at the open source repo. But they don't control their runtime training data, right? Because their runtime training data, which goes back to my analogy earlier, which is your context, is actually the repo they're working on. And so you could have codecs, the same agent, quote unquote, on one code base perform somewhat well and then another code base performs spectacularly because that code base has much better runtime training data, right?
1:00:58i.e. it has potentially good skills or good context about how to operate in the code base because there isn't contradictions and confusions, whatever it is. And so the ontology of your code base, it always mattered for engineers. It matters just as much, if not more, for really good agents over time. That actually, I think, goes up another level if you're thinking about non-coding agents. Because in coding, you don't own the runtime training data. In non-coding, you do, right? Like most of the data, quote unquote, that an agent sees when it's a basis agent, basis owns, right? It's training data that we have to ensure works really well.
1:01:40And again, when I say training data, I mean like effectively handwritten context, right? Or things that are part of your broader progressive disclosure. Some, you know, could be handwritten, some could not be, whatever it is. And because the agent is always starting from scratch, designing that ontology in a way that is ergonomic for the agent is super key to building something that's long running that's both like the data that is like maybe static like skills and whatnot that all the agents have but also once you get the really long horizons if you're talking about you know a stateful agent that's maybe operating over months days to months now suddenly you have an ontology of going back to the memento example of information the agents left for itself which if you're operating for maybe a couple hours could be a couple notes if you're operating for months you're talking folders, right?
1:02:25And you have so much knowledge and context that describes like the lived experiences of the agent that like suddenly this new agent, well, new, that is standing up with, you know, effectively very compacted context has to get it back into the state of mind of its entire lived experience. And so the ontology designed to make it easy for it to do that and the behaviors you encode that, you know, properly ensures it's doing that well are sort of the key to making it work really, really well. What does that even mean, designing an ontology? An ontology practically is a graph database. It's a series of relationships.
1:03:01Yeah, it could be a lot of different. There are different formats. I think the simplest way to think about it is honestly just a file system in which you have some structure. Obviously, most people, you do virtual file systems, and so you have a lot of flexibility there. So there could be other types of metadata associated with the files and the folders, right? There could be connectors in nodes and some graph TV if you wanted. Obviously, the girls would be embeddings. Like, there's so many different, you know, sources of data that you can get to, like, help enrich. And now, by the way, as models are getting cheaper and cheaper, more and more of that actually can just be done using inference instead of using determined, like, things like graphs or things like embeddings.
1:03:41You know, if Luna costs is free, then suddenly you could run Lunas across your entire ontology and summarize stuff for the agent up or things like that. And so that's kind of how I would think about it. And maybe also one more piece, it's not just the, the ontology doesn't just mean the structure of the folders and ontology traditionally, it's the, it's also like the language, right? Like what are the objects and the concepts? Because at the end of the day, if you're being trained at runtime, you need to ensure you're not confusing concepts together and that like things kind of generally make sense.
1:04:09And so you're, that's what I mean by defining the world, right? You're, you're defining what the agent can expect to, to, to, to see as it goes and explores the kind of world around it. You mentioned somewhere that internal documentation for agents has to be treated like a code base. Yeah. Delete a crucial paragraph and you break the agent just like deleting a line of code. Is that documentation something outside the ontology that the agent goes search like a tool call? How does that all work and what are the best practices? Yeah, that's a good question. So just to quickly separate, just now when I was talking about the ontologies and like historical, I kind of was referring more to, you know, inside of like the basis product and like the kind of like production product in terms of like internal use of agents, like let's say coding agents or maybe other internal agents that we might make.
1:05:03There it's kind of interesting because you're in, depending on what you've set up, your environment might be less controlled by your ontology because to your point, you know, in the real world, you have to go and access like linear and gong and pylon and all these different things. And so I think one of the keys for internal agents is having a very keen understanding of what is canonical versus what is not canonical. And so just like if I'm a human who joins an organization, I could go and read all the gongs, but what is our sales strategy today? If you watch the gongs from two years, you'll get a lot of context, but you won't know what your current sales strategy is.
1:05:46There must be some canonical piece of documentation. In practice, a lot of times humans learn this by just like talking to people and you kind of learn stuff. But with agents, it's hard to get that. And more importantly, if you want real organizational intelligence, you don't want an agent hearing one thing from one person, another thing from another person, or having them have kind of different written records of like, what is the current, you know, sales pitch? Or how do we make our decks? Or how do we make our emails? You need one canonical source. And that's why I think, you know, for true agent native companies, especially in the future today i think it's still quite early but especially in the future having a clear understanding of what your company canon is and organizing that in an ontology that makes sense and ensuring that that is like kept up to date just like code in some form i think ends up becoming one of the most important parts of a humans inside of a company and within the basis teams you're hiring for jobs that quite literally did not exist two years ago like language architects or agent managers?
1:06:42Who are those people? What do they do? And what's a good background for them? Yeah, we need a lot of them. So if you're listening and you want to join, please, please hit me up. Great question. We're still figuring that out. It's not easy. Here's what I know. You need, I think this has probably gotten somewhat in vogue recently, but I think one One skill that really matters is good systems thinking. And where does good systems thinking come from? It comes from people who have had to think about some abstraction, some system, something, and design it in such a way that it performs in a plethora of situations.
1:07:31so obviously if you're a really good engineer like engineering is systems thinking now I think the majority of engineering historically has not been really systems thinking based you know it's been a little bit more execution oriented but if you think about like the hardest engineering like hey I'm trying to design like what systems are going to look like or I'm trying to like you know create the right abstraction that really is high on systems thinking but it's not the only profession that's like that I think law actually is kind of like that you know in many ways I think about I think maybe the founding fathers would have been really good context engineers or agent managers, because you had to, you had to write, you know, a piece of English that was going to be, you know, interpreted at runtime, literally millions of times by lawyers and judges and whatnot.
1:08:12And so if you're writing a law and I don't mean like, you know, some politician, but if you're like actually trying to write a law and trying to write a well, you're trying to like somehow write something in English that will abstract at just the right level across the universe of situations. And you have to have like theory of mind over the judicial system to think about like how they'll interpret it. You know, it's funny, you'll look in certain airports and, you know, sometimes they'll have these signs. It's like, you know, don't bring a gun, don't bring a, you know, don't bring a sword, don't bring blah, blah, blah.
1:08:40And it's like, you list out 30 things to your point on brittle rules. And so same thing, you can list out like 50 brittle rules, or you can like write the right abstraction that somehow covers it just perfectly. So I think anything where you need to think in abstractions in that way, I think is good practice for being a good systems engineer. I think, you know, people who've had to manage like the most complex Excel models in the world is honestly not that dissimilar either. So I think there's a lot of potential, potential backgrounds for it. I think I did that driven when you spoke, you were talking about your deployed intelligence team and you were saying deploying agents at a firm was like onboarding 300 brilliant alien employees who have no context.
1:09:19Hence the deployed intelligence team. So what do those people do? Are they still around or is that concept evolved? No, no, no. Of course. The DI team is awesome. I think to this, I need to go Google it, but we definitely came up with the term deployed intelligence. I know because if you Google it, we are the first company that comes up. It's a cool name. I don't know if other people have taken on the name. I don't think it's actually caught on yet, but the idea is it's actually, it's not FDEs. So it's not like engineers who are coming in and building something custom for you. It's also not this kind of like agent PMs you see now at some companies that I won't name where it's like, you know, these like PMs, quote unquote, are kind of coming in and like building an agent sort of for you using an agent builder.
1:10:01It's actually neither of those. Um, it's, uh, it's people who, uh, have a deep empathy and understanding of the profession of like what it takes to put in place good process. Um, and what it takes to be like successful when suddenly you can start to offload certain things to, to agents. Right. And so the DI team's job is to come work with our accounting firms to help them transition into this new era, right? We're giving them magic, but if we don't teach them how to leverage the magic, not just in the day-to-day, but how it changes the nature of the firm, how does it change what kind of business they can take on, who they hire, how can they scale to be that firm of the future that everyone wants to be?
1:10:47That's what the DI team works really closely with people to do because it wouldn't be fair to ask them to go and learn that themselves or do that themselves. Instead, they bring a lot of the knowledge about their firm, about how things have worked, obviously their people. And we can combine that with our knowledge on how to deploy agents and I think together get to something where it can be a really, really frontier accounting firm. Speaking of frontier, maybe as we start getting to the end of this conversation, I would love to spend a little bit of time on that. You know, obviously a big topic in 2026 is the concept of self-improvement.
1:11:23Where does that fit in your picture at Basis and with autonomous agents? I think you talked about agents developing a theory of mind about other agents. So at a system level, paint that picture for us. What does self-improvement look like? Yeah, we actually internally, one of our, in the thread I talked about some of our research directions, one of our big research directions is how do you actually close the loop, as we like to call it. Going from, hey, an agent, you know, made a mistake or not performing or whatever to, you know, We've gone in and improved the system to do that. And I think that closing the loop is going to happen pretty fast.
1:12:06I think you'll have like, I don't know about the entire loop being closed, but I think you'll be relatively close by end of year. I think that, so as agents are getting better theory of mind over themselves and therefore other agents, you have two things happening. One is at runtime, they're being better at orchestrating sub-agents and also regulating their own environment for themselves. But it also means they're becoming better context engineers, right? They're becoming better harness engineers. Right now, they are far, far, far worse at engineering agent systems than they are at engineering most software.
1:12:41Far worse. Because by definition, like that kind of work, which is so novel, has not seen a large amount in their training data. And so they have very bad intuition. Actually, I think a lot of the mistakes a lot of agent builders make is, you know, they have this weird intuition that like slop in your context or your agent is somehow more acceptable than slop in your code. And you'll see people like freaking out over, you know, a code file that isn't abstracted properly. And yet their context is like total shit, which is hilarious because the context actually affects the performance at runtime.
1:13:15The organization, the code does not affect the performance at runtime. time. And so I think a lot of, you know, obviously because of like built up, you know, behaviors, a lot of engineers, they treat the code as more precious than the English, when actually the English is more precious because the English affects the performance. The code does not affect the performance, right? If the logic, assuming the logic is the same, it does not affect the performance. And so I think as agents get better and they get better at this type of engineering of like, how do you build agent systems? You'll start to actually be able to close the loop because in order to close the loop, right, you need to take the signal that you are getting and you need a lot of signal.
1:13:54One of the pieces of behavior, one of the points of behaviors is a way to get more signal. Whereas if you only have outcome-based evals, your signal is pretty sparse. So how can you take all of that signal and actually now use it to improve your system, right? You proliferate it throughout the system and that could be done by an agent that is like updating the context, like changing the nature of the tool, updating the harness, et cetera. And I think you'll probably, in a way that is generalizable, that doesn't overfit. You can do it today if you want it to overfit to like some signal. But if you want it to do it in a way that generalizes, you need, the agents need to get a bit better still.
1:14:28And there's some more work to do there. What about self-improvement at the model level? So I think you mentioned earlier that you guys don't do yet much reinforcement learning on the model itself. you've done, you're doing mostly harness work, if that's correct. But that's the next big bet to start doing actual reinforcement learning on behavior adherence. Yeah, it's a great question. The way I think about it is that the hard part of reinforcement learning is deciding what is your reward function and then deciding how you are going to allocate that reward over whatever occurred, right? So we are, that is like the active research that we are doing.
1:15:11Because that's what behavior is, it's a source of signal, right, that you can use to craft into a reward function. Same with, like, some of the other kind of production monitoring work and some of the, like, evals that you build. And so deciding how to build good signal and doing the research and understanding what does it mean, especially in a non-perfectly verifiable domain, that's the work we're doing. Whether you take that signal and then like proliferate it through the weights, you know, through formal RL or you proliferate it through the harness, through whatever you want to call it, like informal RL or like harness engineering, I think is like a separate question from the development of the signal.
1:15:44But I think the development of the signal is really the hard part. And that's the part that we're really focused on. And today we don't go directly to the weights. And primarily the reason we don't do that is because a lot of the advancements the models are having when it comes to orchestrating themselves yield far more performance gains than benefits you would have of like updating models directly. I think that might asymptote, we'll see, but I think that's one of the reasons that you don't go to the weights yet. I do think that is a very interesting research direction as well that we'll start pursuing.
1:16:18And so for us at our level, at the kind of like applied application level, that's the kind of research that we're doing. But I suspect that in order for the models to actually do real work in the economy, that's the only way you can get there. I don't believe that if you were to train a model, you know, and you scale up the amount of pre-training computing, and you scale up the amount of post-training from perfectly verifiable rewards, that suddenly will output a model that will do a tax return reliably. It might do one that's really good. But the question is not does it do really good. It's that does it do it at the level of quality, reliability, scalability that would be expected from somebody operating with it?
1:17:01Do you worry about the better lesson, though, that you mentioned earlier in this conversation? Do you think that all this work that you guys brilliantly and others are doing at the harness level are going to be eventually swallowed up by the model? Oh, I assume it will be swallowed up. Yeah, I'm not worried about it. That is absolutely the future. And in fact, if you look at my thread, as a hint to this, I think I said the behaviors don't get shown to the agents yet. And the reason is because if you are truly bitter lesson-filled, then in the future, this whole idea of like trying to context engineer, it will just go away and you'll just specify, hey, I want these behaviors.
1:17:42It'll just work. And so I definitely think it'll get swallowed up. No question. I think that how long it will take to get swallowed up, I don't know exactly. You know, I think it's probably sub five years. I don't think it's sub two years. I think it's probably sub five years. And so for us, you know, we're in hyperscale mode. Like we can't wait for the bitter lesson to arrive to like perform tax returns accurately. And so that's why I think that'll be one of the keys, you know, to doing it. I do think that, I don't know, this is now total speculation, but I do suspect that maybe some types of process rewarding for this type of work might end up being pretty important, you know, for optimizing the compute that the labs even use over time.
1:18:34because I don't know if you want to move 37 like every single profession in the economy. Like we have a lot of learnings already and like there's kind of no reason to do that would be my guess. So do you think of doing your own RL as some kind of moat against being swallowed up by model performance? I mean, as you think about applied AI companies of the future, will they all be RL labs of some sort? So generally, and if anyone's currently trying to found a company, I recommend thinking this way. Technical moats are not real moats. Like, there's no portion of basis's long-term terminal value that stems from some, you know, secret RL trick we found that nobody else found.
1:19:18So that doesn't matter. What matters is that right now we are obviously very good at building long horizon agents that can be reliable and deployed in production. And we'll continue to be the best at that. and that allows us to win market share and get deeply embedded. And that's why we move really fast because the work we're trying to do is to go and proliferate maybe before you get to AGI, whatever you want to call that. So I think that most of the moats that will exist will be business moats. That is true in the AGI era. I would argue that's also been true in the pre-AGI era. Like I don't think that, you know, Salesforce can write a better SQL query than I can.
1:20:02The mode that Salesforce has is not related to their technology. It's related to their business position. It's the powers. It's the workflows that they own. It's so many of these different things that come with being embedded in. And that's what matters, not the technology. The technology is a temporary dislodgement that allows someone like us, who obviously didn't exist three and a half years ago, to now suddenly be able to do all this kind of work. I do think, though, for a long time, maybe to your point, that it's not like it will be a total commodity because just like you could have a bunch of genius humans doesn't mean that all the genius humans are equivalent in being able to do a tax return.
1:20:45Because at the end of the day, there are certain things that you care about for that that are subjective. And I do think that building up the competency of the work we're trying to do is actually really important for us to be able to deliver really good quality for, I think, a foreseeable future. All right, Mitch, it's been absolutely brilliant. To close, any advice for AI builders, anybody like building agents today, in addition to everything that you've talked about, like the do's and don'ts and lessons learned and anything that comes to mind? I think maybe the biggest lesson I would say is it's easy to look at the world and how fast things are changing and say like, oh, things are just, you know, one thing here, one thing here is like, you know, ADHD on Twitter.
1:21:36It's crazy. You know, the Chinese labs are at least something every day and feel powerless to, uh, make first principle decisions. Whereas I actually think if you treat the new world as like an underlying paradigm shift, you know, in the same way that like a shift to cloud or something is, you know, some underlying paradigm shift and you try to extrapolate, okay, you have these things. What does it mean? You know, if, you know, the intelligence became X better or if it didn't, I think you'll build a lot more coherent systems and make a lot better strategic bets, both at like a technical and at a business level, things are changing, but it's not like the paradigm is changing, at least not that dramatically.
1:22:20And so I think really understanding it is very, very key for being able to build in this world. Which was absolutely fantastic. Thank you so much for sharing all of this. We really appreciate it. Of course. Thank you so much for having me. Hi, it's Matt Turk again. Thanks for listening to this episode of the Matt Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing if you haven't or leaving a positive review or comment on whichever platform you're watching this or listening to this episode from. This really helps us build a podcast and get great guests. Thanks and see you at the next episode.
From the publisher
AI agents can write code for hours, but ask them to do real work in the real economy, and they break. Mitch Troyanovsky is co-founder of Basis, a unicorn AI company whose agents run autonomously for hours — sometimes days — completing complex tax returns end to end. His answer to the reliability problem: stop grading outcomes, and start supervising the process.
This is a definitive, reference-style conversation on building long-horizon AI agents. Mitch walks through the full history — from ReAct and the AutoGPT crash to reasoning models and RLVR — and explains why the industry abandoned process supervision in 2023, and why it's now coming back at a completely different scale. We go deep on behavior specs, the open standard Basis just released with Braintrust for defining and evaluating how agents behave across entire trajectories, with no ground truth required.
Along the way: why context is really runtime training data, why your documentation must be treated like a codebase, ontologies as "worlds for agents to live in," the judge-as-agent architecture, why Basis hires philosophy majors as Language Architects, deploying agents as "onboarding 300 brilliant alien employees," and Mitch's prediction for when the bitter lesson swallows the harness.
(01:09) Why Everyone at Basis Whispers to AI
(04:12) Accounting as "an Intelligence Over the Economy"
(06:11) What Makes an Agent Truly Long-Horizon
(08:24) Inside an Autonomous, Multi-Day Tax Return
(10:19) Agents That Hand Off Like Senior Engineers
(11:17) A Brief History of Agents: From ReAct to Today
(12:33) Why LLMs Have No Long-Term Memory
(14:13) Why AutoGPT Failed
(15:51) The Three Breakthroughs: Opus 3, o1, o3
(17:07) Why Reasoning Models Unlocked Agents
(18:23) "Let's Verify Step by Step": The Road Not Taken
(20:32) Pushing Back on the METR Chart
(22:09) Why Coding Agents Won First
(25:14) Why Real-World Agents Are Harder
(26:55) How Accountants Verify Non-Deterministic Work
(29:18) You Can't Scale Tax Returns Like Math
(33:16) 100 Evals Pass — So What?
(35:53) Right Answer, Wrong Process
(36:37) Behavior Specs, Explained
(39:58) How Specific Should Behaviors Be?
(42:18) Context Is Runtime Training Data
(44:21) Who Judges the Judge?
(46:45) The Move 37 Objection
(50:02) The Magic Box Mental Model
(52:41) "Nothing Has Changed Since o3"
(54:56) Open-Sourcing Behavior Specs with Braintrust
(59:45) Ontologies: A World for Agents to Live In
(01:04:20) Documentation as Codebase
(01:06:33) Why the Founding Fathers Were Context Engineers
(01:09:05) Onboarding 300 Brilliant Alien Employees
(01:11:10) Self-Improving Agent Systems
(01:12:50) The Context Mistake Agent Builders Make
(01:14:29) RL on Behavior Adherence
(01:17:01) Will the Bitter Lesson Swallow the Harness?
(01:18:46) "Technical Moats Are Not Real Moats"
(01:21:03) Advice for AI Builders
