In short
Unify’s AI agent platform for go-to-market (finding leads and drafting outreach) and how it cut AI agent costs 90–95% in two weeks by fixing prompt-caching and agent execution patterns.
Guest backgrounds
Connor Heggie, co-founder and CTO at Unify. Unify builds outbound “growth” agents powering $900M in pipeline; he discusses agent harness architecture, cost engineering, and sales-rep UX.
Key claims
Sub-agent “main agent” design plus prompt caching is essential because model providers won’t optimize for your traffic distribution. Prompt caching is best-effort and depends on throughput (e.g., ~15 requests/sec). LLM-as-judge can cause distribution mismatch and “mode collapse/groupthink.” Company knowledge should live in structured “memory” for sales reps, not just engineer-authored skill files.
Notable examples
KYC research via web agents (checking terms of service/HTML for KYC signals). Email generation with human approval and batch editing (review first ~15–20). Memory ingestion from Gmail to learn “email voice,” including “talk to me like a pirate.” Prompt-cache warmup via hashing user IDs across 16 buckets.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOCost Optimization in AI Agents
0:18 to 1:12
Discussion on achieving 95% cost optimization in AI agents and the challenges faced.
“We were running so many sub-agents, and that's way too expensive.”
Evolution of Agentic Systems
1:12 to 1:41
Exploration of the evolution of Unify's agentic systems over three years.
“You guys have been building agentic systems for almost three years at this point.”
Go-to-Market Strategy with AI
1:41 to 3:13
How Unify approaches go-to-market strategies using AI to identify potential leads.
“um uh so founded this company unify what do we do uh we essentially do ai for go-to-market right is kind of like what we say.”
Scaling AI Research Agents
3:13 to 6:15
Understanding how Unify scales its research agents for efficiency and effectiveness.
“Wow, that's so embarrassing if we go back and look at it now.”
Chat Product Launch and Impact
6:15 to 8:14
Discussion on the launch of Unify's chat product and its implications for sales reps.
“And just to maybe talk about the scalability briefly, and this gets into like a little bit of like the UX of these agents.”
Agent Harness Design Considerations
8:14 to 9:38
Insights into the design and engineering of the agent harness for Unify's platform.
“do best, which is like talking to people, right?”
Implementing and Optimizing Data Functions
9:38 to 14:00
Overview of how Unify optimizes data functions for better agent performance.
“engineer in their back pocket to do all this work for them in a scalable way?”
Challenges in Data Processing
14:00 to 14:48
Explore the difficulties in running functions over large datasets with security considerations.
“So if you want to run a snippet of code over every row in the table, that's actually really pretty hard to do.”
Implementing Strong Tenancy
14:48 to 17:12
Discover the importance of strong tenancy in data access and manipulation.
“was really important to us, which is like absolute P0 is strong tenancy.”
Common Agent Functions
17:12 to 18:25
Learn about the typical functions an AI agent performs on data tables.
“So you have a thousand rows, even if each one is one cent, that's 10 bucks.”
Show all 29 chapters
Prompt Caching Explained
18:25 to 20:08
Understand how prompt caching works and its significance for performance.
“And when you're pulling people or companies, you're pulling them into this table.”
Impact of Cache Hits on Costs
20:08 to 22:09
Analyze how prompt cache hits affect operational costs and efficiency.
“they make no guarantees on any of this, by the way.”
Sub-Agent Architecture Considerations
22:09 to 24:16
Examine the architecture around using sub-agents for enhanced processing.
“And so you want to keep basically, there's like this game between keeping enough of them warm versus not.”
Forking Sub-Agents for Research
24:16 to 28:00
Discuss how forking sub-agents can streamline research tasks in AI applications.
“But now it's really important because you have this game of, let's take the sub-agent case, for example, which is kind of the most illustrative of this prompt caching issue, this 15 requests per second issue.”
Exploring Fork Subagents for Email Automation
28:00 to 29:20
Learn how fork subagents can streamline the process of sending personalized emails efficiently.
“So let's say I'm researching, you know, you, I'm researching Harrison and, or let's say I'm researching five founders, right?”
Human in the Loop: Balancing AI and User Control
29:20 to 30:50
Discover how incorporating human oversight can enhance AI-generated email processes.
“Why can't you use a regular subagent and prompt it with all the knowledge previously?”
Optimizing Email Approvals and User Experience
30:50 to 33:00
Understand the strategies for improving user interaction with AI-generated content.
“Like I imagine I would want some form of human in the loop if I had an agent sending emails for me.”
Personalization in AI-Generated Emails
33:00 to 35:30
Learn how sales reps can personalize their email communication using AI insights.
“Um, one, you had to be able to see the full email and you had to be in context, right?”
Memory Management in AI Systems
35:30 to 37:50
Explore the mechanisms behind effective memory management in AI applications.
“Because I imagine every seller probably has a different tone or different things they like to mention.”
Structuring Data for Effective AI Communication
37:50 to 42:00
Discover the importance of structured data in enhancing AI communication capabilities.
“Like what is this data structure for you guys?”
Memory Management in AI Agents
42:00 to 43:15
Learn how AI agents handle memory and user preferences.
“And I guess I could see why it merged those two memories.”
User Engagement and AI in Sales
43:16 to 45:18
Explore user engagement strategies and AI's role in sales.
“Like, I'll ask you what you remember about me.”
Cost Management in AI Usage
45:19 to 47:20
Understand cost management strategies for AI tools in businesses.
“Let's make sure everyone uses that skill.”
Evaluating AI Performance
47:21 to 51:22
Discover methods for evaluating AI performance and metrics used.
“month per user and you get some allocation and then that kind of helps like guardrail you directly.”
Testing and Sandbox Environments
51:23 to 56:01
Learn about testing strategies and sandbox environments for AI applications.
“And we'll iterate on one of those at a time, usually.”
Exploring Open Source Frameworks
56:01 to 58:27
Learn about the Monty open source framework and its application for AI agents.
“Do you guys have just the, do you guys have the VMs and repels or just VMs?”
Optimization Strategies for AI Costs
58:28 to 1:01:28
Discover how to optimize AI agent costs and improve efficiency.
“It's the same thing with like the ability to like host, like inject in host functions.”
The Role of Model Selection in Cost Reduction
1:01:29 to 1:05:04
Explore the impact of model selection on operational costs in AI.
“Is within a skill file, is it consistent?”
Evaluating Open Source Models
1:05:05 to 1:07:18
Examine the challenges and considerations in using open source AI models.
“So I have to say that they like are great partners.”
Transcript
Automatic transcript. May contain errors.0:00Before, you might have a team of 10 sales reps. And instead, with Unify, with Unify TM, you would have one growth marketing person who would spin off a million agents a month.
0:09Connor Heggie:Today, I'm talking to Connor Heggie, co-founder and CTO at Unify, the outbound agent platform powering$900 million in pipeline. We actually got a 90 or 95 % cost optimization from two weeks before we launched to the day that we launched. We were running so many sub-agents, and that's way too expensive. So how do you move to doing this main agent that's smarter? He gets into the prompt caching cost issue and why the model providers won't solve it for you. We're at like a 95 % cash hit rate or something. It's really important that you get these cash hits or otherwise you're cooked. It can't be built into the providers because they don't know the distribution that's coming in.
0:45But for them, they don't care, right? They charge you either way.
0:47Connor Heggie:He explains why the model you pick to judge matters more than you'd think. When you have an LLM as a judge, it can never be the same distribution as the original model. You have this mode collapse of agents talking to each other, the equivalent of groupthink and humans. You want an almost adversarial. And he asks where a company's own knowledge should actually live. Engineers love writing things like skill files, but most other people don't. Why not just have it be memory about your company? Welcome to Max Agency, the podcast that goes deep into how the best agents are being built by builders like you.
1:20Connor Heggie:You guys have been building agentic systems for almost three years at this point. you recently launched a new version of your platform and agent how has your agentic system evolved over time so maybe maybe a little bit of background right so rewind man yeah we've been doing this a long time huh our first thing was also built on top of uh on top of uh linksmith we we did tracing on there originally too yeah just starting right off with a plug you're welcome um uh so founded this company unify what do we do uh we essentially do ai for go-to-market right is kind of like what we say. That's kind of like the pithy line.
1:54The formulation of the company is essentially go to market as a search problem, right? So you find people in companies that have a problem that your product uniquely solves, right? You know, some people think that the job to be done of a salesperson or like a marketing person is persuasion, right? Is to convince you to buy their product, right? When you talk to great ones, right? What they really say is, how do I find the people who have a problem that is so bad that they'll buy it no matter what I say? And when you formulate it like that, right, it's search over huge amounts of unstructured, semantically rich data, right?
2:28It's like all the information on the internet, you know. As it turns out, we have a system now that's really good at parsing all the information on the internet. It's, you know, LLMs. And so we built this first system that was basically scaled web agents to do research. Back in the days where you had to do the reflect, you know, loop, tool call loop. we essentially built this web agent that would take in a question with a structured set of options, you know, structured output, a set of, you know, questions like, does this company perform KYC, right? So let's say that you sell a tool to do KYC, right?
3:03You want to know, does this company perform KYC? And if they do, they probably need a KYC tool, right? If they don't, right, you're not going to try to sell them, right? And so we built the system that would do that web research, right? now this is like feels like a very trivial example but three years ago it was really cool it was really really novel uh to run this for our customers i think our first version of it was using like maybe 4.0 was the was the first version of it it was like 4.0 it would like call a tool it would reflect and then it would call another tool um the big improvement for us also was uh when o1 came out which was very slow really expensive we gave it a a planning step to start off with so So basically, we might have even done a blog post about this together.
3:47I think we did, yeah. Yeah.
3:49Connor Heggie:Wow, that's so embarrassing if we go back and look at it now. We were way ahead of our time. And so the first step was this planning step with the reasoning model, which we thought was mind-blowing. It would come up with ways that it might go find this information, as well as pitfalls that it might run into. Massively improve the quality of it. And then we would run it over our company, our customers' whole TAM, basically, their total addressable market. sort of the world of companies they might sell to or might want to sell to. And on that, you know, KYC example, you might think that it's something really obvious, like, oh, well, think about this company.
4:22And if they're a bank, they definitely need to perform KYC. And yeah, like, those are the really obvious ones. But the interesting part is in the long, long, long tail, which for most companies is actually, you know, a double digit, like high double digit part of your TAM, of your market. it. And so we would actually, that agent would, you know, spin up, it would go look at, you know, like, let's say we were running that on you guys, right, on Langchain. We would spin up, we would go look at your terms of service. If the terms of service mentions that you get access to social security numbers, that's a pretty good indicator that you're going to need to perform, that you're performing some sort of KYC, you know, et cetera, et cetera.
5:00We would look at the HTML of the page to see like, hey, is there mentions of, you know, KYC or identity resolution, things like that. And that's work that an, you know, an SDR, an AE, a salesperson would be doing. And our job was to make that 100x more scalable or 1000x more scalable, you know, more cost effective, faster, and sort of bring leverage to one person, right? Whereas before you might have a team of 10 sales reps who would be going and doing this research and, you know, figuring these things out and selling to those companies. And instead with Unify, with Unify TM, you would have one sort of growth marketing person or demand gen marketer, sales leader who would spin off a million agents a month.
5:42We would run it for you over your, over your, your market and then do the work that that team would do again, like much more scalable and, and effectively or, or cheaply, you know, quickly, et cetera. People internally will probably laugh. The three things I always say is we wanted to build repeatable, observable, scalable systems, which doesn't play very well with AI, as it turns out. But those were always the tenets that we were going for. And so how do we bring scalability, repeatability, observability to go to market, which is again over semantically rich, unstructured data, using AI agents.
6:15Connor Heggie:And just to maybe talk about the scalability briefly, and this gets into like a little bit of like the UX of these agents. Do people interact with your agent via chat or does it operate over almost like a table of leads and do thousands or millions of them in parallel? Yeah. So this original version that we launched three years ago, iterated on a bunch. It was part of this workflow builder that we had. So we would do all this work to join first and third party data for you. We would look through your CRM. We'd look through, you know, potential companies that you might want to include in your CRM.
6:45Then we'd run it in this big async background system. And so actually what it looked like for a user was I would do the classic, you nodes and edges, sort of drag and drop workflow builder. Then I'd kick it off on a table of a thousand, a million, whatever number of records. And then I would come back in a day, right? And I'd like stream in, trickle in. You would refresh the page because it was a nice dopamine hit. But these agents were really meant to be these async, long running background kind of things. And so the thing that we really optimized for then was how do we get these costs as as low as possible, to run them as high scale as we possibly can.
7:22The transition to where we are now, so that wasn't through chat. The thing that we just launched last week, very exciting last Wednesday, was a chat product, right? And so how do we bring all these primitives that we've built for this, again, scalability, repeatability, observability, to a chat product that an individual sales rep can use to get all that value that a growth marketer, a demand gen marketer, go-to-market engineer was getting before, but now given in their hands, right? There was kind of this funny notion that's been tried over the last, you know, two years now that, you know, everybody thought, like, one of the big mainstream narratives was we were going to automate away sales reps, right?
8:01Replace sales reps. And there was a bunch of companies that tried to do that. They took big swings at it with different levels of success. You know, we always thought that we wanted to make the sales reps more effective. We wanted to take the busy work off their plate to get them back to doing what they do best, which is like talking to people, right? You know, I don't think humans are going away in that part at all. I certainly don't want to buy from an AI agent. You know, if you guys start to sell to us with an AI agent, I'm going to lose my mind.
8:26Connor Heggie:You want to give humans more leverage. You want to give humans more leverage. Exactly. And so we always really had this thesis. But the thing that the models have gotten really good at is engineering, right? Is all of the work that supports people, at least in the sales context, like the Rev ops kind of work, the go-to-market engineering kind of work. And so instead now we're giving every sales rep kind of an engineer in their back pocket that they access through chat to go and hit APIs, to build tables dynamically. And one of the things that we've learned now is like the frontier models have gotten so, so, so much better.
9:00Usually instead of running a million agents over all this, you know, data, yes, you might run, you know, 10 ,000 or 100 ,000 over part of it. But the bulk of that distribution, sort of the head of that distribution, you can get by stringing together a bunch of APIs really intelligently, like you or I would go do it if we sat down and were like, okay, I really need to pull this data for a million companies. Okay, I'm going to sit and think, how can I go pull from this data vendor, this sort of, you know, output classification, and then I'll just run like a lightweight classification on it to get the outputs, you know, et cetera, et cetera.
9:31And so now that's what this new chat product really does. It's kind of like, you know, how do you give an AE or an STR, a sales rep, you know, an engineer in their back pocket to do all this work for them in a scalable way?
9:43Connor Heggie:Super interesting. I mean, one of the things that we've talked about is how a lot of agents used for a wide variety of domains start to look or kind of look similar to coding agents. It sounds like that is very much the direction that you guys are headed in. So like I have to ask, how similar does the harness that you guys use look to a coding agent harness? Hilariously similar. So our, we did a lot, a lot, a lot of work in this harness. We've talked a lot about the harness too. A lot of learnings, I think, have come out of it. The thing that's different about our agent compared to a coding agent, just like in principle, the output of a coding agent, the artifact of it is sort of a modification to a code base, right?
10:25To an existing set of files. For us, the output isn't a modification on a set of files or file system. It's more about, you know, a list of records in a database that it's, you know, modifying a series of emails that get written and then sent, things like that, which aren't quite, you know, that like edit to a file system. And so you have to do, you have a different like premise or a different goal, you know, different outcome of the trajectories. But everything else is all the same. We have it write code. The way that it gets to those ends is still by writing code. Does it run in a sandbox? Yeah.
10:57So this was, oh my gosh, this was like a whole thing that we, again, we've, we talked a lot lot about. We, so the two sort of like sets of evaluations. Okay. So, so maybe like pause and like take a step back, right? Like what are the, we thought a lot about what criteria, what sort of, what sort of like features we wanted of these, of these, like, you know, the agent harness, like what would make a great agent harness for us specifically? So some criteria of that one, we run these in the cloud. So, you know, engineers run them on their local machine. So they get all the permissioning and durability and stuff like that from the local machine.
11:29We don't have that luxury we run in a cloud you know web web app also with their sales reps that that we're doing they
11:36Connor Heggie:don't want to run things on their machine like it's just non-starter so it needs to run in the cloud um so part of running in the cloud is it needs to be durable so if it gets killed midway through it needs to restart really gracefully we need to have sort of a like uh some system that runs this in the background because it can't run in like a server route right because you know uh our alb is time out after 60 seconds blah blah blah etc right and you have to be able to hook back into it. So things like that. So you need to be able to basically like unhook from this and hook back in. So it's running, you know, durably, even if there isn't a UI on top of it.
12:06Another criteria that we had for us internally, we wanted the code that it was writing to be TypeScript. So our whole code base is in TypeScript. I know a lot of people's code bases are in Python. And so there's a lot of, and the agents are great at writing Python. Another piece afterwards, again, with this cloud aspect, we needed to have environments that would be really, really cheap for us. Cause again, we run a lot of these needs to be able to spin down really quickly. We also wanted it to spin up very quickly. So there's kind of a bunch of considerations there. And the last one was around a file system.
12:34We needed it to have access to files. But in our case, we didn't need it to have strictly a file system. So we could kind of have like, you know, in the world of like, you want everything right there. That was the one where we're like, okay, well, like,
12:45Connor Heggie:what does that mean? Like, what type of files did you want to have to have access to? And how would it have access to those if not through a file system? Right now, we use the native OpenAI files API. So it just kind of like reads them in. We also do some wrapping around it. So if you upload a CSV, for example, which is kind of the most common data that our customers will send us, we'll create a native sort of table artifact that we call it. Oh, that was another criteria as well. It needs to be able to operate over tabular data really, really effectively in like a really native, durable way, scalable way.
13:17Are models good at operating over tabular data? No, but they're really good at writing pandas-like code, which is really good at operating over tabular data. So yes, we had a whole thing around how you would basically make it accessible to the model. We ended up with this sort of class that wraps a data store that we basically implemented. It's like, imagine like kind of a, you know, we virtualize the tabular data. It's actually just like in the database for us. Then we re-implemented a bunch of pandas-like functions on top of that class to let it do. In TypeScript, though, to let it do things like filters and, you know, fetch rows, ag rows, re-sort columns, change data types, you know, map, map rows.
14:02That was a really important one. So if you want to run a snippet of code over every row in the table, that's actually really pretty hard to do. And so how do you do that over a thousand, 10 ,000, 100 ,000 rows, things like that?
14:14Connor Heggie:Were there any functions that you implemented that were like not part of Pandas that you implemented specifically for this kind of like LLM access purpose? Yeah. So we didn't do a strict like rip over the Pandas interface. We debated a bunch about that, which in a funny way, the fact that we were debating about literally rewriting the entire Pandas API, but in, you know, types and there's pollers. There's like, there's TypeScript native versions, but we needed this again, sort of over this, you know, virtualized table that we have to like do a bunch of modifications of the database itself in order to access this data.
14:47Just to put a pin in it, we also had the last criteria that was really important to us, which is like absolute P0 is strong tenancy. So any data between our customers, our tenants can never touch each other. And in fact, that needs to be not only deterministically, you know, encoded, but it needs to be, you know, provably impossible for the agent to get access to the other things. So there's a big security point on data access, which plays into this, right? So then when you're accessing over the database, a couple of things that we considered, one was re-implement the whole Pandas sort of interface.
15:19Another one was to give it access to just straight up SQL. So you could imagine having it like write SQL. There's a bunch of considerations over that. I think we have like a kind of cute solution that we ended up on that we didn't ship yet, but I think is a good idea that gives you kind of like, you know, row level security to do like, you know, strong tenancy and things like that. But then where we ended up was basically like, the agent actually has 10 things it commonly does. It like reorders columns to make it easy for the user to see them. It maps a function over all the rows. It adds rows.
15:52It deletes rows. It filters the rows.
15:54Connor Heggie:Can the function that it maps over the rows be like an LLM or a subagent itself? Yeah. And so our formulation is that a subagent is just a function column. and because all of these so this is another important aspect so it's just an arbitrary function call and then it can pass all sorts of params into the function call right so it can pass the prompt it can also pass the model that it wants to run on the you know reasoning budget etc etc and then it just awaits those we chose our architecture on that as well as the like the mapping of mapping over all the rows all of these are sort of this async pattern that pretty much all the coding agents have moved over to where it kicks off this job, kicks off the subagent, gets a handle back to that running, you know, job by ID.
16:42And they can either await it directly, it can pull it directly, things like that. And that was another important point, any sleeps that we had, we needed to be non-blocking. So much of our stack runs on temporal on the back end. So you need this, like they have this like non-blocking await that you want, or non-blocking sleep. And so that was like, you know, this like first class thing that we needed. And so the map rows can basically run sub agents over and then it can use it to fill in a row. But the great thing about that is if the agent is writing code, you know, running sub agents is really expensive.
17:12So you have a thousand rows, even if each one is one cent, that's 10 bucks. I do that math, right? That would have been embarrassing if I didn't 10 bucks. And our base plan is$20. So you would run out of your whole plan, you know, basically immediately. So instead, by giving it access to it as one function call in a, you know, map that it might write over, you know, with code over a whole data table over the rows, it can call one API and then waterfall down and do another API and then waterfall down and do another API and then only if it's exhausted all the other options, call the sub-agent and use it as sort of the fallback.
17:48So that's a really common pattern that it uses.
17:50Connor Heggie:How many of the agent runs end up manipulating this table in some way? Most of them, like 80 % probably. Might be higher than 80%. Our product has two main jobs to be done of it. It's identify, which is find the people and companies, right? And that includes kind of enriching data on them, finding more information about them, filtering it. And then the second piece is engaging with them. So actually sending them email, sending them LinkedIn messages, things like that. And so in order to get to the second piece, you have to find the people that you're going to send an email to. So almost all the time you're pulling people.
18:25And when you're pulling people or companies, you're pulling them into this table.
Read the full transcript
18:29Connor Heggie:You mentioned prompt caching. That's definitely a topic that we see a lot of people talking about. How did you guys implement it and what were the effects of it? So all of our agents right now run on OpenAI models through the OpenAI API. We also have some capacity on Azure so that if there is an incident on OpenAI, you can kind of like, you know, gracefully fall back. It's like a bunch of things around like, because we use the responses API, there's a bunch of stickiness between it that you have to like worry about. And so like, what is prompt caching? You know, I know that you know, but just like for anybody who doesn't, basically the outputs to us is when I send you the next step in a, you know, a model thread.
19:10So, you know, we're still doing reflect tool call, but it's just trained into the model now. So, you know, the model runs, it does a tool call. And then when I send the result of that tool call back to request the model to run more, you want to get a prompt cache, which basically says, read the KV cache, read the activations that you already had that you were just running for me, but just add a little bit more. And it's 90 % cheaper through the OpenAI API than it would be to send the whole thing. So it's like binary makes your product work or not work, basically, as well as speed it up. So it's really important that you get these cache hits, or otherwise you're just, you're cooked.
19:48By default, what OpenAI does is it takes the first, I think it's like 70 characters or 60 characters, some number of characters and it hashes it and then it routes based on that. That's really great if you have low volume. If you have high volume, that prompt cache key, I think it's 15 requests per second. So if you are above 15 requests per second, they make no guarantees on any of this, by the way. It's all best effort. So you can be mathematically perfect, which I'm not saying we are, but like we're, we did the math and we're pretty close and we're at like a 95 % cash hit rate or something, which is fantastic.
20:22Connor Heggie:So I didn't actually realize this. You're saying basically like if I send a hundred requests, even if all of them like match the prompt prefix, only like 15 of them roughly will like hit it and the others will not hit that prompt cash. Yeah, so let's take this case, right? So let's say you send one request with, you know, 100 ,000 tokens or something and that goes down. and then the next second you send 100, you will get 15-ish, right? Interesting. Yeah, prompt cache hits. But those other 85 will route to different machines on OpenAI's backend on the cloud that will then get the cache themselves.
21:01So the next time you send it, you will be more likely to get a cache hit because you have a bunch of machines warm, right? A bunch of the GPU racks warm, essentially.
21:11Connor Heggie:actually realize that's how it worked. I'd assume that if you sent it, you would always get a cache hit. No, that's on the first message. And then after that, they use the response ID if you're using the responses API, which everybody should use the responses API. If you are not using the responses API, you are missing out on 20 % or 30 % of the quality because you don't retain the thinking trace between different calls. So the response ID is actually used to route then to the correct machine. You can still get prompt cache misses there. But again, if you, for example, let's say forked a chat with subagents.
21:43So we have two types of subagents. We have child subagents, which is a full fresh context window. And we have forked subagents, which takes the current chat and forks it, obviously. And you pass the previous response ID and you spin up 10 or 15, you're probably okay. But if you spin up 20, you get a full cash miss on your entire context, which could be huge. And if you're using 5.5 is really expensive, right? It's like 80 cents or something. Again, on a$20 a month plan, that's like a big, big, big deal. And so you want to keep basically, there's like this game between keeping enough of them warm versus not.
22:16Another interesting point. So there's the user message and then the agent, you know, assistant message will have many model calls inside of it with the tool calls. And then the next turn, right? So then it gives its final answer and you're like, cool. You know, I ask it, how much money has Unify raised? and it says this much money. And I go, cool. Okay, thank you. What about this other company, right? The next message down, by default, right, the prompt cache key is hashed, right? You can pass in your own prompt cache key to basically say, hey, nudge, use this, right? Even if you get that perfectly right, on the next message down, they remove from the previous turn all of the old thinking tokens from between the messages.
22:58Does that cause a prompt cache miss? It causes a prompt cache miss. Why do they do that? Because it ends up being fewer tokens overall. And they think that on average, people want that. Do you want that? We don't want that. So if you do the math for us, the next message down, getting a cache miss. So let's say I sent a request and something really complicated, right? It took 15 minutes and 100 tool calls or something, right? Maybe not 100 because you'd have compaction and that messes it up. Let's say it's like 20 tool calls. in our case that you know maybe there's a couple thousand thinking tokens between that you know maybe it's a few tens of thousands of thinking tokens right that would get removed out that causes that whole thing to be missed right the issue with that is our output from tool calls is much much bigger than that 10 000 tokens of thinking state so we're paying for all that when we could have gotten you know the 10x reduction and so instead they they i think last month released an API param that lets you retain those immediately increases your prompt cache hit hit rate or it's not that the hit rate it's that the amount that you get in your prompt cache hit
24:08Connor Heggie:like the amount of cache tokens is much higher did you find that that also helped performance because now you're seeing these thinking tokens whereas previously they were removed i know earlier you said that it really really mattered does it matter in between turns or is it largely just for that for that single turn we found it mattered some to be honest so hard to measure that like you know does it like they say it doesn't really help like you know i look at 100 examples 200 examples i'm like it kind of helps like who knows the thing that's really important though is it's much faster for our users it's much cheaper as well because you don't have to reprocess all this tool call input at full api price input and so that that was one of these big improvements.
24:47But now it's really important because you have this game of, let's take the sub-agent case, for example, which is kind of the most illustrative of this prompt caching issue, this 15 requests per second issue. So let's say I'm going to spin up a thousand sub-agents, right? Which, by the way, value of our product, right? You can use Claude if you want to spin up three sub-agents that can do that. If you want to spin up a thousand, they can't do that. You can't do it on your local machine. You really need this durable cloud-hosted thing. Cool. So I'm going to spin up a thousand, a thousand sub agents.
25:17Um, let's play the, the like prompt cash game, right? So what do I do? I send my first one, I get a full cash miss. Okay. Dang. So my next 15, I get full cash hit. Awesome. So do I just process 15 at a time? Like that kind of sucks. Okay. So what if I process 15 in different, uh, you know, prompt cash keys, and then I can send 15 times 15-ish, right? So you can kind of like fan out. Well, that's great, okay, for a single user. I can maybe get, you know, you can do the math for like what the right amount of fan out is to start off with to get, you know, the optimal sort of like gradual warmup. But we have this great benefit, which is most of, for these subagents, most of the cost for us on input is this huge upfront context of system prompt and tools and things like that that are retained across users.
26:07So how do I play the game of, you know, I have a thousand users that are sending messages in a day. How do I distribute my cache, my warm cache across all of them? And to make a very long story short, the way that we kind of ended up tuning the knobs, you take the user ID and you hash it. So we have 16 hashes of user IDs. And so if you're user X and user Y and you hash the same thing, you'll share this first part of the prompt cache key. and then we just have a random number between one and 30, right? So for us, you know, we average at our peak, whatever that is, right? It's like 30 times 16 or something.
26:49You know, we average that about at our peak throughput on subagents. And so we want to have that maximal distributed across all of these different prompt cache keys so that they're warmed across a day to then do it, right? There's like, you know, go ask Fable to do the math for your like specific numbers, does it perfectly.
27:05Connor Heggie:Do you think that a lot of this prompt caching logic will just be built into the providers over time? Or do you imagine maintaining this? It can't be built into the providers, which is the interesting thing about it, because they don't know the distribution that's coming in. We have a bias, right? You know, like back before it was cool to do AI, we were doing ML, right? You know, traditional machine learning. And the thing that we talked, one of the things that we talked a lot about was inductive bias. And so what is the like, you know, inductive bias that we have or the bias that we have in our prediction mechanism?
27:33Well, we know that this is a specific user. We know this is part of a specific subagent batch of runs. We know that users on average have similar, you know, prefixes of their prompts. We know that users in the same tenant on average have the same prefix up to some point. And the model provider doesn't know that until they get that and get a bunch of the data. maybe over time you could imagine sending enough metadata to them that they then you know predictively look at all the different attributes that you're sending and then and then do it which would be an interesting product but for them they don't care right they charge you either way right like you know and like yes it's like more efficient because they'll push more tokens through the system but it really like kind of falls on you as a developer as it always does right like you know ai is amazing and all these like tools and stuff but like you still have to do the like hard engineering work.
28:25You mentioned fork sub agents. What's the use case for that? So let's say I'm researching, you know, you, I'm researching Harrison and, or let's say I'm researching five founders, right? Cause it needs to make sense to fork it. Doing like a ton of research there in chat. And I'm like finding really interesting stuff. I'm finding these interesting similarities between them. And then I want to go write each of them an email, right? okay like the you know easy way to do it would be to write the email out one at a time and so write the first email for the first person then write the second email for the second person the third email for the third person that's just really slow you know and now let's say you want to get with 20 people like that's really slow and then your context window gets really big and then you know like it actually bails out like midway through because these models are like kind of lazy you know they just go etc um although that's not not really a problem with the latest gen of models and with a fork subagent, you give them all one clean task and you say, hey, we've been talking about all these people.
29:21Just go write an email for Harrison.
29:23Connor Heggie:Why can't you use a regular subagent and prompt it with all the knowledge previously? I have my take on it, which is mostly from intuition. And there's two different pieces to it. One of them is you have to summarize all this stuff, right? And it's going to be a really lossy summary no matter what. Or it's not a lossy summary and you put a ton of information in this prompt. Okay, like I was saying, we've optimized prompt cache hits a bunch. Okay, now you're getting a full prompt cache miss. Like I could just be getting a cache hit and then 500 tokens out for the email. But now I have to go summarize all of the stuff so I get a big output and then pass it into the input of a subagent and then have that parse through all the input and then output, which is really messy.
30:07But on more of a like thesis-y level point, the distribution, you know, like imagine, okay so these what are these llms like these llms are like functionally modeling language right they model the distribution of language over like some trace or some trajectory like some chat right and i've gotten it into just the right spot with all the things i'm saying to it right because i'm i got i know what i want and i got it right to the right spot that i want like it's so obvious what the next the next thing is the just the perfect email for me and so it's in the right sort of vector space and it's like, you know, language model.
30:40And I just wanted to like output that. Whereas, you know, with the subagent, it's like, got to come up in some random fucking thing.
30:47Connor Heggie:You mentioned like sending emails. How do you think about human in the loop? Like I imagine I would want some form of human in the loop if I had an agent sending emails for me. How do you guys handle it? Both like maybe from like a UX point of view, but then also like, yeah, anything interesting technically under the hood? I'll do this maybe in reverse order. um technically under the hood we do a bunch of managed deliverability for our customers so um one of the great things you know um like i feel very fortunate i feel very lucky that we started the company three and a half years ago sort of pre ai wave because we had a bunch of this really strong like standard architecture engineering that you like you know to build on top of right you know we had this amazing deliverability system and durable you know durable email sending, as it turns out, like guaranteeing an email sends exactly once is really hard to do.
31:35Like, you know, you can get it to be like on average once pretty easily, but like not sending it, double sending it and always sending it is not trivial. And so we had all of that built already, which was awesome. On the UI side, the UX side, totally don't want emails ripping out of this agent sight unseen. What we ended up on was this sort of proposal and then approval system. So inside of our agent, we have an artifact, you know, UI, UX, very similar to what you have in Claude, where the right side is sort of a full screen or almost full screen view. We'll put a table in there most of the time when you work with a table or for email sequences for emails.
32:17We have this sort of quick approval flow where you can scroll down and read it. And then we, you know, and that's kind of the easy version, right? So that's like V1. And you go in, you can edit the email, you can look at it. But then one of the things that we found was that all of the sort of second order interactions end up being really important. So the highlight a sentence and then ask it to change it, not just on this email, but across 100 enrollments.
32:40Connor Heggie:Yeah, I was going to ask, like, I get how you might show one email on this right hand. And I want to talk more about UX because I think UX is pretty under invested in and under talked about. And so I love that you already went there. So like, I imagine what that could look like for one email, but you're talking about like maybe a thousand emails. Like how do I, like, how do I insert myself in the right position there? For us, um, a couple of things that were, that were really important. Um, one, you had to be able to see the full email and you had to be in context, right? So you needed to see who I'm reaching out to, the job title, what company they work at, a bunch of information.
33:11Connor Heggie:For all like thousand rows in the table. For all, for all thousand rows, right? So we played a bunch with, do we show, like you're saying in a table, the full email, and then, you know, columns off to the left. We tried it. Ended up being really noisy. People I found, myself certainly, and seeing other people work with products, you're much more effective going one by one really quickly than you are looking at a bunch in a row. We had this internal tool. I used to work at Scale, which is a data labeling company. We used to have this tool internally there called Speed Audit, which basically let you queue up a bunch of work and then you would just hot key central, look at it, move on to the next one, look at it, move on to the next one.
33:50And it was really effective because you could zone in, you could work through it one by one, right? Sort of like, you know, the modern day, like Cal Newport deep work of just like focus on the thing and like do the thing really effectively. So instead what we opted for was show one email with all the contacts you need very quickly. But then if you want to modify things, give you tools to modify it across the whole batch. Although I might be looking at one email at a time over five or six or seven or eight, by the time that I've gotten to the 15th one or the 20th one, I can give them modifications and then I'm probably ready to send all of them.
34:25Connor Heggie:So if there's like a thousand emails, you generally see people maybe like reviewing the first 15 or so, giving enough feedback where it applies to all of them where they're comfortable setting, but they don't go through all the thousand. It really depends. Some people, this is, if humans were going to spend time in one part of the system, this is the point to spend time in, which is kind of the, you know, funny thing. You know, SDRs, the salespeople that we sell to, this is their secret sauce. You know, it's not like Google searching until I find the right company. So like we totally automated a bunch of that work, right?
35:00Gave them the scalable things on that, but they really want to make sure that they send the perfect message, right? There's not an m-bash in there, right? They're talking about how they surf and they saw that this person surfs and like they were just, you know, surfing at, you know, blah, blah, blah. You know, I just went to Hawaii and it was so nice, you know. That like human connection point is actually their alpha. And so, yeah, we'll make it really easy for them to do that part of their work. And then if they want to like ship all of them, like, you know, command A ship, like that's fine.
35:27Connor Heggie:On that note, how do you guys think about personalization? Because I imagine every seller probably has a different tone or different things they like to mention. And great, so they can go through this flow and correct it. Do you remember that over time? do you let them set preferences like upfront? Like what does that personalization story look like? Yeah, I'm so glad you asked that. Great, great question. So we have two flavors of this. One of them is memory, which is built into the agent. The other one is sort of more of this Unify specific thing that you're not going to get in Cloud or ChatGPT because we specifically are building for sales reps.
36:02And that's when you integrate your Gmail mailbox. We look through your whole inbox. we come up with the ways that you like to write emails and not just like the generics of your tone and the generic you know like the things that all sorts of products you know might do when you integrate your inbox to try to replicate you but we learn how do you talk about your product and why and what types of people do you say what types of messages and is that is that different than memory it's a different process to get it in but then we push it into our memory system so that it can be recalled at any time.
36:35It can be pulled into the email, you know, content writing pieces. But the system to get it in isn't a, hey, remember that I like to write XYZ kind of email. Or even, you know, we do have the, okay, now edit this email to say X, which is great. And there is some signal there that we push into our memory system. But maybe I just like this one email wanted to say something different, you know, when I say, oh, hey, like make this kind of funny. Maybe it's because I'm sending, you know, my friend an email, not because I want all of the emails I send as a sales rep to be funny.
37:10Connor Heggie:So I want to dive into this because I think this is like pretty similar to some of the other things we see in memory where there's maybe this like upfront kind of like just knowledge extraction. Sometimes like we wrote a bunch about this this past week around like this like wiki creation. I saw that. It looks sick. Yeah. I mean, we see this in like coding agents, you have like deep wiki from cognition and things like that, where they do a bunch of upfront work, they take this raw material, in their case, the code base, in your guys's case, it sounds like their email history, and they produce some like condensed version of it that's like, useful for agents kind of like going forward.
37:42Connor Heggie:And then there's also this other like, as you're interacting with the agent, how do you how do you take learnings there and update? So I actually want to talk about both of them. Question, like, what is the underlying data structure? Is it a set of files? Is it a wiki? Is it a knowledge graph vector store? Like what is this data structure for you guys? So it lives, lives in our Postgres database. It's essentially, let me maybe like explain the overall extraction and it'll end up like with the right data store. There's the email, there's the bootstrap step, which is kind of what you're saying, which is this initial knowledge extraction.
38:12And then there's the ongoing system, which both proposes, right? I'm really big on this, this like thesis or concept of generative and then discriminative, right? You're convergent, divergent sort of thinking where we propose, propose over propose a bunch of things of, hey, like in this message, I think it's every third message or something you send right now, we will take, you know, proposed memories out of it and say, hey, this user might have said blank, this user might have said, you know, might like like this might like that. In this kind of scenario, do this kind of thing. They're all structured.
38:44So they're, you know, we have email voice, we have a general user preference, we have hard user preference.
38:50Connor Heggie:So you basically set like a schema of this data structure that you're trying to ingest things into. Yeah. And so it's, it's arbitrary strings. So it's, you know, natural language for the content of it, but then we give it some structure above that to give it like different classifications so that at recall time, we always recall at least some of the right ones. And then we can do the, you know, semantic search over the rest of them. Okay. A few questions. Yeah. So, so like the, the keys, not the values, the values, arbitrary string, the keys, the structure you're imposing, is that a fixed structure?
39:25Connor Heggie:Can it like add more keys? Can it add more structure? No, it can't. So we gave it, I think there's like six or seven of these. And they're, like I said, they're, again, we sell the sales rep. So we have a good structure on top of it. And we have, you know, email voice, right? Like rep voice, attributes about the rep, attributes about the company that they work at, user preferences, you know, soft user preferences, hard user preferences, a few things like that. And then we have an other category, which is kind of our bucket all, you know, for that. We decided not to do the arbitrary keys and have the agent sort of self-manage, largely because every added dimension of complexity is multiplicative instead of, you know, additive.
40:10And so having it manage both the keys by which it'll classify as well as the classifications, you know, felt scary to me. Maybe we end up there, you know.
40:19Connor Heggie:I'm curious whether you'd agree with this. I think we generally see that the hardest part of this is like getting, I'm assuming you're using an LLM to reflect on both the raw data and the conversations. The hardest part is usually getting the LLM to decide what the right thing to remember is. The data structure is not that hard. As you said, it's a key value store. It's not that hard. The hardest part is getting the LLM to remember. And so, yeah, like enforcing the structure and like saying, hey, we only care about remembering these seven things, I think makes a ton of sense if you if you can do that reliably it's all about that like inductive like what is our inductive bias like what is the thing that makes unify special and it's that we care a lot about sales reps and what they need from a chat product and we know what they should you know what the agent should remember about them totally so so the ingestion of like the raw emails i'm assuming there's like some agent that's doing that basically and deciding to write to those keys and then for the conversations you said like every three messages you basically run this this process Lightweight proposal.
41:10And so what'll run is it'll overgenerate proposals. And then we have a system that'll run on a cron in the background, which takes all of the proposed memories, as well as all of the historical active memories. and then we do a sort of justification step or a processing step where we either promote a draft memory to an active memory, we merge memories, we deprecate memories, you know, and then we actually give it a pretty structured set of lineages to process then over time where it can't, it doesn't just decide what to do, you know, here's my current memories and here's my new memories. It has to say, okay, memory A and memory B, the op I'm going to run is merge these two and then here's the new memory c and then output that or drop memory d completely right and then it's marked as dropped or this memory actually supersedes this one and so you know this one's superseded and then you end up with a new set of active memories why does it matter that it like
42:11Connor Heggie:marks something as dropped like do you use that like history in any way or is it just yeah why does it matter observability right for us we can go in and a lot of a lot of the LLM engineering is just black box right you say you know and we have evals that are great right we love evals you can you can learn things over time but so much of it is a black box and just like imposing structure in the right way to make it observable to you so that when we as engineers go and look at the system in three months and we say well that was pretty dumb like what was it thinking there, at the very least, there's like some structure to say, oh, well, you know, it merged those two memories.
42:53And I guess I could see why it merged those two memories.
42:57Connor Heggie:Speaking of that, like how much do you expose to the end user? Do you let them see their memory? Do you give them control to say like, don't remember this, remember other things? Do you let them approve changes to their memory? We don't show them in app today. If you ask it, what are your memories about me? It'll tell you. We don't have a structured UI. We've gone back and forth on that. We might add it. I don't, as a user, want to see it. Like, I'll ask you what you remember about me. You know, like, I don't want to look at a bunch of bullet points about me. I also get like a little bit of the, you know, like, I get like embarrassed about what the LM knows about me or what the agent knows about me.
43:32So how much do we let the user see? You know, not that much. The thing that we do take, though, is one of these classifications that we have of types of memories is strong user preference, direct user preference. And so anything that the user says, hey, remember blank, will always be inserted as this top level kind of like first class concept. The funniest use case for this is there's a bunch of people that are say, you know, remember to always talk to me like a pirate. And it remembers it every time, which is so funny. Or talk to me like a dude, bro. And it talks, you know, and like so many of our customers are, you know, their sales reps, they're like usually oftentimes is their first job out of college.
44:10You know, they're like pretty new. They're like, you got to have some delight and fun in the product. And this is totally one of those ones that ends up being a delight.
44:18Connor Heggie:How advanced do you think AI penetration is into kind of like the sales world? Like, it's probably not as far along as coding. So I imagine you get people who are maybe earlier on and trying things, but like, where is it? Is it super early on, middling? Yeah, I would say everyone's using it somehow. Teams aren't standardized or structured today. And so one of the things that we hear all the time from sales managers, sales leaders, even C-suite execs, is how do I figure out what the most effective AI usage for my sales reps are and get them all to do that? In fact, it doesn't even need to be the most effective use.
44:59It just needs to be like a pretty good use. And then how do I get them all using it? Then how do I hold them accountable to using it? How do I, you know, get them using it in similar ways? How do I track them with AI? How do I make sure that they're not spending a million dollars or a thousand dollars a month or whatever your budget is per person?
45:17Connor Heggie:A few questions there. I imagine one version of this could be like, hey, someone's got a really good skill. Let's make sure everyone uses that skill. Do you guys have a concept of skills in your platform or are you guys basically like, hey, the workflows we have, these are basically souped up skills that are way better and it's just used this? Yeah, it's 100 % that today. We've done so much. We have dozens and dozens and dozens of skills and skill files and guidance and things like that across different data sets and data providers and the way that you would access these data providers and whatnot.
45:47So today it's all sort of like Unify curates all of it. We have all these best practices. You know, we work with some of the best go-to-market teams and we take, you know, all the things that we know and, you know, like have worked with them on and make it really accessible for everyone. Will we add skills? I think we will because different companies do have different theses. A great example is, you know, when you're accessing one company's Salesforce instance versus another company's Salesforce instance. I know the most riveting thing on the planet. All engineers love talking about Salesforce.
46:14It's very important. Really different on how to do it. And so those, you know, we would want. But whether or not we make them skill files or not, I'm a little bit torn on. Engineers love writing things like skill files. You know, it's like writing a runbook. But most other people don't. Why not just have it be memory about your company, right? Like my company, our Salesforce operates like this. Okay, well, that's like just a memory.
46:38Connor Heggie:You mentioned people caring about costs and how much people were spending on this. I saw one take on Twitter recently that I liked, which was like, every AI product will have some form of like cost controls built in for their end users to control that. Have you guys started doing that already? Or are you still in the phase where it's just so early on you're trying to get people? because like I'm thinking about coding up until six months ago, we were just like, yeah, use whatever you want. And, and, and recently it's become where we need something. And so it switched somewhere recently. And so I'm curious, yeah, like is, is that built into the product yet or is it on the roadmap?
47:13Yeah. When you sign up, uh, there's per seat limits. And then, uh, what we've seen is that's largely good for everyone. You sign up, you, you spend 20 bucks a month, 60 bucks a month per user and you get some allocation and then that kind of helps like guardrail you directly. And then you can pay for sort of pooled credits for people who go over very similar to like a, you know, Claude or ChatGPT model. We don't have the per user guardrails where, you know, within this then shared pool, let people, you know, guard within that. I think we'll build it really soon. We need something like that. It is really interesting though, because the important thing will be that it's per user guardrails.
47:52Because what we see is some people are really trusted. Like if we look internally at, you know, at Unify even, I think I spent like$7 ,000,$8 ,000 on LLMs last month. And like, I trust myself. I don't know. I trust that spent. Like I was crushing PRs last month. And then the next highest on PR or on, you know, token spend on, you know, dollar spent was like 5 ,500 bucks. And I also really trust him. He was cranking like that was well worth it. But there's other people on the team that if I would not trust to spend$5 ,000 a month, I would trust him to spend$500 a month. if that goes well like totally like let's rip open a thousand dollars a month maybe but there's kind of this like um you know like you kind of earn the right you earn the trust to be able to do that
48:37Connor Heggie:we also see i mean we see this in coding that like some of it's just like accidental spent yeah and i imagine if you're spinning up like thousands of sub-agents that accidental spend could also kind of like creep up on you but i think like you know again we didn't really care about costs until maybe six months ago and it sounds like sales is like earlier on so it's I think, as you said, like people want just more adoption and more standard adoption of best practices. And I think that makes perfect sense. One of the things you mentioned earlier, evals. How do you guys do evals? We use Langsmith to run our evals.
49:07Yeah, how do we do evals? We do a couple of different kinds of evals. So a lot of my thesis on evals comes from my time working in self-driving. So before, right before I started Unify, worked at scale on the ML team, the mapping team before that. And then before that, I worked in a small self-driving company. It was like 15 people, 11 of them were math PhDs. Like it was run like kind of a research lab. At a self-driving company, the product is the car that drives. And like the thing that makes it drive is the vision models, right? So how did they run it? How did, you know, how did we run it? Of course we had evals.
49:42Of course there were metrics that we looked at. And that was very important. But the gate to get a model on the car was this big grid. it was this DQA video this like video this like hour and a half or two hour video that had you know seven second clips it was this table of you know six rows down five rows wide or five columns wide and each row was a different model run different like training run and each column was a different checkpoint along that training run and we predicted the same video on all of those different checkpoints of all those different models with the current best at the bottom you know semantics It was a semantic segmentation at the time.
50:21And we watched it. We sat down as a team and we had a projector in the living room equivalent of the office. And we sat and we watched like 30 minutes, an hour of video. And we took notes and talked about it and said, you know, oh, like it actually gets this right there and this wrong there. And there's just no better eval than looking at 100 examples or 1000 examples. So we do a ton of that, a ton, a ton, a ton of that, oftentimes filtered down to specific distributions of things and issues. And so we'll classify every message as a certain user intent. So we have, you know, this user is trying to find people, this user is trying to send emails.
51:00And we might look through a bunch of examples for that to see how the model's doing or how the agent harness overall is doing. That's a bunch of it. And that's really helpful to get a current feel of like what's good, what's bad, you know, where can we improve? But then you need regression testing basically to make iterative improvements and so we'll pull traces from that into a data set which you then run an eval over we have different basically buckets of it we call them dqa sets dedicated qa sets or or like you know they're basically just groupings of these examples that are meant to cover you know some specific distribution of usage right whether it's our hero use cases that we just know should be rockstar, you know, rip out, or kind of adversarial use cases of people trying to prompt inject, or, you know, weird use cases where somebody is speaking Spanish midway through and we didn't expect that, or things where they, you know, etc, etc, right?
51:57And we'll iterate on one of those at a time, usually. So let's take, you know, today, for example, we have this specific data vendor that the agent is just hilariously over-calling, right? Like you show up and you ask for 10 companies and it calls it like 500 times.
52:13Connor Heggie:They did some good agent engine optimization. They crushed it. They did. And we spent a lot of money on that API. And we should not be because most of those results don't get to our customer. And so we have a DQA set, which is 40 examples where it went off the rails and called this a bunch of times, pulled those into a data set. And then we iterate on it with a bunch of metrics, a huge number of metrics of number of tool calls, the tool call efficiency, the trace efficiency, the credit cost that we would give to our customers, the LLM cost, kind of all these metrics that you would look at. How do you test services like that where it calls, where it costs money to call it?
52:51Connor Heggie:And then maybe also like emails, like you're not going to send emails as part of, well, maybe you will send emails as part of some tests, but like, yeah, how do you think about calling these services during tests that are either cost money or like take actions? So when sending the emails, the actual end result of the email sending doesn't matter to the agent. That one's easy. You mock it out. It like fake sends. Okay. Easy peasy. A harder one is ask user questions. So you ask the user some questions like, what do you do in the eval harness? So we like have another LLM play the user. But now you're starting to stack distributions.
53:23One thing that I always like, a drumbeat that I always hit with the team is when you have an LLM as a judge, it is, or it's interacting in this user interview point, it can never be the same distribution as the original model, right? So what does that mean? If we are running GPT 5.4 in our agent, we have to be using an anthropic model to do the judge, to do, you know, things like that. Probably for a bunch of obvious reasons, but just to say it out loud, right? You have this, like, you know, mode collapse of, like, agents talking to each other, this, like, you know, overlapping distributions where it's, you know, the equivalent of groupthink in humans.
54:00And you just, you badly don't want that. you want it almost adversarial, right? You want it in a completely different distribution, you know, like language modeling distribution. And so we do that for a bunch of them. On the, you know, APIs that we might call, we have two modes. One of them mocks it out and will replay sort of old ones. That's not a perfect test though, because let's say this new one, it actually just calls the API once, but with really good params to get the single result. Well, that's not going to get replayed. So it's not a perfect fit. so oftentimes we'll just actually take the cost hit and run it for real we are already spending a lot of money on the llm costs to go run these so the data cost is kind of marginal on top of that
54:43Connor Heggie:so i want to ask about that we've talked about cost a few times it sounds like you've got cost of llms cost of some of these search providers cost of sandboxes i presume in some form what do do for sandboxes? It sounds like you maybe have more kind of like stateless, quick ephemeral sandboxes, or is it incorrect and you actually have stateful, long-running ones? So they are stateful, but they're not long-running. This is so great. There's like the, you know, anthropic blog that was, what was it, like split the hand in the heads or something, the brain in the hands or something. I know what you're talking about.
55:14Connor Heggie:I think it's, yeah, brain in something. Brain in the hands, whatever it is. Basically run the agent outside, have a sandbox that executes things, but it's not running inside there. But it's not running inside of. Very important. Strongly agree with that. Especially important with this cloud execution environment that you might want because you need, again, durability. If it crashes, you need to be able to handle that really gracefully and things like that. Then this other point around how do you ensure tenancy to make sure that the data is always well-scoped to a single user, a single customer's data set.
55:50And so where we netted out, we don't have a full VM. So one of the really common options is you create, you spin up one of these, you know, like Daytona VMs, LinkSmith or LinkChain has a VM. Do you guys have just the, do you guys have the VMs and repels or just VMs?
56:04Connor Heggie:No, just repels. No, we have VMs. Yeah, you have full VMs. And so you would spin one up, spin a full VM and then, you know, inject commands into it. Yep. Which is awesome because you get statefulness, you get a file system, you know, you get a bunch of really great stuff. The place that we struggled with that was you then have to write a bunch of CLIs. So maybe you wrap a bunch of your code in CLIs, but then the CLIs you need to like, you know, ensure tenancy over them, which like then it can actually write arbitrary code inside, which was really scary to me. Then you have to deal with all the networking, which is also really scary.
56:39How does it access if it needs to access other services? Like, do you have it then pop back out and hit your API and then your API hits your internal? Like it's, you know, really, really complicated really fast. And so we stumbled upon this open source framework called Monty. you've seen. It's in Python, which is great. And essentially what it does is it gives a fake Python interpreter or real Python interpreter that the agent can spin up and run and actually execute real code inside of. Except you bind functions into it that once you hit that, you suspend the whole REPL and the whole code execution environment.
57:13And then you break and you say that your host process, your backend code gets basically a function name and some args and says, hey, it tried to run this function. You get to run that function however you want. Then we inject a tendency at that level. We say, cool, we're going to run this in the scope of this specific user, in the scope of this specific thread. We'll add a bunch of billing stuff inside of it, et cetera, et cetera. And then when you're done, you resume that REPL backup and inject the outcome, the output of that function call basically into it. Like I was saying earlier, we really want it to be writing TypeScript because we're fully TypeScript shop.
57:50So we rebuilt that internally to do TypeScript. And then that runs in our actual backend process. So we're running workers that are actually calling the LLM, calling the agent, and then spinning up these really ephemeral repls that add to over time. So even if it's calling it, you know, once, then, you know, again, similar to a Python notebook where it can access variables that it had previously, but, you know, now at the end.
58:14Connor Heggie:We just launched something in Deep Agents using QuickJS. Amazing. Which is, I was about to make a funny comment about how you guys are a TypeScript shop using kind of like using a Python REPL. But that sounds like it's not the case. What did you guys do with QuickJS? It's the same thing with like the ability to like host, like inject in host functions. Yeah, yeah. It's super useful. Super useful for in particular kind of like programmatic calling of subagents. Yeah. So that's the big thing we launched. basically write a script that has, we have like a task function, which is basically the same as just our subagent function.
58:50Connor Heggie:Just kicks it off and we have some nice prompting to like tell it. It's like a combination of like Anthropics, like dynamic workflows thing. Yeah. And like the RLM paper. I'm obsessed with RLMs. We haven't come up with a way, it's way too expensive to run for us, for customers, but it is brilliant. It gives you like basically merge sort, but over with LLMs. It's all just programmatic calling of subagents. But yes, to your point, like, yeah, we did some benchmarks and yeah, it gets better results. It's like, it is a lot more expensive. It's so expensive, but it's so sick. Yeah. Cause like, think about this, right?
59:20Today in Unify, you can go run this query and you can say, pull my book of business, pull the hundred accounts that I am assigned to, and then score them. Tell me the best one for me to reach out to today. Okay. How do you do that? You either have the main agent iterate over each one and remember, yuck, or you have it right, which is what we'll do right now inside of Unify is it'll do some research on all of them, create a scoring function, then order by the scoring function, which is very, very lossy, very noisy, right? How do you do that? So really what you want is basically like semantic merge sort.
59:50You could do that with RLMs, which is crazy, which is so sick, yeah.
59:53Connor Heggie:So it sounds like because of this, your sandbox code interpreter costs, zero. Yeah, great. Okay, so you've got LLMs, you've got search. How do those compare? Depending on the day, different. We actually got probably a 90 or 95 % cost optimization from two weeks before we launched to the day that we launched. It was a very big effort internally. We were very concerned because what that would have meant is you show up to Unify and you send one message on your$20 a month plan and you're done. And that's insane. What'd you do? So many things. Well, so we were running so many sub-agents and that's way too expensive.
1:00:33So how do you move to doing this main agent that's smarter, that can write code that it maps on top of the table? That was like a big set of optimizations. A lot of the optimizations that we found, like, you know, you go and you just like, really the answer is we dig into the traces. And you bucket things and you find ways that it's going off the rails where it's being dumb or not doing the right things. And then you optimize the prompts for it. you look through all the, you know, one thing that we did that is very obvious, but of course, like, you know, made a huge difference. We looked through all the skill files, we looked through the system prompt, and we made sure that there were no contradictions.
1:01:10Because every little contradiction meant that it messed up a tool call, it messed up like the ordering of something and did it. You know, I think we need this, like, we don't have this today, somebody should build it as an open source project, or maybe you guys should build it as a product. We need a set of like semantic linters over our skill files. Similar to how we do with evals that do things like self-consistency. Is within a skill file, is it consistent? Is every combination of skill files, if loaded, consistent with instructions? Please tell me you guys are building this.
1:01:38Connor Heggie:No, but I've thought about it because we have Context Hub, which basically stores kind of like skills and agent.md files. I hadn't thought of it for, I thought of it more generally because I think another use case is if you're updating these automatically or even not automatically, you don't want PII in them. Like that's just generally bad. and so having some like lint rule for that but like yeah i think like ensuring consistency or just like is it too long i don't know totally yeah like you want both like program at like programmatic versions of it which are like the hard you know like are there m dashes and like we don't want it to write m dashes so we shouldn't have m dashes in our you know instructions so boom right like easy one or like we try to enforce a rule which is we don't say do this don't do this never do this, because the more that you add that, models have just gotten so smart and smart enough that it should, if you give it the why behind that, they're really good at figuring out what to do and coming up with the right trajectory.
1:02:32But they also will follow your instructions no matter what you say. So if you say always blank or do not blank in the 3 % of cases where that's wrong, it'll never do it.
1:02:45Connor Heggie:So going back, you got costs down. We got costs down a ton. What percent of that is model versus kind of like the search? It's hard to tease out how much of it was model versus not. I would say probably of those reductions, a grand majority of it was fewer token costs. So actually bringing token costs down very, very dramatically. A large part of that also, we were using 5.5 instead of 5.4. We brought five, you know, we like brought that down a bunch. A lot of it also was just inefficient tool calling. So it was doing a bunch of tool calls to vendors it didn't need. So we did a lot of optimizations of how do you make sure that you are always calling functions that you will actually use the results of.
1:03:27That's like a big part of our evals, wasted tool calls, something to drive down. It's kind of the number one thing to drive down for us. Sub-agents, a ton, a ton, a ton there. But actually, like I would say on a maybe like a practical takeaway thing that really like moved the needle for us was adding in a really robust planning step at the beginning. You know, going back to our three years ago, adding a one, right? Or two years ago, adding a one to our, as a planning step to our four O agent, just tell the model to pause, whatever the user asks you, pause, think about it for a little bit, come up with the ways that you might go solve it, and then ask the user a couple of questions to clarify it and then go and do it.
1:04:12And then as part of that planning step, it'll scout out a couple of trajectories. So let's say you say, hey, find me 30 companies that aren't in my CRM that are hiring for AI engineers and have raised a series B in the last two years, right? It's actually like a pretty hard set of criteria to like find all of in one go. And so it might scout out and say, okay, cool. I'm going to try and call this data vendor and get some companies that are series B. I'll try to call this data vendor to find some people hiring for AI engineers. And then I'll look and say, oh, well, this one's higher recall, but this one's really high precision.
1:04:49So let me, I'll actually start with that one. And then from there, you know, go run the full suite of finding 20 or 200 or whatever it is.
1:04:56Connor Heggie:It sounds like you guys are mostly on open AI models. How do you think about model choice and model selection? Yeah. Yeah, we are. Honestly, it was a pretty easy decision for us to start off with. OpenAI was also our first investor. So I have to say that they like are great partners. They're amazing to work with. And the models are really good. They're really cost effective. I'm like clearly talking a lot about costs. And, you know, somebody I remember me five years ago would be listening to this and being like, oh yeah, you know, haha, costs are so important. Like, I just want this thing to be sick.
1:05:30And like the sick part is that our users can log in and send 30 messages or 40 messages, 50 messages, get a ton of value out of it instead of using Fable and it killing it on one message and being like, okay, well, here's another 50 bucks. For the price performance, it was really, really easy for us to make that call pretty early on. We evaluated both of them. We've been using both of them for a long time. Open source models at all? So we evaluated open source models. We did this big, big, big cost reduction in Q4 of last year where we took our old agent costs, that research agent, brought it down to sub one cent, which was awesome.
1:06:10So, you know, like we passed all of that value, by the way, onto our customers. We 10x reduced our credit costs to our customers. And actually now on that old agent, tool calls actually dominate the cost there. And this one in this new agent that we have, we try to think about it a lot as the value that we're giving to our customers is in large part, not entirely, but in large part, the data that we're bringing to them and helping them define and craft and filter and curate. So how do we keep the token cost to be a percentage, a small percentage of that data cost that we're bringing to them to then output?
1:06:48And so open source models would be a fantastic fit for that. The issue that we ran into when we evaluated them in Q4 of last year is they're so much less tool efficient that even if they are 10x cheaper, which they're not, but even if they were 10x cheaper on the token costs, the tool efficiency is actually not worth it, net net. That being said, I've been hearing great things about GLM 5.2. I've been, you know, trying it out on the side and it seems great. So maybe that is like the saving grace, right? And we will certainly fine tune and model ourselves very soon. You know, all of that good stuff.
1:07:24fable came out you know maybe not to date this this you know podcast but like fable just came out again yesterday and people are rerunning all their evals on it and it's actually a pretty cost-effective model because it uses way fewer tokens than certainly sonnet 5 which was using way too many tokens and so that trade-off is really non-obvious and again to go back to the evals point the only way that you can do that you know rigorously is with great evals all the evals that we run, we have a pass case, so we'll run like, you know, five runs of each, you know, trace to evaluate to see how well they perform, you know, rather than, you know, one at a time.
1:08:04But I really, I really badly want open source models to be there. They're just not yet.
1:08:08Connor Heggie:Thanks for listening to Max Agency. If you liked this episode, leave a review and subscribe. Send feedback or questions to maxagency at langchain.dev. We want to hear from you.
1:08:23Thank you.
From the publisher
Connor Heggie spent his early career on a fifteen-person self-driving startup run like a research lab, then moved to Scale AI's mapping team, before becoming the co-founder and CTO at Unify. Unify builds agents for go-to-market teams, and for the last two years, one of AI's big mainstream narratives has been automating away the sales rep entirely. Connor and his team built the opposite: an agent that gives every sales rep "an engineer in their back pocket." He walks through how Unify's harness evolved from million-agent batch jobs to a chat product, and unpacks the engineering that makes it cost-effective to run at scale.
–
We also discuss:
- How Unify cut 90-95% of costs two weeks before launch
- The 15-requests-per-second ceiling inside OpenAI's prompt cache
- Why your LLM judge must be a different model family
- The Speed Audit: why one at a time beats a table of 1,000
- Why Unify's subagents are just a function call
- What working on self-driving taught Connor about running evals
–
Timestamps:
00:00 Introduction
01:30 "Go-to-market is a search problem"
06:10 The old workflow: drag-and-drop nodes over a million-row table
09:40 Why the harness is similar to a coding agent's
10:50 Running durable agents in the cloud without a full VM
13:15 Giving models pandas-like superpowers over a live table
15:50 Why Unify's subagents are just a function call
18:25 The 15-requests-per-second limit hiding inside OpenAI's cache
24:20 Optimizing for prompt caching hit rates
28:20 Fork versus child subagents
32:35 The Speed Audit: why one at a time beats a table of 1,000
37:50 Locking memory to keys instead of letting the agent freestyle
44:15 Why AI isn’t taking over sales
49:00 What working on self-driving taught Connor about running evals
52:46 Why your LLM judge must be a different model family
54:45 Ditching full VMs for Monty, a Python REPL that suspends
58:55 Semantic merge sort: why Connor is obsessed with RLMs
59:53 How Unify cut 90-95% of costs two weeks before launch
1:01:15 The case for "semantic linters" over skill files
1:04:55 Why Unify runs mostly on OpenAI, their first investor
1:06:12 Why 10x cheaper tokens still lose on tool efficiency
1:07:24 Why open-source models don’t make economic sense (yet)
–
Referenced:
- Anthropic
- ChatGPT
- Claude
- Claude Fable 5
- Claude Sonnet 5
- Context Hub
- Deep Agents
- GLM-5.2
- GPT-5.4
- Helm.ai
- LangSmith
- OpenAI
- QuickJS
- Recursive Language Models (RLM)
- Salesforce
- Scale AI
- Scaling Managed Agents
- Unify
–
Where to find Connor:
–
Where to find Harrison:
–
Where to find LangChain:
–
Send feedback or questions to maxagency@langchain.dev




