In short
Inngest’s durable “step functions” execution layer for AI agents, focused on reliability (automatic retries without losing state), deterministic replay, built-in observability (agent traces/trajectories), and outcome-based evaluation using real product signals for safer iteration, A/B testing, and cheaper model routing.
Guest
Tony Holdstock-Brown, from Inngest. Background described: previously led engineering at a healthcare company with highly event-driven, auditable workflows (patient forms, doctor approvals, prescriptions). He also spent months rebuilding durability infrastructure on Kafka/queues around 2019.
Key claims
- AI agents need durability and deterministic execution; otherwise failures force whole-run retries.
- Inngest abstracts queues/state/events and retries failed steps exactly where they failed, preserving prior context.
- Coupling infrastructure execution with trace generation enables deterministic replay and faster self-improvement loops.
- “Local evals” (unit-test-like) aren’t enough; best practice is grading outcomes from product events at scale.
- Product signals can replace expensive “LLM-as-judge” for many evaluations.
Notable examples
- Healthcare flow: checking whether a doctor approved a prescription within 24 hours; Inngest would reduce months of work to days.
- PR review agent: wait for product events (e.g., PR merged without new commits) to score agent runs as good/bad.
- OpenAI thumbs-up/thumbs-down analogy: using user/product actions as success signals.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding Ingest: The Basics
0:06 to 1:58
Tony explains what Ingest is and its significance in providing durable execution.
“I think probably best place to start, what is Ingest for people who don't know?”
Challenges Without Ingest
1:58 to 3:56
Discussion on the challenges faced in building infrastructure without Ingest.
“Everyone uses classic Xambo is SQS and really, really complex to work with.”
The Evolution of Ingest
3:56 to 5:08
Tony shares the evolution of Ingest and its importance for AI applications.
“harness yourself, but it doesn't go as far as you need in order to give you the durability, the reliability, concurrency controls, constraints for each of your users.”
AI and Step Functions
5:08 to 6:58
Exploration of how step functions benefit AI processes and their importance.
“So the APIs and the concrete ideas from ingests around an event comes in and it runs a step function.”
AI and Step Functions
7:53 to 8:55
Exploration of how step functions benefit AI processes and their importance.
“It's the AI native private bank for business owners.”
Observability in Step Functions
8:56 to 11:12
Discussion on the importance of observability in AI step functions.
“A question here is like, what happens if you wanted to replay your 20-step function and change step 18 to see if that improved the outcome?”
Testing and Production in AI
11:12 to 14:00
Exploration of testing practices in AI development and the need for reliability.
“if the infrastructure just gave you the tools to do that.”
Ingest's Positioning and Flexibility
14:00 to 16:43
Learn how Ingest bridges application code and infrastructure for flexibility.
“Which is what we've been doing in engineering for forever.”
Evaluating AI Performance
16:43 to 18:33
Understand how to evaluate AI agents based on product signals and outcomes.
“Pretty much what people have been doing.”
Evaluating AI Performance
18:38 to 19:15
Understand how to evaluate AI agents based on product signals and outcomes.
“It connects agents to thousands of tools, handles permissions and LLM routing, and lets teams move faster without building it all themselves.”
Show all 34 chapters
Product Signals as Performance Indicators
19:15 to 21:55
Explore how product signals can indicate AI effectiveness and guide improvements.
“And somebody was like furious at OpenAI on Twitter, just like tearing into them.”
Building Better AI with Observability
21:55 to 24:48
Learn about the importance of observability in optimizing AI performance.
“And it gets better because then you have all the traces and agent trajectories.”
Evolution of Ingest and User Needs
24:48 to 28:00
Discover how Ingest evolved based on customer feedback and usage trends.
“You're not paying for like hell of RAM that you don't need because you're not putting an orchestrator in the sandbox for thousands of copies paying for like so much extra RAM.”
Building Applications with SDKs
28:00 to 29:46
Learn how using SDKs can simplify application development and reduce costs.
“Based off of SDK, no Terraform, just write some basic application code using your SDK and auto provision.”
Understanding Bare Metal Infrastructure
29:46 to 32:48
Explore the advantages of bare metal setups for AI and compute efficiency.
“But in order to do that, we're saving a bunch of state every time a step runs, only to delete all of that state when your function finishes.”
Future Trends in AI Infrastructure
32:48 to 36:46
Discover emerging trends in AI infrastructure and the role of NeoClouds.
“I think it's interesting and multiple people have pointed out something very similar which is everything is converging.”
Simplifying Coding with Agents
36:46 to 38:26
Understand how coding agents can streamline the software development process.
“and you're like, I want to fix this bug is, how would I do that?”
Challenges with AI and Infrastructure
38:26 to 40:20
Discuss the complexities and costs associated with deploying AI solutions.
“Can you expand on that a little bit more?”
The Future of Coding Agents
40:20 to 42:01
Explore potential advancements in coding agents and their impact on development.
“and everyone building crazy agents that like go out and paralyze on 10 different tasks and just YOLO merging stuff.”
The Future of AI Models and AGI
42:01 to 45:25
Explore the evolution of AI models and the concept of AGI.
“And it's going to completely change how fast, how often you can check a model's output.”
Startups vs. Big Companies in AI
45:26 to 48:22
Discussion on competition between startups and large tech companies in AI.
“And no matter how good the models get, I think the concept of one shot is just never enough.”
Startups vs. Big Companies in AI
48:26 to 48:45
Discussion on competition between startups and large tech companies in AI.
“When you're running a startup, you guys have raised, I think the public number is like 20, 25 million bucks.”
Startups vs. Big Companies in AI
48:48 to 49:26
Discussion on competition between startups and large tech companies in AI.
“Monaco's AI native platform replaces your legacy CRM and sales point solutions.”
Navigating Competition and Building Infrastructure
49:27 to 56:00
Insights on dealing with competition and building robust AI infrastructure.
“The first time it happened when Cloudflare copied us, like I won't go into the process of how they copied us or why the communication that we have with them.”
The Reluctance to Engage on Social Media
56:00 to 59:30
Explore the guest's perspective on social media usage and its impact.
“Like most people, like it's like, hey, I'm starting a dev tools product.”
Building a Product Without Twitter Presence
59:30 to 1:02:30
Learn how to grow a dev tools product without relying on social media.
“GTM strategy for DevTools is Twitter, you don't have it, what did you do?”
Content Strategy and Experimental Marketing
1:02:30 to 1:06:30
Understand the importance of diverse content strategies in marketing.
“There's so much good stuff that we've learned in the company, infrastructure-wise, that would be really cool to talk about.”
The Infrastructure Behind Ingest
1:06:30 to 1:10:00
Discover the unique infrastructure solutions that Ingest offers.
“I mean, the sad part is like a lot of people, a lot of people do AI.”
Building AI Without Worrying About Infrastructure
1:10:00 to 1:12:20
Learn how to build reliable AI products with minimal infrastructure concerns.
“that make my life easier and save me money.”
The Importance of Truth in Development
1:12:20 to 1:16:30
Understand why truth and adaptability are crucial in product development.
“But like we were like, oh shit, this is pretty cool.”
Navigating the AI Landscape and Industry Changes
1:16:30 to 1:20:40
Explore the challenges of adapting to rapid changes in AI technology.
“Well, even too, when you think of NVIDIA, Jensen didn't start it in 1993 saying like, LOLM, we need to build the infrastructure balance.”
Evaluating AI Performance: The Challenge of Evals
1:20:40 to 1:23:30
Delve into the complexities of evaluating AI performance and the costs involved.
“I think like there are some really, really good tools that if I were back 100 % engineering full time, I would love to use, you know, I would love to use more sort of like reinforcement loops.”
Evaluating AI Agent Performance
1:24:00 to 1:26:56
Explore the complexities and costs of using LLMs to evaluate AI agent outputs.
“But I also think it's absolutely craziness in that you're basically asking LLM, did you do the right thing?”
The Importance of Production Tracking
1:26:56 to 1:27:50
Discuss the significance of tracking AI agent trajectories for improvement.
“I think that's really hard for people right now in general, you know, and that's not to say that every eval company is wrong.”
Transcript
Automatic transcript. May contain errors.0:02Turner Novak:Tony, welcome to the show. Thanks for having me. Good to see you again. Good to see you too. I think probably best place to start, what is Ingest for people who don't know? Yeah, so Ingest is a small SDK that gives you durable execution on any platform. The TLDR is you build step functions and they automatically retry. They will succeed every single time your functions run. And we abstract all of the infrastructure there. So no queues, no state, no events. It just kind of works. So somebody who has never heard any of those words before, you just act like, why is that all important? Yeah, totally.
0:34So the TODR of this is when you're running something like agents or you're running some sort of process like order shipments, non-AI workloads, you might be interfacing with API providers that fail. You might have systems that fail because, for example, your database is down. And typically the way you get around this is building this crazy system of cues and events and spending a lot of time, aka months, on infrastructure. When you use something like Ingest, you write some really basic code that says, in this step, we are going to call Anthropic. And in this other step, we are going to call a tool that Anthropic's LLM wanted us to call.
1:12And if Anthropic fails because of rate limits, we will automatically retry that step exactly where it is in the function without losing any of the previous context because we stole all of its state. and you basically do zero work to get durability out of the box. And that's really cool because then any engineer who wants to create these step functions, whether that's an agent harness or if that's something like e-commerce orders or health insurance, you get all of this out of the box without spending any time on infrastructure. Do people spend a lot of time on it if they're not using Ingest?
1:45Turner Novak:Yeah, yeah, yeah. That's actually where Ingest comes from. I spent months rebuilding the same infrastructure on Kafka and then on Qs and it was like terrible. This was like in 2019 ish. Yeah, this is a while back. And like, you know, Qs still exist. Everyone uses classic Xambo is SQS and really, really complex to work with. Even harder with AI because you're going to be building really, really complex chains of logic for your agent harnesses. And if you're not doing that, then And when your agent harness fails, you're essentially trying to retry in memory and or your entire agent is going to just fail, which sucks.
2:27Turner Novak:So how do people get around this today? Like if I'm not using something like Ingest and I'm just like straight up, you know, shooting this all with my own code, building this all myself. How do you do it? You either don't and you live with the fact that an agent is going to get halfway through, die, and you're going to retry the entire thing, which is a ton of wasted effort time. or you're trying to string together this architecture yourself, building queues and so on. And that is going to take you a ton of time and be relatively inflexible. And that's kind of the antithesis of AI. With AI, you need to be able to update your code, your agents, your harnesses, your prompts, your models really quickly because models are going to change relatively frequently.
3:07You're going to learn a bunch as you run things in production and you're going to take that learning and put it back in your application code to continue improving your harness. so in general it's sort of got to the point in which you need to develop using some sort of harness and your harness is almost always going to be using Dribbble execution so that it can generate the traces retain state, retry correctly yeah
3:30Turner Novak:and is this not something that's just like built into the AI products already? Like it just kind of sounds like it's so table stakes as somebody who's not in the weeds everyday like you I'm just like oh this is like that should have been a solved problem because it's so obvious. It should have been a solved problem and it's so obvious. I think there's been some attempts at trying to fix this. For example, agent SDK is a thing, but it doesn't go as far as you need. It allows you to create some sort of harness yourself, but it doesn't go as far as you need in order to give you the durability, the reliability, concurrency controls, constraints for each of your users.
4:05And this is also relatively new in that we've been going for like four or so years. These are all relatively new things in the scope of infrastructure. So I think a lot of people have also been trying to experiment with what an agent SDK could look like, like an agent framework itself. And some of those learnings are that each particular LLM call is in a step, and that LLM will return some data, which will maybe ask you to do some tool calls, and you run this in a loop up until you hit some goal. TODR is that still needs steps, and that still needs durability.
4:39Turner Novak:And to your point you mentioned earlier, when you first started ingest like no one was using like this kind of capabilities like it wasn't we didn't need to do this so how did it kind of start and then how did it evolve over time yeah so um i'm going to go through like how ingest started um and then ai and then the evolution post ai um but real quickly i used to run engineering for a healthcare company which is super event driven because it's a good audit trail so what are some examples of stuff that you know when a patient filled out a treatment form or a doctor approved a treatment form or a patient asked for a prescription you have to do like a bunch of different things which is a series of steps and you also have to check whether or not specific things have happened in that flow like for example did a doctor approve a prescription within 24 hours and if not you've got a follow-up on that particular request that was literally one of the worst things to build and it took us months to build these really basic flows and it turns out that if we had ingest with its step functions it would have taken a day or two.
5:38So the APIs and the concrete ideas from ingests around an event comes in and it runs a step function. You can pause that step function up until something happens in your product and then it will automatically resume or timeout and it will automatically resume and you'll see that that particular thing didn't happen. All of those came from trying to build this in the first place. And it turns out that step functions are really, really good for AI because it is a literal loop of steps that we'll call an AI LLM and it will do something afterwards with that response over and over again until that goal has been achieved.
6:13And so step functions are a great way to build this. So we in some ways got lucky that AI needed our infrastructure, but then it turns out that our infrastructure is basically in the perfect place for running AI. And I'll quickly talk about what that means. When you're running AI and you're running these step functions, I think everyone knows this by now, you want to store all of those agent trajectories as traces so that you can see what your AI is doing, if it's doing the right thing, has it produced the right outcomes. And right now, a lot of people are doing this by storing traces separately from their infrastructure,
6:50Turner Novak:which is interesting. Is that a bad way to do it? Yeah, I think like from a first principles point. Absolutely. What's so bad about it? This episode is brought to you by Numeral. Numeral is the fastest, easiest way to stay compliant with US sales tax and global VAT. It's easy to set up and they automatically handle all registrations, ongoing filings, and their API provides sales tax rates wherever you need them with all the integrations you need. Their solution combines AI-driven automation with human expertise to manage global sales tax compliance end-to-end. Numeral supports over 3 ,000 customers, including companies like Brex and Character AI, and they pride themselves on white-glove, high-touch customer service.
7:31Turner Novak:Plus, they guarantee their work, and they'll cover the difference if they mess anything up. If you want to get compliant, check out Numeral at their new domain, numeral.com. That's N-U-M-E-R-A-L dot com for the end-to-end platform for sales tax and VAT compliance.
7:50Turner Novak:This episode is brought to you by Flex. It's the AI native private bank for business owners. I use Flex personally and I love it because I use AI to underwrite the cashflow of your business, giving you a real credit line. The best part is 60 days afloat, double the industry standard. Flex has all the features you'd expect from a modern financial platform like unlimited cards, expense management, bill pay that syncs with your credit line and their new consumer card, Flex Elite. Flex Elite is a brand new ramp-like experience for your personal life. A credit card with points, premium perks, concierge services, personal banking, cars and expense management for your family, net worth tracking across public and private assets, and a whole lot more fully integrated with your business spend.
8:33Turner Novak:One card for your businesses, one card for your personal life, one card for everything. To skip the waitlist, head to Flex.1 and use my code Turner to get an additional 100 ,000 points worth$1 ,000 after spending your first$10 ,000 with Flex Elite. that's flex.1 and code Turner for$1 ,000 on your first$10 ,000 of spend. Thank you, Flex. And now let's jump in. A question here is like, what happens if you wanted to replay your 20-step function and change step 18 to see if that improved the outcome? If you're just recording what happened in your step functions and you're not doing this at the execution layer, it's particularly hard to replay steps 1 to 18, change step 18, and then see if it improved the outcome.
9:17because you don't have that deterministic playback. But step functions give you deterministic playback, which is really good for building self-improving agents and checking whether or not changes to those agent trajectories actually work. And so it turns out that observability data is essentially a derivative of your infrastructure and your application code. And if your infrastructure slash step functions can create these traces, then they're coupled directly to how the infrastructure runs these steps and you get a ton more information, a ton more context, which allows you to build this self-improving loop way, way faster.
9:50Turner Novak:So you essentially need the observability to be built into the infrastructure, basically, in order for the self-improving system to self-improve. Exactly, it's far easier. And one like two, I suppose, examples here are, yeah, if you wanted to replay something and change step 15, much easier if you've got something deterministic to play back steps 1 to 15. and then you can see if changes to step 15 improve the outcome. Similarly, if you wanted to, for example, A-B test two individual harness changes without changing your infrastructure, that could be particularly difficult. But if you built this into the infrastructure such that you could A-B test different steps or different groups of steps, then everything is really, really simple and it suddenly becomes really easy for you to change prompts or models or harness changes to like 1 % of your users because the infrastructure is aware of what's happening.
10:46It can fork two different paths, two different steps or groups of steps. And then you can safely roll out, I don't know, two deltas between Anthropic and OpenAI to 50 % or 20 % of your users and check to see that the outcomes were the same, but that the token costs were cheaper. And so either way, you're going to be building this sort of stuff directly into your step functions and your flow of the harness. And it would be really nice if the infrastructure just gave you the tools to do that. So otherwise you're going to be rebuilding the same stuff over and over again.
11:19Turner Novak:So how would you do that if Ingest doesn't exist? Or I guess you guys do. You guys build it and then other people. But how do you actually make that possible? Or how would I do it if I wanted to build this myself? So this is a really interesting question. Right now, to be frank, when we speak to a bunch of our users and users that just push out AI into production that don't use us, people do local evals to check do the vibes seem good with the current AI changes that I've made with the current skill changes or prompt changes this is like the current industry standard yeah yeah yeah am I okay to push out these changes and then you're gonna I don't know, change your prompt locally change your harness maybe run local evals on 100 test cases see if that works if it looks good YOLO it into production because LLM as a judge is relatively expensive, you're probably not going to be checking 100 % of your agent production runs.
12:17I actually haven't met a single person that does that for 100 % of your production runs.
12:21Turner Novak:You basically just have like, here's the 100 most common things and we just check those. And if they're good, then... Here's my golden test case of a absolutely stochastic kind of random system. Let's hope that it looks good and push it to production for 100 % of users. And maybe I'll spot check using either LLM as a judge or human review, a sample of what happens in production and I hope that it works. Which is bananas. Like, absolutely crazy. So, isn't it probably kind of okay and we've been able to get away with it because it's like there's like these aren't necessarily life and death situations necessarily in a lot of cases so far.
13:03Turner Novak:In a lot of cases so far. In a lot of cases so far, if you're building a Vibe coding agent and you're doing something in which you're processing HTML, building a React app, You're like, oh, maybe this harness change is going to improve things because we're moving TypeScript 5 to TypeScript 6 or 7 was just released today. You know, let's move to that. Fine. But if you're doing something in which you're maybe we've got a bunch of users that use this, one of which is doing therapy and mental health using air, those circumstances are particularly important because you want to make sure that you're abiding by specific regulations.
13:34You also want to make sure that any changes to models or prompts
13:37Turner Novak:aren't inducing harm to patients. That sort of stuff is particularly important. And also generally speaking, if you're changing a prompt, there is some intended outcome that you're looking for. You've probably changed it because you've seen that the agent hasn't done the right thing in some particular use case. That's when you bring something back into your local models, but you know that this is just going to happen over and over and over again. So it'd be nice to be able to test outcomes predictably and make sure that the changes you're pushing to production actually do the right thing. Which is what we've been doing in engineering for forever.
14:08And we've sort of lost with AI, which is craziness.
14:11Turner Novak:You think we lost it? It was just because the tools weren't there to continue to do it? Yeah, I mean, I wouldn't necessarily know if we've lost it or we just don't know how to do it with AI right now. That's fair. And so how did it work then that Ingest was so well positioned to be able to do some of this? Because that's usually not the case. Totally, totally, yeah. when we started in Jest there was always this particularly interesting concept of us sandwiched in between your application code and the infrastructure. And by that I mean if you've got a step function that has five steps and you write this in TypeScript running on AWS you could rewrite your function in Python and move to GCP and your function will pick up where it left off even if it's halfway through.
15:03Because we retain the state, we call your function in GCP instead of an AWS, and it picks up where it left off with all the state from the prior three steps, and it will finish running in Python, and you get cloud migrations. And this means that we've distracted compute, and suddenly compute is completely fungible, including maybe the programming languages, which is really, really cool. And that's awesome, because then we create this deterministic environment for running your application code, which is also really good for AI, because it turns out that if you're trying to run an LLM and a step and an LLM and a step, which might be a tool call, that deterministic environment for running AI just automatically applies.
15:44So we're in this particularly good position because the fundamentals just enabled that. And from there, we can build really, really interesting primitives that do things like A-B testing variants. And then also using this tool that we have called step.wait4event to wait for things to happen in your product that prove whether or not AI did the right or wrong thing. An example is like you build a coding agent or a PR review agent, and there's a PR. It reviews some code, tells you that there's a P0. If you see that the PR was merged without any new commits being pushed, maybe that P0 was an incorrect flag.
16:19So you can listen to that webhook using step.wafer event, and then you can take that webhook information and then automatically score your agent's run as either good or bad based off of product events, which is really sick. no one can do that right now and that allows you to grade outcomes on 100 % of your production agents using the same primitives we've had for like four years and I think like thinking about things from first principles when we were originally building the system allows us to do a ton whether or not you're building with AI or it's just like regular infrastructure just kind of all fit together which was both lucky but also part of how we engineered
16:56Turner Novak:the system for flexibility so I mean it sounds like I mean since this is kind of like so important at this point should you not be building some of the stuff yourself because it helps you understand it better or like why would somebody build this themselves or use something like ingest to fix it yeah so firstly I think a lot of people build their own jank eval harnesses which is okay because evals really locally are just unit tests So that's pretty much what people have been doing. Pretty much what people have been doing. Yeah, it's craziness, yeah. I don't think many companies have the ability to check whether or not their AI or agents are performing well based off of product outcomes.
17:38How do you do that? You have to consume basically every signal from your product, and then you have to specify whether or not the signals indicate that AI is doing the right or wrong thing.
17:50Turner Novak:I mean, how do you even gauge that? Yeah, yeah. Like, what is being measured? When you can build anything, Amplitude lets you know how to build the right thing. Use human language to get complex answers about your products. No more manually selecting events or building charts or dashboards, just to ask. Use agents to sense changes in customer behavior, decide what's causing them, and ask you if it's okay to fix it, continuously in the background while you work. Get the answers you need while building directly in the tools you are already in, like Claude, Cursor, Lovable, and more. And for the first time, understand if your agents actually work, measure quality, debug failures, experiment, and measure their ROI with agent analytics.
18:28Turner Novak:Amplitude. With AI analytics, all you have to do is ask. This episode is brought to you by Merge, the connective infrastructure for production AI. The hardest part about building an agent is everything around it, connecting to the tools your team and customers rely on, letting agents take action with the right permissions, and keeping everything reliable and cost-efficient once you're in production. Merge handles that all for you. It connects agents to thousands of tools, handles permissions and LLM routing, and lets teams move faster without building it all themselves. OpenAI, Dropbox, and Ramp all use Merge to move faster and build AI right.
19:05Turner Novak:Visit merge.dev slash Turner to start building for free. That's merge.dev slash Turner to try Merge for free. So there was this interesting tweet from OpenAI where someone was complaining about OpenAI recording whether or not you copied and pasted from chat. Oh, interesting. Yeah, crazy. And somebody was like furious at OpenAI on Twitter, just like tearing into them. So how did that go? This PM replied saying like, well, dude, no one hits the thumbs up, thumbs down. And so we have to take signals from the product. Like, are you copying and pasting some of our responses? So then they know that it worked.
19:40To figure out whether or not. Exactly, yeah. Because if chat says something good and you copy and paste that, then the chances are that's a pretty good outcome. Yeah. And if you don't copy and paste anything, maybe it is, maybe it isn't, but it's very ambiguous. Yeah. So OpenAI themselves use product signals to indicate whether or not their chat has done the right or wrong thing. And if you're, say, for example, something like Xavier do something similar, if you enable a workflow that was generated by AI, then AI did a good thing.
20:06Turner Novak:Yeah. It's kind of like if you churn out of your session, it was a success. Yeah, in some ways, in some ways. And so depending on what your agent does in your own product, you can classify particular product signals as either good or bad markers. and then you can rate your agents using, honestly, like super cheap events. You're not paying the crazy LLM as a judge on every single agent trajectory. You're just saying like, there was a patient, patient came in, AI gave an answer through chat. Did the patient follow up with an appointment? Did they not? And based off of that particular outcome, maybe agent did a good thing or maybe it did a bad thing.
20:43Depending on your own product, you'll have signals that you can gather and then you can start automatically rating agents which then allows you to do more advanced things in the future. Like for example, A-B testing to see if the same outcomes were generated with cheaper models or open weights models. And that way you can safely roll out new open weight models in your products knowing that you have the same outcome and the same efficacy of your agents overall but that the cost is like way cheaper. And this is interesting because this is like where the world is moving.
21:10Turner Novak:Yeah. Like this is like a big thing on Twitter. I'm not the biggest Twitter user but it's still a big thing on Twitter. people were talking about how good GLM-5.2 is, which is this open-waste model. Everyone loves it. And if you're somebody building with agents right now and you're paying super high token costs to one of the frontier labs, maybe internally you think, I can take GLM-5.2 and swap it in for part of my agent loop, and that would reduce costs dramatically by like 10 to 20x. And a question would be, how do you track the outcomes to make sure that it's as effective? and using something like detecting whether or not your product is doing the right thing is one way of doing this at scale, cheaply, 100 % of your agent runs, which is cool.
21:55And it gets better because then you have all the traces and agent trajectories. You have product signals to know that you have known good trace trajectories and agent trajectories, which then you can take to post-training for open-weights models so that you create your own model that's relevant for your own sort of product and then swap that in using the same experimentation method. And then you get better models than you would from frontier labs that are generic. Really, really, really good with a ton of intelligence
22:27Turner Novak:but super expensive to something that's much smaller but suited for your use case and also way cheaper and more efficient to run. And so is there like a whole other, I don't know, like dashboard insights and there's like an observative observability layer of ingest that people are also getting and using. Yeah, for sure. And we had to do this for just step functions, driven execution anyway, because you need to see what steps are running. So it's like observing your own product performance. Exactly, yeah. So we give you the same sort of observability you'd expect for AI. What steps are running?
Read the full transcript
23:01Which LLMs did we call? What were the tokens? What was the TFP, the time to first buy? And how much did it cost? what were the total tokens consumed, input and output, and so on. We give you all that information. And we also tag things with metadata, group things by session in case you have many runs but one single chat session. And all of this combines to give you basically a complete overview as to what's happening with your agents using basic step functions. So it's kind of out of the box. Yeah.
23:35Turner Novak:And then how did the product evolve into this over time? Like what kind of pull were you getting from customers? Yeah. Yeah. So again, like generically, we started as like a basic, I'd say basic is really complex to build, but basically over specific pieces of infrastructure, queuing events. I mean, that's usually how infrastructure companies are initially started. It's like some like really boring, kind of like lame, not fancy thing. Yeah, yeah, exactly. Exactly. And so like we allowed you to build step functions in a much nicer way than any company had previously done. And I was like, cool, people liked it.
24:10Since then, a lot of people started using us for AI. And the AI observability piece made sense. Using ingest step functions, there's some really fancy things you can do that aren't possible. Like Golang has this defer keyword that will queue up a function to execute when your parent function finishes. And we basically run that to TypeScript. So you can use defer to run a function to evaluate and judge your agent runs once the agent finishes. so we saw like a bunch of different use cases and we figured we already give you the observability data and we already manage the process of calling LLMs and we already orchestrate your entire function therefore it would be super easy for us to build A-B testing so that you can check two specific models because we already do the orchestration and if we're doing A-B testing you need to be able to score particular variants to see which one won and you can do LLM as a judge and we can do that for you as well but that's expensive an easy way of checking whether or not something did the right thing is just did the product do the right thing which we can do using our previous primitives of step way for event so it all kind of came together and that customers just really needed this our users really needed this and it's really really really hard to build yourself we're seeing the same thing around like just general compute like literally every one of our users right now that uses something like sandboxes has orchestration and they either do orchestration out of the sandbox or in the sandbox which is an absolute pain and they also have to manage the sandbox life cycle themselves like starting stopping suspending it would be really sick if you could just do step function have something like group.sandbox and inside that particular sandbox if you ever said i want to sleep or wait for the specific thing to happen in my product the sandbox automatically suspended you didn't pay for any compute because you weren't running anything you get automatic durability and that orchestration lives outside of the sandbox.
26:03You're not paying for like hell of RAM that you don't need because you're not putting an orchestrator in the sandbox for thousands of copies paying for like so much extra RAM. You just have one copy of the orchestrator that sets outside managing many sandboxes, which can be much lighter weight. You get like a much better experience. The sandbox in this case is spin up an environment that then goes away when you don't need it anymore. Just running arbitrary code. So a bunch of our users do stuff like that, like Code Review. You want to git clone someone's PR. So you git clone their code, check out the PR, and then you have an agent that does a git diff, looks at the code changes, and analyzes the code base to see whether or not I did the right thing.
26:43But that's like an ephemeral environment. You want to do that in a VM that is completely isolated from other customers. You might be doing code review for thousands of customers,
26:52Turner Novak:and you don't want to mix people's code bases together. So you have a thousand different... In theory, you might need a thousand VMs, virtual machines, constantly on and constantly running, which would cost much more than only using it when that customer is using the... Totally, exactly that. And you're managing all of that yourself. You're managing the orchestration. And so, like, if you're building Drupal execution, you can just build a much nicer primitive for this overall, because you have the context of when your function starts and stops. You can start a sandbox using sort of, like, item potent APIs built in without you having to manage the sandbox lifecycle.
27:26you can suspend and resume. Like for example, if you did build a code review agent, create a sandbox, run some steps to review the code, use step.waitforEvent to wait for the review to come back, the PR either approved or rejected, or new code to be pushed. But when you do step.waitforEvent to wait for that feedback, the sandbox automatically pauses. You don't pay for any active CPU,
27:48Turner Novak:you don't pay for, you don't really need to do anything. There's like a ton of stuff that our users have asked for that we're essentially building and releasing because it just makes sense. So we're going both deeper on the infrastructure to give you compute, sandboxes, lambda, runtimes, VMs, that sort of stuff. Based off of SDK, no Terraform, just write some basic application code using your SDK and auto provision. Plus up the stack for observability so that you can see exactly what your application is doing. Get trajectories and then build better agents. And our view is that they're really combined because one is a derivative of the other.
28:23Turner Novak:aka observability is a derivative of what your product does and if you get them both for free and everything is really nice it sounds like it's you're helping them build better product but also decreasing the cost that it takes like maybe maybe there's faster speed in there too yeah if i'm interpreting this all right yeah basically basically like similar to similar to what basically everyone will say in every infrastructure company um i've used on things are that if you have an SDK that defines what your code needs to do, aka run step one, two, three in a loop, then that's far better than you provisioning Kafka, provisioning queues, provisioning servers, and then managing everything yourself.
29:04You write basically five lines of code. An agent, an AI can literally just
29:07Turner Novak:churn this out in one pass and you're ready to deploy on any infrastructure, good to go. And we run everywhere. We don't care where you host your code. You can run it on our compute soon, or you can run it on Railway or Render. It doesn't really matter. We don't care. You mentioned that you've got this bare metal bat. What does bare metal mean in this case for someone who doesn't know what that means? This is nerdy and very deep on the infrastructure level. Firstly, I'm going to talk about how we work, and then I'll talk about what we do and why, and why it wouldn't work on public clouds, and then I'll talk about what that means for the future of AI.
29:47the short story is you have a series of steps that run if you imagine you've got five LLM calls first to classify second to get some context third to get some information about whatever the user's put in any one of those steps can fail and so we have to save all of the information from step one all the information from step two so that if step three fails we can restart your function step three
30:12Turner Novak:and you're at exactly the same point. Super deterministic, yeah. Like 100 % determinism. But in order to do that, we're saving a bunch of state every time a step runs, only to delete all of that state when your function finishes. And so we're getting tons of data from you, which might be encrypted. We don't care what the data is, but still it's super bandwidth heavy. And AWS, GCP, and all the other big clouds are extremely expensive when it comes to both compute and bandwidth. and so for us it's kind of untenable to make this really cheap for our customers on public clouds because of the amount that they charge so the only way to build a really good durable execution company and step functions in general is to either sell it at a really high price because you've got to eat the public cloud cost or do it yourself on your own servers which is bare metal so we run out of multiple dcs we've got our own racks we do everything from the switches to the firewalls to the machines the only thing we don't do is the power and the connectivity to the internet and we we run everything ourself which is a lot of work for a smaller company but the costs are like 20 times cheaper which is insane and the performance is way better and that means that if you were to use for example our sandboxes or compute with us you would be running on bare metal extremely close to where your workloads run the queue the the execution the that function.
31:37And also because we run our own machines and we manage our own connectivity really super cheap. Like super cheap. Way cheaper than you get from anyone else, which is basically selling public clouds.
31:49Turner Novak:And so that just makes it so if I'm essentially using public clouds, there's a baseline of what I just can't go below that price because I need to still make money if I'm selling it to you. Exactly that. So a lot of folks that are like the NeoCloud I know Railway took an early bet on building on bare metal so they're one of the few not to but a lot of the other NeoClouds specifically just resold AWS for a long time and that means their costs have to be higher than AWS because otherwise they'd lose money so that's really, really tough so we took an opposite bet and that was our view since we started the company we had to build on public clouds to begin with because that's just how you get started but we quickly moved off of that to bare metal and that's been really, really good for us.
32:36Turner Novak:How do you think the next year or so is going to go in AI infrastructure? I don't really know what the best thing to ask your opinion on is because you probably have opinions on a lot of it but how do you think the next 6 to 12 months will look like? I think it's interesting and multiple people have pointed out something very similar which is everything is converging. Cloudflare and Vasell and Railway and all the other folks sort of look similar right now. And AWS released their own versions of step functions called Durable Functions, which looks very similar to ours recently, which is cool. Everyone is converging on the same sort of concepts because it turns out that when you're running code, step functions are a great way to do that.
33:20They give you the observability and it's a really, really good abstraction so that you get all of these trajectories. I think it turns out that like, as you've probably heard before, no one wants to mess around with infrastructure. Agents don't want to do that either. You don't really want to mess around with Terraform and have this really awful provisioning policy. And so things are moving much more lightweight, like SDK first. And infrastructure for AI sort of looks like this self-reinforcing loop in which you get step functions, agent trajectories, and then you take all of that data that runs your own production systems, do post-training on lighter models so that you can get your cost down.
33:56And then you have this inference layer that will run your own models, which is particularly interesting. I think a lot of AI infrastructure is basically moving to that direction of sort of inference hosting to be super lightweight. And NeoClouds are in this really good position to take over, which is really, really, really cool.
34:15Turner Novak:To capture a lot of that inference. To capture a lot of value. Yeah, to capture a lot of the inference spend and compute spend. And I think this is also the case with the proliferation of people that are becoming engineers. and I don't think many of them would like to deploy to EC2 by creating Terraform. I don't even know if many of them would know what a VPC is, whether or not you're using public or private IPs in your VLAN. And so NeoClouds are in this really good position to capture this new wave of developers and people don't want to think about that really. Yeah, because in theory you could say, shouldn't AI, couldn't you just say, hey, Claude, just do the backend for me?
34:52Turner Novak:Yeah, yeah, yeah. It'll do it. Is that not the case? It might do it. And you'll probably approve it. And you'll probably be like, cool, that looks good. Firstly, it does the right thing. Secondly, it introduces a bunch of complexity. And thirdly, it massively increases cost is a huge question. And in some ways, I think basically throughout engineering, we've learned that reducing complexity is good. And the delta between having all of that work and then just writing six lines of code and have it automatically work is huge. Because then there's far, far, far lower chance of mistakes if all of this is handled for you than there is if you're writing a Terraform policy that has EC2 and then you're attacking security groups and you're managing IAM.
35:36And then you hope that whatever model you're using knows how to run all that and deploy it all correctly. So I think generally speaking, yeah, models both get smarter, but you also want to focus specifically on your differentiator, which is running the product and building the product rather than managing said infrastructure, which kind of sucks.
35:58Turner Novak:Yeah, I guess there's like the how specialized is the thing that you're doing and like should you do it yourself versus if somebody, if like if it's a shared problem by everyone else, there's probably like a shared provider. And I think if you don't get any specific edge or differentiation, like as a company that's using one of these, like you should just, they'll fix the problem for you. Like don't waste your time. Yeah, yeah, yeah. It's already solved. Go figure out other things that no one else has solved. And one question that you maybe have as well is the future of AI. I know recently everyone talked about loops and software factories.
36:36That's what a step function is, right? You're just running loops. You guys have been doing loops for four years. Yeah, exactly. One question you have if you've got a loop or a software factory and you're like, I want to fix this bug is, how would I do that? Some interpretations are like, use Fable. use this great model and it will read the code and analyze everything and it will use extra high thinking and then you'll understand exactly what's happening and maybe you'll fix the bug like no joke sometimes it will just come up with a bunch of crap and it won't be the right issue um the right way to do things would be to take a look at the logs take a look at the errors trace that code path to fix the bug and then have a better understanding of exactly what went wrong couldn't ai do that in order to do that you need everything to be set up correctly right like you need to be taking logs from your ec2 service and from the applications and make sure you've got the right log train sorted and you've got to make sure that all of the steps to for all this observability was properly tracked or you could just use something that allows you to yolo six lines of code and you get all of that out the box so if there was a failure it was reported you see exactly what step failed you have an api endpoint for listing those errors and then you You can mark them as resolved really easily.
37:49Your agent would know how to do that because skills and MCP is a thing. And all it has to do is make one query to get all of the errors. And then it can tie that directly into the code context because it will see which function failed and why. And that process of fixing things in your self-reinforcing loop of software factory becomes super simple. And so I think there's also where the park is moving with agents and infrastructure, that you want things to be as simple as possible so you can get the right outcome in as few steps as possible with agents. Because the fewer steps you have, the less chance there is for failure overall.
38:26Can you expand on that a little bit more? Yeah, I think like, if you've got an issue in your code base and you're like, hey, there was this race condition
38:33Turner Novak:and these two things happened, you could have your coding agents using maybe Cloud Code or OpenCode or whatever you want. Go ham, read the code base, spin up a ton of sub-agents to read each particular package, hope to figure it out. This is kind of what people do, right? It's kind of what people do, yeah. That's kind of software factories overall. And that's super expensive. And you're hoping that the agent, your models, have enough knowledge to piece together things through the context that it generates to find the right issue. And you're hoping that that creates a process. It's basically just spin up to a ton of things, just get as much context as you possibly can, and then thread.
39:08Turner Novak:That'll get you the answer, because you've just you just got all this stuff. So you pay for super expensive models that have a bunch of knowledge and you hope that it threads the needle to get the answer, which is cool. And it usually does kind of work. Yeah, I mean, or sometimes. And all these new models are getting better and better. So that definitely works. But it's also like hell expensive, crazy expensive. And this is the whole thing about Fable right now. Everyone's like complaining that Fable's going to cost a bunch of money. It's token use. It's not included in my subscription. All that craziness.
39:35Turner Novak:I still think it's crazy that we went through this era of token maxing. Like, I mean, I get it. like learn it, see what it does, but also like, who's paying for all this? Like, I mean, that was just like crazy, those headlines. Like, I think it was Uber spent like a billion and a quarter or whatever. So I'm like, that's nuts. It's crazy. Absolutely insane that that happened. I mean, I, it's probably good for them. Like, I mean, that's a drop in the bucket for them, whatever. Like, they're, they're fine spending a billion. And then they probably learned a lot, whatever, however they want to spend this.
40:05Turner Novak:But I'm also like, how is that? That's just like classic, like, like top of market type behavior. Like money doesn't matter. Yeah. Just spend it as much as you can. It is crazy. It's crazy. And I think like the, the kind of over overview of like how agents work and everyone building crazy agents that like go out and paralyze on 10 different tasks and just YOLO merging stuff. And then you're paying for agents to review the agent's code is like, cool. Maybe we'll probably get there someday. Right now it's like super expensive to do that. And it works fairly often, which is good. I think it's kind of like playing the lottery you know and you get addicted to the win like oh my god this PR landed it shipped I didn't have to do anything and I just gave it a little bit of guidance and some PRs you're like oh my god this thing is an absolute dumpster fire which is terrible and so like yeah it's kind of like playing the lottery which people love you know like that sort of gambling aspect of will it do the right thing or will it not I feel like it's part of it though when you're like when you're using like clock you got that little ink blot that's like expanding it now so what's it going to do like how's it gonna work what am i gonna get out of this and then the first line will be like i just analyzed the information i'm doing this okay cool it's on the next step it's like what's the next output going to be and it's like now i just did this you're like oh yeah it's getting closer like what are we yeah exactly and then you're like it's kind of like that dopamine hit after a couple minutes of like you're just waiting you're like oh this is actually pretty good it's also like why i can't wait i was talking about this with somebody from nevius and like I don't know, the future of AI and the future of coding agents, the future of all these harnesses is super interesting.
41:40Especially if you consider diffusion models, which give you 1 ,000 token a second performance. It's like you've got this waiting and you're running all these agents and you've got this crazy loop and it's doing all this post-processing. And if you can just spew out the entire response in a second, really quick, or sub-second, that's craziness. And the efficiency you get from that is going to completely change how we build things. And it's going to completely change how fast, how often you can check a model's output. So I'm super interested in where things go, both with traditional LLMs, with diffusion models, with harnesses, with this engineering.
42:15But it all sort of revolves around the same thing, which is how can we track whether or not it's doing the right thing overall? How can we track outcomes in your product? I think this is also why coding models are so easy. You've got unit tests. Did it or did it not do the right thing? You're saying coding because it's just determined. It's very rules-based.
42:31Turner Novak:It's just like, yes, it's worth it. It's so easy. in comparison see whether or not agents are doing the right thing or your model is doing the right thing and it's much easier to train than it is doing something that's very hard to gauge um and so like i'm super interested in the future and about where things go all of that said i think overall you know harnesses loops all this sort of stuff will ultimately remain because the only option to not have these loops is to have a one-shot prompt that gives you this golden answer which is literally AGI and if you don't have AGI you're going to have to massage and control the LLM to generate thought, tools, context and so on and that's always going to be some sort of loop so Dribble execution, step functions and workflows must exist up until AGI hits and at that point all bets are off yeah I've kind of in terms of the whole AGI discussion I've always just been like, is it even a thing?
43:36Turner Novak:I don't know. Because with self-driving cars, I feel like it was like 10 years and everyone like, it's two years away. It was never here. And then all of a sudden, it's here. And it doesn't do 100 % of it, though. It's not 100%. It's not 100 % self-driving. In SF, all Waymo stopped because of the fireworks recently. Oh, did they? Yeah, the tons were being towed. It just completely failed. Wow, that's crazy. Which is super interesting. yeah super interesting so it's like the same thing like when is are we 50 years away from like true agi but then the other side it's like you kind of when i when i think back three years ago four years ago when this entered the discourse like you need to be able to raise money you need to be able to convince people to come work for you are you who you're going to use as your ai provider it's like stodgy old 25 year old company or this like you know we're building agi like you're going to use the A.
44:29Turner Novak:So it's like I get why it was done but I'm also like I kind of hate that all the like we've been 18 months away from AGI for a couple of years. We got Fable which is marginally better than Opus. It's pretty cool. It's good for coding. But like yeah I both agree and also think that you are still when you're building products going to need to provide context. You're going to need to give the model more information when it asks for it, you're going to need to build some sort of loop, some sort of harness over what the model needs. You're going to need code to interface with other parts of your system if the agent what it needs.
45:08You're going to need to track that so that you can guarantee it's doing the right thing, whether or not the model is extremely good or whether or not the model is like a year old from now.
45:17Turner Novak:You know, give me. Because you could make the argument to your point of right now people just kind of like give it to the model. It solves it. It's good enough. It costs a lot of money. but if we say if the models continue to get better you can just keep doing that and it will get cheaper and go down there's no point of using all this custom orchestration that you get with ingest because that could be another argument the models keep getting better but even if you're using something like there's the new OpenAI model coming out there's Fable you run it, use Claude Code locally and you can imagine Claude Code is basically a harness which it is and that harness is going to do something very similar to if you are building your own harness in your product.
45:59You're going to throw a prompt in, it's going to generate some answer, it's going to think a little bit more, it's going to call itself to get more context, it's going to maybe spawn a subagent which is another function core, another series of steps and it's more than likely going to request that you run some tools to get context about the code that you've written. Like for example I want to set some stuff, I want to read some files, I want to write a unit test to see if this particular bug exists and I want to run that unit test and I want to feed the results of that unit test back into my context.
46:28And no matter how good the models get, I think the concept of one shot is just never enough. One shot is never enough M &M. So even
46:39Turner Novak:when we have AGI, it's still going to need to be self-improving. Still need context. Because fundamentally, I think it's impossible to answer a question without more context. This is the same with humans as well. Yeah, so even if a PhD level acts around something, they're like, can you clarify the question a little bit more? They'll probably ask more questions than a dumber model, you know? PhD level is going to be like, I need to know so many specifics about the question you're asking to give you the right answer. And so you're going to have more and more loops and more and more context so they can narrow down on the right thing and link the relevant concepts.
47:13And that happens now. That happens with Fable. That happens with like, every time models get better, I've noticed, like I've noticed like it would just run more and more tools yeah
47:23Turner Novak:which is crazy well that's the thing too is like when I think of like three years ago four years ago you like used ChatGPT back in like you know December 2022 and you say it's just like it gives you an answer and it's like kind of wrong yeah and you're just like whatever but now today you know they'll ask like three qualifying questions and you're like oh shit I forgot to tell you that like yeah yeah that is helpful for me to give you yeah exactly like oh this is actually ambiguous and I was going to give you a completely misleading answer, but instead I had thought enough to clarify this. And so the whole thinking thing that they've done is pretty cool because thinking will direct the model in specific ways as well.
47:58And yeah, I just can't see that we can't live without this sort of harness-like system, which must exist if you're building with AI. And it's also particularly good if you're not building with AI because a bunch of people use this for deterministic workflows, just like e-commerce, insurance, all that sort of stuff as well.
48:15Turner Novak:You mentioned that AWS launched Step Functions. I think you've also had another massive public company sort of took your product as inspiration. So you have like, how does that typically go? When you're running a startup, you guys have raised, I think the public number is like 20, 25 million bucks. I don't know. You have like one millionth of the resources in AWS. What happens when a big public company, one of the biggest companies in the world, is like, oh, this is cool. we're going to do it too. This episode is brought to you by Monaco. Monaco is the first revenue engine built specifically for startups.
48:52Turner Novak:Monaco's AI native platform replaces your legacy CRM and sales point solutions. It has everything startups need all in one place. Monaco helps build your TAM, generate demand, run outbound, capture every interaction, manage pipeline, and automate follow-ups all in one tool. They pride themselves on an effortless onboarding, white glove activation, and get you to value in days, not months. And the product practically runs itself with built-in agents that are always working for you. Start growing your revenue faster with Monaco. Try it now at monaco.com. The difference between that happening now and the first time it happened is huge.
49:30The first time it happened when Cloudflare copied us, like I won't go into the process of how they copied us or why the communication that we have with them. But the first time they copied us, I was like pretty annoyed, you know, personally. because I'd spent a lot of time building and developing the SDKs with the team and the team have done a great job on what we've built and defining the DX. And then seeing somebody just like rip it off was like really annoying, especially somebody that we work closely with. And that happened a few times. So on a personal level, if you're a founder, it can be a little frustrating, but it's also like really, really good.
50:07Now we're really happy when AWS copied our stuff, used it. I won't say copied. When they use similar DX approaches, I think that's a fair way to say. It actually made us really happy. And I was talking to somebody from AWS last week at the AI e-conference.
50:20Turner Novak:They were like, oh, like durable functions. And I was like, yes, which is very similar to our API that we made four years ago. And they were like, I had no idea that we did that. Like, it's just kind of funny. It's nice, it's validation that we've done the right thing. And it's validation that this DX is really, really nice for engineers. At the end of the day, like if you're a founder, getting caught up on what your competition does is maybe worst case. I think like we'd mentioned it before and you were like, how was Ingest in a good position to run AI to begin with? And it turns out that because we'd thought about the engineering from first principles and the way that the SDK worked and the way that we work with particular systems and that you can rewrite things from Go to TypeScript and it'll pick up where it left off.
50:59That's what first principle thinking allowed us to build a really flexible platform that continued into AI. And if you just think about what your competitors are doing and try and copy your competitors or poke holes in it, maybe you don't have the same first thinking principles and maybe you're just like
51:16Turner Novak:not doing the right thing for your own product really um so i don't care what our competitors do i don't care what other people in the ecosystem do i don't even look anymore um i don't care i just think about what our product can do what are the use cases for it what are our customers doing and what would make that process better and if other people copy us like cool that's actually great. They don't know how hard it is to build these super high read, high write, high delete systems at scale. They don't know how hard it is to build the infrastructure to make all of this work. They don't know all of the pitfalls that we've run into.
51:48They don't know the future of what we're doing, which I'll honestly just broadcast in podcasts like this, aka traces, agent steering, trajectories, sessions, taking all that information so that you can do post training on open waste models so that you can take your production data, generate your own custom models and then run them so that you can free yourself of the dependency of frontier labs and run the same sort of ai quality for like one twentieth of the cost like all of that stuff might not be in their roadmap and if they listen to this and it is like cool good doesn't matter i really care because we have a job to do and we'll get it done um so this is quite good actually i quite like it so then like what's so hard about doing all those things because couldn't i just
52:30Turner Novak:take the transcript from this episode and throw it in the cloud code and be like, build this, no mistakes. Especially the models are getting better. Yeah, I'm pretty sure that somebody has. I'm pretty sure that somebody out there has a system that's like ingest for AI, build AI specifically and also don't care. Also don't care. You might be able to do that for your own personal stuff. Maybe, honestly, maybe you can have a personal version of ingest that you spin up and you have an agent work on and then you work on your products on the side the question would be do you want to do both probably not because if you're working on the infrastructure you're not working on your product and that means other people that just build the product are going to do a better job than you so then you've got opportunity costs and you're going to lose when you're building your products and secondly if you're doing that to build infrastructure then i don't i i don't think it will scale the challenges are so extreme when you're building at the scale of like millions of events per second doing so much qps on your data stores um that just still get it'll be tough so is the issue that you i could technically vibe code this but then there'd be a lot of issues i'd have to probably like fix manually and i'd have to go in and really repair like fix figure out how to how to serve these millions of events per second to the point of like, it can't just be like a vibe code thing.
53:58Turner Novak:You need like a big team that's constantly Yeah, basically. I think you've got two paths. You vibe code this for your own product and that means you're scratching the surface of what we give you for free. And you're not going to get the same self-reinforcing loop. Oh, so I can use you for free. Yeah, exactly. In that case, you could say like, why would you vibe code it? Yeah, exactly. Why would you vibe code it? Anyway, just use this for free. Secondly, like you're not going to get the entire self-reinforcing loop to the same degree that we would give you for free. That's just going to be really hard for you to do.
54:28And if you do decide to build your own infrastructure to compete with us, that's super cool. Just very difficult.
54:35Turner Novak:I think you still grew, I think the number, you've grown like 35x, maybe it's more than that. How much have you grown since the very first time you got copied? Thousands. Yeah, and that's great. At first I was worried because I was like, Like, oh, this is the end of the world for us. You know, this guy is falling and so on. Well, it's interesting. One of the biggest companies in the world is doing the same thing as you. There's a pretty high bar for them to do something. So if they do it, it means that there's a really big opportunity stack ranked along all the other things that they're doing.
55:12Turner Novak:Like this was worth pursuing. So it almost like validates the market is there. It does. Yeah, it does. And in order to compete with us to some degree, they can do the basics, but it's going to be extremely hard for them to go deep on all of the things that we do. For example, as of yet, no one else does A-B testing of different steps within your step function so that you can safely roll out changes. Nobody else does tracking the outcomes of those particular changes. Nobody else has things like deferring so that you can run cleanups or do sagas or score based off of outcomes of that particular agent run.
55:48People are still relatively far behind,
55:50Turner Novak:but it's like total validation that we're doing the right thing, which is cool. I think one thing that's super interesting about like Tony, the person, especially being a dev tool, like infrastructure founder, you just don't have a personal brand. Like most people, like it's like, hey, I'm starting a dev tools product. Like I need to start building up my Twitter persona or something to like start getting customers. So why have you not done it? Because it's like the most common, like you should, you should do it. Like you should be doing this. You and other investors of mine told me that I should do this every day.
56:21Turner Novak:I mean, I'm like, whatever you want to do, whatever your strength is that's what you should do but I feel like probably every day people are like oh all the time yeah and when I tell people that I don't use social media really their initial reaction is that's insane that's crazy why you should do this yeah so I am I have pretty strong views on social media from before I started my company that it just wasn't good for people as in like a the addiction the health addiction the health comparison that sort of stuff I much prefer talking to my friends to catch up with what they've done than look at their IG stories overall back when I had IG and I guess there's still technically a profile that I've logged myself out of two-factor that I would like Facebook to delete so if you're listening please delete it oh like they're trying to force you to open your Instagram to log into things I could technically ask them to recover my two-factor but besides that I much prefer talking to people and catching up and I just viewed social media as as sort of generally bad for people's mental health.
57:26And I didn't like it. And I also viewed things as like sort of, there's some really, really good social media. There's some really, really good things you can learn. And that's really cool. I appreciate that from a bunch of people, but because of my previous views, I just didn't really want to do it. And also like I much prefer learning, building, doing than I do talking about that sort of stuff. So like we have this entire FoundationDB backend, which is super sick. And it's actually up to the engineers to talk about and we'll probably talk about this in detail about how we built this FoundationDB backend.
57:59And I much prefer talking about that rather than how we did it and why. But I know that that's important and this is what other people in the team are for. So I'm really grateful that they do it. I just spend all of my effort on what to build, why, how, where we're going and making sure that we do the right thing. And my co-founder Dan is also way better a tweet and some banger tweets than I am.
58:21Turner Novak:Yeah, I was going to say, there's probably three other people on the team I can think of off the top of my head that have a bigger social persona than you. For sure, yeah. Then I sort of know that it's partly required. You take a look at G from Vassil, and he just lives on Twitter. Craziness. And that's cool. Then you take a look at Ali from Daily Bricks. I don't even know if he's got Twitter. Maybe he does. He probably tweets a little bit, but I haven't seen him. I've seen some tweets, yeah. You've seen some tweets? Maybe. Not a lot. Yeah, not a lot. They're probably written by someone else. Yeah, yeah, yeah.
58:55Exactly. Exactly. And so, like, there's a Dicost me. We probably need to do it. My investors will ask me to do more. I will probably give my Twitter handle to some folks in the team so that they can do stuff on my behalf, which sucks, but probably. Or they will eventually convince me and I will have time to do it. But as it stands, like, there's so much to do. and we already have enough growth that that sort of distribution of a personal founder level just wasn't top of mind. It was doing the thing. I much, much, much preferred doing the thing.
59:30Turner Novak:So if the primary GTM strategy for DevTools is Twitter, you don't have it, what did you do? How did you approach getting people to use the product if you weren't doing that? Yeah. Oh, man. Wow. So way back in the start, we noticed this missing piece, aka queues, could not run on serverless. You couldn't use Kafka in a serverless function. The only way would be to create this jerry-rigged SQS, SNS, Lambda-type deal, which is extremely ugly, very hard to build, zero local testing. It's like pure acronyms, the thing that you just described, which tells me it's like an insane setup. It's like an incantation, you know, like you're casting some spell, but you're doing it with like crazy code and terrible.
1:00:18It's just like awful. It's the worst thing. And I remember like people trying to do that and people had this like crazy architecture diagrams because you're an AWS consultant. And I was like, these people are fucking insane. This is nuts.
1:00:29Turner Novak:So it's AWS consultants coming to people like explaining how to do this. Basically, yeah. You need to know the infrastructure very deeply in order to tie these specific things together. And it's just like the worst thing. and so that realization was like well there's a wedge here we can build some sort of event driven queue some sort of step function and we can make it run on serverless and that would be sick because then we give it to everybody and serverless is like both really good also the lowest common denominator of running functions you know like stateless kind of crappy short lit functions but super good at scaling and zero DevOps and that's cool because there's this unmet need of people trying to ship stuff fast sometimes YOLO out in API endpoints and we give them durability out of the box.
1:01:13Meeting that unmet need was good because then organically we would talk about it. We'd work with different frameworks that people could write their code in. We'd work with different cloud platforms that people could deploy to. Can you discover it through the marketplace on AWS? Just the zero integration, zero marketplace. People would just find out about it. We'd maybe do some guerrilla marketing and we'd like talk about it on Reddit or Twitter or stuff like that, like my co-founder would. But we would just like grow, which is cool. And then yeah, partnerships, marketplaces are good. Eventually like the idea was just so obvious that people realize this is something they need to do.
1:01:53And then we do lots of blogs, lots of content. I think like one underrated thing in every GCM strategy or not even underrated, everyone knows, just content, right content. Write content for blogs, write content for SEO, write content for AI to scrape up, like just content, all the things.
1:02:11Turner Novak:So you have to do all of it. It's not just AI, SEO. It's like... Yeah, yeah. Blogs are so, so good. Just write more blogs. Everyone should write a blog. Because there's so much useful information and there's so much information about why you have solved a particular problem a specific way, which is really interesting. And so we did a ton of that and we still do. We have so much to write. There's so much good stuff that we've learned in the company, infrastructure-wise, that would be really cool to talk about. So how do you decide what's worth talking about? If I'm listening to the founder, how do I come up with ideas?
1:02:44Because a lot of people, they don't know what to say. Yeah, talk about everything. Literally everything and then see what works. That's the best thing. Talk about everything from an infrastructure point of view to how you built it, to what you can do, to case studies, to examples. Just talk about everything and see what sticks. That's the only way, really. Because your audience, every audience is different. You could build an exact replica of our company but maybe you target a different demographic and maybe that means you can't talk about infrastructure maybe that means you need to talk about really really really basic things because they're not necessarily the most technical people and that's okay so I just talk about everything
1:03:18Turner Novak:I know one book it's called Traction that you got a lot out of so what did you learn from that? The DDG founder, I can't remember his name wrote this book called Traction it's basically an experimentation framework where it talks about different GCM strategies, which is very similar. Just write about everything and see what sticks. Do a bunch of different GCM strategies like talk, Twitter, socials, blog posts, webinars. Pick a few. Test them over the course of a few weeks. See if they work. See if they resonate. And if they do, keep on doing it. And if they don't, drop it temporarily. Don't drop it forever.
1:03:53Drop it temporarily. Come back to it at some point in the future.
1:03:55Turner Novak:So it sounds like running loops of your strategy. Loops for marketing. Yeah. Loops for everything. So yeah. My whole Twitter stuff is like, I mean, Twitter's good. It's noisy, but it's good. So I can understand why individual dev tools founders would prefer a presence on there. Totally get it. Yeah, and you do get pretty quick validation of like, let's say you get a million views on your thing and like a hundred replies of people discussing. Like, it's pretty hard to argue that that didn't really quickly have a tangible value provided to your product or your brand. It's a dare. Versus if you do a webinar, someone might convert two months later.
1:04:38Turner Novak:Versus on Twitter, you see it. Spam all the stuff as fast as possible. Yeah, totally. Also, there's the whole 1 % contributes to 99 % look at type deal. You get 100 people looking, but there's thousands of people that have read that that are thinking something, either good or bad, about your product and what you've said. and so that's really really good so I can totally see the value I can totally see why people would like to do that maybe maybe I do tweet maybe you should maybe I should get on there in Croft I mean because the interesting thing is you probably there's probably a lot of things that you write in Slack to the team or email like you explain something, some new framework and you make a blog post and you put it on the blog and then it just sits on the blog versus you could literally copy and paste the chunk of it, put it in a tweet and just tweet it.
1:05:30Turner Novak:It literally took you an extra minute maybe versus you already put an hour, maybe five hours or whatever to writing this. I don't know. That's how I tell people is a good way to just get started if you're kind of like, I don't have time to do this. It's like, well, you had time to write this two hours of thing that you made. It doesn't need that much more to get it out there. And then it's like, what's the upside? What's the downside? Like, downside is just, like, no one read it, whatever. Like, you didn't spend that much time on it. You already made the thing. And then the upside is, I don't know, like, you know, Jeff Bezos, because he uses Twitter, comes across it, reads it, and he's like, oh, Ingest, this is pretty cool.
1:06:09Turner Novak:Like, we'll use it for Prometheus, our new AI thing that we just raised, you know,$12 billion for. And, like, you know, you're a board-level vendor for this, for Jeff Bezos' company. And, you know, they're going to pay you, you know,$100 million to be a customer. Like that's pretty good outcome from literally copy and pasting a Slack message and putting it on Twitter. You know what? Maybe I'll just get an agent to write my tweets for me in a loop. That would be it. Just create some bang of tweets. I mean, the sad part is like a lot of people, a lot of people do AI. It's just a fucking slop, man.
1:06:37It's slop. It's slop everywhere. Yeah.
1:06:39Turner Novak:And I think like then you just need to make sure is you're using it for like idea generation, but not necessarily like the whole thing. Yeah. Is AI. There was this really, really, really good post from somebody at Microsoft, maybe the director of AI at Microsoft. Also, I can't remember their name. He was talking about self-reinforcing agents, learning feedback loops, and so on. Really, really, really good tweet. And it contained a lot of the same principles that we built Ingest on, which is observability, determinism, reinforcement, based off of the production trajectories, using that particular framework to run agents.
1:07:15Super good. Everyone loved it. It was like a huge tweet. and I copied and pasted it
1:07:21Turner Novak:well the link to the tweet in Slack because I was like this is exactly what we've already said we were working on this is complete validation of our entire system and so like we should we should just do the same thing we should do the same thing and that happens a lot I feel like where you know you I mean because that was probably like some kind of memo that was written internally or something and then they posted it publicly and I mean if it's good like people will share it and use it so it's like you're almost like really kneecapping yourself by not sharing. I mean, you want to be careful. Like, here's our roadmap.
1:07:55Turner Novak:Here's like, maybe you don't give a shit. Like, maybe it doesn't matter. I don't care. But like, there's definitely like a, there's definitely like a, you don't want to share everything, but like there's a lot you probably could. Yeah, for sure. For sure. And again, like we mentioned previously that all roads are converging. Many infrastructure companies are looking very similar, you know? And because of that convergence, the roadmap is fairly easy to predict. You know, like computers are thing. running ingest functions is a thing because it's bananas that right now you have to choose to host ingest functions on some other provider.
1:08:27We'll just let you do that for cheaper than other providers because we already run the bare metal. We can do better sandbox DX because it's integrated into the durable execution framework inside your already existing harness. So the roadmap is pretty easy. It's pretty easy for people to guess where we're going. I think it's pretty difficult for people to understand the nuance of what we do. For example, I think it would have been pretty difficult for people to have guessed that we were going to release something that allows you to score 100 % of your production agent runs using product events.
1:08:59That's something that's so special and unique to us that nobody else could have thought to do that anyway. We're very unique in that only we can give that to everyone. You're right. I don't know if people know. we're all working on very similar problems at this point and a lot of the stuff that we've done right now that we're working on is also ground that's been trekked on before whether or not you use Firecracker or use Cloud Hypervisor to provision your VMs it's all the same stuff that people have been doing for 10 years and a lot of the infrastructure that already exists has been paved by a bunch of other prior companies and clouds doing the same stuff in open source so it doesn't matter everyone right now is doing micro VMs because micro VMs for sandbox is like the new harness.
1:09:46So it's all the same.
1:09:48Turner Novak:So it sounds like if I'm a customer trying to decide who I should go with, the reason I would consider ingest or the reason that I should go with you guys is if I appreciate that you will launch new features that make my life easier and save me money. Our entire thing is allow you to safely build reliable AI or good products without worrying about infrastructure. And that could be without worrying about compute, without worrying about queues, events, observability, tracking that everything worked correctly, and making sure that your product does the right thing. And you can do that using a few lines of code.
1:10:25Everything else is completely abstracted. That means you can focus on what specifically you need to do instead of building anything else. Zero infrastructure required. That's our entire anti-infor thing.
1:10:36Turner Novak:Just don't worry about the infrastructure at all. What are some of the values you guys have as a company? Yeah. firstly uh truth i think like have you have you read principles by ray dalio oh no i've not read it oh cool really dry really dry book really good um so like both good and bad at the same time and he talks about a few different things and he's got a few principles and one of them is truth and i really really agree with this one so why truth uh if you're building the wrong thing and you don't respect the truth and that people are telling you that it's wrong and they're not using it correctly, then you're going to be misguided and you're going to continue down the wrong path.
1:11:15If you built the right thing, but the world changes around you with AI and you don't understand and appreciate the truth that the world has changed around you and you continue down that path, you're inevitably going to be building the wrong thing because the truth of the situation is that a lot of stuff has changed and you need to change with the times. And so if you can understand the truth, are we building the right thing? Are we on the right path? do our users appreciate what we are doing or is it actually wrong then we can basically get to the to the as close to the right answer as possible and it also keeps us from like it keeps us in some ways like free of ego from thinking like oh we had dribble execution and we built this step function sdk and the apis that everybody else copied and we must we we this must be correct because everyone else has covered us like if we just agree that finding the right thing finding truth is the right thing to do then like turns out in five years time that's wrong then then
1:12:13Turner Novak:no one is too attached in the way that we've built things we're really okay to change everything that we've ever done because we've learned more information um so truth is like super important um context uh openness is also super super important important for us as well like context nuance that sort of stuff of what we're doing but yeah i think like truth is such a fundamental thing that if you avoid something that's not true it inevitably comes back to to screw you over has that happened to you before um probably accidentally never deliberately we've never avoided something because it was true by thinking like oh we know best um i can't necessarily think of anything off the top of my head I was going to try and make something up around like AI.
1:13:01Turner Novak:But like we were like, oh shit, this is pretty cool. We should just adjust to do this. Did you, did it take you like a couple weeks? It probably took us too long to really adapt to the whole AI thing. So it should mean like the day ChatGPT came out, you should have like instantly gone. Yeah, that would be great. So why don't you think you did? Like why did you not instantly change everything? I feel like sometimes it's hard to know what's the fad or what's real, you know? Like sometimes it's very hard to predict the future. and so you're forced to make bets when you run a company like you're forced to make bets on the complete direction of the company are we going to go all in on AI or are we not are we going to go semi in on AI and you could argue that like you know one is better than the other and that's very true one is always better than the other but you never really know and so we had a bunch of people start using us for AI and it turns out that was really really good and so we became much more AI first and saw the e-mail stuff that we just released but at the same time being general purpose infrastructure.
1:13:57There's this big component of ours that you can use this to build whatever product you want. And that dichotomy is hard to turn the line on. But foundationally, I think we
1:14:10Turner Novak:probably could have done it faster. I feel like we were just coming off literally a month prior FTX had failed. And we had gone through this whole Web3 era. I don't know. Every podcast you listen to, it's going to like mint an NFT that like other people might want to trade your your podcast you know listen to NFTs I forgot about NFTs the most ridiculous thing yeah it was crazy and like it was all I was talking about like NFTs are the future of like the economy and you're just like what so then like this new thing comes up and you're like I don't know I'm like whatever you guys just you guys just told me that like the whole world was going to be on the blockchain and that was like wrong yeah yeah and then now like there's this new thing I don't know it's is like you guys are just crazy yeah i feel like personally i probably got a little bit like thrown off by that initially like where but i feel like it's sort of how the industry works where like you go all in on things um you see if it works and if not like yeah which is like in a bad like when you think about like the reason like a startup can be successful is like a startup raises 20 million dollars like all of that capital is basically r &d they're like a tax problem when you think of like there might be a competitor you're competing with that has a billion dollars in revenue or something like that but then what do they actually spend on true r &d like like deep tech quote-unquote like very much like high risk you know capital is being work being put to play it might be zero like they might not actually be doing any r &d yeah it's like as a startup your point the reason you exist is to like go after this like pretty highly risky probably potentially won't work like you're going on an adventure like it's like adventure capital it's like adventure and like it's like a truly crazy thing that we're trying to like fix and solve this problem yeah and so I don't know it's like it kind of should be a little bit crazy it should be a little crazy yeah totally I think like also the infrastructure that we built just didn't exist before we built it you know and whether or not it was 4AI or not it's foundationally the same infrastructure we never really pivoted we never really did anything different it was always that infrastructure but it just straight up didn't really exist in the way that we built it which is cool and it turns out that that's actually really, really effective and will help people build this new wave of products that must exist with AI.
1:16:29And so, yeah, really, really, really fortunate that that's the case. But yeah, everything's a gamble.
1:16:36Turner Novak:Well, even too, when you think of NVIDIA, Jensen didn't start it in 1993 saying like, LOLM, we need to build the infrastructure balance. I don't know if it was barely a word. Artificial intelligence was like science fiction. And it was basically like gaming, graphics, cards or whatever. Cards, numbers, matrix, multiplications. Yeah. And even then, when you think of back in like 2021, I mean, it's basically like a crypto company. Majority of NVIDIA's revenue was like Bitcoin mining. Yeah, yeah, yeah. Cuda, man. Cuda, the best bet that NVIDIA ever made. Cuda is the best. Do you have a favorite like CEO or founder or company just like throughout history, whether it's like who's still around or like currently operating?
1:17:17Turner Novak:Like who do you, do you have anyone that you've gotten a lot of inspiration from? I guess the reason I brought up Ali from Databricks is I really rate Ali from Databricks overall. I think Ali from Databricks is, I don't know why I call him Ali from Databricks. I could just call him Ali from now on. I think that Ali from Databricks is great. He's great. He's a really, really good person. He's really, really ruthless, but good at what he does. If you take a look at the team that's got around him, everyone has dicked for such a long time, which must speak to a lot of his work and the way that he operates his company.
1:17:47Don't they have seven co-founders? They have a lot of co-founders, right? It's an interesting company for sure. It's a lot of academics.
1:17:54Turner Novak:It's all academic co-founders, which you don't see as much. Yeah, and I think that Ali is, in general, their company is interesting and cool. So, they've done a really, really, really good job. And also, interestingly, they're not so big on Twitter like you were saying. They've done a lot of really, really, really good things. And they built this really interesting technology that's pretty cool. And the way they operate is quite good as a company. So I really respect that. I really respect that. I really like the way that Cloudflare was built. Copying and that entire thing aside, it's an interesting way to attack the problem and then basically get all this leverage by running so much bandwidth through all of your pops that you can do a ton.
1:18:41So I think that's particularly interesting from like just a company building perspective which is cool. Personal favorite, like I also love the folks from both Planetscale and Railway. You know, Sam is great, Jake is great. They're both really, really, really good people as well. So huge fans of them personally as people as well as their companies. Yeah.
1:19:03Turner Novak:Your favorite like new AI tool they use or like what does your like stack look like I guess? Oh man. Alright, so coding I was the biggest holdout in the company I was just like writing manual code myself in NeoVim and I still use NeoVim for everything so you don't use AI to code? I've started using a little bit of AI I've started using a little bit of Codex and Clawed but I manually review every single trunk that comes out of them like a crazy person why do you do that? every time I look at it I'm like something was wrong we can simplify this
1:19:3640-60 % of the time I'm like let's change this and let's improve the way that this particular thing works. Sometimes it's pretty good. Sometimes I don't care about the problem,
1:19:45Turner Novak:like front end, for some internal tool. And I'm like, but if it's in our code base, I'm like reviewing each chunk manually. You got to be training the code base on your changes. Training the AI on what you changed. Yeah, exactly. Other than that, I think I've just started to use voice-to-text, especially for PRDs. especially for communication about what we do and our vision and why things need to exist. Context around the company, context around what we're building and why. And I YOLO voice to text in Notion like a crazy man. It's the best. Do you get it to where you'll just talk for 10 minutes and it will just succinctly rephrase what you said?
1:20:29Turner Novak:I do paragraph by paragraph in text. And I have a hockey combo that will start. I'll speak maybe like eight sentences and I'll stop and then I'll do a quick check to make sure it look good make sure it reads okay and then I'll continue so that's actually pretty good so you use notions built-in AI no no I use something on the Mac I use local models on the Mac for that I think Nvidia has this Parakeet model which is pretty cool so I use that with some some local stuff to make it work it's pretty good and that's not because of our new principle I just was mucking around and I was like, this is cool.
1:21:08So it just stuck, you know. I think like there are some really, really good tools that if I were back 100 % engineering full time, I would love to use, you know, I would love to use more sort of like reinforcement loops. I would love to use more feedback from errors or stack traces or production stuff directly into the code base. And I'd like to have this entire setup that would be great to use with multiple agents that I know some people in the company have, but I'm unable to do that day-to-day, unfortunately.
1:21:40Turner Novak:Is that because you're just on calls with customers, recruiting people? Yeah, calls with customers, talking to people, talking to marketing, talking to sales, talking to product, taking a look at the bare metal builds that we have and infrastructure stuff. Just like day-to-day, so much changes. So I wish I were more knee-deep in stuff, you know, because the world has changed so much. and I look at it and I'm like, this is great. How do you stay on top of it then if you're just knowing what direction to go with how fast things are moving? Yeah, so even if I'm not building the product day-to-day, I'm still pretty close to what happens.
1:22:16Well, I'm super close to what happens in our products, what we need to build, why. I'm really close to our users' feedback that we get. I'm really close to what our users are doing and why they're doing it in specific ways. Also, what other things they're doing and why. And also pretty plainly thinking about the problems we have in the company and how we can improve overall as a company and where we need to be. And also just looking at the trajectory of things in the past year and what that implies for the next year so that we can continue to build for where things will be in six months and 12 months' time.
1:22:49Otherwise, we presumably might be doing something wrong. So all of this means that often the work that I'm doing isn't like day-to-day engineering, even though I do that sometimes. it's a lot of that context to propel us in the right direction. And that's like, honestly, more frequent than ever because so much changes so consistently with AI. But I think like it would be much slower and the world was moving at a much slower pace even five years ago than it is now. In the tech world, at least. And so you need to really keep on top of things and think about where you're going and how you adapt to the world, which comes back to that truth principle that we have.
1:23:31Turner Novak:in terms of like building observability into the product, there's this thing called evals. We hit on a little bit. Can you maybe really quick give us a slightly more in-depth explanation of kind of how that works and why you think it's kind of crazy the way people do it? Yeah. Okay, cool. This is maybe a hot take. I think the way that we do evals is absolutely batshit insane. Not in that it's terrible. I think it's like a pretty good first step. Yeah. But I also think it's absolutely craziness in that you're basically asking LLM, did you do the right thing? You're asking the criminal, did they commit the crime?
1:24:09You know? And that's also super expensive with the context you need to pass in. If you've got this agent trajectory that's taken like 10 sub-agents and 100 steps and you pass in all that context to say, did it get the right answer?
1:24:20Turner Novak:It's like hella expensive. And so agent evals are essentially unit tests over input. Did it give me the right output? And that's cool. You can do that both in code programmatically. You can say, like, I expect the output to be, you know, 95 cents given this particular input. And you can do that using LLM as a judge, which is the insane part. But also I can totally understand why it's necessary. All of these things mean that it's really hard for you to test whether or not AI is doing the right thing in production. Because LLM as a judge is insanely expensive. You've got to put all of that context in to another LLM to ask it to evaluate whether or not it thinks the previous calls were correct.
1:25:04That's crazy because token costs are really expensive. That's slow. And that's also really expensive. So most people are not doing production evals. And if they are, they're doing it on a sample of their production stuff. And if not, they're using humans to review some of their aging trajectories by doing sampling too. And the idea of not knowing what your crazy black box that's super non-deterministic is doing and whether or not it's doing the right thing in production is insane. And so our views on things were like specifically, how can we make this as close to deterministic as possible? How can we make this work for every agent run in production?
1:25:45And that means using product signals, which is why we ended up building the whole product signal part of our eval alongside allowing you to do LLM as a judge when you think it's necessary. Like, for example, the code review thing. If you have a code reviewing agent and you wait for the PR to be rejected, you might want to use a product signal to wait for that rejection and then use LLM as a judge to rate the feedback that somebody gave when they rejected that particular PR.
1:26:15Turner Novak:Cool. All makes sense. But you're using product signals to derive whether or not you end up asking AI for more information about that rejection. and that way you can basically sample 100 % of your production trajectories and use cases. You can get all of these product signals which are basically free because events are cheap and we've been doing that for decades and then you can sparingly use the LLM as a judge to get more information and score agents in greater detail when you think it's necessary rather than doing it at the sampled rate which is super expensive and you can make your LLM as a judge much more specific to give you much more detail about whether or not I did the right thing overall, which is super cool.
1:26:56So overall, I think that LLM as a judge is necessary, but I think that we can improve the way that we track production AI and that we must do that if we want to track every agent trajectory and then use that to build self-reinforcing agents that get better because that data set must be good for you to take that and do post-training. I think that's really hard for people right now in general, you know, and that's not to say that every eval company is wrong. That's to say that the way that we do it is a superset of the way that other observability companies do it because we can do the same unit testing and eval staff, of checking the outputs, plus listen to product events and so on, which gives you a lot more capability.
1:27:49Well, cool.
1:27:50Turner Novak:This has been a lot of fun. Thanks for coming on the show. Dude, thanks for having me. Yeah, it's interesting, you know, and really, really great to talk to you. So yeah, it's always good chatting. Good to speak with you as well, yeah. Cool. And thank you for listening. Thanks again to this episode's sponsors, Flex, Numeral, Amplitude, Merge, and Monaco. If you enjoyed this, please like, comment, subscribe, and share the conversation with your friend building AI agents. Make sure to check out the back catalog of over 100 episodes with investors like Gary Tan, Alad Gil, Chathan and Eric at Benchmark, and the founders of companies like Robinhood, Sweetgreen, and Mercury.
1:28:22Turner Novak:Tune in over the next few weeks for conversations with Anatoly Yokovenko, founder of Solana, Naveen Chata at Mayfield, Michael Tannenbaum, the CEO of Figure, the first blockchain-based lending company, and my friend Shensi Dang at Merch. If you don't want to miss any of these, subscribe to my newsletter, The Split, linked in the description to get each episode plus a transcript emailed directly to your inbox every week. Thanks again for listening. See you next time.
From the publisher
Tony Holdstock-Brown is the co-founder and CEO of Inngest, the durable execution platform that quietly powers your favorite AI agents.
We get into why agents work in a demo and die in production, building their own cloud to get 20x lower cost, growing 35x after AWS and Cloudflare copied them, growing a dev tools company without a personal brand or Twitter account, why he thinks evals today are like “asking the criminal if they committed the crime”, and the thing they built to score 100% of your production agents without paying for LLM as a judge.
Thank you to Numeral, Flex, Amplitude, Merge, and Monaco for supporting this episode.
Numeral: Sales tax on autopilot https://www.numeral.com
Flex: Premium banking, 60-day credit, 0% APR https://home.flex.one/referral/bananacapital
Amplitude: AI analytics https://www.amplitude.com
Merge: Every model, one API https://www.merge.dev/turner
Monaco: The revenue engine for startups https://www.monaco.com/
Timestamps:
(0:00) The hidden infra layer every AI agent runs on
(1:46) Building complex chains of logic
(3:31) Why agent SDK's don't go far enough
(4:49) Healthcare was the original event-driven nightmare
(6:32) Storing traces on your infrastructure enables self-improving loops
(14:26) Why Inngest was already in the right place for AI
(15:49) Score agents off product events, not LLM's
(17:31) The OpenAI copy-paste signal
(21:24) Swap in LLMs and cut costs 20x
(23:44) How customers pulled the product forward
(25:41) Orchestration belongs outside the sandbox
(29:48) Building a neocloud to cut costs 20x
(32:09) Most neoclouds just resell AWS
(32:54) All AI infrastructure is converging
(34:49) Why Claude can't just build your backend
(36:44) How to build a software factory
(39:12) Agents are a lottery you get addicted to
(42:44) Loops must exist until AGI hits
(45:38) If models keep getting better, why orchestrate?
(48:28) When incumbents steal your features
(52:30) Why you can't vibe code infrastructure
(55:54) Why Tony has no personal brand
(59:38) Dev tools GTM without Twitter
(1:03:20) Lessons from the founder of DuckDuckGo
(1:10:39) Truth as a company value
(1:13:08) Taking too long adapting to AI
(1:15:10) Startups are 100% R&D
(1:17:19) Ali from Databricks
(1:19:03) Writing his own code, Voice-to-text with local models
(1:23:53) Evals are batshit insane
Referenced
Inngest: https://www.inngest.com/
Principles by Ray Dalio: https://www.amazon.com/dp/1501124021?lv=shuf&channelId=500&plpRedirect=mhFallback
Traction - How Any Startup Can Achieve Explosive Customer Growth: https://www.amazon.com/dp/1591848369?lv=shuf&channelId=500&plpRedirect=mhFallback
Follow Tony
Twitter: https://x.com/itstonyhb
LinkedIn: https://www.linkedin.com/in/tonyhb/
Follow Turner
Twitter: https://twitter.com/TurnerNovak
LinkedIn: https://www.linkedin.com/in/turnernovak
Subscribe to my newsletter to get every episode + the transcript in your inbox every week: https://www.thespl.it/




