In short
How Hex (podcast guest Izzy Miller, AI engineer) builds “data agents” that reason like human data analysts, emphasizing context harvesting, tool orchestration, and evaluation/verification.
Guest background
Izzy Miller is an AI engineer at Hex, shipping data agents since before most teams focused on agents. He works on agent orchestration, context pipelines, and an internal eval system.
Key claims
- Agents need more project-level context than single-cell prompting; conflicting context can trigger long “pondering” and collapse behavior.
- Hex’s notebook agent is the foundation; it evolved from AI operating within notebook cells to agents operating across cells.
- Hex is unifying separate agent experiences (notebook, Threads, semantic authoring) into shared “capabilities” bundles (tools + static context + prompts).
- Verification in data work is harder than in coding; Hex uses semantic models, admin governance, and observability/LLM-judge feedback loops.
- Their eval philosophy favors small, handcrafted “failure mode” sets (often 30–50) over huge generic benchmarks.
Notable examples
- An eval where Sonnet 4.6 reaches ~24% by day 90 in a 90-day simulation.
- A “fan-out” internal dashboard bug makes agents confidently claim massive AE quota success; they usually miss it unless prompted to sanity-check.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe Evolution of AI Agents at Hex
0:45 to 5:05
Discussion on how Hex developed their AI agents and the challenges faced.
“People would be like, oh, why can threads do this?”
The Role of the Notebook Agent
5:05 to 9:10
Exploration of the notebook agent's features and its importance in data analysis.
“and they just want to know the number to like build me a crazy predictive model of how our vehicles drive in snowy conditions when the temperature is like this.”
Challenges in Data Analytics with AI
9:10 to 10:40
Izzy explains the unique difficulties of using AI in data analytics.
“I want to talk more about that evaluation step later on, but maybe one last high level question.”
Diverse User Needs in Data Questions
10:40 to 12:50
Discussion on the variety of questions users ask and how Hex accommodates them.
“But all of the code has been kind of abstracted and hidden away.”
Exploring Different AI Agents at Hex
12:50 to 13:50
Overview of various AI agents developed by Hex and their functionalities.
“I think it's like the capabilities that the agent has.”
The Future of AI Agent Integration
13:50 to 14:00
Discussion on the integration and unification of Hex's AI agents.
“When I think of an agent, the simplest form of it, I think, is like an LLM running in a loop calling tools.”
Understanding Context in AI Agents
14:00 to 16:48
Explore the concept of static and dynamic context in AI agents.
“The opinionated stuff is mostly around context.”
Challenges of Long-Running Tasks
16:48 to 20:00
Learn about behavior tuning for long-running AI tasks and technical choices made.
“complicated system to basically map between short references and IDs.”
Overcoming Technical Debt in AI Development
20:00 to 22:32
Discuss the difficulties in managing technical debt and evolving AI systems.
“Do you need to repeat the same thing in every one of those tools?”
Managing Tools and Functionality in AI Agents
22:32 to 26:24
Delve into the complexity of tool management and user experience in AI agents.
“model, like maybe that's just boom, one shot, lickety split.”
Show all 25 chapters
The Importance of Verification in AI Outputs
26:24 to 28:00
Examine the challenges of verification and ensuring accuracy in AI-generated results.
“I'm assuming some combination of all the above, but what's the distribution that you see people using it for?”
Philosophical Challenges in Data Work
28:00 to 30:00
Explore the complexities of data work and verification methods for businesses.
“But when you're doing this for a business stakeholder who doesn't know SQL or doesn't know Python, it gets much more, much more philosophically different.”
Feedback Loops and Human in the Loop Systems
30:00 to 32:20
Learn about the importance of feedback loops and human oversight in data systems.
“Observability actually like baked into the platform.”
Challenges of Memory in AI Agents
32:20 to 34:20
Understand how memory functions in AI agents and the implications for data accuracy.
“Your agent, you know, Claude Sonnet46, what we were running this on, and it actually runs fairly efficiently and quickly up to a point.”
Observability and Evaluation Tools
34:20 to 36:40
Discover how observability tools are used internally for agent evaluation and improvement.
“warehouse context etc and then that that feedback loop kind of churns with the human memories are this confounding thing that can sort of you know inject all kinds of clutter into that loop at the user level today.”
Overhauling the Evaluation System
36:40 to 39:10
Examine the updates to the evaluation system for better usability across teams.
“So you guys are able to see that because you run some LLMs of judges, which cluster tag things, and then maybe you cluster those clusters or something.”
Creating Effective Evaluation Data Sets
39:10 to 42:05
Learn what constitutes a good evaluation data set for AI performance assessments.
“to run evals on changes that they've made.”
Evaluating AI Data Sets for Human-Like Reasoning
42:05 to 46:32
Explore how to design effective evaluation data sets for AI that reflect real-world challenges.
“And so I've really handcrafted artisanally all of these failure modes, traps, I've called them basically that I want the model to fall into.”
The Importance of Model Upgrades
46:32 to 47:05
Learn about the considerations and challenges when upgrading AI models for better performance.
“And we have some really interesting evals for this.”
The Role and Journey of an AI Engineer
47:05 to 51:41
Discover the career path and daily responsibilities of an AI engineer in a tech startup.
“but also upgrading the harness around these new things in the API.”
Finding the Right Talent in AI Engineering
51:41 to 56:00
Understand the qualities and experiences that contribute to effective AI engineering teams.
“I think my official job title is AI engineer.”
The Role of Domain Expertise in AI Visualization
56:00 to 58:27
Explore the irreplaceable nature of domain expertise in enhancing AI visualization.
“agent and think well enough, oh, if I were to solve this, this is how I would do it.”
Evaluating AI's Performance Over Time
58:27 to 1:01:06
Learn about the importance of long-term evaluation in AI performance assessments.
“You know, if the models get twice as good as they are today, what turns out to have been sand that can be washed away and what is stone and remains and is very meaningful.”
Benchmarking AI Agents: Insights and Challenges
1:01:06 to 1:04:36
Delve into the complexities of benchmarking AI agents and the value of iterative evaluations.
“And I think this gets into some of the value that I think Hex provides that we were talking about, about like day zero performance versus like day 90 performance once the flywheel has churned.”
Creating a New Benchmark: Metric City
1:04:36 to 1:07:54
Discover the concept of Metric City and its potential role in evaluating AI models.
“It's just evaluating the model's behavior.”
Transcript
Automatic transcript. May contain errors.0:00Overnight, internally, we released this and everyone was just freaking out. We were like, oh my god, this is it.
0:06Izzy Miller:Today, I'm talking to Izzy Miller, an AI engineer at Hex. They've been shipping data agents since before most teams were even thinking about them. And it was increasingly feeling like we had this high horsepower beast driving 25 in a school zone. It's like the thing we are harnessing here needs more context, not just from the cell, but from the whole project. Izzy describes one of the failure modes of building agents and how access to more context in user behavior unlocked big gains. And if you inject in conflicting context, it will spend 30 minutes pondering, going, wait, let me see, hmm, actually, and enter this crazy collapse mode.
0:38Izzy Miller:Getting agents to reason as well as a human data analyst requires some cleverer approaches. I dig into how Hex started with separate agents and why that's changing. People would be like, oh, why can threads do this? But why can't threads write Python? But the notebook agent can. And so unifying these all together, which has had interesting code level impact on the way that we organize our tools, prompts and skills into bundles. Right at the end, Izzy walks through the eval test suite he's building, a 90-day simulation where agents learn from context and improve over a longer arc of time. If the agent is demonstrating the skills and behavior that we want it to by day 90, all of the questions and the tickets are carefully crafted such that it should get 100 % of the questions right.
1:19Sonnet 4.6 gets 24 % on day 90. It's very expensive to run. I need credits, please.
1:27Izzy Miller:Welcome to Max Agency, the podcast that goes deep into how the best agents are being built by builders like you. Izzy, really excited to have you on. I think Hex has been one of the companies that's pushed the boundaries of what's possible in agents. You guys have launched a bunch over the past year. We use a lot of them internally. I think the notebook agent is one that I hear rave reviews about. What is that and how did that come about? The core of Hex has always been this notebook interface in which you can, you know, intersperse cells of SQL or Python or text or charts and construct very complex or simple analysis in a way that's very scalable, approachable, understandable, like literate programming.
2:08I think we were the first real product to ship a text to SQL feature to real paying users that they were using.
2:17Izzy Miller:What was that product like at the time? Was it typed in, you get SQL? It was cell contained, right? So we had this notebook that had all these cells in it, and we scoped all of our AI features to operate within cells. And this is when way before anyone was thinking about agents. It was just like, you know, you open a SQL cell, you ask for a query, the query appears. This is running on like GPT 3.5 turbo or something. And we were like, oh my God, which all feels very quaint to think about now. And then for the next like year or so, we were very focused on this single shot type workflow, which I think is kind of in many ways, like uniquely cursed in data analytics, because it is this like iterative domain in which you get an answer.
2:59And it's like, interesting, I want to follow up. And so I think that agents were not this thing that, I mean, maybe with the exception of you, I feel like people weren't like, this is coming. And so it took us a long time to like go from AI operating within the cell to AI operating across all of the cells. And we tried. And frankly, I think the models like weren't there the first couple times we tried. And so we like looked elsewhere and built all this other AI stuff. And then there was kind of this moment where we were just like, the models are here and we need to do this again and it's going to work.
3:30And we put a sidebar in the notebook and we gave it all the same features or the same tools that like a user could access in the notebook. And like overnight internally, we released this and like everyone was just freaking out. Oh my God, this is it.
3:45Izzy Miller:Do you remember if there was anything that made you guys realize that it was time to come back to this idea? I think it was two things. I think one, tracking model capabilities and not necessarily where they were at the moment, but like the obvious trajectory that they were increasing upon. And it was increasingly feeling like we had this like high horsepower beast driving 25 in a school zone. It's like the thing we are harnessing here needs more context, not just from the cell, but from the whole project. You know, it needs to be able to take two steps, three steps, etc. So one side of it was the models.
4:16And I think the other was that like it just wasn't working, candidly. Data, I think, is a uniquely difficult environment in which to try and do good single shot text to SQL or text to code or text to answer generation because it is iterative and it's about refining and sort of evaluating an answer and changing and pivoting and rabbit holing. And so I think like we worked really hard and realized we were stalling out and like this more flexible approach was obviously what's necessary.
4:45Izzy Miller:And what types of questions were people asking or wanting to ask, whether it's in the single shot or in the agent that came after? Was it to build a notebook? Was it to get an answer? Like, what did you see people doing? Yeah, it's all the above. I mean, this is a cool thing about working in data in general, I think, but also about working with such a flexible product like Hex is that, you know, people are asking questions ranging from how many widgets did we sell last week? and they just want to know the number to like build me a crazy predictive model of how our vehicles drive in snowy conditions when the temperature is like this.
5:18And like we have such a wild variety of customers and they're all building crazy complex stuff in their domain. And Hex like the I don't know how early and usually you were, but like long before all the AI stuff, the core promise of Hex was like you do this notebook, you construct all of your logic and And then everything there becomes a building block to build this app that is like much more flexible than a dashboard. And so it was kind of fun when we started working with the agent because Hex has always been this tool that was come and ask your hardest, craziest, most flexible questions and go deep.
5:49And I think that because of the model and the harness limitations, the AI stuff was always like so tightly scoped compared to the whole workflow that users wanted to answer. It's like you come in, you have this huge thing in your mind and it's like, OK, I'm going to work with AI and I need to ask it like these teeny little bite sized bits. And I think that's the fun thing about the notebook agent is that like, yeah, customers come in and they're like, OK, like we just launched this new feature. Build me a report of how the users that have access to it are liking it. And like that's the remit. And the notebook agent will work for like 20 minutes and come up with the answer to not just that, but, you know, many of the sub questions that come up along the way.
6:25Izzy Miller:What does it do and how does it do that? Does it write notebook cells? that are then kind of like executed? The notebook agent operates in the notebook almost exclusively with one exception is that we allow it to run scratchpad SQL queries that do not show up in the notebook. But all of the other work that it's doing is sort of contributing to this artifact that A, the user can follow along with as it's happening and B, can ultimately use to turn into some kind of report that they might use or send to someone else or something like that. I'm forgetting the exact word, But Barry, our CEO, went to one of those like VC summer camp things.
7:01And I think Satya Nadella was there and said something inspirational about how many buttons Microsoft Word has in all of its toolbars. And it's like the agent needs to be able to use all of those, like all the power of this very complex tool should be just like a prompt away. And I think that we came into Hex or to building the notebook agent specifically with that same thing in mind, where it's like this is a complex power tool for like technical users. and it has a million little buttons and things to do. It's like, can we just not change any of that, not introduce any new capabilities? Can we just make all of that accessible through the simple natural language entry point?
7:36Izzy Miller:I feel like the agents that most people use are coding agents, and those are also kind of iterative, right? Like, how do you view the differences and similarities between data agents and coding agents? I think there are two big things here. What I was referring to before is that a lot of the way that I code, especially now that models are much more capable, You lay out the plan up front and it executes. You know, you write the spec, build it, and it builds it, and then you get it and you're like, amazing. You're kind of either like, yes or no. And I think that when you do good data science or good data analytics, there are many decision points along the journey of a project in which you're not just saying, good job, bad job, but you're saying like, oh, that's interesting.
8:18Like, should we look at it this way? Or like, oh, like, I wonder if we, you know, break this out. do we see a different trend? And I think that that is something that, again, really does not lend itself to just being like, question, answer, done. And so I think agents really solve that. And I think that you can see with the notebook agent, the adoption of it, and if you've ever tried it, it's a pretty remarkable experience for doing that kind of work. I think the thing that's still frontier difficulty problem is that all of those decision points are very hard to verify and validate the outcome of.
8:50And I think like coding increasingly is like a very verifiable task, right and this is why the models are getting so much better at it because they can train on it with this verifiable reward and you say make a dashboard with these five buttons and it does it and you're like there are five buttons and they work and maybe the code is slop behind the scenes but like the task has been accomplished correctly and i think with data that's much less clear it's actually a very big challenge to how do you validate the accuracy of that question and so how do you train the models to be very good at this and how do you use them in your harness and So I think that that is still kind of like the interesting part of the puzzle, whereas I think agents versus single shot prompting kind of resolved the earlier part of the like interesting AI data puzzle.
9:30Izzy Miller:I want to talk more about that evaluation step later on, but maybe one last high level question. What are the other agents besides notebook agents that you guys are building? And when we were talking earlier, you said that these are all maybe melding together in some way. We'd love to hear more about that. This is all very in flux in our code base right now. Now, we started with the notebook agent, which we just talked about and was very much this kind of like, here's something that already exists and our users love and get a ton of value out of. And can we just make it easier to work with basically and automate it with this agent?
10:01The immediate next step was to build something. And so that's a product for technical people who, you know, are comfortable working in a notebook, a SQL and Python. And of course, once you make it more accessible, non-technical people start doing that as well. And so we have a lot of actually non-technical folks that work in the notebook now to build complex projects. But there's this other demographic in data, which is the true self-serve contingent of folks who want to ask a question and get an answer. And so our second agent was what we call Threads. It's really very, very, very similar to the notebook agent, but it's kind of a more abstracted conversational interface.
10:35Izzy Miller:So just like a chat. Exactly. It looks much more like a chat GPT conversation. But along the way, you know, you're getting these data artifacts that pop up and are interactive and explorable in the same way that would be in any other part of Hex. But all of the code has been kind of abstracted and hidden away. And it behaves slightly differently. It's not question answer. It's like question a ton of stuff, maybe follow up branching rabbit hole and then finally an answer. But it's designed for the kind of thing where it's like you go in with a question and you want an answer. Whereas a notebook agent is much more about perhaps building some complex unknown report or diving super deep into something.
11:09We also introduced some semantic authoring capabilities earlier this year where you can write a semantic model directly in hex as well as import them from, you know, dbt or what have you. And we also introduced a model there. The interesting thing about that model is not necessarily the visible capabilities it has because it just, you know, writes YAML docs, which is pretty simple. But all of the context harvesting that it does in order to be able to accurately contribute to your semantic models. So we talked about unification, but like we very early on gave the semantic modeling agent the ability to read a bunch of other artifacts from around the Hex workspace that it could, you know, uncover interesting things people are doing and upstream them into the model.
11:52And that's now capability that's being rolled out to all of the agents in a first party capacity.
11:57Izzy Miller:And when you say other artifacts, is that like other notebooks and stuff that people have created that might inform? Yeah, notebooks, other threads, conversations, all kinds of pieces of context inside of X. This is, I think, our big information architectural challenge is laying out this graph of all of the context that's like warehouse source, semantic model, some user saying something in a thread, the admin saying something, how that all comes together and is synthesized. So we're actually behind the scenes somewhat. We have a context agent that helps synthesize all of this. We have the chat with your app agents.
12:31If you built a data app on top of your notebook, you have a very sort of threads like experience with that as well. And then, yeah, these are all basically trending towards being the same thing.
12:40Izzy Miller:Same thing from what perspective from UI UX, from the underlying kind of like harness and architecture? I think that the UI UX is a piece that's more up in the air and interesting right now. I think it's like the capabilities that the agent has. If it turns out, big surprise, if you're interacting with something that looks pretty similar in the same product in different places, you expected to have the same tools and capabilities and access. And our initial agents were all totally separate things. And so people would be like, oh, why can threads do this? But like, why can't threads write Python, but the notebook agent can?
13:12And why can the semantic authoring agent like, you know, read my other projects, but the notebook agent can't? And so just unifying these all together, which has had interesting code level impact on the way that we organize our tools and prompts and skills into bundles that we're calling like capabilities now.
13:28Izzy Miller:Great transition. Let's dive into that. Like, what do these agents, maybe we pick the notebook agent, like, what does this look like under the hood? There is kind of this high-level workflow that runs of the notebook agent with a bunch of sub-activities. We've built our own orchestrator to run all of this stuff. We can maybe later talk about how we eventually migrated from that very single-shot-oriented world towards this very, like, long-running agentic world. When I think of an agent, the simplest form of it, I think, is like an LLM running in a loop calling tools. Is that basically what you guys have or do you have more opinionated workflows around that?
14:02That's basically what we have. The opinionated stuff is mostly around context. Great.
14:08Izzy Miller:How does context work? I think like we break our concept of context up into dynamic and static context. And then the different agents you expose have different capability sets, basically. For now, although we're kind of exploring how it works when they have very similar looking capability sets. But yeah, the capabilities bundle together tools, static context, prompts, also some weird esoteric finicky stuff like final turn behavior. How should the agent work when it's wrapping up its project? I think that like to what I said earlier, the actual high level agent loop, I mean, it's a fine piece of engineering, but it is it does kind of boil down to just run the LLM in a loop.
14:50and the interesting stuff is really about the context harvesting pipeline and then like all of these finicky little things about how does it work on the last turn how does it know when to wrap up how long do we let it go for all of that interesting stuff how does it work on the last turn i think that this is like one of the things that we've always been trying to tune and tweak the behavior of to get it to wrap up appropriately i think actually like it's kind of interesting with longer running tasks, more capable agents, compaction becoming a thing, longer context windows. Like I run my coding agents forever.
15:25Same. Like I hope I wish there was no last turn. I want it to just keep going. And today we still enforce, you know, like a hard ceiling on the number of iterations. And so if it's in the middle of something, when it hits that, we force it to wrap up in a certain way. Why?
15:40Izzy Miller:Out of curiosity. You know, you make these technical choices when models are at a certain level of capability. And when you have five of those, it's easy to keep advancing them along as the ball rolls. But I think when you have 500 of them, I feel like every day we wake up and realize, oh, there's this thing holding the agent back. And like, why? It's like, oh, well, you know, it used to help the agent. I think a remarkable amount of things that at one point we felt very proud of building and were in fact necessary are now hobbling the agent, actually. Like just this week, we've been working on figuring out if we can remove this very complex system that we built for changing the way that our agents see the static IDs of artifacts in Hex, whether they're cells, projects, data connections, etc.
16:27You know, we have a million of these things. They all have a static ID, long, unique numbers, and earlier iterations of models, the ones that we built the first agents on, had a kind of ceiling after which they started hallucinating these IDs or making them up or swapping them around. And so if you had a notebook with like more than 50 or 60 cells, you'd start to get crazy behavior. We built this crazy complicated system to basically map between short references and IDs. And this is like a core fundamental part of the agents. And just earlier this week, someone ran an eval and was like, yeah, we don't need that anymore.
17:02They're fine. And we're always fixing bugs and dealing with stuff about this like ref registry compensation system and so there's probably like five more examples of that this week and this is like when you get into the good and the bad of building your own thing we've built our own very custom orchestration system it allowed us a year ago to compensate for this like severe model deficiency and build at a very low intimate level this mapping that allowed us to build and release the notebook agent and scale it and it to become this incredible thing and kind of change the trajectory of the Hex product.
17:36But now it's like tech debt baggage unnecessary. And it's very hard to get rid of tech debt once it exists.
17:42Izzy Miller:When you do these changes to the harness, are they incremental or has there ever been a point where you're like, hey, we're just going to completely rewrite this harness. It's going to be agent V2, agent V3. And I guess like related at those points, did you consider adopting an off-the-shelf harness or not building yourself? And what was the considerations there? And how What do you think about that? We looked at a bunch of orchestrator tools and technologies and made this decision primarily on the basis of things moving really fast and us wanting to be able to continue to move fast with them on our own terms, even if it meant taking on a lot of this maintenance overhead on a, you know, ongoing basis.
18:21In the background, kind of unbeknownst to our users, we migrated everything from this very single shot bulk-queue oriented flow to something based on temporal that kind of has proper long-running workflow orchestration. It was a tremendous task. And the worst part is you have to maintain two things at once. And, you know, we're always going through two things at once as we introduce the new version of something. But late last year, we kind of moved to this new, better world that has since enabled us to build all of these other agents and stuff. It wasn't a true blocker, but it was kind of a self-enforced blocker to building other agents was we want to actually be operating in the proper framework before we do that.
18:59Izzy Miller:You mentioned earlier that the notebook agent needs to be able to do everything in a notebook that human would be able to do. I imagine that's a lot of tools. And now you're talking about combining multiple agents. How do you deal with so many tools? Oh, it's so many tools. I think it's too many tools. How many tools is it? Not quite 100 ,000, but something like that tokens worth of tools in the tool kit, Which is too much, to be clear. It's too many tools. I'm not proud of it. One thing here is to reduce the number of tools and simplify and consolidate. But another thing is to implement some kind of tool search or like tool retrieval, which becomes a big thing with the primary coding agents that have to reach out to a bunch of MCPs.
Read the full transcript
19:38I think if you have a bunch of MCPs installed, Claude Code can theoretically use like, you know, thousands of tools. So we implemented tool search, which helped take that pressure off. I don't necessarily purport to be an expert on this, but we only have kind of like our empirical evidence. But like we had a create chart cell tool, an update chart cell tool, a delete, all of these sort of very normalized tool sets. And it's always a little bit interesting and unclear unless you evaluate and test. Do you need to repeat the same thing in every one of those tools? Can you say it in the system prompt?
20:07Is saying it in one tool as good as saying it in all? And it's like, kind of. But we basically see we haven't gotten around to it everywhere because it's just kind of work. But I did a test, ran an eval, and we can consolidate this all into just chart tool and it works fine. So there's that kind of just meaningless tool explosion. But there's also just a lot of temptation to introduce lots of little tools when you have a product that doesn't operate at the command line coding layer.
20:33Izzy Miller:Have you thought about making it operate at the command line coding layer? Because you guys also already run it in the notebook environment, right? We do, which isn't truly the command line. It's running in like an IPy kernel type, not exactly, but similar environment. But it has a coding environment with functions. It can run arbitrary Python code. And yet we built it a tool to check if a certain package is installed in the environment or not. And part of the reason for that is do you want the agent to run a quick little ephemeral tool or do you want it to write Python code that, you know, becomes a cell and is visible to the user?
21:10Like, oh, should we let it run ephemeral Python code? And then you get into all kinds of interesting behavior around like, once you let the agent, especially modern agents, run secret ephemeral SQL queries in Python code, they really like these days to be pretty sure of themselves before they get back to you. And so like, especially GPT-5 series models, if you ask GPT 5.4 a question and it's the wrong question and it turns out to be very complicated depending on how it's feeling it might run like 50 ephemeral SQL queries just to be really sure before it actually starts doing any real work and so these things have trade-offs so yeah we could give it you know very generic tools to just run code but introducing very specific tools allows us to introduce these kinds of behavioral guidance around when to use this tool or when to take this action, I suppose, that you don't necessarily have if you're just like, here's a Python tool and use your best judgment.
22:06Izzy Miller:So I didn't realize this, but all the code it writes basically shows up as a cell in the notebook. For the notebook agent, almost all the code it writes. But it does have this ephemeral SQL tool. I'd be curious to hear like why that and like how do you see it using that? It comes back to what I said before about like just how important it is to be able to do a lot of work when you are answering a data question. If you ask, did we have more users of this product feature this week than last week? And if you have a well-modeled, beautiful semantic model, like maybe that's just boom, one shot, lickety split.
22:36But a lot of people don't. And I think a lot of people that think they do also don't. And you know, you might need to run some quick checks and be like, okay, does this table actually have the information I want? I have to join across these two tables. What format is the data in this column in? And there are all these things that you could discover iteratively by error-driven debugging. Like you could throw out a SQL query and get an error. It doesn't look right and refine, refine, refine. But we found that it's a more efficient workflow and better for the user if instead of doing that, it kind of does these very small little atomic ephemeral investigative queries.
23:07And then once it's gathered enough information to write the primary query correctly first go, then it will write that, pop it in the notebook. And it's a much better experience for the user than having to refine and iterate on that query multiple times. But again, as with most things you give agents, like it does have kind of unintended consequences where we're constantly battling against the case where a user asks a really simple question and it's just like tells you the answer because it ran a little secret SQL query and the user's like, where's the chart? Like, where's the proof? The agent's like, take my word for it.
23:40And so figuring out these little behavioral edge cases is always the thing when you introduce a like new tool that isn't just right code. But also like, you know, I use Codex a lot. I like to use Codex to code. And Codex has like a particular pension or GPT 5.4, I guess, for running like crazy bash. And like, I've noticed it writing like Perl commands now. It's like, I kind of understand vaguely what's going on when it's writing bash, but not really. But it's right running Perl commands. Like, I have no idea what it's up to. And I think that, again, is one of these things where it's like, okay, maybe to do that for the like technical notebook agent audience.
24:17but we have these agents that are operating in different contexts for different users. And I think a lot of it does boil down to the UX of knowing exactly what the agent is doing, because they take a while and they think and they ponder and they run a bunch of tools. And I think that engineers, the hands are like coming off the keyboard and it's like, yeah, run for 50 minutes and I don't care what you do as long as you do it right. And I think that maybe the like business stakeholders will get there soon too. But I think more than you think, they actually are interested in following what's going on.
24:46Izzy Miller:What do you show these business stakeholders? Do you show them every command that's running? Do you collapse them and show the little quad code, like, you know, noodling or whatever, whatever they have? I'm like, perhaps regarded internally as like a sort of class clown jester type guy. And I hate the little noodling thing. It's like my biggest Grinch, like Scrooge McDuck take. We like add some version of it. And I was like, no, no, this is like, I don't know. for some it's there's no explanation i'm just a grinch about it um we say thinking very professional while the agent is working we show you what it's up to expanded and then once it's done we collapse it which i think is for now kind of a nice paradigm but i don't know if you've played with gpt 5.3 codex spark or whatever this the open ai is like newest the really fast one the cerebrus test basically it's so fast there's no beyond no need to show the user what's going on there's no way to realistically show the user what's going on when it's like and so i i am operating under the assumption that this is going to be one of those things that we felt really stoked about today but in a few months they're going to be desperately figuring out how we rip out the you know user following along behavior because we want to unhobble our agents to do more and you know, work more verbosely or do, you know, a million tool calls in parallel without overwhelming the user.
26:12And I can already see that this is going to be a like UX capabilities friction point soon.
26:17Izzy Miller:What about the end results that the agents create? Is it an answer? Is it a notebook cell or multiple notebook cells? Is it a chart? I'm assuming some combination of all the above, but what's the distribution that you see people using it for? Without being too much of a Kool-Aid man, I think that this is one of the cool things about Hex is it's such a flexible product. It's kind of like always been such a flexible product. I think that I used to work for Looker before Hex, which is a tool that lets you, you know, build dashboards and run explorers that are basically charts that look the same. I love Looker, actually.
26:49But one of the things that was most thrilling to me about Hex is the diversity of report and answer you could construct. And I think that we have carried that through to the agent world, like the Threads agent, You know, notebook agent, you can build truly anything you dream of. Interactive apps that you click and they run some workflow and update crazy forms. You can build some really wild stuff. But even in threads, it really depends. And I think that flexibility of if you ask a question, it's very simple. I'm always trying to get the agent to give you the least cognitive overhead answer. And so if you're like, hey, quick question, how many users does this company have?
27:26So you just want to know. You don't want to report. Do you think it's important for the agent to like show its work? This is why the notebook is so great. And I think why we got so far so fast with the notebook agent, because famously notebooks are literate programming. This is sort of their entire point is to be able to self-document the work that's being done in a way that is remarkably well suited to allowing agents to write code. And I think also there's like a sense that when you are writing code with AI for a technical user that there is kind of this expectation that they can follow along and can verify and validate.
28:02But when you're doing this for a business stakeholder who doesn't know SQL or doesn't know Python, it gets much more, much more philosophically different. Right. This is codex running a Perl command that I cannot follow along with. I run codex with the outputs off, so I don't even know what it's up to. I've moved past that. I think that the concept of showing your work is not the right concept for data work in terms of citation and verifiability. I think it needs to be some stronger form of verification or confidence or accuracy. This is something that we're working on right now and trying to figure out.
28:36Izzy Miller:What does that mean? What are the ideas you have there? We don't have a massive amount of this figured out yet. There's one really obvious amazing way to do this in the data world, which is to use a semantic model. If you define an amazing semantic model and you let users use it within reason, you can probably be relatively sure that they're getting correct answers. People always figure out how to shoot themselves in the foot. Like I would always shoot myself in the foot with Looker's perfect semantic models. But that is one way to know. And then you don't necessarily need this like, you know, separate clever verification loop because you're basically always operating on like trusted governed context as one half of it.
29:15We do this as semantic models, we do this with sort of admin endorsement or verification of various assets, etc. But the models love to write SQL and code. And sometimes you need to go beyond the semantic model, because that's like the whole point of the task. And I think then you get into this interesting world where verification and verifiability becomes very difficult for data things. And often it's very vague, like what truth even is, depending on how two teams define a metric or, you know, a data pipeline update changing the number out from underneath you. And so I think it's just a really tough domain.
29:52I think the way that we're gonna tackle it, the way that we're already tackling it is with a ton of feedback loops from the user backup to the data team. Observability actually like baked into the platform. We have this new context studio that basically gives admins on the data team a bird's eye view of what kinds of questions people are asking, what answers they're getting. We flag out using a separate LLM as a judge cases where we think something might have gone wrong or where the agent might have been confused or there's a warning that conflicts with your semantic model. Today, all of this stuff runs as a sort of post-action and gives admins the opportunity to improve the context that allows the agent to do better.
30:34So it's kind of like human in the loop system for the model does something, the user perhaps gets a wrong answer. It's flagged up, the data team becomes aware of it, they can improve the guide or the semantic model. So it doesn't happen again, maybe, you know, let the user know what happened, etc. Once you've built the like human in the loop system, it's like, okay, how do we begin to automate this and scale this and make it work better, faster with AI and agents in the loop.
31:00Izzy Miller:Would that look like an agent updating the context itself? Yeah, we've been working on this context agent internally and prototyping what it might look like for the agent to do this. Is this what you guys would consider memories for hex agents? A little bit. I think that the interesting thing with memory and how it plays into this is that memory, as users think about it, I think is a very user level concept. your chat GPT memory remembers your conversations and that you told it you have a dog named Rover. And so when you ask, my dog is sick, it's like, oh, no, sorry to hear about Rover. I think that's less impactful and almost scary.
31:37Actually, I know, not almost, very scary to data teams and admins who are worried that the exact opposite of what we just talked about might happen within a user's little memory where the user asks a question and gets a wrong answer or something or the user tells the agent, no, that's wrong. This metric is actually defined this way. And maybe that user is wrong or it's outdated and that gets stored in memory. And then you have these kind of conflicting levels of context.
32:03Izzy Miller:Because the different levels would be like user level, team level, or how many? Yeah. Yeah. Potentially even more. Potentially. And the data team, you know, sets down these guides and semantic models, reconciling. Like, I think one thing that agents and models do very bad with today is contradictory, like dissonant situations. We actually ran an interesting evaluation on this, and it wasn't actually necessarily looking at accuracy with regarding how can it resolve a conflicting advice situation, but just the amount of time it spends thinking. Your agent, you know, Claude Sonnet46, what we were running this on, and it actually runs fairly efficiently and quickly up to a point.
32:41And then if you inject in a piece of conflicting context that goes against some other information it has, it will spend 30 minutes pondering going, wait, wait, but let me see, hmm, actually, and enter this crazy sort of collapse mode. And I think this happens a lot on a very smaller level. And so I think without having the secret sauce of the answers, I think this is something that we're just now starting to reckon with of like admins need to be able to provide very strong governance and guidance. One way they can do this is a semantic model, and then everything is just on Rails. But if you do that, you miss out on this glorious world of more flexible work.
33:16There's kind of this like, oh, like thing on the horizon. It's like, you know, what if you can just tell the LLM stuff and it's careful about it? That's basically what people do with guides and rules files, et cetera, skills.
33:28Izzy Miller:So do your agents have a concept of skills in that way? They do. We call them guides, but they're modeled exactly like skills with sort of progressive disclosure, the agent, when it's running, can see all of the guides that are available from that workspace, and then it can retrieve them and read them as it's operating. And is that how this memory gets passed in to the agent? Or does some of it also get inserted into the system prompt and what determines what goes where? This is all like bin testing at the moment. For now, we're keeping memory as a separate thing from these other sources of context that are much more data team admin driven.
34:03I think that like the the core feedback loop of hex improving today is driven by the data team it's by users doing work their work or sort of the exhaust of their work being surfaced to the data team in an interface that allows them to you know notice mistakes improve the guides the models the warehouse context etc and then that that feedback loop kind of churns with the human memories are this confounding thing that can sort of you know inject all kinds of clutter into that loop at the user level today. And so we're being very thoughtful about how we roll it out. I think we have more testing to do.
34:39Izzy Miller:You talked about observability and evals in the form of like LM as a judge a few times already. How do you think about the observability that you guys have as developers of the agent versus what you expose to like admins or people in, I think it was agent control or... The context studio. The context studio. Yeah. Is that the same that you guys use internally? Is it different? What are the similarities? What are the differences? My hot take internally that people always argue about is I think they should be the same. So I'm guessing they're not the same. That's a hot take. They're not the same.
35:12I've built an internal observability and experimentation system. I sort of dream of a beautiful utopic future in which we use very much the same tools as our users do to understand what the agents are doing in the Hex product and how we might make them better. And I I mean this for observability and evaluation. We're just now starting to think about how we expose evaluation tools to users. We have a kind of very rudimentary setup for this now. If you edit your context, you can kind of run some tests. I'd like for these things to be convergent. Today, the observability tools that we built, I kind of built the first version of them in the beautiful before times before we launched this to any users.
35:56And it was only internal usage. And we had, you know, full Panopticon privileges. And it was open season on all the data. And we still have that to some extent for our internal usage, which is really helpful to be able to see. Even just your own local development usage to be able to like really deeply introspect it all. I think that A, obviously we can't do that for real user stuff. B, I don't know if admins or data teams want their job to become pouring through everyone's conversations and being sort of synthesizers of all this like that to me feels like an agent's job. And so when I think about our internal observability tools and where they're going and also where the context studio feedback loop oriented tools that we expose to our users are going, I think they ought to probably become more agentic or at least higher level.
36:45The things that we can see about usage data from production are like the clusters of issues that are occurring, the kinds of failures that are occurring in the agent and what those clusters are and how they're shifting over time.
37:00Izzy Miller:So you guys are able to see that because you run some LLMs of judges, which cluster tag things, and then maybe you cluster those clusters or something. And then that's the data that basically you guys have access to, which treats all the underlying traces. You don't have that raw data. You just have kind of like the clusters or the insights from the LLMs. That's right. There's a blog post by Anthropic. It's about this. Clio, I think. We actually built this into Langsmith. Yeah, I think it's a great idea. Yeah, privacy preserving something, something, something. Yes, exactly. I think that this, again, is another one of those situations where like we built our own observability and evaluation stack here because of this.
37:42I don't know if I'll comment on whether it's truly correct or not, but this kind of like perceived need to be able to move very, very, very rapidly. And this uncertainty about what the agent or the product will look like or the model will look like three months in the future. and a fear of limitation if we didn't build our own thing. And like the tax is tremendous. I have spent the bulk of last week actually doing very little besides updating and refactoring and working on our evaluation system to make it more user-friendly. Our eval system internally, we call this system the shoebox, which I tried, I tried.
38:17How did you come up with that name? Because I didn't want it to last. I wanted it to be like the shoebox where you just stuff all our shit, all our receipts into and put it under our bed. And then it caught on and is still around. The shoebox, which exposes some evaluation capabilities, was always kind of a like high priesthood tool available to just the AI engineers that were like actively iterating on the agents and prompts. It was very difficult to use and had a terrible UX and was kind of the kind of mini foot guns such that you needed to know how to use it in order to actually make good use of it.
38:47But now the line between AI engineer and just any other kind of engineer has very much blurred internally. We have all these agents. Most of our new features are agentic in nature or somehow related to AI. And everyone wants to run evals and validate if the feature they're working on is improving or hurting things, or, you know, you're doing some big refactor and you just want to double check. It doesn't mess stuff up. And so we are overhauling the eval system to make it really, really easy for anyone at the company to run evals on changes that they've made. The two questions you want to ask are, did my change have the desired effects and did my change have any undesired effects?
39:25And being able to answer those two questions is the goal for anyone at the company.
39:29Izzy Miller:What exactly does that process look like? Is there one big data set of a thousand examples and they all have a ground truth and you run it against that data set and then compare it to the ground truth with an LM as a judge? Or is it different than that? And does it branch off from that in some ways? Nailed it. Okay. How many examples are in your eval data set? Like this is something we hear a lot of people ask, how big do I have to get my eval sets to be before? And maybe curious to hear if this has evolved over time as well. I have a lot of opinions here. I think as a general rule, most eval sets are bad unless they are being actively worked on.
40:05I think almost, or at least in the data space, pretty much every time I've cracked open the hood on some data benchmark or evaluation, I've been very disappointed by what I saw. I've seen. I don't want to slander anyone, but just sort of. But the rule applies to everyone that I've seen. Bad ground truth, incorrect ground truths, problematic grading. Some folks try and do deterministic grading, which is a valiant quest, but your script has bugs or doesn't accept a percentage as a decimal point. And like all of this stuff, some questions are just bad or not representative. There's a very popular benchmark set out there that we actually do use a slightly altered flavor of.
40:48And it's quite difficult unless you know the like one secret to the benchmark. And many, many, many of the questions in this data-related benchmark revolve around whether or not the agent correctly treats an empty array as the same thing as a null or not. And there's like this esoteric tiny little rule in a manual somewhere, and that's make it or break it. And this is not like this isn't a data benchmark. This is like a needle in a haystack, like context attending benchmark. And I think that a lot of benchmarks out there conflate that. Talking specifically about analytical reasoning and data capability, I think that a lot of benchmarks conflate SQL syntax and retrieval and like this kind of needle in a haystack stuff with actual desirable analytical behavior and the kind of scientific reasoning that you would want to be able to do this, that process I described at the very beginning, where you get an intermediate answer and now you have a decision point of like, do you accept it?
41:51Does this tell you what you want to do next? Do you reject it? That, I think, is behavior that's very hard to evaluate and that most of the eval sets available publicly do not evaluate at all.
42:04Izzy Miller:What does a good eval data set look like? Well, I have some opinions. I think one is that a good eval set should be actually small enough that you as the interested party can can sort of hold it in your mind this may be very controversial i don't know i like to be able to know why all of my eval sets like we have these aspirational eval sets that our agents do very poorly on that all all agents do very poorly on like opus 4.6 max gets like 20 on super hard i think that that eval set is much more useful if i know why, you know, G7 to eval like fails or it's like there's these four reasons. And so I've really handcrafted artisanally all of these failure modes, traps, I've called them basically that I want the model to fall into.
42:52And there's a lot of benchmarks out there like this one I was talking about with the empty array. It's like 390, 470 ton of evals. Most of them are just variants on this one gotcha. and I think that it's more interesting to have a couple of gotchas that you can maintain in your head and then just run a ton of repetitions on them rather than fan it out to a bunch of different
43:13Izzy Miller:cases and make things more complicated. So that's one opinion. Our eval sets are like 30 to 50 for these very very difficult cases but again we run multiple repetitions on them. I also think that a lot of data eval sets now are no longer representative of the kind of work that users are actually doing. They're much more of the thing I was talking about before actually this like single shot pub trivia type question where it's like, how many users did we have on April 16th, 2019? Can you syntactically rearrange this English into SQL? It's like evaling a coding model on tab complete when everyone's trying to ask it to write full on things.
43:48Yeah, exactly. And so a lot of our most interesting evals are notebook agent evaluations that begin in the middle of a very complicated notebook that has already been aggressively built out and the user says something like that's weird that's not the number i expected to get and the eval has been carefully crafted such that there's like a chain of three bugs in the data and the sql queries that the agent needs to kind of unravel and explore and just looking at the state that the like notebook is currently in tells you about the first bug but it actually obscures the second and third bug and So it needs to be this actually like thorough process in order to actually resolve all of the bugs and get the answer.
44:31Izzy Miller:Do you take those from real trajectories or things that you see or are those completely synthetic hypothetical made of ones? We have the luxury of having a ton of internal hex usage for like real stuff. And I model a lot of it off of that because data is hard and we make mistakes internally also. My favorite eval that all current models that I've tested it on have failed, though the newer ones take longer to fail on it, is an internal dashboard, real internal dashboard about our sales like AE quota attainment. And I took the dashboard. I intentionally introduced a fan out bug that makes it look like all the AEs are dominating.
45:13Like everyone is at at least 900 percent of quota, like killing it. best quarter ever and then you ask the agent how are the top performers doing this quarter
45:23Izzy Miller:and every agent is like oh my god like it's the best quarter ever like your business is popping when you compare this to last quarter it's like step change you know like josie has 1200 percent of her quota like they're all stoked and i actually tested you can ratchet this up into the thousands of percents of quota before the models start to be like, there may be a data accuracy pipeline error or something. And almost none of all, literally none of them catch the bug itself. But if you then say, that doesn't seem right, take some 10 seconds to catch the bug. This is what I think is most interesting to evaluate for.
46:04I'm being very sort of grandiose. We also have just a totally normal set of like normal person evals that help us know if our product is like good day to day and we try and maintain a reasonable pass rate on them and it kind of helps us prevent regressions and lets us evaluate new features etc just like normal data questions that have a right answer and maybe there's a messy warehouse or whatever but then I also think it's interesting to maintain this like very aspirational set that measures the things that the models are all very bad at right now and there's a number of cases similar to that where it's just having the human in the loop for some of this stuff still makes a very big difference with regards to the agent's ability to like catch mistakes or have that kind of like, interesting, like their ears don't perk up the same way a human analyst's ears perk up.
46:52And we have some really interesting evals for this.
46:55Izzy Miller:Speaking of the models, as new models come out and get better, but also introduce kind of like new capabilities into the APIs, how do you guys think about upgrading these models, but also upgrading the harness around these new things in the API. Data analytics and data science is a task that requires general intelligence rather than like domain intelligence. And so I think just smarter models is great for us and for our customers. So we always try and be on the smartest, newest models. Which ones do you think those are right now? OpenAI and Thropic? Opus 4.6 and GPT-5.4 are extremely capable in our domain.
47:32They have all these knobs and widgets to tune and sliders. And so it's harder than I think people might assume to be like, this one's the best, that one's the best. Like, well, GPT-5-4 like does a little better, but like takes twice as long. And it's like, oh, well, you know, you can ratchet down the effort and then it's like takes actually half as long. It's like, oh, but then it doesn't do quite as well. And I think effort is this kind of new thing since Opus 4-5, I think. And GPT-5-3, both of those labs have built in this concept of effort. I think OpenAI calls it juice internally, which is funny.
48:07And this is something that I do not yet fully understand. We ran an internal test. We had a like effort picker and the feedback we got was, is it working? It does not seem to actually be having an effect because sometimes on low effort, the model gets confused and goes into this spiral. And sometimes on high, it, you know, just answers a question like that. So to answer your original question, we always try and be on the latest, smartest model. now we always try and run it at a high enough effort that we see evaluates well without making our users wait 10 minutes to get the answer to a simple question what's increasingly interesting and challenging i think is that like maybe this is what you were referring to all these other little add-ons keep coming with them like tool search api or like server-side compaction and these are also all things that are very valuable to make use of and also kind of a pain in the ass to maintain, especially when we've built all of our own stuff around it.
49:03So what do you guys do? Do you use them? We do. Well, we use them where they help. I think that part of the benefit of having our own thing and having evals is that we can actually say, this helps. Does the 1 million context token window help or hurt? It's like, well, it hurts at the margins because you can't fill up that window without incurring pretty severe intelligence or intelligence, almost the wrong word. It's like model starts doing weird stuff at the ends of that window. So it's okay, we want to like compact very early. But you know, maintaining access to the 1 million token model via some beta header is actually valuable.
49:37So we can compact at, you know, 300 ,000 instead of 200 ,000, because we found that's the best. So we do support all of these things whenever possible. I do think though, that like talking to the labs these days, I think that they have a vested interest in you using their proprietary locked in stuff and the kind of technique, the psychological warfare technique that I feel like they're using on us to try and get us to use this stuff is the words in distribution. And I don't yet totally know how I feel about this, where they're like, you know, if you use the Cloud Code agent SDK, like you'll be in distribution for what we've trained the model on or open AI is like, you know, if you use it like stateful server side execution or whatever, you'll be very in distribution for the new models.
50:19And I'm very open to that being true. I wonder how quickly that goes out the window once you like throw in your custom tool to build a chart or a SQL query and it's like are you suddenly out of distribution again or do you wind up in a weird codexy case where it says screw your tools I'm going to use like Perl. So I don't know totally how I feel about that but it does feel like that's where it's trending right it's like there's all these little bits and bobs that are layered into the new models and some of them are very helpful and we try and adopt them. I think it's very difficult for us to think about moving from this like a la carte adoption to actually centralizing into the full-on SDK or into like some, you know, stateful programmatic tool calling like server side thing.
51:02Izzy Miller:Because you would give up a lot of the control that you like. I think because we would give up a lot of control. And I think just at this point in time, the benefit quantitatively is a little unclear to me. I don't yet have a strong intuition for just how much it matters to be in distribution. Because these models are very good at in-context learning. And I don't know, I mean, the model uses our harness pretty dang well. And it's unclear to me what being more in distribution might get us versus the trade-offs of being slightly more locked in. Would you describe yourself as an agent engineer, as an AI engineer?
51:38Izzy Miller:What is it that you do on a day-to-day? And how did you get there? I think my official job title is AI engineer. I feel like I never really had a job title at Hex anyway. I was an early-ish employee joined to do marketing, to do community, dev rel, technical, developer marketing, all that jazz. And when you join to do that at a very small company where the only marketer go-to-market-y person, you wind up doing a whole lot of stuff. And I did that for like four years. And I think I always wanted to be an AI engineer. It's just the job didn't exist. I was always technical. I was always a practitioner first, a data practitioner, not necessarily a full-on engineer.
52:19But my kind of my dev rel philosophy was always you shouldn't be a marketer. You should build stuff and then talk about it. And you shouldn't talk about anything you can't build because that's what I call lying. I prefer a more honest form of marketing. And so I was always technical, always very interested in AI stuff. I helped actually build our first version of those old single shot text to SQL workflows just out of interest. And the honest truth is like, I don't think I could have been an AI engineer until GPT-01. I don't know when the models got capable enough to allow me to be a productive member of the engineering team and actually like contribute valuable, meaningful, clean code.
52:59but uh sometime in the last year and a half the models hit the point where my experience with hex my opinions my like user empathy gathered over four years of doing dev rel work for hex paired with claude or whatever has basically allowed me to become very rapidly an engineer which is sweet and um i think there was like i was just reflecting on this there was like a very brief window where I was writing all of my code by hand. I was like, I'm an engineer. Like, this is so cool. I can't believe I'm an engineer. And then I started like copying and pasting code from ChatGPT more and more. And now it's like, I'm kind of back in where I was before.
53:40I feel like I've forgotten how to actually code. And I'm right back to square one two years ago. I mean, I'm really good at reading code, but I don't write that much code anymore. It's interesting.
53:50Izzy Miller:What do you think the right profile for an AI engineer at Hex is? I think there's a lot of different work to be done. We hire people from very diverse domains. We have mathematicians, all kinds of folks, marketers. I think I feel now similarly about engineering as I do about DevRel marketing work, which is the best person is someone who really cares about the problem that you're trying to solve. that was always what I told people to hire for in DevRel and it's like you can teach them how to do marketing stuff like they're bad at writing I don't know you can probably teach them how to be a little better but you can't teach them how to like really care about your domain enough to like stay up late and build that cool feature that's going to be the one that like goes viral on Hacker News or whatever and I think that was just something that didn't really exist for engineering because like well you have to be a good engineer if you want to build incredible things you can't just have opinions and really care you gotta also be able to do it and I think it's still very important to have senior engineers.
54:45And I think I read and review and manually edit all of my code. But mostly I think my job is just like to care and have an opinion. And so I think if you can somehow find those people, I think that's really interesting. One of our most productive members of our engineering team is that it was a user. And he gave us enough feedback finally that one day we grabbed him and threw a hood over his head and put him in a chair and we're like, you work for us now. And he's killing it. But he was a data scientist before, not an AI engineer or any kind of engineer. And he's, you know, phenomenally productive, primarily as a function of his opinions and his instincts and intuition for the things that are going to be useful for our users.
55:25And it's kind of funny, like I almost model in my head, maybe this is anthropomorphizing or worse, but the notebook agent and the threads agent is a data scientist, a data analyst. We tell it, it's like, you are the hex threads agent, like you're a data analyst. And so having that user empathy, I think actually makes people better at building the agent itself because you know how it should work and you can tell when it's doing weird stuff that might be an indication. Not that it's like in distress or whatever, but that something is holding it back or something is keeping it from this golden path that it should be allowed to run on.
55:58Izzy Miller:They can put themselves in the shoes of the agent and think well enough, oh, if I were to solve this, this is how I would do it. They would care enough about that. I think so, but I think the domain expertise is irreplaceable. I think a A very specific place I think it's irreplaceable is visualization, which is actually a place where I don't have that many opinions. I have lots of opinions about like the, you know, basics of how to not commit crimes, but not at an incredibly deep level. But we have two people on the team, Madeline and Nicholas, who I don't know what level is beyond expert, but they're like the like shokunin level of visualization experts.
56:37experts. And increasingly, I think their job, I hope they don't mind me saying this, is to just have incredibly strong opinions about how this stuff should work. And half of the time, they're directly encoding those into the agents. Half of the time, they're telling other people that they did something wrong or that they should have built it this way and helping other people build in a way that is more opinionated. When you're building products in the age of AI, it's always a funny dance of your opinions, the user's opinions. Users sometimes expect now to be able to just type something in and the product looks totally different.
57:06But I do think that with data and visualization and stuff, it's important to have these strong domain expertise driven opinions. You can codify into the product to make it delightful. And I think specifically with regards like the viz and stuff, I think we get a lot of feedback that's like, Dex is mind blowing, like the Hex threads agent is amazing. And sometimes when you like drill on it, it's almost hard for people to say exactly why they're like, I got the answer I wanted. It was way better than, you know insert other tool here and then you like really drill on why and why and why and a lot of the time at least when i talk to users they're like yeah it just was like better and i think that that is like a very interesting je ne sais quoi because like accuracy whatever but if the answer is like a long report with five charts and these bullet points and this offer to look at this next that's the kind of thing i mean that is impossible to evaluate i lose sleep at night trying to figure out how to evaluate that vagueness of what it means to be great at data outside of just getting the answer right in a very, very complex situation.
58:04Izzy Miller:Do you think that's the main value you provide is helping the agents be great at that and get those great answers, even when it is tough to evaluate quantitatively? I think so. I think that this is something we always need to be reevaluating is like how much of it is the model and how much of it is the harness. And I I think Barry, our CEO, has gotten really into saying what's sand and what is stone about the product. You know, if the models get twice as good as they are today, what turns out to have been sand that can be washed away and what is stone and remains and is very meaningful. And I think as the agents get naturally more capable at data science and analytics and stuff, I do think that for a long time, at least I won't say always, but I think that for a long time there will be value in the opinions and the domain expertise.
58:49being encoded alongside the tools that will make Hex better than a snowflake connector with a
58:55Izzy Miller:markdown document that you run on your command line. I will say, though, I think that's like this much of the pie. And I think most of the other pie is the context exhaust that we talked about earlier in that like feedback loop of if you work in Hex and you're producing artifacts and getting answers and building projects, and those are all becoming sources of information, potentially verified for the agent to, you know, validate or guide its future work. That is most of the value the product provides. And I mean, it's impossible to provide that value on day zero. So I think like the day zero version of Hex is better than your rando connector.
59:36I think mostly in just the form of being opinionated and user friendly. And like, I would actually, maybe it's a hot take, but I do think that on day zero, clawed with a snowflake connector is probably like, oh, probably just as good. But the value is over time. It's not just a tool. It's like a platform that everyone works in and the flywheel causes it to improve over time. And so by day 90, you're operating a totally different ballgame.
1:00:00Izzy Miller:What are you most excited for next? There's one thing relating to evals that's kind of interesting. The way that I constructed the warehouse that we use to power our internal eval sets, it's done very aspirationally. I've like left the door open for much more to be done in terms of evals here. So today we have a Snowflake warehouse with a realistic representative amount of business data in it about this made up company, Shorelane Commerce, that does, you know, tchotchke sales. And I very carefully hand injected all of these terrible data quality problems into the warehouse that cause the agent to have a difficult time and have to like push through a bunch of nulls and messed up columns and joins that don't quite work.
1:00:47all sorts of misleading, confusing things. That's like V0 of our benchmark is kind of all of these questions that are very realistic of what users ask of a very realistically messy, like half modeled, but half messy, filled with quality issues, realistic warehouse. In addition to this snapshot, the most interesting way to evaluate data agents and data work has to be long horizon. And I think this gets into some of the value that I think Hex provides that we were talking about, about like day zero performance versus like day 90 performance once the flywheel has churned. It's kind of unfair to make Claude answer a bunch of hard questions right on the first try.
1:01:31And it doesn't ever have an opportunity to try again. And that's what most eval sets are. It's like, who's the president of Malaysia? And you're like, good. It's like, wrong. And then like you never revisit that and no value came out of that. I think that Hex is a product is the kind of product that should get better over time every day. It should get better and better and better at your tasks. And so if we're not evaluating our product in a way that allows for the agent to, you know, actually demonstrate its ability to compound, we're only evaluating this like point in time part of the system.
1:02:04We're not evaluating the full flywheel. And so the full vision for this evaluation set is that day zero snapshot of the database basically is used to run a battery of benchmark questions, which are the evals we already have. Many of these are borderline impossible to answer with the information available to the agent on day zero. But I've built a 90-day simulation during which the clock ticks and every time the day turns over, dbt models run and alter the warehouse to keep it sort of time shifted up to date. New rows are coming in, things are breaking, new products are launching, fraud is happening, and tickets are coming in from stakeholders to the agent in the form of email tickets it needs to answer.
1:02:51and they're asking data questions and they're telling it information. And along the way, it's kind of discovering things about the warehouse. And once it's replied to a ticket, the evaluation doesn't end or the conversation doesn't end. It gets told, you have replied to the ticket. You now have access to an end turn tool. But if you'd like to do a little bit of proactive knowledge work before ending your turn, you're welcome to. You can follow up on loose threads from this conversation, document some of your findings clean up the wiki etc and we let the agent sort of proactively do its thing and then end its turn whenever it wants and this happens every day for 90 days and the conceit is that if the agent is demonstrating the the skills and behavior that we want it to by day 90 all of the questions and the tickets are carefully crafted such that it should get 100 of the questions right and so today it's very expensive to run i need i need credits please.
1:03:50Sonnet 4.6 gets 24 % on day 90.
1:03:55Izzy Miller:What does it get on day zero? Like 4%. But again, uninteresting to me is what it gets. Like, if it got 100%, I'd be like, well, my benchmark sucks. But again, I think that, like, for these very difficult questions, it's not a productive number to know day zero performance because it's not realistic. It's not what the real world actually looks like when they're evaluating this stuff. This is sort of like, mad science laboratory state right now. But in my spare time, I'm tinkering with making this something much more legitimate. So we actually can evaluate our, and this is model based right now.
1:04:31This is running in sort of its own harness. It has very simple tools. It isn't evaluating the hex agent at all. It's just evaluating the model's behavior. Part of it is that I want to observe. I want to see, thinking about like in distribution, how does the model like to organize its wiki? What kinds of information nuggets does it like to store? How does it retrieve them? What does it find? What does it miss? Like it's almost a fact finding, like a research project so that we can make our harness better.
1:04:56Izzy Miller:Do you run the hex agent on this benchmark as well? The before and after? I haven't figured out how to yet. It's very, like I said, this is a mad science laboratory at the moment. But I very much plan to. Very cool. That's the honest evaluation, right? I think any kind of one try eval is uninteresting. A lot of people report like pass at K or whatever, but that just means in parallel. Like as far as I'm aware, the only people that are doing really interesting long horizon evaluation work are Anthropic and Andon Labs at the Vending Bench, which is kind of what this is inspired by. And that has sort of like this long horizon reward.
1:05:32Izzy Miller:Does the hex agent right now have a tool to like make notes in real life? Not at the moment. So right now everything happens via this synthesis step after basically. I think maybe it should. But I think it's an interesting thing to consider. Like I was arguing with a coworker about this actually recently about is there value in the context agent or the synthesis agent? Like being able to see 20 threads before it saves a piece of information versus the point in time saving of information. It's like there are strong arguments both ways. There are pros and cons to both. The thing that he had implemented that I was kind of pushing on was the like synthesis step where it sees everything.
1:06:13And so it can actually like validate something or just dedupe it, etc. I think that scans, but I also, I wonder if there's some suggestion mechanism that needs to happen in real time during the thread as well. Like I sort of alluded to, like these make things very difficult to evaluate. If you are allowing the agent to dynamically change the system during the evaluation, and then you have to run these more simulation style evals, judging becomes really difficult. even just orchestrating this simulation becomes a total pain in the butt. I don't yet have all the answers here, but I do feel very strongly that the evals we do today that are just one try, I think are very interesting, very helpful.
1:06:55They help us understand agent performance, but I don't actually think they're an honest evaluation of a system like Hex. It is this platform that is designed to compound over time. So there's something missing there that I'm interested in digging at.
1:07:09Izzy Miller:What was the name of the benchmark? Vending Bench? Vending Bench is what Anthropic has built. I've called mine Metric City. Is Vending Bench open source? Vending Bench is closed source. I think it's actually done by a small company called Andon Labs, which is this weird Swedish company. I don't know if you read about like Claudius, the vending machine. Yeah, yeah. This is sort of related to the vending bench. Andon runs the vending machine too, actually. Well, it's not open source. I want to run on it. Right. I would totally run on it. too, but I'm just surprised that no one else has spent, although I haven't released mine yet.
1:07:40I hope to open source this Metric City benchmark. I'd love to try that.
1:07:44Izzy Miller:Yeah. I mean, I think it's really tied into memory. We think a lot about memory and our agent harnesses and there's, yeah, it's hard to evaluate them. And so if they're, I mean, yeah, we might make one as well, but like, I think something like this should totally exist. Thanks for listening to my conversation with Izzy Miller. This is the first episode of Max Agency, a podcast where I talk to the builders behind the agents and get into the details, the architecture, the improvement loop, and what's working or not. If you liked this episode, leave a review and subscribe. Send feedback or questions to maxagency at langchain.dev.
1:08:14Izzy Miller:We want to hear from you.
From the publisher
Izzy Miller is an AI engineer at Hex, an AI analytics platform that was one of the first companies to ship data agents to real paying users. Today, Hex runs a multi-agent system with nearly 100K tokens of tools, and Izzy is building a 90-day simulation to evaluate whether those agents actually get smarter over time. In this conversation, he walks through the harness decisions that shaped their architecture, the failure modes Hex is seeing at scale, and what it takes to build an eval that no current model can pass.
We also discuss:
- Why data agents are harder to verify than coding agents
- Under the hood of Hex’s agents
- How Hex is unifying separate agents
- Why most eval sets are bad
- The 90-day simulation for long-horizon evals
- How Izzy went from marketing to AI engineer
References:
- Andon Labs
- Anthropic
- Barry McCardel
- ChatGPT
- Claude Code
- Claude Sonnet 4.6
- DBT
- GPT-3.5 Turbo
- GPT-5.3 Codex Spark
- GPT-5.4
- Hex
- LangChain
- LangSmith
- Looker
- OpenAI
- Opus 4.6
- Satya Nadella
- Snowflake
- Vending Machine
Where to find Izzy:
Where to find Harrison:
Where to find LangChain:
Send feedback or questions to maxagency@langchain.dev
Timestamps:
(01:35) Where Hex's notebook agent started
(03:46) The moment Hex knew it was time for agents
(07:36) Why data agents are harder to verify than coding agents
(09:30) How Hex is unifying separate agents
(13:28) Under the hood of the notebook agent
(15:41) The harness features that are now holding the agent back
(17:41) Why Hex built their own orchestrator
(18:59) Managing nearly 100K tokens of tools
(20:49) Ephemeral queries and agent behavior trade-offs
(24:46) The UX problem with showing agents' thinking
(27:28) Why verification is harder than transparency for data agents
(31:00) Memory, context conflicts, and collapse modes
(34:38) How Hex built their internal eval system
(39:29) Why most eval sets are bad
(44:30) The 900% quota eval that every model fails
(46:55) Model upgrades and the "in distribution" debate
(51:34) How Izzy went from marketer to AI engineer
(59:59) The 90-day simulation for long-horizon evals




