How Listen is building a system of AI Agents & subagents for specialized tasks | Florian Juengermann, CTO

23 Apr 2026 · 48 min · 26 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

How Listen’s CTO describes building an “agent-first” system for analyzing hundreds of multimodal user interviews/surveys/focus groups using a table-like data model, subagents, and quality/evaluation loops.

Guest

Florian Juengermann, co-founder and CTO of Listen. Background: builds agents that extract “signal from noise” from large sets of qualitative research media (transcripts, video, audio).

Key claims

Listen rethought its pipeline from hard-coded workflows to an agent architecture that represents responses as a virtual table (rows=responses; columns=questions/features). The main agent creates new columns (e.g., sentiment) and spawns constrained subagents (MapReduce-like) for classification across ~500 interviews, then aggregates results. They use contextual prompt engineering, sandboxed Python for long-tail analysis (E2B), and an asynchronous “reviewer agent” to check reports against criteria (e.g., claims must be backed by citations/data). Eval runs after each analysis; live mode uses latency-optimized parameters.

Notable examples

emotional understanding from video/audio; generating charts and PowerPoint slide decks via a code-executing subagent; cutting highlight video clips; rerunning analysis at thresholds (10/20/100+) while keeping numbers verifiable via placeholders.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding Listen's Agent Structure

0:45 to 1:48

Discussion on how Listen's agents analyze data from interviews and surveys.

“We also go deep into their approach on quality control for the end report the agent writes and their eval system to ensure accuracy.”

Operational Features of Listen's Agents

1:48 to 3:19

Explaining emotional understanding and quality control in agent output.

“And it really works with the artifact on the side.”

Different Types of Agents in Use

3:19 to 5:01

Overview of various agents used by Listen and their functionalities.

“Is there some big background job that runs off after all of these?”

User Interaction with Research Agents

5:01 to 6:58

Exploring how users interact with research agents and data analysis tools.

“We can actually look, have a small model, you know, GPT-mini or Haiku or something similar.”

Agent Architecture and Workflow

6:58 to 8:27

Discussion on agent interactions, workflows, and customization options.

“It's more, we think of it more as a table.”

Python Integration and Custom Code Usage

8:27 to 10:50

Explaining how agents use Python for advanced data analysis and visualization.

“you did, 325 were positive, you know, 100 were, you know, it returns some default values that you can then use and give the answer.”

Sandboxing and Security Measures

10:50 to 13:00

Discussion about the sandbox environment for executing Python code securely.

“actually laid out differently, but you can kind of synthesize it as a table.”

Future of Agents and New Subagents

13:00 to 14:00

Exploring future developments in agents and the introduction of new subagents.

“And the way we implemented it so far is it's a sub-agent to our main agent.”

Exploring Tools and Skills in AI Agents

14:00 to 15:00

Learn about the role of tools and skills in enhancing AI agent performance.

“You know, based on the transcripts, it selects these are some interesting quotes.”

Sub-agents and Contextual Prompt Engineering

15:00 to 17:25

Discover how sub-agents are utilized for specialized tasks and feedback mechanisms.

“of like, if it's a study that has the specific data structure, then we include a specific how do you compare concepts in the prompt, those kind of things.”
Show all 26 chapters

Evaluating AI Performance and Reliability

17:25 to 19:38

Understand the importance of evaluation in improving AI agent reliability and performance.

“it should not have any claims that are not backed up by citations or by data.”

Iterating on AI Insights from Production

19:38 to 20:43

Learn how insights from production usage lead to iterative changes in AI functionality.

“and the customers, the first customers who tries it will use it in a way that you didn't anticipate.”

Architecting AI Agents for Better Performance

20:43 to 22:05

Explore the evolution and re-architecting process of AI agents for enhanced capabilities.

“Like where have you seen the distribution of things go?”

Optimizing Tool Calls and Agent Workflows

22:05 to 23:58

Find out how to streamline tool calls and workflows in AI systems for efficiency.

“I mean, I think the first iteration was just like before kind of agents really worked.”

Trace Analysis and Observability in AI

23:58 to 26:02

Learn about the importance of trace analysis for debugging and enhancing AI performance.

“Can I like call it over a subset of rows?”

Challenges and Lessons from Using Sandboxes

26:02 to 28:00

Explore the challenges faced while implementing sandboxes in AI development and the lessons learned.

“And that's usually in the development, like you just run it on your own computer or on your cloud dev setup.”

Analyzing Customer Insights

28:00 to 29:10

Explore the complexities of analyzing customer insights and knowledge sharing.

“they will be applied for setting up new projects.”

User Experience and Onboarding Agent

29:10 to 30:44

Discuss the user experience of the onboarding agent and its iterative design process.

“And maybe here we can zoom out to the three or actually even like four different agents you have.”

Voice Interface Challenges

30:44 to 33:36

Examine the challenges faced with voice interfaces in AI applications.

“And then you realize like sometimes chatting with it is not the fastest way to modify things.”

Handling Real-Time Data Analysis

33:36 to 35:52

Learn how real-time data analysis is managed and its implications on user experience.

“And oftentimes these real-time services, and last time we evaluated them at least, I think the models are now getting pretty good.”

Optimizing Data Retrieval Techniques

35:52 to 42:05

Discover the evolution of data retrieval techniques and their impact on processing.

“We do like a full analysis and kind of rethink all the hypotheses.”

Evolution of Retrieval Techniques

42:05 to 43:14

Learn about the changes in retrieval methods over time and their impact.

“So then we're actually using like another layer to summarize that.”

Challenges in Semantic Search

43:14 to 43:54

Discover the difficulties of implementing effective semantic search in natural conversations.

“Have you implemented that over your transcripts?”

Building a Product Engineering Team

43:54 to 45:18

Understand the importance of product sense in engineering for AI systems.

“We're a relatively small team, but one thing I look for in hiring engineers is kind of this product sense.”

Role of Engineers in AI Development

45:18 to 46:34

Explore the responsibilities of engineers in AI projects and the importance of end-to-end ownership.

“Do you have any non-engineers contributing to the agent, whether that's a product person or a design or some subject matter expert?”

Qualifications for AI Engineering Roles

46:34 to 47:16

Learn about the evolving expectations for engineers in AI fields and the need for relevant experience.

“I think you can pick this up on the fly?”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00We've been struggling with this for a bit where we used to have our own hard-coded pipeline, now we completely rethought it. Today, I'm talking to Florian Juengermann, co-founder and CTO of Listen. They're known for their agents that can analyze hundreds of interviews, surveys, and focus group feedback to pull out the signal from the noise.

0:16Florian Juengermann:Right now in the main agent, it's not directly a file structure. We think of it more as a table. So the table is every row is a response and every column is kind of like a question or a feature that we extract. And then the agent can basically create new columns. Florian describes how their agent works to put structure to interview responses and gather implicit signal from media-rich user conversations. We have been relying more on contextual prompt engineering. There's one feature we have, which is emotional understanding based on the video and the audio and not just the text. We also go deep into their approach on quality control for the end report the agent writes and their eval system to ensure accuracy.

0:52Florian Juengermann:Basically, we have this sub-agent, which is this reviewer agent that just knows what a good report looks like. So that's what we run it using the asynchronous runner. And then in the live runner, we use it as an evaluation system. Listen has figured out how to run agents at scale, solving some tricky problems on breaking up tasks so that they can be massively parallelized. We have this hard-coded workflow. If you call this tool once, we spawn those 500 agents and then we aggregate it in a very specific way and then it returns it. Welcome to Max Agency, the podcast that goes deep into how the best agents are being built by builders like you.

1:26Matt, listen, you guys actually have a bunch of agents that make up your platform. Would you mind talking about the different agents you have and what they do and how the user interacts and sees them?

1:35Florian Juengermann:So I think our platform is pretty broad. There's a lot of use cases and everything is now agent first. The first step is you actually create your product, you create your study. And there we have this agent that is, you know, this interactive creation agent that, you know, we call Composer. And it really works with the artifact on the side. It can modify this, but it's really like human and AI interacting in the same documents. There's some very interesting UX challenges there. And what's the artifact on the side? So the artifact is your discussion guide, which will kind of be, those are the questions we will ask in the interview.

2:04Florian Juengermann:And this then actually goes as an input to the next agent, but actually does the interview. So this is the AI that has a conversation with our interviewee and goes back and forth. And you'll do this with like thousands of candidates? So yeah, so not necessarily candidates. It's more users or customers we'll talk to, talk to hundreds or thousands in parallel. And yeah, this agent, it's a little bit less interesting from like a, you know, agent's building perspective because it doesn't have as many tools. It's more like a regular conversation, but it's also like multimodal. It has like image input and you can do screen sharing and those kind of things.

2:38Florian Juengermann:Is it voice as well or? Exactly. So that's voice based as well. Yeah. So that's the second agent. And then the third big step is on the analysis side. And that's, I think, where we spend most of our time so far. It's building this what we call research agent. And that's really like, imagine you have, you've done, you know, 500 interviews now and those exist, like you have the transcripts, you have the videos, but now you want to explore the data. You have questions about it, you can ask. And this research agent is really powerful. It can create everything from, you know, charts. It can, you know, obviously summarize things.

3:08Florian Juengermann:It can even like cut video clips for you. It can even now create PowerPoint slide decks in your own company template. And that's, I think, where we spend most of our time so far. Okay, so let's focus on that agent. So how do people interact with this? Do they chat with it? Is there some big background job that runs off after all of these? And how many interviews are there? You said 500. Is that a typical amount? Do you see hundreds of thousands? Yeah, so interaction, it's both. It is the, you know, we do run like one analysis run up front and that can take like 30 minutes. Like we use the same agent architecture for that as for the live interaction chat.

3:43Florian Juengermann:there's some different parameters we want to optimize for latency versus you know just quality what are those different parameters yeah I mean sometimes use different models so for some like live calls we use like faster smaller models or we tune thinking parameters so you know in the live run we have like minimal thinking whereas in the long run we have like medium or even more thinking and then there's some there's some things of like if you have more than 500 you have like thousands of interviews you can't always look at all of the interviews live because you know rate limits and and so on even if you spawn a lot of subagent there's some limits to that versus in the asynchronous workflow you can actually look at most of them so then we're doing some things like maybe we just sub sample some and we show you some results earlier so there's some slight differences but the same architecture um the same kind of agent tools overall for both of those tasks and for the asynchronous run like how do you decide what to kick it off with?

4:35Will the user set that ahead of time or is that a standard template that you guys have that you run?

4:40Florian Juengermann:Yeah, we know kind of what works best. So there's a standard template to it, but it's also something you can customize on your kind of organization level. So if you have a specific format, if you have specific background information, you can input that before and then it will take that into account. I mean, I think there's like different outputs. Sometimes you want to have like a long written document. Sometimes you want to have like multiple documents or multiple sections for maybe you have something that's multi and also like a study that's run in multi-country and you want to have like a comparison between those countries you can specify that up front and we can have a tailored report to you but again if you don't have what you need you can also just chat with uh with it after and how do you specify that up front is that in natural language or are there some boxes to check it's all natural language natural language so you write basically a paragraph of hey i want a detailed research report there should be one for each country yeah exactly and the ai is really good at understanding i mean we used to have some some more rigid things but i think we've all it's all gone to natural language now okay so what does this agent look like under the hood what is it doing and if there's differences between the live and the more long-running one maybe let's focus on the long-running one specifically but it sounds like they're pretty similar yeah no overall they're pretty similar i mean how does the agent work underneath um to actually build their own harness which is something you can talk about and it basically has access to these you know the main level is the transcripts of these 500 conversations and then the goal is to you know really make that understandable and it has you know a bunch of different tools we're in this interesting domain where it's not infeasible to look at all of these individual interviews again at least with like cheap lm we're not in like the millions of conversations where that's infeasible so we have some tools that can do like a you know more recursive summarization or even a classification of like okay if you want to know how many interviews mention this specific thing.

6:27Florian Juengermann:We can actually look, have a small model, you know, GPT-mini or Haiku or something similar. Actually look at all the interviews, classify it, and you actually get robust quantitative data out of these very open-end conversations. So how are these files or how are these interviews presented to the LLM to the agent? Are they presented as files? Are they presented as variables in some programmatic environment? Yeah, so we've been iterating a little bit on that and we're constantly thinking about if we should change it to more file structure Right now in the main agent, it's not directly a file structure.

6:58Florian Juengermann:It's more, we think of it more as a table. So the table is every row is a response. And every column is kind of like a question or a feature that we extract. And then the agent can basically create new columns. So it can say something like, you know, what is the user sentiment towards this specific topic? And then please, it could be like an open-ended, like a summary of the user sentiment, or it could be a categorical value, like, you know, positive, negative, neutral. and then it will basically add this column into this, you know, table and, you know, fill in the values for each one of them.

7:29Florian Juengermann:And then you can use things like Python or things like, you know, like to chart the data basically based on that. Super interesting. And so, and when it fills in the values for each row, it kicks off basically a small subagent or a small L, are those small LLM, small subagent, are those the same things? Yeah, you could call it a subagent, but it's really like a very constrained agent. It doesn't have to have, you know, room to decide what to do. It just does like a classification call. So that's one of the tools it has. It's basically like this. You can think of it more as like a MapReduce call.

7:59Florian Juengermann:And it's something that we've hardcoded as one of the tools. So you can call it subagent or you can call it just LLM. So it's got this table. It's got a row for each of the transcripts, each of the interviews. It creates different columns. You mentioned like Python or plotting things. Like does it also have access to code? Can it write code? So by default, we give it some, let's say you call this, you know, this classification tool, like with the sentiment, right? And by default, this tool actually returns a statistic on, you know, out of the 500 interviews you did, 325 were positive, you know, 100 were, you know, it returns some default values that you can then use and give the answer.

8:37Florian Juengermann:Most of the time, that's actually enough. Or it can use that, you know, we have some, you know, sophisticated way of creating charts based on columns. So you can say, please create a column, a chart based on this column or a chart based on this column segmented by this column. So we have some like logic should create charts and visualizations out of the box. But then obviously there's like a long tail of unlimited things people would want to do. And that's where the Python really comes into play. So you can also write custom Python code. And it's a little bit, you know, it takes more iteration.

9:05Florian Juengermann:It will not look as nice because it's not in our UI, but it can do more sophisticated, you know statistical analysis great very custom chart types that maybe we don't support in our platform and then compute either compute values or even create specific charts that we then return and show to the user so like a mat plus lib or something like that what percent of queries do you find need this more open-ended long tail kind of like just raw python when we see it using a lot of python it's sometimes worth to give us give it a specialized tool to do that task so i think the numbers have been going up and down.

9:41Florian Juengermann:The long-running default reports, it actually doesn't use Python that often, maybe in like 20 % of the times. So you would think it's maybe not as important, but actually it is super important even if it's still 20 % of the time because we have this dynamic engine as well, which is, I have a specific question and the given tools don't support that. So I fall back to Python and actually run their own analysis. And that's very powerful. So even if it's not in the majority of cases, it's very powerful to have it. So you can drill down if you want to. So if you're writing this code, does this code run in a sandbox?

10:10How are you guys managing that?

10:12Florian Juengermann:Yeah, so this code runs in a sandbox. We're using E2B for executing that code. There's some challenges on kind of making sure we get all the data because it's not tons of data, but it's still like a decent amount of data to spin up the sandbox, load the data fast enough, especially in the live view. So we do some like pre-warming and like setting up the sandbox ahead of time and those kind of things. So you have the sandbox separate from the agent. So the agent's kind of like running, doing its own stuff. It's got this table-like thing, which is, yeah, actually, is that a real table under the hood?

10:44Like, is it?

10:44Florian Juengermann:It's purely like a representation. So I think the agent, we tell the agent there's a table, but, you know, in our database, it's actually laid out differently, but you can kind of synthesize it as a table. And then if it actually runs the Python, then it basically gets this pandas data frame, which is basically this table. But it's never stored as like a table, like a CSV file or something. is just stored in our Postgres in a different format. Yeah, we've been experimenting with some stuff like virtual file systems, and it seems like this is a virtual table. Exactly, exactly. And so you've got this agent running, it's got this table-like thing, and then it calls this tool, and then that spins up a sandbox.

11:20So it's not like the agent's always running in the sandbox. It's got like a sandbox as a tool.

11:24Florian Juengermann:In this approach, exactly. So most of the time, it can actually just run in our backend, and we don't need the sandbox because it's more hard-coded regular tools. And then only for the SEC backup, we need to go to Python. We also do have some newer agents or subagents that are using the sandbox more natively and definitely thinking about where's the future going. It definitely seems like the future is going to more, you know, agents that have it, you know, are built based on code and run continuously. So it's definitely always something we're experimenting with. What are some of those newer subagents?

11:55Florian Juengermann:Yeah, so one thing, and we actually just posted a blog post about this, is the PowerPoint generation. So we've been working on the, you know, again, if you could take this stack back, Like for our customers, oftentimes the end result is often like a PowerPoint slide deck they want to present. And PowerPoint is like weird format where it's not, you could just create like an HTML, like React page that looks like slides, and then you can create a PDF out of that. And that's great, but you can't really edit it. Like our customers can't edit it. They want to have something in PowerPoint. So we've been struggling with this for a bit where we used to have our own hardcoded pipeline where we create templates.

12:31Florian Juengermann:and then we basically have like text replace in this PowerPoint slide decks and image replace and find the right templates and so on. So we used to build that pipeline. Now we completely rethought it as we've seen the agents being able to use tool calling and especially code generation to create PowerPoint. So basically what we have now is this Cloud Code agent SDK that runs and writes Python code and can then modify the PowerPoint file and kind of do a lot of iterations in that. And the way we implemented it so far is it's a sub-agent to our main agent. So the main agent decides, okay, I want to create a slide deck.

13:07Florian Juengermann:It should have this content. It gives all the data. And then the sub-agent kind of iterates on the code level with a specific skill of how to create a slide and then returns the PowerPoint in the end. And that's this agent that actually runs in another sandbox. It sounds like definitely the sub-agent, maybe the main agent, they can create artifacts in addition to giving the final response. response how how does that work yeah i mean artifact in some ways it's just you know json output of a specific thing so if it's like the basic artifact we have is a chart which is again in our virtual table we basically say create a chart like it's just a json tool call create a chart with you know this column and this column and then we render the chart to the output there's some other artifacts you can also create like memos or like you know like basically like a written output that combines charts combines other things can even create things like you know cut video clips from those interviews together.

14:01Florian Juengermann:You know, based on the transcripts, it selects these are some interesting quotes. You have some, like, retrieval pipeline for that, and then it wraps that in a specific tool call, and then we create this wheel, and it can reference that in the output, and then we render the wheel. You mentioned tools and sub-agents already. Do you also use skills? Is that a concept that's made its way into these agents yet? As you can tell, like, really getting into, like, the longer tail of use cases, where we've been struggling with, you know, the model's getting smarter and the prompts are getting longer, and That seems to kind of hold the balance.

14:31Florian Juengermann:But if you want to go into the more detailed rare instances where we know this is how it's supposed to work, but the agent maybe doesn't want to, we don't want the agent to reinvent it every time, then we're using skill. I think in this specific instance, we haven't used skills that much yet, but we have been relying more on contextual prompt engineering, I would call it, where like if you want to use a specific tool, or I guess there are two things. We have the sub-agents, which we rely pretty heavily on, and then the contextual things of like, if it's a study that has the specific data structure, then we include a specific how do you compare concepts in the prompt, those kind of things.

15:08Florian Juengermann:And then we rely on sub-agents. So for example, we don't have the instructions how to create the PowerPoint in the main agent, but we call the sub-agent and that has basically the skill loaded, like preloaded by default with all the context. We have some other agents that maybe do use skills a little bit more based on what you do. And I think the real opportunity is, can you reduce the context? Can you make it more dynamic? But it also comes with some challenges. For the contextual prompt engineering, just to make sure I understand. So that's basically like, you'll look at the study that was run, and based on those properties, you'll just insert different things into the prompt.

15:41And so the system prompt isn't really the same system prompt for every study or every agent. It kind of varies a little bit depending on the study.

15:49Florian Juengermann:Yeah, exactly. I mean, I think the main thing is just if you can cut out like a big chunk or like there's there's one one feature we have which is emotional understanding based on the video and the audio and not just the text but then sometimes some of our studies don't have video and audio right and then we can remove all of these instructions from the prompt and then just have it less be less confusing again there's no magic trick because in in the worst case you have all of the cases in there and you don't save anything so it doesn't really help you in that case but oftentimes we see that you know it can help a little bit how How many different tools and how many different subagents do you guys have?

16:22Florian Juengermann:For this research agent, it's probably, it's more than you think. It's probably like 15 tools or something like that. And are most of these tools like running something over the table, like a classification or no? There's probably one or two that run everything across the table. There's maybe one or two different retrieval modes. There's then maybe one for creating PowerPoints. There's one for creating highlight reels. There's one for outputting a specific chart in a specific way. this one for creating like a different artifact. So I think most of them are either output related or like processing related, like either computes new data or create some artifact that you can then use in the artifact.

17:00Florian Juengermann:And one other thing we also have is this, which is pretty interesting, is this feedback tool, which especially for the long running task, you can't really do it in the live one, but in the long running task, it can then self-request feedback. And so it's basically we have this sub-agent, which is this reviewer agent that just has a clear context, doesn't have all the complicated instructions on, you know, and not necessarily sees all the history, but just knows what a good report looks like. And then has a list of criterias, like, okay, it should not have any claims that are not backed up by citations or by data.

17:31Florian Juengermann:And then it kind of goes through the report and gives feedback. And that's actually the loop that, you know, runs, you know, you can run it quite often and it will actually catch a lot of things and will make things much better. So that's what we run it using the asynchronous runner. now. And then in the live runner, we use it as an evaluation system. So basically, I mean, we have to make some slight adjustments, but basically you want to know how many issues does this evaluation agent find in the report? And can you, you know, at least the first step is just at least know how bad it is or what the common issues are.

18:03Florian Juengermann:And then if you make a change to the problem to the model, to the architecture, or maybe some change you think is not even related, you'll see if it actually increases things either on our benchmark set or even in production of like, if there's a spike in a specific problem. When will you run this eval over the live things? Is it like offline evals that you run before production? Are you also running it over production data? And like, if so, is that happening at the same time as it's chatting? Or like at the end of the day, you do a big cron job and score a bunch of things. I mean, I think that's something where we're still optimizing right now.

18:32Florian Juengermann:We just kick that off after each run in like an asynchronous runner. So it doesn't like block things, but it just gets evaluated live. I mean, it's probably some cost optimization we can do on batch inference, but hasn't been the biggest priority right now. But it's just for us to kind of have internal metrics. And we might also move to subsampling, but the cost is not the biggest concern for us right now. I think it's really, can we make the product better? And yeah, it's not the biggest concern right now. What does make the product better mean to you? Is it just increasing accuracy on the existing things?

19:02Is it spreading out to more challenges and more domains?

19:05Florian Juengermann:It's something we also thought about a bit recently because in some ways, if you have the signal of these are all the things that are broken, And obviously you want to fix them and make progress on those. And that's kind of improving the reliability. That's very important. But sometimes you also want to take a step back and then maybe just optimizing this local minimum. And maybe sometimes you actually want to, you know, change the architecture overall or change the approach overall. And then maybe the criteria is even change. This hill climbing is very important because otherwise you ship a product that is aspirationally great, but then in practice doesn't work.

19:37Florian Juengermann:Again, you can only test it in so many cases yourself. and the customers, the first customers who tries it will use it in a way that you didn't anticipate. So I think it is very important to try those evals in production. Are there any fun stories of how people were using it in ways that you didn't anticipate when you first launched it or that you're discovering now? There's a bunch of things. People definitely very soon try to like, not break it in like a prompt injection way, but like really test the limits of like, can you do like a cluster analysis? Can you do like, can you run like a, you know, classification model in the Python and so on and then You know, usually the Python leverage image we have is limited.

Read the full transcript

20:15Florian Juengermann:So it can't do everything. And then it tries to like create, code things up from scratch. Like, okay, sure, let me write like a classification algorithm like myself in my own Python environment because it doesn't have access to those tools. And then we're like, okay, maybe we should give it some more tools. So it doesn't have to do those things manually. It's always interesting and it happens much sooner than you expect. When you're like iterating on these insights from production usage, are the most common things you're iterating on, like changing the prompt or giving it more tools or updating like the environment that it's running code in?

20:45Like where have you seen the distribution of things go?

20:48Florian Juengermann:I think obviously the updating the prompt is the lowest lift. The hard part is, you know, if you change the prompt, it hopefully improves the thing that you want to improve it on. Does it then, especially if you like change it to all caps and repeat it three times, does it then mean something else regresses, right? Because there's less attention to some some other introductions. I think that's the hard part. How do you gain confidence in that right now? Yeah, I don't think we have a good solution. I think obviously as we deploy it and sometimes we roll it out slowly and see other number of issues.

21:18Florian Juengermann:But I mean, we have like an example test set of like maybe 20 different products. We can always run it on. We can evaluate. We can see at least in that small sample size is anything. And that's what we actually manually look through the outputs and see, okay, it still seems good. But it's really hard to compare one five-page output with another five-page output. They're different. People would probably argue this one is better. Some people would argue this one's better. Like, I don't know. But at least you will see if something clearly breaks. How many times, if at all, have you guys completely re-architected this agent?

21:48It sounds like maybe you're thinking of that now or starting to think of that.

21:51Florian Juengermann:I've done it a bunch of times. I mean, over the last two years, I would say. And what's been the main impetus? Is the models just getting better? or I don't know, you hear a great talk from someone at LinkChain maybe or something like that. But no, like what's been the impetus for the architecting things? I mean, I think the first iteration was just like before kind of agents really worked. Like the first version was just like a simple rack bot, which I guess probably most companies start out doing and that's still a primitive that we have. But I think both like smaller models getting good enough that you can actually do, like look at all the data, classify things live, that's going to change.

22:26Florian Juengermann:and then the big models getting smart enough that it can really orchestrate those things. It's been the biggest change. And I don't think it was like a single thing. I think we've been talking about this a lot and it's been a big investment on our side to re-architect this. And we definitely did a lot of explorations before trying it out. And yeah, now, I mean, obviously, you know, most people are moving towards like a file system-based agent, something we have explored, you know, a couple months ago. And there it felt like, especially in our specific use case, it wasn't they kind of still want to have these specific tools is that largely because of like the table that you guys have and the ability to quickly run over exactly the table is pretty powerful but there are arguments against that especially if like the model companies start to you know post uh you know post train on the specific harness or the specific tool calls it feels like you're kind of fighting against that movement if you have your own harness i think right now all these like coding agent harnesses are pretty bad at like calling subagents programmatically like they can call them one at a time but what you guys really want it sounds like is to basically call it 500 times and have that always happen and this is similar to some of the stuff in like the recursive language model paper but i don't think any of we're thinking a lot about this how can we take those ideas and put them into like these coding agent harnesses and there's some stuff we're thinking about but i don't think we or anyone has really nailed that yeah exactly i mean you know obviously the way we do it is we don't actually the model doesn't actually call it 500 at a time.

23:50Florian Juengermann:So it just calls it once and then we have this hard-coded workflow. If you call this tool once, we spawn those 500 agents and then we aggregate it in a very specific way and then it returns it. Can I like call it over a subset of rows? Could it pass in some like filter criteria to filter things out? Exactly. So that's one of the use cases. And maybe also coming back to your question about what have we changed? Like it's prompt changes, but then something like, oh, oftentimes it would have to filter. And what it used to do is run it on everything and then write a Python script to filter out. Instead of giving the aggregate results of all of the data, it will just with Python filter out the aggregates of those.

24:22Florian Juengermann:We're like, okay, this seems a little bit cumbersome. Like, let's add the specific additional field in the tool call of like filter column, right? Then you said this column equals this field to filter. It works in like 90 % of the filter cases. There might still be some advanced filters based on like a combination of columns. It might need to write a Python script one, but that helps a lot. So that's one of the other things that we obviously see as we deploy those. And that's, but that's really something you have to look at traces yourself. You have to have good observability. You have to really go deep and see what did it do?

24:50Florian Juengermann:Does it actually make sense? Sometimes even look at the reasoning traces of the model, like why did it call this thing? And then see, oh, I wish I had this tool more or less. And you're like, okay, maybe I should give it that tool. What does trace analysis, trace observability look like for you guys? How do people do it? Who's doing it? Is it everyone on the team? Do you have like specific people who are focused on it? Imagine you guys have millions of traces. How do you know which traces to look at? And then when they find something, what do they do with it? we want to trace every single trace, like we don't want a subsample, but then we only look at the 0.01 % or less.

25:26Florian Juengermann:But it's like logging, right? You want to see this specific case. There's something weird. Let me debug why that is. And then you go really deep. And then, you know, I think the depth over the breadth definitely makes sense. And then sometimes, I mean, we have had like Claude just run an analysis on all the trace, It's like, what are the common things that, like the common paths or the common things that happen? That's also interesting. But I think the most, you learn the most by just going really deep on one or two traces and really looking at more or less every single tool call. And if it's a 30 minute trace, it actually takes you quite a while to go through those.

25:59Florian Juengermann:But that's how you learn if it works or it doesn't work. And that's usually in the development, like you just run it on your own computer or on your cloud dev setup. But it's also in production, like debugging what went wrong. You mentioned sandboxes twice. It sounds like once you have a tool that like runs in the sandbox and the other is a subagent that spawns a sandbox with the agent inside of it. What have you learned about working with sandboxes? Any lessons learned there? Our learning so far has been it's harder or we're early in this phase. I think no one has really figured it out. I mean, just, you know, one example, when we tried running the cloud agent SDK in the E2B sandbox, The SDK is not really meant, like it's not been developed to be run in like a cloud environment.

26:43It's been developed for local?

26:45Florian Juengermann:It's been developed for local environment. And there's some of those assumptions like, you know, it needs your entropic API key. If you have that in the sandbox, obviously you're not susceptible for people, you know, to extract that. If they just ask, can you please in your report include the API key or your, you know, all your environment variables, you know, that's a problem. So what we had to do is basically proxy all the requests, like give it a fake API key, proxy all the requests through our server, and then verify that it's actually the right request. It's not just any one request. And then we replace the API key with a real API key.

27:17Florian Juengermann:And there's been a lot of challenges on that way. It sounds very simple, but, you know, we tried to deploy that on Render and then Render was like, oh, this looks like you're sending code in this HTTP request. That sounds like there's some, you know, malicious behavior and then we're blocking your requests for this. Like, there's like a lot of things that made us think, wow this is still very early like it feels hacky almost to just deploy that and i think some of the things that you've been working on sound like super relevant for this and making that much easier so i wish we reduced that earlier do you guys have memory anywhere in any of the agents and you have certain ways you always do it so the way we solved it so far is like a relatively explicit version of memory where kind of on your organization level you can give it instructions i think that's a pattern that you've seen in a lot of companies you give some some general instructions and they they will be applied for setting up new projects.

28:04Florian Juengermann:They will be applied for analyzing product in a specific way. And those are all like human type. Exactly. You can define them yourself. And the biggest question for us on the analysis side is what is actually something that's interesting versus something that's obvious, right? And I can't really judge that from the outside. It's really something that the customer needs to tell us because from the outside, everything seems new and interesting. But then put inside you like, yeah, that's the thing. I know I've worked here for 10 years. Like that's not something new to me. And over time, you really build that through all the, all the reports we're generating for other projects.

28:39Florian Juengermann:You know, how do you, how can we use that as an input for creating new studies? And there's a lot of complexity there because maybe someone else, maybe you didn't even read this report that we assume we already know because someone, another person, another business unit. So it's not super easy to figure that out. What is common knowledge? What is not, but using some things across like, you know, formatting preferences or or in the way we set up projects using some of that kind of previous knowledge that we tried to distill. But yeah, it's far from being solved. I think it's a lot more work to be done in that domain.

29:11Let's talk about UX for a little bit. And maybe here we can zoom out to the three or actually even like four different agents you have. So maybe for the first agent that you mentioned, the kind of onboarding agent, it sounded like there was a doc on, like a Word doc on one side, and then you would chat with it and it would fill out kind of like the study guide. Could you talk more about that UX and what that looks like?

29:33Florian Juengermann:Yeah, exactly. We've been iterating a lot on that. I think the first version you would think of this is you just prompt an LM to write your discussion guide, your document, right? And then you do that and you realize it never gets it right 100%. Not to blame the model. I think the model is great. But just like you give it one sentence, you expect it to write like a whole page out of it. Like that's probably not going to work. So there needs to be some interaction. and then basically the second version we built is can we just have the ai can you just chat with it right and you can kind of make modifications to that right and then you have the two basic principles that you have to decide between it's like either the alm rewrites your entire document and that works for shorter documents or it works if you're actually making some like changes like can you change the tone or something like that or do you say like you have some kind of edit functionality that just either string replaces or you have some ids that you replace only specific IDs, which usually works better, but sometimes can also be confusing.

30:28Florian Juengermann:If you're making a lot of changes, you stack them and you say like ID number two is now ID number three. And then you insert a new element here and then like, it's also not perfect, but that's kind of the approach that we picked. Can we just make, because our document's pretty long, can we just make targeted changes and so on? That's pretty cool. And then you realize like sometimes chatting with it is not the fastest way to modify things. Sometimes you just want to delete that and then telling the AI, please delete question number three. feels a little bit cumbersome or sometimes you just want to reformulate it yourself so you do you want to have a way of also manually modifying it and you know open ai and entropic they have some version of that i don't think the ux is like supernatural and they've also iterated a lot on those i've seen like they rolled back some of the changes they did and so on i think we have a pretty good solution now where you're both working on the same edit history so you can actually undo undo changes and you kind of have this change log you can compare things and you can manually make changes and then you can also make changes with the chat and it kind of knows about the changes you made, so it doesn't undo the changes immediately.

31:28Florian Juengermann:How does it know about those changes? Are those inserted? How does it know about those? Basically, the way it works is every change we make is formatted as an edit operation. So you see a log of all the edit operations. So then it knows, okay, you just modify this question. Because the main problem is you can't always have, like, this is the old document, this is the new document, or somehow you need to modify the div. So we kind of formalize that in an edit operations way, and then the model knows, okay, if you just touch that or if in the history you see that you haven't touched that and you probably don't want to rewrite that.

32:00Because those are past. So if I'm chatting with it, it gets a document, I go into the document, edit something, and then I chat with it again, that edit is passed in prior to my message in some way?

32:11Florian Juengermann:It's the same. The LM writes out edit operations and the human edits also edit operations that kind of fit into the same kind of message history. So that's the approach that worked pretty well for us. Cool. So that's the onboarding agent. then there's the interviewer agent. And that sounds like you've got voice there. It's multimodal. It's more just like a chatbot style thing. Anything interesting there that you guys have been playing with? The interface, like the voice interface is still not quite solved. And we've seen even like OpenAI, like the ChatGPT app, it's been going back and forth, right?

32:43Florian Juengermann:They used to have this like blue bubble level speaking. And now that their voice mode is actually, you see the text and it writes out text as well. That's kind of the approach that we have been taking for a while as well, because you can actually read text faster than you can listen to it and sometimes can actually be annoying. You're like, you know, I wanted 2x speed or something. But then sometimes 2x speed is too fast and I actually want to go back. So I don't think it's fully solved yet. And same time for our use case, we really don't want to interrupt people in almost all cases. We just want to, we're a listen company, right?

33:11Florian Juengermann:We want to listen to the customers. We don't want to interrupt them. Even if they may be saying something that's, you know, rambling or maybe going slightly off tangent, oftentimes there's a reason for that. And in only very rare cases, we actually want to interrupt. So even if I take a break for a second to think about something, I actually don't want the eye to jump in. And are you guys using kind of like the speech-to-text, text-to-speech sandwich or using the real-time APIs? Right now we're mainly using the speech-to-text, text-to-speech pipeline just because it's so important to have the smartest models.

33:42Florian Juengermann:And oftentimes these real-time services, and last time we evaluated them at least, I think the models are now getting pretty good. but at least when we evaluated them, they were like one or two tiers, you know, faster and dumber than the, you know, top tier Opus and so on models that it is so important to ask. It sounds very simple to just have a conversation, but then asking the right questions. I mean, I guess that's what you're doing today. It is actually a pretty hard task. And that's why we don't want to compromise on that and rather compromise a little bit on the, you know, real-time aspect of it.

34:14Yeah, you've got the fast and dumb interviewer here today. I feel like most people, myself included, have largely stayed in kind of just like the text domain of agents. When you think about adding on voice, like how much extra work is that? Is it easy? Is it hard?

34:30Florian Juengermann:It depends. I think the hard part is not necessarily the AI. I think it's more like, you know, we're collecting. For us, it's the interview runs on like millions of people's devices, right? And the more modalities you have, the harder it becomes from like a compatibility perspective. and you can have to see all kinds of issues on like somehow the microphone stopped working or you know how like everyone used to have like trouble with like zoom like microphone not getting recognized and those kind of things so we see all of that across the globe right globally so i think those are more the challenges with the multimodality for us than like on the ai side i think the ai models are getting pretty good i think there's still some challenges on you know transcription you'd think that it's a solved problem and you know models are getting pretty good but there's still something like you release a new product uh tomorrow the transcription model doesn't know about you know that name and it probably uses some other name and you know the only way we could you know solve that for now is actually having an lm that has a context on the interview and maybe even know some of the terms that might come up correct the transcription model basically on the fly just to give this additional like smartness that the transcription model itself can't can't use those are some of the challenges and there's obviously like traditional things of like now he's storing like you know thousands of hours of video data and you should you know do that but those are more like traditional infrastructure problems and then going on to ux of the final research agent in in the two different modes it has so when it's running kind of like long in the background how long does that take and do you kick one off how do you kick one off does it happen like automatically after all 500 interviews are done yeah that's that's a big it's a big challenge so again it runs for like 30 minutes and so on so the cost is like you know significant it's nothing crazy and you know we're usually like a higher priced uh offering so it's not the biggest concern but you don't want to re-kick it like after every time there's a new response or someone updates and you know we have 500 responses at the same time we do want to show your results early right you get the first 10 responses that that's a magical moment and that might happen like 30 minutes after you launch the study you order to get the first 10 people to respond which is very magical so we do want to give you something um the way we do it right now is we run it at certain thresholds of interviews.

36:43Florian Juengermann:We do like a full analysis and kind of rethink all the hypotheses. And, you know, if the data has changed from 10 to 20 or from 20 to 100 interviews, we actually want to completely rerun this. And that's one part. But then now you have, we run it after 100 interviews and now then 100 and first comes in, 100 seconds come in. Like you don't want to rerun everything. But at the same time, we have like numbers in our report, right? We have percentages, we have charts and all these kinds of things. And those things we constructed in a way that we can actually replace those things. So the LM never outputs those numbers.

37:16Florian Juengermann:It only outputs placeholders. And then we can run all the classifications, all the Python code again, but keep the core thing the same. So this way, our numbers are always verifiable and updated. So you can always click on them. You can see the data that's backed up and there's a new response coming in. Or you remove one response because maybe you don't like them or it's like low quality. Those things will update immediately. Obviously, there's like a limit to that. But if you say like, if in the text it says like, this is definitely the best idea because 80 % of people liked it and then more and more responses come in and suddenly it's no longer 80%, it's only 20%, then actually the qualitative takeaways change so that you do need to run it occasionally.

37:54Florian Juengermann:But yeah, that's one of the things we work with. And at the end, you usually, in our use case, you usually say, okay, now I'm done or you haven't seen any new responses for two days and then we run like a full new analysis. But I think the real-time component, And I really believe in this delayed gratification, like the faster you get your results, like the more, the better the user experience is. So we really want to do that. Interesting. And then for the real time kind of like asking, if you ask questions, how does that work? Because I imagine it's doing a bunch of tool calling under the hood.

38:23Do you surface those to the user? Do you hide those? Like how transparent are you about the agent's work that it's doing?

38:30Florian Juengermann:Right now we do show like an abstracted version. We don't, like our customers don't really care about what happens under the hood too much, but they do want to see something's happening and maybe things are like, okay, I'm actually now looking at all the responses again and that's why it might take a while and we have a loading state, those kind of things. The more tricky thing there is, what happens if it now starts writing Python code? Because in some way, if you ask it a complicated question and says, you know, the answer is 42, you're like, okay, I guess. Usually all of our findings are very traceable.

39:02Florian Juengermann:Like you can, as I said, you can click on the numbers, You can see the breakdown. You can even go down to the individual level of this. Everyone we classified and why we did it and can really kind of explore the data. But the Python, that's no longer the case. And at the same time, you know, no one wants to read Python. And in our case, our customers will probably also not understand the Python. So can we make it, can we build the confidence that it's the right answer? And obviously there's always been, there's always assumptions going into that. The instructions are never completely clear. So, you know, basically we're currently summarizing exactly the assumptions we're taking when we're executing Python in like a little box that you can expand.

39:42Florian Juengermann:And if you want to look at that, but we don't show you the raw Python script, which is something we're thinking about. It's kind of in the middle ground. It's not perfect, but it gives you some confidence that what it does is actually right. And how are you summarizing? Is that another LLM call that's like looking at the trace so far and generating something? Exactly. So after the Python code is run, it will, like, first we show it as a message, just like, you know, running Python. I don't think we're actually saying that anymore. We're saying, like, running some computations or something. And then after it's written the script and while it's executing the script, we actually summarize it to have that text.

40:14Florian Juengermann:And that's more, it's less for, like, the status of what it's doing right now, and it's more for, you know, where does this come from? Let me go deeper and look at where it comes from. When you guys are running, like, these subagents or small LMs over the 500 documents, what What if the documents are like really big or really massive? Do you do any chunking there and like further kind of like subsetting the text and chunking into three things and then running a small LLM overall three things? Or do you always just treat it as one big thing? Yeah. So there's multiple layers. Our interviews are what we call semi-structured.

40:46Florian Juengermann:So we have a rough idea of what people are talking about. So there's different sections that, you know, in this section we'll talk about, you know, this concept, in this section we talk about this concept, or we're talking about different ideas. And so basically the interviews are annotated that way and we can filter to a specific relevant section. So that helps. And that filter would be part of the filter that you pass into like the... The main agent would decide, you know, for this question, we don't need to look at the entire interview. We just need to look at their background information or something like that.

41:15Florian Juengermann:And then we just cut it to that section. It's not perfect, right? Sometimes people might say like, oh, by the way, I forgot like the very different setting. I was like, oh, I actually forgot. I changed my mind, whatever. So again, not perfect, but I think that's a pretty reasonable assumption. We do use some chunking and some retrieval for if you ask a question like, oh, did anyone mention something like this? Or can you find clips where people talk about this specific topic? Then, of course, in theory, you could run this, you know, map reduce function call over all of them. But in practice, it's usually faster, especially if you're going to like the thousands and tens of thousands interviews to use a retrieval, like a semantic search on chunking.

41:51Florian Juengermann:We do some hierarchical summarizations as well for this extraction step, because sometimes if you imagine doing like a thousand summaries on individual interviews, that's still a thousand times maybe 200 tokens. That's still a lot of text. So then we're actually using like another layer to summarize that. Has your use of retrieval changed over time? So it has changed a lot. We're still using retrieval. Again, the first version of the research agent we built like two years ago, it was just a rack. just a semantic search and then the second thing we added was can we add some robust filters? Can we just, if you ask, can we just filter for man, like we extract metadata, we filter based on that.

42:30Florian Juengermann:That was the first version. The second version we were like, okay, we saw all the problems and we actually completely moved away from retrieval. We were only doing these like small elements looking at everything. We didn't have any retrieval pipeline there at all. But then we realized for some cases, again, especially if you're scaling to larger samples, it's still useful and it's sometimes faster and in some cases it's better. But it's definitely less critical than it used to be. But even if it's not that critical, it still means we embed everything. And the cost is not prohibitive. So we just do that.

42:58Florian Juengermann:And especially if we're working towards a platform where you can search through all of your findings, not scope to a specific project, but kind of over time. And then, of course, the data becomes much like a much larger corpus. And I think retrieval will continue to be important. In the coding world, people now just use keyword search or grab. Have you implemented that over your transcripts? We haven't really implemented that. I think the main thing is, in code I guess you have symbols that are like types like strongly typed and you can actually search for exactly that string. In like a natural conversation, it's much harder.

43:33Florian Juengermann:If you want to think, if you want to retrieve all cases where people are frustrated, that's maybe a semantic search but it's like really hard to, sure you can search for a list of adjectives that people might have used but it's really hard. So in our use case we've been sticking with the regular retrieval. On a completely different note, what does the team that does all this agent engineering look like for you guys? It's engineers right now. We're a relatively small team, but one thing I look for in hiring engineers is kind of this product sense. Because I think there's maybe two types of engineers that do well in this world.

44:07Florian Juengermann:I think there's one that is really good at building large-scale systems and have seen that, have a good taste, and maybe there's something that at LEMS, at least right now, can't do super well. and then the other side is these you know product engineers that really you understand the customer iterate fast and you know try out things I think it's super hard to whiteboard this is how the agent's going to work and you know these are going to be the problems and this is how it's working you need some part of that you need to have some some idea of where you want to go but then you need to try out things how reliable is it you need to adapt and for that I think it's very important that the engineer itself is the one that that is evaluating that is talking to the customer and and hearing how it's going, looking at the logs and the traces themselves.

44:49Florian Juengermann:You mentioned, you see, is someone else looking at the logs? No, it's actually the person that built the system that's looking at the logs and trying to understand, does the model do what I want it to do? And I don't really believe in people that just prompt, like just write the prompt, because again, it's a very nuanced thing of, if you change the prompt, then you also need to change the tools. And then for the tools, you need to understand, you know, some of the infrastructure and so on. So I think this end-to-end ownership is how we've been building the product. And I think that's where I think it's going to stay the same, even as we grow the team.

45:22Do you have any non-engineers contributing to the agent, whether that's a product person or a design or some subject matter expert? I don't know.

45:32Florian Juengermann:Not directly. So I think the problem I see is it's relatively easy to change something like the prompt. They're like, oh, yeah, just add the sentence to the prompt, you know, and people definitely want to do that. But then from my perspective, I'm like, okay, sure, it's easy to change it. But what about the validation? Who's actually going to take the blame? And if it breaks, who's actually going to fix it? Right now, like the number of PRs we have in our organization has exploded. It's so easy to write code and everybody's like, oh, can I get into the code and write code? But, you know, I don't want the engineers to just be the ones that, you know, review code and improve changes.

46:05Florian Juengermann:And then later on, the ones that fixing it. I think that's a little bit of the problem that I see. And we're obviously working closely with, you know, customer facing people. and getting feedback. But then it's typically the engineer that will consolidate with that with all the other requirements and try to improve that and then get feedback from them again. Do you care if people joining the agent engineering team or the team that works on it, do they have to have previous kind of like AI or agent experience? Or is that something that, hey, if you're like a good software engineer, I think you can pick this up on the fly?

46:36Florian Juengermann:I changed my mind a little bit on that. Maybe like a year ago, I was like, we just found very smart people. And if they haven't worked with AI, I think that's fine. They can learn that. Now, I think we're at the time where like, it's a bit strange if you've never worked with AI. It's been like three and a half years. Exactly, right? Or like, or even if you haven't worked on your job because like for whatever reason, you're a company, you know, you're not doing it. Like you should at least be intellectually interested and curious about how does it work behind the scenes and build something on the side or kind of at least know what is Cloud doing if I'm asking it to kind of build my PowerPoint.

47:07Florian Juengermann:Like how does it actually build that? So I think if you don't at least have that level of experience, I think it's no longer fit. Yeah, I think I'm kind of the same way. I also updated some beliefs. Yeah, I think this is great. Awesome. Thank you for that. Thanks for listening to Max Agency. If you liked this episode, leave a review and subscribe. Send feedback or questions to maxagency at langchain.dev. We want to hear from you.

From the publisher

Florian Juengermann is the co-founder and CTO of Listen, an AI startup that turns qualitative research across hundreds of interviews, surveys, and focus groups into structured, traceable insights. Listen's agents analyze responses at scale, and Florian has rearchitected the system multiple times to get there. In this conversation, he walks through the virtual table architecture at the core of their Research Agent, how small models run map-reduce classification across thousands of open-ended responses, and the self-reviewing feedback subagent that catches errors during long async runs.


We also discuss:

  • The three agents inside Listen's platform
  • How Listen rearchitected from a simple RAG bot to a multi-agent system multiple times
  • Why the PowerPoint subagent was completely rebuilt using Claude's code SDK
  • Contextual prompt engineering as an alternative to skills
  • How Listen keeps report numbers live as new interview responses come in
  • When to trigger the long-running agent vs. showing early results
  • What Florian looks for when hiring agent engineers


References:


Where to find Florian:


Where to find Harrison:


Where to find LangChain:

Send feedback or questions to maxagency@langchain.dev


Timestamps

(00:00) Introduction

(01:25) The three agents inside Listen's platform

(03:15) Live chat vs. long async runs, and how Listen tunes for each

(05:33) Under the hood of the Research Agent

(06:37) Listen's virtual table architecture

(07:34) How small models classify thousands of open-ended responses

(10:05) Running code in a sandbox: how E2B fits in

(11:52) Why Listen rebuilt the PowerPoint subagent from scratch

(14:11) Contextual prompt engineering instead of skills

(16:32) The feedback subagent that reviews its own reports

(18:14) How Listen runs evals in production

(19:47) Unexpected ways users push the agent to its limits

(21:42) How many times Listen has rearchitected, and why

(24:59) Trace observability: depth over breadth

(26:10) Lessons from running Claude Code SDK inside E2B

(27:42) Memory: what's solved and what isn't

(29:10) The Composer agent UX: co-editing a document with AI

(35:50) How Listen keeps report numbers live as new responses come in

(43:47) What Listen looks for when hiring agent engineers


More from Max Agency

All 11 episodes
How Listen is building a system of AI Agents & subagents for specialized tasksMax Agency · 48 min
Listen in VO