In short
Podcast Summary: AI Agents for Data Analysis with Shreya Shankar - #703
Host and Guest
- Host: Sam Charrington
- Guest: Shreya Shankar, PhD student at UC Berkeley
Episode Overview In this episode, Shreya Shankar discusses her research on DocETL, a declarative framework designed for building and optimizing data processing pipelines leveraging large language models (LLMs). The episode delves into the architecture of DocETL, the challenges of designing agentic systems for data analysis, the current landscape of benchmarks for data processing tasks, and the need for robust evaluation methods in human-in-the-loop LLM workflows.
Key Concepts and Discussions
Introduction to Shreya Shankar
- Shreya is a PhD student focused on interactive and intelligent data processing.
- Her background includes working as an ML engineer, emphasizing the significance of data quality in machine learning applications.
DocETL Framework
- DocETL is a system aimed at optimizing large-scale document analysis tasks using LLMs.
- Core Features:
- Declarative programming constructs for data processing.
- Ability to specify high-level operations such as map and reduce, powered by prompts.
- Interactive nature allows users to iterate and refine prompts based on initial outputs.
Challenges in Data Processing
- Data from sources like police departments is often unstructured and massive.
- Traditional methods (e.g., hiring interns) are inefficient for annotating and analyzing such data.
- The stakes of incorrect data handling in sensitive areas (e.g., police misconduct) highlight the need for reliable systems.
The Need for Robust Evaluation and Benchmarking
- Current benchmarks for data processing tasks are lacking.
- Shreya emphasizes the need for tailored benchmarks that focus on the complexity of data processing rather than traditional reasoning or coding benchmarks.
- Discusses the importance of user-defined evaluation criteria, such as precision and recall, which can vary based on subjective task requirements.
Agentic Systems and Fault Tolerance
- The complexity of the DocETL codebase (14,000 lines of Python) stems from the need for multiple agents and fault tolerance mechanisms.
- Discussed strategies for handling agent failures, such as automatic retries and human oversight.
- Emphasis on developing reliable agentic systems that can effectively navigate complex tasks while managing uncertainties.
Future Directions
- Development of user interfaces for better interaction with the pipeline.
- Continued exploration of optimization strategies for LLMs in data processing contexts.
- Expanding the research agenda to include the creation of meaningful benchmarks for data processing tasks.
Real-World Applications
- Example use case involves analyzing police misconduct data to identify patterns.
- Highlights the iterative process of discovery in data analysis, where initial outputs inform subsequent prompt specifications.
Conclusion This episode provides a comprehensive look at the intersection of AI, data processing, and user interaction. Shreya Shankar's insights into DocETL highlight the ongoing challenges and innovations in creating reliable, efficient, and user-friendly data processing systems powered by large language models.
Further Resources
- For complete show notes, visit: [TWIML AI Podcast Episode 703](https://twimlai.com/go/703).
Key Takeaways
- Importance of Human-in-the-Loop: Human feedback is crucial in shaping the interaction with LLMs to ensure task alignment and output quality.
- Need for Tailored Benchmarks: Existing benchmarks do not adequately address the specific challenges of data processing tasks.
- Agentic Complexity: Building effective agentic systems requires careful consideration of fault tolerance and user interaction strategies.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00So just to give you a sense of the complexity of the codebase, it's 14 ,000 lines of Python. So the reason that it complexes basically a lot of agents and a lot of fault tolerance for these agents, every point where you insert an agent into the system, you need to handle the case that it failed. There are multiple ways to do fault tolerance. One is to ask the human, another one is to automatically retry, and then there's task specific ways to do it. It's super dependent on the situation. But the main point is you have to have fault tolerance for every agent. And you have to build that in from the beginning.
0:46All right, everyone, welcome to another episode of the TwiML AI podcast. I am your host, Sam Charrington. Today, I'm joined by Shreya Shankar. Shreya is a PhD student at UC Berkeley. Before we get going, be sure to hit that subscribe button wherever you're listening to today's show. Shreya, welcome to the podcast. Thanks, Sam, for having me. I'm excited to be here. I'm excited for the conversation as well. We're going to be digging into your research and, in particular, the project that you just released yesterday, Doc ETL. But you've got tons of experience working with LLMs and LLM evaluation, and I'm sure we'll touch on those topics as well.
1:25To get us started, I'd love to have you share a little bit about your background. Yeah, so I am a PhD student, fourth year at UC Berkeley. I'm interested broadly in interactive and intelligent data processing. And this is inherently multidisciplinary. I'm very excited about data management, data processing. How do we unlock insights from so much data that we have in this world? I'm interested in AI because I think that AI provides all of these capabilities to do analysis that we just never, never had before. And I'm interested in HCI. When humans have to work with these inherently unreliable systems, what are the right interfaces to help humans gain trust and iterate quickly on the analysis tasks that they want to do?
2:13So it's such an exciting time, I think, to be doing research in these areas. Before my PhD, I was an ML engineer at a startup, did a lot of structured data, tabular data, machine learning, a lot of XGBoost. And then before that, I was at Stanford doing, I don't know, learning to code, I guess, doing my undergrad at CS. Now, your research or the areas of interest you described are pretty broad. How have you structured that into a research program for your PhD? Yeah, that's a great question. My primary area of research is in the data management community. I was drawn to it, actually, because my job as an ML engineer, I found 99 % of my problems were data problems.
3:02Maybe my data was not clean, the training wasn't good, but it was because the data was bad. The inference wasn't good, but it was because of the data was bad. And I was learning so much about data quality. I was learning so much about large-scale data engineering, like how do you build these robust ETL pipelines, that I thought, do people do research in this because it's such a hard job and nobody could do this well? If you go to industry meetups, everyone has these problems. And yeah, no better way to learn, I guess, than do a PhD in Berkeley Database Group. and kind of along the way I think I've gotten I've always been interested in AI and machine learning especially coming from Stanford and along the way I really picked up this interest in HCI especially with generative AI and large language models these days are fundamentally changing kind of how we interact with content and data and produce content and the experiences just don't they feel so like day zero we're not even we're so early and it doesn't feel great like I'm writing something with AI but I open a chat to like tell the AI to write something but I'm writing it and I have it's just it doesn't feel right the vibes are completely off I mean that's what's exciting to me about it it still amazes me that the thing that cracked this all open was slapping a chat interface on top of an LLM that's nuts absolutely absolutely it's such a fascinating thing and you can think about like the marketing behind it too like TikTok chat to bt videos were all over tiktok in 2023 and i'm sure that contributed to this and it was always things like oh here's this app that'll do your homework for you and you have these like high school students who are like bring me an essay on napoleon or whatever and did that blow it all up for us i don't know and so you know someone who spends time thinking about interfaces and interactivity in HCI?
4:59What's your kind of glimpse into the post-chat dialogue-oriented interface to these types of systems? How do you think about where that's all going? I think we're inventing new interaction paradigms with these AI models. At least in HCI research and literature, people kind of did work on these physical devices that came out. So when computers came out, people did research on cursors, on inputs, on text boxes. These were the primitives that kind of defined how we build UIs. When the phone came out, people did a lot of gestural research. If I drag down, it will refresh the app. Somehow we intuitively know these things.
5:44And I don't think we've built the intuitive primitives for interacting with ML models, like the inherent non-determinism of it. How do you edit content indirectly, but still have that direct feeling that you edited content? That's just a problem I've been thinking about all of today. We're building an interface for people to interactively write data processing pipelines with an LLM. Like writing in a chat, extract all of these fields from a document, and then it synthesizes the code, and then you run it. it just doesn't feel right it's not very iteratable especially when you look at the output and you say oh you missed these fields are you going to type in the chat you missed these fields yeah yeah like that seems it's yeah it's i don't know it seems such a far cry from what we could do unfortunately i don't have a better answer but we're definitely exploring these kinds of questions yeah i remember having the feeling uh i think it was in the context of like i was at the Google Next conference and they rolled out their like enterprise agentic vision, which kind of looked like, you know, all of their old application silos for sales and marketing and customer service were replaced with like agents now.
7:07And it was interesting, but I also remember having the feeling that like dialogue boxes are not the right interface for everything. And if, you know, all enterprise software suddenly became like a chat box and not like these interfaces that we've refined over time. That seems like a step back, and it sounds like you agree with that. I agree with that. I think it's difficult, right? Chat is universal, so maybe that's why people use it. And you kind of do need some sort of domain-specific things. But one thing I really like and I think about, now Notion AI actually did this, But the ability to highlight content and then edit it, but edit it through an AI.
7:54So you highlight the content you want edited and you instruct in the text box how to change it and the AI changes it. That seems so much closer to direct manipulation hearings. Yeah, there's something about like, I don't know, I'm thinking about state as like a concept and wanting the LLM to be able to like directly manipulate some state, but also the human to have whatever the native way of manipulating that state is, you want them to have that also. Absolutely. Absolutely. That's 100 % the core of it. And then there's the added dimension of what's the context that you need to manipulate the state.
8:35What we see in our vision, everything that you're seeing on the screen is the context that you're using, plus like stuff in your head you haven't told anybody. But LLM can't read your mind. And it can't possibly have all the context on your computer. So like, what's the right way to to indicate or at least shape the context here, I think is so hard. So a lot of this speaks to some work that you're doing around Doc ETL. We kind of jumped ahead into the future. Let's pull it back and talk a little bit about Doc ETL. And maybe a good place to start is kind of the way you think about the problem that you are solving with this project and kind of how it came about.
9:19Great question. And Doc ETL is a declarative framework for building and optimizing LLM-powered data processing pipelines. So that's a lot of words. I'm going to illustrate this with an example use case here actually at Berkeley. So there's a big project going on that's interdisciplinary, spanning the lab that I'm part of, the Epic Lab, the Berkeley School of Journalism, and even the state of California. The state of California is funding it. But the goal is to collect data from all the police districts in the state and identify patterns of officer misconduct. So that can be kind of excessive use of force.
10:05That can be lying, manipulation, lots of different definitions of misconduct and see if there are repeat offenders. California has this problem of sending an officer from district to district, even if they might have done something bad, whatever, wipe the slate clean, send them to the next district. So all of this data from the police departments is fully unstructured. And the way that they would solve this problem is to hire interns to annotate this data, things that could possibly be police misconduct, create a big knowledge graph. The data can be, you know, thousands of pages long. It can span multiple cases, multiple officers, all sorts of things.
10:52So the question really is, how can we use, how can we program LLMs to help solve this problem for us? What are the programming constructs that we need here? What is a data processing DSL look like here? And how do we make sure it's reliable? It's something to say, to submit your 100 page PDF to chat GPT and ask for officer misconduct. Is it right? Probably not. It's going to miss officers. It's going to get misconduct wrong. It's going to hallucinate things. The stakes are too high here for this kind of task. So Doc ETL is really, the goal is to solve these kinds of data processing tasks. Users simply have to specify prompts for operations at a high level.
11:34We have different types of operations, like a map operation, a reduce operation, very similar to kind of common data processing frameworks. But they're powered by prompts. So you just specify the prompt and then Doc ETL will rewrite these prompts chunk up the data, chunk up the task into smaller units that can be executed accurately by the LLM and then stitch the results together. And the implication of ETL being that this is something that happens as kind of a batch operation as opposed to an interactive, you know, a lot of the dialogue box we were talking about earlier. You are, yeah. The ETL, it's a pattern of, you know, writing a pipeline to analyze a large amount of data.
12:21Now, it has to be interactive because what we found through preliminary studies, and I've even written about this in some of our validation work, is that humans have no idea how to write the prompt until they see the initial output. They have to see what the LLM came up with the first time that they wrote the prompt and then say, oh, like, actually, like, they changed the task. They like, they redefined misconduct. They redefine what it means to use excessive force. And so it has to be interactive in this way. The question becomes, how do you allow them to iterate fast, not how do you completely remove them out of the loop?
13:00And you mentioned the DSL is the idea that the LLM is exclusively used to create this artifact that defines some backend processing. or are LLMs also used in the backend processing for things like classification, text extraction, what have you? Those aren't clearly not mutually exclusive, but I just want to make sure we understand where LLMs come into play. LLMs come into play in both the execution of the operators. So I have the LLM do the extraction from the document as well as orchestrating. So LLM agents determine, hey, this input document is too long. Let's figure out the right chunk size to process this subtask to get good accuracy.
13:50So that's what we call the optimization stage. You define your pipeline at a high level. We optimize it, which means break it down into units that we know the LLM can do well based on automatic validation and so forth. And then you can run the optimized pipeline. This is maybe more of an engineering question, but with these ETL systems, often like one of the biggest challenges is that everybody wants a different connector to either their data source or data sync. Like, have you started dealing with all that? Where is all the data coming from and going here? Yeah, we it's such a research prototype right now.
14:27I think people just have like PDFs and folders. A lot of people don't even put unstructured data in the data warehouse because it's designed for structured data. unstructured data lives in like document clouds in jasons and i i totally see like once like in the future we're certainly going to have to like connect directly with the data warehouse and so forth but right now i think it's so novel to kind of run these kinds of analysis tasks with llms that doing it in a folder is a lot and so is the user experience you point this thing at your folder, it slurps up all your documents and then you're interactively kind of building out this processing DSL.
15:14Like, do you see like a spreadsheet kind of interface or you get like, how does the user as part of that interactivity, how does the user see the results of what they are trying to do from a document processing perspective? Great question. Right now we don't have an interface. We're building it. the user kind of writes a YAML file that specifies the operations. Each operation has a type and a prompt. So that part you kind of do in your text editor and there's a command to optimize the pipeline and a command to run a pipeline. But the future we see is we're going to have a UI. We're building the UI.
15:57I like the idea of a spreadsheet, but people just get very turned off by large documents and spreadsheets. It's not the right way to visualize it. So we're exploring, you know, how do you create one operation at a time, run that operation, look at the outputs for a sample, and then have that, you know, inform your next specification of that operation. Once you're feeling good about that operation, moving on to the next operation. A lot of these analysis tests across domains follow a map reduce like pattern. For each document, map it to a set of one or more topics, where topics can be misconduct instances, pretty much like anything domain specific.
16:40Then canonicalize those topics, because LLMs have variability, they might extract names differently, so forth. Then group by the topic and apply some summary. And is summary like a textual summary or like an aggregate in the database sense? A summary is simply an aggregation. You take in many inputs, you output one thing. So it's an LLM-powered aggregation. So it's anything, your prompt, you can write anything in your prompt. You have access to all the inputs and you return one thing. So summaries are just a natural way to think about aggregating a lot of text. But you can certainly imagine, like, you know, find the most common type of misconduct.
17:25That could be a prompt. find anomalies, find officers who committed the most misconduct. Like all of these, it's very flexible in that sense. But yeah, a bunch of pipelines follow this map reduce like pattern across domains. Are there any particular models or types of models that you have built this around? Did you emphasize open-waste models like LAMA? Or did you start with the open AI models? How have you thought about model support? Good question. We started with open AI models because I think they were just the first to support tool calling. And we do support tool calling in our map operations.
18:13But, I mean, any LLM can be used, I think. I do think that there's a big need for supporting open source models for sure, especially when you're dealing with sensitive data in the medical domain. I find it very fascinating. Some people say, oh, we're fine using OpenAI. We're fine using Azure. We're already on Azure, so we're going to use Azure. And then other people will say, no, we absolutely have to use Mistral. So, I don't know. All sorts of... I think we need to support both. Now, extracting data slash text from a PDF, you know, sounds very simple when you say it quickly, but that can be arbitrarily complex.
19:01And there's lots of maybe adjacent projects that are doing things like using vision-based tools to identify regions in the PDF and dealing with the visual structure of the PDF or the text structure of the PDF? Are you dealing with any of that? Or is that outside of scope for you? Not yet. We want to. We just naively run OCR and take Azure OCR tools and then take that as input. The kinds of tasks that we think about, the kinds of extraction tasks we think about are really intelligent in extraction. You need to have context or domain knowledge to know. No, for example, extract all medications from this document.
19:48You need to know what a medication is to be able to extract it. It's not extract all instances of Tylenol. Like that, you can kind of figure out some other way. um but these kinds of that's just a simple intelligent extraction task like the officer of misconduct um that one's a really tough one like llms can do it well when prompted correctly um sometimes you know the data might add a lot of misconduct in it um and then you need to look at smaller chunks at a time to extract all instances of misconduct if it's a very large document and it's like a needle in the haystack problem, there's only one instance of misconduct in there, the LLM might get that automatically.
20:32So it's very data-specific and very task or prompt-specific. And these are the kinds of interactions that people who are writing pipelines, whether they are journalists or they are expert AI ML engineers, everybody has to do experimentation to figure out what's the best way to make these reliable. To me, that's the problem of an optimizer. Users should be empowered to see outputs, make change to prompt, and then everything else is the job of a system to help them achieve that. Are there things that this particular problem has shown you about writing prompts that you haven't come across, you know, generically or with other problems?
21:15Yeah. The more you're able to have users look at the intermediate outputs, say the extracted cases of misconduct for a sample of docs, the more complex the prompts get. When people, and that thing is very fascinating to me because I think there's a lot of work in automated prompt engineering, optimization, and so forth that really puts the human side of the line. The SPY and the like? I really like the idea of DSPy making it easy to declaratively specify things. But not having a human in the loop is very... You have to have a human in the loop here. Because if a human is seeing the intermediate and would do something differently than what the system will do, then already you're not aligned with the intent anymore.
22:07So that's been surprising to me. Like the more humans iterate, the more it diverges from what the system actually would do if you took them out of the loop. Is there a particular example of that that you've seen that comes to mind? Yeah. So in this example we talked about in our EvalGen paper, which we'll be presenting in a couple of weeks at WIST, we've been doing research on how do we have human guided AI assisted evaluation of outputs. And something that's interesting is that humans look at more and more outputs and change their mind about what makes a good output based on the LLM's behavior.
22:51They might say something like, don't specify a hashtag in the output, you know, if they're mining tweets. But then later on, they might realize actually they do want important hashtags in the outputs. And this is not something an automated eval assistant would ever be able to ascertain. The human looked at it and changed their mind about what the task is, which I thought was very interesting. A similar case of this was somebody was using Doc ETL to mine reviews for a particular course to see what are common themes for why people signed up for this class and so forth. And when they saw the intermediate outputs where they were mapping reviews to themes, they were like, this is way too many themes.
23:48I don't want this many themes. But if you ask an LLM to kind of refine that prompt further, the LLM won't go in that direction of reducing the number of things. It won't know to. So it's always a question of how do you get the human to steer it really, really easily and without too much overhead. What did you come up with in that paper as some ideas for how to do this? It sounds like just letting the human kind of have at it in their little chat box and bang their head against the problem. Is there more to it than that or have you come up with some interesting approaches? interestingly that paper um so we built an interface that allows humans to grade outputs as good or bad and then uh use as an llm to find implementation like predict whether the human will grade it as good or bad um and then in the study in the user study we observed this kind of criteria drift so it's we were i don't have a solution for that because it kind of emerged as part of the study.
25:02The other thing that I think we were able to, we were uniquely able to observe that because our interface supported really fast iteration. If we're able to show a human 100 outputs and get their feedback on it in the span of a minute, like you can observe, we can observe things in that kind of study. Then say, if you ask a human to only grade three things, and in between each time you ask them, it takes a really long time. So speaking more broadly about evaluation, like if you're, if you've got the LLM in this planning loop and you're trying to build out and evaluate, first of all, what does evaluation even mean in the context of data processing?
25:45Is that well-defined or is that to be defined? And then once you stick an LLM in there, like how does that change the way you need to think about evaluation? Yeah. Good, good questions. So first I'll explain kind of how.etl's optimizer architecture works. The user specifies a high level pipeline. And then we have generation agents that take the operator and rewrite the operator into a sub pipeline. You can imagine writing a map operation, which just one to one document to output into its own little pipeline that chunks up the document, applies a map to each chunk and then aggregates the results.
26:28That's one type of rewrite. There are many types of rewrites. You could do a chain of thought like rewrite where you break up that task into an intermediate task that's helpful and then the task. And that could be arbitrarily long. Or if it's a, this is very common, people want to extract 40 fields of information from documents. You could break that up into 10 operations, each extracting four fields. So you can imagine there's like all this number of rewrites that we have an LLM come up with. And this gives us so many candidate plans. Then we can run these pipelines on samples. Is the LLM coming up with these candidate plans in natural language or DSL domain?
27:18In our, it's very guided. So we have a DSL of these operators and we go through all of our predefined rewrite rules and we ask the LLM to come up with prompts for the new operators in those rewrite rules. And we have no idea how well it's going to do until we try it on a SAP. But it's kind of a constrained domain that it's working in. Yeah, exactly. And it's nice because this paradigm is very flexible across domains. Like you can do this for processing legal documents. You could try a bunch of different chunk sizes, or you could try a bunch of different task decompositions. And the agents will kind of come up with it.
28:03They'll come up with the prompt. We'll try a number of chunk sizes. Like here we have 100 plans. Like, great. All the varying accuracy. Who is going to know? How is anyone going to know how well the candidate plan will do? And this is where the evaluation then comes in. If we run all of these 100 pipelines on a small sample of data, which outputs are the best? So in the police misconduct case, or any case, first we synthesize a task-specific validation prompt. So in the police misconduct case, you might have, are all instances of misconduct extracted to associating recall? For each instance of misconduct, can you actually trace it back to the source document?
28:54Is every extracted instance actually misconduct? Like, does it satisfy use of force, so forth? So we'll have like kind of three bullets that are very specific to the task. And then we'll have the LLM evaluate the planned outputs with respect to that. So everything hinges on having a good validation prompt and a good kind of ranking algorithm here. So ideally, we would do pairwise comparisons. Like that's pretty well known in the ML literature. If you want to find the best model or find the best plan, you compare outputs side by side. But that doesn't scale when you have hundreds of candidate plans.
29:36So we kind of do a coarse grained rating for each one and then take the top K plans and then do pairwise comparisons there. So if your validator prompt is good, this will give you, if you trust your validator prompt, this will give you the best plan. And in practice, the validator prompt is a really good first start. It gets precision, it gets recall, it gets accuracy. where it fails is in these like when the human wants something that is a little bit ambiguous and can vary from human to human like maybe you want to extract coarser grained themes from your product reviews than i might want i might want hundreds of themes and summaries of all the quotes around that and you might only want two themes so like that that's where the knobs have to change on validation and i think that's also where you get an interesting ui problem of how you can get the human to guide that and so is that uh evaluation framework uh let's say of precision recall accuracy is that kind of a hard-coded aspect of doc etl etl or is that an example of an evaluation regime that a user could specify somehow.
30:56It's not hard-coded by a doc ETL. We use an agent to come up with that validator prompt specific to the task. And the agent's prompt is, you know, you're doing a data processing task. Here are a common data processing tasks. Synthesize an instruction related to precision. Synthesize an instruction related to recall. Come up with one or two other task-specific instructions that relate to the quality of the output. And then that's like the validator prompt. Got it, got it, got it. And then in talking about this evaluation stuff, the whole concept of agent and agentic has come up a few times. You know, kind of zoom back out and talk about the system from that perspective.
31:38Like, did you, is it a multi-agent? It sounds a little bit like a multi-agent kind of thing. is it and if so like what are the main agents like are they orchestrated from the ground up like you know did you write the orchestration or are you using some open source thing like you know what can what have you learned about building agentic systems and what can we learn you know from the way that you've done this oh man agentic systems are so hard to build we started building this prior to any of the frameworks that are out now. just to give you a sense of the complexity of the code base, it's all in Python just because it's easy to work with LLMs in Python, but it's just disgusting to build a system in Python.
32:28Anyways, it's 14 ,000 lines of Python, which I think is a pretty huge Python-based system compared to other code bases. The reason that it's complex is basically a lot of agents and a lot of fault tolerance for these agents like every point where you insert an agent into the system like you need to handle the case that it failed um and one one place one one way to handle it is to surface it to the user right so already we have like an interactive system where we like somehow are passing that context everywhere that will stop the console and prompt the user if that's one way you want to escape from the failure mode.
Read the full transcript
33:14For all agents, we can't do that, right? That might just be too much effort. There are multiple ways to default tolerance. One is to ask the human. Another one is to automatically retry. And then there's task-specific ways to do it. If this is not critical, right? If this is just one type of our, we have 13 different logical rewrite rules that we could apply to operators. if this is simply one of those 13 like maybe it's okay to toss this guy out and we'll consider the other 12 right so a lot of it is very um it's super dependent on the situation but the main point is like you have to have fault tolerance for every agent um and and you have to build that in from the beginning i see so many people building their agents and then like not thinking about fault tolerance the first time they build it and then like at some point it's going to break but then there's so much complexity into like trying to recover that um yeah so i think maybe that's the biggest lesson not really framework specific um another thing is really constraining the domain of the agents so all of the agents that we write output like structured um outputs and the structured outputs are no more than two things because agents just cannot do too many things at once.
34:36And they're always like new prompt or like yes or no, like should I do this? Rather than have an LLM immediately synthesize a new operation, we always ask it, given this data, given this task, is it a good idea to make this new operation, to decompose in this way? And then if it says yes, we ask for the prompt rather than one step, is it a good idea? If so, give me the prompt. That's a little bit much to expect out of one LLM call. So a lot of just really unit decompositions. This itself kind of creates like there being 100 agents, 100 or more agents in the database, or sorry, in the code base.
35:17But from the perspective of like, what do these agents do? There's really only two types of agents, like validator agents who do validator type tests, and then generate agents that rewrite pipeline assembly. So. And if you had to guess at how much of your code base was, you know, essentially kind of agent orchestration infrastructure versus agent behavior, prompt oriented, like business logic-y kind of stuff, like do you have a sense for what that might look like? I don't think there's too much prompts for the agents. to be honest. The prompts are pretty small. At Bust, we have one example of your agent.
36:02A lot of the code is really in the logic of the operators itself. How do you execute operators fast? I've only talked about logical optimizations of pipelines where I rewrite pipelines to be different, or to be logically equivalent, but there's also kind of efficiency - Like performance optimization? Yeah, performance optimizations that there's a lot of code for like got it got it so i kind of split that up i am implicitly in like two levels but there's multiple levels and uh i guess that at the heart of the question was like do you think you like you know wrote the essential you know wrote like a crew ai or something like of that breath like Like, did you write a lot of code just to do low-level agent orchestration, or is it mostly in what I'm now thinking of as like a middle tier of like managing the business logic specific to your app?
37:08Yeah, good question. I have one LLM agent class that does the, it has some logging, and it has a wrapper around whatever LLM that I want. and then anytime I want to invoke the agent, I instantiate this object and I specify the agent prompt as well as the schema and then I specify what to do on how many retries and that kind of stuff. Or like the service. That's awesome. I mean, I had a lot of these conversations about the agent frameworks and often there's this tension between like using some off the shelf thing. But then like you have hidden complexity that you can't really see. And when it comes down to it, you know, in many cases, the agenticness isn't all that complex.
38:03It's all the stuff that you want to do with the agents. And it sounds like that's where you've ended up. It's really the fault tolerance. If I had to copy paste that code everywhere, I'm thinking about how like somebody else might build this and run into an issue. Like they would write a bunch of, if-else statements around every single agent call. And that will really blow up the complexity of the code base. But if you just centralize that logic and have specified policies, I think it's really not that hard. And now that I think about it, this LLM agent class that I have, it's not in its own file.
38:39There's other code in that file too. You know, when we talk about ML systems, we're often talking about benchmarks and that kind of thing. Is there any concept of a benchmark for like these data processing by blinds? I'm baited. You've baited me. No, nobody has built a good benchmark. We really should build one. We've been talking about it here. And I think it's partly because it takes a data person to know challenges of data processing. And there's just a very few data people who are also very well versed in kind of generative AI right now. I think the benchmark or the whole agenda in the ML research community has been to focus on a specific kind of task, like reasoning based tasks or really hard math problems or really hard coding problems, which are obviously very hard.
39:36But it's not the same kind of hardness as a data processing task. Like, a task where you could hire college interns to annotate the data, but it would take them 100 hours. And they have to be college interns. They can't be five-year-olds. They don't have to be Olympiad gold medalists, but they require, like, a base level of intelligence. These are hard tasks for LMS to do. They just get tripped up with so much context. and a lot of these math problems, coding problems are not that long. They're like a paragraph long. But the more we only focus on those problems, the less we're able to do complex reasoning over data.
40:25I was just going to ask, do the reasoning-oriented challenges not quite get you there? Is the core challenge that or the core aspect of an LLM that you're relying on to build these pipelines? Is it reasoning or is it more narrow than that that you would want a separate benchmark? Or how would you focus a separate benchmark on this problem? This may be a better way to ask the question. Yeah, yeah, good question. The first thing is maybe like why are data problems different? And I think the way that you use attention. So like all these transformer models are using attention. the way that we use attention for like solving coding problems is different from the way that we use attention for solving data processing tests like there's so much attention that you need to like maintain throughout the entire context and you also need to take actions you basically need to read the whole book if you're feeding in a book as an input and at every point you're reading you need to be able to make some reasoning like decision this is this is a different challenge this is not a needle in a haystack problem like if you're doing a needle in a haystack task it's kind of like an anomaly detection task like you're kind of just like looking along and you're like spiking whenever you see the pattern that you're supposed to have and then you're done and you know you're expected to spike a few times you're not going to spike twice in a row sort of thing um but data tests are different like identifying misconduct um requires like if you're in the middle of a police a transcript between a police officer and a suspect and they're saying he said she said he said she said and the officer name is 10 paragraphs ago right like this is not like the attention problem that the llms have been trained well to do i think the benchmarks just don't have it um yeah so in that sense i think like data processing requires its own set of benchmarks where the tasks, I think ideally, it's not specific to a single LLM call.
42:38That's another thing. The LLM benchmarks are focused on what can you do with an LLM call or a loop of LLM calls, but the input is fixed. These data processing benchmarks should allow for flexibility in how you decompose the data, how you orchestrate the LLM calls, that's like a departure from traditional benchmarks. Another thing is in data processing, you can have infinitely many correct answers. And you have subjective correct answers. You might want to do theme extraction and have more themes. I might want fewer themes. And a benchmark should have two flavors of that theme extraction task. And simply based on the flavor, your system should be able to do it correctly.
43:34These are just the kinds of thoughts we've been having in terms of like what makes data processing different. Have you explored any of the state space models or long context LLMs as a way to try to address at least that reliance on attention and locality that you referenced? We should. I know that they're getting better and better. We should definitely look into it. It's been on my mind. In terms of future directions, you already mentioned that one of the areas that you're focused on is building interfaces. And we spend a little bit of time up front talking about that. Where else do you see this going?
44:10Yeah, we needed to build a benchmark for data processing. That's one big thing. Interfaces, of course. I think there's a lot we can do in optimizer reliability, like the agentic reliability. We do a lot of fallback to asking. There's certain kinds of rewrite rules that agents are bad at determining whether they're appropriate in the situation. An example of this is a chain of thought like decomposition. If you have a task, technically, you can come up with a chain of thought decomposition. Or you can come up with any chain of thought decomposition for any task. But the question is, is this a good chain of thought decomposition?
44:55Or is it necessary? LLM agents are really bad knowing whether a chain of thought decomposition is necessary or not. So we do, if it recommends one, we ask a user like to validate whether they want to explore this plan further. And I think like decisions where we need to query the human there are just like a matter of like maybe our agents aren't very good. We should improve that. Along those lines, have you looked at like the 01 series of models, these more kind of reasoning oriented models? Yeah, we're thinking about it now. I mean, it just came out. Yeah. Definitely. Definitely. It's not very interactive.
45:43I think that's my thing with O1. I'm sure somebody will come up with a O1 that actually shows this chain of thought. Maybe from a latency perspective or introspection into the chain of thought and the reasoning process? Both. Like the worst thing from a user experience is to like wait 50 seconds for like something that feels like it shouldn't take 50 seconds. So, and then like to also have no ability to like see why it took 50 seconds. So, I think we might wait for a more open model to kind of be trained on this chain of thoughts stuff. Yeah, but kind of there's that. And then there's a bunch of interesting ideas of like we've come up, the paper will come out later this week, but we've come up with these rewrite rules for data processing pipelines.
46:34And I think there are more. I think there are more in terms of performance and efficiency optimizations. We can do rewrite rules there. Well, Shreya, thanks so much for jumping on and sharing a bit about what you're working on with DocETL. Thank you. Is it docetl.com? Is that where folks can find out or do they have to search for it? I can't believe I got the URL. I just looked it up. It was$20 or like$15. I just looked at it and I was like, wow, does that really exist? Yeah. I was like, cool. Yeah. That's awesome. And then someone commented, why isn't it docetl.org? And in my mind, I was like, because I could get docetl.com.
47:17That's why.
47:22Just because of it. Awesome. Awesome. Yeah. Cool. Thanks for having me. This was fun. I talked to your ear off. No, this is great. Yeah. This is great. Thank you.
From the publisher
Today, we're joined by Shreya Shankar, a PhD student at UC Berkeley to discuss DocETL, a declarative system for building and optimizing LLM-powered data processing pipelines for large-scale and complex document analysis tasks. We explore how DocETL's optimizer architecture works, the intricacies of building agentic systems for data processing, the current landscape of benchmarks for data processing tasks, how these differ from reasoning-based benchmarks, and the need for robust evaluation methods for human-in-the-loop LLM workflows. Additionally, Shreya shares real-world applications of DocETL, the importance of effective validation prompts, and building robust and fault-tolerant agentic systems. Lastly, we cover the need for benchmarks tailored to LLM-powered data processing tasks and the future directions for DocETL.
The complete show notes for this episode can be found at https://twimlai.com/go/703.




