In short
Notes on The TWIML AI Podcast - Episode #700: Automated Design of Agentic Systems with Shengran Hu
Episode Overview In this episode, host Sam Charrington interviews Shengran Hu, a PhD student at the University of British Columbia, discussing his research on Automated Design of Agentic Systems (ADAS). The conversation covers the motivation behind ADAS, its key components, and the implications of using large language models (LLMs) to generate novel agent architectures.
Key Concepts
- Definition of Agentic Systems
- Agentic Behavior Spectrum: Agentic systems exhibit a range of behaviors that enable them to operate autonomously and effectively.
- Key Components: Includes prompts, control flows, tool use, memory, and potentially reinforcement learning, which can all be optimized for better performance.
- Motivation for ADAS
- Evolutionary Inspiration: Shengran’s interest stems from understanding evolutionary processes and how complexity emerges from simple algorithms.
- Agent Popularity: The increasing effectiveness and applicability of agents in various domains prompted the exploration of their design patterns.
- Automation of System Design
- Learning from Design: The goal is to automate the design of agentic systems to discover and learn useful components autonomously, which can be applied to real-world tasks.
- Iterative Design Process: ADAS relies on iterative cycles to refine and improve the design based on previous iterations, akin to natural selection in evolution.
Key Components of ADAS
A. Search Space
- Code Representation: The search space for agent design is represented in a programming language, allowing for comprehensive exploration of possible building blocks beyond basic prompts.
- Flexibility: This approach enables the exploration of numerous configurations and combinations.
B. Search Algorithm
- Meta-Agent Design: A meta-agent is employed to explore the search space, leveraging its knowledge of coding and agentic system principles to propose effective designs.
- Sample Efficiency: This method aims for high efficiency in discovering new agents with fewer iterations.
C. Evaluation Function
- Performance Metrics: Generated agents are assessed based on accuracy, speed, cost, and robustness, leading to multi-objective optimization rather than a singular focus on performance.
- Trade-Offs: The evaluation recognizes the need for balance between different performance aspects, enabling adaptable solutions.
Discussions and Insights
- Complexity and Emergence
- The ADAS framework allows for the emergence of complex behaviors and design patterns that can evolve over iterations.
- There’s potential for agents to show meta-behaviors that reflect higher-level cognitive processes, akin to how human organizations function.
- Collaborative Systems
- The podcast discusses the possibility of utilizing multiple agents, each with different specialties, to enhance robustness and effectiveness.
- Real-world applications might benefit from a collaborative framework where agents can correct and enhance each other's outputs.
- Practical Applications
- Initial user feedback indicates that the ADAS framework has been successfully adapted by developers for real-world tasks.
- Future applications could involve continually refining agentic systems based on performance feedback and evolving requirements over time.
- Open Questions in Research
- The podcast raises questions about how to balance exploration versus exploitation in design iterations and the potential for emergent properties to arise from simpler design elements.
Key Takeaways
- ADAS presents a promising approach to automating the design of agentic systems, merging principles from evolutionary computation with modern AI techniques.
- The iterative and flexible nature of the framework allows for the discovery of complex agent behaviors that may enhance the performance of AI applications in diverse fields.
- Further exploration could lead to deeper insights into the nature of intelligence and cognition, potentially informing both AI development and our understanding of human-like reasoning.
Conclusion The episode highlights the innovative work of Shengran Hu in the field of automated system design and its implications for both AI research and practical applications. The discussion underscores the importance of exploring complexity and collaborative approaches within agentic systems to optimize their functionality in real-world tasks.
For more information, check out the complete show notes at [TWIML AI Podcast](https://twimlai.com/go/700).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00We have so many building blocks in the agent X system. We have prompt, we have the control flow, we have different tool use, or we may have memory. Some agentic system will even use reg. And all those things, if we only focus on prompt, all those things will keep fixed and cannot be learned. In this paper, we are trying to argue an approach that we can learn more things, and actually other things in agentic system.
0:43All right, everyone, welcome to another episode of the TwiML AI podcast. I'm your host, Sam Charrington. Today, I'm joined by Shengren Hu. Shengren is a PhD student at the University of British Columbia. Before we get going, be sure to take a moment to hit that subscribe button wherever you're listening to today's show. Shangren, welcome to the podcast. Hello, Sam. Thank you very much for having me in the podcast. I'm excited to dig into our conversation. We are going to be talking about your recent paper, Automated Design of Agentic Systems. The work is related to work that we talked about on the podcast with Jeff Klune.
1:27That interview was episode 602, almost 100 episodes ago, back in December of 2022. Before we dig in, though, I'd love to have you share a little bit about your background and how you came to study in the field. Yeah. So I think my research motivation, a lot of them is actually come from the theory about evolution. I'm very curious about the evolution in our biology, in different creatures, and also how those complexities emerge from a very mechanical, purposeless algorithm like Darwinian evolution in our nature. So during my undergrad, I was doing a bit of research in evolutionary computation, an optimization approach that gets some biology inspired from the natural evolution process.
2:21And then that led me to read Jeff's work, and many of his work really impressed me. And very fortunately, I'm now working with Professor Jeff Kuhn in the University of British Columbia. That's awesome. So tell us a little bit about the origin of this effort. How did you start down the path of attempting to automate the design of agentic systems? Yeah, yeah. I think, first of all, we are really curious about why agent is so popular and why agent is so powerful in many different applications. We are very curious about the design pattern that emerged from this agent and what kind of things is different from an agentic system design to a single query to LM or we call foundation models.
3:31And I think all thinking is actually starting from an analogy between how humans complete the task and how foundation model will complete the task. For example, if we query a foundation model, a single query, a single Q &A, ask it a question, say, how to solve P equals NP? And then we only allow it to reply it in a single turn. No extra iteration, no extra reflection or searching online. We feel that no matter how advanced our foundation model will be, this will be really, really hard. And that's not only true for our models. That's also very true for humans. If you ask me, what's your opinion on P equals 2MP?
4:25If I'm not allowed to think about it, jack about it, search around it, my quality of response will be quite low. And not even the task, the very hard question like P equals MP, just some very simple question. is today a very good day for traveling to, say, Vancouver, the city I live in? Well, if I directly reply you, then I would just say, oh, it's generally a nice place. But if you give me more time to think about it, I may start planning. Say, okay, to answer this question, I need to think about what aspects should I talk about when we're talking about is it a good traveling place? we will say, how about the weather today?
5:13How about the schedule? Can I spare some schedule to travel with you? How about booking some new restaurant? We can get some good food. So with all those planning, and then we need to use some external tools to check for more information. And then we can actually do this task very well. So from that analogy, we believe that agent tech system could be more powerful than a single query of the foundation models. It's a good analogy. And I think at its core is this idea that most of our interaction with the world, at least the more higher level interactions with the world, require some degree of thinking as opposed to kind of a one-shot action.
6:09but maybe a counterpoint is that well and you're also saying that that approach is more emblematic of like a agentic approach as opposed to a single query response from an llm but a potential counterpoint is all the work that's gone into chain of thought which is through prompting asking LLMs to do that same kind of thinking and planning before producing the ultimate response. How does that fit into your analogy? Yeah, I think the trend of sub-prompting, it's actually a bit of agented. If we can rate the so-called agentedness level, like we just directly ask the question no we don't have any requirement no sync step by step that that could be defined as zero and then if we ask it to do more plan it's essentially to give it more chance in in its response to uh right rewriting down its almost all its own planning and it's kind of automatically conditioned on its previous planning and then solve the task.
7:30So it's starting to get more and more agentics with that channel-soft technique. And then if you involve more techniques like self-refinement or using different tools or using and maybe have memory mechanism and also maybe you have iterative interaction with the environment, the system is getting more and more agentic. Yeah. I've got to remind myself to start all of these conversations with how do you define an agent because it is a spectrum and there are several different definitions. Is there a kind of a threshold for you or what really captures the essence of agenticness when you think about it?
8:25Yeah, I think I don't have an explicit threshold. I think for every different task, we might find a sweet spot in this wire spectrum. But I feel like we are noticing that there are many, many emerging or evolving building blocks in this spectrum. We have many very cool stuff getting developed from the community. And also people are investing many, many different ways to combining them into a new system and then applying them to new applications. And we feel that is very, very interesting. To what degree do you think about collaborative systems of different types of agents or agents with different specialties?
9:18How does that figure into what you're looking to do? Yeah, I think that's a very cool question. we will recently think about. I actually have another analogy here. People have actually recently have a very high expectation on the foundation models. That what we are expecting them to not make many mistakes, to have very good reasoning skill or planning skill. Actually, the current state we have for foundation model, if we put it back to like five years ago or 10 years ago, people may already call it AGI. But now we seem to have more expectation for what AGI looks like. But when I think about it, when we think about how humans solve tasks or how humans collaborate with each other as an organization to solve the task, sometimes we expect the whole organization or the whole structure or whole architecture become stable or robust.
10:23But we will not expect every single human in it will never make any mistake, right? Because everyone will make mistakes. I am making mistakes every day. So the important thing is we have this kind of organization. We have this something called workflow or SOP, Standard Operating Procedures. And actually one very cool example is just the checklist you have given me before the podcast. You require me to check different things, make sure I have this device, make sure I'm in a good location so that we can make less mistakes. So I think that kind of thing is quite important in our agentic system. because it's very hard to expect our foundation model will never make any mistake.
11:24But if we have this kind of SOP or workflow, that foundation model may can correct each other so that we can more likely to have a more robust system. Your argument there is essentially that even as LLMs get better and their propensity to make mistakes decreases, there's still a place for these multi-aging collaborative systems because the robustness of the system as a whole increases through the collaboration, whether that's checking each other's work or providing guidelines or orchestration of some sort. Yeah, yeah, exactly. But we actually are very... So during the writing of this paper, we are designing on many decisions on the terminology.
12:16And we are very careful about the terminology of multi-agent. We are still not very sure that what kind of system we should call it multi-agent or what kind of system we should call it single agent. What's the boundary of one agent? Because when we think about it, one agent inside is still not also a single LM query. Inside one agent, if we have some definable boundary, they still have multiple modules inside of it. They have reflection, they have planning, they have maybe their memory. So the boundary could be the rule definition. We define some set of modules, say, okay, you're under one rule, so that's the boundary of an agent.
13:12Or so we use different LMs for different agents because we know that the diverse perspective from different people, I believe, is a source of why a system can be robust because I can make mistakes, everyone can make mistakes, But if we are from a very diverse perspective, then it's unlikely that we all make mistakes at the same time, at the same moment. But we are now experiencing a time that we are using mostly very similar models or even exactly the same model in our whole system. I'm using GPT-4 series for all the development of my paper. So that's a bit of not very robust, if you think about it.
14:08If we find a way that can hack in the GPT-4, or we can find a way of successfully inject the prompt to let it do something that is not so aligned, And that might be not very good for AI safety perspective. But overall, I think it's very interesting of this idea of multi-agent, but we kind of need to think more about it on what's the definition and how does it differ from our definition on the agent's concept. It may end up being another spectrum. Yeah. So a core idea behind the ADAS paper is this idea of AI generating algorithms. Can you give us a brief review of that idea and research direction and how it leads us to the ADAS work?
15:10Yeah. I mean, I think the AI generating algorithm. So this paper is published by Jeff during 2019. And I think the origin of this approach or this direction is from the observation that in the machine learning community, we have a recurring theme. That we are replacing the handcrafted design of our AI system with a more efficient learned design or solution. For example, a very, very early example is in computer vision. What we do is we handcrafted some representation or some feature. Feature detectors and all that. Yes. And then we realized that we can actually do end-to-end with the rise of convolutional neural networks.
16:09We found that if we do all the things end-to-end, we don't need to design any features for specific tasks. We just define a very powerful architecture to learn them all. And more recent example is actually we can put more data and compute resources in learning more components in AI system. For example, in AI generating algorithm, we will say there are three main parts. One is the learning of architecture of the AI model. One is the learning algorithm of the AI system. And then to generate many learning environment or learning data for our AI system. And some recent example is like recently people trying to actually meta-learn the training function or the loss function for the RLHFLMalive.
17:11And people find that the learned loss function is actually better in some tasks, better than the hand-designed, the state-of-the-art, say, DPO loss function. And this kind of advancement in AIGA is very inspiring in us to our thinking in agent tech system direction. Because when we think about it, there are so many building blocks it's inventing by the community at the moment in the agent tech system. And then people are combining all different kinds of building blocks in a new way to create new agentic systems for all different kinds of applications. But the question is, how long does it take for us to discover all possible or all useful building blocks?
18:07How many more should we discover or design? And even if we discover most of the useful ones, how many effort we need to take from human to actually combine them for every single use case in every of our application. so inspired from that consideration we believe that the next step could be the automated design of the agentic system to let it learn from the from the design itself and then we can meta-learn the agenting system design where we put in more compute and more data and try to learn more component, this AI system. Talk a little bit about what it even means to design an agentic system. In other words, for the purposes of your research, how did you define the elements of a design that you're attempting to automatically optimize?
19:11Yeah. I think in this paper, we try to describe this newly forming research area, automated design of agent tech system, ADAS. And we believe actually it has three key components. The first key component is exactly what you described is the search space. What is the space we are designing our agent system in? So there are some previous work we can consider as some attempts to ADAS. say we are optimizing the prompts for the LM when it's trying to do some tasks. But we will realize that the prompt design is actually a quite small part in the whole agentic system design. Because we have so many building blocks in our agentic system.
20:12We have prompt, we have the control flow, how different LM calls or foundation model calls interact with each other. We have different tool use or we may have memory. We may have some agent system will even use reg. And all those things, if we only focus on prompt, all those things will keep fixed and cannot be learned. So in this paper, we are trying to argue an approach that we can learn more things and actually other things in agentic system through a search space in code space, in programming language. Because we are representing the agent in a piece of code and the code space actually is complete.
21:07So actually any possible building blocks and the way of combining them can be represented in this code space. So this is actually a key point in our paper. And that's the first component we believe that is very important in ADIS. That sounds hopelessly open-ended. Yes. Yes. And that's very, very exciting, we believe. And the second key point we try to argue in this ADAS work is in the ADAS research direction is the search algorithm. How can we explore such a large search space if we define it in a very flexible way, say in the code implementation? There are endless possible implementation can happen in the code.
22:07So people are exploring all different kinds of ways to doing that in the history of AIGA or even in the history of some early attempts to aid us. But we believe that if we are going to search in a code space, then a very good option is actually have something called meta-agent to program or invent new agent in code. because our foundation model, they are very good at coding. So we can actually take a use of this prior to enjoy its existing expertise of the knowledge of coding and also the knowledge about the foundation model and the agentic system itself. because it already knows a lot about coding and the agendian system.
23:05So it can actually propose very novel design or very effective design in code in a very short amount of iteration. So we believe that this kind of approach is very interesting because it can have quite good sample efficiency. and it itself is a very interesting self-referential method because the meta-agent that's trying to program and search for new agent, it itself is also an agent that can be improved through ADAS. So this is quite interesting future work direction we are looking at. And also our final key component in Ada's work, we believe, is called the evaluation function. How can we evaluate our generated solutions?
24:08Are they good? Are they bad? In what sense are they good? How can we improve them? So a very simple way to do that is just test it, test the performance of it and get a scalar performance reward. So that's the standard way we are doing currently in this paper. we have a target task, say, MMLU or MASS or GSM-AK, and we test our generated agent on it, and then we collect the accuracy or F1 score or any performance matrix. But looking ahead, we will feel like there are a lot of possible work actually can be happening here and lots of interesting possibilities. For example, in real world deployment of agentic system, we will not only focus on the performance, actually.
25:04We will consider many different aspects, say, the cost of the agentic system, how many API calls involved, say, the latency. Even though you have many API, if they can be parallel, then the latency could be better. And in some cases, the latency is very important. and the robustness or safetyness. Like we can test our generated agent X system on some benchmark of safety or we can test how it performs against some prompt injecting attack. So, and then I think it's very interesting that we can have some kind of, say, multi-objective optimization algorithm. That's something I have been working on during my undergrad.
26:00It's like an algorithm that not only optimizes for one single objective, but multiple objectives for the same time. And then what you get is not only one optimal solution. It's actually something we call parietal front, a trade-off front between different objectives. Say in some application or some scenario, you want the system to become cheap or have less latency. But you kind of say, okay, I can sacrifice a little bit of performance here. And then you can choose from the front of the agentic system that have a faster inference speed. Or in some case, you may want, okay, this place, the performance is really important.
26:56So I want to take more time, more money in it. And you can choose from the other extreme. So I think there are many, many different future work direction we can investigate in ADAS. I commented earlier that the search space being any Turing complete program is seems hopelessly open ended and very difficult to tackle. But it does strike me that in some ways, one approach could be to constrain your search space to known building blocks and then have your search algorithm do simple iterations over the known design components. right another uh approach though is to have a very open-ended search space but then implement constraints in your search algorithm like the components that you're expecting to build these agents out of or just the approach that you search and it uh i'm wondering if you can explore like that trade-off at first blush it sounds like oh you're tackling this wide open search space but if you then just are shifting the constraints over to your search algorithm, like, what's the difference?
28:13And why even have them as separate components? Yeah, I think that's very cool thinking. And actually, we think quite a lot about that because in theory, we can allow our meta agent or our search algorithm to construct any possible agent in programming language. Cerectically, that's completely possible. But in practice, that may be very inefficient. I think you'd have to give it some direction. Yes, yes. Because we have so many basic functions in our agentic system building, like how to QEA OpenAI API, say, or Cloud API, how to format in the prompt, how to manage the information flow between different modules.
29:08So in this work, actually, we are doing a very simple investigation that we give it a very, very simple framework that's less than 100 lines. And then it did help it to define the function of Q-reader foundation models or LMs and formatting the problem automatically. And then it can only define something we call a forward function, and then to describe the behavior of a genetic system using those existing function cores. But also it can develop its own function core, as we observe from the experiment. It will define its little bit, say, many of these meta-agents generated agents, they like to have an ensembling technique.
30:08So they have implemented a little majority voting function in their design of agent. And we find that that is quite an interesting way to explore our code subspace. And also, we believe that this is actually a benefit of our search space to define encode, because we can use the existing human efforts building blocks that are available there. Say we have a lot of existing LAM or agent framework. And human already built a lot of cool stuff in there. all different kind of tools, maybe an implementation of React system, how to automatically format the prompt. And if we define our agent in code, then that means that we can just simply take those existing efforts from the human and then put them in the search space as the existing building blocks.
31:21And then our algorithm can actually start using them. And that will be a lot more sample efficient than build everything from scratch, although building everything from scratch is very fun for research purpose. One of the diagrams you show is one of the agents that was discovered in applying the framework to ARC.
31:44displays a fair degree of complexity. Like, you know, there's a bunch of chain of thought. There's several different types of experts that presumably this agent, the ADAS system discovered and built into this agent. You know, et cetera, et cetera. There's like more complexity. And so now I'm thinking about, Well, I guess I'm trying to wrap my head around where the turtles stop, like the turtles all the way down thing. Is the ADAS system itself similarly complex? Is it actually very simple? Does it just have access to search and an LLM, and it can build all of this complexity from very simple building blocks?
32:31Does the complexity need to co-evolve? How does that all work? Yeah. So as we just described, I think a very, very cool future direction is to have, say, agent all the way down, that we have a meta learning of the meta agent itself that we can have meta meta agent. But in this world, we are trying to demonstrate our simple idea first. Our meta-agent currently in our implementation is a hand-designed algorithm. And it's quite a simple agent where we just prompted some related information, like how the given framework looks like, what's your task, and what's the previous discovered agent. And then we ask it to build new agents.
33:28And then we allow a few rounds of data on the reflection to make it help to improving its own result. And basically that's it. We are applying a quite simple design of meta-agent. So it's a relatively simple framework that is heavily relying on the built-in intelligence in some external LLM. that given a rough code framework and an objective can spit out some code that kind of fills in the blanks and essentially is the design for a new agent. Yeah, yeah. And actually, for the results of, we are actually getting some results that we are very excited about because we actually observe, as you said, the complexity evolved from a very simple set of building blocks.
34:25So what we provided as a seed is like the existing manual design and very classic design patterns like chain of salt or reflection. And those seeds we give them. It's no, there are not many very complicated things in the seed. But then, as you mentioned, one of our experiments is in Arc. Arc challenge is a very popular reasoning challenge these days. And it tries to access the general intelligence of AI systems. And our funnel agent or the best performing agent we discovered in Arc through this ADAS algorithm is actually quite complicated. It has a lot of, as you said, a chain of thoughts, and it has a very complicated feedback system that they have multiple experts.
35:26They are experts in different aspects, like readability or simplicity of the answer. And then they also have something called a human-like reviewer. The problem is actually, say, okay, you are now emulating a human giving feedback to this solution. And we are very curious how those complex mechanisms evolved. Is it come from nowhere? The LM just programmed it? So when we look at the history of the discovery, and we are very surprised actually, this kind of feedback mechanism is a combination of three different stepping stones that occurs previously in the history. So that is something we are very excited about.
36:22So because we are doing a research direction called open-endedness, we try to learn in an open-ended way and discover a lot of building blocks. and then we can combine those building blocks into a more powerful solution. And then what we exactly see here is a very good, true example of how this kind of open-ended algorithm can discover all different kinds of stepping stones. Because when these stepping stones appear, they are not the best performing architecture. They themselves alone is not performing so good. and then but our algorithm automatically realized that although this currently does not perform so good but this has some potential if we combine them cadet together combine them previous discovery or step-by-step together we could get some more powerful agent and it tries so and it actually get quite good results.
37:28So this is something we are very surprised about. This design process is an iterative process, and with each successive iteration, the search agent has access to the prior history of discovered agents. Do you find that complexity in this environment in a setting is like monotonically increasing or does it exhibit for example like the ability to identify abstractions and like simplify and maybe an intermediate question is like how many iterations is it even stable like does letting it go for more iterations necessarily produce better results or does it is there a tipping point after which it gets unstable and it really only works for three or four iterations?
Read the full transcript
38:30Yeah, I think that's a very cool question according to how we design the search algorithm. There are many, many different ways we can design our search algorithm. One point is that we can... So there is a classic problem called exploration exploitation trade-off in optimization. So in one hand, we want to explore more, to explore around the search space, finding new design. Maybe that's not something related to previous discovery, but we want to try some novel things. But on the other hand, we want to get some more promising solution from the experience we learned from the previous discovery. So as you said, things might get a little bit more complicated with combining those previous discovered stepping stones.
39:29And for now, what we are doing is actually quite simple. We just, in a very fuzzy way, we ask it in the prompt. You can decide whether you want to explore some new idea or you want to learn from the archive to combine some existing stepping stones. And we discovered that in our experiments that you will kind of alternatively do exploration or exploitation in a quite fuzzy way. And sometimes you will want to explore some new ideas. sometimes we want to build the better solution from the previous discovery. But we will anticipate that one could investigate more about the design of this search algorithm.
40:20Do we want to have some problem designed for exploration or have some problem designed for exploitation? And how can we trade off between this thing? And this is quite interesting. And for the third generation, it's actually surprisingly very sample efficient. For our Arc challenge, we only do 25 iterations of design. So in total, it only designed like 25 different agents. But it already outperformed the state-of-the-art handcrafted agent by a substantial margin. and we think this is this efficiently thanks to the prior that foundation model have about foundation model themselves and agents and coding so that they don't need to during the search to learn those prior from scratch they already have those prior and so that they can from the generation one it's already starting to generate some interesting agent.
41:31And does each iteration result in an agent that performs better than previous iterations, necessarily? Not necessarily. Okay. We'll find that because in our problem, actually, we encourage it to do more exploration. Sometimes you will try to explore some new idea, And that's unlikely to have an immediate improvement on performance. But as we just described, all the discovery become the stepping stone for future new invention. Yeah. It's interesting that you mentioned that in your prompt, you're suggesting to it that it can combine previous solutions or not. and that it's not the combination that you described prior wasn't like some emerging property.
42:31It just did it on its own. It's part of the way you're prompting the system to use the history that you give it. Exactly, exactly. And actually, from the open-endedness research from our lab, we found it's a very straightforward and effective way to encourage the open-ended learning of a foundation model system is that we, in the prompt, we explicitly ask it, is this solution you generated novel? Is it interesting? And then let it reason about why this is novel and interesting. And then through this way, we can encourage them to discover some novel, interesting, and more exploratory solutions.
43:20So all that said, did you observe it discovering novel approaches? Not novel solutions. We just talked about that. But novel approaches, like were there any kind of emergent behaviors or observations or things that it did that weren't explicit in its prompting? Yeah. We actually discovered quite a lot of interesting design patterns emerging from our discovered agent. One design pattern we found is actually very useful in ARC challenge. Because for ARC challenge, the question is, you are given some input-output example of some grids. and then you need to find the patterns or find the transformation rules for those input-output grids.
44:19And then I give you a new grid and you need to answer me what the output grid looks like. So a common practice in this domain is actually not directly let the AI system answer what output grid looks like. It's actually let them to generate the code that implement the transformation you will learn from the future example, like two examples. And then we can actually test the generated code on these given examples and then apply them to the test question. And one design pattern emerging from the design of the ADAS algorithm is that they will try to generate a diverse set of a lot of program or a lot of transformation rules they predict.
45:15And they write different kind of Python program to try to predict what kind of transformation rule is actually going on here. And then they can evaluate on the given example input-output to see how good is the generated Python program. Because a good solution, a correct solution, should pass the existing example. And then it can pass the test example. So an emerging pattern is that in many, many different ways, they will generate a lot of candidate solution, and then they store it in the memory and then they rank it through the success of the input-output patterns we get as the examples. And then we can do a huge example among those good solutions.
46:20And once those patterns emerge from the search, it becomes the dominant patterns in our archive and every following architecture or agentic system try to follow that pattern. And we believe there's something quite interesting observation from our experiments. That is very interesting. I think I was also trying to get at like higher level, you know, meta behaviors or things that maybe a way to frame this question better is like, you know, you started off talking about your research interests in kind of the evolution of cognition and that kind of thing. Like if you were to map this current project to, you know, those kind of objectives and learning about how, you know, reasoning and cognition has evolved, like, you know, what might you do different or like how might you instrument your meta agents or what do you think there is to learn about the way that this system is working, if anything?
47:27Yeah. I think one very interesting thinking we are having is actually through this ADAS approach, it could be, well, one way is it could be a way that we can better understand our current foundation model through its automated design. because that's also happening in many AI-generating algorithms practice, say in neural architecture search. By observing different kinds of neural architecture that emerge from the search, we can actually learn more about some insights from the convolutional neural networks. And also the same here, because when we are applying the algorithm to different foundation models, because different foundation models, they have different things they are good at or they are not performing so well at.
48:24And then we might have different agentic system search from those foundation models they are using to evaluate those agentic systems. One example we discovered is quite interesting here is when we search. So all of our search for saving the cost, we use CPT 3.5. It's a relatively weaker model, but faster. And during the search of that, we discovered, as what I described, it has a very complex feedback mechanism. And that's the optimal solution we found during the search for GPT-3.5 model. But we zero-shot transfer that model to some more advanced models, say GPT-4 or Cloud 3.5. And we realized that this no longer be the optimal system design.
49:24The optimal system design becomes something that generates more candidate solutions and focuses more on the refinement steps and the ensemble steps. So that might hint that GPT-3.5 may have a weaker ability in evaluating themselves' solution so that they need a complex feedback loop to enhance that capability. And we will believe that if one tries to understand more about the behavior of foundation model, the aiders could be an interesting direction to explore. Another thing, as you said, my personal motivation of research, I'm very interested in how complexity involved in a system. For example, in our society or many of our human organizations, we have very complex patterns.
50:28I'm very curious, can this kind of pattern also evolve from a bottom-up approach? And I believe actually ADAS is a cool way to investigate this kind of observation because we will find a lot of relationship between the Asian work and the architecture of human organization or the human society. Say there is a work called MetaGPT. They try to emulate a software engineering company. They have different character or rules of Asian acting in the software development team. They have team leader. They have software developer. They have tester. They have designer. And then all those so-called agents collaborate with each other, communicate with each other.
51:25They try to design a software. Or we can see some artworks. I think it's from 2023. They try to use ChatGPT to simulate a human village. the interaction between different NPCs. And all those things, actually, they are behaving in an agentic system. They have a design of agentic system that how each module or each component interacts with each other. And if we can, if in ADAS we can allow the evolution or emergence of more complex design or interaction between different components, maybe we can understand more on how human organization or how human society evolved those complicated organization or architecture.
52:23Yeah, I'm finding the, as we discuss this and as I think about it, the idea of the flow of complexity or the evolution of complexity in the system to be a really interesting idea to explore. And also how that complexity might be abstracted by the system in some way. And if that helps it deal with more complexity, like maybe is there a way to prompt it to kind of consolidate what it's learned about past systems by creating some DSL for agents and then starting to think in the context of these predefined components that it creates or this DSL that it creates or how the system might manage the complexity beyond just keeping lines of code in memory.
53:27But then maybe, on the other hand, And maybe that's the advantage that the LLM has that we don't, is that it can keep a bunch of code in memory and not have to create those abstractions. It'd be interesting to see how that plays out. Yeah, exactly. Exactly. We believe that there are... I think we keep describing a lot of exciting direction. We can open up through this early investigation to the ADAS. How close do you see this... being to practical application? How far down these future paths do you need to go in order for this to have some utility? Yeah, I think, well, for us, we are, in this paper, we are more like academic orientated work.
54:16We are trying to demonstrate this scientific argument more. So all our algorithm design, all our framework, I tend to keep it as simple as possible. But recently, to my surprise, and we are very excited to see many developers in the agent community actually giving us some good feedback about they try to use our algorithm in their application. And one of the developers on Twitter, they said, quote, they get very, very interesting results after applying AIDA to their system And another developer posted that. They said, quote, after one hour of adaptation, we can simply apply the ADAS approach to their system.
55:14And we believe that that's very, very interesting to see more people trying out this approach in the practice, more practical application. And we are very, very excited to see how they go. Do you know if that means that they use an ATIS or an ATIS-like approach to design an agent and then they kind of extract that agent and apply it to their task? Or is the implication more like an ongoing optimization using ATIS as part of their solution? I'm not exactly sure what our people are doing, but what we anticipate is actually for one possible use case is when there is a new use case, when there is a new application, we can just use ADAS to automatically design the agent tech system to apply to those specific application.
56:25We just need to give it some basic building block and see how it goes, if this ADAS algorithm can build some workable solutions to effectively solve some of tasks. And then one other application we think is quite interesting is that if we have a working solution, like a genetic system that already can work on some tasks, and aiders can become something similar to continual learning because during the deployment of the authentic system, we will get a lot of feedback from the practice. Well, I think a very cool analogy I'm thinking this morning is that I'm reading the checklist you sent me about the podcast.
57:14And I actually wonder, what's the first version of the checklist? gets that. You must experience something in each bullet point so that you will ride those bullet points. So I think that kind of thing is very important in agentic systems or even in any human organization. That's the way we evolve. And that's exactly where the question was coming from. There's a clear place and design time. Like just, here's my problem, spit out an agent that will solve this or an agentic system, let's say, that will address this issue. But then also, so much of machine learning AI is like closing the loop with the data and continuously improving the system.
58:03And it seems like there would be a role for that as well. But also some risk, like you have stability risk, like you're continuously introducing new code, like you have to vet that code. just seems like there are really interesting challenges in there, but also opportunities. I think so. I think so. Very good. Well, Shungren, thanks so much for taking the time to share a bit about your research and all things AIDIS. Super interesting work. Thanks so much. Thank you so much for having me, Sam. Thank you.
58:46Thank you.
From the publisher
Today, we're joined by Shengran Hu, a PhD student at the University of British Columbia, to discuss Automated Design of Agentic Systems (ADAS), an approach focused on automatically creating agentic system designs. We explore the spectrum of agentic behaviors, the motivation for learning all aspects of agentic system design, the key components of the ADAS approach, and how it uses LLMs to design novel agent architectures in code. We also cover the iterative process of ADAS, its potential to shed light on the behavior of foundation models, the higher-level meta-behaviors that emerge in agentic systems, and how ADAS uncovers novel design patterns through emergent behaviors, particularly in complex tasks like the ARC challenge. Finally, we touch on the practical applications of ADAS and its potential use in system optimization for real-world tasks.
The complete show notes for this episode can be found at https://twimlai.com/go/700.




