AI Trends 2025: AI Agents and Multi-Agent Systems with Victor Dibia - #718

10 Feb 2025 · 1 h 45 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The TWIML AI Podcast Episode #718: AI Trends 2025: AI Agents and Multi-Agent Systems with Victor Dibia

Episode Overview In this episode of The TWIML AI Podcast, host Sam Charrington sits down with Victor Dibia, a principal research software engineer at Microsoft Research. The discussion revolves around the future of AI agents and multi-agent systems, focusing on key trends expected to shape 2025 and beyond.

Key Concepts Discussed

  1. Understanding AI Agents
  2. Definition of Agents: AI agents are defined as systems capable of reasoning, acting, communicating, and adapting.
  3. Agentic Foundation Models: The emergence of models that integrate multi-agent capabilities natively, enhancing the dynamics and interactions of agents.
  1. Developments in AI Agents
  2. Interface Agents: Introduction of agents that simulate human interaction with interfaces, moving beyond traditional API calls.
  3. Complex Task Management: A shift towards handling complex workflows rather than simple task chains, leading to the creation of frameworks that allow for generalist systems.
  1. Multi-Agent System Design Patterns
  2. Architectural Patterns: Exploration of design patterns for autonomous multi-agent systems, such as graph-based and message-driven architectures.
  3. Actor Model Pattern: Discussed the advantages of using the actor model for creating loosely coupled systems that communicate asynchronously.
  1. Challenges in Agent Evaluation
  2. Benchmarking: The complexities of evaluating agent performance and the need for end-to-end benchmarks that consider the trajectory of actions taken by agents rather than just final outputs.
  1. Future Directions (2025 and Beyond)
  2. Emerging Trends: Anticipated advancements in autonomous agents, including improved reasoning and memory capabilities.
  3. Human-Computer Interaction (HCI): The need for evolving user interfaces to accommodate proactive agents, allowing for interruptible interactions similar to those with human co-workers.
  1. Skills for AI Integration
  2. AI Literacy: The importance of software engineers developing skills around AI integration to leverage the potential of AI tools effectively.
  3. Adaptation: Understanding when to rely on AI tools versus traditional methods based on the context and complexity of tasks.

Key Takeaways

  • Reasoning and Planning: Reasoning ability is a significant differentiator between traditional software and agentic systems, enabling dynamic problem-solving.
  • Task Complexity Framework: A framework to determine when to employ multi-agent systems based on planning, expertise diversity, context management, and dynamic environments.
  • Implications for Workforce: A shift in job roles, particularly for junior engineers, as AI systems are increasingly capable of performing basic coding tasks.

Quotes

  • "Your task exists in a dynamic environment requiring adaptation."
  • "Not all actions are equal; some require human oversight."
  • "AI software engineering literacy creates stark differences in productivity."

Final Thoughts The podcast emphasizes the evolving role of AI agents in various fields and the necessity for individuals and organizations to adapt to these changes. As AI technologies become more integrated into daily workflows, understanding their implementation and limitations will be crucial for maximizing their potential.

For complete episode notes and further details, visit [TWIML AI Podcast Episode Page](https://twimlai.com/go/718).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00At some point, the agents were looking for some information. They were supposed to conduct a Google search and they failed. And you can imagine what they did. They wrote some code to send an email to request an FOIA, freedom of information. And email that organization to request that data. It's like, hey, we're conducting this research. We need this information. We can't find it. And by law, we're supposed to have access to it. And they crafted this email and they were going to use like an email API to sort of send it.

0:43All right, everyone, welcome to another episode of the TwiML AI podcast. I am your host, Sam Charrington. Today, I'm joined by Victor Dibia. Victor is Principal Research Software Engineer at Microsoft Research. And we've got a great conversation lined up for you today. we'll be reviewing Victor's take on the most important AI agent innovations in 2024 and what we should expect to see in the year to come. And of course, we'll also discuss Victor's work on multi-agent frameworks in general and Autogen in particular. Victor, welcome to the podcast. Thank you, Sam. It's great to be here. I'm super excited for the conversation.

1:22I know in your world as in mine, I guess more so in your world even than in mine, Agents is a hot topic that comes up all the time, and I know you've got a lot of interesting takes on that topic. Let's get started by having you share a little bit about your background. I'm a research software engineer. I work at Microsoft Research. I work specifically with a group called the Human AI Experiences Group. And essentially, by design, we are interested in scenarios where a human works in tandem with an AI model to solve tasks. In terms of background, my training is mostly software engineering, some work in HCI.

2:02So I have a master's in computer science information networking from Carnegie Mellon University. And I did a PhD in information systems at City University of Hong Kong. And my PhD is mostly focused in human-computer interaction, user behavior psychology, and how we can conduct a set of experiments that help us understand how people make decisions as they use technology tools and interfaces. And the whole idea is that we take all of that knowledge and we sort of apply it in designing better interfaces. Your group at Microsoft sounds like a traditional HCI research group, but you ended up kind of building Autogen.

2:48And I don't know if you own that as a product. I guess my point is it feels very productized as like a software infrastructure product coming out of this HCI group. So how did that all come about? Yeah. So I could talk about my own like personal path to agent. And I could talk about a little bit about the history of some early stories around like Autogen. So I started out right after my PhD. I started out as a postdoc at IBM Research over at New York, Yorktown Heights. And then I stayed on as a research staff member. And one of the things I really, I worked on there was, I did, I was with a HCI group and we worked very closely with a core mission learning group.

3:32And at the time, IBM had just come up with the cognitive service APIs. And we're building all this complex multimodal demos around like speech to text, text to speech, image recognition. using all of that together in like physical room scale experiences. One of the things I did back then was that I trained perhaps the first model for automated data visualization. So I don't know if you remember the sequence to sequence models. And so they were typically used for language translation and they were like the state of the art back then. And we showed that if you could represent visualizations in JSON, Vigalite, and you could represent like data in the same JSON specification.

4:13Then you could learn translations across the two. And so at runtime, we gave this system, this model, like some text. We sampled a couple of rows from the JSON dataset directly and it will generate a bunch of visualizations that were grounded on that. So that was really interesting. And if you think about it, there's a bit of action there, right? And so we generate like Vegas specification, we compile it and we give a visualization. And so after that, I spent some time at Cloudera, a traditionally big data warehousing company, but it had like a machine learning group and we built a lot of prototypes, did some customer consulting.

4:51Then after that, I joined Microsoft Research. And there I did some work on a tool called LIDAR. And so LIDAR is, again, an automated visualization tool, but had like a bit of a pipeline. So first of all, we got some data. We got like an LLM to generate a summary and rich summary of this data. Based on that, we got an LLM to generate a sort of hypothesis that made sense for this data. And then for each of these hypotheses, we could get an LLM to write code. On the background, we did like pre-processing, post-processing, we got code, we executed it. And we gave folks a bunch of visualizations. And so, and this was in 2000, just pretty early, I think late 2020 before chat GPT.

5:39And so it was an entire interface that did all this pipeline work on the backend. And it used the codec set of models. And for the most part, you can see some agentic behavior there. And so it's not just LLMs generating stuff. We're compiling code within preprocessing and post-processing. And so once you create a lot of these pipelines, you start to see some broader patterns. And so, you know, the next question is, can we go from that to a system that like, without having to manually sort of build out the exact steps in the pipeline, instead we define the set of agents and we give them a task and they sort of collaborated autonomously to sort of solve this problem.

6:23And one instantiation of work like that was Autogen. And so for the most part, Autogen started out like exploring the theory that maybe we could explore a new way to develop applications. And so instead of building specific pipelines to express the problem, to express the solution to a problem, we could instead define a set of agents with fairly broad capabilities. We give them a task and they could sort of collaborate to solve a problem. and there were a few really clever people within our broader group started to explore this, did a bunch of experiments, wrote a paper. So I wasn't an original author on that paper, but I worked very closely with that group and essentially that's sort of like what led to Autogen as a framework.

7:13And it's been about a year and some months, a lot of things have happened. We've got a bunch to dig into here And I don't necessarily want to belabor things by talking about like defining agents. But you mentioned that you heard me talking with Chip on that topic in a recent interview. And you had your own take on how agents are defined. I'd love to have you share that. Yeah. So I think a simple definition works. I think a lot of people are converging on the idea that like, if we take an LLM and we give it access to tools that let it take some action. So essentially this LLM can now act, then we have an agent.

7:57And from the software engineering point of view, that's like the basic instantiation. So you take an LLM, you give it some tool calling capabilities, you give it the ability to execute the results of the tool calls, and then you have an agent. And I feel like I'm happy with that base definition. In practice, it can be, if you want it to be a bit more precise, I think there might be a few other things. So you want something that has the ability to reason, has the ability to act with your tools, has some adaptation capabilities, in this case, perhaps memory. And then finally, some ability to communicate.

8:35And so it can send messages to other agents or to humans. And so it's reason, act, communicate, and adapt. And so this is, I think, like the four built-in blocks I would say sort of make up an agent. It seems like that reasoning ability is a key differentiator between, you know, traditional software system that uses large language models and an agentic system in my mind, in that it is what unlocks the ability for it to be dynamic as opposed to like a statically defined workflow. Yeah. So reasoning, yes, I do agree. I think one way to think of it is from the perspective of, let's say, planning.

9:25And so you get a task, you decompose it into a set of steps. And the idea is that if you succeed at executing each of the steps, you go from a state where the problem is unsolved, the task is unsolved, then you arrive at a step where the task is now solved. And I think you touched on the idea of dynamic. I think the interesting bit here is like, do you, you know, as you take each of these actions, you might, you know, the problem might exist in a dynamic space or a dynamic environment. And each time you take an action, it changes the environment. And in some cases, these changes might lead to errors or failure conditions.

10:08I think the key part is like a good like autonomous or good multi-agent system should have the ability to sort of recover from that and sort of either abandon the current plan or make adjustments and sort of keep making progress. and a really simple example is that let's say you're trying to solve a problem your agent writes some code, executes it there's some errors there it looks at the code, pays on the error it sort of modifies the code, executes it again it might be missing libraries, incurred arguments and if you do have this sort of behavior where every action has some results, some outcome and you can respond to that and then keep making progress then I feel that sort of speaks to the dynamic aspect of a multi-agent system.

10:56So we're going to dig into that, I think, in a lot more detail. But I think for this first part of the conversation, you put some thought into kind of what you, you know, from your perspective, the most important developments in agents over the past year or so and, you know, and how those set the stage for the upcoming year. And I wanted to start by digging into some of those. So I think the first thing on your list is about adoption. Yeah, let's start right there. So over the last five months, I figured it would be great to sort of keep track of what's changing. So what I did was that each time I saw like a research paper or a new product or a new tool, I kept a bunch of notes.

11:49And at the end of the year, December last year, I sort of figured, you know, what are high level categories here? And in terms of adoption, I feel like a lot of enterprises and teams adopted the 10 agents, but they did that with like some caveats. And so for the most part, what most people deployed last year or in the last year was mostly an LLM as a thin wrapper around existing APIs and tools. And so as opposed to like fully autonomous behavior where like at runtime, the agents can explore on unknown or like unscripted paths. Essentially, they just took the existing APIs, very, very tight structured action space.

12:37And all the LLM can do for the most part is to sort of make calls to these APIs. And this is a really good game plan because you get a lot of reliability out of that. I think the second thing on that little list was the rise of agent-native foundation models. And so about a year ago, we mostly had models like Gips 3.5 and the equivalent and Google from Google and Anthropik. And most of these models mostly focused on language modeling. And so they were writing text. They were writing code. and for the most part, they were mostly text in or in some cases, multi-model text in, image in, but only text out.

13:25And one of the things we saw in the last year was that these models were increasingly integrating multi-agentic capabilities just baked right into the model. So some of the capabilities like the ability to reflect things. So if you remember the React pattern where the idea is like You get the LLM to come up with the thoughts, get it to reflect on that, and then take more actions. And so we're seeing that with things like the O1 model family, the ability to just think and reflect is sort of just lifted up into the model itself. And, you know, you give the model a task, it does all this internal introspection and reasoning before you get like a result out.

14:06And also we saw things like natively model to model, in and out model. So I think the Gemini 2.0 model, so these things can take in text, image, video, audio, and the same model can sort of spit out results across all three modalities. And so I think that was the second interesting thing we saw in 2024. The third thing had to do with interface agents. And some other people have referred to them as computer use agents. And the idea is that as opposed to agents just calling tools, your API is a code. We now see like sort of a shift towards agents that act by simulating what humans do with interfaces.

14:51And so examples of that are agents that sort of solve tasks using a browser. And so I think about two days ago, we saw OpenAI release the operator agent. And the whole idea is that you could tell things like, you know, book a flight for me. and I'll go to, let's say, flights.google.com or click around, put in all the information, dates, source and destination, locations, all that stuff, and then probably come back at some point, get some feedback, get some confirmation, and get things done. Fun fact, on the origin land, we've built out systems or tools like this. And I think one of the excitements of the last two days was that once the operator came out, we said that he has like 40 lines of code and you could implement your operator using Autogen.

15:41Did the browser control framework already exist in the Autogen world? Yes, that's an excellent question. So in Autogen world, we have a bunch of presets. And so we have like a preset assistant agent that like, it's a classic. It just has an LLM model, a set of tools. But we also have this preset called a web surfer agent. And underneath, this agent drives a Chromium web browser. And it uses a multimodal, any multimodal LLM model. And so essentially, it has an action space about how to get work done on the browser, text input, clicking around, navigation, all of that. And essentially, it pretty much just acts by sort of driving and manipulating this browser.

16:30And it's a really nice, well-designed agent. It was done by one of my colleagues, a really brilliant fellow, Adam Fonny. And essentially, all you have to do is plug in this preset into your multi-agent team and you get all that capabilities. Nice. So interface agents? Yes. Yes. So interface agents. So we have that with Autogen, the web server agent in Autogen. We also have like tools from Anthropic. You know, they have like a computer use implementation. And there are a bunch of other like tools that sort of exist in that space. And so we saw a few of those sort of like advancements in 2024. The third thing had to do with complex tasks and frameworks.

17:17And so I did see that, you know, as a community, as a field, Langchain got us very, very far. So Lanchain showed how you could sort of get a set of deterministic steps, put them together and chain execute them. But, you know, there was a bit of appetite for more complex workflows. And we essentially more autonomous kind of workflows where like the task, you want a system that can address any task. And in fact, it reminds me of an article that like Bill Gates wrote about, I think, a year and a half ago, talking about how today, you know, back then, if you wanted to sort of, let's say, write an email, you went to like an email processing app, Outlook.

18:01If you want to do CRM stuff, you went to a CRM app. And if you wanted to do music stuff, you went to a music app. And he talked about the idea of an everything app, a unified interface, where you just express your task in natural language. And the system just, if it needed to manipulate or reach out to other systems, it did that. And so I feel there's a lot of value, a lot of time-saving, effort-saving value proposition there. And I think the community sort of started to resonate around that. And I think organically that has led to the design of frameworks like Odigen, LandGraph, CREI, Lama Index.

18:42because we want to figure out ways to provide good presets that help people build this sort of like generalist kind of systems. And I think, again, I might touch on what I mean by complex tasks, but I think that's one of the shifts that we saw in 2024. So beyond scripted deterministic pipelines to more autonomous like systems that could sort of address multiple disparate tasks. And then the final, I said the final update had to do with moving beyond just benchmarking models independently, but essentially extending to just end-to-end like evaluation of agentic systems and tasks that require action across multiple domains in the real world.

19:34And I think one of my favorite benchmarks there is the Gaia benchmark. And essentially, I think if I recall correctly, it's about 300 problems that as at the time of release, it looks really simple. These problems are really simple, really easy for humans to accomplish. But the best models at the time, I think was GPT-4, just failed really bad. I think they had like a 10 % pass rate there, if I recall correctly. What are some examples of the Gaia tasks? Yeah. So it might be things like,

20:09how long would it take Aliud Kipchoge to run across the Earth, let's say 50 times? Now to do something like that, you need to figure out, oh, who is Aliud Kipchoge? What's his maximum? He's a marathon record holder. So you need to figure out what's his speed. And if you figure, you know, what's the circumference of the Earth? then you need to do that little math that sort of puts everything together. And it might be things like, you know, like what did Sam Charentin say in the 78th minute of his 2025, I don't know, January 1st podcast. Now to do that, you need to go to YouTube, find like who's Sam Charentin, find the exact YouTube video that's being referenced, extract the transcript, then go to the 78th myth and figure it out.

21:01Now, as a human, this is really straightforward, frankly. But how does a machine go about stuff like this? If you really think about it, there are all kinds of ways where this machine might fail. And I remember a group, again, led by one of my colleagues, developed like a generalist agent system called Magentic One. And essentially for a long time, It held like the state-of-the-art performance on tasks like that. And so I think, and all of that process was extremely instructive. You know, we learned a lot about like how these things could feel, how, you know, the differences between how humans think about problems versus when machines, even the ones driven by sophisticated elements, try to address the same tasks.

21:48And so I think that was like the third and more interesting, the third interesting of these for 2024 or the fifth, sorry. So let's dig into these. I have a bunch of questions across this list. Maybe let's start with these kind of agent-native foundation models, you call them.

22:15Talk a little bit more about the you know kind of the way you think of them as agents i think i don't the way i i guess my personal experience is that i i originally thought of them in a very agentic way but then um you know as we've seen with like deep seek you know showing you the you know the thought tokens it it's it seems less agentic in a sense. Does that make sense? I guess it's like, it seems more like a straightforward, but slightly more complex application of, you know, traditional LLMs in some way. Yeah. So I guess, I guess what you hinted at is like, you know, if you didn't see the thought tokens, then it looked like it was doing something more clever.

23:13But when you saw the thought tokens, it was just an autoregressive model. So it's just predicting the very next token. So I guess the interesting thing is what is different when the model is primed to explore like you know the iterative thinking process as opposed to just um um generating the next likely token um i think from the from the human behavioral psychology perspective and you know i say this with caution llms not humans um they're not like humans in any form but if you if you think about it, you know, there's the whole concept of system one thinking, system two thinking, thinking fast and slow.

24:04And there's a whole idea of like, you know, for things that are simple, as humans, we've adapted to use heuristics, right? So if I did ask you, Sam, what's your birthday? You don't need to think about it. You tell me. Or like, you know, like what time is it? Or is it morning and evening? You know the answer to that. So you can rely on heuristics. So these things are like right there at the top of your mind. But if I did ask you, like, you know, like the question earlier, how long will it take a marathon runner to run around the earth, like 50 times? Now, this requires a bit more investment, a bit more computation and effort.

24:40And a lot of people did complain earlier that, like, say about a year and a half ago, if you ask the model these two questions, you'll take exactly the same amount of time to give a response. And there's just something not right about it that like some two problems so different, so complex, we are assigned about the same effort. And I think a lot of that has informed some of the work in test time compute. And the whole idea here with the O1 Resin and Diff6 family models is we want a way to communicate to the model or the system that like some problems perhaps require a bit more investment, more compute investment than others.

25:23And it turns out that it does work when you design the system that way, you just get better results. underneath is still an autoregressive model doing autoregressive stuff. But it just turns out that the setup, the problem setup sort of results in better results for thinking and reasoning style problems. And so the advantages of those types of models for, you know, like complex information gathering and presentation, report generation, those kinds of tasks is pretty clear. Are you seeing those reasoning advantages play out in terms of like the planning style of reasoning that's important in making complex agentic systems work?

26:15Yeah, yes. So one of the good things about like, let's say a tool like Autogen is, let's say you could decompose your problem and express them as agents. So you could have an agent that's explicitly just focused on planning. So for example, I mentioned earlier the Magentic One paper. So the way that that setup was done was that like we had, no, we argued for the design of a generalist system that can address multiple different types of tasks. And we tested them across multiple agentic benchmarks, the exact same system. So nothing was fine tuned for a specific system. And they were composed of four agents.

26:52So the first was an orchestrator or a planner. So all it did was it took a task and it would decompose into a plan and assign steps in the plan to other agents. And there were, I think, four other agents, something called a coder. All it did was write code. There was one that was a computer terminal. All it did was execute code. There was a Web Surfer agent. Essentially, the task needed interaction with the websites. And then there was a File Surfer agent. So if you need to open things like video files and image files or PowerPoint presentations, that sort of thing. And the core idea is that for, let's say, the orchestrator, for each of these things, you could assign them different models.

27:35And so for the orchestrator, I think my theory is that something like that that's meant to, like, reason through the problem, do some sort of task decomposition, assign steps to different, like, agents.

27:54an agent like that really would benefit a lot from like some of these sort of like test time compute or reasoning models now the other is something like let's say file software all it does is just has a bunch of tools that lets it interact with files probably not a lot of benefit there and is that intuition or have you seen benchmarks that you know take a system like a a Magentic one and insert a reasoning agent for that orchestrator step? We haven't released any results. Let's say, let me use the word release. We haven't released any results yet, but early experiments and some of, some of, I think, I don't remember other papers off the top of my head, but I think I have seen a few where just dropping in the reasoning model did give like a significant, significant boost.

28:47I guess I wanted to poke at like there's maybe a nuanced difference between seeing a significant boost and like unlocking a whole new area of capability. Do you see these types of models doing the latter? I wouldn't see exactly a whole new type of capability. I think it's just maybe performance improvement. I think a lot of the types of problems we're thinking of solving the system is still the same class of problems. I think the more interesting thing here is if we look at this system from the perspective of failure modes, bad plans or the ability to come up with good plans on the first try is like a significant like performance issue for like this sort of like autonomous systems and from that perspective if you get something that like reasons well comes up with the plan or good plan on the first try then you get some benefit there but the type of problem hasn't changed it's not like we're suddenly, I don't know, but suddenly doing new types of things.

Read the full transcript

30:06It's just that like we're getting like reliability or performance, maybe even safety improvements where possible. So with the Magentic one work, you mentioned that one of those agentic types was a coder. Was that primarily used in the context of coding problems or was it code that was generated in the process of solving other types of problems. I've come across several different papers that use code as this intermediary for planning and other things. And I find that a really interesting and compelling way to use code. Yeah, it's a really, really, really good question.

30:53So it brings me to like how I think about tools. I think there are two types of tools. So there are like tasks specific on our domain tools, and then there are general purpose tools. And I feel like a code interpreter is a type of general purpose tool. Some problems, a lot of problems can be, the solution to a lot of problems can be expressed as code. and getting the key point here is getting the orchestrator to figure out, okay, the support problem could be solved really well when expressed as code and then getting the coder to sort of write that code and then getting the code interpreter to execute it is like an emerging pattern.

31:41So to answer your question, it wasn't just solving like software engineering task type of problems. It was mostly like, hey, no, here's a task. So it was more code interpreter than code generation. Yeah, yeah. That would be like a good focus. And of course, there are caveats there. So if you have a system that has this very wide action space, then it can do a lot of interesting, maybe even unusual things. So an example that we, a funny example that we like to talk about, like in the Magent, and we talk about in the Magent, to go on paper. At some point, the agents were looking for some information.

32:22They were supposed to conduct a Google search and they failed to find that information. And you can imagine what they did. They wrote some code to send an email to request an FOIA, essentially to send freedom of information and email that organization to request that data. It's like, hey, we're conducting this research. We need this information. We can't find it. And by law, we're supposed to have access to it. And they crafted this email and they were going to use like an email API to sort of send it. I'm imagining the agent sending it to like trying to send it to Google as opposed to like a government organization or something.

33:03Yeah. So it's a fun fact, but it is true if you don't constrain like the action space of what these models can do. because code is just this really expressive thing. They can take any kind of action, express it as code, and then it could lead to like things that like, and the way you solve this is that like the orchestrator has like, can make some high level decisions as to like, is the task being stalled? Is it going the wrong direction? Sort of metacognition, right? As these agents act, the orchestrator is sort of inspecting the progress and is saying things like, you know, are we stalled? Was our maximum stall count?

33:41And... Should we reset, modify the plan, abandon this route and take a separate route? So, yeah, there are caveats to using like general purpose tools. And the way you talked about those aspects of the orchestrators, maybe a segue into talking about these complex tasks and frameworks, which was one of your items. Specifically for those types of parameters you were describing, are those things that the user of a framework like Autogen, like, are they thinking about them? Are they setting parameters? Are they coding them? Like, how do you manage the level of abstraction that someone working, you know, trying to build an agentic system to tackle complex tasks has to deal with?

34:33So a question like this, you know, know ties into like slightly how how do you design frameworks um how the developers think um and i could tell you a little bit about how like as an hci guy i feel like i've opened the box oh yeah um and i i could i could be a bit more practical tell you about how we are pushing with other gen and so in other gen they're currently there are two api levels so there's a core api and the idea is that like, it mostly just provides you with the bare bones capabilities for things like just message delivery. So anything could be an agent. You could define anything as an agent, inherit from base class.

35:20And the only thing you're required to do is to modify a method that says, you know, the agent has received the message. What does it do? And so it could be as simple as it receives a message, it does nothing, or it sends back the exact same message and that's all. no opinion to whatever the agent does when it receives the message. The developer is responsible for that. And essentially all we guarantee is that there's a concept of a runtime. When you define your agent, the runtime spins up, creates instances of this agent, enables message delivery, and these agents might live across multiple machines, they might be on the same machine, and that's all.

36:03but for most developers, this is still too low level. And so we have another API called Agent Chat. And Sam, you're from the old world. You probably remember Keras, where Keras was like this high-level abstraction, very intuitive, but on beneath, it could run a TensorFlow backend or a PyTorch backend or a Jax backend. And so think of Agent Chat as Keras of this world. And the kind of presets we have there is things like a basic assistant agent. And so this thing is what I think is the fundamental representation of a basic agent. And so it can take an LLM, it can take a list of tools, it can take a list of memory banks, and essentially that's the standard interface or the standard definition there.

36:55And so model client, a list of tools, which could be functions, It could be anything. And a set of like memory banks, so just to enable a rag or like just-in-time retrieval of the information. And we have another preset, which is like the web server agent, which essentially just all the things you need to drive a web browsing and sort of accomplish tasks using that. And then we have, I think, one or two other presets, not very important. So that's at the agent level. Then how do these things sort of collaborate? So we have the concept of teams. And so think of them as containers that you put these agents into, and it mostly governs the order in which messages flows across these agents.

37:39And so we have a preset, something called a round robin team. And what it does is that once a task comes in, it just sort of sends messages across each of the agents until some termination condition is met. And then the final abstraction we have at the team level is a termination condition, which can be really, really tricky. It's like, this guy said exploring a task, how do they know when it's done? And so we have abstractions like text message termination. So it means if any of the agents, you could define in their prompt, their behavior, they might constantly sort of inspect the state of the task.

38:17And if the task is done, they might respond with a word like terminate. And so let's say you scan for that in the messages, you decide that things are done. It could be budget-based. So timeout based or a maximum number of tokens used or maximum number of steps. It could be some external signals or something external, just monitoring the state of the task and then sends like a termination condition. And you can compose all of these things in all end combinations. and if you took all these presets, agents with all the stuff inside, teams, round robin, selector group chat, graph-based selections, all of that, and termination conditions and you put all of that together, we are seeing that that has been a powerful way to express autonomous multi-agent systems.

39:09Popping up a level, you started talking about the way developers think about problems and things like that. And it sounds like at its core, the way Autogen is organized is kind of message based and there are other approaches that are graph based, other approaches that are, you know, I don't know of specific other ones, but there seems to be like, you know, messaging graph is like one big, you know, difference in paradigm. You know, are there others and, you know, why do you think message is better, you know, historically, like, or traditional software, it's like loose coupling is an advantage of message as opposed to other things.

39:53Like, talk us through that whole abstraction, you know, thinking. I'll talk about two things. So first of all, right now, the current version of AutoGen is based on the message passing like architecture and how we got there. And the second thing I'll talk about is like some of the emerging patterns. I alluded to some of that, but I'd like to structure it a bit more. Emerging patterns we're seeing for building sort of autonomous multi-agent systems. So about a year and a half ago, when we released the first version of Autogen, the interesting thing is, you know, it was built on this idea of conversational programming.

40:33And so the idea is that like to solve a task, we just get these agents to each act. So each time they act, they have like a shared conversation history or list. And LLMs were actually being fine-tuned and there's still a paradigm, the chat completion paradigm where like every time you ask a question, you append it to like a long chat and the model just gets, you know, it's been trained to use all the context. That's where its context is coming from, yeah. Exactly, all of that context. So essentially, solving tasks is just all about building context. It just extends. But the problem was that this thing was a list that lived in memory.

41:14And so all the agents had this list, literally. There's this list and they all appended stuff to it. So each time they executed code or they responded to a message, they all appended into that list of memory. Now, if you do that, you can build systems where the agents live on multiple machines. and this is a really common production requirement. Like you want the agents living in a separate machine. They have like security boundaries, have access to information that like nobody else should be able to have access to. But, you know, if your design is at like everything lives in memory, you can do that.

41:48In addition to that, you also want an asynchronous stack because each time the agents act, an action can take an arbitrary amount of time and you want a scenario where like these things can run in the background while the rest of the system sort of continues. And so quickly we ran into all of these sort of issues and a solution to that is the actor model. And so you sort of reference that where, you know, you treat every element in the system as an actor and they're loosely decoupled from everything else and they only communicate via messages. And you send these asynchronous messages, the messages can be delivered in arbitrary order.

42:38And if you do that, you can compose this message sending behavior into all kinds of complex patterns. You can have agents or systems that live on multiple machines as long as they can connect to the same communication layer or message passing or message delivery layer. And even from the application development point of view, you can build this async applications, websites, UIs, communicating with things like Teams, Slack, where like, you know, messages are just sent and received asynchronously. So things like, you know, enabling true like distributed agents is one of the driving principles there.

43:24And also just application, just better integration with external applications are one of several benefits that you get from a message-driven, async, actor model kind of paradigm. So hopefully that sort of provides some background in why message passing is a good idea. Of course, it has its own complexities, message delivery, ordering. It's really hard to think about and debug async code. But that's what a framework is there for. It's meant to help with a lot of those issues. So that's one thing. The second part I want to talk about is, you know, I mentioned the idea of graphs and chains and all of that.

44:08So I feel like there are two high-level patterns that we are sort of observing in the multi-agent space. So the first is control flow patterns. And so if you have multiple agents, how do you determine the order in which they act or the other message delivery across each of these agents. And the simplest version of it is a deterministic chain where like you say, you know, I want something that like generates, like, I don't know, finds me some news articles every day. So you might just keep it simple. It has three steps. You have an agent that like takes in my query. It takes in my request, generates like a web search query.

44:49You have another one that makes a request to Bing or Google. And you have a third one that downloads all of the data and all the results and summarizes it into some simple outcome. Extremely simple. A simple chain like that is also a graph, but just a very simple version of a graph. And the core idea here is the developer already has a clear idea of exactly what they want the system to do when you run this graph. Now the graph could get a little bit more complex Which is what tools like line graphs support So you enable like conditional edges You enable loops By the end of the day you can still make very clear Deterministic predictions about what would happen Before you execute the graph So you know that this graph is directed It's acyclic And it will always come to some end goal Or one of several end goals which are all good things.

45:46And so that's like the first two patterns, so simple chains, graphs. And then the third piece is more implicit planning or group chat kind of thing, where we model the solution to the problem, not as graphs, but just by a conversation history, which is where Odagen started. and this agent sort of, there's no predefined path, but there are some structures around like, you know, you might have some round rubbing communication flow or we might get an LLM to decide just in time which agent speaks next or takes another turn. And I'll say that's the second plan. And this implicit plan, group chat, shared context kind of thing.

46:38It can still work across like multiple distributed agents. It just is a bit more non-deterministic and we still see a lot of failures. And I think as a research group, that's the part that's really, really interesting because like we want to get a point where we understand this thing really, really well and make it work for everyone because this is how you get like truly autonomous, like new step function, increases in system behavior, at least in my opinion. And then there's the opportunity to implement things like metacognition, to just thinking about thinking. And so I mentioned magentic ones, an early version of that.

47:20There's a concept of an inner loop and an outer loop. And so every time an agent takes a step, just the orchestra sort of takes a pause, is a task making progress. If yes, okay, let's keep on going. If no, we increase some stall counter, like, hey, you know, we're stalled. And after some threshold, we just reset all of the agents. We agree that like we are on a bad trajectory. We might be summarized what went well. We keep it. We use that formula to new plan with some notes and then we sort of explore new trajectory. So a bit of metacognition, reflection, that kind of thing. And then the final set of patterns are around task management, around like first human delegation.

48:08And so not all actions are equal. And so if there was, if an agent came and said things like, hey, I'm going to download this file to disk, no big deal, right? It downloads the file. But if you said something like, hey, I'm going to send, I don't know, I'm going to transfer some money to Sam Charrington's account using ACH. Now, I want to know about that, right? Right, right, right. I want to know about that. So there's the idea of like... I want to know about that too. Yeah, you want Sam to know about it, right? So he doesn't get worried like, hey, you know, where did all this money come from?

48:40So there's the idea of like the inherent risk associated with actions. So we need to be able to quantify that and maybe figure out if we want to delegate. And also there's the idea of how do we know when the task ends? And I've hinted on that earlier. So those are like the two high-level patterns, control flow and task management are patterns that I'm sort of seeing. And kind of putting it on the hat of someone who, you know, is thinking very pragmatically is, you know, maybe at a startup and is building something or wants to build something. Is it too reductive to translate what you said or a little bit of what you said into like Autogen is a research project and, you know, they're doing things the way to make it, you know, interesting and complex for research and open up new avenues of research.

49:31and I might want to go a different direction if I am really just trying to solve a problem. So maybe six months ago, that would have been correct. But I think two weeks ago, we released a new version of Autogen. If you search for it, I know this is a little bit confusing for some people who have used Autogen, but there's a new version called Autogen 0.4 and it's based on this new asynchronous message delivery behavior I mentioned earlier. And we've put a lot of effort there to make it production ready. So like I mentioned, there's a core API, right? And so I'd say my suggestion is that if you're a startup trying to build like a production level application, just take a look at the core API.

50:21There's a chance that like your business logic is so niche that maybe some of the presets we have in the higher level API might not be directly or immediately applicable. But take a look at the core API. Use it as a core building block because at the end of the day, you still need like message delivery capabilities. You need some of the utilities around tools, around like model clients, that sort of thing. So I'd say the core API is where you want to sort of invest time in. And also another fun fact that I probably will tell you is that when we build Autogen, you know, again, as a research group, we're more interested in autonomous behavior.

51:07And this is a good thing. Everything starts out as research. Somebody has to, like, explore it before, like, we solve all the bugs and make it work. And so, yeah. And so we started out with the autonomous thing. But then we saw that, like, this is just my own assessment, but like I think about 70 % of the people who come here to use Autogen, they already know what they want to do. They know like my problem has like six steps. And essentially what they're trying to do is they're trying to shoehorn Autogen to just do that six steps that they're there to do. And we're working very hard to make things like that possible.

51:45And these are all great ideas. I think a lot of the reality is just that a lot of the use cases, So for a startup, right, let's say you're in finance, you know exactly what the problem domain is like. It's pretty structured. You're not trying to build everything up, right? And so from that perspective, what you want to build is some type of workflow or pipeline or chain or graph, maybe with some dynamic behavior here and there, as opposed to something fully autonomous. And so go ahead, try out the core API, use it to express your business problem. And I think that's where you get the most benefit from.

52:23And as the space grows and we figure out like really how to get this autonomous thing to work really, really well, then maybe at the future time, the higher level API or the autonomous exploration kind of agent might be a better fit. I think just being pragmatic, this would be my thought process here. Going back to, you know, our review of 2024, you mentioned interface agents and computer use. I think that is, you know, really starting to capture a lot of kind of maybe energy or imagination is maybe even a better word because like you can like really visually see the computer like doing tasks that I don't want to do.

53:12book my flights, you know, do my grocery shopping.

53:18I'm thinking a little bit of another recent interview with Dan Jeffries, where he talked about actually how hard it is to do those kinds of tasks. You know, he used an example of like trying to get an agent to book a flight on Google Flights and like pulling up the calendar and you've got to click the numbers and like getting the agent to get, you know, localized on those numbers is just really hard. Like I think, you know, to maybe kind of open up this part of the conversation, I'm just curious, like, your take on, you know, where we are with regard to computer use and, you know, what needs to happen to evolve it for it to be.

54:03Like how practically useful is it, you know, in your opinion, like, just kind of what's your rundown of like the state of play with regard to like agents controlling browsers? So if you've ever done like a user study where you get like people to get, you get people to use these agents to accomplish tasks. One very funny thing that tends to happen is that there's something really unsatisfying about a human just sitting and watching an agent struggle is like, okay, I'm going to open this browser. Then I'm going to click this, click that. The divisor on some parts, some people like, hey, that's cool.

54:41That's really fun. But some of the people like, hey, you know, like I could do that in like a quarter of the time, you know, that sort of thing. And then sometimes, just like you mentioned, like it feels that some funny like step and then it has to restart and all that. So I think this is all growing pains. As a researcher, I have very high tolerance with stuff like this. Of course, as a product designer, this is not the case. But I think there are two things, right? So we are seeing like specialized models just for like just UI interaction. So some of my colleagues, also at Microsoft Research, released a model, I think, a couple of months ago, something called OmniParser.

55:28And essentially, it's a multimodal model. And it's just fine-tuned to predict the bounding boxes with fairly high quality of interactable elements on any screen, any UI. And again, if we treat this as an engineering problem, as we gather more data about like how humans like sort of interact. Think of it as like all the good work that OpenAI and Co. did on supervised fine-tuning. So you just invested a ton of time in just figuring out like, you know, getting a bunch of experts to interact with the GPT models, assembling this SFC data or for reinforcement learning from human feedback. I feel like we are going in that direction where like these models develop intuition similar to how human beings sort of act.

56:26And so there's some things you'll do. If you see an ad, you close it immediately before you do anything. So like you don't try to like click other things while an ad is open. You know, oh, that's an ad. I got to click it. It's a cookie pop-up banner. I got to get it out of the way. And so we need more of those types of examples. Some icons are really small. There's some models optimized for that. I think also about a week ago, there was a new model from ByteBand, something called UI TARS. And also just significant, I see a lot of progress. So the benchmarks are just getting better. And I think we'll see more of that.

57:03So, and models like that, they have to specialize like traditional, just generic multi-model models into that world there. So just having specialized like UI interaction models is becoming a thing. So I think that's how we get better there. I think there's also a lot of work around like memory and adaptation. And so we need to figure out ways to get the model to do things like remember preferences automatically. Say anytime we give feedback or anytime like it explores a path that seems like it works well. we need ways to sort of serialize and recover and retrieve that sort of successful trajectory in the future so short answer I think that accuracy will increase latency will still be a problem and from that perspective I doubt that the right model to use these things is to have a human sit and watch I think you really want this thing doing back office stuff like you know um i i i am a i don't know i am a customer sales consultant uh a new ticket comes in just without me being there i want this thing to like open up salesforce open up our custom i don't know ticket management software do all the data transfer open up an excel sheet write all these things down and then just let me no one is done.

58:38I'm not sure that the right model is that like, no, the human actually actively supervises it. Maybe there could be like one training or one like learn by demonstration fees and then every other thing should be like sort of in the background. And the truth is if an agent takes a minute to do something that I can do in 30 seconds, I think this is fine because 30 seconds of my time is worth 100 times of 30 seconds of one minute of the agent's time. I think people get hung up on that. 30 seconds of your time is really worth a lot more than all the compute that an agent needs to use, even if it uses just slightly more time.

59:27Yeah, yeah. Yeah, I used to kind of think that... bit. And maybe it was more true than when we were further from seeing computer use and things like this in a while. But like, you know, for any problem that was, you know, sufficiently valuable enough to, you know, point an agent at it, you know, you may as well just like build the API integration or something like that. Your Salesforce and ticketing example is a great example. but you know it's is clear that as you know the cost and capability and complexity of like deploying an agent that works through a web browser you know goes down and they become more reliable like there are way more integration problems out there than we have the ability to tackle all of them so yeah there's certainly a place for this back office type of work that you're describing even though as an engineer it's painfully less efficient than just like hitting an API.

1:00:30I completely agree with you. This whole thing is an anti-pattern.

1:00:40There's something not right there. The least efficient way to solve this problem.

1:00:49I agree with you. I think maybe reduction in costs and maybe increased reliability and the added flexibility, right? There's just not enough. The surface of all the things you need to build APIs for, it's pretty large. So, yeah. Yeah, I guess another take that I've had is that historically, a lot of integration problems aren't like technical problems they're like political problems the example i always use is like you know forever you couldn't get southwest flights on google flights and that's not because they couldn't integrate them it's because southwest didn't want their flights there um do you see um do you see pushback starting to happen you know on agent systems, like, you know, in the traditional, like in the search engine world, there's robots.txt, like, you know, keep your bots away from my content.

1:01:53You know, do you see a thing like that starting to, or do you see a future in which people are trying to prevent robots from, well, I guess they're already trying to prevent robots from accessing their sites, but like, how does it, I guess, how does it change given the kind of technology that, you know, agentic bots are based on. Yeah, you bring up a really important point, and I think it's something that we all should... Well, I mean, I guess technology, from the technology standpoint, we'll all adapt to it in some way. But I feel I agree with you at some point, we'll have something like agents.txt, just have robots.txt.

1:02:33And it might be things like, hey, no, we don't want robots or interface agents, like sort of interacting with any of our interfaces, or it might be just a specification for what we expect the agents to do or the right way to sort of interact with our interface. And for that matter, you know, beyond agents.txt, like, you know, we'll have Cloudflare and, you know, all the other services that people use to try to prevent, you know, non-human access to their things? Oh, we already seen stuff like that. We already seen a lot of websites that just try to auto-detect if you're using a playwright instance or if it's actually human.

1:03:20And essentially, the website is just blocked completely. And there's a chance that like a big chunk of the internet will just, which is because it's in some ways, it's like costs for the folks running the websites. And I think at some point we need like some clear language, a standard around like web interfaces or any interface in general that's designed for human versus an agent. And it brings me to something that I know I've been thinking about the concept of like agentic noise. and agent technology is just thinking through like, you know, at some point we'll have like a lot of agents acting in the digital world on behalf of humans.

1:04:07And if these things are not implemented well, they can sort of compete in unusual ways for the human bandwidth. And so as a human, you know, we all have finite bandwidth. And as we navigate the digital world, sort of everything is sort of vying or struggling to sort of get a chunk of that bandwidth. And I think as humans, we want to optimize for like, unregressed interactions or in some way that might also translate to interactions with other humans or interactions with like, I don't know, like just high quality, like high quality, like artifacts. And if we just have like agents sort of loose on the internet, then in some ways they might sort of capture like bandwidth, or I don't know, like they might like, you know, capture attention that like.

1:05:01Is that different or in what ways is that different from like the spam problem, right? It's like as the cost of sending emails and spending, sending text messages, you know, has dropped to essentially zero now, like there's this kind of, you know, attention tax. well, it is this theoretical, at least attention tax in our inboxes and our text message boxes, you know, that kind of vies for our attention, you know, and, you know, per or in line with our previous conversation, you know, that, you know, technology has been created to fight that, right? And so now we don't even think about spam anymore because it's in this, you know, hinterland in our inbox that we never even bother to check anymore.

1:05:45I think the point is we would need to sort of as we design the agents, we need to figure out the technology that curates our interaction with these agents. And following the email example, we have filters that figure out what a good email is and what a bad email is. And so we will need technology that does the same type of filtering as to what are good agentic interactions and what are bad ones. And sometimes these things can have like all these other secondary effects, right? So imagine that like, you know, if you've interacted with some government websites, there's the notion of Whitley saying, you know, to get an appointment, you go to some calendar, you sort of like click around, you get an appointment.

1:06:31And I imagine like there were a couple of ambitious agents that just went and took up all the appointments. Now, that's a real problem, you know, like it's sort of interfering with the social contracts that like we have when we interact with all the systems and and we need to figure out ways to evolve evolve and sort of it might be like some new types of capture just like you mentioned some new types of filtering some new types of humanness confirmation that sort of thing but either way we will need to sort of evolve and get better at that you mentioned the rise of kind of these end to end agentic benchmarks.

1:07:16Are there and you mentioned, I forget the name Gaia, was that the name of the benchmark that you refer to? Are there? Are there frameworks or methodologies that you're seeing people using to kind of benchmark their own tasks with agent performance on their own tasks as opposed to benchmark tasks? So evaluation is a whole, it's a whole like field. It's a whole thing. And yeah, it's a whole kind of what I'm trying to, what I'm trying to ask is like, what, you know, how does all of the kind of energy and work that's going into eval like apply to agents and in particular like end to end real world agentic performance?

1:08:00Yeah. Um, so I think like it, whether we like it or not, I think a lot of the evaluation here is still, still follows a lot of them as a judge kind of thing. So no, you could benchmark things and some, some problems have like objective like results so far on the Gaia benchmark. You might have things like, you know, how long does it take? So the answer is a number. But then the result that the agent comes up is... I don't even know that... I don't even know that... I don't... Go ahead, go ahead. I was just going to say, like, one thought that I had earlier really questioned the objectivity of the result.

1:08:48Like the marathon runner. Yes, like I can divide the speed of the marathon runner by the circumference of the earth. But like that gives me one answer. But do I expect the agent to take into account wakeful, you know, time versus sleep time? Do I expect the agent to take into account routing? Do I expect it to take into account transportation? Like, I don't know that that objective number is actually objective. Yeah, exactly. So, so you can have one run where the agent says, you know, I'm going to assume that we run exactly around the equator. right? And we use that exact distance. You might have another agent that's a bit more clever and say, hey, no, if we look at the map and make estimations around like from this exact point to this other exact point, this is the only land travel route.

1:09:43Completely different answer. Just like you mentioned, we might have another agent that says things like, okay, I'm going to take into consideration that half the day this person is going to take a nap. Completely different answer and so it becomes very hard to do stuff like that and so the the key point here is you shouldn't just benchmark the final answer you should benchmark the entire trajectory and the best that you can do is to have like an lm as a judge where you define the criteria for evaluation and so it's like it might be something like did this does this result So is it based on like solid sound or reasonable assumptions?

1:10:25Is the calculation correct? It's all this like fuzzy, fuzzy logic sort of evaluation criteria that you probably might adopt to sort of benchmark how the system behaves. Yeah, it's like what you're really trying to do is like you're trying to benchmark its reasoning ability, but in a way that I think is different from reasoning benchmarks and that those are all benchmarking, those are all comparing against an outcome, an answer. And what you want to benchmark is like a thought process. It's like an interview where you're given a really hard problem and you're expected to talk through how you get it.

1:11:04That's what you want to do. And the point is not the final thing you come up with. The point is whether you're thinking in a reasonable, logical manner. I mean, you could have a couple of rubrics, right? So all interviews go into rubrics. So essentially what you're doing as a judge is that you're defining the rubrics. And more importantly, right, how do you interpret the results of these things? Every number itself is not very meaningful, but it's relative numbers. So if you start off with a B's version of the system and you get your first set of numbers. Now, what you want is that as you tweak the system the number increases, right?

1:11:45You don't care about the absolute number. You just care about the fact that like, you know, this number is indicative of progress in some direction. And essentially it's not the number itself that matters. It's mostly like, as I make changes to the system, do I see changes in a direction that like I care about? So I think this is a common thing in this space and designing the right, like, structure. And as we get, like, cheaper open source, like, reasoning models like DeepSeek and co, we get a chance to be a bit more creative in how we sort of evaluate, like, you do, like, LLM as a judge. And I think a lot of people are beginning to sort of integrate, like, these sort of approaches to their own, like, business problems, you know.

1:12:37Invest a bunch of time, come up with those rubrics, structure it well, and then optimize, you know. I think Jason Liu, who's also been on the podcast, has talked about like, you know, being a bit creative and coming up with like low-level metrics that are not as expensive to compute. Also, again, this can be like good, like relative numbers that you can use to sort of make sense of the direction of the impact of changes to your systems as you sort of iterate. Yeah. We talked a bit, quite a bit about multi-agent systems and some of the, you know, we talked about architectural considerations and abstractions.

1:13:22And I'm wondering if we've, if there are other aspects of that that are worth digging into. I feel like we kind of dug into specifics, but we didn't really talk about kind of broad motivation of multi-agent, you know, and when the complexity of multi-agent systems is warranted from a use case perspective. Yeah. Yeah, so the idea of multi-agent or autonomous systems is really attractive. And one of the downsides is that you might see teams just hurry, just like rush to try to apply these things, even when it might not be the best tool for the task. Of course, choosing what to use should always be a careful scientific process.

1:14:14Like what is your business problem? and does it fit the parameters of the tool? So at the end of the day, the focus should always be solving the business problem or user problem. Now, in determining when to use like an autonomous multi-agent system, I have found that like a good framework to use is something I call like the complex task framework. And so my thesis is that like multi-agent, autonomous multi-agent systems are sort of good if your task is complex. And what does that mean? I think there are four high-level areas that I sort of ask people to sort of think through. The first is planning.

1:14:53Will your task benefit from some sort of just careful planning stage? You know, can you take the task? Can you decompose it into a bunch of steps such that, like, you know, successfully completing each step in whatever order would take you from a state of unsolved to solved? Now, if your problem really doesn't have that, maybe you don't need an autonomous multi-agent system the second is for each of these steps um does it make sense to sort of um at this step sort of distinct enough that like they benefit from like multiple expertise or tools or specialized knowledge and the idea is that if they do then you can represent each of them as like agents so kind of like domain driven design where let's say if you're writing some piece of software you need like i don't know uh someone that translates the user requirements into a set of like product requirements and they need something some ui engineer that designs the user interface and they need like some back-end api engineer that like creates the back-end and they need some software engineer that like builds out the front end you need some integration work and then final and you can think decompose each of these things into steps these are independent expertise these guys can do all their work and you can map them to individual agents.

1:16:13Another property here is does the task require consuming extensive context? Again, if we look at the software engineering example, to write code, sometimes you might need to read a bunch of documentation. You might need to figure out like API references and arguments. Now, putting all of that across multiple domains in the same model, the same agent might be challenging because, you know, we all know about like, you know, as context just gets long, like LLMs might lose context. So it just, it does make sense to sort of isolate some of that context and all that extensive context processing within individual agents.

1:16:53And I know that like there's all these arguments around like long context and all of that. But again, does your problem have this parameter? And then the final piece is, does your problem exist in a dynamic environment? So dynamic here means that, let's say, you take a step or an action and the environment changes. And those changes could lead to errors that you need to then recover from. And so, and in that case, you need something that can adapt, something that can sort of explore, like retries, branching, and adaptation logic. And so across this four areas, planning, diverse expertise, processing, a ton of context, and then the task existing in a dynamic environment requiring adaptation.

1:17:38I think if your task just fit like four of these things, then maybe like you're probably, you've landed with a problem that would really benefit from an autonomous multi-agent system. And are there, you know, beyond kind of use cases as kind of fitting into those patterns, are there specific use cases where you found folks getting the most bang for their buck with multi-agent? And therefore, like high level areas that I've seen a lot of people sort of explore like multi-agent systems. And these things might not always be like fully autonomous, but on the spectrum between like, let's say something with some deterministic chain with some complex retry logic.

1:18:26It's a little bit of what I was getting at because like, for example, the canonical example of, you know, multi-agent system is like a researcher. Like I'm going to, something's going to go grab some context on the web. Something's going to write something. Something's going to evaluate and edit that, and it's going to be a loop. But I don't know if that's because that's the best way to build that system or because that's the easiest way to demonstrate that system. And there's a difference. Yeah. So it's on a spectrum. So what I'm saying is like, you know, some complex graph or loop. And then on the far right, it's more autonomous behavior.

1:19:05But software engineering, so things like dev in, magic code, back office tasks, so like process automation kind of tasks, legal and finance, customer service and sales agent. So these are like four high level areas. And yesterday, just before this call, I sort of pulled the numbers. And so I pulled data from Y Combinator. I mean, it's not a perfect representation of everything, but it's a good sample. And I looked for all the companies that mentioned like AI agents explicitly in their task description. So in 2022, there were 17 companies that sort of mentioned AI agents. And in 2024, about 92 companies, a 441 increase.

1:19:47And across all of these companies, you know, most of the value proposition that they had was around like the automation of tasks that were previously reliant on human labor, things that required a lot of repetitive processing, things like data analysis, things that required communication across multiple systems. so the core idea is like if your task no has repetitive processes we can automate it using elements or like agentic systems if your task requires like you know say individuals actual human sort of having to coordinate across multiple systems and we can automate some of that and that's how we sort of provide provide value so um I think it's also instructive to sort of look at that list and see what like you know those companies are sort of doing um but um i think these are like we can also categorize all of them into these four high level areas like software engineering back office tasks in some cases it's like even health or dentistry management systems um in some cases it's just legal like hey no we'll help gather all documents required for your case preparation uh will generate briefs will save your lawyers a lot of money or a lot of time um And in some cases, just customer service, like we triage all the information around the customer.

1:21:08We'll like come up with some easy like resolutions as we're saying, we'll have a human in the loop. And so this is kind of like what I am saying. Yeah. One aspect that, you know, comes up all the time is like, do I need a framework to, you know, build an agentic system? You talked already a lot about what the framework is providing. and, you know, a lot of that sounds complex, but, you know, especially when you're talking about dealing with the, you know, challenges of message passing systems at scale and distributed computing in general. But, you know, talk us through, like, you know, use a framework versus, you know, build it yourself and that whole thinking.

1:21:57Yeah. Yeah, that's a great question. And I like to think back to the early deep learning days, you know, say five, six, seven years ago. Frameworks like TensorFlow and PyTorch were just coming up. And the truth is at the time, a sufficiently skilled machine learning engineer could take a model and represent it using non-Py matrices. And they'd write down the sort of represent their weights using matrices. they could write their own like automatic differentiation like library to uh implement gradient descent they could put all of that into like a loop write the training loop but the problem is that like half the time you make just a single mistake and all the numbers are wrong and it takes weeks or months to sort of debug that stuff and so as a community you know the machine learning sort of community sort of coalesce to organize into like good abstractions so for example like we want some good abstractions for automatic differentiation done.

1:23:00We want some good abstractions for a forward pass, a backward pass done. We want some good abstractions for like an optimizer done. And it turns out that if you can compose all of these abstractions, then you can represent almost any neural network architecture, any type of training loop and that sort of thing. And I feel the same will apply to multi-agents or autonomous agent systems. If you're sufficiently skilled, you probably can write things from scratch. And if your setup is relatively simple, you just have a simple set of chains, you probably don't need a framework. However, if you want to build something that's autonomous and you wanted to think through, you know, what is the control, the right control flow?

1:23:57How do we define when the task is to end or is completed? You want to define, you know, how do we figure out when to delegate to humans? How do we express what pattern do I use? it gets pretty involved. The configuration space for the systems, they sort of interact and they can get sort of combinatorial in some sense. And at that point, it's helpful to have a framework. So the whole idea is like, so for example, like two days ago, openly I released the operator agent that sort of explores tasks by driving web browsers. And with the Autogen API, the high level API, you could implement about the same functionality in about 40 lines of code.

1:24:53And so this sort of being able to take building blocks, sort of put them together, enables like accelerated development. And then a lot of known or stabilized patterns that just get baked into the library. And so these are like good reasons to like sort of use a framework. Yeah. And of course, one caveat is that like if you use a framework, there's some level of indirection. And so frameworks have defaults. They are default system messages. There are some default transformations to the, I don't know, as message flows through the network. And so maybe what hits OpenAI is really different from, let's say, the input that the user provided.

1:25:42And there might be some behaviors that, you know, some assumptions that are made underneath. And so from that sense, you know, as you use frameworks, it's always a great idea to sort of get familiar with exactly what happens underneath. So that like your debugging process and just making sense of what your end system does is just better. And does the, do you see the framework moving towards giving the user more visibility and control over some of those underlying assumptions and transformations that you mentioned? Yeah.

1:26:26So I think the right way to go about this is to have two levels of APIs. And so have a low-level API where like if people are comfortable expressing just anything they'd like to, it's possible. Is your core versus chat in the case of Autogen? Yes, yes. Core versus chat in the case of Autogen. And in situations where you really, really need to be in control of everything the system does, definitely go with the low-level API. and with the higher level API, the abstractions, I think a good framework should have a strong observability story. So first baked into the developer experience. So for example, in Autogen, there's the idea of as the agents sort of interact, they sort of yield these asynchronous messages that tell exactly what each of the agents are doing at any specific time.

1:27:27And you could take those messages, display that in the UI, right? It's some login system. In addition to that, we also emit like open telemetry events down the stack. And you could have like your own open telemetry endpoint, sync, data sync, just store all of that. And it's just great as a way to sort of debug and sort of review exactly what went to the API, exactly what came back from the API just down the stack. So I think observability is one way to sort of like open the box. And the second has to do with just the flexibility to either use a low-level API or a high-level API. Got it, got it.

1:28:07So not necessarily a world in which the developers overriding the system prompt assumptions that the framework is making or those kinds of things? Well, by design, everything is parameterized. So for example, even in the agent chat high-level API, To define an agent, you can supply your system message directly. It's an argument. You can supply the list of tools. Again, an argument. The memory interfaces you want this thing to use. Again, an argument. And there are a bunch of other stuff. So there are good defaults, but they all can be overreading. And again, you can also override or overload those classes, just classic software engineering, and then implement your own core behaviors.

1:28:51So everything is extensible. So I was just, I was speaking in terms of like the developer that really doesn't want to do anything at all. You make zero choices about that. And then you have to live with those assumptions. And the observability gives you some visibility into that without you having to necessarily specify everything. Yep. Correct. So let's jump into your thoughts for 2025. So I think there are a few things, you know, that I think will happen in 2025. And a lot of this is informed by my perspective on work done with Autogen. So in 2025, I expect to see like continued improvement in models.

1:29:35I talked about how like, you know, agents, especially in autonomous mode, might explore like trajectories that might be suboptimal, that sort of thing. I feel like this is a really ripe reinforcement learning trajectory. And so let's say you ask an agent to book a flight. There's a chance that we figure out ways to collect enough data, do some sort of fine tuning that gets the agent to just go through the most efficient trajectory to get the work done every single time in the first try. I also think the same kind of thing could apply to memory and adaptation. And for the most part, it's figuring out what to remember,

1:30:28what type of previous experience to retrieve in order to have a better chance of solving the correct problem. Is this kind of like dynamic optimization of the context? Is that the way to think about this? Yeah. Yes. Yeah. The right pattern, the current right pattern, you know, it sort of makes all these assumptions that like you have the right thing in the database and you can sort of retrieve it and use it just in time. But there's also the other part of like, how do you know when to put things into the memory bank or into the database, to your vector database? Does this always have to be explicitly by the user?

1:31:11Is this something that you can do in some automatic dynamic manner? And I think it should be a combination of both. And making this like a reinforcement learning problem, I feel this will make us like, help us make progress. um another thing i think another thing that i think will happen in 2025 is the consolidation of agentic patterns and i talked about like control for patterns and task management patterns and i'm hoping that like we all as a field or a community sort of align well on on what works and the goal is that like it gives us some shared understanding some shared language as to like you you know, for this class of problems, then here's the right pattern for multi-agent system that works best.

1:32:05I've also been working on the idea of declarative agents or declarative multi-agent systems. So imagine that we could 100 % specify the entire multi-agent system as, let's say, JSON file or something like that. And then if we do that, then we can rapidly get to a point to where the construction of these systems could even be dynamic. So as opposed to a developer having to say, you know, here are like, imagine take one. Here's like an orchestrator for agents and these things are going to, it's going to be what solves the task. Maybe we can even just pop up a level more. And whenever we get a task, we construct this declarative representation of the entire agent system that might work best for this task.

1:32:50And then we sort of instantiate and run this thing and maybe even optimize along the way. And then finally, I think we'll make progress on the UX for agentic interaction. And so a lot of people have made this, drawn this parallel between an agent and let's say junior developer or an intern. And so an intern should be proactive and we need that interface to do the same. So like go get some work done, come back and notify the user. It should be interruptible. and so you know just like an intern if you get them to if you see they're going down the wrong path you should be able to like sort of interrupt them provide feedback and then get them to keep going um we also talked about how like the a big chunk of the digital world needs to adapt to just agents becoming a part of this interaction it might be agents or txt it might be new types of filters and captures that kind of thing and then there's a whole idea of like figuring out how to ensure that human bandwidth is not completely overrun by just agents sort of interacting.

1:33:58A lot of those point to control of agents in the wild and not necessarily control on the part of the people who are publishing the agents, but other actors in the world. Yeah. Yep. Correct. And then there's a final piece that might not be technology or agent focus, but there's also the consideration of how the workforce will change as we have more just agents out there. And will there be, will there, a lot of people talk about the concept of hiring agents instead of software engineers. Would we see these sort of parallels? There's companies advertising that now. Yeah. And what does it mean? Of course, you know, as technology sort of emerges, like a lot of things change, But I feel like, you know, at least on the minimum, we should be having like a lot of conversation around like how like agents will sort of impact like the software engineering field and all of that.

1:35:00So, yeah. So these are like the things I'm thinking of when I think of AI agents in 2025. Do you have a personal take, you know, with regard to software engineers as an example? Yeah. So I actually wrote an article about like how AI might impact the software engineering career. So my summary of that take is mostly around. And it probably will not replace software engineers one-to-one because there's a lot of other things software engineers do, especially senior or above software engineers. There's just a lot of things that these engineers do that are beyond writing code. Yes, communication and context and translating human requirements that are usually severely underspecified into like software technical systems.

1:35:56Those things are iterative, require a lot of effort back and forth with a human, actual human that like, yeah, might not be too well. However, junior engineering rules, like they kind of like, hey, build a web page in React that shows a company's logo or something like that. I think jobs like that are gone forever. or things like write a script that like, I don't know. I mean, we used to have interns that would just write one script that like did one thing. I think jobs like that are gone forever. Another thing I started to see is like, I've been in teams where in meetings where there are some engineers that appear more productive and probably more capable than another set of engineers.

1:36:53But essentially what's happening is that these engineers, the first set of engineers, they're just using AI. They're leaning and using AI really heavily. They have their IDE set up. They have their workflow set up. They figure out how to take problems, write design documents, give that to AI, get like nice modularized implementations. They've learned to write tests. they've learned to be really vigilant about the kind of mistakes that like these models will still make. And they've learned to integrate the entire thing into PRs that are error-free or bug-free that the rest of the team sees. Now, this is a skill, you know, just with how internet literacy or digital literacy was a skill that everybody needed to cultivate, how to use Google search, which helps navigate the web.

1:37:43I feel like software engineers need to sort of invest in that skill that lets them effectively integrate AI into your workflow. And the difference is really stark. So these two groups of engineers, I know for a fact that, like, they're pretty capable about the same individual capability, but just one has invested in just going through that integration process and ensuring they can come out with correct high-quality code while the others just haven't done that yet. So that's sort of like my high-level take. There won't be like one-to-one replacement, but some jobs below some level are probably gone.

1:38:22And then second, like investing in just AI software engineering literacy really creates like significant differences in productivity across engineers. So what do you think, Sam? Yeah, you know, I find myself frequently struggling with, you know, the classic difficulty seeing exponential change, right? It's like I use AI coding agents, you know, not professionally because I'm not building any software system, you know, of any significant scale. But I, you know, I've got, you know, cursor set up and I use it, you know, fairly heavily. and you know there are times when I think it's like magic and incredibly productive and I can get so much further so much faster and then there are times when like I get into the loop of like banging my head against this thing and it's not making progress and it's like you know change you know reorganizing the the chairs on the deck of the Titanic like you're asking me to like change things around that have no consequence.

1:39:35And, you know, I think, you know, in some ways, I think, you know, it's probably like, you know, self-driving cars and that, you know, it's going to take a lot longer than people think because like it's easy to see the 80 % progress, but the 20 % progress is going to take years and years and years and years and years. So, but I think what you mentioned, I think there's something to what you mentioned in that, you know, there's a skill to using these systems and it, there's not just a skill, but I think there's a level of investment in building, you know, structures around these systems that, that I may be, I may not see in my personal use, but if I'm doing this at the scale of an organization that has thousands of software engineers, I'm able to invest in the degree of customization that, you know, takes some of the frustration I see out of it.

1:40:39I don't, I need, you know, if you know anyone I can talk to, to, to dig into, you know, how folks are using this stuff at scale, or if anyone, you know, has a recommendation for me. That's an interview I'd love to do. Cause I, you know, the, it's weird because you, on social media, it's like, yeah, it's all bullshit. The stuff doesn't work. Like, you know, and then, you know, it's like, you know, there are doomsday, you know, doomsayers and cheerleaders. And I think, you know, the reality is in the middle somewhere. And, you know, certainly what you're saying about skill is an important piece of that.

1:41:19Yeah.

1:41:23Half of the people that like complain that it doesn't work. So I think half of that is a skill issue. They just haven't come up with like a structure for using these things. If you're pretty efficient, it's like, hey, there's a class of problems where don't even bother. Like don't even bother. and you build intuitions as to like if i get if i ask oh one to do this stuff it's just going to get confused and reorganize the deck and like or if i ask it to do this without mentioning this really important context you'll completely make a mistake so there's all the stuff that like i feel some type of tacit knowledge you know there's a there's tacit and explicit knowledge So tacit knowledge is like the kind of thing that like is very, you know it, but it's really hard to express in words.

1:42:15It's like if someone asks you, how do I swim? You can't really describe it to them and then they jump in the water and swim. And I feel similarly like just building the right intuition as to when and when not to use it just comes from practice. So, yeah. I think there's when to use it, when not to use it. And I've also found that there's an intuition around when to go to fundamentals, meaning, you know, so I get a lot of value out of Cogen when I'm working with, you know, libraries or APIs that I've not used before. And, you know, I just want to do something quick and dirty. I don't necessarily want to go to the docs and read them top to bottom.

1:43:04You know, but at some point, you know, you get a sense and my ability to sense this has refined over time. And so now I'm much quicker to recognize that I'm at this point where I just need to go read the docs and understand what's happening around this thing that I'm trying to do. Because the models, you know, I'm not maybe asking the question the right way or the model hasn't seen enough in the training data about this. and I need to help it get over the hub. So, but that is also a, you know, a skill or a feeling that, you know, is really useful in this. Yeah, yeah, absolutely. Awesome. Well, Victor, it has been wonderful both catching up with you personally.

1:43:52It's been a long time and chatting with you about all this stuff. You know, great conversation. I really appreciate the time you've taken to go through this with us. Yeah, absolutely. Thanks for having me, Sam. Pleasure. Thanks so much.

From the publisher

Today we’re joined by Victor Dibia, principal research software engineer at Microsoft Research, to explore the key trends and advancements in AI agents and multi-agent systems shaping 2025 and beyond. In this episode, we discuss the unique abilities that set AI agents apart from traditional software systems–reasoning, acting, communicating, and adapting. We also examine the rise of agentic foundation models, the emergence of interface agents like Claude with Computer Use and OpenAI Operator, the shift from simple task chains to complex workflows, and the growing range of enterprise use cases. Victor shares insights into emerging design patterns for autonomous multi-agent systems, including graph and message-driven architectures, the advantages of the “actor model” pattern as implemented in Microsoft’s AutoGen, and guidance on how users should approach the ”build vs. buy” decision when working with AI agent frameworks. We also address the challenges of evaluating end-to-end agent performance, the complexities of benchmarking agentic systems, and the implications of our reliance on LLMs as judges. Finally, we look ahead to the future of AI agents in 2025 and beyond, discuss emerging HCI challenges, their potential for impact on the workforce, and how they are poised to reshape fields like software engineering.

The complete show notes for this episode can be found at https://twimlai.com/go/718.

More from The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

All 156 episodes
AI Trends 2025: AI Agents and Multi-Agent Systems with Victor Dibia - #718The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) · 1 h 45 min
Listen in VO