#214 Ece Kamar: Why AI Agents Are the Next Big Thing in Tech (Microsoft Research)

17 Oct 2024 · 1 h 1 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Episode Summary: Eye On A.I. - Ece Kamar: Why AI Agents Are the Next Big Thing in Tech

Overview In episode #214 of the Eye On A.I. podcast, host Craig S. Smith interviews Ece Kamar, VP of Research and Managing Director of AI Frontiers Lab at Microsoft. The discussion revolves around the transformative potential of AI agents, particularly in enterprise settings, and the ethical considerations surrounding their development.

Key Themes

  • Definition and Importance of AI Agents
  • The Shift to Agentic Workflows
  • Multi-Agent Systems and AutoGen
  • Ethical Challenges and Responsible AI Development

Episode Breakdown

00:00 - Introduction

  • Podcast sponsored by RapidSOS; discussion on AI's role in safety.

03:11 - What Are AI Agents?

  • Definition: AI agents are autonomous entities that perceive and interact with their environment to perform tasks efficiently.
  • Significance: Moving from traditional AI models to agents represents a paradigm shift in how technology interacts with users.

04:43 - Building Responsible AI at Microsoft

  • Emphasis on the importance of developing AI responsibly amidst rapid technological advancement.
  • Challenges include managing ethical concerns and ensuring safe practices in AI development.

10:55 - The Rise of Agentic Workflows

  • Explanation on how agentic workflows automate tasks by breaking them into subtasks assigned to specialized agents.
  • Importance of these workflows in enhancing productivity in various industries.

12:30 - Multi-Agent Systems and AutoGen

  • Introduction of AutoGen, a multi-agent orchestration library developed by Microsoft.
  • AutoGen allows for the creation of teams of agents that can collaborate on tasks, enhancing scalability in complex environments.

18:04 - Scaling Multi-Agent Systems

  • Discussion on the capability of multi-agent systems to scale up to millions of agents.
  • Potential applications in enterprises, including simulating customer interactions for enhanced service delivery.

20:22 - The Creation and Evolution of AutoGen

  • Insights into how AutoGen was developed and its ongoing evolution based on community feedback and use cases.

23:07 - Real-World Applications of AutoGen

  • Examples include automating tasks such as expense reporting and customer interaction simulations.
  • Highlighting the diverse applications emerging from the use of AutoGen.

25:52 - Large-Scale Simulations with AI Agents

  • Potential for AI agents to conduct large-scale simulations to study societal dynamics and scientific discoveries.

27:36 - The Role of AI Agents in Scientific Discovery

  • AI agents may revolutionize scientific research through improved data analysis and hypothesis testing.
  • The combination of reasoning capabilities and agent society paradigms can drive innovation in scientific methods.

31:20 - AI Agents and Complex Reasoning

  • Discussion on how agents can engage in complex reasoning and decision-making processes.

36:49 - Challenges in Defining Agent Boundaries

  • Concerns regarding the operational boundaries of agents and the risks of unauthorized actions.

39:12 - The Risk of Agents Interacting with Each Other

  • Potential dangers of agents working collaboratively, including unforeseen consequences.

43:59 - Building Trustworthy and Safe AI Agents

  • Importance of establishing safety protocols and accountability measures in AI agent development.

48:44 - Learning from Human Factors in Automation

  • Emphasis on understanding human factors and incorporating them into AI design to mitigate biases and errors.

50:50 - Why Speed and Coordination Matter in AI Development

  • Urgency in developing safety measures alongside rapid advancements in AI technology.

55:08 - The Future of AI Agents in Enterprises

  • Predictions for the increased adoption of AI agents in various enterprise applications and sectors.

57:47 - Low-Code/No-Code Development for AI Agents

  • Potential for low-code/no-code platforms to democratize access to AI agent development, allowing more users to create customized workflows without extensive programming knowledge.

Key Takeaways

  • Transformative Potential: AI agents have the potential to revolutionize industries by automating complex workflows and personalizing user experiences.
  • Ethical Considerations: As technology evolves, ensuring responsible AI development is crucial to mitigate risks associated with autonomous agents.
  • Future of AI in Enterprises: Enterprises are likely to be the first adopters of agentic workflows, leveraging automation to enhance efficiency and reduce operational costs.

Conclusion The conversation underscores the urgent need for interdisciplinary collaboration to ensure that the development of AI agents is safe and ethical while also maximizing their potential to improve various aspects of human life and society. The integration of AI agents into everyday processes represents a significant advancement in technology but requires careful navigation to address the accompanying ethical challenges.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00So when we think about agents, we are thinking about AI based entities that can perceive the world and act in it in a continuous fashion to be able to carry out tasks. Why agents is such an important paradigm shift compared to the models. First, this is really about the value. Why we are building the AI systems at the first place. In my opinion, the biggest question to be addressed right now in our community is what are we doing with these models? What are our goals? Showing demos and just getting some input outputs with text is not the value that we are looking for from the systems that we are building.

0:36We want systems that are aware of their environments and that can do things for us so that as users, we can get the right value. Okay, AJ. Yeah, wonderful to meet you. I'm particularly interested when this came across, you know, everyone's talking about agents. And so I wanted to hear, and I do some ghost writing for Boston Consulting Group, and I just wrote a piece on agentic workflows, and I didn't mention Autogen, but it seems like you're involved in that space, and I thought maybe we could talk about that. And then I saw that you also were involved in evaluating GPT-4 before it was released.

1:29Yes. And I have some questions about that. I have a pet peeve about the TaskRabbit anecdote. I don't know if you know much about it. Which one is that? That according to OpenAI, GPT-4 was asked to do something. and it had to solve a CAPTCHA and it went on TaskRabbit and hired a freelancer. Yes, yes, and I know about that, yes. And that anecdote gets repeated everywhere. I just saw Yuval Hariri is out promoting his new book and he's been telling that anecdote. And, I mean, you and I know that GPT-4 did not, you know, on its own log on to TaskRabbit, TaskRabbit, log on to TaskRabbit, you know, set up a job, talk to a, you know, it's, I hate stuff like that.

2:34I mean, just be honest about what happened. It's, it's interesting. And someday that may be able to work, but that's not what happened. You know, that story is actually relevant to the discussion on agents as well, Craig, because when we think about agents, we are thinking we think about ai based entities that can use tools and that can take actions in the world that's right and that doesn't mean consciousness that doesn't mean like what they are thinking about like this conscious entity discovering things completely brand new in the world at all yeah but we are and i know we are going to get to that conversation but we are increasingly building ai systems that can use tools and a tool can be also going to task rabbit and putting a job posting in there so like i think we need to call it what it is but also be aware of like what that really means not for consciousness not for agi but like really preventing some issues that may happen there soon yeah um yeah no i i i agree it's just um you know people take an anecdote like that and they make it sound it's it's a little bit like uh my other pet peeve is that uh talking head um that ben gertzel runs runs around with uh sophia which you know i am going to destroy all humanity i mean it's if there's someone in the back like typing in this stuff it's not you know here is my real issue with all of this i spent a decade helping to build responsible ai inside microsoft yeah i work with many different product groups building different ai technologies where they were serious risks and there still are and that work is time consuming and takes a lot of effort from people who are involved and they have to a lot of times conquer new territory to be able to do it right.

4:48And that work is already under-resourced compared to all of the AI developments that are happening. If we are not creating, if we are not focusing our attention to the real problems that are happening in the world or what is on the horizon and kind of fooling ourselves with things that that may not show the real state of AI, we are actually not utilizing our thinking and resources properly to be able to address what is most important for society. And this really comes up, you know, when some of the AI leaders write letters and say like, let's pause everything for six months or, you know, kind of make these calls.

5:26That is like, the good thing about these conversations is the impact of AI on society happening now or happening in the future is going to be so important. So we need everybody to think about it. But we need people to think about the right problems, to realize where this technology is, and not distract the conversation to hypotheticals. So that's kind of my pet peeve. And the other thing I'll say about that is, pause is not the solution. You have to invest in the science, in the research, in the development of responsible AI and safety. And that's not going to happen by pausing. It's going to happen by prioritizing those topics as much as you are prioritizing the frontier models.

6:13Yeah, I agree. Yeah, yeah. Okay, so why don't you introduce yourself? I may use some of that stuff. Sure, sure, happy to. But why don't you introduce yourself and then talk about how you got to Microsoft and then we'll talk about agents and agentic workflows sounds good hi Craig thanks for having me at the podcast my name is Ecek Amar I'm a VP of research at Microsoft and I also I'm the managing director of an AI lab called AI Frontiers we started AI Frontiers a year ago with the mission of understanding where AI is headed and pushing the frontiers of AI technologies by always going after the question of what is the most important thing we should be studying at the given time.

7:10So I have a team distributed between Redmond in Washington and New York. And yeah, we have been doing a number of interesting projects, especially aligning our work on agent so i'm excited to talk to you today about those topics yeah what were you doing before you got to microsoft research before i got to microsoft research i was actually a phd student at harvard this is between 2005 and 2010 and i did my phd at a time where these topics especially general ai was not a popular topic at all as you know the ai field has gone through many ups and downs to its history. And during the time I was doing my PhD, my PhD topic was actually human AI collaboration, advised by the wonderful Barbara Gross, who is actually now recognized as being one of the godmothers of AI, because there are actually quite a few women that has been trailblazers in the AI field.

8:15And when I was studying the topic of human AI collaboration, And in fact, my PhD thesis was also on multi-agent systems, multi-agent AI systems. Those topics were not top of mind for industry, and the academia was actually quite scattered in terms of the study of these problems. So when I was doing my PhD and working on these problems, some of the work that was coming from the real world, building systems and trying them out in the world, was coming from Microsoft Research. In fact, Eric Horvitz's team. So as a PhD student, I was looking for opportunities to kind of get my hands dirty, apply these ideas to real world problems.

9:04So I did some internships at Microsoft Research. I became a research fellow. And, you know, it's kind of interesting how you get to, at the place like Microsoft Research, where there's this history of research and there's the freedom to think. You can actually get some ideas before their time. So one of my internships project was actually ride sharing. Before there was an Uber, we were actually imagining like how to bring all of the drivers and the passengers of Microsoft shuttles and cars commuting to the campus every day and how we can gain a lot of benefits from that kind of a marketplace. So that was kind of one of my internship projects.

9:48And then I started at Microsoft Research as a researcher right after grad school. And I got to work on a number of interesting projects over the years and then found really a passion of mine which is making sure that ai is developed responsibly and that became kind of the passion of my first decade at microsoft just just pushing the company around that really important work um and then now my second passion is really bringing the talent in MSR together in this pretty crazy adventure we are having in AI. Yeah. And agentics or agents, AI agents, which are suddenly sort of entering the economy, they're no longer research projects.

10:47I mean, there's certainly there been AI agents before, but these generative AI driven AI agents and then networks of agents and workflows of agents is a pretty exciting topic. And there's a lot of sort of hyperventilation about it, as there is with a lot of these developments. But I can see that indeed, as you are able to get agents working together on whole workflows and then talking to agents in other institutions or domains, that you will build kind of a fabric of AI agents that support human activity. So can you talk a little bit first about the development of agents and the working with multi-agent systems?

12:00How big could those systems be, for example? Definitely. In my lab, the mission we are after is really pushing the frontiers of what AI can do, what AI can do for people, what it can do for society, what it can do for organizations. And all of the work that has been happening on frontier models like GPT-4's clouds of the world, they have been, of course, wonderful to give us a foundation to build the next generation of AI systems with. And all of us have been in this journey, Craig, you and I and many of our listeners that when you actually get your hands into a frontier model, you start doing things with it by prompting it, putting things into its context, and then giving it tasks to do.

12:48And we have been also amazed with, like we have been some of the early adapters of the GPT-4 models before there was a JetGPT in the world and spent a lot of time understanding and prompting it. And soon enough, if you actually play with these models enough, you realize that the way you want to build systems with these models is not only prompting these models. And this is, in fact, a big piece of a new computing paradigm to be built for our world, where foundation models become important processors, but they are not the only thing. In fact, when we are building AI systems, we want these systems to have memory.

13:32We want these systems to be able to remember things. We want these systems to be able to personalize to us, remember things about us. But not only that, we actually want these systems to be able to interact with the world, interact with us, and do things in the world for us. This is where we actually now get the agents. So when we think about agents, we are thinking about AI-based entities that can perceive the world and act in it in a continuous fashion to be able to carry out tasks. Why agents is such an important paradigm shift compared to the models. First, this is really about the value, why we are building the AI systems at the first place.

14:13In my opinion, the biggest question to be addressed right now in our community is, what are we doing with these models? What are our goals? Showing demos and just getting some input outputs with text is not the value that we are looking for from the systems that we are building. We want systems that are aware of their environments and that can do things for us so that as users we can get the right value out of them. And agents are how we are going to be pushing for that value. But second, agents are also how we utilize these really beautiful models for getting a lot of things done. But then there's also a question about how we engineer increasingly more sophisticated and complex AI systems.

15:00And on that one, the reason we actually started, my lab started doing the work on multi-agent systems. Last year we released AutoGen, which is one of the popular libraries in the multi-agent space. is because the multi-agent framework provides a really nice programming paradigm for increasingly more sophisticated AI systems. What I mean by that is imagine you have a difficult task to do. Now, if I'm only using models, I'm not using agents. The way I would program that system would be trying to put all of the instructions, all the context, all of the memory into the context window and say, do it for me.

15:41And you have no idea how the model may or may not be able to do that. But what you can really do with multi-agent orchestration instead is you can take the task and decompose it into pieces. And then define this task should be done by these agents where each agent specializes in this kind of capability or this kind of knowledge or this kind of expertise. And then once you have that team of agents where there is a particular specialization in them, then you can orchestrate them. And that orchestration can either happen by a human creating a workflow, like how you would create a workflow for human teams.

16:24Or it can happen through having an orchestrator agent that actually uses LLM-based orchestration and planning to be able to guide the work. or based on our now new work, it can also happen by new approaches that can use the best of those two. But this is so exciting for us because this is where we think AI is going to be really have a big impact on people and the society because it can do things for you. Yeah. And, you know, my question was how large can a network of agents be? I mean, because what we're talking about initially are breaking a task into subtasks and assigning an agent to each subtask.

17:20But is it possible for that workflow to spread and handle a much broader series of tasks? Yeah, that's a great question. This is really about how our understanding of the agents are evolving in time. What I describe you in terms of using multi-agents as a programming paradigm is how we came to the idea of building Autogen in the first place. But I completely agree with you that that's not where we are stopping. And in fact, where we see with many of the customers using Gotogen and our internal teams as well is they are not using multi-agents only as a programming paradigm anymore. They are using multi-agents to represent, we call it a society of agents.

18:14You may call it the network of agents, or we are calling it a society of agents, to actually start thinking about building persistent and composable teams that now become part of the way we work, that now become part of the way an organization is being built, now become a part of a way that consumers may be able to interact with technology. So, you know, coming back to your question about like how big these teams might be, I think it may become as many as an organization's have people or more. And we actually don't see a bounce in that numbers. And we are seeing some really interesting auto-gen use cases right now.

18:57This is coming from some of the sectors that are really interacting with consumers around the world, is that they are actually thinking about building an agent per customer. And for big organizations, this may actually mean millions of agents. Wow. Because now they are thinking about scenarios where agents may actually simulate their consumers. And with this, they may be able to test the systems they are building, or they may be able to simulate some of the consumer scenarios up to a point where the user may take off and then continue the rest of it. And this diversity, the scale we can get to an agent, I think is getting us to really think about what we are building from its core.

19:42Just to give you an example, when we built Autogen, Autogen was built to be like a centralized multi-agent system where there would be like a top agent that tells every other agent what to do through doing a central broadcasting. For the last year, we have been interacting and working with many users of AutoGen around the world, both professional organizations as well as anybody who would like to build something with multi-agents. And after seeing these feedback of how they are using Autogen, that they actually want to have millions of agents interacting, they want to have agents that have persistency, that do not go away after doing a task, but they become part of the organization and capturing like what that organization needs to know about.

20:33We have actually completely redesigned Autogen. Just last week, we have released a new version of Autogen that we are calling Ag Next, where Autogen is completely designed from scratch again, so that we can build it in a way that it can scale to millions of agents, that agents can be persistent and can be discovered in the future towards this vision that we are imagining as a society of agents. Yeah, and I know what Autogen is, but listeners may not. So can you stop and introduce Autogen to listeners? Autogen is a multi-agent orchestration library we've developed. I think we released it exactly a year ago in the October of 2023.

21:20And this was mainly a library to get us and the world, both academic world and the industry, to start experimenting with multi-agents because in our research lab, again, we were doing a bunch of work with models, but then to push the frontier of what you can do with them, we started experimenting with multi-agent ideas. And then we said, okay, we think this is really cool. Can we actually push it to the world and see what people do with it? So we released Autogen as the multi-agent library. And what you can do with Autogen is, Again, it has some functions in it to be able to define tasks, to be able to decompose tasks, to be able to create agents and get agents to collaborate with each other to be able to do tasks.

22:09And people have done wonderful things with Otogen. It has been one of the most exciting adventures that we've been on, just watching the community, the open source community doing things with it. So this feedback between us and the open source community has not only taught us a lot about where Autogen needs to go to kind of become a better library for the diverse set of scenarios people are using it for, but it's also giving us wonderful ideas in terms of what new research problems we push in agentic AI. Yeah, Autogen itself is an open source initiative. Yes, it is. It's an open source project by Microsoft.

22:48but it is it is open to of course many collaborators we have that are contributing to the code um with microsoft yeah uh you mentioned watching some of the exciting things that people are doing i don't know what you're authorized uh to talk about but can you give us some examples that of multi-agent systems that are being built with autogen yeah we can talk about the things that people are freely sharing on discord autogen also has a pretty lively community that loves to share with each other to be able to learn more we are seeing people using it in really large-scale applications to be better understand their consumers and to be able to reflect their consumers inside the applications they are building.

23:43We are seeing some automation work going on, like taking tasks that are very, very time consuming for people that people love, hate doing, and then getting agents to do that. For example, one scenario my team is quite excited about is automating the expense reporting because everybody needs this. There are certain tasks that when you actually say like, I'm going to have an agent or an agent team that's going to do this for you, everybody's eyes light up. One of them is always expense reporting. So those are some of the cases that people have been successfully trying. But the one that I think I'm surprised about, and I would love to see more work happening is large-scale simulation and all the discoveries we might be able to do about the society and science by just doing large-scale simulation of systems and the dynamics and seeing what can happen with that.

24:44So I guess the most interesting part of the AutoGen adventure has been like just seeing the diversity of cases and also seeing how some of our users are taking it so much further than we can as a research organization and kind of pushing the frontiers much faster than we can and then we're like okay now we have to learn from you in terms of how we are using it teach us not us teaching them but them teaching us about how to use autogen yeah I'm curious about large scale simulations can you can you uh give an example of what that would mean I mean what simulations of what kinds of systems. I think a wonderful example of this came from Michael Bernstein's team at Stanford last year.

25:32I cannot from the top of my mind remember the title of the paper, but they were actually, they built some kind of a sim simulation using agents. You remember the game Sims where people would interact with, not people, computer agents would interact with each other, have conversations, or just live their lives through interaction. So they actually built a version of that using multi-agent systems last year. And that paper is really interesting. It has interesting insights. And that was done as a research study then, but now you can actually, if you are building agents to represent different demographics, different political views, or other kinds of considerations and create a simulation of how these different agents may interact with each other you may you might be able to scale the dynamics of large organizations similarly now with our sister organization at microsoft research that looks into science and incorporation of ai models into scientific discovery they are now looking into like what agents mean for scientific discovery so there's a whole different horizons of applications that we are working on or we are seeing others to do that is that is really interesting yeah i i think i think one of the really interesting areas to watch in the next few years is going to be like how the way we do science is going to change with ai yeah and how ai is just going to change the nature of scientific discovery i think that's a space that has so much promise but also a bunch of risks that we have to be careful about yeah no absolutely and on uh on simulations and on uh scientific discovery i mean you you now have um gpt-01 uh but presumably everyone's going to follow uh doing a sort of chain of thought reasoning and uh and and uh and that sort of thing so so the the output is much uh more human-like in in its uh uh in its uh in in that it's evidence-based and and that sort of thing that seems to me that that would apply to uh scientific discovery or research and then what you're saying is you could have these reasoning agents working together uh to to what critique each other or or to each solve a discrete problem within a larger problem Hi, this episode is sponsored by RapidSOS.

28:34Terabytes and petabytes data is exploding in our daily lives. How many connected devices do you own? You ever wonder how all this data could be used in emergencies? By 2030, we'll have over 32 billion IoT devices worldwide, double what we have now. Despite this data abundance, there's a critical safety gap. Vital information isn't reaching emergency responders in time. This gap results in reactive rather than proactive emergency response, putting lives and property at unnecessary risk. Many of us have had emergencies in our homes where we depend on EMS to respond quickly. Rapid SOS is closing this safety gap with their AI-powered intelligent safety platform.

29:31They connect life-saving data from over 540 million devices to more than 21 ,000 public safety agencies across six countries. Despite$200 billion in annual safety system investments, enterprises often lack timely, actionable information during emergencies. Rapid SOS is changing that. It's not just about avoiding problems. Safety investments are linked to increased customer satisfaction, employee retention, and long-term firm value. Close the safety gap and transform your emergency response with RapidSOS. Visit RapidSOS.com slash IonAI. That's IonAI, E-Y-E-O-N-A-I, all run together. Visit RapidSOS.com slash IonAI today to learn how AI-powered safety can protect your people and boost your bottom line.

30:42That's rapidSOS, R-A-P-I-D-S-O-S, dot com slash IonAI. Visit them today. So what we are really imagining to happen in sciences is a lot of the scientific process is complex reasoning, right? Like looking into different sets of evidences and being able to detect patterns in that and being able to apply those patterns to new problems and asking questions about, you know, why and how and creating new hypotheses that may turn into experiments. So as you've pointed out, the O1 models and particularly any reflection and chain of thought patterns or a graph of patterns like people have been developing, these are all getting models to take their time with thinking.

31:40instead of just doing simple pattern recognition saying, take your time, run things through, reflect on what you're seeing, provide the feedback. And in fact, a lot of the earlier autogen use cases we have developed and we have seen from the community was applying these patterns to models, doing chain of thought, doing self-reflection with agentic technologies. And those getting models to do better reasoning, more complex reasoning and doing self-reflection are of course very important, but that's not enough. What we are imagining for like a scientific agent or a consumer agent is systems that are able to really perceive the world in deeper ways.

32:32and then use, of course, those reasoning capabilities to be able to guide people to kind of think about what they do next and that way discover new information and put that information back into the system. So with respect to now thinking about the scientific applications, again, reasoning is an important part of it, but we believe these notions of society of agents or creating specialized agents are going to be crucial for helping scientists with their tasks. For example, we will have agents that really become experts on a particular literature. Right? So scientific discovery is not only about running experiments, but it's also using past knowledge and being able to reflect on the past knowledge.

33:19We will have agents that will be able to brainstorm with people about what are the most important experiments to run and such. So we believe the combination of the society of agents paradigm plus the reasoning, plus these agents now having access to new tools to be able to, for example, run MATLAB code and actually simulate things. Those are gonna be, I'm really excited about how those pieces are gonna come together now to kind of create the full package rather than like some vertical demos on things. Yeah. You mentioned safety. And at the beginning, you were talking about responsible AI. Is there a danger that these agents interacting with each other could use data or share data that shouldn't be shared and, or, or, uh, uh, you know, there's, there's all kinds of scenarios you can imagine where, uh, once the agents are activated and interacting, uh, outside of the, the view of, of humans, uh, that they could either get, you know, stuck in a blind alley, or they could do things that are unauthorized but had been unforeseen by their creators?

35:01Yeah, so I believe we don't yet understand the safety considerations around having agents in the world yet as a community. And I believe this is where we have to have some coordinated efforts to really gain a level of understanding comparable to the level we have right now with LLMs, for example. It's very early days, but we are already seeing examples that this is going to be an important space of concerns with respect to agents. Here, I want to highlight three particular issues I'm seeing right now. So first, let's take an agent that is not working with any other agents, that's not connected to a network of agents, but it's just an agent that can interact with the world.

35:53And this agent will have access to a bunch of tools to be able to do things in the world. So when we think about tools, we can think about an agent that can do things on a website, that can do things on a mail server, that can do things by, I think you brought this up earlier, like maybe able to call some people in for help that can be a tool running a calculator or writing code might be an access to a code and the particular thing about agents why we are excited about agents is because you can give a test to agents and agents will discover ways of getting that done but this is where the dangers are also coming from unless you can define operating boundaries around agents, they are going to do whatever they know how to do to be able to make that task happen.

36:51So what I mean by that is, let's say the agent I have, I actually tell the agent that it should provide some information for me. If the agent needs to get into my emails to provide that information for me, it's going to try to go to my email. If it doesn't know my password, let's say, it will think about ways of resetting my password. Not because it has an intention of resetting that password, but it's going to get creative in terms of finding ways of making that task happen because I ask it to do that task for me. And if I haven't defined the operation boundary around that agent, that it should not be resetting my password without asking it to me, it's just going to go and do it.

37:48So there is this, the first concern I have is how as a community, we can define those operating boundaries for these agents. and with now agents being able to interact with the web interact with online resources really freely enumerating all of the right and wrong for these agents is not a trivial task at all so that's that's my concern number one my concern number two is now the agents interacting with people in many of the applications we develop we kind of throw away the task of accountability to people We say, okay, I'm going to, you know, my agent can do whatever, but we are going to show things to people.

Read the full transcript

38:34People can validate, people can verify, people can stop the agent. But we all know that as people, we don't always have the attention to check everything, right? So this creates a false feeling of control. Just saying the human will figure it out, the human will check that, the human will do that, right? So that's not something we can 100 % rely on. That's just a misconception. So that's my third warning, the second warning. The third warning is what you're saying, Craig, which is now I actually have these agents connected to each other. What are the dangers of that? What I can tell you is I don't think we can even understand what that means.

39:20No. can these agents start coordinating and when they do something wrong it can now scale to much bigger risks possibly yeah and this is why we need to work on this and we need to work on this pretty fast yeah i mean on the and without going to the science fiction scenarios which of course you could easily imagine but just for example on uh on the password issue uh you you have a you know a system of agents that are are doing tasks and one agent needs to access your email so maybe you have it set up so that it asks for the password you give it the password but you You don't know what the agent, now that it has the password, what it's going to do with that.

40:22You know, maybe another agent will ask for that password and that will be unseen to you. And so that's the, the, the, when I was saying about data, once, once the agents are accessing data, uh, you, you don't know where that data is going to go. I mean, frankly, we already don't know where it's going, but it'll be used by, you know, in a much more active way, or it may be used in a much more active way. Is there talk of sort of visualizations so that, and maybe there'll be like an agent master who is in charge of kind of watching what the agents are doing and what kind of data is being shared and, you know, or maybe that's another agent that does that.

41:21Yeah, I think we have done quite a bit of work building multi-agent solutions where the roles of some of the agents are around compliance and watching what is really happening. Just to give you an example, we've done this for the issue of hallucinations, for example, where you can imagine some of the information created by agents in a workflow may not be always correct or truthful, but I can have some agents kind of watching for this and giving corrections and saying like, can you show me the citation for what you're seeing here? And then looking at that citation and validating and coming back and saying, no, you know, you cannot use this because this doesn't really exist in that document.

42:11Can you take another pass? I think we can imagine having some privacy proxy agent for the user that basically watches for what is the information that's being shared and being kind of a gatekeeper for the user's information. And this is something we are thinking quite a bit, again, in the notion of society of agents. If I'm going to have my agents, those agents will know things about me and will have information about me. How can we define the boundaries of what they can share with other agents? And for the things that they cannot share, how can they partner with other agents to still give me the value without giving the information away.

42:59So I believe the multi-agent scenarios and the bringing agents that actually represent me, whose job is representing my concerns and my values and my privacy considerations are going to be an important part of how we engineer these systems. But again, I think where things are actually falling short, Craig, right now, and this is true for everything we are doing in the AI space, is that we can engineer these systems, but we don't have guarantees about them yet. Yeah. These systems are still stochastic, and maybe there's a 99 % case that my privacy agent is going to do the right thing, but what about that 1 %?

43:47So I think this is where a lot of the interdisciplinary research needs to come in now. For example, at Microsoft Research, one of the things we have going for us is we have people working on systems and verification and the AI experts. And I think the part we are trying to push right now is like, are there ways the techniques of the past decades can be translated into the world of LLMs and multimodal models? And maybe there are some verification techniques from the past that can now come us and give us some guarantees right because i think we all want guarantees we don't want to work on probabilities when it comes to our privacy and safety yeah and and uh to that point one of the things i've been thinking about recently is a couple of things one is is automation bias that you mentioned about uh whether or not people have the attention or to to track what agents are doing part of the problem is people tend to trust machines if it's done the right thing 10 times and you kind of let it run yes and and then as uh as The other thing that I'm interested in, this is with regard to coding, as auto-generated code grows and code bases of auto-generated code get larger and larger, there will come a point, maybe we're already there, where no human really understands the code base because it was auto-generated.

45:38they can look into it and and uh you know read it in in discrete uh chunks but no one and then you become beholden to ai or to ai uh agents to to understand the code base for you in the same way with these agentic workflows if you know if if it you know you you play it out if you end up with the bureaucracy of human society, all this stuff that needs to get done, being taken care of by this fabric of agents working together, at a certain point, we're going to be beholden to them because no one really understands everything that's going on. But maybe, I was talking to my wife, Maybe it's like fire and electricity.

46:39I mean, you know, everyone was, well, we're going to become over-dependent on electricity, and we are. And maybe that's just the way it'll go. We will become dependent on AI. And when the AI goes down, we're going to all be hands in the air. I mean, what do you think about all of that? You know, the part I don't accept is pretending that we have no control over what we are building. We are controlling what we are building. We should be putting in the time and the effort into understanding what we are building. And we should not be pretending that, oh, these things are going to be so complicated.

47:22One day we are going to wake up and we are going to realize whether it's electricity or fire or something else. Right. You know, this is why I've spent so much time working on the field of responsible AI, because there's actually a lot of things we can do, from the design to the development to the deployment of these systems in terms of asking the questions of like, what are we doing and what can go wrong? and then bringing in all kinds of mitigations and controls into the picture to make sure that we are happy with the outcome we are creating. That doesn't mean, look, that doesn't mean that I can predict everything that's going to be happening in the space of agents in the next few years.

48:02But we should be taking our time and asking these questions and we cannot just pretend that, oh, it's technology, one day something is going to happen and, you know, whatever happens, happens. So with respect to this issue of automation bias that you are bringing about, there's actually quite a bit of literature from human factors in terms of what happened when autopilots came into the picture for the aviation industry. And they have developed practices in terms of how to make sure that there are the right checklists or the attention on the transparency layers. And not only that, but people are asked to do these things once in a while so they don't forget their practices.

48:42Again, I don't know what they exactly look like, is automation is getting into everywhere in the way we work and whether those practices are going to generalize. But as a field, we have to take the responsibility of asking the questions and then not only assume that everything that came before is throw away now. Go back to the literature, learn from what worked and what didn't work, do the job of translation and for the open problems, actually do the work of doing, solving those problems. Yeah. And I believe there is quite a bit we will be inventing in the next few years to get this right. And part of it, I think, is going to be really looking into the question of human factors and asking the questions there.

49:30Part of it is going to be building the right transparency layers in a way that we can be careful about how much attention we are asking, but really going to the humans when important. And the third one is, I think we are also going to get more clever in the way we are architecting the AI systems. If the stochasticity is something really bad for human factors, can we engineer our systems in a way that we can reduce stochasticity as much as possible? For example, if something really worked for you for the task before, like, do I have to redo it completely? Or can I remember those patterns and do them again in the way that you expect them to be?

50:11This is part of the AI engineering I think we will need to, again, get familiar with in a way that is not kind of throw away, but we are building the layers needed to be able to satisfy the user's expectations and needs. Yeah. And you mentioned earlier that you have to do this fast. Because while you're doing the research on the safety and constraints and guardrails or whatever, the tools are out there and people are building agents and they're building networks of agents, a society of agents. I mean, presumably, you know, it won't get so far ahead of the safety research that it can't be retrofitted.

51:01But do you worry about that? I worry about it all the time. I worry about it all the time. The reason is that the field is moving very fast. And this is like the old sayings, like, this is the best of times, this is the worst of times. you know like it's the best time to be a researcher because the topic you are working on is one of the most important topics in the world and talking about industry the way that product groups are building things right now like just looking at microsoft this is my dream company right when i started like everybody is going 100 miles an hour in terms of putting things into the world and the kind of the line between what is research and what is product and what is just like people building things that is blurring and everybody is innovating all the time.

51:54Like all we could do with Otocan is amazing because none of that would be possible without the whole community building together. So those are the amazing pieces of the puzzle. But now, again, speed. How do we do things in the right way when everybody is going 100 miles an hour? And I think that's the conversation we should be having. like instead of talking about oh ignoring the challenges or talking about sci-fi scenarios that we don't know if we are ever going to get there can we actually focus the conversation to the real problems that needs to be solved and coordinate our efforts to solving them yeah i think that's the real conversation i would love to see in the ai community yeah and and you mentioned when we were talking earlier the pause letter uh and and as unrealistic as that was uh and i remember at the time thinking that it was ridiculous it did focus public uh uh opinion and and more importantly, government opinion on paying attention to AI safety.

53:10So it had some benefit. I agree with you. I always appreciate everybody speaking up, especially people who have a big voice speaking up to call attention to the risks in the space. My caveat there is we should not stop there. We need some coordinated efforts then to say, what are we doing about the challenges and the issues we know about? And how are we going to come together in coming up with solutions for them? because raising the alarm bells is useful, but it doesn't really solve the problems that he has. Yeah. So coming back down to what's being done today with AI agents and agent workflows, you know, without talking about Armageddon or something, do you is this entering do you see this entering uh enterprises uh that enterprises are are building these discrete agentic workflows to to as you said expense to handle uh you know expense reports or that sort of thing yes i believe where things are going to happen first in terms of agents coming and changing things is going to be enterprise scenarios.

54:52And the reason is that I think that is where there's a lot of value. That there are all of these processes that are actually really prime for automation because they are well-defined and you can create workflows for them. And for big organizations, if you can start automating them, that's real value. that has good dollar values associated with it, which is creating the right incentive structure for people to push forward on the technology. And we are seeing a bunch of autogen customers that are enterprises that are looking into automating. And the one thing, without revealing much, the one thing that has been surprising is that some of those enterprises are not the big technology enterprises you think about.

55:45They are all different sectors that are looking into gaining more efficiencies by using agentic technologies. And they teach us how to do things. So they are actually like these unexpected experts are emerging in enterprises that are kind of coming to you with their agentic solution. And I'm like, wow, I'm amazed that a non-technology company is able to develop such solutions yeah this is kind of the beautiful thing we have with llms and agentic solutions that the bar for developing something is really dropping now right right that you can start putting something together pretty easily without really needing a phd in agentic systems right so i come i completely believe that the first frontier of these agentic systems are going to be the enterprises where there are these really interesting workflows to be automated.

56:46And I'm already seeing that change. Yeah. Even with Autogen though, to build an agentic workflow, you have to be familiar with code, with programming. Do you see any no code solutions that coming that where a consumer can move blocks around of an agent to read your emails, an agent to answer certain kinds of emails, an agent to do something else? Yeah, the low-code, no-code space is going to be very interesting when it comes to agents. My team has developed a library called Autogen Studio, which complements Autogen in terms of bringing a low-code, no-code development experience into it. It is more of a researchy prototype that we have developed, but it also gained quite a bit of popularity in the open source community.

57:49And then at Microsoft, there is the Copilot Studio product that they are developing for the low-code, no-code. Not because we are just predicting, we are seeing. It is evidence that there's a lot of demand in the no code, low code part of the agent development. And I expect that space to grow further and further. The other thing we are now putting some research thinking into is like, are there ways to create agents that are reusable and composable? What I mean by that is, let's say somebody has created an agent that is really good at interacting with the web. Yeah. some other person may need it for their application.

58:35And instead of creating that agent from scratch, they may actually now say, get that agent and put that agent into their team. And if that is true, and we are getting some research signals that we can indeed start building like these composable teams and that can be reused for different purposes, then we can actually start unlocking some ecosystem thinking. What is that ecosystem that people can start sharing their agents and eventually start hiring these agents? And could this become like a marketplace in the future where, let's say, you need some text agent, you actually go to TurboTax or whatever the company and kind of hire that agent on demand.

59:20And now that becomes part of your team. This is kind of the society of agent thinking that we are doing all the time. I'll mention one more direction that we are pushing for. So my lab is also doing a lot of work on small models and particularly customizing and specializing small models for different capabilities and domains. And that's another area where I see a lot of promise because once you understand like what these agents are going to be specializing at, they don't necessarily need to use a very big model anymore. because now this agent's focus is becoming so specialized and well-defined that I can actually start teaching a small model to power that agent really well.

1:00:10So the other thing we are seeing in the community right now is this kind of the synergy between model development and agents, and these do not have to happen independently. There could be like a synergy between how we are training the models and how those models are powering agents in return. And just zooming in so that we can be a lot more efficient about this development path.

From the publisher

This episode is sponsored by RapidSOS. Close the safety gap and transform your emergency response with RapidSOS.

Visit https://rapidsos.com/eyeonai/ today to learn how AI-powered safety can protect your people and boost your bottom line.

 

 

In this episode of the Eye on AI podcast, we dive deep into the world of AI agents with Ece Kamar, VP of Research and Managing Director of AI Frontiers Lab at Microsoft.

 

Ece shares her unique insights on the future of AI, discussing how AI agents are reshaping the way we interact with technology and perform tasks.

 

Throughout the episode, Ece explains the groundbreaking potential of AI agents, describing how they act as autonomous entities that can perceive, learn, and carry out complex tasks in real time. She discusses the revolutionary shift from traditional AI models to agentic workflows, highlighting how multi-agent systems like Microsoft's AutoGen are creating scalable solutions for industries and everyday life. Ece also shares her thoughts on building responsible AI, touching on the ethical challenges and safety concerns that come with the rise of autonomous agents.

 

We explore how multi-agent systems can scale to millions of agents, and how they are transforming enterprises by automating complex workflows, personalizing customer experiences, and pushing the boundaries of AI development. Ece’s perspective on the future of AI in scientific discovery, as well as her work in responsible AI, offers a thought-provoking glimpse into what lies ahead.

 

Don’t forget to like, subscribe, and hit the notification bell to stay updated on the latest in AI, automation, and ethical tech!

 

 

Stay Updated:

Craig Smith Twitter: https://twitter.com/craigss

Eye on A.I. Twitter: https://twitter.com/EyeOn_AI

 

 

(00:00) Preview and Introduction

(03:11) What Are AI Agents?

(04:43) Building Responsible AI at Microsoft

(10:55) The Rise of Agentic Workflows

(12:30) Multi-Agent Systems and AutoGen

(18:04) Scaling Multi-Agent Systems

(20:22) The Creation and Evolution of AutoGen

(23:07) Real-World Applications of AutoGen

(25:52) Large-Scale Simulations with AI Agents

(27:36) The Role of AI Agents in Scientific Discovery

(31:20) AI Agents and Complex Reasoning

(36:49) Challenges in Defining Agent Boundaries

(39:12) The Risk of Agents Interacting with Each Other

(43:59) Building Trustworthy and Safe AI Agents

(48:44) Learning from Human Factors in Automation

(50:50) Why Speed and Coordination Matter in AI Development

(55:08) The Future of AI Agents in Enterprises

(57:47) Low-Code/No-Code Development for AI Agents

More from Eye On A.I.

All 266 episodes
#214 Ece Kamar: Why AI Agents Are the Next Big Thing in Tech (Microsoft Research)Eye On A.I. · 1 h 1 min
Listen in VO