In short
Notes on The TWIML AI Podcast Episode #632: Modeling Human Behavior with Generative Agents with Joon Sung Park
Episode Overview
- Host: Sam Charrington
- Guest: Joon Sung Park, PhD Student at Stanford University
- Topic: Generative agents and their ability to simulate believable human behavior
Key Concepts
- Generative Agents: AI systems that exhibit human-like behavior through interaction.
- Empirical Methods: Approaches used to study and evaluate the behavior of generative agents.
- Human-Computer Interaction (HCI): The study of how people interact with computers and to design technologies that let humans interact with computers in novel ways.
Guest Background
- Joon Sung Park transitioned from running a startup to exploring AI research after moving to the Palo Alto area.
- Worked under Professor James Landay at Stanford, focusing on the intersection of AI and HCI.
Main Discussions Research Motivation
- Exploration of AI capabilities beyond traditional improvements (e.g., better classification).
- Generative agents aim to resolve longstanding HCI challenges by creating agents that can interact in human-like manners.
Generative AI Evolution
- The rise of large language models (LLMs) has reshaped the approach to creating believable agent behavior.
- Generative models trained on diverse data possess a nuanced understanding of human behavior.
Worldview and Common Sense in AI
- A core question is whether generative models possess a worldview or common sense.
- Ongoing debate exists regarding the interpretation of "worldview" in AI and the nature of their common sense.
Believable Agent Behavior
- Key focus on the believability of agent actions and interactions.
- Agents are evaluated on their ability to behave in a plausible manner within specific contexts (e.g., simulating daily routines).
Memory in Generative Agents
- Memory is crucial for maintaining context during interactions.
- Two types of memory:
- Long-term Memory: Stores experiences and knowledge.
- Short-term Memory: Retrieves relevant information for immediate tasks.
- The retrieval function is based on recency, relevance, and importance, enhancing decision-making capabilities.
Evaluation of Agent Behavior
- Human evaluation methods used to assess agent believability.
- Statistical methods will be utilized for larger collectives of agents to analyze emergent behaviors (e.g., information diffusion, relationship formation).
Significance of Research
- The research showcases the potential of LLMs to create agents that can realistically simulate human-like behavior without cumbersome manual coding.
- It opens new avenues in the field of AI and HCI, enabling the design of interactive agents that can address real-world problems.
Conclusion
- Joon Sung Park's research emphasizes the importance of generative agents in modeling human behavior and the challenges that lie ahead in refining these systems.
- The ultimate goal is to empower users through AI, transforming how we interact with technology and each other.
Closing Remarks
- The episode highlights the collaborative nature of AI research and the exciting possibilities that generative agents present in both technical and social contexts.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:07All right, everyone, welcome to another episode of the TwiML AI podcast. I am your host, Sam Charrington, and today I'm joined by Joon Sung Park. Joon is a third-year PhD student at Stanford University. Before we get going, be sure to take a moment to hit that subscribe button wherever you're listening to today's show. Joon, welcome to the podcast. Thank you for having me. So, Joon, tell me a little bit about your background. How did you come into AI? Right. So when I graduated from college, basically, I was running a startup. That wasn't really going to pan out, but one of the really nice things that happened during my first year out of college was I actually moved to Palo Alto area, which is right next to Sanford.
0:46So I was living here for about a year. And during that time, one of the things I got really passionate about is sort of this research scene that was going on. So I've been running this startup for about half a year. I have one other co-founder. Me and my co-founder sort of realized it wasn't really going to pan out, but we sort of along the way found a lot of different passions. And one of the things that I found was basically something like research, especially in AI field and in computer science, seems to share a lot of similarities as to creating a startup or creating a piece of art. And that we sort of delve into it.
1:18We have an idea that we want to pursue and you take responsibility over it. And you basically try to express what your thoughts are about that particular topic. So it was around then I joined Stanford's computer science department as sort of this visiting student researcher or visiting researcher. So I was already I had graduated from undergrad and I started working with a professor here named James Landay. So James is now the associate director of the institute that we call H.A.I. It's H.A.I. So that's Human AI Institute at Stanford. And it was the era where when we were just thinking of creating that.
1:54So James and Fei-Fei and some other faculty members here were really getting excited about this new opportunity that we had with AI technology that was just coming about. And I was sort of watching James from side as I was working with him and other faculty member here, Jeff Hancock, and a PhD student back then who is now a professor at UPenn at the NIMATACSA. So I was sort of working with them and seeing from sight what they were getting really passionate about, because that's one of the ways I like to find new interesting areas to work on. Basically finding really passionate people who are really smart.
2:25It's trying to see what they're seeing in the future, in the five to 10 year time horizon. So that's how I got interested in AI. And in particular, sort of the angle that I got particularly interested in was back then, as you can imagine from this Human AI Institute, this was an institute that cared in particular a lot about human and our role with technology and the way we use technology. The underlying philosophy here is we seem to have this really powerful new piece of technology, AI, that can do a lot of different things, right? And the question then became the uninspiring answer to what we're going to do with, of course, with this technology is, hey, we'll just do better classification or slightly better content moderation that improves the performance by a touch.
3:11That could be still interesting, but from the interaction perspective, a little bit less inspiring. So what I wanted to answer in that area in particular was basically ask, well, what can we really do with them? And so that's how I got into AI. A part of that journey is trying to envision what our future might look like with AI technology, but also getting a little bit more technical and go deep into what's actually possible and where we think there is a limitation, see if we can push the limitation and boundaries so we can leverage this new interactive opportunities. And is your particular area of interests from a research perspective at kind of the intersection of, like, do you incorporate traditional HCI, human-computer interaction types of approaches into your research?
3:58Or is it more grounded on the computer science side of things? Not that, I don't know, maybe that's a false dichotomy, but I'll pass it over to you. Right. So it's an interesting question. It's a bit of a mix. And looking back, it's one of those things that we don't really think about too much in that we basically use a tool that seems to be the best for the moment. And ultimately, the goal here is to create something that people enjoy and can empower people. So that's sort of what we're trying to achieve. The way this sort of comes up in sort of my conversation, one way that I think that defines a lot of my research, regardless of what technique we use, is we like to identify some of the existing problems or challenges that lasted for many decades sometimes and trying to see if we can solve them now.
4:44And a lot of the challenges that we identify are challenges from your sort of traditional HCI and AI field. So certainly with generative agents, one of those main challenges actually was, can we create believable agents? Agents that can interact with other agents and other users in a human-like manner in an open world over a long period of time. And that was sort of the vision that HCI certainly had for a long time. So think back to people like Alan Newell and Herbert Simon when they were creating cognitive architectures and cognitive models in the very early days of both AI and HCI when these two fields used to be a little bit together, right?
5:22We had this cognitive psychology folks who were really trying to pioneer that area. So we kind of look back to those eras and try to see, they had a lot of really interesting visions that we sort of forgot about and try to see, well, can we crack them again? And that's how a lot of sort of HCI and traditional AI literature informs me. It informs me of what are sort of the lost vision that we had in the long arc of our journey here, and what are the new techniques that can really solve them. And of course, the techniques are much more grounded in more recent progress we had in AI. Yeah, I was just going to ask about that.
5:56When you got started, generative AI was not the same term as it is today. How has the evolution of generative techniques impacted the way you approach your research? Right. So in a more narrow sense, if you're thinking about something like generative agents, generative AI, especially I'm talking about large language models in this particular instance, although other multi-model models can also come into play here in the future. the part that really interested us in the early days was this observation that these models seem to have been trained on broad data that include the social web wikipedia and so forth so they actually encode something really deep and meaningful about human behavior so if you poke at it the right angle it kind of knows the way we behave the way we like to talk to each other and so forth and that presents a new opportunity in terms of how to design these npc characters these chatbots, and so forth.
6:54In the past, if you look to how we used to do this, it was a lot of manual coding or manual authoring. And even in game industry today, that's what you often see. Kind of like rules engines or things like that. That's right. The rules, behavioral tree, even state machines, these are all the techniques that we use today. What's interesting about these models, though, is because they are aware of so many different contexts, It knows how an artist might behave before, let's say, a deadline or how a student like myself might behave when there is a paper deadline coming up. It knows a lot about this or what we might do to finish our morning routines, like waking up, brushing our teeth, cooking breakfast and so forth.
7:38So we have a single model that seems to encode all this. So we can just ask this model how somebody with this characteristics and experience might behave around this. And that was sort of the new opportunity that we saw in this instance. So one of the kind of big questions that's come about since the rise in popularization of large language models, chat GPT, etc., is this question of whether they have a worldview, which seems particularly germane to not just what you've talked about, but the context of open worlds and games and things like that and how an agent might behave in those kinds of worlds.
8:18based on your research, how do you think about and answer that question? Right. So this sort of comes up. So when you say worldview, I'm looking more towards the direction of what are sort of the values that were encoded in these models and so forth. Is that sort of the angle that you're looking at? I think that's part of the open question is what exactly do we mean by worldview? And different people answer it in different ways. I think sometimes the question is taken to mean, do these things have like, does the large language model have what we might call common sense? Sometimes it means, does it have a perspective on the way the world operates that's similar to a human-like perspective on the way the world operates?
9:05It seems like the premise of your research in many ways is that these models do bring a worldview to the agent that you're trying to create. Right. This is an interesting question. And this is something that we sort of ask ourselves all the time. In the field, of course, recently, there seems to be paper after another where one paper seems to say, hey, these agents or these models actually have a worldview, they have common sense, and they have all this theory of mind. And another paper comes in and say, actually, if you change this parameter a little bit or change the prompt a little bit, that all goes away.
9:36So maybe not. Yeah, exactly. Both the question and the answer here seems a little bit open-ended for me. I can sort of say like what we in this particular project assumes and kind of go beyond that a little bit and try to see what I'm sensing in this. So we do assume that these agents or these models, chatGPT4, chatGPT and these models, encode human behavior. And I think that much I think is fairly concrete. Now, the question is whose behavior, what types of behavior do these language models encode? I think it's a really important question that we continue to need to ask and answer. But I think the case that these models encode human behavior because we can actually prompt it and we can extract it.
10:20And there's been our work, certainly, but there's been another small group of work in this domain that basically tries to much more empirically answer this question. And they seem to be getting sort of a similar answer. So you can get these models to behave like human participants in certain contexts, and it will give you the responses that sort of match the human behavior in many ways. So that's what we assume. Now, there is sort of a deeper question here as to do these models have common sense? Do these models, let's say, have theory of mind? And it's one of those things where one thing to note, of course, is all the model's capacity oftentimes is sort of emergent, or that's the term that we've been using recently, where unlike a lot of traditional AI systems.
11:09Isn't there just a paper? I think there's some paper that just came out that talks about there not being emergent behaviors in LLMs. Right. I am trying to remember. So I think that was the paper by my colleagues here, led by, I think the faculty was Sam Lee, which was sort of interesting. Sam Lee and I briefly overlapped at UIUC, which is where I got my master's in computer science. And he was a professor back then. He was a really cool professor that I knew a lot of his work and I looked up to the things he was doing. And I never really got to chat with him when I was there. And now we're both here and it's kind of cool connection.
11:48So there was certainly that paper. So one of the arguments I believe that paper is trying to make was a lot of this emergent behavior, the argument that we're trying to make is if we were to basically measure a certain performance of some task on these models, the performance all of a sudden goes up when it passes a certain threshold in terms of, let's say, the number of parameters or nodes it has and so forth. And I think what this new paper basically was trying to say is, well, you know, you can actually, depending on what metric and how you use them, how you scale them. you can actually map this basically emergent behavior in a much more predictable manner.
12:23So it becomes a much more linear function, which is, again, much more predictable. And much less emergent. Right. Much less emergent. It's more like a scaling law than emergence kind of thing. Exactly. So I'm still trying to wrap my head around. So, of course, this paper had come out very recently. I saw this kind of trending both at Stanford and on Twitter like last week. So I think we as a field is sort of trying to figure out what that means. I think a part of it still remains true. At least for now, it does appear to be the case. There's a lot of surprise element. So whether these behaviors truly emergent, I think, is going to be another really interesting question to ask.
12:59What appears to be true is at least we as a community didn't expect a lot of this behavior. And whether they're predictable or not, we are finding out about them in ways we sort of, with our classic AI systems, we didn't have to. So certainly one thing that is changing is our relationship with these models and AI systems. Where in the past, we sort of modeled them and we had a very specific task in mind that we wanted to tackle. But now it is still the case that a lot of these behaviors and capacities we are finding out after the fact, after we have sort of developed these models. And that leaves sort of this interesting questions as to what do these models contain?
13:42Do they contain a worldview? To some extent, they might. To some extent, they don't. And the way I've been sort of poaching in a lot of my work is go about it in a very empirical manner. And this is where I think some of the interaction community likes to differ compared to some of the more, much more AI focused communities. And I think both have a lot of values to offer here. For the interaction community, sort of the problems that we're trying to solve is very much grounded in the human problem that we are aware of. So for instance, if we have a machine vision system that can do really good image recognition, that can help people with blindness or sight impairment to basically take photos and try to understand what's in their surroundings.
14:29So there has been really interesting work from people like Jeff Bigam who was trying to tackle those kinds of tasks. And in that kind of environment, in that mode of thinking, what you care the most about is, is the AI system good enough? for that particular idea or task. And that's sort of the mode of thinking that I sometimes like to take, where we had a task for generative agents. These are agents that we wanted to see if they can exhibit human behavior. That's sort of believable. And believable is the term that we've been using because we're not actually promising with these agents. And I think this is sort of not just our work, but I think that's been the case with a lot of this human behavior-related work around large language models.
15:13Right now, a lot of it plays around this idea of believability. So can we predict, let's say, one of the ideas that we present with generative agents is we can simulate a human society or the small society of NPCs. Now, do their behaviors predict human society and their outcome remains to be seen? What appears to be the case is when we look at them, when a human participant looks at them, they are sort of believable in that they seem plausible. Like we can imagine something like this happening. And that is sort of the mode of thinking. Is the thing that you're looking to imagine happening, the behavior of the agents, the responses of the agents to stimuli, the evolution of the world that the agents create?
16:02like there's all kinds of things that you could apply a believability test to in this environment, which I understand is kind of like the Sims or inspired by the Sims. Which of those are you focused on from a believability perspective? And then how do you even go about measuring that so that you can baseline benchmark all the things that you'd want to do as part of a research study? Right. This still is an interesting question to us too. So in terms of kind of like what aspect of believability do we really want to understand? There's sort of two parts to this. One is much more straightforward and one is much more ambitious.
16:37The straightforward one is, do these agents create believable behavior in a very narrow segment in time within a very small setting? So let's say Jun woke up in the morning and he needs to get to the office. What does he need to do? And believable behavior there might be, well, Jun just woke up, so maybe he will quickly finish his morning routine, pack. In this particular instance, maybe I walk to campus. Highly believable. Now, what wouldn't be believable is, let's say, June gets on a space rocket and then flies to Stanford campus. That's a little bit out there. So that's clearly not believable, but something like morning routine, walking short.
17:17And that's sort of the level one of believability. And that large language model seems to do really well. And I think our study and many other studies, or I wouldn't say many other studies, but studies that are now becoming interested around this idea, we're showing that that is certainly the case. Is believability in this context constrained by or a product of the environment? Meaning, do you have a rocket there that the agent could use or does the agent have a shower or a coffee maker and a toaster and therefore the things that the agent does is interact with what's in its environment and its kind of believability is conferred by the environment as opposed to the agent's decisions.
17:57Right. It definitely is. The context and sort of the agent's environment plays a huge factor in believability. Now, there are some boundaries here. Just with our common sense, it would be pretty unreasonable for somebody to get on a plane to go to five minute distance. Maybe somebody does, but it seems quite unlikely. So there's a little bit of probability distribution playing in here as well, but the environment also is a huge factor. And that's sort of the level one. And going a little bit beyond that, there's a lot of interest around how these believable agents act as sort of the stateless personas.
18:34So one of the sort of the precursor paper that we wrote before generative agents was called Social Simulacra. And there we wanted to ask, here's one setting in a social media conversation. Let's say it's a community for artists where people talk about their artistic ideas and new tools to use and so forth. what would this person say? What would another person reply to that post? Or what would a troll do? So we can basically teach developers or moderators what could happen in such a community. And Lawrence Langshmore was actually quite good at reproducing a lot of this behavior. And that was interesting.
19:08And in terms of evaluation, that's sort of where we saw the most progress. We were able to create this behavior. We've been evaluating it. And they seem believable to human participants. Now, the more ambitious goal that we have in this line is in a multi-agent setup, for instance, can these agents exhibit believable emerging community behaviors at scale? So what that might mean is we demonstrated 25 agents in our particular demo for the work that we're talking about, where these NPCs were walking around and communicating with each other and going about their days. Now, sort of the simplest emergent behavior.
19:48Now, there's been some debate around whether it is emergent behavior is sort of the right term to use in this particular instance. But for now, I stick with it and then we can discuss what that might mean if that comes around. But here's sort of the emergent behavior that we describe in three forms, information diffusion, relationship formation, and action coordination. So do these agents spread information amongst each other? So if there's an election going on, do they hear about this election and who's running in it? Do these agents form relationships? So after a day, is there new agents that this agent now knows that they interacted with?
20:27And do they remember? And do they coordinate? So if there's a party going on, do they actually send invitation and actually gather for the party? And these are sort of the initial sets of emergent behavior within this community where what's interesting is, now this actually to some extent, the philosophy here actually is heavily related to agent-based modeling, which is still a popular technique, especially in computational social science, where we have a very simple equation that models the behavior of a single agent and tries to understand sort of the large-scale behavior of a collective. And that's sort of what we're interested in in this particular instance as well, where what we are designing is a single agent.
21:09Here's an agent that behave in certain ways, they remember, and they act in this way in this game world. That's the only thing we've designed. So we didn't say anything to the system or our AI agents about these collectives, but they can still gather around and actually exhibit the three forms of emergent behavior. And we, in particular, focus on these three forms of behavior because that's what we know from social science studies that humans also exhibit. And what we find is you can design this single agent really well and replicate the known human behavior as a collective. So that's sort of the next layer that we want to get to.
21:48Now, within our study, we sort of showed that again, in sort of this kind of like, I was still considered to be like maybe level 1.5, right? Where we basically simulated these agents for two days. We see information to fusion, action coordination, and relationship formation. That was sort of very interesting for our paper. And when it happened, we were really excited when we saw that. Now you can take this further. It's sort of what I'm sort of excited by right now. This is not the perfect example. And I'll say why in just a minute. But imagine if we can run a simulation that lasts many decades.
22:26And let's say at the beginning, we set it to prehistoric era before all the monetary system and whatnot. If we just let these agents run, do they come up with a monetary system on their own so they can trade with each other and so forth? So that, if we can see that, that's sort of the next level of this sort of emergent behavior that will be truly fascinating. Now, the reason why I'm saying this is maybe not the perfect example, and this is going back to this idea of evaluation for our work and a lot of work in this space that we're still trying to tackle and wrangle with, which is when we see a behavior, like believable behavior that are generated by a large language model, it is sometimes unclear.
23:09whether they are succeeding in this evaluation because what they're generating is believable human behavior, or whether because they sort of memorize something in their data set. So if they create monetary system, is it because that truly emerged from their human-like behavior, or is it because in the training data, there was information about monetary system, and that was sort of a natural path? And that is an interesting question for us to keep on asking. We have, there is some groups of people now trying to tackle that. And we are also thinking critically about that problem. But this is sort of the flavor of research.
23:45In some sense, it's really super interesting if they could effectively implement a monetary system, even if they didn't generate the idea uniquely, if they could somehow implement that in this world, that in and of itself is interesting. but it's not evidence of them kind of collaborating and coming up with the idea of trade and building a monetary system from scratch. It's kind of two separate tasks or dimensions. Right. And from interaction perspective, I think both are interesting and useful in that if you want to see how a troll might behave in a social media site, then we are actually explicitly leveraging the fact that these models have seen a lot of this troll behavior and that's really useful because that lets us think critically about what could happen.
24:29in known communities that we need to design around. So that skill is interesting and useful. But as you're mentioning, if we want to really show that these are emergent and human-like behavior, then we also need to think about, well, can they create new behaviors that we can predict, but we haven't seen in our training data? So does this, going back to this question about measurement and benchmarking, does this mean that this kind of direction of research is kind of fundamentally very subjective and you know you set these agents out and then you know you as a researcher look and say oh you know they look to be believable versus not you know and if that's the case how might you compare what you're doing to another method like or is that even the point right so there's two layers to this so first layer so right now and when i say two layers let Let me answer this from more of an individual Asian perspective or from a small group perspective, like small groups of Asians.
25:31And then there's this perspective of large collective behavior. And the way we go about studying both of them slightly differ, even in social science, when we study humans, like actually humans in our human communities. So I'll take them from there. Now, when we're trying to sort of study the smaller groups of individuals, what do they do in the morning? Do they cook breakfast? And a lot of those assessment or evaluation to some extent is human evaluation. Now, there are rigorous ways to achieve that. So in a lot of human-centered studies or human computation studies, basically, we bring in human participants under certain conditions.
26:10It's a lot like psych experiment in some sense, where we try to put human participants who are not us. So we recruit them from online. So it's any kind of like any app, any person with experience living as human. They look at these agents and try to decipher themselves. Can I see, can I, in some, one of the studies that we ran was sort of a dispersion of Turing test. We bring in a random participant. Can that participant tell whether this is human generated or machine generated? And we can get statistical power from those kind of studies. So that's the accepted sort of the method in a lot of human computation studies.
26:50And we rely on them as well. So generative agents certainly relied on that kind of evaluation. And our prior work, social simulacra, also relied on that kind of evaluation. And to some extent, yes, it is human evaluation again. So it is how our participants thought about these agents, except we can run, again, the statistical test at scale. So that's how we study these smaller groups, one agent, how an agent might behave in the moment and so forth. If you scale it up, however, the question I think becomes a little bit different. The way we study large collectives of people oftentimes is through modeling and statistical methods.
Read the full transcript
27:31So one of the things that we see, of course, most often, or most recently, the most sort of a salient form of study was for, let's say, pandemic, right? How would virus spread across a populace? That becomes much more of a modeling question. Can we model that behavior of that diffusion? And can we predict its diffusion? And that's how we often study collectives. And my guess right now is once we are able to simulate really large collectives of these agents and we want to see whether their behavior is human-like as a collective, a lot of those modeling and statistical methods will come back. And we'll be comparing how information diffuse, let's say, amongst people versus amongst these agents.
28:19And what kind of coordination emerge? that can also be answered to some extent, right? Because we have, for instance, a lot of information about polarization in our communities. Do these agents also form those kinds of polarization within their communities? And that's something that we can much more concretely answer with something like a statistical model. And I think that's an interesting area of study in the future. Early on in this conversation, you alluded to the role of memory for these agents. Can you elaborate on where memory comes in and what the role is when you're thinking about this kind of plausibility task?
28:57Right. So there's two parts to memory. So we have what we call the memory stream, which is basically a long-term memory module that includes everything an agent has experienced. And then what might be sort of considered as a scratch memory or a short-term memory space. In the paper, we basically just describe it as retrieve memory. And the way they play in is, let's say, in this conversation that we're having right now, I should remember what my research area is. I should remember the paper I wrote so I can create this meaningful conversation. And that's sort of what the memory module, to some extent, is trying to do.
29:37It's trying to make sure that these agents maintain long enough context so that it can be used in the future. Now, if we didn't have memory, what that basically would mean is we would need to feed in everything an agent has ever experienced into the prompt. And there's sort of two or three reasons why that's not really a good idea. One, right now, it's really not possible. Now, there's been some really exciting work recently that basically extended the prompt, the context limitation. So there's a number of characters that can go into a GPT-4 or a chat GPT prompt. There's work that's coming out that raised it to a truly significant amount.
30:17A million tokens is the number that I'm hearing. And that's a lot. That's way beyond what chat GPT and these models provide right now. So I think there is some interesting work to be had there. And I love that paper. And I think it's fascinating. Are you referring to this million token paper? That's the one. Can you point me to that one? I haven't seen that one. Yes. And that's sort of an interesting way. And that's certainly one way to get it. Every time an agent has to make a decision, they would basically prompt this language model with their entire life experience. Right now, that is not possible, right?
30:49With the known sort of with the APIs that are available today. Now, in the future, with papers like this, would that become possible? Maybe. Even million token, I'm going to posit, is actually not enough to contain one person's life experience. So I think there's going to be some limitation, but I think the vector is pointing us towards a future where the token limitation will truly relax if we need it to. But there's another question here. Do we want to relax that in the first place? And I'm not so sure about that in the sense, so there are applications where relaxing this token limitation will truly be beneficial and transformative for sure.
31:26If your goal, however, is to build human-like agent, it's less clear to me that that is the case for two reasons. One is we make decisions and we process information at an extremely high rate. Every sentence I say is new information coming in. Every sentence I hear is new information coming in. In each of those instances, I need to make decisions on how I'm going to behave. I drop my cup. I need to decide on what to do. So it's going to be highly inefficient from the time perspective to have to reconsider all my life choices when I need to decide what I'm going to do when I drop my cup, right?
32:03And there's also the computational aspect here as well. The cost also increases. So that's one. So I think that's one reason why having sort of this long-term memory module is sort of an interesting way to talk about this, where it is the way to bypass the need to basically put in all of the life experience of a person into a prompt to reason about. And one more element that I quickly mentioned here is there's been a lot of interesting work. So think chain of thoughts, paper, and so forth, or think step-by-step paper, where the idea here is these language models often get distracted by unnecessary information and they perform better, especially when we make some of the reasoning processes explicit in their output and in their prompt.
32:53And that's what, to some extent, long-term memory allows you to leverage, right? So you could put in base clip. So in our conversation right now, I could be thinking about what I ate for breakfast two weeks ago. Now that's not exactly relevant. And now it may have become relevant because now I'm talking about this. But up until now, that's not really relevant. And it is certainly distracting, especially when we're trying to prompt a model with such information. On the other hand, you can much more selectively retrieve the portions of memory that matters the most for the moment. Then that's going to make the reasoning process much more explicit and much more narrow to up the performance of these agents.
33:37Those are the reasons why I think this, we were particularly interested in implementing this memory module. And the idea that we had to retrieve certain pieces of memory from long-term memory and put it into short-term memory, that was our sort of the motivation behind all that. And is the specific, the actual mechanism that you use to implement that? Is that an interesting discussion point or was just some kind of text blob that the agents would have access to? Right. So I think there's a few things that I quickly highlight here. I'll highlight two things. One is sort of what does that long-term memory, like that database look like, because it's different than what we had in the past.
34:17And how do we retrieve it? Now, what's unique? The opportunity that we found with this work is when we are creating this memory stream, which is, again, the long-term memory module. We went in more from AI researchers' perspective today. So we decided at first, hey, can we use knowledge graph? Because that's really popular and powerful. So knowledge graph is this graph, of course, that has object, predicate, and subject that connects different nodes. It's kind of a compelling idea. You have these memories, you want them connected to each other in different ways, and you want to be able to traverse that to maybe create an experience or recall an experience or something.
34:56Is that the idea? That's exactly the idea. Yeah. And especially because a lot of this information that we need to retrieve our associative, right? So knowledge graph certainly made a lot of sense at first. So we went in, what we found was knowledge graph still can work. So I think it's entirely possible that some other researchers might come in and try to implement generative agents with a knowledge graph. And it could still work well and maybe might have different strength. But what we found in our work is using a very structured database actually downplays the opportunity or the strength we have with a large language model.
35:35A large language model can basically parse natural language really well. So we can actually create a database that's really simple. It's just lines of text, just lines of natural language text, and that is it. And we just retrieve different lines during different moments as we need. And we found that to work surprisingly well. So that is the structure of our database. And this is kind of a theme for this entire architecture where the core medium that connects all different modules, memory and planning, reflection, everything, is everything can run in natural language. It's text. Yeah. The world itself, the interactions in the world itself were fundamentally text-based, or did you have to go from text to some command structure or some interaction structure or something else?
36:26Right. So that was also mostly in text. So the grounding of these agents happened in text. So they will reason about everything in natural language. So all the way from, I want to cook breakfast. Where should I go? Or what are the locations that I'm aware of? That reasoning happens in text. And at the last moment, we translate that into an action in the game engine. So if an agent is in their, let's say, bedroom, and they need to go to the kitchen, let's say, to cook, then we can use a game engine to move this agent from one location to another. And there's a little bit more detail that I think is a little bit more of implementation detail where each of these agents basically contain a tree or a really basic form of scene graph that lets them know what the object is where so they can reason about that.
37:14But all the reasoning itself is done in natural language with a large language model. So that was sort of the interesting binding or philosophy that we have where, well, can we really leverage natural language to an extreme? One last thing I'll quickly mention is the retrieval function. I sort of mentioned two things about the retrieval function, which is another interesting aspect of how we develop this memory module. So that's the one that takes things from long-term memory into the short-term memory. That's what decides what's important in the moment. Our retrieval function has three aspects, or it's a combination of three variables, recency, relevance, and importance.
37:51Recency is just an exponential function or exponential decay function. So the more recent elements are more likely to be retrieved. That's something we are familiar with. The relevance function is another one that's sort of all familiar to us. It's basically cosine similarity of embeddings, the text in memory and the current context. So we're doing this podcast interview about generative agents, then I'm going to bring in relevant pieces of information that's about generative agents. The last element was sort of interesting. It's called importance. How salient or how poignant is a certain event?
38:26So for instance, what I ate for breakfast two weeks ago, Yep. I already forgot about what I ate. Not really important. It's somewhat trivial. So those will be scored really low. But I broke up with someone or I graduated. That's a really salient moment for a lot of people. And that scores really highly. And we basically use a large language model. And so is this implemented as like a novelty score that the LLM produces or surprise score or something? That's basically right. It's a prompt that literally asks the agent or the model, this happened, how important is this to you? Okay. And that worked surprisingly well.
39:06And that was sort of the last opportunity that we found in there. So that's our retrieval function. One thing I will note is with this particular work, what we wanted to present is more of a framework of thinking about such a retrieval function. And we think those three variables are going to be sort of important. We think those are sort of the sensible variables to use. And certainly recency relevance, these are variables that we as a community already know a lot about. Going forward, however, there's a huge literature, there's an entire community of information retrieval community or experts in this kind of tasks.
39:41And I think there's a lot more value that can be added to something, a retrieval function, if we can leverage the community's knowledge. And in our work, it was more of a framework. So we basically implemented retrieval function to be something relatively simple and easy to communicate. But there's a lot more you can do in this space is our perspective on this. So you're calling out the importance of the retrieval function. You're not saying that these are the three end-all be-all components of it, but rather they get you pretty far and kind of we can pull in from adjacent areas to figure out if there are other useful elements of it.
40:16Yep, that's right. With an asterisk, we do think these three elements. So our opinion right now is that these three elements are going to be important. So a lot of the future ritual function, if we had to guess, will have them. And things like the importance function that we have here is interesting because it's a prompt that's sort of new or that feels fresh. So we think these could be interesting, but rest is exactly as you said. Yeah. The importance thing makes me think of conversations that I've had with folks that are working on the cognitive science side that talk about how human memory is kind of conditioned by, you know, some degree of surprise or novelty.
40:54So that is consistent there. But my immediate thought was maybe you're doing some kind of probability distribution and implementing it based on something like that. And just, let's just ask the LLM. It's awesome. It's been interesting to see how much we can offload to a large language model and how well they actually do in practice. Do you ablation studies around the particular LLM that you used, or did you use Alpaca or something like that? What did you use and how specific to the LLM was your research result? So we use ChatGPT. So for the API, that's basically, I believe, GPT 3.5 Turbo is the one that we used.
41:38Now, it is an interesting question though, as to whether the language model that we use is going to impact the way these agents behave. To some extent, they will. And to some other extent, they may not. So we originally implemented this entire architecture first with GPT-3 because ChatGPT wasn't a thing when, or was just becoming a thing during the early stage of this project implementation. And then during the last month or two of the project, my advisors, Michael and Percy mentioned, hey, can you change this with ChatGPT? Because we're getting a little bit behind. And we did. So there was one or two all-nighters where we tried to reimplement everything with chat GPT, but that's sort of the process.
42:18But that is to say, it works with GPT-3 as well. Was that kind of a buzz capture bandwagon thing, or was it intuition, or was it principled like, hey, the RLHF process is going to aid the performance of the agent? Or both, I mean. Right. Honestly, it was a little bit of both. But one certainly was we've been using ChatGPT in different contexts. ChatGPT was much better at listening to a lot of instructions and perform better on some benchmarks that we cared about. And we thought, hey, if we can implement this with ChatGPT, it would be better. And that was the case. It was certainly better. And then there was this element where, well, we want to demonstrate this system with sort of the most popular model.
43:04Sure. And the most popular model was becoming ChatGPT. So it made sense. Got it. Just to recap, it's an open question as to the sensitivity of the results to the LLM because you've not tested others yet. That's right. You need another all-nighter or two to introduce another model. So just to kind of recap, like this research kind of explores broadly the behavior of LLM-based agents in sim-style environments, which interestingly, they're, you know, graphical environments, but the interactions are fundamentally text-based in nature, it sounds like. And you found that the agents' behaviors emerge, whether we call them emergent behaviors or whatever.
43:50They communicate with one another about events and other things that suggest kind of social behavior that comes from the language model as opposed to you programming it into these agents. And what is the ultimate conclusion of the research in your words? Right. The conclusion or the message that we want to have is we believe that we can actually now start thinking about building these agents, building human like a believable agents. This was something that we as a community wanted for quite some number of decades. And in the past decade, what we've seen is we sort of abandoned it for the right reason.
44:27We didn't have the right angle to tackle this. It was too difficult. It couldn't be manually authored. That was sort of the main way we had. we see a new opportunity. And I think it's an interesting area that's opening up for us. Got it. So the upside here is that for the first time with LLMs as kind of this fundamental tool, we've got agents that exhibit behavior that is kind of plausibly like human behavior without, well, not without, like we've not been able to do it before because rules are hard and cumbersome. And how do you encode human behavior and rules anyway? But now like we're kind of there with LLMs and the next steps are, for example, the retrieval function.
45:10And, you know, what are some of the other, you know, memory, how do we implement memory? And like, there's a bunch of mechanisms around this LLM core that we, you know, are yet to kind of figure out, but you've kind of demonstrated that LLMs can do this thing that we've been trying to do for a really long time. That's right. Awesome. Cool. Well, June, thanks so much for taking the time to share this research with us. And kind of talk through this exciting area. Thank you for having me. This was a lot of fun. Thank you.
45:42All right, everyone, that's our show for today. To learn more about today's guest or the topics mentioned in this interview, visit twimla.ai.com. Of course, if you like what you hear on the podcast, please subscribe, rate, and review the show on your favorite podcatcher. Thanks so much for listening and catch you next time. Thank you.
From the publisher
Today we’re joined by Joon Sung Park, a PhD Student at Stanford University. Joon shares his passion for creating AI systems that can solve human problems and his work on the recent paper Generative Agents: Interactive Simulacra of Human Behavior, which showcases generative agents that exhibit believable human behavior. We discuss using empirical methods to study these systems and the conflicting papers on whether AI models have a worldview and common sense. Joon talks about the importance of context and environment in creating believable agent behavior and shares his team's work on scaling emerging community behaviors. He also dives into the importance of a long-term memory module in agents and the use of knowledge graphs in retrieving associative information. The goal, Joon explains, is to create something that people can enjoy and empower people, solving existing problems and challenges in the traditional HCI and AI field.




