⚡️How Claude 3.7 Plays Pokémon

4 Mar 2025 · 38 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Summary: ⚡️How Claude 3.7 Plays Pokémon

Podcast Information

  • Title: Latent Space: The AI Engineer Podcast
  • Episode: ⚡️How Claude 3.7 Plays Pokémon
  • Host: Alessio, CTO of Decibel
  • Guest: David Hershey from Anthropic
  • Date of Episode: 2024
  • Description: David Hershey discusses "Claude Plays Pokémon," a project where the AI model Claude attempts to play Pokémon Red on Twitch. The episode dives into the development process, the challenges faced by the AI, and insights into the capabilities of AI models in gaming contexts.

---

Key Highlights

Introduction to Claude Plays Pokémon

  • Concept Origin:
  • The project began as a personal experiment for David Hershey to test the capabilities of AI agents.
  • Inspired by the nostalgia of Pokémon and the interactive nature of Twitch Plays Pokémon.
  • Current State:
  • Claude 3.7 is currently in a live Twitch session attempting to complete Pokémon Red.
  • As of the episode, it has been navigating through Mt. Moon for 52 hours.

Development Insights

  • Technical Framework:
  • The AI uses a custom harness built to allow it to see the game screen, navigate, and remember game facts.
  • A Slack channel was created for sharing updates, resulting in a community following.
  • Implementation Challenges:
  • Initial versions struggled with navigation and game comprehension.
  • The current iteration of Claude utilizes a "navigator" tool to enhance its spatial awareness.

AI Interaction with the Game

  • Game Mechanics Understanding:
  • Claude has some knowledge of Pokémon mechanics (e.g., types, weaknesses) but occasionally misinterprets them, leading to comedic errors.
  • It learns and adapts based on its experiences within the game.
  • Emotional Connection:
  • Claude exhibits attachment to Pokémon it names, demonstrating a rudimentary emotional intelligence.
  • The AI shows care for its in-game characters, enhancing the overall experience for viewers.

Technical Architecture

  • Prompt Design:
  • The conversation history and prompts are designed to facilitate long-running tasks.
  • Uses a knowledge base with a token limit to optimize memory usage.
  • Token Consumption:
  • Each interaction consumes a substantial number of tokens, with costs potentially reaching thousands of dollars for extensive experimentation.

Future Potential

  • Model Development:
  • The conversation suggests that future models will improve in visual recognition and navigation, making them more effective in gaming scenarios.
  • David emphasizes the potential for real-world applications of such AI models beyond gaming.
  • Exploration Beyond Pokémon:
  • There are discussions about extending Claude's capabilities to other games, including Magic: The Gathering.

---

Key Takeaways

  • AI's Learning Behavior:
  • Claude exhibits a learning curve, where it becomes better at interpreting game mechanics over time.
  • The emotional connection it forms with Pokémon illustrates the possibility of AI developing personality traits.
  • Development Challenges:
  • The project showcases the limitations of current AI models in visual comprehension and spatial understanding, indicating areas for future improvement.
  • Community Engagement:
  • The Twitch integration and community involvement highlight the growing intersection of AI and interactive entertainment.

---

Conclusion The episode provides a fascinating insight into the intersection of AI technology and gaming through the lens of Claude's experiences in Pokémon Red. David Hershey's work demonstrates the potential for AI to not only perform tasks but engage emotionally and socially in interactive environments, paving the way for future advancements in AI applications.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:01Hey everyone, welcome back to another Latent space lightning pod. This is Alessio partner and CTO a decibel. there's no swix today we got a special co-host vibu which if you're a part of the latent space community on discord you've definitely seen um welcome people as a co-host first time what's up guys and then we had david hershey from entropic today who's the person behind club plays pokemon it's funny i saw we first dm'd about playing magic the gathering together and nsf i really and then people are like on all of the different nerd angles you can get me uh glad and then people are like david is the person doing this and i was like okay i'll dm him and then um yeah it was cool we already had a touch point so welcome to the to the show this is our second entropic episode we are eric schlons from the sweet agent before so welcome thank you glad to be here excited to talk pokemon yeah so let's give a little background on this so sonnet 3.7 came out a couple of weeks ago.

1:01I don't know. Time goes by this week. I don't know, man. It feels like two weeks ago. And then you had this Cloud Plays Pokemon thing that kind of went viral where if people remember, there used to be this thing called Twitch Plays Pokemon where people could go on Twitch and kind of type in the chat and then busy like figure out what the next section that the emulator would take us. What you've done instead is given it to Cloud and basically have Cloud figure out how to walk through it. I'm looking at it right now. So far it's been stuck in Mount Moon for 52 hours. Poor guy. I probably met a 15 ,000 zoo vets.

1:34So yeah, let's talk about what gave you the idea for it, kind of the origin story that we can go through the implementation. Totally. Yeah. So I actually started working on it in like June of last year for the first time. And for me, so I work with customers at Anthropic and I just like really wanted to have some way for myself to be able to like experiment with agents, like in a real way, some framework, some harness where I could actually just like go to town and try some different things and see what actually worked to get Quad to do like pretty long running tasks in general. And so I like had that in one hand and then I was like, okay, what is the thing that will make me the most addicted to making this work?

2:11Like how will I grind the hardest actually trying this? And Pokemon was like a pretty clear answer. Someone else at Anthropica actually like tried once to hook it up. So I had a little bit of like the shell of what I needed to actually put together and to like kick off what became an obsession a little bit in the coming months so yeah like i played with it in june in this like trying things out this was like sauna 3.5 came out in june of last year was just when i started kicked it around it's very good but like you know you could see like kind of signs of life but like not much really happened um and then ever since then as we released new models it's sort of been like the way that i get to know one of our new models a little bit right so we released the new version of sauna 3.5 in october and like use this to like really kind of see like what's it better at and it got better like you could see it start to like it could get out of the house somewhat reliably which was not always true and it got a starter and it like even named it sometimes like it was like doing stuff not great but like it it could move um along the way too like i'm just like we have a quad place pokemon slack channel like i'm sort of just like giving people updates so over time as i'm like posting gifs and up parties updates like i'm this is like slightly growing a popularity of a cult following internally of people who are somewhat interested but then like uh you know a couple weeks ago i was bashing an early version of sonnet 3.7 and it just like you could just tell it had like it was a little different it's clearly not still good as you said at the top like it's it's in mount moon for its 50 somethingth hour this is a little bit worse than average from what i've seen so far by now but like this is like you know about on brand it doesn't really have a great sense of direction it's pretty bad at seeing the screen stuff like that but like it plays the game you know like it gets pokemon it catches pokemon like it caught its first pokemon it got out of the radio the first time like a whole bunch of stuff happened for the first time or like could squint and see a thing play in the game and yeah like posting updates obviously internally it was very fun like people were just like kind of going wild at the fact that this was actually happening finally um yeah and it was like entertaining enough that i could kind of see it and the other side is like we kind of just like got finally a sense that this was like an actually useful way to measure what was going on with this model you know what i mean like there's one thing that it's like fun and fun follow along but like internally like i think we got more of a sense that like you could actually use this as a bit of a measuring stick for what's going on in the model i've spent you don't know how many hours i spent staring at quad play pokemon i've so much i have to have seen and read like millions of words that claude has generated in the course of playing pokemon over the last eight months so like you can kind of get feel for like what's actually going better or what's it getting better at and that kind of thing and with this particular release like i think the fact that it got this much better at this kind of reflects a lot of things that we wanted to be true about the model to begin with and uh and those sort of lined up were like okay maybe this is like an interesting way to actually tell people about what's going on here for a crowd that maybe doesn't like quite know as much about software engineering and all the other ways we've told people about agents in the past yeah uh were there any other games that you consider to me seems like pokemon is good because it's like you know isometric you know it's kind of like flat so you can use cord and it's it doesn't have too many hidden facts about objects you know kind of like everything it's described did you consider anything else or was pokemon just kind of like by far and away the first choice i didn't but it's mainly because like pokemon was the first game i ever got as a kid right it's this is like purely coming out of my own nostalgia um but also like the twitch play pokemon like i was also something that i cared a lot about well a decade ago or whatever that was um please tell me it's not a decade ago i think it's actually ago i'm sorry yeah painfully um and 11 years ago yeah february 2014 yeah that is nuts okay pokemon red is 20 years ago oh my god 20 25 at least so yeah i for me it was that like since then there have been a lot of people in the lobby We were like, oh, we can do this.

6:06We can do this. We can do this. I think there's like a lot of fun things you can do. Pokemon's actually really nice because like if you don't do anything for five seconds, like there's typically not a consequence by the nature of like doing inference on a model every like a snapshot of time. It's actually a pretty good game to be able to do this with. But yeah, it was mostly just like my love for Pokemon coming through here. You put together a very nice architecture diagram. Do you want to screen share that so people on YouTube can follow along and then we'll put in the show notes um if you are just listening um you got it i know that vivo had a bunch of questions on the on that too yeah let's do it very very straightforward questions basically can we just double click into all of it yeah yeah it's easy i found it off twitch and like no one was talking about it so i started sharing it around and i lost the original source but basically everything in here is like pure gold the memory is a little interesting but yeah if you want to just go through high level yeah you got it yeah i i want to like preface that i do not claim this is like the world's most incredible agent harness in fact like i explicitly have like tried not to like hyper engineer this to be like the best chance that exists to beat pokemon i think it'd be like trivial to build a better computer program to beat pokemon with claude in the loop this is like meant to be some combination of like understand what claude's good at and benchmark like an understand quad alongside a simple agent harness.

7:32So what that boils down to is this is like a pretty straightforward tool using agent from my perspective is how I would frame it. So at the end of the day, like the core loop is just like having a conversation that rolls out and it's essentially like you build the prompt, including like everything we've had up till now. You call the model. It sends back some tool use. typically you resolve those tools and then talk about summarization but like basically some a few different mechanisms to maintain the information you need to do something long running inside the context window um so like what this boils down to is like when you think about what an actual prompt looks like it rolls out kind of like this you've got tool definitions which describe three tools that i'll get to in a second a short system prompt it's like pretty boring it basically tells the model how to use the tools and like there are about six facts about pokemon that i give it and like a few corrective things that i've seen it do like really horribly wrong i'm like hey you might want to consider doing this a little bit better but it's like really not a lot of system prompting going on uh we have that knowledge base which referred to you i'll talk about this is the main way it stores like long-term concepts and memories as it's operating over time um and then the the bulk of things is this conversation history which is it's like a chain of tool use there's no like user interjections at all for the most part so it's like go and then the model uses the tool and then it gets a result back and then it uses another tool it gets result back so uh pretty straightforward uh feel free to like cut me off too if you've got questions along the way but otherwise i'm gonna keep rocking yeah yeah go ahead cool uh okay so most of the money of this is just like in the tools themselves when you think about what's going on it's really like it can press buttons and it can like mess with its knowledge base and that's about it i'll talk about navigator separately because that's like a patch for how it actually can deal with some of its vision deficiencies um using the emulators basically like execute a sequence button presses it'll say like press a b left right whatever it gets back a screenshot and screenshot overlaid with coordinates of the game these coordinates are used for this navigator tool that i'll describe in a second but it's just basically like help quad get a slightly better spatial sense of what's going on on the gameboy screen i've been through it a lot with the sorry does that come with the emulator or are you adding those in i add that in okay i have somewhat extensively reverse engineered Pokemon Red by this point to like extract roughly every bit of possible information from it.

10:14I don't use most of it, but like I have essentially everything you could know about the current state of the game I have exposed programmatically to be able to tinker with it at this point. I was just reading this diagram like, yep, you just get what spaces are walkable based on what's stored in RAM, and I'm like, oh, you definitely reverse engineered this little animal. Yeah. The good news is we also released Claude Code this week, if you saw that. And that has been This would all not be possible without the help of having Quad also go figure out how to do all of this for me. Because I could have done it, but there's a lot of tedious, here are addresses in memory, map that to a Python program that I had no interest in doing.

10:52So thank goodness for Quad Code. So yeah, it gets these two screenshots. It gets like a small blurb of state, which I read straight from the game. There's a lot of this here. Actually, funny enough, the thing that matters is location. Claude will like pretty aggressively hallucinate that it succeeded in transitioning between zones if you don't like tell it it did not uh this just comes down to like literal vision issues and and so like most of the patching of extra help i've given it been like attempts to make it so they could still play despite not being very good at seeing game boy screens in particular um and then it gets like a handful of like reminders this is this reminder says a decent amount of work but it's like things like, you know, remember to use your knowledge base occasionally.

11:40And we tell, we tell if it gets like stuck, for example. So if you detect that like hasn't moved in 30 spots or 30 time steps, I once saw it, see like a red box on the screen that was like the doormat and think it was a text box and spend 12 hours pressing a overnight to try to clear the text box, which you see that happen once and you add in some, some helpful reminders to not do that. How much knowledge does the model have about the game itself? You know? So for example, types, right? Yeah. Doesn't know about types, weaknesses and things like that, or how much are you trying to put into it? Yeah.

12:17If you go to quad.ai, like it, it, it will tell you about like some stuff. I have not yet decided if the knowledge that it has about Pokemon is helpful or harmful towards it playing the game. um like half of the time when it's like oh i know this about pokemon it then like uses that to hallucinate something so for example beginning of the run on twitch you saw like go out of the lab and see like this npc in the bottom of pallet town and be like it's professor oak i found him and it's like very much not professor oak but like the fact that it has like indexed on this concept is like a little stuff like that that it's like unclear to me where it is but it clearly has some information about it.

12:56There's like a million game guides about Pokemon sitting on the internet. It's unsurprising that like, there's a decent amount of information there. I don't really give it a lot of extra information. It picks things up. I watched on the stream the other day, like it tried to use Thundershock on a Geodude and it failed. It's like, hmm, I forgot about that. That does not work. And so like, clearly there's like, it knows some stuff. It's not perfect. It picks some stuff up as it goes through the run. ideally for me like i think it's just interesting to see like what it actually learns as it's playing so the more it does that is the more i'm like actually interested in it yeah the one of our discord members not junkie had a good question about the sense of self yeah like sometimes it gets confused who is the actual playable character in the in the scene like how do you steer that uh yeah i think like sometimes it gets confused uh it can be applied to many things in quad playing pokemon in particular when it's trying to like look at the screen and understand what's going on so i have like attempted to prompt it all sorts of ways like you are at this exact coordinate and you're in the middle of the screen and you're wearing a red hat and things like that and like that's all neat but quad doesn't particularly understand like the middle of a gameboy screen and a whole bunch of concepts like that which means like you can prompt all around everywhere but like this kind of like spatial awareness and where something is with perspective something else is something that quad still just like not great at in its current incarnation so one of the side of this is it sometimes loses track of who it is on the screen and thinks there's something else there i'll keep tracking through this so i hinted at this like other tool that i give it called navigator and this is just like the only other patch that i have for the the vision issue so navigator basically what it does is like quad can say it wants to go to one of these coordinates that we provide in the screenshot.

14:50And then we automatically press the buttons to get there. It has to be something on the screen. I'm not trying to let Claude just navigate a whole map by asking too politely. But one thing you'll notice if you run it without this tool is if Claude wants to get from one side of a wall to another side of the wall, it happily just tries to walk through the wall repeatedly because it doesn't quite have the concept of what's between it. And I spent a lot of time like prompting around this and it just like, isn't, it's just not, it's one of those things not very good at. So in order to make it somewhat fun to learn from quad playing Pokemon at all, we use this navigator tool, which like helps it actually get around a little bit better.

15:28So since we covered a bit about the different tools, the prompting and the strategies, I'm curious how many tokens all this is using. Like there's a part to conversation history and truncating parts of the most citizen state, but like, yeah, at a high level, how many tokens is this using? And then can we kind of go into where those are coming from? What's being truncated? Yeah, you got it. Uh, when you like think about the, the prompts here, uh, essentially like every step, something that looks like this gets sent. So if we just go through what each of these looks like, uh, everything in the system prompt is probably like a thousand tokens, pretty small, like handful of paragraphs knowledge base.

16:09I let get up to like 8 ,000 tokens. right so i put some like arbitrary cap on it so it doesn't go to like claude will write put a whole bunch of bs in there if you just let it keep writing stuff so like the cap helps constrain it to like try to think about what's actually important a little bit and then the conversation history i haven't like kind of finicky but it basically rolls out um 30 messages that's actually like something you can tune i've tuned it to be 30 messages about like the best performance i've gotten. And so what that means is it basically like uses a tool, get a response back, use a tool, get a response back.

16:45It's allowed to do that 30 times. And then at that point, it triggers the summary, which takes that conversation history, summarizes it, makes it the first user message, and then we kind of roll back out again. So the bulk of the tokens end up being in the conversation history once it's its longest. In fact, like this, the bulk past that ends up being these screenshots which are scaled up a decent amount to to fit in i do actually like i allowed to see a number of the previous screenshots but not all of them because you start like it ends up being a ton of context if you let it see like even 30 turns worth of screenshots so i trim out a few that's where the bulk of the actual tokens are so in practice this rollout ends up like at max ending up around 100 000 tokens i think is where it is like the longest message you ever send to the api on one of these turns and it will it will fluctuate in like summarization depending on state of knowledge base probably between like 5 000 and 100 000 tokens and is that like per action state of the game and roughly do you have like a high level ballpark estimate of how long this how much and how long it costs to run this like let's say people want to compete and yeah yeah like how much i think you'd really want to think about running this as a side project in terms of the impact on your personal wallet and how much you care about pokemon it's not clear to me that without the blessing of anthropic i would have decided to take on take on this project for my own wallet's sake uh especially if you want to like experiment and like try 10 different things i mean it's it's costly i don't know like i haven't spent a lot of time on the exact number it's not that hard to estimate if you like i just told you a bunch of numbers, you can kind of back it out.

18:31Uh, but like, I think to like do a lot of experimentation, there's like at least thousands of dollars of tokens being consumed. So it's not a, it is not a, uh, a cheap rollout. Yeah. But yeah, in the scheme also how some people use tokens, it's not terrible. How many turns are you keeping in memory before you summarize? It's 30 right now. Yeah. I've tried more and less. i think like one thing you see a lot when you talk to people building agents is there's like some effective context length that actually like has the model be the smartest um and that seems to vary slightly model by model but but for this model for whatever purpose like this 30 message work better than 20 and better than 40 so uh kind of plot in between those that it worked pretty reasonably yeah does that change based on location like how many would you want to give it to get it out of mon moon so hey we gotta we gotta bring it and plot home we can't let him say yeah yeah for another 57 hours i actually am not sure it does like i i've i've tried posting like i can have a ton of screenshots like 20 or 30 screenshots at a time be able to see and it's like not obvious that like that temporal concept is actually super relevant relevant to it and again this is just like trust me as someone who has spent like a lot of hours obsessing over this uh you can try to prompt quad a lot of different ways to understand how to navigate better and anything short telling it exactly what to do does not improve it's like actual navigation it's just like not a skill it's great at it's like good enough to to like random walk its way through some of the complex mazes and in like good easy areas it's pretty good at popping around but yeah i think i i can tell you if there was like a way to prompt this slightly different that uh would navigate better i would believe there is something but it is not like it is not an easy lift yeah yeah i asked the i just asked claude ai right now how do you get through a mountain moon in pokemon red it does have it does have a plan but i don't i don't i don't know i don't know if it's the right i don't know if it's the right plan i have seen it come up with a lot of answers to that question and most of them right this is part of the pain when i talk about i'm not sure if its knowledge is better or worse like usually fixate like oh i know the exit is on the eastern wall and it just like spent 12 hours trying that um and i yeah it's like unclear to me that that we're actually not just like harming it by having it think it knows the answer yeah i think that's the interesting part right like you don't want it to just know the answer yeah like model clearly knows a lot about the game there's like ev iv maxing pokemon were very very extreme but like if that's what you wanted, we could just hook it up to a knowledge base, like hook it up to a guide.

21:21If you know how to be Pokemon red, but the interesting piece here is actually like, can it figure out what to do without just memorizing the path? That's exactly right. Like, that's part of why, I don't know. Part of what I've realized putting this out in the world is people will draw their line of where purity is anywhere on the spectrum. Like, is it, is this cheating? Like, yeah, maybe, um, who knows? um like frankly like i don't particularly care the the main insight that i have is like when we put this out like you learn a lot about what the model is good and bad at by staring at it and that's kind of what i like about it so evaluating the model is kind of separate than your emulator and how it can use an emulator right like we can always improve those things i'm curious um as you switched from 3.5 to 3.7 and sort of reasoning models were there any degradations there like did it did it kind of get worse at anything and was the prompting somewhat consistent like a lot of what we've seen with different reasoning models is like you kind of prompt them differently right you tell them what to do let them figure it out but um yeah any any insights there yeah yeah that's a good question one thing that's nice about three words on it is it's like this hybrid reasoning model so like it kind of can do the old thing and the new thing and it's actually pretty good at just like being an out-of-the-box model and having this like thinking mode where it can spend time reasoning so i i didn't like really run into any like serious degradations the one thing i'll say is like literally every model that has come out with pokemon like the the main change that i have made to this agent is deleting prompt stuff like there's a whole bunch of like band-aid-y prompt stuff i've added in the past it's like trying to like steer it away from doing a lot of the things that it got horribly stuck doing in the past and as the models get better i found that just like making sure it's as simple as possible and giving them as much sort of like free reign to try to solve a problem as possible is useful and like the way i think about this is i'm like less confident over time that i understand exactly how a model is intelligent right like it's capable of all of these like ridiculous things it does phd level stuff in some ways and like is unable to screen see a screen as well as a four-year-old in other ways but like my confidence in like exactly what i need to tell it to do to be smart at playing pokemon is actually like really small right now if i tell it this is the way you need to solve this problem that might not actually be the best way for 3.7 side to solve this problem it's like just different than i am in terms of how it thinks about these things i found that just like kind of like pulling some of the unnecessary instructions where i tried to like use my intuitions about what would make the model better out of the prompt over time is the thing that's just like sort of consistently as models got smarter gotten more juice out of this i was watching the stream yesterday or the day before and it was a very tense battle i think they were like down to like 2 hp each and like the opposing pokemon like missed a scratch or something and it didn't die and like you could tell like i was like wow it was like very dramatic and i was talking about the game how i yeah is there any thought being put into like trying to have it more like do you prompt it to be more rational to let it know that it's not a real life that it's a game it's like it feels like it gets very distressed when they're actually the pokemon are actually gonna die it's funny they um it knows it's pokemon it's like you're playing pokemon red like it does know that and it has a sense of that but it clearly gives us some attachment i'll tell you a fun story we tell it to nickname its pokemon now it will occasionally do without it but it's like more fun if it nicknames its pokemon so that's like in the prompt is like it's fun if you nickname pokemon you should consider it and one thing we found when we started doing that is it got more protective of the pokemon it nicknamed like it's pretty obvious like when it catches a pokemon now that it has a nickname it will like go heal it right away if it's hurt and that did not ever happen before which is pretty like so there's some cute little things cute quirks about quad who really wants to protect its uh precious nicknamed pokemon which is great so i will say it's kind of normal like like when i was five playing pokemon red and you know i had two hp in amidst the scratch that meant everything that was existential i agree i agree completely how about skilled transitioning so one question that i had so you're playing pokemon red right say you want to play silver or gold next um have you thought about how models can kind of learn from these games and like store these learnings and then use them again in the future i'm sure it's not part of the project today but curious your thoughts i've thought about it only a little bit which is like i think there's some like interesting when you actually read one of the knowledge bases that it has gained like on some of the longer rollouts when they're good like there's actually some like pretty decent tidbits about how it should act and try and do things and like some of the ways it succeeded in and actually like one of the things that's most unique about 3.7 sonnet that i've seen is like it will have like meta commentary on what it's good at and bad at and it's knowledge base like i misperceived this thing and so like i need to be careful doing that again you occasionally see show up there which is um which is pretty cool so like i could imagine there being some way to like translate that knowledge base from one game to another.

26:49I think my knowledge base is frankly kind of kludgy of an implementation right now. It's more or less a Python dictionary that's appended to the prompt. And I think you could find better ways if your goal is to transfer across games and things like that to manage a knowledge base that Quad can actually use more or well in different scenarios. is um but there's definitely pieces there that like i think it would get be off on a better foot on the next pokemon game if it had that or even if like i were to restart the stream it would like have some some tidbits that it would probably like uh speed up if it like had access to things that i learned in the past that it's interesting yeah yeah i always think of that in card games you know like you have the idea of like temple in a card game and it's like you know it's the same magic as it is and you know star wars flesh and blood all these different things i feel like games is similar where like learnings you get from pokemon you can bring over to similar kind of like open world games and i think it's also like particularly interesting for some of the things that are like how quad learns how to play a game in general where it's like pressing too many buttons at once is a bad idea like i watch what's going on that kind of thing like definitely is stuff that it has learned that is like interesting in a meta way uh that it's like hard to give it that sense of self necessarily in training i think sometimes like it's hard for it to know like what it's getting bad at in some scenarios but it's interesting to think about how it can learn across things well like uh some of this also is due to a simulator right so a lot of what's learning is how do i use a simulator what am i good and bad at but the model internally should know quite a bit about pokemon right like if you've played pokemon going from Pokemon Red to Emerald to Diamond, having played the first one doesn't help you that much in the second, right?

28:38You kind of get the general concept. You get what types are good against other types. And the model knows a good bit of this, right? But it's still interesting to show. This is more so like it shows that knowledge bases kind of help with understanding how to use the emulator, right? Like it struggled and then it figured it out. So, you know, with Pokemon, it's like this thing can now learn how to use them. Yeah, which is pretty cool. That has been like part of what's been fun, seeing all of my progress on this thing. I had a bit of a follow-up question to the last one with Alessio. So if people want to blow thousands of dollars and want to, you know, improve this a little bit, is there anything else that you'd want to see done, whether that's like improve emulator, try different stuff?

29:19Is this just anything that like anyone watching this, you'd kind of hint them towards what you'd want to work on, what they'd want to work on? Yeah, no doubt. If I had to guess, like the biggest lift that exists around this is probably something around the memory, which I don't think is like hyper-optimized right now. The nice thing about the memory is like, it's always in the prompt. Like it's, it doesn't go away. Like some, sometimes if you leave it up to quad to try to like read and load and save to memory bases, like it will underutilize it or forget things. But I think there's probably something there.

29:57I will say all of the many, many hours I've spent tweaking around the edges of this thing, nothing quite does it like a new model though. Fundamentally, I think the limitations right now are some smart things. I've seen, and I mean this in the kindest way, but I've seen a lot of people in Twitch tell me about ways that they could fix the navigation capabilities with a better prompt. People would be welcome to try, but I would guess that would be a somewhat fruitless avenue. I don't think, I think it's just not very good at understanding. At the first time, I'll give you a very quick anecdote, which I think is like my favorite for like why this is particularly hard.

30:35I have this clip of Quad leaving Oak's lab and being like, great, I left Oak's lab. Now I need to go up to the north end to go to route one. And it just like hits up on the D-pad and go straight back into the lab. And it's like, shoot, I'm back in the lab. I need to leave. And it hits down. It's like, great, I'm out of the lab. Now I can go up to route one. It's straight up. It just like goes up and down 12 times. And it's like, you're not, you're not fixing that with a prompt. It just literally doesn't get it. It doesn't understand. And so it's pretty hard to make like little around the edges changes that like make a huge, huge difference.

31:09Yeah. I mean, I've always been fascinated by the fact that Twitch plays Pokemon actually beat the game. Yeah. From a, you just look at it and you're like, this cannot possibly work because you have people trying to sabotage it too in the chat. Not everybody's trying to solve it. what what so i i just like that up it took 16 days and seven hours for twitch plays pokemon to be red how how close do you think we are to a model that can beat it in less than 16 days and do you think it needs like some core like model really big jumps or like do you think it's like we're close i think i think there is model stuff at least from quad like i am confident there's model stuff that needs to happen for it to be like really capable i i have like four spots in the game stuck in my head it's like i think there's literally no hope it's going to get through that so i think there's like a gap that's mostly around like its ability to like see and navigate and remember visually like what's going on that i just don't think is like we've figured out yet so to me that's like a pretty big gap i do expect like i i think it's going to keep getting better like i have no reason to believe that this is not just like a fundamental like ability to scale learn and understand problems thing that i think is getting better as we train models to be more capable of sort of these like long horizon tasks like i actually do think this is like a pretty reasonable proxy of that and i think it will continue to get better for a little while i don't know if there are like affordances around images and videos and stuff like that that we need to figure out to make it work it's like unclear to me if that's true or not um but yeah i think we have a little ways before we can beat the game in 16 days i do not have a lot of faith that the current stream is going to be standing in Victory Road in 13 days.

32:50What's been your favorite moment from building this to thinking of the idea to just seeing it play? Any major highlight? I think the hypest I have been is when it beat Brock the first time, where I was just like, I've been doing this for eight months, and then a few weeks ago, I kick off a run, wake up the next morning, and it's like, oh, my God, oh, my God. and it was the other good thing about it is like i woke up at 8 a.m and i checked my i i have it send me updates to slack i'm this is like ridiculous things but um it's like literally like about to start the brock battle like i opened my phone it's like oh this is like happening right now and it's like a pretty hype way to start a day i think that was my uh my highlight i have a lot of like other cute things like some of the cute nicknames it's done over time and things like that are are endearing but but that was like the peak hype for me it was like we beat a gym leader like we've got a badge like quads doing it, you know, a bit of a follow-up.

33:46So I noticed that you mentioned it eventually started beating multiple gym leaders. Were these all the same run? Was it different ones? Was it? Yeah, I have like the run that you saw that's like on the graph we put out alongside like in our research blog is like a single run that I have watched like get through at least Serge's gym. And then it got a little past that. And the reason that that's where we stopped reporting is because that's like the physical amount of time that occurred between when I started it and when we launched the model. So that's like, uh, that was a very hyper, uh, hyper up to date graph on, on the best run we had.

34:23So awesome. Um, I know we're running out of time. My last question is, are we going to work on magic on cloud place magic next? Or maybe we can do like the magic arena intro. Yeah. Uh, funny story. there was a project I did right before I joined Anthropic that was like training uh an open source model to like slightly be better at picking draft or cards in a draft like I was training it on like the 17 lands data that exists to like learn how to how to pick cards out of a packs a little bit better uh and I I did talk about that in my interview to get hired at Anthropic so so I've put time into this I'm ready I am ready for that project too that I have that code sitting around as well somewhere i really get it in all my nerd nerd ml slash uh slash gaming hobbies here yeah no i'm ready i don't know if you're planning on open sourcing any of the pokemon stuff but if you want to work in open source on the magic stuff i'll be happy to collaborate awesome we talked about it i don't i don't know yet what the plan is there's like a certain amount of like this is not my day job that i have to figure out how i want to uh yeah deal with that uh we'll see yeah um awesome david any parting thoughts anything people have missed no i think like the one thing i do like to drive home when i i've been talking about this is like i really do think like this is just demonstrating like a thing that is going to make agents better with this model you know like this is a very fun way to see it but like i think the thing is that it like has some ability to like course correct update and figure things out a little bit better than models have in the past.

36:00And even if there's like stuff it's dumb at, like it tends to have an ability to like power through it in a new way. And so I think what it's exciting to me is just like, I think there will be some real world stuff that comes out of this model once people play with it. And I'm pretty excited to see like how people take the skills we put on display a little bit here or lack thereof in some cases and figure out how to turn them into actual agents that do stuff. I have a quick last question on that. Actually, is there any guidance or any way that you like quantitatively measure the evals of this system like a lot of it is vibes a lot of it is how far it gets where it gets stuck but like are there are there any lessons or any specifics about how you measure how it actually does so i've done a lot of like little small tests of like put it in this scenario and see what it does but i like frankly the best test i have is just like run it 10 times on this configuration and like see how quickly it progresses through milestones of the game.

36:53I mean, it's the best thing about games, right? Like it's why the games are such a useful thing. There's literal like benchmarks of gym badges that are moments of progress in a game, which are like ways to evaluate what happens. And so I think like how quickly it's able to make progress is actually a pretty recent or a reasonable like eval, if a slightly expensive one to calculate. It's an integration test, not a unit test. Awesome, David. Thank you for joining. Thank you for filling in on the whole side too. Yeah, my pleasure. Thanks for having me guys. I appreciate it. Awesome. Good to see you.

From the publisher

Special lightning pod with David Hershey from Anthropic, the person behind Claude Plays Pokémon. Sonnet 3.7 is currently trying to complete Pokémon Red live on Twitch thanks to a special harness that David built so that it can see the screen, navigate through it, remember facts about the game, and more. (Since recording, it has successfully escaped Mt Moon! You can follow along on Twitch: https://www.twitch.tv/claudeplayspokemon)



Get full access to Latent.Space at www.latent.space/subscribe

More from Latent Space: The AI Engineer Podcast

All 247 episodes
⚡️How Claude 3.7 Plays PokémonLatent Space: The AI Engineer Podcast · 38 min
Listen in VO