In short
```markdown
Google DeepMind
The Podcast Episode Summary Episode Title: A Whistle Stop Tour of AI Creation with Paige Bailey Host: Professor Hannah Fry Guest: Paige Bailey, AI Developer Relations Lead at Google DeepMind
In this episode, the discussion shifts from a deep dive into AI research to a practical exploration of various AI tools currently available for public use. Host Hannah Fry is guided by Paige Bailey through the latest AI advancements, including prompt generation, video creation, and audio production. They explore hands-on demonstrations of tools such as Gemini and the recently launched VO3 model, showcasing how these innovations are reshaping creativity and content creation.
Key Concepts and Discussions
AI Tool Evolution
- Transition from Theory to Practice:
- Discussion on tools that have evolved from their early iterations to become user-friendly applications for the public.
- The episode emphasizes the importance of understanding both the "what" and the "how" of AI technologies.
Featured Tools
- Gemini:
- An AI model that helps users generate prompts and craft better input for AI applications.
- Real-time demonstrations of generating video content with Gemini.
- VO3:
- The third iteration of a video generation tool.
- Capable of producing high-quality visual outputs with synchronized audio.
- Compared to earlier versions, VO3 includes features such as sound generation, improved physics grounding, and character consistency.
- Flow:
- A tool designed for filmmakers, allowing for video stitching and specialized controls to enhance the creative process.
- AI Studio:
- A platform for experimenting with AI models, featuring capabilities for text-to-speech and audio generation.
Creative Potential and Accessibility
- Expanding Creativity:
- The promise of AI tools enabling individuals across various disciplines (e.g., science, history, art) to express their ideas and projects digitally.
- Encouragement of non-experts to participate in creative processes traditionally reserved for specialists.
Ethical Considerations
- Safety Measures:
- Discussion on the introduction of safety filters within AI models to prevent misuse (e.g., generating harmful content or deepfakes).
- Implementation of watermarks for AI-generated content to distinguish it from authentic footage.
Audio Capabilities
- Steerable Audio:
- Introduction of a Gemini text-to-speech API capable of generating audio in various emotional tones and languages.
- Interactive demonstrations of creating audio responses with specific emotional cues (e.g., romantic, angry).
Integration of Multiple Modalities
- Multimodal Learning:
- Importance of integrating various data types (text, audio, images) to create a cohesive and rich output.
- Discussion on how this is fundamental to the advancement of Gemini models, enabling outputs that feel immersive and realistic.
Key Takeaways
- The episode provides a practical exploration of AI tools that enhance creativity by making advanced technologies accessible to everyone.
- There is excitement around the potential for these tools to democratize content creation, allowing more voices and ideas to emerge in the digital landscape.
- Ethical considerations remain a crucial part of the discussion, especially concerning the implications of powerful generative tools.
Conclusion The episode encapsulates the ongoing transformation within the AI landscape, showcasing tools that bridge the gap between advanced technology and everyday users. As AI tools become increasingly capable and accessible, there is a growing expectation of a significant shift in how creativity is approached and executed across various fields.
Further Listening
- Listeners are encouraged to check out previous episodes for deeper insights into AI research and future trends in technology.
- Feedback and suggestions for guests or topics can be shared through podcast platforms.
```
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:03Human creativity is about to have this explosion of progress and there's this promise of everyone being able to become a creator and not just a creator in one specific discipline, but to also be able to expand that out into many other disciplines. Welcome to Google Deep Mind, the podcast. I'm Professor Hannah Fry. Now, one of the things that we've always done with this podcast is to bring you access to the people who are working on some of the biggest breakthroughs in AI. And a lot of the time, the researchers, they're talking about techniques and technology that underpins big ideas. But now we are at a stage where more and more of the tools that we have seen the early iterations of here are now live.
0:49They are out there in the world for you to interact with. So what we wanted to do in this episode is just to pause, to look at the array of tools that have been released, talk about how they have changed since we first encountered them, and to explore the myriad ways that they can be used. And if that is our objective for today, Well, there is no one better to show us this progress than Paige Bailey, AI developer, relations engineering lead at Google DeepMind. Paige, welcome to the podcast. Thank you so much for having me. I'm so excited to talk more about what we've been building at Google DeepMind.
1:23The thing is, is that we get to see a lot of the early iterations of this stuff. Yes. And last year we had Doug Eck on the show and he was showing us, I think the very first iteration of VEO. Yes. which now with the launch of VO3 is, I mean, quite a different beast. Exactly. Like the first implementation of the VO model was still, you know, just visual only, not including all of the really enriching sound qualities that we see from the VO3 model. And you also had to give it pretty significant guidance in order to get the model to produce something that looked photorealistic or even like something that you might see in a cinematic film.
1:58But we've come a long way. I would be really curious to see what that first video looked like. Yeah, so I think we have it actually. Yeah, let's do it. If I were to describe what you should see, we're going to come in from the top. It's a tracking shot. We're going to come down from the tracking shot. We're going to have this neon hologram of a car driving at the speed of light, cinematic. And then the car leaves the tunnel back into the real world city of Hong Kong. So we should expect a kind of transition back to Hong Kong. All right, okay. So we're starting off. There's no other way to describe it than what the prompt said.
2:33You've got these buildings covered in neon lights. It's very smooth, this tracking shot. And then you speed up and zoom in closer and closer. You're in between the buildings now. And then we have this car racing through the streets. You can see the neon lights reflected in the wet pavement below. There's other cars jostling for position around. And it's almost like everything is blurred because you're just going so fast. But it's really consistent. Now it's gone through a tunnel. there are these big lights overhead and it's come out of the tunnel into an extremely realistic modern scene. It's a wow moment.
3:12I mean that is really good isn't it? It is so good. So some things that I notice now it is quite blurry and I guess that's sort of part of the vibe that it's going for but you're not seeing this pristine detail on the car. Not at all. I think if we also looked really closely we might also see that some of the the physics that's expressing the shots is not quite right. And the way that's kind of the light gets reflected on things is also not necessarily quite consistent. So let's see how well the new VO3 model does for this. You're using exactly the same prompt here. Exactly the same prompt. And here we're in the Gemini app generating this video.
3:48It can take up to two to three minutes in order to get it to have the right outputs. Well, let me ask you then, one thing that was really noticeable was the way that Doug's prompt was very poetic. It was lovely. It was gorgeous. It was a little mini movie in its own right. Yes. What are your tips for generating these prompts? So interestingly, we have a new feature with the VO3 model called prompt rewriting, which gives you the ability to give your input sentence to the API the way that we interact with model and to get a response back that makes the prompt a lot more detailed, a lot more aligned with perhaps what you're imagining.
4:23So you don't necessarily have to go through all of the mental work and all of the terminology that would be appropriate to kind of describe this thing that you just imagined. And I often use Gemini for something quite similar, is you can give Gemini the prompt that you're thinking in your objective, and then ask it to craft a prompt for a large language model or a video generation model in a way that would make the prompt much more likely to produce the optimal output that you were expecting. That's quite a good tip then, actually. If you are not that good at writing prompts, get Gemini to write the prompt better.
4:57Yeah, and so you can just invoke it naturally through the Gemini app, which is probably easiest for people to try. But we also have the prompt rewriter hyper-specified for VO3 available through the API if people are a little bit more comfortable with programming languages. So how long are the VO clips? The VO clips are the ones that we released publicly, So the ones that people can try, they're all around eight seconds in size. So very short. Why does it need to be eight seconds long? So it's eight seconds made available publicly. Internally, we have models that are capable of producing much more long form content.
5:33But we find that eight seconds is really good to kind of give you full creative control over that first clip. It's also useful in the sense that you can kind of get an idea of the style. You can start experimenting with language. And it's also wonderful in the sense that you can start like putting to life things that you might have been imagining before. I know that all of the Internet is very enchanted with memes. And now you can have memes that are much more long form that are actual videos as opposed to just a single snapshot. Absolutely. OK, it's ready. Amazing. Let's do it. Oh, my gosh. All of a sudden.
6:07This is so cool. OK. Oh, wow. Goodness me. Yeah. OK, it's like suddenly someone has switched on HD. So the neon lights have changed from these sort of pink, sort of horrifying neons to full-on billboards, which you can't totally see the detail of as you go past them, but there's structure to them. The car is now sort of painted with light almost, running through the scene. But also look at the lighting on that bonnet. Oh, my gosh. That's extraordinary. It is gorgeous. And that level of detail in all of the buildings. This really does look like a city in a futuristic Hong Kong. What I'm noticing here is there's a spotlight.
6:49A car comes out of the tunnel. There's a little lamp post above it. And the spotlight perfectly tracks where you would expect it to be along the bonnet of the car. That's amazing. Do you hear the sound? The screeching tires. Yes. Oh, sirens in the background. But the sound perfectly matches the frames then. It definitely does. And the background tracks, the sounds for these videos, they get generated with VO3. They're actually able to match not just things like the cars, but also capable of giving you background music if you wanted to have kind of cinematic music coupled with the audio or coupled with the background noises.
7:27All of this kind of stitched together into a single video. So let me understand then, what is new in VO3 that we didn't have before? What makes it better? Yeah. So VO3 is better for a few reasons. One is that it has the ability to produce sounds, which has been really enchanting a lot of people, I think. Another way that we've improved VO3 is that the video outputs are grounded in more physics understanding. So as we look at videos, we can kind of spot ways in which the light or gravity really does seem to align with the physical world. And then there's also been a lot of improvements around character consistency.
8:03So some of these things that, you know, perhaps would have been not possible a while back, you can kind of see them expressed and brought to life. Well, OK, but how are these additional features possible? DeepMind is really, really paying attention to how the data is curated to use to train the models. When you think about it, you can see the word tree. You can see a picture of a tree. You could have kind of the 3D representation of trees or any of the other things that you would expect. A sound of wind blowing through leaves. Exactly. And it's like a video of somebody panning around. All of these things are still associated with that one entity, but it's all very different modalities describing the same thing.
8:47You know, historically, folks have been concentrating on just one modality, so text or code or something similar, when as humans, we experience the entire world in very different ways, you know, everything from seeing to hearing to touching, like all of these things. And so I think the team has put a lot of time and energy and effort into being able to couple together not just the video footage, but also the sound that composes the video footage, detailed descriptions even at the frame by frame level. And then also kind of stitching all of that together into a full representation of the training.
9:23So whereas a language-only model might have the word tree and it's closely associated with the word branch or twig, the multimodal version not only has all of those embedded within it, but additionally has audio, images, video, all of these different layers. Absolutely. I think this is one of the things that makes me most excited about the Gemini models, right, is that we're really the only model family that also allows you to output text and code, but also images to edit images to have output audio as well as steerable audio. So being able to say like speak softer, speak more loud or speak in a different language.
10:01All of these other model families kind of relied on stitching together different trained experiences as opposed to baking it all into one innate model. So, OK, I know a lot of the buzz around this, a lot of the new thing about VO3 is the audio. I mean, how is it being generated in order to correspond? Is it that it's generating something that matches the visuals or is it like there's a context which produces both the visuals and the audio? I think it's because the training data has all of the different modalities associated with the thing that it sees. So it's not just seeing a video. It also has the transcript.
10:36It also has the frame by frame level description of the video and what's happening. It also has the description of the audio if there's any background tracks. And so all of that kind of brought together simultaneously is capable of generating these much more immersive and natural sounds and natural responses. Because there are certainly instances where if you listen to a song, you could read the sheet music and you could kind of hear the different tones displayed. But it could also make you feel a certain way. And so if I describe that I want to hear a song that reminds me of being in a rainy cafe in Japan at nighttime, that's something that you would really want to have all of these different modalities coming together to understand as opposed to just guessing at what a rainy cafe might sound like.
11:29It's like a two-dimensional version of a fully immersive experience. Exactly. That's nice. I love that description of it as well. We're getting closer and closer towards something that feels very close to reality. Like a simulated reality. Yes. And I don't think that was possible previously. Okay. So that thus far then is the Gemini app. Yes. But if you are a professional filmmaker or you want to take this a bit more seriously, there is another place you can go to, correct? Yes. It is called Flow. It's built by our colleagues over in the Google Labs team. And they've been partnering directly with filmmakers to really build an experience that aligns with their expectations.
12:09So is the idea here then, you still have the eight second videos, but you can stitch them together? You can stitch them together. You can kind of style them. There are even camera controls associated. Oh, wow. So it really does give you a lot more creative control as a filmmaker. And we find that, you know, just as you would have specialized environments for musicians to create their electronic tracks or CAD designers, you would probably want a really, really dedicated and focused UI for each one of these use cases that can really hyper optimize for the things that you would care about as a filmmaker.
12:42I think the point about this is that you then have this absolute open door for creativity. Yes. I've seen people take videos as though the Spartans were Instagram influencers. Yes, absolutely. Like reporting on their siege. Absolutely. And also character consistency across different experiences, no matter what the lighting might be. Like you might have a little monster character that you want to have swimming through the ocean and then you also want to have him climbing a mountain and you want to have him singing on a stage. And it's able to keep that same character consistent but to change all of the dynamics around it, which is pretty magical.
13:18I do wonder though, about putting these tools in the hands of people, There are also concerns about it too, right? I mean, deep fakes, but also scams, tricking people into thinking that news events are happening that perhaps aren't. Where do you stand on that? Yes. So we do have safety filters introduced within the VO models themselves. And relatedly, for all of the VO models that are generated through the Gemini app, there's a specialized watermark that gives you the ability to kind of know that this was AI authored as opposed to being something that was just shot via raw footage out in the world.
13:54But we also have special constraints in place around not being able to generate images of things like children or special entities. There's also a constraint in place such that government officials are people who are significantly present in the public sphere for policy or for science or for any of the kind of notable figures in the world. We can't generate video content about them. and even the models that we experiment with internally, they still have these constraints. Well, if one of the key things about VO3 is the audio, can we go into the audio a bit more? Yes. We even just released a Gemini text-to-speech API that allows you to generate audio, including steerable audio in multiple languages.
14:37So without the images to go with it, just audio only? Just audio only, but really, really expressive audio. And you can also have multiple speakers in different languages. So I believe you might have seen the Notebook LM before with the podcasts that were generated. This allows you to create customizable and similar experiences using multiple speakers or a single speaker. Well, let me bring you back for a second, because we did actually get to talk to the researchers who were working on WaveNet. Oh, yes. This is like now, I mean, only four years ago. And this is where they were at that point, because what they did is they trained a model.
15:15This was neural networks, right? Using my voice. And this is where they got to. Hi there. I'm a mathematician, author and podcaster who's fascinated by artificial intelligence. How have things changed since then? Yes. So things have changed significantly. When WaveNet, which was pioneering at the time, was first created, you would need dedicated single task models for each one of the things that you were trying to do. And so I started doing machine learning, I think, around 2009, 2010. And it was extraordinarily painful because you had to acquire all of these special purpose data sets. You had to get them cleaned up.
15:51And in general, once you got your training data, you trained your model, you would have to monitor for things like data drift. Like if anything changed over time, you would have to retrain the model from scratch. And so WaveNet was a dedicated single task model for generating these really realistic at this time sounding voices. But it couldn't do other things like steerable audio. Like you couldn't say, give me this audio clip in this kind of style and do it in German. Whereas with our recent models, they're a lot more steerable by design. So you can give instructions about style, about the language that you're speaking, about pause instructions or speak quickly, speak slower, all sorts of things.
16:32So how much of the WaveNet, I mean, code even, has actually ended up feeding into this model? Or is it sort of a, you kind of started again once large language models and transformers came on the scene? A lot of the code that was used to create the WaveNet model, the architectures are a little bit different for our Gemini family. But the data that was used is definitely repurposed. And so all of the examples created, or this is the text input, this is the audio output. You can also enrich those kinds of data sets with the descriptions of the style of the audio or the tone or the temper of the voice.
17:06All of that is incredibly useful for Gemini. Go on then, show me how it works. Give me some examples. So if we go over to AI Studio, you can see here this great playground for experimenting with and trying out the latest Gemini models as soon as they're released. And so to create audio, it's the one that looks like an audio wave on the side. So interestingly, for the text-to-speech model, you would go to generate media, and then you would go to Gemini speech generation, and you're launched into this text-to-speech UI where you can specify the different speakers, the different voices, and then also the style instructions for each one of the speakers.
17:41Let's think of a prompt then, because I'm particularly interested in this different emotions. So what about if you get it to say something like, I was waiting for you, and then we try different emotions? All right. So we're going to specify in the system instructions, speak in a friendly tone, like you're greeting a relative who just came home. Nice. And then start typing a prompt. I was waiting for you. And then hit run. I was waiting for you. Amazing. Friendly. Can we try something different though? Can you make it more romantic? Yes. So speak in a romantic, hushed tone, very breathy, and only say the words that the person puts into the prompt.
18:29I was waiting for you. Saucy. Very saucy. I also love that you can see in this UI the thoughts associated. So Gemini's thinking process as it's going through the path of creating this audio response. What has it said? It says, I've processed the input. I've pinpointed the exact phrase I need to use. The task core now centers on delivering the specific phrase. I've meticulously identified the romantic, hushed, and very breathy tone required. And the next step is generating the audio. So it really gives you like step-by-step instructions of how to incorporate all of these responses. Can I do some more?
19:03Can we do angry? Yes, you can definitely do angry. So let's change the system instructions again. And so let's say... Someone's late for the date. Yes. And so let's try. This is so much fun. Yes, it is. I was waiting for you. Oh, that. She's angry. She is very angry. Yeah. Can you do like grieving? Like a. Grieving, yes. Like a lost loved one. Yes. So let's try another stream and speak in a grieving tone. Like you have just lost a loved relative. Only say the words the person puts into the prompt. I was waiting for you. Hit run.
19:46I was waiting for you. And so I feel like that was a little bit less grieving. So you could probably experiment a little with the prompt. Now, can you make them French? Let's try it.
20:01Je t 'attendais? I don't speak French. Je t 'attendais? Yeah, I mean, sure. Yeah, excellent. So this is all in AI Studio. This is all in AI Studio. Which is available for people to play with. Yes, it is available today for free for people to experiment, to try things out. And even cooler, if you're a developer, there's a little button called Get SDK Code that whenever you click it, it gives you everything that you just did in the UI, but in code form. So if you wanted it in Python and TypeScript and any of the languages that you're using in your personal project, all you have to do to replicate what you just did in AI Studio is click Get SDK Code and you're kind of off to the races.
20:41Absolutely extraordinary. Are there other things that you can do in AI Studio that you can't do in the Gemini app? So Gemini Live is kind of Project Astra baked directly within the UI of AI Studio. And Project Astra, of course, we did a whole episode on this with Greg Wayne. This is the real-time visual understanding, the first attempt at a universal AI assistant. Absolutely. It can see anything that you see. It can talk to you in real time. It can talk to you in multiple languages, and you can also add additional tool calls to this process. So if you wanted Gemini to be able to pull out and extract up-to-date information, if you wanted it to be able to interact with any of the apps or any of the products that you use every day, something like Calendar or Docs or Sheets or Gmail, you can give it access to be able to use all of those features within the context of the Gemini Live API.
21:33So, okay, give me an example then of this working. Yeah. So let's turn on grounding with Google search and then also share screen with Gemini. So you can share your screen, you can share your webcam, and of course you can talk with the model. And let's take a look at that video that we saw with Doug Eck and see what it has to say. Hey there, Gemini. Could you tell me what you see on the screen?
22:01I see a YouTube video playing. A woman with red hair is speaking, and subtitles are displayed on the screen, saying, Hannah Fry, welcome to Google DeepMind the podcast with me. The video appears to be about AI and creativity. Amazing. So tell me a little bit about what you think the video might cover. And also, what is Google DeepMind's latest models for creative teams?
22:29And so you can see here it's expanding out its thoughts. And it's realizing that it needs to make a search. Based on the video title, AI, Your New Creative Muse, and the speaker's introduction of the Google DeepMind, the podcast, it's likely the video explores the potential of AI in creative fields. It may feature discussions on how AI can inspire or assist creative professionals. Regarding Google DeepMind's latest models for creative teams, they've introduced various tools, including Veil 3 for video generation, Imagen 4 for image creation, and Lyria 2 for music composition. They've also developed Flow, an AI filmmaking tool, and are working on world-modeling AI models with applications in gaming and entertainment.
23:08These advancements aim to empower artists and creators. So this stuff, this Gemini Live stuff, you also have it on Android phones, right? Yes. So actually, I spotted this on my phone a little while ago, and I've been using it extensively. Let's give it a go. Hang on. What can you see right now, Gemini? okay i see a woman with blonde hair sitting at a desk with a microphone in front of her she has her arms crossed and is smiling there's a computer keyboard and mouse on the desk and a lamp in the background it's amazing what else can you tell me about that has been so useful in diagnosing my sick houseplants i can't tell you oh wow like it's and it's also just been so cool to have something that feels like a real life commentator on you know as you go about your daily life.
23:53There's been something really, really fascinating too about using it as kind of a pair programmer. One thing that I commonly use this for is we have a tool at Google called Google Colab, which is a notebooks-based environment with Python code. You can run code, you can execute it, but Gemini also understands Colab features. So if we want to share this tab instead, You could even say, hey, Gemini, what do you see on the screen now? And if you see any code, please explain what it does. Oh, nuts. Yeah.
24:29I see a Google Colab notebook on the screen titled Making the Most of Your Colab Subscription. It seems to provide tips for optimizing Colab usage, focusing on faster GPUs, and more memory. There are a couple of code snippets visible. The first one appears to check if a GPU is connected and display GPU information. The second code uses the psutil library to determine the total amount of RAM available in the runtime and presents that in gigabytes. So it's like having a helpful assistant that somehow understands every single thing. That you could possibly be looking at. Absolutely. And that also is capable of answering in an empathetic way all of the questions that you might have.
Read the full transcript
25:10So this is good. I mean, I already use Google Colab with my students. Oh, amazing. So thank you for that. That's going to be very helpful. But for people who can't necessarily code themselves, there's another offering within AI Studio. Absolutely. So in addition to being able to take the Gemini APIs and embed them within tools like Cursor or Windsurf or Copilot or any of the other many coding IDEs that you might have access to if you're attempting to write code, we also have something called the new build feature that allows you to build apps directly with Gemini. The build apps with Gemini section is unique in that all of the code that gets generated within this UI is hyper-optimized for the latest SDKs, the latest kind of code that we have for interfacing with the Gemini models.
25:53So just to put that another way, even if you have never written a line of code in your life, you are going to have slick code if you just prompt Gemini here to build you an app. Absolutely. And it will be using the latest models, the latest features from the models. It will be generating really, really robust TypeScript code, and it will also be resolving any errors along the way. So if the model hits any problems, it's able to cycle back through, fix the error to get the model's kind of implementation of an app in a clear and coherent state. Self-healing code. Self-healing code. And all with you just having to describe what you would like to see.
26:29Okay, so I am someone who has built websites in the past, right? And honestly, even with the tools that are supposed to make it quicker, weeks and weeks, weeks and weeks and weeks this would take. Absolutely. And my websites were rubbish. What does this mean for developers, though? I do feel like for developers, this means that they can focus more on building and ideation and kind of the product experience, as opposed to the daily life of a developer right now is a lot of things that aren't necessarily the most exciting in the world. So you might be upgrading a code base from one version to another, or you might be adding typing to a repository, or you might need to periodically check for security vulnerabilities and make sure your code base aren't susceptible to them.
27:14And all of these things, you know, they aren't joyful. It's kind of like tidying up an apartment to make sure that your app is kind of sustainable and maintained. I think one of the promises of these models and these kinds of capabilities is that there are so many more opportunities for software developers to really build more ambitious systems. And even more importantly, it opens up the door for more people to kind of learn and get inspired by this whole process of creating and getting something out into the world. Do you think all of these tools together then, do you think that they will sort of fundamentally change the way we think about human creativity?
27:49Absolutely. Like human creativity is about to have this explosion of progress. And there's this promise of everyone being able to become a creator and not just a creator in one specific discipline, but to also be able to expand that out into many other disciplines. So I think that there's real magic in having people who are scientists, maybe in the physical sciences or chemistry or biology, or people who are historians or musicians, being able to suddenly get all of their ideas or their projects into a digital form that can be shared with others. Paige, absolutely fascinating. Thank you so much for joining me.
28:29It was awesome to be here. Thank you so much.
28:34with these new tools it feels like suddenly everything has clicked together you know until now we've spoken to researchers about every individual element of these new releases audio video language but there is something so different about having all of those elements integrated and working seamlessly together and i have been coming here for years right i have spoken to the people all along the way. But even still, seeing these tools come to life and imagining the possibilities, I still feel completely wowed. And to be honest with you, I'm now itching to get on the train home and just try out all of the things that I've always wanted to build but never had time to.
29:15You have been listening to Google Deep Mind, the podcast with me, Professor Hannah Fry. If you enjoyed this episode, then do subscribe to our YouTube channel or leave a review on your favourite podcast platform. And of course, we have plenty more episodes on a whole range of topics to come. So do check those out. See you next time.
From the publisher
This week's episode is a slight departure from our usual deep dives. Join Paige Bailey, DevRel lead, as she guides Hannah Fry though some of her favorite AI tools. Having spent years understanding the 'what' of these models, Hannah finally gets to experience the 'how’ – from generating prompts and 'vibe coding', to creating her own version of the infamous spaghetti meme.
Learn more and try the tools yourself:
- Gemini: https://gemini.google.com/
- Google Labs: https://labs.google/
- AI Studio: https://aistudio.google.com/
- Veo 3: https://deepmind.google/models/veo/
- Flow: https://labs.google/flow/
Thanks to everyone who made this possible, including but not limited to:
- Presenter: Professor Hannah Fry
- Series Producer: Dan Hardoon
- Editor: Rami Tzabar
- Commissioner & Producer: Emma Yousif
- Music composition: Eleni Shaw
- Audio engineer: Richard Courtice
- Production Manager: Dan Lazard
- Studio Manager: Nicholas Duke
- Video Director: Bernardo Resende
- Video Editor: Bilal Merhi
- Audio Engineer: Perry Rogantin
- Camera and Lighting Operator: Robert Messere
- Production Coordination: Zoey Roberts, Sarah Ellen Morton
- Visual Identity and Design: Rob Ashley
- Commissioned by Google DeepMind
Please leave us a review on Spotify or Apple Podcasts if you enjoyed this episode. We always want to hear from our audience whether that's in the form of feedback, new idea or a guest recommendation!
Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.



