Google I/O Afterparty: The Future of Human-AI Collaboration, From Veo to Mariner

3 Jun 2025 · 54 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Episode Summary: Google I/O Afterparty: The Future of Human-AI Collaboration, From Veo to Mariner

Overview In this episode of the *Training Data* podcast, hosted by Sonya Huang of Sequoia Capital, leaders from Google Labs discuss exciting new AI developments showcased at Google I/O, particularly in the realms of creative tools, productivity workflows, and human-AI collaboration. Featured guests include Thomas Iljic, Jaclyn Konzelmann, and Simon Tokumine, who explore the implications of their latest projects—Whisk, Project Mariner, and NotebookLM.

Key Topics Discussed

  • Generative AI and Creative Tools: The merging of filmmaking and gaming through AI technologies.
  • Human-Agent Interaction: How Project Mariner is changing the way users interact with their browsers.
  • Personalized Information Platforms: The evolution of NotebookLM into a comprehensive tool for personalized content engagement.

Episode Timetable

  • 00:00 - Introduction
  • 02:12 - Google's AI models and public perception
  • 04:18 - Google's history in image and video generation
  • 06:45 - Where Whisk and Flow fit into Google's strategy
  • 10:30 - Ideal tools for creative crafts
  • 13:05 - Blending of movie and game worlds through generative AI
  • 16:25 - Introduction to Project Mariner
  • 17:15 - Functionality of Project Mariner
  • 22:34 - User behaviors with Project Mariner
  • 27:07 - Contextual memory and its applications
  • 27:53 - The future of Project Mariner
  • 29:26 - Use cases and capabilities of AI agents
  • 31:09 - Impact on e-commerce interactions
  • 35:03 - Evolution of NotebookLM
  • 48:26 - Predictions for the future of AI

Detailed Insights

Google's AI Models and Public Perception

  • Discussion on the rapid shift in public sentiment towards Google’s AI capabilities—highlighting recent advancements and product launches.
  • Emphasis on Google's long-term investment in generative AI over the past three years.

Creative Tools and Generative AI Whisk and Flow

  • Whisk: Focuses on image and video generation for consumer use, allowing users to remix and create visually engaging content.
  • Flow: Targets professional filmmakers, providing advanced tools for video creation and storytelling using generative AI models.

Project Mariner

  • Overview: Serves as an intelligent browser assistant, designed to remember user contexts and manage tasks across multiple tabs and applications.
  • Functionality:
  • Operates through user-defined tasks and understands multi-tasking capabilities.
  • Offers flexibility and efficiency in managing online activities.
  • User Interaction: Users can monitor tasks in real-time or allow Mariner to operate autonomously, with options for oversight.

NotebookLM

  • Evolution: Transition from a simple audio overview tool to a comprehensive platform for personalized content creation.
  • User Engagement: Focus on long-term projects, enabling users to accumulate information and adapt content to their needs.
  • Current Developments: Launch of mobile applications and international audio reviews, significantly enhancing user experience.

Predictions and Future Directions

  • Generative AI Landscape: Anticipated growth in video and content remixing capabilities, with the potential for AI to facilitate new forms of storytelling and user engagement.
  • AI in E-commerce: Projected increase in conversion rates driven by seamless human-agent interactions, enhancing user convenience in online shopping.

Conclusion The episode sheds light on Google's innovative approaches to AI, reflecting on the intersection of creativity, commerce, and personalized user experiences. As the technology continues to evolve, the implications for creators, consumers, and businesses are vast and promising. The leaders express optimism about the future developments in generative AI and its transformative potential in various sectors.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00You know, I was talking to a founder, he gave me the analogy of, you know, you want the user to almost be like, the way that a director would direct the cast and crew of, you know, change the lighting here. Like, can you say this with a little bit more of an accent there? And like, almost like natural language, the way that a director would direct a cast and crew. What do you think is the right way to mold the, mold the plane? I still think it's show and tell everywhere. So I don't think you do everything through text. I think it's kind of actually counterintuitive to have to transcribe everything.

0:26So I think there's a lot of like showing and acting and mimicking or giving a reference just as inspiration in addition to the text But but the one thing that's starting to become more clear at least for me is kind of Video generation simulation games. They're kind of like the same thing in this new world And what that means is basically you're kind of world building you're saying this is the stage These are the assets. These are have things are supposed to look and then you shoot in it and they can reshoot and refine and pause and correct something and go back in time and regenerate. I think that's where this is heading.

1:00And you guys are going to be fairly novel. Yeah. Yeah.

1:19Welcome to training data. Fresh off of Google I .O. we're exploring some of the exciting AI updates with three leaders from Google Labs who are the leads on Google's product experiments around generative video, computer use, and notebook. Thomas Illgic of WISC and VO reveals why the future of content isn't just about generation, it's about remixable experiences with a line between movies and games blurs, and where your creations become starting points for others. Jacqueline Consulment of Mariner explains how computer use agents will fundamentally change e -commerce by removing human friction from purchasing.

1:51And Simon Takamine of Notebook LM shares why personalized AI content, designed for an audience of one represents a completely new media category. You'll discover why these teams feel like a new chapter is starting for AI at Google. Enjoy the show. It's been exciting to see Google's just cooking in AI. And I .O. last week was very exciting. And it seems like the court of public opinion has just turned on its head so quickly. And right now, everyone's just like Google's out in front in AI. Why do you think that is? Why did the public opinion change so quickly? I mean, the models start with something They have a big thing to do.

2:27Good. Good answer. Definitely the models. And I think just the number of products that we have in seeing all of this breakthrough in technology and AI come out into all of those products. But also all the net new products that we're launching and the net new experiences. It just it was a lot last week and even not just at IO, but like the week leading up to it. I think you're in a big moment the day before. Yes. Yeah, I did. I did. Yeah, it's definitely validating to see the public opinion on the models and Google's position in AI changing maybe recently. It does feel to us on the inside at least that it's kind of it's the it's a result of a lot of work though.

3:08So it feels like we've been improving to me at least for at least the last three years in this area of Gen AI and maybe what we're seeing externally is people seeing what we've been up to. It helps with number one on many of the leaderboards and it helps some of the stuff that the models can do in the state of the art. And I think it's only possible with some of the Google models. But I think it's only it just feels like the end of chapter one and the start of chapter two. Wonderful. So here's what I'd love to do today. We have three of the leaders from Google Labs in the room with us, for those in the audience.

3:43What I'd love to do is spend a little bit of time on each of the topics that you are responsible for. And then we can round up with some overall thoughts in the United States. That's good. three topics we'll cover. We'll go into Whiskflow in VO with Thomas, which is you know Google's video models and kind of creative image generation playgrounds for lack of other words. We'll go into Mariner, Google's computer use agent with Jacqueline, and then we'll close on Notebook with Simon, and everyone knows Notebook, so that needs no introduction. Awesome. I would do that. Okay, Thomas, let's start with you.

4:20Tell us about the history of how you all have been cooking and building and experimenting in the creative image, video generation space and how long have you been experimenting with these products and what have been the key milestones so far? Sure. It's been a really exciting space. I think that's a very long question. So I'll probably rant a little bit. I think the, I mean, we've had for a long time good imagery models. There was like image and there's been dally obviously externally, et cetera. But something like two, three years ago is when, at least for us in labs, when we were thinking about products, we had the control net paper for people who remember.

4:53So it's kind of like, how do you take the model and start channeling it where you want? So it's not just like a push button thing. You can start saying, I want the pose to be like this. So it's seem to be like this, that was one. And then the second thing was Laura's where, like, you can kind of show the model a range of things. And then some of you are able to kind of like, you know, pull from the image and be like, what's the range of possibilities for that particular piece? And so that like iteration, the sense that you can start controlling the outputs that felt like the right moment for us to start for exploring the creative process.

5:20When was that? Probably two and a half to three years ago. Okay, we were really happy. And so then a lot of stumbling and trying things and failing, I think we had things where we trained with a bunch of our people internally to see what they could do with confiwi type workflows. We even had a little animation thing going on where we could have an episode with artists and we published the things that not so super villain and if you want to check it out on YouTube. And then more recently, we ended up with a bunch of convictions out of that exercise. So we had things like, creation has to be iterative.

5:52So we need to build these controls next to the models. Media comes with the blueprint, which is this idea that if I generate something, you're able to pick up where I left off. And then the third one was like, it should be show and tell. So basically, this driving force was instead of just telling the model with very long prompts, I can actually show you images, say, it should do kind of like this. And we can build off of that. So this is where we started with WISC on the consumer side for imagery and flow for everything that's high end filmmaking exercise. Really cool. And do you imagine WISC and Flow will be end consumer products in the Google portfolio of a billion user scale consumer products eventually?

6:36How do you, is your playground for kind of testing model UX and, you know, how best to bring this magic to users? Yeah, I think we see it as a spectrum. So I think whiskey is kind of our play in the, you know, really consumer space and thinking about like Everybody now has this visual language at their fingertips. They might not have it necessarily like The most advanced ideas in terms of like storytelling But they can quickly remix each other's things and so we're trying to see what those dynamics look like So I think that's kind of our exploration space with risk. Yeah. We'll see how it picks up.

7:07I think a lot of the lessons will probably also graduate in just how we deal with user inputs and treat those across multiple surfaces. And then flow is the other side of like, you have a vision. You know what you want. And it's kind of like how do we give you all the tools to create the best version of this in video. Yeah. OK. Super cool. Who's the ideal user, do you think, for Flow and Lisc? For Flow, I think it's pretty clear for us. We're starting with AI filmmakers. And the reason is we want to build this kind of, we call it the generative AI camera. Like, you know, you're doing world building, then you're shooting inside this world.

7:44How do we actually develop the DSLR camera of generative AI video? Yeah. And then we'll distill kind of the Android version of the pixel camera very out of it. Whiskey's much more consumer. There's a wide range of audiences. You know, is it you creating something funny with your friends in a chat, is it kind of more, you know, inside the company you're trying to create some visuals for slides. It's kind of like all this whole range that we're exploring will see where it lands. Yeah, so cool. Okay, you said AI filmmakers. Is that a thing? Are people calling themselves AI filmmakers now? And does it tend to be existing filmmakers that are, you know, looking to be more AI savvy?

8:22Are you seeing, you know, net new creators come in and try to create feature films? I think it's certainly an ill -defined term, but the reason why I like to say AI filmmakers versus filmmakers is I think If you take the extreme end of the spectrum, these are people who need very bespoke tools They have like entire workflows and processes and you need to develop very specific ideas There's one tier under that which maybe I classify as AI filmmaker what potentially is you know Previsualizations where you're trying to quickly get like a version out and I'd be then you do the full process or people who just don't have the budget.

8:53So they're like, I don't have a hundred thousand dollars to put my idea out there, but now I can actually take a shot at it. And so those people are interesting to us because you can really start from the ground up thinking of like if you had this generative AI camera, what would the user feel, look like? Like how would you fit those pieces? Your answer to my initial question, the models are the reason that the Corridor Public Opinion has flipped so quickly. It's been amazing to see VO's progress and VO3. And for me, I don't know what EVA was you all look at performance, but for me, it's the Will Smith spaghetti eating test.

9:29And we seem to have fast that. So are we at video AGI, or how do you think about the quality and the performance and what's ahead? There's still some room that is pretty cool. The GDM has done really great with VO3. I think the drug plus week was that it beat VO2 in the ranking, so it's kind of VOV. The videos of people were very happy about this. I think it's, you know, adherence is going up. Yes, we don't have the six finger problem. Physics are getting pretty good. There's still things where like, you know, if you want to have, for example, multiple characters and kind of choreographed of characters have like full consistency across multiple scenes.

10:03Like that's where there's still a lot to come. How do you refine your output? Can you propagate changes across clips? Yeah. There's going to be still a lot of like improvements, but in general, yeah, a huge step up and the biggest reveal this time was audio. So be able to picture, could generate audio with the video that brings kind of like, you know, an image is what like a video is more than an image and a video is sound is way more than a regular video. That's when it has opened up like a lot of virality. Do you think the You know the R &D left to do to make the ideal tool for the craft? How much do you think is in the product and in the UI and how much do you think it's gonna need to happen the model research layer and things like sterability.

10:43I think it's both but at least, and I'm sure people will have a wide range of opinions, but it's almost like we're at a state where everything we imagine in terms of controls, I think we have visibility in how it can be built. You know, you want to have consistency of characters, of scenes of location, there's like different ideas around this, you want to reshoot. So that part, I think the part that's hard is still the abstraction of all of it. So how do you put this into what are the inputs that you want from users? In the context of for example, why do I define the voice, how do I touch the voice to the character, how do you find the mannerism, how do I propagate.

11:15So I think there's going to be a lot of working that's abstraction layer on top of the models and on top of the controls. Oh, so interesting. So you think most of the model kind of R &D is almost a solve problem, it's maybe too strong of a word. Not solve, but I think we know how to do it. It will happen. I think it's pretty clear that it's moving very fast and you know we see a lot of things just like week after week coming up. But how we do the connective tissue on top, I think is still pretty much open. Audio is one of those new frontiers, for example, of like, should I be talking and driving the audio, then changing my voice, should I be typing the text, how do I do diorization?

11:50There's a lot of what are the inputs, how do you let people mold clay with all these models? What's your guess for how that future is, for how people will mold clay? And I was talking to a founder, he gave me the analogy of, You know, you want the user to almost be like, the way that a director would direct the cast and crew. You know, change the lighting here. Like, can you say this with a little bit more of an accent there? And like, almost like natural language, the way that a director would direct a cast and crew. What do you think is the right way to mold the blame? I still think it's show and tell everywhere.

12:21So I don't think you do everything through text. I think it's kind of actually counterintuitive to have to transcribe everything. So I think there's a lot of like showing and acting and mimicking or giving a reference and it's just as inspiration in addition to the text. But the one thing that's starting to become more clear, at least for me, is kind of how would I say it. The video generation, simulation games, they're kind of like the same thing in this new world. And what that means is basically your world building, you're saying this is the stage, these are the assets, these are have things are supposed to look, and then you shoot in it, and it can reshoot and refine and pause and correct something and go back in time and regenerate.

13:01I think that's where this is heading. You guys are going to be fairly novel. You mentioned games. I wanted to ask about this. It feels to me like the existing way that we consume games versus movies is because there's such a tremendous fix upfront cost of producing a movie. If you imagine that in a world where every movie frame is generated, not pre -rendered, And that entire story arcs can unfold. It does feel like the movie and the game worlds start to merge. How do you think that plays out? I think with, I mean, so for example, we have the genie model that's been really interesting. So you give an image and you can kind of move your character and the world builds in front of the ice.

13:44But what's going to be really interesting is how you ground it. Like games are fun because there's like very set constraints. Movies are good because there's like very small details that matter, you know, the expression and the moment and the timing. And so I think it's almost about the constraining of the capabilities towards what we need. So I don't know. I think the other thing that strikes me and I think a couple of people on the team is like it's not clear that we think in terms of the static formats that we have today, like an image, a video and a game, is there something in between almost and what does that mean?

14:17Where is that going to be distributed and interacted with? Like I can share an image with you but you can instantly turn into a scene that you're walking into. So am I sharing an image or am I sharing an experience? Lots of questions I guess. I do feel like the story is almost the common thing that makes a game and the movie good. That's different from an image is just a visual, right? Yeah, it's actually the setting, the constraints, the... You define the rules of the game basically and then you let other people enjoy themselves. Really cool. My understanding is that video is still expensive and somewhat slow to generate.

14:54Is your sense that that's getting solved quickly and like will we have everybody's going to be able to generate you know two hour films and their you know in their pocket. In a couple years time or is your sense that this is a you know longer. We got a lot of efficiencies that we need to build in order to make this kind of cost practical. I think I mean we've seen in imagery and we've seen in video kind of like the same speed of cost reductions that we've seen in other places. both the hardware is getting better. I think the efficiency to your point, we have the regular models, and then we learn how to distill them so that they just take less processing to get to whatever you asked for.

15:33I'm actually pretty optimistic that the cost are just going to keep coming down, and the speed is going to increase. Kind of aligned with what we're seeing with other models. Yeah, got it. Fantastic. What do you think is ahead for AI in the creative space, at least from a Google labs perspective? Well, we just launch flows. We have a lot of things to do to just like deliver on that promise of like keeping you iterating. I think that's the first thing refinement of like outputs and like keeping there going there and like insertion editing reshooting. I think it's really interesting to us But I think the holy grail will be some of these new formats and experiences like what does it mean as a creator to share something with you that you can interact with That's something that we want to explore really cool.

16:15I want to be able to talk to Will Smith as he's eating the spaghetti

16:20Really cool. Thank you so much for sharing what you all are doing over in the creative sphere. Of course. Okay. Jacqueline. Yeah. I would love to talk about computer use and mariners. Maybe first off, why is it called mariners? Great question. So we wanted to give the project a name that really embodied what we were trying to do with this space, which was enable users to just go out and explore, enable agents to go out and explore. and Mariners sort of this whimsical open -ended name that just sort of embodies the spirit that we have on the team right now. I love that. You all actually have really good product names across Google.

16:56These are all really whimsical fun. I'm so trying to get rid of the LM bit. Apart from that. You're the challenge. I'm pretty happy with Whisken Flu. I think you're going to be sitting here. We're evolving our approach to naming. That's what we evolved at IO naming. That's the way we go. That's the statement. Yeah, that's funny. Can you say a little bit about how Mariner works? Like is it computer vision model behind the scenes? Like it just feels like pure magic in a box, but give us a peek under the hood. I will take pure magic in a box any day. So the way it works is really leveraging the power of Gemini.

17:33That's kind of, you know, it's an action -tune model on a recent version of Gemini. But what that means is that we have all of the multimodal capabilities that Gemini gives us. So it's able to plan in reason when a user enters in a task. We're able to understand that. We're able to come up with a plan on how we should actually fulfill that task. And then the way it actually works is taking that and understanding the screenshots. So this is where the multimodality of the Gemini model really comes in handy. We're able to continue to take screenshots, continue down the trajectory of what it is that we're trying to achieve from the user's task that they gave us and bring it all together that way.

18:09Yeah, got it. Super interesting. What's the history of the project and when do you anticipate you'll be rolling it out on this? So the project initially started last year, shortly after this time. Actually, if we go back at IO last year, we kind of graduated the Google AI Studio and Gemini API out of the labs team onto the developer team now. And that free -dissup to start exploring what we thought was coming next. and that happened to be agents that could actually take action on behalf of users, not just answer questions or generate content. So the team started working on it. At that point, we started grouping up with a bunch of different folks across Google to kind of bring together what we launched in December last year, which was Project Mariner as a Chrome extension that took action on your browser.

19:00And then we continued to iterate on it based off of a lot of the feedback that we got from trusted testers of that initial launch. So I actually had a large group of trusted testers that we would be talking with regularly and understanding what was working well for them, what wasn't. And we took that feedback and iterated on the most recent launch of Project Mariner, which we announced last week at Google IO. Really cool. What was some of the feedback? And like, what are the people, what are the magic sparks when people really are like, this is a game -changing product for me? Yeah, great question.

19:30So it's funny. One of the initial kind of magic moments that everybody had was watching Project Mariner take control of the mouse on the browser and being able to click scroll typing text into text boxes actually felt net different when you realize it was an agent doing it. But quickly as you were using the initial version the feedback became this is super cool. Can I please use my browser again? Like I'd also like to be able to do work. Which makes a lot of sense. So that was one of the big motivations behind moving towards this idea of users entering a task in the web app that could then Run in the background on virtual machines.

20:12Okay, exactly But one of the key things that we did also try to keep true to the initial vision was how can we Start to think about bridging the context that a user had on what they were doing in their current Environment to the task that they were sending to the VM and Mariner executing in the background And the way we tried to do that was if you install the companion extension now, it'll actually be able to see all the tabs you have open. So when you're giving Project Mariner a task, let's say you happen to be looking at a recipe on a recipe site, you're like, oh, wouldn't it be great if I could canonical use case add all these ingredients to my Instacart cart.

20:48Now when you go to Project Mariner, you could say, hey, add all the ingredients from this chicken recipe to Instacart and you can select the tab that you have open with that chicken recipe. and Marino will understand that context will be able to revisit that site on the VM and complete the task with the context that you had in your local browser as well. And it's almost superhuman in a way because as a human I only, it's like hard to context switch between browser tabs. Yes, and you're able to kind of see everything in the tabs all at once. Yeah, I think a big a big net win also was the ability for Project Mariner to do 10 tasks at once, not just one.

21:24And that was really a big net unlock. I was using it the other day and I'd just come back from running an errand and there was a bunch of stuff on my mind that needed to get done. And the first thing I did was open up Project Mariner, enter in three different tasks for it and then just send them off to start making progress. And I was able to jump back into the document that I happened to be working on. And it was this like magic moment of just, okay, not only is progress being made on these things, but I just got it off my mind. Yeah. I didn't have to keep thinking about it. Do people want to see the computer mouse moving around first for a while before they're like, OK, I trust that thing to go off and do things for me?

22:01If they do, they have that opportunity in the current project mirrorer experience. You can go into full screen mode. You can see the agent moving around and clicking on things and entering text. You could also pause the task at any point and be able to take over it. So giving the user the ability to take over and or provide oversight on these tasks is something that we think is still very important when we have an open -ended platform like this or an open -ended experiment like this that really lets it up to or leaves it up to the user to try out different things. And what's the user behavior you're seeing?

22:36Like are they like, please just take the wheel. I don't want to deal with it. Are they actually want to, you know, backseat drive and watch the agent and make sure it's doing what it's supposed to be doing? That's a great question. I think it initially watching it is this fun element, but also it develops a comfort for knowing how the agent is thinking and what it's doing. But one of the pieces of the feedback we also got from the initial launch was at the end of a task being complete. We just saved the entire conversation history and it can get quite long. And what users ended up wanting was just a summary of like what did Project Mariner do to complete this task so I can make sure it did it correctly.

Read the full transcript

23:14And that really kind of points to the question you're getting at, which is, I wanna just hand the task off to this agent, but then I wanna be able to just verify what it did at the end of the task, not sit there the entire time and watch it. Yeah, yeah, so interesting. What do you think of as solves and the unsolved technical problem so far with computer use? Because computer use still feels like to me, or maybe in the Will Smith, you know, the spaghetti is still sort of disappearing a little bit, phase, and maybe that's non -fair characterization. But I'm curious where you think we are on the e -vails and the performance so far for computer use and whether the unsolved problems right now I think that's actually a totally valid comparison There's a reason we launched this as a research prototype with the experiment level on it right now I think we've seen really big gains from December to what we launched last week That said there's definitely still model quality improvements to go I think there's also just application level improvements to go there's more seamlessly being able to have the user provide context upfront, which will make the agent more capable of understanding what it is it should be doing.

24:17And then there's just more planning and reasoning that we could do, like at inference time, or at the application layer time, that sort of in addition to the model improvements, you know, improved system instructions, improved checks and calls to different models. And then of course, right now Project Mariner entirely completes a task by actuating or taking action on a browser. You want an agent that has more skills than that. You want an agent that knows when to call the right tools that has memory that's able to you know, take advantage of a lot of the other stuff that we already see out there.

24:51So I think it's just integrating a lot of that in and starting to innovate and climb on that. And then of course, right now Project Mariner, it's in the browser. people use computers. So you know we call this computer use. So there's that entire dimension as well that I think we're going to continue to see innovations in. Really cool. Were there any contrarian opinions you all took in building mariners? So for example I think some people have said screenshots it's going to be too slow. It's not going to be fast enough you should use the website DOM or whether like any contrarian bets you guys made.

25:26So the reason we went with the screenshot is we wanted to make sure that it was a skill that we could develop that could be applied across things that aren't just websites. I think the other aspect of that is like DOM versus accessibility settings or accessibility pieces and other leverage. We're kind of betting on this one right now, but I would say everything's evolving. So we're just willing to take pivots if in one it makes sense. Yeah. Yeah. Makes sense. What is it capable of doing it today and what is the speed? Like if I tell it to go, you know, the canonical go order me a pizza from Domino's.

26:06Can I do that and how long does it take? The speed is definitely an area that we want to keep held climbing on is what I would say. But it's interesting you say that because one of the things that I, so I was recently using Mariner to help me complete a task which was come up with, let me take a step back. I have a three year old at home. She is gonna be four soon. Part of that means organizing a birthday party for her and being able to figure out loop bags for kids at a four year old's birthday party. This task, because you can imagine involves understanding what to put in the loop bag and then actually buying all of those things or like finding links somewhere to go buy them.

26:43And I gave Project Mariner this task and it was basically a personal research that turned into an action taking task, which is finding me the links and save them. And the thing that really resonated the most with me on that one is as it was performing this task first did a search for good ideas to go in a loop bag. And then as it just remembered those five items, that's something any of us could do like that itself was an impressive. But the first one was I think temporary tattoos. So then it started looking for temporary tattoos. It found a great link for it. Instead of having to copy that link and paste it in a doc somewhere else, it could just remember it.

27:16It could remember this massive URL. And then it moved on to the next one. And at the end of these five items, it just gave me all five URLs that had been able to inherently store. So when we talk about speed and efficiency, I think there's two dimensions. One is just the model calls on the, you know, taking action and like, how do we improve it with different tool use? But then the other one is, how can agents just do things in a different way that are inherently faster than the way we would do things? And I think we're going to continue to see improvements on both dimensions. Yeah, I wish I could remember five URLs.

27:47Oh gosh. Okay, good point. Let's see. What do you think is ahead for Mariner? Where do you see the evolving from here? I think there's a couple things. Number one, we had a bunch of announcements last week around project Mariner like capabilities, making their way into different Google products. And I think that this is a kind of core capability that you'll start to see emerge everywhere from the Gemini app to AI mode in search. So I definitely see a lot more coming to Google products with the stuff that we're doing right now in Project Mariner and kind of paving that path forward. And then I think for Project Mariner itself, I actually like to think of things in three categories.

28:28There's the agent itself. I think that's going to get smarter, that's going to get better, that's a better model, that's tool use, that's memory, that's context. Then there's the environment. We talked about how in December it operated on your local desktop and your Chrome browser, so that's in the foreground. Then we moved towards this idea of Project Mariner operating in virtual machines, which meant that is now operating on VMs. I think there's this middle layer, which is an agent that can still operate on your device, but in the background, and there's a bunch of reasons and types of tasks where that becomes a really important kind of way for the agent to operate.

29:00And then of course, there's all the other devices. But really what you want is a capable agent that's able to operate in a way that is omnipresent across all your devices locally on VMs. And then the last one is the ecosystem part, which is where you start to get into the agent to agent interaction and how does your agent interact with all of the things that exist outside of its own world, essentially. Yeah, so cool. I think the canonical examples for computer use are book me a flight or order me a pizza. Is that your sense of what computer use agents will actually be really good for? Or whether you think, I'm sure you've seen a lot of time thinking about what applications will actually be the bulls eye here.

29:42How do you think that shapes out? So I think we default to those because they're just easy to understand. The travel planner, I mean, literally it's a travel agent. It couldn't be more analogous when you think of agents right now. But no, the way I like to think about it is on a spectrum where you have tasks that are sort of in what I would consider do it with me, where you have your agent alongside and you can easily offload certain tasks to it, but it's really working in in unison with you. And then you have these like do it for me tasks, which is hey, I just want to give my agent a bunch of stuff to go do and it will run it in the background.

30:18I think part of the reason we see these tasks being used is twofold one. They're just incredibly easy to understand and everybody kind of gets what that use case is and they're usually starting from scratch. Like there's no context you need upfront. You can just send on an agent out to go do it and the demo as a result is pretty easy. to put together. Yeah. And then the other one is just where the capabilities are at today. And so as agents get more capable and you start to have more of these realizations on what they are actually able to do, you'll see much more advanced use cases or much more complex use cases.

30:55And that also requires the user having more trust that they can give to the agent. So I think that that will evolve over time and we'll see people come up with even more interesting use cases that they're willing to give an agent good to do on their behalf. Yeah, it's all like, It's also going to require, I guess, it's going to inspire, I think, a shift in business model, right? Basically, if you have a bunch of agents going off and browsing, you know, trip planning, for example, they're not necessarily looking at the ads and, you know, the first things that show up. And so it is, I think it's going to create some business model evolution as well.

31:31I agree. I think there's a lot of evolution that's going to happen across business models, across how websites work across how users will always want to use the internet going forward. Like there's a lot of joy. I think we all get in it from content creation to consumption. But there's also a lot of other tasks that it's just right for disruption in a lot of ways. Yeah, yeah. Like, I'm thinking humans are suboptimal in some ways. We see that we get excited, distracted, and I go and buy the dress. And my agent, maybe I can instruct it to ignore the ads. Maybe it actually knows it's going to find the best content regardless of what's showing up on the page.

32:12So it's kind of interesting to think about how that future plays out for agents to do more of our browsing. It's super interesting. I will say that the dress that maybe you got distracted, I always get distracted by things to and end up purchasing stuff that gets done my way. But I'm most happy with it by the time I do end up purchasing it. So I think that there's new opportunities to think about how do you actually involve agents in this new business model ecosystem and hence that third bucket of like, there's gonna be a lot of evolution happening in that space. I think that that's where we need to evolve as an entire ecosystem and it's not just like one player that's gonna say this is how it's done.

32:51So it's been interesting just talking to different companies and different people who are also thinking in that space right now. Yeah, really cool. I mean, I do think also just as a user as well. So, you know, I often don't buy things on the internet because it's such a pain. Oh, I've definitely dropped off of it. I can't navigate this thing. I don't understand it. That happens quite a lot. Or it's just like, I've just not got time. That happens as well. Or I can't be bothered. You know? Maybe it's just me. But I'm not a fan of shopping. Let's put it that way in the real world and online. But I'm a fan in what I get.

33:26You know, I'm a fan in the outcome. and so I don't know, I kind of feel like I might do more, I would probably do more online shopping, I think. If I didn't have that barrier of actually having to do the shopping bit, I don't know, that would be me though. No, I agree with you. What's really interesting is, I don't know about you. There are certain stores that I'll go on to and I'll just like accumulate stuff in my cart and I won't wanna, like pull the trigger until a little bit later on, I don't know how to chance to think about it. Yeah, yeah, yeah. But then I end up with a bunch of half -built carts across a bunch of different websites.

33:59And part of me also wonders, is there a world where my agent is at Universal cart essentially? Where I'm at all this stuff to create this aggregate area of all the items that I might be interested in buying. And it can be across any site at this point, because the agent represents me. And it can remember which sites to go on. And then when I'm ready, it's sort of like, okay, one click, make this entire purchase basically and it can go and check out on all of the different sites or all the different stores. So that'll be an interesting area to think about. Yeah. Okay. When I just heard from you guys, is e -commerce conversions about the skyrocketing?

34:34I mean, on my computer it'll go up. That's what I'm saying. I don't know about anyone else. It's the diversity as well. Like I go to the same old sites. But I would love suggestions. Yeah. Yeah. So yeah, it's like once you kind of democratize computer use, than the laziness of humans to get through, check out, there's no longer the determining factor of which e -commerce companies will do well. It's just like the best product wins. It's, yeah, so interesting. Okay, cool. Thank you for sharing. You're welcome. Okay, Simon, you're last. Hi. Notebook, notebook LM or notebook? We'll go with notebook LM.

35:09We're still in that book. I think it's been so long now that it's definitely notebook LM. There was a period where we were like, okay, it's now the time. Yeah. But I think we've gone through that multiple hockey stick moments, which you can talk about. Yeah, it's going to be hard to remove it. I like it though. I mean, maybe every product that this kind of like has an acronym or some weird letters after it and there are a couple of them in the AI space regrets that. But at the same time, they become part of the team and the identity. Totally. Yeah, it's nice. I like it. I love that. Okay, so notebook LM was one of the biggest, one of Google's biggest viral hits last year.

35:47Last year? Yeah, it went viral last year. Yeah. But you know, the team had been building it for a while before it took off. Totally. Yeah. Tell me about how it's evolved in the last year. Yeah, yeah. Well, so firstly, the viral moment, you know, so my way into notebook at LEM was through audio overviews. So me and the team had, we were also exploring the future of content of my different angle. I think. Well, notebook was the perfect balance of user control, but also the power of the technology. Our hypothesis was that there was an opportunity for personal content. So not content that it is for everybody, actually a content that's from an audience of one.

36:35maybe two, maybe three, small group maximum. And that was kind of how we shaped the product. We didn't think it was gonna, looking at the notebook user base back then, we thought that it was a great place to kind of test PMF, just kind of iterate on the product. We were totally unprepared for the massive success of audio overviews, and then through that notebook, I'll end as well. So it was honestly the first couple of months was really just kind of hanging on for dear life. Firstly, it was making sure that the TPUs don't fully melt. So as a riser, I think, had a gift out back then. But there was also just a lot of iterations and fixing things and improving things.

37:20And that was really the first couple of months. I think since the start of this year, maybe we've managed to take stock. So at the end of last year, we launched the Join mode, the ability to join in a podcast and what you owe you, I should say, and talk with the host and ask questions and all this kind of stuff. But at the start of the year, we kind of took stock. And we've really been thinking about, what is a notebook for the notebook users? How are your notebook users really leaning into notebooks once they've come in the front door through audio reviews? And we've started to think about, and Jack, you kind of touched on this, I think the criticality of context in really enabling these AI systems to be genuinely useful for you.

38:00And we've found that a lot of users, when they're using Notebook, they use them for these kind of more longer running, almost like projects that they have. So either they're hobbies or if they're in the world of work, these, they can be ongoing projects or they can be projects with a goal. You know, like I've got to prepare for a presentation or something like that. And so a lot of what we've been really doing is retooling, you know, how we look at Notebook and also, you know, building a strategy as well that leans more into, And I think there's more sort of longer running opportunities that we see in the internet book user data.

38:37Of course, we've done a whole bunch of improvements too, so we've just launched the mobile applications. Finally, so they came out last Monday. And we also launched international audio reviews as well, which was kind of the end of a long road, honestly, of upgrading the underlying AI infrastructure and models away from the very first almost like research grade model that we used for the initial launch to, you know, native Gemini audio. So what you hear now in the international audio overviews that are very least native Gemini audio. And that was a big push for many teams across labs and also GDM.

39:16Yeah, super cool. It feels like audio overview was almost the viral hook. And you guys have been building out a lot in almost like the rag UI. Yeah. And just imagining what that workspace looks like. Yeah. What do you think the actual audio overview podcast thing becomes? And actually, I'm curious how you even ended up on the shape of two podcast hosts talking to each other. It's just like, it's such an engaging format. I'm curious how you even landed on that. And I feel like it's only in its infancy still in terms of I would love podcasts every morning to type me up for my day and things like that.

39:54And so how much of your time is thinking about notebook, the kind of rag workspace environment for lack of other word versus notebook, the podcast killer, the training data is going to be built on notebook in the future. Yeah, yeah, yeah, well, I hope not. But maybe it can help. So the way that we started increasingly to look at notebooks is that comprised of kind of three, they give you sort of like three superpowers. So one of them is they help you really accumulate information over time. And that's a lot of amazing underlying database technologies that we apply that I think lean on first party Google technologies in a pretty unique way.

40:37The second is they bundle in intelligence and when we launched last year, we used the old Gemini 1 .5 Pro model back at that time, but obviously now we've got thinking models and so on. But the third thing is the ability to, for content and information to be adaptive to your situation. And so, you know, podcasts or audio overviews, a conversation, it's one form that information might take, but you can imagine many other forms that that information or knowledge might take as well. So you might imagine it coming at you in the form of a comic book or maybe a short movie or maybe a mind map which you've also launched.

41:20But you can imagine many other types of media that fit the right circumstance and form and function for the moment, for you to understand information, to be able to analyze it, make decisions with it, kind of do except that I think we have when we're thinking about the different, might hear me, I'll talk about transforming information from one state to another. I think that's a fine word. It's a little bit technical to be honest. It's more like adapting to you and fitting you, I think that's really what we're going for. But in terms of just going back to your actual question around audio overviews and, you know, where it's going, there is a huge amount, I think, of room left in that technology.

42:08So, you know, I enjoy audio reviews and I use them a fair amount, but I also, you know, every now and then I'll be like, that's weird. You know, why do they say that? Well, they've kind of lost the plot there. I didn't quite get the right narrative. It's sometimes it, like the uncanny valley or the illusion is broken, you know, when you're listening to them. And, you know, well, it might seem like there's a small amount of work we might need to do to, to fix that last step. There's actually a ton of work that we've got to do. And so there's a lot of effort being placed into all of the various components that you'll need to make the experience feel like something where you suspend your disbelief more completely.

42:50And alongside that, there are many other different show types. We've kind of had one show type for a bit too long, I think, actually. And we're bringing more out. So we're actually working on some really cool things a lot of them inspired by users honestly So one of the things that we saw users do right back at the start But you can keep on seeing it is Users putting in that LinkedIn You know, they're putting a little in or why well number one It's kind of fun to like hear people talk about you, but a lot of users are using it to get feedback You know like to kind of I to understand from from another person's perspective who they may not have access to you If feedback is truly a gift, like real feedback is hard to find.

43:32So, how would somebody else look at me? How would somebody else talk about my strengths? And how might somebody else talk about areas to improve? This is something that we see users already using audio overviews to kind of access that sort of content or sort of information. We think we can make that easier for people. So a lot of what we're thinking about now are different show types that are in into some of the more viral successes that we've seen are users, you know, explore online. And also think about, you know, brand new formats as well. I think it's going to be fun. Okay, so we're going to have training to the comic strip.

44:10I'm not saying we're definitely going to have it, but I mean, it's, I think, I think not everything has a story, you know, and so applying different adaptations will almost be sort of context dependent, I think. But oftentimes it does help. You know, so one of the things that we were looking at the other day was we were looking at a 150 page PhD dissertation on it was invasive invasive wolves I think in some part of Europe and Yeah, you could have looked at a my map. You could have looked you could have maybe listened to an audio overview If you had like 10 15 minutes of spare, but actually getting a kind of a comic book rendition of of that PhD was really helpful just to kind of understand the overall narrative within it.

45:01So, you know, we're still working on things like that, but I think there's a lot of opportunity there, and of course, you know, comic books are very similar to storyboards and that intersection work. I was just thinking exactly that, yeah. The Thomas is doing too as well. So, yeah, there's a lot of interesting ways that I think labs projects intersect and we'll continue exploring. So you can create the heroes journey comic book of somebody's LinkedIn career arc. I mean, for an audience of one person and one person only, that's probably going to be the most awesome movie ever seen. So maybe, maybe.

45:35That's awesome. Really cool. Where do you see the notebook going from here? Yeah, like I said, where we're really, I think our focus is aside from a whole bunch of different adaptations. We're really thinking about how we can be more useful to our users over their more longer running projects. And so both in the world of the knowledge worker, but also in the world of students these are kind of like core users I think. The project is really the an area where those users both need the most assistance, but it's also the point of highest value I think for them. So if you're in the world of work, the project is where value accumulates.

46:16It's not like you need to work. Yeah, right. It's a real unit of work. It's a, we actually call it units of knowledge, but it's a great way of putting it. And the same for a student as well. The project, if it's a project with a goal, passing a test, that's a big deal. Or if it's an ongoing lifelong learning thing, that's also really important as well. So I think really focusing on use cases in those domains is something we're thinking a lot about. I'll say the other thing is, I think one of the things I'm personally very excited about. I've been in the consumer product space for many, many years.

46:50And I guess one of the things that we did at Google when we went mobile first in the mobile first era is we moved a lot of our desktop products to mobile. And if you look at those mobile products, many of them are the desktop products shrunk down to a small screen. And that's okay. And I think because we were one of the first, because we built Android and a lot of our big products basically got mobilified at that point. We found it hard to change at that point forwards but I've always been really interested in thinking about if you have a desktop experience, what is a companion mobile experience that doesn't have to just be a carbon coffee of the desktop experience that maybe leverages the form factor, the sensors, you know, the actual, the fact that it's on with you all at all times, to deliver an additive experience on top of the desktop experience.

47:42So, you know, we've just launched the mobile experience after a fair amount of time in development, it's fair to say. But what I'm really most excited about there is the opportunity to actually iterate on that kind of novel mobile experience going forward. Like, for example, wouldn't it be cool if I was, you know, maybe I'm in a discussion with some amazing, really smart people and I've popped my notebook down, I've opened a voice, it's native voice recorder and it's just able to record the conversation for me. And then I can transform that to later days and accumulate them and all this kind of stuff.

48:15That's the thing that is probably gonna be weird if I open my laptop and push record on my laptop. But the mobile device is the perfect opportunity. –Cobely. –Yeah. –Really cool. Thank you for sharing. –Yeah, nice. –Okay, we're gonna close it out with some predictions on AI as a whole. Please jump in. Hot Takes, welcome. Let's see. Let's start with what are your favorite Google labs projects that we didn't talk about today like what what is under What is the what are the gems right now that you're most excited about the unreleased all parts that we know last talk about Not then released You guys just announced like 50 things they have to be others beyond the three we talked about Yeah, yeah, I have one which is kind of still in this like video and image space, but I think the virtual try -on on stuff that we presented, like this, a lot of preparation.

49:02And I think that went to me really nice because I think it meets a real direct user need. It's the strength of Google. Obviously, we know we have all the inventory and how to connect this. It's just so fun to just see things on your site. So I'm very excited about that one. I think this says like a good thing. That's my favorite as well. That's so funny. OK. I think that one's a good one. Stitch, I think, is really cool to be able to just talk to the product and describe what design you want and have it actually come out with that front end design. I'd been using it a little in dog food before it was launched and so it's just I want to spend more time Using it now that it's it's actually live really cool without you Well, mine's gonna be stitch I guess what areas do you think will be hottest in the application space for a broadly in 2025 like I think coding was maybe the breakout application in the last 12 months.

49:59What do you think will be the breakout application in the next 12 months? Video. I think there's something around these remixable content. You generate something, I take your thing, I just riff off of it. There's something around this that I think is going to pop up somewhere. I hope it's us. But that part feels really interesting. It's going to feel like if you're whisk is heading a bit that way, view obviously can power a lot of this in video. I think that's going to be something this year. As you look back at you know past predictions of what you thought was gonna be interesting And I what wherever you got been really right and where we've been really wrong Just we know we're let's say where we're really wrong altogether.

50:35All right three two one timing

50:41I think that there's been several examples where we definitely felt like we were on to something and we were on to something We were just too early into the space and so it's been fun to see like projects kind of go on pause or you know, stop for a little bit and then some of them are starting to even come back around. Again, at this point and so sometimes we just were a little too early, but it just gives us a jump start when the models and the capabilities are ready. Good problem to have. Well, you think you've been really right on and like sticking to your convictions on. I think this, at least for me, a nice space like the Show and Tell piece.

51:15This idea that like you shouldn't, you shouldn't ask users to kind of write two pages of text to describe, for example, in image. the idea is that you should be able to show and tell like you would do a friend or an artist that's working with you. I think that has stuck in this, it's kind of moving people away from prompting and towards kind of instructing and relying on the intelligence that lives behind. So I think that's one. But when I'm sticking with my guns and I think you're yours there to stay. Yeah. I mean, this is probably obvious at this point, but when we all started in labs, there was no Google LOM API, Google didn't have a functional instruction tuned language model or anything like that.

51:55I believe it will not back then. In fact, I think the general consensus was that these were not really things that are easy to build a business around because of their cost. I think one of the things that was all done actually is we've kind of stuck with the technology and now it's obvious, right? But in the early days, it wasn't like, it certainly was not obvious. So, yeah, we got that bit of timing right. Yeah, yeah. Yeah, yeah. Inferno's costs, just writing that curve and just capabilities up, cost down, and so what will you build assuming that those curves continue? Yeah, exactly. In fact, when we joined, one of the transitions is right to think inside labs that George started actually.

52:36And a lot of the docs that we'd write were around, well, what happens in two years? And of course, yeah, that curve is something that I think inspired a lot of us. Yeah. Fantastic. Thank you all so much for joining to share what you're doing across the creative sphere, the computer use sphere, and the, what do I call the notebook sphere? The podcast killer slash. Let's not say podcast killer, but yeah, we can say knowledge. Knowledge creation, transformation, space. It's just it's really really cool what you all are building and you guys have such a cool job getting to kind of cook in the little test kitchen of Google And thank you for giving a preview of some of the stuff that's coming down the pipeline Thank you

From the publisher

Fresh off impressive releases at Google’s I/O event, three Google Labs leaders explain how they’re reimagining creative tools and productivity workflows. Thomas Iljic details how video generation is merging filmmaking with gaming through generative AI cameras and world-building interfaces in Whisk and Veo. Jaclyn Konzelmann demonstrates how Project Mariner evolved from a disruptive browser takeover to an intelligent background assistant that remembers context across multiple tasks. Simon Tokumine reveals NotebookLM’s expansion beyond viral audio overviews into a comprehensive platform for transforming information into personalized formats. The conversation explores the shift from prompting to showing and telling, the economics of AI-powered e-commerce, and why being “too early” has become Google Labs’ biggest challenge and advantage.

Hosted by Sonya Huang, Sequoia Capital

00:00 Introduction

02:12 Google's AI models and public perception

04:18 Google's history in image and video generation

06:45 Where Whisk and Flow fit

10:30 How close are we to having the ideal tool for the craft?

13:05 Where do the movie and game worlds start to merge?

16:25 Introduction to Project Mariner

17:15 How Mariner works

22:34 Mariner user behaviors

27:07 Temporary tattoos and URL memory

27:53 Project Mariner's future

29:26 Agent capabilities and use cases

31:09 E-commerce and agent interaction

35:03 Notebook LM evolution

48:26 Predictions and future of AI

Mentioned in this episode: 

Whisk: Image and video generation app for consumers

Flow: AI-powered filmmaking with new Veo 3 model

Project Mariner: research prototype exploring the future of human-agent interaction, starting with browsers

NotebookLM: tool for understanding and engaging with complex information including Audio Overviews and now a mobile app

Shop with AI Mode: Shopping app with a virtual try-on tool based on your own photos

Stitch: New prompt-based interface to design UI for mobile and web applications.

ControlNet paper: Outlined an architecture for adding conditional language to direct the outputs of image generation with diffusion models

More from Training Data

All 110 episodes
Google I/O Afterparty: The Future of Human-AI Collaboration, From Veo to MarinerTraining Data · 54 min
Listen in VO