LIVE: Google's Jeff Dean on the Coming Transformations in AI

16 May 2025 · 31 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Episode Summary: LIVE - Google's Jeff Dean on the Coming Transformations in AI

Podcast Overview Title: Training Data Description: Join Sequoia Capital partners as they discuss AI with leading builders and researchers to deepen understanding of evolving technologies and their societal implications. Episode Title: LIVE: Google's Jeff Dean on the Coming Transformations in AI Episode Description: Jeff Dean shares predictions on AI's evolution at the AI Ascent 2025 conference, discussing advancements from specialized hardware to more organic systems.

---

Key Participants

  • Jeff Dean: Chief Scientist and AI lead at Google.
  • Bill Corn: Sequoia partner and former Google engineer.

---

Episode Highlights

Introduction

  • Live interview conducted at Sequoia's annual AI conference in San Francisco.
  • Jeff Dean's contributions to AI include leading the development of Google's TPUs and foundational AI research.

Evolution of AI

  • Historical Context: The AI industry began gaining traction around 2012 with large neural networks solving complex problems across vision, speech, and language.
  • Scaling Models: Emphasis on how larger neural networks yield better results, with the mantra "bigger model, more data, better results" becoming prevalent.

Current Landscape

  • Multi-Modalities: The capability of models to integrate and process various data types (audio, video, images, text).
  • Agent Development: Discussion on the industry's fascination with AI agents and their potential, though some current implementations are seen as lacking substance.
  • Future Capability of Agents: Prediction that agents will eventually perform many tasks in virtual environments, with gradual improvements leading to more sophisticated applications.

Large Language Models (LLMs)

  • Market Landscape: Few major players are likely to dominate due to the high investment needed for cutting-edge models.
  • Diversity in Models: Variations in models will emerge, focusing on different functionalities and applications.

Hardware Considerations

  • Specialized Hardware: Importance of accelerators designed for machine learning computations.
  • Google's TPU Innovations: Development of TPUs to enhance both inference and training capabilities.

AI's Influence on Science

  • AI and Research: Jeff Dean discusses how AI is revolutionizing various scientific disciplines by providing faster data analysis and simulating complex processes.

Future of Computing

  • Changing Algorithms: The need to rethink computational approaches due to the unique demands of AI workloads, focusing on efficiency and speed.
  • AI's Role in Everyday Tasks: Predictions on how AI could assist in routine tasks, making technology more accessible and efficient.

Predictions for AI Development

  • AI Junior Engineer: Dean predicts that a virtual assistant capable of performing at the level of a junior engineer could emerge within the next year.
  • Skill Development: Essential skills for AI systems will include debugging, performance testing, and tool usage, mirroring human engineers.

Closing Thoughts

  • Dean emphasizes the importance of continuous learning and adaptation in AI systems, advocating for an organic approach to model development that mimics human cognitive processes.

---

Key Takeaways

  • Scalability of AI models remains a critical focus, with expectations of increasingly sophisticated agents and models.
  • Investment in hardware and algorithmic innovation is necessary for maintaining competitive advantage in AI.
  • The future of AI may involve a blend of large foundational models and specialized lighter models tailored for specific applications.
  • AI's potential to transform industries and daily tasks is immense, particularly as it integrates into tools and services that enhance productivity and efficiency.

---

This episode provides an insightful look into the future of AI through Jeff Dean's expert perspective, touching on significant advancements, challenges, and the transformational potential of this technology.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Hi and welcome to Training Data. We are mixing it up for this week's episodes and dropping a conversation that was filmed live as Sequoia's annual AI conference in San Francisco with Google's Chief Scientist in AI lead, Jeff Dean. Jeff is interviewed by our partner and Google alum, Bill Corn. We hope you enjoyed this special conversation with Jeff about the future of model development and compute, whether or not he likes vibes coding, hint he does, and his expected timelines for a 24 -7 software developer agent. We have Jeff Dean. And if you read Jess Bio, he's run everything at some point in Google, including overseeing the genesis of this industry and the birth paper that kind of sparked things so many years ago.

0:44And we're very fortunate to have our partner, Bill Korn, who's been about a decade before Sequoia, running most of engineering at Google with Jeff. And so please welcome Jeff and Bill.

0:58Thank you. Jeff, it's great to see you. We got to work together for a few years. Jeff still was occasionally willing to talk to me, which I'm very proud of. We have an occasional dinner, which is great fun. Yeah, no, he's now the chief scientist, I think, and Alphabet. I thought we'd start. Obviously, a lot of the people in the room are excited about AI and what's happened with AI. Google clearly introduced a lot of the tech that the industry is based on transformers and other things. Where do you see things going these days as you look out both within Google but also in the industry as a whole?

1:41Yeah, I mean I think this sort of period has been a fairly long time in developing even though it's sort of come into sort of popular visibility only in last three or four years, but really starting maybe in 2012 and 13 people were starting to be able to use these really, you know, at that time what seemed like large neural networks to solve interesting problems and the same sort of algorithmic approach would work for vision and for speech and for language. And that was, you know, pretty remarkable and and kind of brought attention to machine learning as a way to solve those problems rather than sort of more traditional handcrafted approaches.

2:24And one of the things we were interested in in 2012 even was how can you scale and train very, very large neural networks. So we trained a neural network at the time was 60X larger than anything else. And we used 16 ,000 CPU cores because that's what we had in our data centers. And got really good results. And that really cemented in our mind that scaling these approaches would really work well. And there's been a whole bunch of evidence of that and hardware improvements to help increase our ability to scale to larger and larger models and larger datasets. We had an expression bigger model, more data, better results, which has been sort of relatively true for the last 12 or 15 years.

3:11And where things are going, I think, you know, now the models that we have are capable of doing really interesting things. You know, they can't solve every problem. They can solve a growing set of problems year over year because the model to get better, you know, we have better algorithmic improvements that show us how to, you know, train larger models with the same compute cost, more capable models. And then we have scaling of hardware. We have increasing compute per unit of hardware. And also we have reinforcement learning and post -training kinds of approaches that are making the models better and sort of guiding them into the ways that we want them to behave.

4:02And that's really exciting. I think, you know, multi -modalities and other big things, like having, you know, the ability to put in audio or video or images or text or code and have it sort of output, all those kinds of things as well. It's pretty useful. The industry is, I think, mesmerized by agents right now. How real do you think agents are? I know Google introduced an agent framework. Some of this stuff's not Google's necessarily, but some of the agent stuff seems to be a little bit vapor -weared in me. Sorry folks, I'm a little direct as some folks will tell you. I mean, I think there's a lot of promise there because I do see a path for agents with the right training process to eventually be able to do many, many things in the virtual sort of computer environment that humans can do today.

4:58You know, right now they can sort of do some things, but not most things, but the path for increasing the capability there is, you know, reasonably clear. You get more reinforcement learning going. You have more agent experience that it can learn from. You have, you know, early nascent products that can do some things, but not most things, but are still incredibly useful for people. And I think similar things will happen in sort of physical robotic agents as well. Like right now, they were probably close to making that transition from robots in messy environments like this room kind of don't quite work today, but you can see a path where in the next few year or two, they'll start to be able to do 20 useful things in this room.

5:47And that will introduce pretty expensive robotic products that can do those 20 things. And then learning from experience, they will then get cost -engineered to now have something that's 10 times cheaper and can do, you know, 1000 things. And that's going to engender even more cost -engineering and more improvement in capability. So, it's exciting. It is and it does seem like it's coming even those figures where today. But they, I guess one of the other things that comes up, I think, with a lot of young companies is what's happening with the large models. I mean, clearly, you know, Google has a Gemini 2 .5 Pro and deep research and so forth.

6:30And then there's OpenAI and a number of other players. I think there's an open debate about how many large language models, open source, close source, where things going. How do you think about that? Obviously, Google has a strong position and wants to, I'm sure, dominate in that area, but How do you see the landscape? Yeah, I mean, I think clearly it takes quite a lot of investment to build the absolute cutting edge models. And I think there won't be 50 of those. There may be like a handful. And there are an awful lot. You know, once you have those capable models, it's possible to make much lighter weight models that can be used for many more things because you can use techniques like distillation that I was a co -author on and got rejected from NURP's 2014 is unlikely to have impact.

7:27I've heard that technique may have helped deep -seek. So that's a really nice technique if you have a better model and then you can put it into a smaller scale thing that actually is pretty lightweight and fast and all the kinds of properties you might want. So, I mean, I think there will be quite a number of different players in this space because, you know, different shape models or models that focus on different kinds of things. But I also think, you know, a handful of really capable general -purpose ones will do pretty well. Fair enough. I guess hardware is the other thing that's interesting.

8:06It looks to me like every large player is building their own hardware. where obviously Google has been very public about the TPU program, but Amazon has their own rumors are meta has one rumors are opening eyes building one. There's lots of hardware. And yet the industry seems to only hear about Nvidia. How fairly for I'm sure that's not true in your office. But how do you how do you think about that? How important is specialized hardware for this stuff? Yeah, well, I mean, it's very clear that having hardware that is focused on sort of machine learning style computations, and I like to say accelerators for reduced precision linear algebra are what you want, and you want them to be better and better generation over generation, and you want them to be connected together at large scale with super high speed networking so that you can spread your model computation out over as many computer devices as possible.

9:10I think it's super important. I helped bootstrap the TPU program in 2013, because it seemed obvious we would want a lot of compute for inference at that time. That was the first generation. And then the next generation of TPU's TPU V2 was focused on both inference and training, because we saw a big need there. And I think we're on now. We stopped numbering them for some annoying reasons. and so now we're on Ironwood, which is coming out any day now. And be truly before that. Be careful. That sounds like an Intel chip naming strategy, which has a more dead wall. Small editing distance to my Kenium, which is a little...

9:50Yeah, no, I guess going a little bit off topic, and then maybe we'll open to questions from folks in the room. I have a lot of friends who are physicists. They were a little surprised with Jeff Hitten and his colleagues, one the Nobel in physics. I guess, how do you see AI? Some of the physicists I know are sort of offended that a non -physicist is starting to win Nobel prizes. How far do you think AI is gonna go in various fields at this point? Pretty far, I think. Also this year, my colleague, Demis and John Jumper wanted for - I almost forgot to ask. Yes, so double Nobel Prize celebration Monday and Tuesday or whatever it was.

10:46So, I mean, I think that's a sign that really AI is influencing lots of different kinds of science because at its core, can you learn from interesting data? And a lot of parts of science are about making connections between things and understanding them. And if you can have AI assisted in helping in doing that, one of the things I've seen in many different fields of science is many disciplines often have incredibly expensive, computational simulators of some process, like weather forecasting is a good example or fluid dynamics or quantum chemistry simulations. And often what you can do is use those simulators as training data for a neural net and then build something that approximates the simulator, but now is 300 ,000 times faster.

11:39And that just changes how you do science because all of a sudden, well, I'm gonna go to lunch and screen 10 million molecules that's now possible instead of, I would have to run that for a year on computer I don't have. And I think that that just kind of fundamentally changes your your your process of what you how you do things and will make faster discoveries. I think it is probably the most interesting for their questions from the audience at this point. I have other questions for Jeff but well actually just to quickly follow up on the that Jeff Hinton famously left Google after studying, I guess, the effects of, or the differences between digital analog computing as a future platform for inference and learning.

12:29And I'm wondering, is the future of inference hardware analog? It's definitely a possibility. I mean, I think analog has some nice properties in terms of it being very, very power efficient. And I think there's a lot of room for digital things to be much more specialized for inference as well. So, and it's a little bit easier to work with typically. But I think there is a general direction of how can we make inference hardware that is 10, 20, 50, 1000 times more efficient than what we have today. And that seems eminently possible if we put our minds to it. It's actually something I'm spending a bit of time on.

13:12So.

13:17Hi. I was just going to ask about developer experience versus hardware. I think the TPU hardware is extremely impressive. But there's a lot of, you know, in the side guys about how CUDA or different like, you know, technologies are easier to use than the TPU layer. And so I'd be curious for your perspective on that. And is that something you've been thinking about or getting a lot of angry emails about? Yeah, I mean, I don't connect with cloud, TPU customers, all that much, but definitely the experience can be improved. One of the things we started working on in 2018 is a system called Pathways, which is really designed to enable us to take lots of different computing devices, and then give sort of a really nice abstraction with those where you have a virtual physical device mapping that is managed by the underlying runtime system.

14:09And we have support for that for both PyTorch and Jax. We primarily use Jax in -house, but what we have is a single Jax Python process just looks like it has 10 ,000 devices on it, and you just write your code as you would as an ML researcher. And off you go, you can prototype it with 4, 8, or 16, or 64 devices, and then you change a constant. and you run against a different pathways back in with a thousand and a thousand chips. And off you go, like our largest Gemini models are trained with a single Python process driving the entire thing with tens of thousands of chips and it works quite well.

14:50So pretty good developer experience, I think. One thing I would say is to date, we did not offer that to cloud customers, but we just announced at cloud next that we're now gonna have pathways available for cloud customers. So then everyone else can have the delightful experience of a single Python process with thousands of devices attached.

15:13And I agree that's a much better experience than managing like 64 processors for your 256 chips. Why would you want to do that?

15:26I love using the Gemini API. It would be even easier if it got one API key rather than like the Google Cloud, credential setup. Do you guys have a plan to unify the Google Cloud Gemini stack with the Gemini project setup right now that's more for testing stuff? Yeah, I think there's a bunch of streamlining that is being looked at. It's a known problem, not something I spend a lot of time on personally, but I know Logan and others on the developer side are aware of this friction. We'd like to make I get frictionless to use our models.

16:09Is that working? OK. So it's an interesting time in computing. You've got the confluence of Moore's Law and the art scaling being completely dead with AI, just scaling like crazy. You have a pretty unique position in the world of driving these supercomputers and infrastructure that is being built. And you know how to map the workloads close onto these things, which is a unique skill. What do you think the future of computing is going to look like? What is the computing infrastructure heading towards? From an asymptotic thought experiment level? Yeah, I mean, it's really clear that we will have dramatically changed the kinds of computations we want to run on computers in the last, say, five years, 10 years.

16:56And that was initially a small ripple, but it's pretty clear now that you want to run incredibly large neural networks at incredibly high performance and incredibly low power. And you also want to train them. Training and inference are pretty different kinds of workloads. So I think it's useful to think of those two as you know you probably want different solutions for the two or somewhat specialized solutions. And I think you're going to see all kinds of adaptation of compute platforms for this new reality that you really just want to run incredibly capable models. And so some of that will be in low power environments like your phone, like you'd like your phone to run incredibly good models with lots of parameters super fast so that when you talk to your phone, it just talks back to you and it can help you do all kinds of things.

17:56You're going to want to run these on robots and autonomous vehicles. You know, we already do someone, but even better hardware for that will make those systems much easier to build much more capable, you know, physical agents in the world. And then you want to run them at incredibly large scale and data centers. And then you also then want to use lots of inference time compute for some kinds of problems, but not others. So you have probably, you know, it's pretty clear you want to use 10 ,000 times as much compute for some problems as for others. And that's a nice new scaling knob we have that can make your model much more capable or give you, you know, much better answers or make the model capable of doing things with that much compute that it can't do with, you know, one X as much compute.

18:42But you shouldn't spend 10 ,000 times as much compute on everything. So how do you make your systems work well for that? And I think that's a combination of hardware, system software, model and algorithmic tricks, distillation, all these things can help you make amazing models come to life in small compute footprints. One thing I've noticed as the computer science, at least traditionally, when people are studying algorithms and computational complexity was all off count based. And I think as people are rediscovering hardware or in details of hardware and system design, I think one of the things that's come back into focus is you need to think about network bandwidth and memory bandwidth and so forth.

19:30And so I think a lot of the kind of traditional algorithmic analysis needs to be completely rethought just because of realities of what real compute computational looks like. Yeah, one of my office mates in grad school did his thesis on like cash aware algorithms because the order of magnitude, they go kind of notation didn't account for the fact that some operations were 100x worse than others. Yeah, that's right. And I think in modern ML computing, you care about data movement at the incredibly small level like moving things from SRAM and to accumulators cost you some tiny number, some tiny number of picojoules, but it's way more than the actual operation cost you.

20:15So it's important to have picojoules at the tip of your tongue these days.

20:22One other quick question. Do you vibe codes?

20:30I've been trying it a little bit. It actually works surprisingly well. Yeah, I mean, we've had some nice, we have a little demo chatroom, actually. We have a lot of chatrooms. We sort of run Gemini via chatrooms. So I'm in like 200 chatrooms. And when I wake up and brush my teeth, I get like nine notifications, because my London colleagues are busily doing things. We had one where people can send out cool demos of things I've seen. And one that was particularly cool was you feed in a YouTube educational oriented video. And the prompt is just something like, please make me an educational game that uses graphics and interactivity to help illustrate the concepts of this video.

21:17And it doesn't work every time, but 30 % of the time you get something that's actually kind of cool and related to differential equations or traveling to Mars or doing some kind of cell aspect thing. That's just an incredible sign for education. The tools we now have and will have in the next few years really have this amazing opportunity to change the world in so many positive ways. So I think we shall remember that as kind of what we should be striving for. What do you mind passing there and maybe there? Yeah, we'll look to hear your thoughts about the future of search and especially given Chrome such big distribution, right?

22:05And especially Chrome already know the credentials, like payments and then web signing credentials. Have you thought about like getting Gemini just directly into Chrome, and making the current app, Gemini app, instead of have a separate app. I say this because I'm long -term Googlers, so just think about it. Yeah, I mean, I think there are definitely lots of interesting downstream uses one could make of the core Gemini models or other models. One is, can it help you do stuff in your browser or on your full computer desktop by observing what you're doing and doing OCR on tabs or maybe it has access to the raw tab contents.

22:50That seems like it will be incredibly helpful. And I think we have some early work in this area that we've published public demos of in video form that seem pretty useful, things like Mariner and things like that. So, TBD. you passed. Jeff? Question for you. So thank you for your comments. Very insightful. Earlier you mentioned the number of foundational model players will likely only be a handful. This is largely because of the infrastructure costs and the scale of investment to remain at that cutting edge. And so as this battle for the frontier unfolds, how do you see the What do you see this end game going?

23:40Where does this lead us? Is it just whoever writes the biggest check to build the biggest cluster wins or is it better? You just talked about better utilization of unified memory optimization and different efficient uses of what you already have or is it the consumer experience or like how? Where does this arms race lead us? isn't it just who I guess is kind at first the games over? Yeah, I mean I think it's gonna require really good insightful algorithmic work as well as really good systems, hardware and infrastructure work. I don't think either one of those is more important than the other because what we've seen in say our Gemini progression from generation to generation is the algorithmic improvements are as important or maybe even more so than the hardware improvements or the more, you know, larger amount of hardware we're putting to the problem.

24:45But both are incredibly important. And then I think from a product standpoint, you know, what it's, there's sort of early stage products in this space but I don't think we've collectively hit on what is the thing that or it's probably going to be many things that become the daily used products for billions of people. I think there's probably some in the educational space or in general information retrieval that is search -like but sort of taking advantage of the strengths of large multi -model models. I think probably helping people get stuff done in whatever work environment they find themselves in is going to be an incredibly useful thing and how will that get manifested in product settings.

25:40How do I manage my team of 50 virtual agents that are off doing things and they'll probably be mostly doing the right thing but occasionally they'll need to consult with me about some choice they need to make. I need to give them a bit of steering. How do I manage 50 virtual interns? It's going to be complicated. Hi Jeff. Thanks for being here, right here. Sorry. I literally cannot think of anyone better in the world to ask this question. How far do you believe we are from having an AI operating 24 or seven at the level of a junior engineer?

26:26Not bad for. Yeah. Yeah. Yeah. Is that six weeks or six years or? Every year in AI seems like a dog, dog seven or something. I will claim that's probably possible in the next year. Yeah. Hi, Jess, you talked about scaling pre -training and now scaling RL. How do you think about the future trajectory of these models? Will it be one large model with all the compute or a constellation of smaller models that have been distilled from these larger models, both working in parallel? How do you see the future landscape? Yeah, I've always been a big fan of models that that are kind of sparse and have different parts of expertise in different parts of the model, because from our weak biological analogies, that's partly how our real brains get so power efficient is we're 20 watts or whatever, and we can do a lot of things, but our Shakespeare poetry part is not active when we're worried about the garbage truck backing up at us in And I feel like there's, we do some of that with mixture of expert style models.

27:43You know, we did some of the early work in that space where we had like 2048 experts and showed that it gave you dramatic improvements in efficiency, like 10 to 100x more efficient sort of model quality per training flop. And that's super important. But it feels like we're not really fully exploring the space yet because right now the kinds of Sparsity people tend to do is incredibly regular Like it feels like you want paths through your model that are like a hundred or a thousand times more expensive than other paths And you want experts or pieces of your model that are tiny amounts of compute and some that are very large amounts of compute Maybe they should have different structures and I think you want to be able to extend your model with like new parameters or new bits of space and maybe you want to be able to compact parts of your model, running a distillation process on this piece of it to make it one quarter of the size and then you have some background garbage collection and you think that is now like, oh great, I have more memory to use.

28:49So I'm gonna put those parameters or put those, you know, bytes of memory somewhere else and make more effective use of them somewhere else. And so that to me seems like a much more organic, continuous learning system than what we have today. So the only problem with this is what we're doing today is incredibly effective. So it becomes a bit hard to completely change what you're doing to be more like that. But I really do think there are huge benefits to doing things in that style, rather than the sort of more rigidly defined model that we have today. I think one more question, and then we'll probably wrap up.

Read the full transcript

29:36Hey, I wanted to return to the junior engineer inside a year. I'm curious, what advancements do you think we need to get there? Like, obviously, just maybe code generation gets better. But outside of code generation, what do you think gets us there? Tool use, genetic planning. Yeah, I mean, I think they, you know, this hypothetical virtual engineer probably needs a better sense of many more things than just writing code in IDE. Like, it needs to know how to, like, run tests and, like, debug performance issues and all those kinds of things. And we know how human engineers do those things. They learn how to use various tools that we have and can make use of them to accomplish that and they get that wisdom from more experienced engineers typically, or reading lots of documentation.

30:25I feel like junior virtual engineer is going to be pretty good at reading documentation and sort of trying things out in virtual environments. That seems like a way to get better and better at some of these things. And, you know, I don't know how far we'll take us, but it seems like it'll take us pretty far. Jeff, thank you for coming and sharing your wisdom. Thank you. Thank you, see you. Thank you.

From the publisher

At AI Ascent 2025, Jeff Dean makes bold predictions. Discover how the pioneer behind Google's TPUs and foundational AI research sees the technology evolving, from specialized hardware to more organic systems, and future engineering capabilities.

More from Training Data

All 110 episodes
LIVE: Google's Jeff Dean on the Coming Transformations in AITraining Data · 31 min
Listen in VO