In short
How Google’s TPUs and the software stack (JAX, XLA, Pathways) enable planetary-scale AI training, and why public academic research and new funding models are essential for future breakthroughs.
Guest backgrounds
Jeff Dean, described as a leading chief scientist at Google, present for major AI systems breakthroughs over ~two decades.
Key claims
Specialized hardware is required because general-purpose CPUs become prohibitively expensive at scale; TPU-V7 uses FP4 to cut memory/traffic (claimed ~75% smaller footprint vs BF16) enabling ~4x faster training. Pathways provides a “single system image” across up to 9,216-chip TPU pods and across cities by managing topology and latency. Societal progress depends on public research; examples include TCP/IP, RISC architectures, and PageRank from a Stanford Digital Library Project grant.
Notable examples
TPU-V1 benchmarks (30–70x energy efficiency; 15–30x faster inference); Titan hybrid transformer-recurrent model (published but not used in Gemini); Google DeepMind-style “vertical co-design”; Lott Institute moonshot labs (3–5 years, 3–5 PIs, 30–50 PhDs/postdocs/engineers) and federated learning for healthcare privacy.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding AI's Technical Scale
0:20 to 1:28
Learn about the immense scale needed for AI models and funding reforms.
“that are required to fuel the next generation of real societal breakthroughs.”
Introduction to TPU-V7
1:28 to 2:15
Discover the latest advancements in Google's tensor processing units.
“It's about a concept from the source called vertical co-design and a really interesting strategic vision that's focused on a three to five year timeline.”
The Role of Low Precision Formats
2:15 to 3:29
Understand the implications of using FP4 in TPU-V7 for AI models.
“And this is clearly not just some, you know, incremental speed bump.”
Efficiency vs. Accuracy in AI
3:29 to 5:40
Examine the balance between efficiency and accuracy in AI computations.
“Well, the V7 specifically enhances its capabilities for these extremely low precision floating point formats.”
Google's Shift to Specialized Hardware
5:40 to 7:20
Learn about the pivotal decision to develop specialized AI hardware.
“The remarkable thing is that this whole TPU program didn't start because they were chasing trillion parameter models.”
TPU-V1: A Historical Perspective
7:20 to 8:35
Trace the development and impact of the first TPU on AI.
“Just to deliver a marginal improvement in speech recognition.”
The Importance of Vertical Integration
8:35 to 11:28
Discover how vertical integration enhances AI hardware and software co-design.
“And it's so important to place that TPU-V1 in its proper historical context.”
Pathways: The Software Layer for Scaling
11:28 to 14:00
Explore how Pathways optimizes the software for efficient AI training.
“Okay, so we have this spectacular specialized hardware, the TPU-V7, 9 ,000 chips in one pod, lightning fast because of things like FP4.”
Understanding Pathways and TPU Coordination
14:00 to 16:40
Learn how Pathways optimizes TPU operations and memory management.
“You have to manually manage all the synchronization points across the network.”
Foundational Research and Societal Debt
16:40 to 19:09
Discover the critical role of public academic research in tech advancements.
“It effectively makes geography irrelevant to the training process.”
Show all 18 chapters
The Return on Investment of Academic Grants
19:10 to 20:52
Public funding for research yields astronomical returns, as exemplified by PageRank.
“The deep learning revolution that we are all experiencing today isn't some recent invention.”
The Long-Term Need for Research Funding
20:52 to 24:24
Explore the necessity of sustained public funding for foundational research.
“The reality is that fundamental research often moves too slowly and is far too high risk for even the largest corporate R &D divisions to fully fund without a really clear path to monetization.”
Innovative Funding Models for AI Research
24:24 to 28:01
Examine the Lott Institute model and its approach to applied AI research.
“The focus is on tackling solvable problems using technology that often already exists or is very near fruition.”
Challenges of Data Standardization and Policy
28:01 to 29:12
Learn about the complexities of data standardization and the legal hurdles in AI research.
“And getting this mountain of disparate, non-standardized data into a usable, structured form for training a global model is this immense, non-glamorous, but absolutely necessary undertaking.”
Balancing Corporate Research and Academic Ecosystem
29:12 to 31:19
Discover how companies navigate the relationship between internal research and the academic community.
“It sounds like the technical challenge is maybe even easier than the policy challenge in this context.”
Types of Research Contributions in AI
31:19 to 34:09
Understand the three types of research contributions and their implications for innovation.
“The model learns to compress those 512 tokens into a summary or a memory vector, and then it uses recurrence steps on the sequence of those compressed chunks.”
Internal Ecosystems and Idea Circulation
34:09 to 35:50
Explore how internal conferences enhance rapid idea generation and R&D cycles.
“that means they have to manage their own closed ecosystem of knowledge sharing just to accelerate the pace of invention.”
The Future of AI Funding and Innovation
35:50 to 38:15
Discuss the need for practical funding models to tackle societal challenges through AI.
“So we've taken the full journey today, really tracing the core pillars of modern AI innovation.”
Transcript
Automatic transcript. May contain errors.0:00Jeff Dean:Okay, let's unpack this. Today we are doing a deep dive into, well, the very foundation of modern artificial intelligence. We aren't just here to talk about the latest large language model release, you know, the shiny new thing. We are digging way deeper into the infrastructure that's actually holding all of this up. We're talking about the specialized silicon, the massive software layer that has to manage planetary scale computation, and this is critical, the radically different funding models that are in the world. that are required to fuel the next generation of real societal breakthroughs.
0:35It's a perspective you just don't get very often. And our source material today is, it's incredibly rich. It's drawn mostly from a conversation with one of the leading chief scientists at Google.
0:45Jeff Dean:Someone who's been in the room for basically all of it. Literally. Someone who has been there for the genesis of nearly every massive systems breakthrough in this space for the last two decades. And the insights, they just cut straight through the usual AI hype cycle. They connect the relentless need for hardware efficiency, the TPU, with the fundamental and I think often forgotten role of public academic research. So our mission for you, the learner, is sort of multifaceted, but it's very clear. First, we really want you to grasp the profound technical scale that's required to train and deploy a model like Gemini.
1:23Jeff Dean:I mean, we are talking about computation that literally spans multiple cities. It's mind-boggling stuff. It is. And second, and this is maybe even more important, we want to show you why the next great societal breakthroughs in AI depend just as much on a reforming how we fund smart people in research labs as they do on, you know, building the next generation of chips. It's not just about speed. Not at all. It's about a concept from the source called vertical co-design and a really interesting strategic vision that's focused on a three to five year timeline. And that whole journey, it really starts where the data meets the metal.
1:56It has to start with the specialized hardware. To understand the vision of AI at this scale, you have to begin with the silicon that makes it all possible.
2:04Jeff Dean:So let's start right at the cutting edge of silicon design. Because just recently, we saw the announcement of the seventh generation of Google's tensor processing units, the TPUs. Right, V7. And this is clearly not just some, you know, incremental speed bump. This is hardware that is built explicitly for a world that's defined by massive interconnected scale. That's absolutely right. When you look at the new TPU-V7, what you're seeing is a chip designed from the ground up to handle these immense memory and computation demands of today's enormous, you know, multi-trillion parameter models. Its architecture has to be fundamentally different.
2:40It is. It's fundamentally different from a standard CPU or even a general purpose GPU in how it prioritizes throughput and raw efficiency over everything else.
2:50Jeff Dean:And the scale they're talking about, it's just difficult to even visualize. They connect these chips together into these enormous configurations called pods. Pods. And these pods total an astonishing 9 ,216 chips per pod. That's almost 10 ,000 specialized accelerators, all acting as one single cohesive unit. And you have to understand that number, 9 ,216 chips, that is the new baseline. That's the table stakes for serious large scale AI development. But to get that kind of scale to work efficiently, the hardware has to make some pretty serious compromises on traditional precision. Okay, what do you mean by that?
3:30Well, the V7 specifically enhances its capabilities for these extremely low precision floating point formats. And the big one is called FP4.
3:36Jeff Dean:Wait, wait, FP4 as in a four-bit floating point number? Four bits, that's it. That seems like a massive trade-off. How can you even represent useful information with just four bits? What is the real-world consequence, you know, both good and bad, of dropping your precision that low? Are there certain tasks where that just becomes totally unacceptable? That is an excellent question, and it gets right to the heart of this kind of hardware design. I mean, historically, standard computing used 32-bit floating point, FP32. Which is super accurate. Tremendous dynamic range. Very high accuracy. Then, for the first wave of training big models, we all moved to 16-bit formats, like BF16.
4:12But for training these new models with hundreds of billions or even trillions of parameters, the single biggest bottleneck isn't the processing power itself.
4:21Jeff Dean:It's moving the data. It's the memory bandwidth. with. It's all about how fast you can physically move the weights and the activations onto and off of the silicon. So using FP4 is really just a resource management strategy. It's a data traffic solution. Exactly. When you use FP4, you are drastically reducing the memory footprint needed for those weights and activations. You're cutting it by 75 % compared to BF16, and even more compared to the old FP32 standard. And that means you can either fit dramatically larger models into the same amount of superfast on-chip memory, or, and this is key, you can train your existing models four times faster because you are physically moving four times less data across those critical interconnects.
5:03Jeff Dean:But what about the loss in accuracy? Doesn't that hurt the model's performance? It's a trade-off, for sure. The sacrifice in precision is deemed acceptable for deep learning because these models, they often exhibit a pretty high tolerance to this kind of quantization, especially during the inference stage. But you're right to be skeptical. For high-fidelity scientific computing, like, say, fluid dynamics, simulations, or weather modeling, FP4 would be completely unusable. Utterly useless. But for the specific math of matrix multiplication that basically defines what large language models do, it's the trade-off that unlocks planetary scale.
5:36Jeff Dean:And that relentless focus on efficiency, that's what drove the whole project from the very beginning. The remarkable thing is that this whole TPU program didn't start because they were chasing trillion parameter models. Legal. It started with a fundamental internal crisis way back in 2013, and it was focused purely on efficiency for inference, just the cost of running the models for users. That 2013 origin story is so crucial because it shows the systemic failure of general purpose computers when you throw these specialized workloads at them. Google was seeing tremendous measurable success using early deep learning and things like speech recognition and computer vision.
6:15And every time they trained a slightly larger model with a bit more data, the results just got objectively better. The quality was undeniable. But the cost, the cost was becoming absolutely prohibitive.
6:28Jeff Dean:And the source material shares this critical back of the envelope calculation that the chief scientist did. Can you walk us through the specifics of what they found? It's kind of a shocking moment. It was a startling calculation that literally forced a fundamental shift in the company's entire hardware strategy. They calculated what would happen if they rolled out just one improved speech model. Just one. A model that was only slightly more computationally intensive than the version they were running at the time. And this was for how many users? Just for 100 million users. And only for a few minutes of use per day.
7:00Not even for everyone. Not all day. And the finding was that if they tried to run this single improved application on their existing global infrastructure of standard CPUs, they would immediately need to double the total number of computers Google had overall.
7:15Jeff Dean:Double the entire massive global compute footprint of the company. Yes. Just to deliver a marginal improvement in speech recognition. For a fraction of their users. Wow. That's the moment that this specialized hardware program becomes an existential necessity. It's not a cool research project anymore. It was an immediate, completely unsustainable scale requirement. The realization was just crystal clear. General purpose CPUs were fantastic for many things, but they were hopelessly inefficient at the dense, low-precision, linear algebra that defines machine learning. They were just wasting energy.
7:47They were wasting something like 80 % of their power and their dye space on instructions and capabilities that machine learning didn't even use. The only viable path forward was to build their own specialized hardware, tailored specifically for the machine learning workload.
8:02Jeff Dean:And so when the first one, the TPOV-1, finally landed in their data centers in 2015, the validation of that strategy must have been palpable. Oh, the initial benchmarks were just astonishing. The TPU-V1 was immediately benchmarked as being 30 to 70 times more energy efficient than the contemporary high-end CPUs or GPUs. 70 times. And on top of that, 15 to 30 times faster for the inference workloads they were targeting. This efficiency gain translated directly into operational sustainability. It meant they could actually afford to deploy these new, better models. And it's so important to place that TPU-V1 in its proper historical context.
8:41Jeff Dean:This was before the transformer architecture completely reshaped the world of AI. Precisely. In 2015, these chips were focused primarily on the dominant AI architectures of that day. So, speech recognition, computer vision convolutional models, and sequence-to-sequence models for translation. But even then, they had to be agile. They did. The source mentions a really critical design change they made at the very last minute for V1. They had to scramble to squeeze in support for LSTM's long, short-term memory networks, which were, at that moment, rapidly proving to be crucial for language translation tasks.
9:16Jeff Dean:And that need for last-minute agility, that leads directly into this huge advantage of vertical integration, the co-design of hardware and software together. And this is where it gets really fascinating, because the role of the hardware designer shifts completely. They are no longer just engineers building a chip. They have to become technical futurists. Predicting the future of a field that changes every six months. Exactly. In a field where the dominant algorithm can become obsolete in two years, they have to try and predict where machine learning computation will need to run two and a half to, say, six years from now.
9:48That is an enormous and incredibly high-risk forecasting exercise.
9:52Jeff Dean:But doesn't that vertical integration kind of lock you into certain design choices? I mean, what if a future external chip architecture from a competitor, for instance, or some totally unexpected algorithmic breakthrough just fundamentally reshapes the AI landscape in a way your designers didn't see coming? That is the core risk they take on. Absolutely. But their strategy is designed to mitigate it. Their hardware is designed on the principle of betting on computational primitives, not on specific named algorithms. So they're betting on the math, not the model. That's a great way to put it. The core primitive is the massive, efficient matrix multiplication unit, the MMU.
10:30They deliberately sacrifice some of the general purpose flexibility you'd find in the CPU in favor of making those core ML operations hundreds of times faster and more efficient. So they're making these multi-year bets on algebraic needs.
10:44Jeff Dean:And they hedge those bets. They hedge. If they see an early stage research idea, you know, the early murmurings of attention mechanisms or the potential for using recurrence in new ways, which we'll talk about later, they'll include a hardware feature or a new capability for it. Even if it only dedicates a small, non-critical piece of the chip area, it's worth the investment. Because if that bet pays off. If that technology suddenly defines the next generation of models, they have a massive head start because their hardware is already tailored for it. And if it flops, well, they've lost a minimal amount of chip real estate, but they still have the best-in-class matrix multiplication units.
11:20That ability to control the full vertical stack from the algorithm running a jack all the way down to the individual silicon transistor, it just allows for an optimization and agility that no purely horizontal company can ever match.
11:33Jeff Dean:Okay, so we have this spectacular specialized hardware, the TPU-V7, 9 ,000 chips in one pod, lightning fast because of things like FP4. But as you mentioned, the hardware is only half the battle. If you just connect thousands of chips together, you don't instantly get a supercomputer. You get a massive power draw and a networking nightmare. You do. So now we have to talk about the critical software layer that actually enables that scale. We have to talk about pathways. That's the transition from optimizing the individual chip to optimizing the entire system. Because even with the fastest interconnects inside a single TPU pod, coordinating a training job that might involve 20 ,000 chips running for months and managing petabytes of data, that requires a software stack that can make all of that complexity just disappear.
12:22Jeff Dean:And the sources are pretty specific about the components of that crucial stack. They are. There are three critical layers that have to work in perfect concert for these large scale training jobs like Gemini. So at the very top, at the application level, the model code is often written using JAX. Right. JAX provides all the mathematical primitives and the symbolic differentiation that you need for machine learning. Yeah. But this code, it doesn't talk directly to the hardware. No, it goes through the compiler layer first. Right. Underneath JAX is XLA, which is the machine learning compiler. And XLA's job is to take that high-level computation graph from JAX and just optimize the heck out of it.
12:58It compiles it down into highly specific, super efficient instructions for the hardware, in this case, for the TPU backend. XLA figures out how to optimally schedule all those massive matrix multiplications on the silicon.
13:11Jeff Dean:But neither of those two layers solves the really hard problem, the networking and the distribution problem. And that is where Pathways, the coordinator, comes into the picture. Pathways is the true innovation here. It started about seven years ago, and its core concept was built around providing the developer, the ML researcher, with the illusion of a single system image across potentially tens of thousands of these individual accelerators. Okay, let's use that analogy again, but let's dig a little deeper into the mechanics. How does this illusion actually work for the person writing the code?
13:42So think about typical distributed computing. If you have one server with, say, four TPUs in it, your Python process sees those four devices.
13:50Jeff Dean:And that's it. And that's it. If you want to expand to a second server to use eight TPUs, your code now needs explicit instructions, often written by hand, on how to shard the model, how to split the parameters, and how to distribute the data. You have to manually manage all the synchronization points across the network. Which is a massive headache, and it just introduces all this debugging complexity. It's a nightmare. With Pathways, all of that manual sharding and synchronization, it just vanishes. Instead of your single Python process seeing four devices on one machine, it suddenly sees access to 20 ,000 devices.
14:27Jeff Dean:Just like they were all in one box. Exactly. And Pathways automatically figures out the optimal way to map the model's huge tensor operations and parameter updates onto the physical network topology, whether that's within one pod or across the entire globe. It dynamically manages memory addresses, and it rewrites all the network transfers for you. So the system just naturally works underneath the covers, figuring out all the plumbing and the traffic control for you automatically. Exactly. And Pathways manages this by having a really deep understanding of the underlying network architecture. So within a single TPU pod, those 9 ,216 chips, they use a dedicated, extremely dense, high-speed interconnect fabric.
15:06It's often implemented as a 3D or 4D torus.
15:09Jeff Dean:Super fast, super low latency. Incredibly fast and total proprietary. Data transfers here are almost instantaneous. So Pathways prioritizes this path above all others. But at some point, you're going to hit the limits of a single pod, and then you have to go to a different kind of network. Precisely. When data needs to cross between one pod and another, Pathways is intelligent enough to switch to the standard data center network. This is still extremely fast, you know, high-speed Ethernet and fiber, but it has a higher latency than that dedicated internal fabric. And Pathways manages the congestion and the potential packet loss across that boundary.
15:43Jeff Dean:And the true complexity, I imagine, comes when they need resources that are geographically dispersed, which for these truly massive training runs seems inevitable. That's where the system flexes its muscle in a way that is really unique to organizations that have this massy interconnected global infrastructure. When a commutational synchronization needs to happen across metropolitan areas, so say coordinating two massive pods in two different states or even different regions, Pathways manages the utilization of these long distance optical links. Which means what? It means that a single, continuous, months-long training job can be driven by one single Python process.
16:22And it can efficiently leverage multiple TPU pods located in multiple cities. It seamlessly handles all the different latency profiles, turning this disparate geographically separated infrastructure into one cohesive, flexible, and reliable machine learning supercomputer.
16:39Jeff Dean:It's just a remarkable feat of engineering. It effectively makes geography irrelevant to the training process. The whole system is completely vertical, from the decision to build the specialized silicon through the network management layer all the way up to the model code. Now, if we just take a breath for a second and connect this incredible vertical stack, the specialized TPUs, the Pathways software, the scale required to train Gemini, if we connect that to the bigger picture, we arrive at a really crucial and I think necessary realization. Which is? This entire cutting edge system, which required billions and billions of dollars of investment, rests on a foundation that Google did not invent.
17:19Jeff Dean:And that is the essential pivot in the source material. It's moving the conversation from, you know, proprietary advantage to societal debt. For all the talk of internal R &D and specialized hardware, none of it, none of it would be possible without the bedrock of public academic research. The chief scientist is unequivocal about this debt. The entire company and frankly, the entire modern global technological infrastructure was built on open, publicly funded academic research often conducted decades ago in university labs. We're talking about technologies that are so foundational now that they've almost become invisible.
Read the full transcript
17:54Jeff Dean:They're like the electricity or the plumbing of the Internet age. Absolutely. And they specifically cite some foundational elements. Things like TCPIP, the protocol that defines how data moves across the Internet. That was designed in open academic and government settings. They cite the architecture of advanced RISC processors, which pioneered the efficient, streamlined instruction sets that specialized chips like TPUs now mimic. And then there's the really powerful anecdote that connects directly to the company's own Genesis story. Right. They point to the specific funding mechanism that led to one of Google's earliest and most central innovations, which was the Stanford Digital Library Project.
18:35Now, this was a publicly funded initiative, and crucially, it provided the money for the original graduate student research project that eventually became the PageRank algorithm at Stanford University.
18:46Jeff Dean:So you're saying PageRank, the algorithm that basically defined modern search and launched the modern Internet economy, was the direct output of a publicly funded academic research project. That's exactly right. The return on investment for that initial relatively small academic grant is, it's just astronomical. You can't even calculate it. It's the ultimate evidence for the societal returns of funding basic research. It is. And this history extends deep into the AI revolution itself. The deep learning revolution that we are all experiencing today isn't some recent invention. It is built on fundamental academic research that is 30 to 40 years old.
19:23Jeff Dean:You're talking about the core inventions of neural networks and backpropagation, which were largely developed and theorized back in the 1980s and 1990s. The concepts were mathematically sound, but they just sort of sat dormant for decades. They did. They were part of the connectionism movement, which was often dismissed as being impractical during the dominance of symbolic AI. The mathematical groundwork was all there, but what was missing was the data and the compute. The very things the TPU was later designed to provide efficiently. Exactly. The academic seeds waited for decades for the technical infrastructure to finally catch up.
19:56Jeff Dean:That lag 30 to 40 years between the initial academic insight and the planetary scale commercial application, that's a kind of a terrifying thought when you consider the current rate of technological acceleration. And that realization leads directly to the core advocacy point from the source material, which is that a vibrant academic research ecosystem is absolutely essential. These institutions, these universities, are where the early stage creative, often counterintuitive and riskier ideas get explored without any immediate commercial pressure. These are the 30 year seeds. These are the ideas that lead to the major breakthroughs.
20:33They are the 30 year seeds that grow into global industries.
20:35Jeff Dean:But OK, let me challenge that for a second. If the returns are so vast, why should taxpayer money or public funding organizations like the NSF bear the full risk? Shouldn't these multi-trillion dollar companies who are now profiting from those ideas be the ones paying for the next generation of foundational research? And that's a completely fair challenge. The reality is that fundamental research often moves too slowly and is far too high risk for even the largest corporate R &D divisions to fully fund without a really clear path to monetization. What's more, the outputs of public academic research, published papers, open source code, they are inherently public goods.
21:14When a university publishes an idea, everyone benefits, which maximizes the societal return.
21:20Jeff Dean:So the argument is that we need a robust public funding model. The chief scientist's point is that society needs to ensure a robust funding model for this basic research, because the long-term returns are just far too large and too beneficial to be left purely to market forces. So before we can even talk about funding the next product, we absolutely must ensure we're funding the next foundation. So taking that point about funding the foundation a bit further, the discussion naturally moves to the need for alternative funding mechanisms for more applied AI research. If the traditional short-term grants just aren't adequate for the scale and the timeline you need for modern cross-disciplinary AI challenges, then what organizational model is best suited to bridge that gap.
22:00Jeff Dean:The gap between pure academic research and immediate, large-scale societal application. You need a model that fosters system building, not just paper writing. And the source details a very specific example that's designed to solve exactly this problem, the Lott Institute model. Okay, tell us about the structure of this nonprofit. It sounds like it's intentionally designed to avoid the short-term quarterly pressures of commercial work, while at the same time providing more sustained resources and a clearer timeline than a traditional university grant. It's a very deliberate, highly structured entity.
22:34It operates as a non-profit, a 501c3 in the U.S. structure, and it specifically raises money from successful technologists, so people who have already benefited immensely from the academic foundations we just talked about.
22:45Jeff Dean:Paying it forward. In a sense, yes, and its core function is to run a moonshot grant program that is dedicated to funding full research labs, not just individual graduate students. And what defines one of these labs? What makes it different? They are dedicated to funding research labs over a very targeted three to five year period. And these labs, they're structured to be sizable and interdisciplinary. You're talking about three to five principal investigators or PIs and typically 30 to 50 PhD students, postdocs and research engineers all working collaboratively under one roof. This concentration of talent combined with that clear time horizon is really the engine of the model.
23:25Jeff Dean:And that structure, the size, the interdisciplinary nature, and especially that time horizon, that seems to be key. Why is that three to five year window considered the delightful sweet spot for these ambitious research goals? Well, if the time horizon is too short, say a typical two year grant, just can't even conceive of doing anything truly ambitious that requires building complex systems or negotiating real-world agreements and deploying systems in testing environments. Two years is enough time to write a paper. It is not enough time to build a robust production-scale system. And if it's too long, say, 10 years?
23:59If it's too distant, the project just loses its urgency. The goals become too abstract, and the personnel might cycle out before you ever see the payoff. Three to five years is described as not so distant that it won't have impact, but not so short that it prevents ambitious goals. It forces tangible system-building deliverables while allowing enough time for the team to navigate all the inevitable technical policy and grungy hurdles.
24:23Jeff Dean:And this approach, it shifts the focus pretty dramatically away from the distant theoretical promise of AGI, which seems to dominate the conversation right now, and toward these achievable, high-impact societal goals. Absolutely. The research in these models is explicitly targeted toward measurable impact in areas like accelerating scientific progress, improving health care outcomes, facilitating job reskilling in vulnerable sectors, enhancing civic discourse. The focus is on tackling solvable problems using technology that often already exists or is very near fruition. The goal is to eliminate human drudgery or bureaucratic friction.
25:00Jeff Dean:And the source highlighted the application of AI in health with a real sense of personal passion, setting this incredibly audacious aspirational goal for society. The goal is profound because it links every single data point. The aspiration is, how can we as a society use every past decision that's been made in health to inform every future decision? That level of continuous scaled learning, a true learning health system, it would fundamentally change diagnostics, treatment protocols, and disease prevention on a global scale. That sounds like a technical ideal, but the reality of healthcare data, I mean, it must introduce just massive hurdles.
25:36Jeff Dean:You can't just throw all the world's patient data into a single cluster and train a model. The privacy and regulatory concerns are just immense. And that brings us right to the first major obstacle. The technical hurdle, which is rooted in privacy. Because of strict regulations like IMPA in the U.S. and GDPR internationally, data cannot be moved easily, if at all. And this mandates a fundamental shift in how machine learning is even performed. Right. So instead of bringing all the data to a central model, you have to bring the model to the data. This requires solutions like federated learning.
26:11Jeff Dean:Can you explain how that works in practice? Because the mechanism seems key to addressing these privacy concerns. Federated learning really flips the traditional centralized ML model on its head. In an applied health setting, there are generally four key steps. Okay. First, model initialization. Central server trains an initial sort of simple version of the model architecture. Second, distribution to silos. This initial model is securely distributed to the data silos, so individual hospitals, clinics, or regional health systems. And the raw patient data stays securely behind their local firewalls.
26:42Jeff Dean:Never leaves the hospital. Never leaves. Third, local training. Each hospital trains the model on its own local patient data. So the model learns the specific nuances of that particular population. And fourth, and this is the most important part, secure aggregation. The hospitals do not send the raw patient data back. What do they send? They only send the aggregated updates or the model parameters. So just the mathematical changes the model underwent. And they send that back to the central server, often using techniques like secure aggregation or differential privacy to mask contributions from any individual patient.
27:16The central server then averages these updates to create a more robust, generalized global model, which is then redistributed for another round of local training.
27:25Jeff Dean:The crucial point is that the sensitive patient data never, ever leaves the safety of the local jurisdiction. It provides learning at a massive scale while respecting those privacy boundaries. Precisely. It's a technical solution to a legal and ethical problem. But privacy isn't the only hurdle. The source calls the second big challenge the grungy hurdle. That sounds like the messy reality of data collection in the wild. It is the brutal reality of data heterogeneity. I mean, just think about the countless different ways medical records are kept. One system uses ICD-9 codes. Another one uses ICD-10.
28:00Some records are meticulously structured, while others are dominated by free text, unstructured doctor's notes.
28:07Jeff Dean:It's a complete mess. It's a mess. And getting this mountain of disparate, non-standardized data into a usable, structured form for training a global model is this immense, non-glamorous, but absolutely necessary undertaking. It requires a heavy investment in data cleaning and standardization tools. And finally, you have the ultimate complexity, which is the policy hurdle. The legal and regulatory complexities are incredibly high, and they change across every single jurisdiction. You have different consent laws in different states, let alone across national borders and international regulatory regimes.
28:41And this legal and policy landscape is constantly shifting, and it's often contradictory.
28:46Jeff Dean:So this is why you need that kind of lab structure. This complex environment is the exact justification for why the Law Institute structure is so vital. You cannot solve this problem with a small team of ML specialists. You need groups with mixed expertise, ML specialists, high-performance computer systems builders, and specialized legal, policy, and regulatory experts all working in tandem on a multi-year project to achieve any kind of real-world deployment. It sounds like the technical challenge is maybe even easier than the policy challenge in this context. Often, yes. The policy work, the negotiations, the regulatory mapping, the trust building, that can take far, far longer than the model training itself.
29:25Jeff Dean:So this incredible internal investment building, custom silicon, managing training jobs across multiple cities, tackling these global health challenges, it all raises a fundamental question about the boundaries of corporate research. How does a company with that level of internal advantage manage its relationship with the broader open academic ecosystem that it so fundamentally relies upon? This is a really complex balancing act, especially in the current highly competitive dynamic of the AI race. The chief scientist explains that publishing is not simply a binary choice. You know, we publish it or we don't.
30:00It's actually a nuanced, deliberate continuum based on product cycle maturity, competitive advantage and the nature of the research itself.
30:07Jeff Dean:OK, so let's break down the three defined points on this continuum, starting with type one, the earliest stage open research contribution. Right. Type 1 is open early stage research. And this involves publishing new earlier stage model architectures and ideas that haven't necessarily been proven out or experimented with at a massive scale. The purpose here is purely to contribute to the global research dialogue and to invite external validation and exploration. This is often high risk, high reward architectural exploration that really needs external partners to test and refine. Do we have a specific example of this kind of speculative early stage paper?
30:43We do. The source cites a paper that was published on a hybrid transformer and recurrent model called Titan. Now, the core problem that Titan is trying to solve is the computational complexity of the transformer's attention mechanism.
30:56Jeff Dean:The famous n-squared bottleneck. That's the one. It scales quadratically with the length of the input sequence, which means that making the context length longer is extremely expensive. Right. So Titan proposes a solution using recurrence relations, but instead of applying recurrence on individual tokens, it applies recurrence to chunks of tokens, say a block of 512 tokens at a time. The model learns to compress those 512 tokens into a summary or a memory vector, and then it uses recurrence steps on the sequence of those compressed chunks. So it's a hybrid approach. It is. It's aiming to achieve significantly longer context windows, the ability to reason over much larger documents with far greater efficiency than standard models.
31:39And the crucial detail here is that the chief scientist explicitly notes that Titan is not currently used in the Gemini models.
31:45Jeff Dean:Oh, interesting. But it is an interesting architectural idea for future exploration. And by publishing it, they're making a high value intellectual contribution back to the academic ecosystem. Okay. That clearly shows a commitment to the fundamental research ecosystem. Now let's move to type two, which sounds a lot more focused on commercial strategy. It is. Type two is delayed publishing, prioritizing the product advantage. In this very common scenario, the innovation is developed internally, it's rapidly deployed into a consumer product to gain a market advantage, and then the science detailing how it works is published much later.
32:20Jeff Dean:And the most salient examples here come from their deep work in computational photography. That's something that has just radically changed the capabilities of mobile phone cameras. It is the perfect illustration of applied ML innovation. I mean, think of features like Night Sight, which uses deep learning to synthesize multiple short exposures into this vivid single low light photo. Or astrophotography mode. Or astrophotography, or the popular magic eraser, which relies on these sophisticated semantic segmentation and in-painting models to just seamlessly erase a person or an object that wandered into your photo.
32:57Jeff Dean:That ability to make something feel like user-friendly magic while relying on this incredibly complex ML underneath, that's really the hallmark of great product research. And the process is very carefully timed. The feature is developed internally, it's rigorously tested, and then it's deployed in the next flagship phone, the Pixel N Plus One. They maintain the product lead for a set period, usually 6 to 12 months, and then they submit a SIGGRAPH paper detailing the precise underlying innovations. So they get the best of both worlds. It ensures the innovation eventually enters the public domain, and academia can build on it, but the company reaps the initial commercial advantage.
33:32It's a way of balancing competition and contribution pretty effectively.
33:36Jeff Dean:And finally, we have type 3, which is just the necessary reality of the current competitive dynamic, the secret sauce that has to stay internal. Given the scale of the investment, the custom TPUs, the pathway systems, the massive datasets, there are certain specific architectural details that just have to remain internal and proprietary. The source explicitly mentions that they do not publish the final secret architecture inside the massive Gemini model. In a global race for large model leadership, key differentiating secrets must be protected. And beyond just external publications, the sheer scale of the internal research effort, we're talking thousands of PhDs and engineers, that means they have to manage their own closed ecosystem of knowledge sharing just to accelerate the pace of invention.
34:22They do. They host an enormous internal research conference, which the source cites as having about 6 ,000 attendees annually. This is essentially an internal, massively scaled version of a conference like NeurIPS. Wow. And the observation cited by many PhD students is that papers presented at this internal conference often feel a year ahead of the work that's presented at external public conferences.
34:44Jeff Dean:Why the perceived gap? I mean, is the work genuinely a year ahead, or is it just a matter of presentation and maturity? It's primarily a matter of the bar for maturity. To get a paper accepted at a top-tier external conference like NeurIPS or ICML, The work has to be meticulous, fully proven out, and quite fully baked. And that requires months of polishing, validation, and writing. Which introduces a significant lag time into the whole process. Exactly. The internal conference, however, it accepts a whole range of maturity. They specifically utilize lightning sessions for cool, early-stage results that are just too fresh, too speculative, or too incomplete to be fully proven or written up in a formal paper.
35:25And this allows for a very fast circulation of ideas among internal colleagues without the pressure of external academic rigor or the fear of premature judgment.
35:34Jeff Dean:It creates a rapid-fire feedback loop for experimental ideas. A massive internal competitive advantage that just accelerates their R &D cycle. It really shows that even with a massive vertical stack and trillion parameter models, the ultimate engine of progress is still the human element. It's the creative freedom to share speculative ideas quickly and iterate without the paralysis of perfection. Hashtag tag outro. So we've taken the full journey today, really tracing the core pillars of modern AI innovation. We started with that fundamental crisis back in 2013, the realization that standard compute just couldn't scale the workload.
36:08Jeff Dean:And that led to the specialized hardware solution, the TPU, built for maximum energy efficiency and that vertical co-design. We followed that hardware all the way up to stack to Pathways, which manages the monumental feat of providing the illusion of a single system image across thousands of geographically dispersed chips. It turns a massive, messy network into a seamless computational fabric. And this vertical integration is what allows for the current unprecedented scale of large model training. And then we synthesized that high-tech vertical stack with its necessary foundation. The debt that is owed to public, academic research, TCPIP, RISC processors, PageRank, and the very invention of neural networks all developed 30 or 40 years ago.
36:49And that realization led us to the crucial discussion of the future, the need to shift our funding models away from chasing hype and toward these practical, achievable AI moonshots. We identified that three to five year time horizon as the optimal window for ambitious cross-disciplinary teams to tackle major societal challenges like health, addressing all the technical, the grungy and the policy hurdles along the way.
37:12Jeff Dean:So what does this all mean for the immediate future? Well, the next era of impactful AI isn't solely defined by who announces the next biggest ship or the largest model. It is defined by highly specialized hardware that's built for efficiency, combined with a concerted, collaborative, and properly funded effort to solve real-world problems on a practical three - to five-year timeline. And it means we have to be willing to invest in the unproven today, knowing that the real returns may not arrive for decades. It means we have to remember that today's competitive advantage is tomorrow's public foundation.
37:43Jeff Dean:And that leads to our final provocative thought for you, the learner, to consider. If the foundational ideas for today's planetary-scale AI things like neural networks and backpropagation were 30 to 40 years in the making, and the most impactful near-future projects are now being defined by a maximum five-year timeline, what current early-stage research idea, right now sitting in an internal lightning session or a small-scale academic paper, will truly define the world in 2030? What small seed is being planted today that will require a future hardware crisis to solve 15 years from now? Something to mull over until our next deep dive.
From the publisher
We summarize a recent interview with Jeff Dean, a legendary Chief Scientist at Google who has been leading Gemini, focusing on the **evolution and current state of Google's Tensor Processing Units (TPUs)**, including the recent seventh-generation announcement. Dean explains that the initial motivation for TPUs was Google's internal need to handle the massive compute requirements of scaling AI models, highlighting the **efficiency gains over CPUs and GPUs**. The conversation also shifts to the broader **AI ecosystem, emphasizing the critical need for vibrant academic research funding** as a foundation for major technological breakthroughs. Finally, Dean discusses **Google's strategy for sharing innovation** externally, such as a delayed publishing model for computational photography, and expresses **passion for applying AI to healthcare and improving compute efficiency**.




