In short
Eye On A.I. Podcast Episode Notes
Episode Title
#324 Sharon Zhou: Inside AMD's Plan to Build Self-Improving AI
Episode Description
In this episode, Sharon Zhou, VP of AI at AMD and former Stanford AI researcher, discusses how language models are optimizing their own hardware, specifically through GPU kernel code generation. The conversation covers self-improving AI, reinforcement learning, compute economics, and the implications for future AI infrastructure.
---
Key Themes and Concepts
Introduction
- Host: Craig S. Smith, former New York Times correspondent.
- Guest: Sharon Zhou, VP of AI at AMD.
- Focus: The evolution of AI technology, particularly in creating self-improving systems.
Background of Sharon Zhou
- Transitioned from being an AI researcher at Stanford to leading AI at AMD.
- Emphasizes the importance of access to compute for AI development.
What is Self-Improving AI?
- Definition: AI models that can edit their own components (data, architecture, and code) to enhance their performance.
- Application: Models writing their own GPU kernel code to optimize hardware performance.
Importance of GPU Kernels
- GPU Kernels Explained: Small pieces of software that perform specific tasks on GPUs.
- Optimization: Efficient kernel code can significantly reduce the cost and time of training and inference in AI models.
Techniques Discussed
- AI Agents and Evolutionary Strategies: Using AI to evolve kernel codes for better performance.
- Just-In-Time Optimization: Real-time kernel optimization for ongoing performance enhancements.
Key Challenges
- Catastrophic Forgetting: Occurs in continual learning where models lose previously learned information.
- Access to Data: Importance of having access to original pre-training data for effective post-training.
AMD's AI Strategy
- Integration of AI research into product development.
- Focus on creating optimized infrastructure for AI workloads.
Reinforcement Learning Beyond RLHF
- Exploration of reinforcement learning techniques beyond traditional human feedback.
- Use of verifiable rewards from GPU performance profiling to improve models autonomously.
Synthetic Data Generation
- Models generating their own training data to enhance learning efficiency.
- Importance of synthetic data in training non-frontier models.
Future of Compute Economics
- Discusses how kernel optimization could potentially reduce the demand for chips.
- Current industry demand for infinite computing resources remains high.
Discussion on Diverse Use Cases
- AMD’s focus on language models, diffusion models, and robotics.
- Collaboration with various institutions (e.g., Meta, Stanford) to advance kernel generation techniques.
Educational Initiatives
- Zhou is involved in teaching AI fundamentals to a wide audience.
- Emphasis on making AI accessible and understandable for non-experts.
---
Key Takeaways
- Self-improving AI can lead to faster and more efficient models through autonomous optimization of their own run-time code.
- Kernel optimization is crucial for reducing computing costs and enhancing model performance.
- Continuous learning and data access remain critical challenges in the evolution of AI systems.
- AMD is positioning itself at the forefront of AI infrastructure development by integrating advanced AI techniques into their hardware.
---
Episode Structure
- 00:00 - Preview and Intro
- 00:25 - Sharon Zhou's Background and Transition to AMD
- 02:00 - What Is Self-Improving AI?
- 04:16 - What Is a GPU Kernel and Why It Matters
- 07:01 - Using AI Agents and Evolutionary Strategies to Write Kernels
- 11:31 - Just-In-Time Optimization and Continual Learning
- 13:59 - Self-Improving AI at the Infrastructure Layer
- 16:15 - Synthetic Data and Models Generating Their Own Training Data
- 20:48 - AMD's AI Strategy: Research Meets Product
- 23:22 - Inside the NeurIPS Tutorial on AI-Generated Kernels
- 30:59 - Reinforcement Learning Beyond RLHF
- 39:09 - 10x Faster Kernels vs 10x More Compute
- 41:50 - Will Efficiency Reduce Chip Demand?
- 42:18 - Beyond Language Models: Diffusion, JEPA, and Robotics
- 45:34 - Educating the Next Generation of AI Builders
---
These notes summarize the key discussions and insights from the podcast episode, offering a comprehensive overview of the advancements in AI technology as discussed by Sharon Zhou.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOSharon Zhou's Background and Role
0:45 to 2:39
Sharon Zhou discusses her background in AI and her role at AMD.
“I taught there as adjunct in generative AI back before all this chat GPT stuff.”
Defining Self-Improving AI
2:39 to 3:46
Sharon provides an overview of what self-improving AI entails.
“So there's a lot of different pieces of work around kernel generation and being able to use LLMs to generate these kernels.”
Kernel Generation and Its Importance
3:46 to 6:17
Discussion on kernel generation using AI and its significance for performance.
“because what's really exciting about kernel generation and kernel development is actually we have the profiler.”
The Role of Kernels in AI
6:17 to 8:23
Explaining the function of kernels in AI models and their efficiency.
“So like what the GPU has as memory and storage, as well as the actual raw compute power.”
Evolution of Kernel Development
8:23 to 9:39
How AI is changing the way kernels are developed, moving from manual to automated.
“And these kernels in the past have been written manually.”
Continual Learning and Catastrophic Forgetting
9:39 to 11:31
Discussion on the challenges of continual learning and strategies to mitigate forgetting.
“And I would say that having those two areas of expertise in one person's head is quite rare.”
The Self-Improving Loop in AI
11:31 to 14:00
Sharon explains the self-improving loop for models to enhance their speed on GPUs.
“And when you say almost just in time, is this is being done on the fly or?”
Self-Improving Loops in AI Models
14:00 to 14:25
Learn how self-improving loops enhance AI model performance on AMD GPUs.
“So this is code that makes the language model itself run really fast on AMD GPUs.”
Deployment of AI Capabilities
14:25 to 15:01
Discover how AMD is planning to make self-improving AI accessible to users.
“And is that something then that you deploy with AMD hardware that people can use with their models?”
Editing Training Data for AI
15:01 to 15:28
Explore the significance of editing training data for AI model optimization.
“So we did a Nurex tutorial that presented a lot of this work, and we're preparing for ways to actually package this and make it possible for our customers to use it as well.”
Show all 27 chapters
Synthetic Data Generation Insights
15:28 to 16:18
Understand the role of synthetic data generation in AI model training.
“Yeah, so we're editing the kernel code for the model.”
Balancing Automation and Autonomy in AI
16:18 to 18:40
Learn about the balance between automation and autonomous self-improvement in AI.
“And there's been a lot of work and research and speculation on it.”
AMD's Role in AI Hardware and Research
18:40 to 21:20
Discover how AMD integrates AI research with hardware development.
“There are pieces that you can launch that feel more autonomous, where you could have it like reflect on its own output and then improve that.”
NeurIPS Presentation Overview
21:20 to 23:38
Get insights into a collaborative presentation from the NeurIPS conference.
“And we tried to connect at NeurIPS, and I'm sorry that we weren't able to.”
Profiling and Optimization Techniques
23:38 to 24:40
Learn about profiling techniques used for optimizing AI model performance.
“that we're using and sharing a lot of that.”
Reinforcement Learning in AI
24:40 to 28:00
Explore the impact of reinforcement learning on AI's self-improvement strategies.
“Kind of like when Chachapiti, they were training on math tasks.”
Exploring Reinforcement Learning in AI
28:00 to 29:13
Learn about the advancements in reinforcement learning for post-training of AI models.
“Are there things other than kernel optimization that you're working on?”
Human Feedback and Reasoning in AI
29:13 to 31:18
Discover how human feedback enhances AI reasoning and model safety.
“And what reinforcement learning really does is a couple of things.”
Kernel Generation and Optimization Strategies
31:18 to 33:19
Understand the role of kernel generation and optimization in AI performance.
“um from the profiler of of our gpus that tell us how fast these kernels are actually running on the GPU and that's a number and that's verifiable and it's true.”
The Complexity of Kernel Development
33:19 to 36:16
Learn about the challenges and time savings associated with kernel writing.
“Oh, there's so many parts of the stack that need code assistance.”
AMD’s Focus on Specialized Kernel Optimization
36:16 to 38:27
Explore AMD's approach to optimizing kernels specifically for their hardware.
“so that they could get far more out of their compute.”
AI Use Cases Beyond Language Models
38:27 to 41:25
Examine the various applications of AI in different fields beyond language models.
“But yeah, but today we're definitely very focused on making sure that these kernels can run very effectively on AMD GPUs.”
Collaboration with AI Startups and Educational Initiatives
41:25 to 42:00
Discover AMD's collaboration with AI startups and educational platforms to advance AI.
“So like the, you know, latest type of diffusion models, for example, like those run smoothly on AMD.”
AMD's Role in Robotic Foundation Models
42:00 to 43:04
Learn how AMD collaborates with startups and utilizes open-source models.
“And are you working with robotic foundation models at all?”
Teaching AI Fundamentals Online
43:04 to 44:14
Discover the AI courses offered online for various professionals.
“And so I think that's, you know, that's something that's important.”
Bridging the AI Knowledge Gap for Non-Coders
44:14 to 45:52
Understand how to enable non-coders to leverage AI in their work.
“And so these are things that are not like you're not these people don't.”
Accessing AI Courses and Resources
45:52 to 46:25
Find out where to access various AI courses and their availability.
“And so, so yeah, that's, that's where the other kind of stuff I'm working on as well that might be helpful to your audience.”
Transcript
Automatic transcript. May contain errors.0:00Catastrophic forgetting is definitely a problem, especially in post-training and when you don't have access to the original pre-training data.
0:08Sharon Zhou:How much of that is focused on improving the design of AMD hardware and how much of it is focused on general model development, which AMD is not doing, right? AMD does not develop its own models. I think people want infinite chips, Craig, so no, it doesn't relieve the pressure. I am Sharon. I'm the VP of AI at AMD. And I think about self-improving AI, self-improving LLMs, which we'll get into later. But my background comes from AI research. So I used to be an AI researcher at Stanford, where I did my PhD with Andrew Ng. I taught there as adjunct in generative AI back before all this chat GPT stuff.
0:52And after Stanford, I started a startup, an AI infrastructure for startup doing post training of language models actually on AMD GPUs. This was started a couple of months before chat GPT launched. And most recently over the last several months, we have transitioned now to AMD. So my team and I are there now and very excited to enable more people to use compute and to get access to compute because that really is the limiting factor. one of the big limiting factors for developing AI and being able to enable more people to steer these models. So that's what I'm really excited about. And yeah, that's why I'm here.
1:36Sharon Zhou:Yeah. And I do want to talk about self-improving AI. Can you start by defining what we're talking about when we talk about self-improvement? Are we talking about models that rewrite their own code or something like refining their own training data? Yeah, I think that's exactly it. It is a broad category, but essentially it's the idea of these models being able to edit any part of themselves to improve themselves. So whether that be the data, whether that be the actual model architecture, whether that be how they evaluate themselves. And actually the part that I'm working on is below all of that, and it's actually how fast they actually run on the GPUs themselves.
2:23So they are writing the kernel code that underlies these models to run faster on these GPUs and to run effectively on them. Yeah. And on new hardware, too. So that's been really exciting to see. Yeah.
2:38Sharon Zhou:I just read about kernel evolve. Yes. Who is doing kernel evolve? I've forgotten. Is that Google? Meta. Is that your guy's work? So there's a lot of different pieces of work around kernel generation and being able to use LLMs to generate these kernels. So we're doing some of that, but I think that work might have been from Meta, but there's some work across. Yeah, I think there's work across the board, across the industry that is very important, I think, towards this end because it enables more people to get on different types of compute. What we did most recently was in collaboration with actually a bunch of different institutions like Meta, Google DeepMind, ML Common, Stanford, NVIDIA, et cetera, was a NeurIPS tutorial on generating kernels using AI.
3:34And so we presented that and basically goes through how we're using AI agents to generate these kernels and how we're thinking about post-training these models to generate kernels more effectively. because what's really exciting about kernel generation and kernel development is actually we have the profiler. So we have the ability to actually see how fast the generated kernels are on the chip themselves. And so that's really exciting. My team, we're also working on a kind of more robust production level benchmark to share with the community as well as different techniques to modify the models to do better on this task.
4:16Sharon Zhou:Okay. Yeah. And for, again, I'm trying to make my podcast a little more accessible. Oh, yeah, of course. So kernel in this sense, because kernel is used in... Everywhere. Too many places. But kernel in this sense, you're talking about a small piece of software that lives on the processing unit that performs a specific task. Is that right? And what are the kernels that you're generating from AI? So there are many layers of the stack for AI. there is the model layer where we kind of build out the model architecture. This is where people talk about transformers and attention and kind of write that out in PyTorch or Jax or TensorFlow.
5:13And so people are like building out models there and then they're using maybe hugging face on top or they're using different tools to leverage those models. Now underneath all of that are ways where the model is running on the GPU, and there are many, many layers, but one of the layers is making those models run really fast on the GPU. And sometimes you can break that down into small different pieces because it's a lot of matrix multiplications, for example. It's a lot of different operations happening on the GPU. And so for each of those operations, you can have a way of basically optimizing the speed of that on the GPU itself.
5:56And so for a given piece of hardware, it may have been designed to be of a certain speed, but you do have to write the software to connect the models of today and of tomorrow, ideally, to that hardware. And that's kind of what that layer is doing. It's improving the efficiency. So it's utilizing the GPU effectively, both the memory. So like what the GPU has as memory and storage, as well as the actual raw compute power. So being able to schedule all of that and use that effectively and maximally is what kernels do. And this is really important because it's expensive, right? Like GPUs over a long time and over just a large capacity when you're parallelizing, it becomes expensive.
6:44And so you want to be able to eke out as much as you can on each GPU. Yeah.
6:51Sharon Zhou:And I understand how. And you're using evolutionary strategies to develop these kernels? Yes. I would say it's a combination of evolutionary strategies, agentic strategies, as well as different types of post-training strategies to be able to get these models to actually be able to write that code effectively, or at least assist our internal kernel engineers to do so. And I think one of the most important things is to do so on useful kernels as well, the most important kernels for language models. Yeah. Well, again, for listeners, what are the most useful? What's what's an example? Well, I would say matrix multiplication.
7:36So when you're doing, when you're learning about the math of a neural network, you'll see that there's a lot of matrix multiplications and a lot of those need to be optimized. And, you know, at first you might think, oh, isn't there just one, you know, two matrices multiplying, but actually no, you can optimize different size matrices that multiply together and you can get like even better performance for something of a different size. Like you can optimize that. So there are thousands, possibly hundreds of thousands of, of those kernels that you could optimize, um, for, for language models. And of course, for a particular, um, customer, particular, like foundation model company, um, they may have a certain set that they particularly care about, um, for their architecture.
8:20And, um, those are ones that are, are high priority. Yeah.
8:25Sharon Zhou:And these kernels in the past have been written manually. Is that right? So having there's a tremendous productivity gain and having AI either do it autonomously or do it as an assistant. What sort of impact does that have on AMD? I think it's having larger and larger impact. So I would say that, yeah, so historically, they've been written manually. And just to give you a sense of how much knowledge this kernel engineer needs to have in their head, they need to know about the GPU architecture. So not just general GPU architecture, not just your average architecture class, but like that generation of the GPU and what changes have been made, you know, what is actually, you know, how to actually write the code to run the software on it and all the different, you know, possible new things that have been invented at the hardware layer.
9:25They have to understand all those things. And then they also need to understand at least enough about what this matrix multiplication is doing, right? And to actually then write that optimally in the code. And I would say that having those two areas of expertise in one person's head is quite rare. And as a result, it's very, very valuable, but also it is bottlenecked at a lot of hardware companies or at actually a lot of different companies, including Frontier Labs. They also have people writing kernels to speed up their models as well, since they have more visibility on their own models that may not be fully visible to everyone else.
10:05And so I think it's a rare skill. And if you're listening, I encourage you to go learn about it. But of course, we're also teaching the models about this. Right. And so what's exciting is for AMD, we have the Rockham stack, which is open source. That's the equivalent of NVIDIA's CUDA stack, which is not as open. And so that being open actually helps us from a language model perspective because the models can kind of read that data, right? And train from that data and use that data to then learn about the Rockham stack and then write those kernels for us. And so that's what we're doing here. And that's the impact is massive from both a, I would say, like direct customer lens of what it could do to improve that.
10:50But also, I would say, and the long tail. So there are researchers working on various different models that are not just Chachi PT or, you know, just Grok or just, you know, just a few of these frontier models. But they're working on a bunch of different other models and they're creating, they're inventing the next model. Right. And those also need to be optimized. And I think having something that is almost just in time, ideally just in time, but, you know, like almost automatically give you a optimized kernel for is very essential to enabling just the entire ecosystem to be able to move over.
11:31Yeah.
11:31Sharon Zhou:And when you say almost just in time, is this is being done on the fly or? That is my goal, Craig. That is my goal. Yeah, because you talk about self-improving AI, that suggests that there are models that are evolving over time or improving over time. And that leads to the question of continual learning, which is kind of the holy grail. but there's catastrophic forgetting and expandability of neural nets and all those problems. Can you talk about where the research stands on continual learning? Yeah. So catastrophic forgetting is definitely a problem, especially in post-training and when you don't have access to the original pre-training data.
12:32Because I think we found in the literature at least that even if you include, originally it was like if you include 20%, but actually if you include only 1 % of the pre-training data back in during post-training, you can actually prevent catastrophic forgetting, basically enabling the model to actually like connect back to its representations back in pre-training, or at least help significantly for that. And I think it, again, depends on data access in terms of who is doing some of that post-training. But I think that's what the literature points to. So it's the extent to which you can access some of that data or use some of those online data sets to just bring back some of those data points from retraining that could actually help your model significantly in preventing catastrophic something like catastrophic forgetting.
13:24But of course, I think this is something that needs to be continually monitored. So I think the frontier labs maybe don't have this problem as much, But as an individual who is probably doing some type of post training, that is something that you do have to consider, especially as you get your especially as your workload gets heavier and heavier. Small little bits of fine tuning that doesn't really change as much.
13:46Sharon Zhou:But so self-improving, are you? Yeah, define self-improving for me. I mean, what are you improving in or attempting to improve in a model? So the part of self-improving that we've been working on is using these language models to write low-level kernel code. So this is code that makes the language model itself run really fast on AMD GPUs. So get the model to actually write its own code to run even faster on the GPU so that it can learn even faster. And so this is the self-improving loop that we're enabling the models to do here. That's really exciting because then that'll enable more people to access this compute and it'll enable more powerful models within a shorter amount of time.
14:33Sharon Zhou:And is that something then that you deploy with AMD hardware that people can use with their models? Or is it a separate model that you make available to users of the hardware? Yeah, this is such a good question. Right now, it's being used internally, and we're releasing different data sets and evaluation benchmarks and methods out into the world. So we did a Nurex tutorial that presented a lot of this work, and we're preparing for ways to actually package this and make it possible for our customers to use it as well. Yeah. And then the idea of editing a model, editing its own training data, or why would you want to do that?
15:27Sharon Zhou:And can you give me an example? And how is that being done? Yeah, so we're editing the kernel code for the model. In terms of editing training data, what that looks like is usually this is like a synthetic data setup. So the model is generating data for the next round of training. So yeah, self-improving AI is more general than just having the models generate kernels and write these low-level kernels to make the models really fast on GPUs. It could also touch on generating data for itself, as well as generating evaluation for itself, and also generating different architectures for itself as well.
16:12So basically using the models to improve itself at any stage of the pipeline. I think data generation is really interesting, just synthetic data generation overall. And there's been a lot of work and research and speculation on it. And I think, you know, we have found that using synthetic data to train these models is very helpful, especially if it's actually not a frontier model and you're distilling from a frontier model. It's not quite distillation, sorry, but it is using synthetic samples from a frontier model to then supervise a smaller model that you might be training. That's very effective.
16:48And I think like there's also debate around synthetic data generation and how much that may or may not lead to the collapse as the model trains since it's looking at its own kind of data. But I do think it is an important space to continue monitoring and to continue understanding. Yeah, so I think what's really interesting is this holy grail that everyone's looking towards of having the model fully improve itself and create its own next generation. I think we're both further than we think and closer than we think. And so I think we're further than we think in the sense that there is still a lot of human expertise that goes into how these things are laid out and what we should try next, for example, of what we should add into the model, whether it be new data source or not.
17:43But I do think there's more than we think as well in terms of the use of AI to help us code in general. So the amount of code that AI is writing is pretty enormous, and it is supporting a lot of engineers that are putting these together and creating the next generation of that AI. So I think it's both we're further and closer than we think. Basically, we're further than we think from the like ultimate like okay there's only one person works at open ai now it's only sam there and it's just the end with chat gbt versus um uh we're we're closer than we think because actually every everyone is actually using um uh chat gbt quite significantly um to write that code so i i think uh it's it's in in in that state right now um but it's it's not exactly like autonomously improving itself fully.
18:40There are pieces that you can launch that feel more autonomous, where you could have it like reflect on its own output and then improve that. It can look at the profiler information so it can look at how fast a kernel ran on the GPU and improve that. But it's not quite like, I'm going to let this go and come back in a month and it will be fully solved.
19:04Sharon Zhou:yeah uh yeah well let's can we back up a little bit uh as i said at the beginning amd's a hardware company you're uh in charge of ai at uh at amd part of it yeah part of it yeah is is that uh uh is is that a research function um because uh and and how much of that is focused on improving the the design of of amd hardware and how much of it is focused on uh on general model development uh which which amd is not doing right amd does not develop its own models oh we have actually uh pre-trained some models and um released them but they are small in the um single digit billion parameter range um and yeah we're probably like not developing models in the sense of like not chat gpt um but uh but yeah we we do have some um that are uh open source models um yeah so i focus on i almost call it like product research i think there is a research element to it that we don't know how good um the models are at solving this full task and end um but there is a product component in that this is part of the loop with what becomes customer facing right so then this becomes part of the loop of what we um want to make available to our customers and so uh it's it's shippable and there is a production component to it.
20:43And so I would say it kind of sits at that intersection, which I know a lot of things today sit at. So, yeah, I think that's one of my teams works on that. And the other team works on, I would say like research and education. So we're thinking about how to stay at the forefront of research and engage with researchers everywhere, but also educate the world and evangelize, you know, educate the world about AI and evangelize AMD GPUs that are able to support all of those new AI workloads. And so, yeah, that's a couple teams that I've been very grateful to be leading. Yeah.
21:27Sharon Zhou:And we tried to connect at NeurIPS, and I'm sorry that we weren't able to. And I missed the talk. can you go through what the presentation that you gave at NeurIPS? Yes. So it was a tutorial, two and a half hours. And my team did a huge part of the presentation. So definitely not just me here. And it was a collaboration across a lot of different institutions. So like, you know, Stanford, Google DeepMind, Meta, NVIDIA, ARM, et cetera. And so it was a really, really great group of people. And what we presented in our tutorial was a few different things. First, because it's a NERVUPS audience, AR researchers, they might not be actually as familiar on GPUs and hardware and how the GPU actually works.
22:19And so we kind of stepped through in hopefully very easy to understand ways, the actual GPU architecture and what pieces that an AI researcher might find interesting, like the memory and like how things are scheduled in. Like something I find interesting is, you know, when you change your batch size, why suddenly like your your job takes forever, right? Your training job takes forever. Oh, it's because of, you know, X, Y, Z happening in in the actual GPU. you. And so that's what we go through first. And then we go through kind of different kernels. So like showing you what a kernel actually is in code in a very simple kernel, and then showing you what it looks like when it's optimized and unoptimized.
23:04So you can understand like, oh, when a kernel runs fast or slow, what it looks like. So if something runs really slowly, that means like ChatGPG could answer you in like minutes as opposed to seconds, right? So that's ultimately what it feels like. But then when you go down into like just a single matrix multiplication, what does that look like? Or a single addition, you know, like X plus Y. What does that look like in terms of time and how do you profile it? And then we go into, okay, what are AI agents and methods that we're using for these self-improving AI essentially? But essentially what are some of these methods that we're using and sharing a lot of that.
23:44So like what kind of things can you do with agents? You can have them, like you can write me a more optimized kernel. Okay, it does it. It maybe is or isn't optimized. You have to check the correctness of it. Maybe it's cheating. And then when you profile it, you can actually get the answer back. And you're like, actually, it was only 1.2x faster. I want it to be even faster. and then you can continually evolve it. That was based on Google's alpha evolve paper, but essentially like continually improve it after multiple calls. The profiling information also provides a really interesting way of collecting or creating an RL environment so that these models can learn from essentially a verifiable reward from the profiler itself.
24:32So it's able to give this number back as, okay, this is how fast your kernel was that you generated and tells the model in a very verifiable way. Kind of like when Chachapiti, they were training on math tasks. It was very verifiable whether the math was correct or not. And this helps the model improve in post-training. And so, sorry, I'm trying to make this accessible, but like, so basically we go through all of those different techniques as well and we share that with the community. Yeah.
25:04Sharon Zhou:And this is done at what point in the process? Someone, a customer is using AMD hardware and they have a model that they want to run on it. And so they're optimizing the kernel before deploying the model. Or is it after deployment and they want to improve it? Yeah. So I think it's a combination of those things. So I think you could use it definitely before you deploy something. So before you deploy something, you're like, okay, this isn't fast enough given my budget with this many GPUs and we need it to hit these markers. Right. And so because that basically translates to TCO translates to the total cost of ownership.
25:50And so let me reduce the speed of all, sorry, yeah, let me make all of these faster inside of the model so that the full model runs really quickly. So that could be before deploying something. It could also happen afterwards, too. A lot of different kernels are written afterwards as well to just speed things up even further. So it could be something you've already deployed, and now you want to speed things up to make use of the hardware you have even more. So I think it could occur in both of those settings.
26:21Sharon Zhou:Yeah. And this is, in my mind, this is more automation than self-improvement, right? You're automating the kernel writing process, and that improves the model. But where does self-improvement come in? Yeah, so the self-improvement is, so I view it as self-improvement, not automation, but I view it as self-improvement because the kernel itself is what the model is running on. So it is the model's own code and it enables the model to run even faster on the GPU. But I guess when broken up, it could be seen as just like, oh, automation. And I do think, I honestly think everything is technically automation.
27:11But if we want to view it through a lens that's like more exciting, where it feels like autonomous, that is just a different perspective on it. Yeah.
Read the full transcript
27:20Sharon Zhou:But can the day come where a model will autonomously go in and rewrite its kernel to speed things up? Yeah, I definitely think so. I think more likely it will be a different model that does it. um yeah i mean an external model other than the inference model is that what you mean yeah or i mean it still could be that model um but uh i guess it depends on how you view what that model is whether it's the same instance of that model or not um or a different prompt um but yeah it will it will be optimized whether by that own model or by another model yeah Are there things other than kernel optimization that you're working on?
28:09Sharon Zhou:I mean, you mentioned a few things. Yeah. So another area has been RL research. So reinforcement learning inside of post-training. And so that's been really exciting to track and to work on. It is related to kernel generation since a lot of those techniques can be ported over. but more generally, we've been exploring those internally for different use cases, as well as for ways to actually make sure that all of this does run on AMD's hardware. The type of RL research that we're looking at includes kind of the techniques that started with chat GPT. So the techniques that got us ChatGPT included something called RLHF, RL with reinforcement learning from human feedback, and that used an algorithm called PPO.
29:00And so it used a certain type of algorithm that now has, I think, found successors, as well as, you know, complements to like GRPO, which came from DeepSeek earlier last year. But basically, there's a huge set of research in this area to more effectively use reinforcement learning in the post training of these language models. And what reinforcement learning really does is a couple of things. One area that's really been exciting is that human feedback piece. Right. So that you could actually take human preferences, whether it be like rankings of things or, you know, pairwise comparisons of things of your own preferences.
29:40and you just give a lot of examples of those, you can actually teach the model to mimic those preferences. And those preferences could turn into a more helpful model. It could turn into a safer model so that it doesn't like say harmful things. And so, yeah, and I think another, you know, application of reinforcement learning has been around reasoning. So if you're familiar with thinking in all these models that you're using. That's what reasoning is effectively doing, using far more tokens. The model is basically writing down its thoughts before it responds to you. And that process is done using post-training, definitely part of it being RL.
30:26And I would say like that extended to, let's say, math or coding tasks has become really exciting because they're verifiable, is what we say in the research world. And what that means is that these tasks, you can verify like a math proof or you can verify the code. And that becomes really interesting because then you don't need to collect that human feedback, which might take time. And it might be, it might be brittle because you can't cover every possible use case. But instead, you have something that checks it that can cover everything and that provides feedback back to the model and that's called you know rl with verifiable rewards instead of rl hf from human feedback um and so that's been a really exciting direction as well um and and been a direction that we've been looking into a lot because um we were getting these models to write code um and we have verifiable rewards um from the profiler of of our gpus that tell us how fast these kernels are actually running on the GPU and that's a number and that's verifiable and it's true.
31:29It's not subjective. And so that can easily go back into the model as it's learning. So yeah, it's a lot of different algorithms going in that direction to improve that process and make it more effective.
31:43Sharon Zhou:And how the kernel writing takes place at the customer, on the customer side, I mean, when someone is deploying a model onto AMD hardware or any hardware and they want to optimize the kernel, or is it taking place, do you do this work at AMD and then make those optimized kernels available to people depending on the model that they're running? Yeah. It's the latter right now. So we are basically doing this a lot of this internally and then making it available open source. All our kernels are open and we're that's what we're doing. I think the hope is to actually get it out there. Hopefully sometime this year we'll see.
32:35But the goal is to get it out there and make it possible for our customers or users or any user really to be able to leverage that very easily. Now, today, I think what's really effective is you can go into Cursor or you can go into any of these AI coding agent environments and you can do some of this on your own as well. And there are companies doing this as well. Yeah.
33:03Sharon Zhou:And you're presumably also working on code assistance for use within AMD beyond writing kernels. Are there other parts of the stack that you're focused on? Oh, there's so many parts of the stack that need code assistance. So I mainly focus on kernels and the actual models, but I think there's so many layers of the stack, like we have to integrate with open source libraries so that things just work by default on AMD when a developer in one of those higher level stacks is using us. um so there's a lot of um code at so many different layers of the stack um which um i think i honestly failed to appreciate before joining amd um like that there are so many layers um even underneath the like shiny you know model layer that everyone knows about yeah but you said you do work on models as well yes it's for the kernel generation right right yeah uh and and can you give me a sense of how uh how much time is saved or how many man hours are saved by having a kernel generation software like this um i probably can't give you direct stats um but what i can share is that that these kernels could take a very long time to write.
34:42A very, very complex kernel can take a, at least, you know, let's say a non-expert, but someone who might still be tasked to do it, months to write. Like an expert would take a couple weeks, but that's still a substantial amount of time. It's not like it's a, you know, a couple minutes, a couple hours, right, type of task. So it is a more complex task than throwing up a website and certainly more niche and therefore not as represented in pre-training data so that the models don't just as by default know everything about it. But it's a highly valuable task and one that is both rare and there's like not enough expertise in any company, I think, to do it.
35:26And very just necessary and urgent today, because if you can even shave off like if you shave off like a tiny amount of time it takes for one matrix multiplication, that occurs billions, trillions of times inside of a model. And that can incur billions, hundreds of billions of dollars to a company. And so that's a lot of money if you're doing a frontier model, of course. But even if you're not, it could save a lot of money. And I think what people don't realize is behind a lot of these APIs, for example, like the Together AI or Fireworks AI or Base 10, they are writing kernels, actually, to make these models run even faster than if you were to try to do it on your own.
36:08And they're writing a lot of those, right? Make them faster. So we're doing that. OpenAI, Frontier Labs are doing that. A lot of people are doing that so that they could get far more out of their compute. Because the equivalent is if you can 10x the speed of your kernel, that's the equivalent of buying 10x more compute. And as we've seen, like buying 10x more compute is interesting for the scaling laws, for scaling these models to get to the next level of intelligence. So this is, I think, actually a very ripe layer to continually optimize and actually be able to get scaling even from optimizing this existing layer.
36:45Sharon Zhou:That's interesting. Yeah, it's something that people don't think about. Yeah, I didn't think about before either. Yeah, and the kernels that you're writing, are they optimized for AMD hardware or could you then take them for other hardware? Yeah, excellent question. So it is optimized for AMD hardware. And that's because our hardware has specific things in it so that we need to optimize the kernel to run, specifically use all of our HBM, for example. We have higher HBMs, more memory than the other GPU. And as a result, how do we utilize that effectively, right, for certain tasks? And so I think that's why we need to be very specific to this GPU and what we're eking out.
37:44I think there are people working on kind of more like general kernels. For example, Flash Attention was a famous one. It basically takes the usual attention mechanism inside of all transformers and they do some math to make it so that scheduling it on the GPU doesn't require as many trips to grab from memory, from storage. and as a result speeds things up significantly. And I think that is an algorithm, a kernel algorithm that has now been used across the different hardware providers. So that's like more of a general invention, which I think kernel generation could get to, which is very exciting.
38:31But yeah, but today we're definitely very focused on making sure that these kernels can run very effectively on AMD GPUs.
38:38Sharon Zhou:Yeah. You know, there's been a lot written recently about overcapacity of or looming overcapacity of hardware, both chips and data centers. uh and as you speed up uh the compute uh using these strategies like kernel optimization uh is it is it going to relieve some of that pressure right now uh for uh chips um your question yeah yeah i mean do you have any sense of that impact the macro uh uh landscape i think people want infinite chips craig so no it doesn't relieve the pressure i don't think they found a plateau where they are like we're we're done like if you just 100x this we're done i i don't see that right now.
39:44Sharon Zhou:Yeah, that's right. And are AMD chips, are you focused on a particular use case? So I guess the use case is language models. What's, I guess, crazy is that that's a vertical because people are like, what else could there be? So what else could there be? So there could be, you know, computer vision models. There could be these GPUs. Now they're all optimized for AI, but before they were originally developed for high performance computing, HPC, which includes like weather modeling and prediction, right? So like a lot of different other tasks. And even within or like adjacent to language models, there's a bunch of other types of models as well.
40:32So I'd say that we're very focused on, or I'm very focused on the AI side of things and the language model thrust and kind of what people are developing there. What kind of diverse, so it includes for sure what the Frontier Labs are doing, but also what leading startups are doing. And some of those startups, like I just met with Yann LeCun regarding me and they're developing JEPA and that might look a little bit different in certain ways. I might need, it might not care about that like low latency token inference that, you know, autoregressive language models like Chat2PT care about, but something else, right?
41:11And so that might look different from a hardware perspective. And so we also care about leading AI startups and what they're doing and making sure that all of those workloads do work very effectively on our chips and that we're very knowledgeable about them. So like the, you know, latest type of diffusion models, for example, like those run smoothly on AMD. And so I would say that like, that's something that's really important. So that's kind of where the focus is. And I know it doesn't feel like a focus because everyone's focused on it, but it actually is a focus because GPU is a fairly general type of compute.
41:45Yeah.
41:46Sharon Zhou:Yeah, that's interesting. Yanlakun's JEPA. So you were talking with him about optimizing kernels to run JEPA on AMD hardware? uh we were talking more generally than that probably um but um but understanding what jet bud does is really important for us to optimize yeah we can yeah we can take some of their um models today and be able to actually run it on on amd hardware uh a lot of its open source yeah yeah i just uh had sergey levine or levin he pronounces it on uh just before that was Yeah, hi, right? And are you working with robotic foundation models at all? AMD definitely is. AMD definitely is.
42:39I'm personally, I'm not working directly with some of those startups, though I've chatted with some of them. But we have folks very focused on that. So I teach online. I teach about a million people online, including developers, but also executives and different professionals. And lately, we have a part, we built up a partnership with Deep Learning AI and During's platform to teach different courses. And we launched a post-training class on, you know, reinforcement learning and fine tuning of these models. But actually the first module I think is accessible to everyone and enables everyone to just like double click and get one level deeper on how these models actually learn how to behave, how to chat with us, how to act safely and add guardrails, how to hallucinate a little bit less and how to stay more focused in a long conversation.
43:31And so I think that's, you know, that's something that's important. And I'm actually working on a book that matches that class, too. And then the other courses that I've been preparing kind of jointly between AMD and deep learning AI also are around like why compute matters and a more general course on transformers and understanding that. Before a more general audience, I've actually been working and spending time at Harvard University to produce courses as well. So this is a collaboration where I'm teaching a few fundamentals around AI for a professional audience. So that's anything ranging from vibe coding.
44:16So basically enabling more people to write code, especially people who can't write code today, and be able to build things that might help them in their professional careers, whether that be building a board deck or like analyzing, you know, something for M &A. And so these are things that are not like you're not these people don't. If you don't view yourself as a software engineer, this is like the right class for you. And then, of course, just generally like AI for leaders and leadership and any kind of like manager who's thinking about how to change behavior within their organization with AI and understand things a little bit more fundamentally and how to use AI that's more practical.
44:58For example, if the models are producing different outputs and you find that frustrating and you find it not trustworthy, maybe we can flip the script a little bit and think about it from the other side of, okay, actually, this is why they act this way. This is why they produce variation. And now that you understand why, you can actually take advantage of this and you can run the model multiple times or run multiple models and be able to then take that analysis with you and use that in your day-to-day more as a tool. And so it's doing that, but like 10, 20 times in examples and being able to just disseminate that knowledge that I know I have as an AI researcher, but want more people to have when they're using this technology so that they can view it more as a tool that's useful to them, other than like something that they're a little bit worried about and don't trust.
45:53And so, so yeah, that's, that's where the other kind of stuff I'm working on as well that might be helpful to your audience.
46:00Sharon Zhou:Yeah. Where, where would people access those, those courses? Yeah. So the deep learning AI courses are available for free on deep learning AI. If, if you want a certificate, I think you do have the pay, but without that, they're free um and then uh the harvard ones have not been released yet but um we've been working on them so wow that's exciting yeah yeah
From the publisher
AI is not just getting smarter. It is getting faster by learning how to optimize the hardware it runs on.
In this episode, Sharon Zhou, VP of AI at AMD and former Stanford AI researcher, explains how language models are beginning to write and optimize their own GPU kernel code. We explore what self improving AI actually means, how reinforcement learning is used in post training, and why kernel optimization could be one of the most overlooked scaling levers in modern AI.
Sharon breaks down how GPU efficiency impacts the cost of training and inference, why catastrophic forgetting remains a challenge in continual learning, and how verifiable rewards from hardware profiling can help models improve themselves. The conversation also dives into compute economics, synthetic data, RLHF, and why infrastructure may define the next phase of AI progress.
If you want to understand where AI scaling is really happening beyond bigger models and more data, this episode goes under the hood.
Stay Updated:
Craig Smith on X: https://x.com/craigss
Eye on A.I. on X: https://x.com/EyeOn_AI
(00:00) Preview and Intro
(00:25) Sharon Zhou's Background and Transition to AMD
(02:00) What Is Self-Improving AI?
(04:16) What Is a GPU Kernel and Why It Matters
(07:01) Using AI Agents and Evolutionary Strategies to Write Kernels
(11:31) Just-In-Time Optimization and Continual Learning
(13:59) Self-Improving AI at the Infrastructure Layer
(16:15) Synthetic Data and Models Generating Their Own Training Data
(20:48) AMD's AI Strategy: Research Meets Product
(23:22) Inside the NeurIPS Tutorial on AI-Generated Kernels
(30:59) Reinforcement Learning Beyond RLHF
(39:09) 10x Faster Kernels vs 10x More Compute
(41:50) Will Efficiency Reduce Chip Demand?
(42:18) Beyond Language Models: Diffusion, JEPA, and Robotics
(45:34) Educating the Next Generation of AI Builders




