In short
The MAD Podcast with Matt Turck: Episode Summary
Episode Title
The End of GPU Scaling? Compute & The Agent Era
Guests
Tim Dettmers (AI2) & Dan Fu (Together AI)
Episode Overview In this episode, Matt Turck hosts a discussion between Tim Dettmers and Dan Fu about the future of Artificial General Intelligence (AGI) and the evolution of computational hardware. The conversation centers on two contrasting viewpoints regarding AGI: Tim's argument that AGI is constrained by physical realities, and Dan's belief that we are underutilizing existing hardware and models.
---
Key Themes and Discussions
- The AGI Debate
- Tim Dettmers: Argues that physical limitations, such as the von Neumann bottleneck and diminishing returns on compute resources, will hinder progress towards AGI.
- Dan Fu: Posits that there is significant untapped potential in current models and hardware utilization, suggesting that today's models are lagging indicators of hardware capabilities.
- Understanding AGI
- AGI is defined variably within the community, often lacking a concrete definition.
- The discussion touches on the cognitive versus economic implications of AGI, emphasizing the need for practical applications.
- Progress and Hardware Limitations
- Tim's Argument: Claims that current GPU technology has peaked, with no meaningful advancements expected in the near future.
- Detailed the impact of memory movement on computational efficiency.
- Dan's Counterpoint: Highlights that existing models are underutilized and that new hardware advancements (like the latest GPUs) could yield significant performance improvements.
- The Role of Agents
- Usage of Agents: Both guests stress the importance of AI agents in automating tasks, with Dan suggesting agents have surpassed significant milestones in tasks like coding.
- The analogy of managing interns is drawn to illustrate how to effectively leverage agents for productivity.
- The Future of AI
- Predictions for 2026 include:
- A shift towards more efficient, specialized smaller models.
- Continued hardware diversification with a focus on optimizing inference efficiency.
- Exploration of architectures beyond classic Transformers, including state-space models.
---
Practical Takeaways
Actionable Insights for Listeners
- Utilization of Agents: Listeners are encouraged to adopt AI agents to improve their productivity. Tim emphasizes that understanding how to manage and utilize agents will be crucial to staying relevant in the workforce.
- Adapting to New Technologies: The importance of remaining adaptable and learning to work with new AI-driven tools is highlighted.
Recommended Practices
- Focus on practical automation of repetitive tasks, which can be achieved without coding experience by leveraging AI agents.
- Develop a mindset of continuously exploring how AI can enhance productivity in various workflows.
---
Closing Thoughts The episode concludes with a discussion on the potential for future developments in AI, emphasizing both the excitement and uncertainty that lie ahead. Tim and Dan express optimism about the possibilities of AI technology, suggesting that while we may not reach AGI soon, significant advancements in practical applications are on the horizon.
Resources Mentioned
- Tim's Blog: [Why AGI Will Not Happen](https://timdettmers.com/2025/12/10/why-agi-will-not-happen/)
- Dan's Response: [Yes, AGI Can Happen](https://danfu.org/notes/agi/)
- Together AI: [Website](https://www.together.ai)
- The Allen Institute for AI: [Website](https://allenai.org)
---
Episode Details
- Host: Matt Turck
- Guests: Tim Dettmers, Dan Fu
- Duration: 1:03:34
- Release Date: [Insert Release Date Here]
If you enjoyed this episode, consider subscribing or leaving a review to help grow the podcast community.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOAI's Current State and Future Prospects
0:00 to 0:29
The discussion begins with insights on the current state of AI and the potential limits of growth.
“We have max out the features, we have max out the hardware, that's what we get.”
Exploring AGI Perspectives
1:06 to 3:29
Tim and Dan present contrasting views on AGI based on their recent blog posts.
“So Tim, a few weeks ago, you wrote a great provocative blog post entitled Why AGI Will Not Happen.”
Defining AGI and Its Implications
3:29 to 7:10
The conversation delves into the definition of AGI and its potential impact on industries.
“And then in the second part of this conversation, we'll talk about agents and coding agents and your thoughts there, because I want to make sure we cover that.”
Diminishing Returns in AI Development
7:10 to 14:01
Tim discusses the concept of diminishing returns in AI and the physical limitations of computation.
“And it starts with software engineering scale going up pretty significantly.”
Exploring the von Neumann Bottleneck
14:01 to 16:12
Understand the challenges and limitations posed by the von Neumann architecture in modern computing.
“von Neumann bottleneck in a major way on the ship.”
The Evolution of Model Training Efficiency
16:13 to 20:42
Learn about the current state of model training efficiency and the potential for future improvements.
“All right, Dan, what is your perspective on all of this?”
Pre-training vs Post-training in AI
20:43 to 24:46
Discover the differences and importance of pre-training and post-training in AI models.
“We just started writing the papers about how to do a four bit training run properly.”
The Importance of Usefulness Over AGI
24:47 to 27:36
Understand the significance of practicality and usefulness in AI development over the pursuit of AGI.
“So I think that's a whole other computational aspect of it.”
AI's Role in Various Sectors
27:37 to 28:00
Examine AI's potential impact across multiple sectors, including healthcare and robotics.
“And there's been these analyses, this book by Dan Wong going around about the manufacturing economy, the engineering economy versus the more orderly economy.”
The Promise of Robotics in Daily Life
28:00 to 29:10
Discusses potential impacts of robotics on everyday household tasks.
“Tech is certainly also a large portion that's kind of leading the stock market and driving the stock market.”
Show all 25 chapters
Evolving Hardware Landscape in AI
29:10 to 30:26
Explores the future of multi-chip hardware and AI models.
“I won't because I personally like driving.”
The Shift in GPU Programming
30:26 to 33:06
Examines the emerging trends in GPU programming and software support.
“where the NVIDIA chips are really strong, really reliable.”
The Rise of Coding Agents
33:06 to 35:06
Details the productivity boost from using coding agents in software development.
“When you're trying to hire for people who can write kernels, it's very hard.”
Integrating Agents into Everyday Tasks
35:06 to 36:36
Discusses how coding agents can automate and enhance various tasks.
“Where are we in that arc where agents are transitioning from being excellent at code to useful for the rest of our lives?”
Navigating the New Reality of AI
36:36 to 38:38
Explores the balance between AI-generated and human-edited work.
“I would say this is the main thing that has been changing.”
Practical Automation Strategies for Everyone
38:38 to 42:00
Provides actionable advice on how to approach automation in daily life.
“It's still sort of a phase of experimentation.”
Productivity Insights from Automation
42:00 to 43:50
Learn how automation affects productivity and task management.
“The more nuanced part is the part that I learned in the automation industry.”
Managing Agents Like Interns
43:50 to 45:20
Discover how managing AI agents parallels managing junior team members.
“And then on your end, what have you learned or observed in terms of agents?”
The Role of Expertise in Using Agents
45:20 to 47:20
Understand how expertise enhances the effectiveness of agents.
“And that's because you can work at such a higher level of abstraction.”
Training Future Engineers with AI
47:20 to 51:20
Examine the challenges of teaching students to balance domain knowledge and AI usage.
“You don't want to just say like, throw the agent out of the problem, walk away and never look at it again.”
Innovations in Coding Agents
51:20 to 54:40
Explore upcoming projects in coding agents and their efficiencies.
“If we let people use agents, they perform very poorly on basic knowledge.”
Maximizing Hardware Utilization in AI
54:40 to 56:03
Learn about enhancing hardware efficiency for AI models.
“Yeah, so together, I think the major question that we're trying to answer today is we have all these powerful AI models that can do all these amazing things.”
Exploring Megakernels and Their Impact
56:03 to 58:07
Learn about the concept of megakernels and how they optimize GPU performance.
“So let me dive into the megakernels first.”
Future Innovations in AI and Model Training
58:07 to 1:01:08
Discuss the future of AI, including predictions for 2026 and the role of small models.
“Yeah, so I think they're both sort of split.”
Architecture Evolution in AI Models
1:01:08 to 1:03:28
Understand the evolution of AI architectures and the rise of state-space models.
“So we're starting to hear a little bit about Rubin, the next generation of NVIDIA GPUs.”
Transcript
Automatic transcript. May contain errors.0:00We have max out the features, we have max out the hardware, that's what we get. By almost any definition anyone could have written down, let's say five years ago or 10 years ago, we basically have the vision of AGI that we had back then. Everything that grows exponential will level off because if you need resources, the resources will be exhausted. You can see up to two orders of magnitude more compute available, 100x more compute. If you don't know how to use Agents well, you will be left behind. How far can you push it? I think it's never been a more exciting time to work in AI. Hi, I'm Matt Turk.
0:30Welcome to the Matt Podcast. Today, we have a special reality check on AGI with two guests who are very close to the computational reality of AI, Tim Detmers of AI2 and Dan Fu of Together AI. In this episode, Tim argues that we are hitting diminishing returns and running into hard physical constraints, while Dan argues that we are still leaving huge performance on the table and that today's models are lagging indicators of hardware progress. Then we shift into a fun, practical discussion on how to use agents and what to expect from AI in 2026. Please enjoy this fun conversation with Tim and Dan.
1:04Tim and Dan, welcome. Thanks for having us. So Tim, a few weeks ago, you wrote a great provocative blog post entitled Why AGI Will Not Happen. And then Dan, a few days later, you replied with your own blog post, equally fascinating, entitled Yes, AGI Will Happen. I'd love to go into your backgrounds. You both have the very interesting characteristics, You think of having a foot in industry and a foot in academia. So Tim, if you want to start with yours. I'm an assistant professor at Carnegie Mellon University, machine learning and computer science department, and also research scientist at the Allen Institute for AI.
1:44My past research has been mostly on efficient deep learning, quantization. That means model compression. Take large models, compress them down from like 16-bit to something like 4-bit. Key research has been there. Qlora, for example, that's a very efficient fine tuning, a compressed 4-bit, use adapters on the model, and then use up to 16 times less memory than if you have dense both handling. And now I'm working on coding agents. There we have a very exciting release in about two weeks. I'm set of the out agents. You can quickly specialize to private data, get strong performance on any cookbooks that you like.
2:23And yeah, that's very exciting. Great. Dan? Hey, so I'm an assistant professor at UC San Diego and also my total VP of kernels at Together AI. So in industry, I focus a lot on basically making models go fast. So GPU kernels are the things that actually translate the models to how they run on the GPU. You can think of them as basically specialized GPU programs. A lot of my research in my PhD, in my lab focused on that. So I developed things like Flash Attention, which was an efficient kernel for one of the core operations of a lot of the language models that we use today. I also did research on sort of alternative architectures to transformers, things like state-based models and things like that.
3:07And together, I'm really focused on how do you make the best language models that we have today? How do we make them go faster? I think, you know, as of this morning's recording, we actually just released a blog post with Cursor about how we accelerated a bunch of their models and helped them launch Composer 2.0 on NVIDIA's Blackwell GPU. So that's a bit of a flavor of what I do. So let's get into this AGI discussion. And then in the second part of this conversation, we'll talk about agents and coding agents and your thoughts there, because I want to make sure we cover that. AGI, obviously, it's a term that everybody uses, and I think we can all agree that nobody really knows what that means.
3:46But for purposes of this discussion, what is a useful definition of AGI from your perspective? Sure. Yeah, I think so. One of the things that we kind of discussed back and forth in this set of blog posts is sort of what AGI means. For me, I think one of the things that I've been thinking about recently is that if you took where we are today with the models that we have today with the language models, and I think we'll probably talk about this a bit more later with the agents. By almost any definition anyone could have written down, let's say five years ago or 10 years ago, certainly when, you know, Tim, you and I started our PhDs.
4:21We basically have the vision of AGI that we had back then. We have things that can write code. They can write, you know, human text, even though, you know, maybe they use too many M dashes or something like that. But they can do these really, really amazing things. I think one of the things that I think about is at what level does this kind of become a new industrial revolution where you can where this technology is really going to change a lot of what the way that we do things today and have a huge, really, really great economic impact. In terms of software engineering, I feel like we're already there or almost there.
4:57There are things that may be super specialized. I don't know if they're going to be able to write the best Fortran and COBOL code in the world. But for web development, even a lot of low-level systems engineering, they're already really great. One of the reasons that I wrote my blog post was, if you think about where we are today, we maybe already have AGI or some form of AGI. And if not, then certainly the next generation of models, the models that today are training already. If they're at all better than what we have today, then we've already hit something that's really amazing and pretty wild.
5:33When I wrote my blog post, actually, I noticed like, oh, I forgot to put the definition of AGI in my blog post. Even though my blog post is very much about AGI. And I think that sometimes sort of reflects how we think about AGI. We don't think carefully about the definition. I mean, they're sort of, and I thought about it before, and I think they're sort of different kind of definitions that have sort of advantages, disadvantages. I wouldn't say, I mean, as you said before, there's not one definition that people agree on, just to mention sort of a couple. And I think one that's sort of quite widespread is to see AGI as sort of cognitive abilities, cognitive tasks.
6:13What can you do cognitively? And software engineering, very cognitive. Writing, very cognitive. moving a robot in space, that's more kinetic. You could also say like, hey, you also need to think about how you move. That's also part of cognition. But I think most people would separate that and say everything digital is kind of cognitive. And if you have physical, that goes beyond that. What I think makes sense is this economic angle. Can we get another industrial industrial revolution. What it means is, is AI useful? And it's just so broad and useful that you want to use it everywhere, kind of.
6:53It accelerates all kinds of things. Similar to when computers were introduced, productivity increased. Not initially, productivity actually went down and you need to do diffusion in the economy to pick it up again. We might see something like that with AGI more broadly. And it starts with software engineering scale going up pretty significantly. But yeah, I think that is useful. Let's jump into the heart of the argument, Tim. I was amused by what you said about where all those ideas of AGI and superintelligence come from, if you want to talk to that. Yeah. And to sort of lay out sort of the entire narrative, there are certain thoughts about AGI and that is sort of very much rooted in a certain kind of thinking.
7:41It comes from effective altruism communities of the rationality communities. I was part of these communities a long time ago. That is now 15 years ago. If you look on Twitter, there's always like, oh, we get AGI in two years. And then one year later, oh, we get AGI in two years. And then one year later, we get AGI in two years. I feel like it's a little bit of lazy thinking, a little bit of being in a bubble and not being exposed to different ideas. And that was one of the main motivations for me writing a blog post because I feel like there are some things, some ideas that if you think about them, they might provide a counterpoint to a lot of the thinking that's out there.
8:19Yeah. And your core thinking is that there is a tension between those ideas and the computational reality. Is that a fair way to put it? Yeah. There's a physical component and then there's an idea component, but there's a very similar structure. And this structure is basically diminishing returns. Everything that grows exponential will level off because if you need resources, the resources will be exhausted. Resources can mean different things. And if you look at the physical aspect, it gets just more and more difficult to advance technology. That is the case almost within any field of research or development.
8:57Things get more and more easily. You need more resources to make further progress, and the progress sort of goes lower and lower. And so if you look at the physical reality of computational devices and also computation itself, it has a particular structure. And so basically computation, useful computation is two things. The first is you need to gather data from one location and aggregate it in a certain location where you then put this new information together to compute a transformation of that information. You basically want to combine known things and compute some new things that you didn't know before.
9:44Useful information. Useful information needs to be transformed from information that you already know. If you move a lot of information around, but you don't transform it, you can't make new information. If you do a lot of computation on the information that you already have, You miss out on the long distance inside, the indirect inside. I think a lot of this actually maps to the neural network architectures that we have. The beginnings, we had convolutional networks. They are very effective. And what they're doing, they don't move much memory. They do a lot of computation. And that means your device needs a lot of flops.
10:21And memory bandwidth is not that important. Once you go to very dense computation, very large matrices, then it goes in the direction of recurrent neural networks. But there you have still sort of this component of your recurrent basically paying attention to previous states. But because it's recurrence, the memory reuse of that computation is minimal. And with transformers, you basically then have these large matrices that compute basically, that transform the incoming information from the previous layer. And then you had attention that now computes information across time or space. And I would argue these are the most two fundamental ways of computing information.
11:09You want to relate information to itself or transformation of that information. But then you also want to basically relate information to distantly related sort of information. So you want long-term relationships and you want transformations based on what you already know. And you say this is slowing down, right? In your blog post, you have a pretty striking sentence where you say GPUs will no longer improve meaning fully. We have essentially seen the last generation of significant GPU improvements. Yeah, so this has two components. And so one is also sort of a very fundamental thing. And it's physical in the sense that I mentioned these two components, memory movement and computation.
11:55And so computation can only be useful if you move memory to this sort of local neighborhood where you do this computation. Now, this is a geometric problem. You need to have a large store of information and then use this large store to move information closer to where you do much to the computation. And we have figured out how to physically do this optimal. We have a large, slow memory, that's DRAM. Then we move it to a cache. If you look at the geometry, that's how you do it fast. If you have a certain size of computation, this is optimal. If you have a different size of computation, matrix multiplication, then you want to use not a CPU, but more like a GPU, which has higher latency, but more throughput.
12:40You can move more data, but more slowly. And yeah, if you look at all of that, you can push around a little bit how you structure everything, like the caches and how large they are and how much core solid are shared. But in the end, the fundamental problem remains the same. You have a geometric problem. You can only fill the space in a certain way. And that means you always have certain access pattern with certain latencies. And the biggest latency is a big block of DRAM. That is the major bottleneck. This is also called the von Neumann bottleneck, based on computers, almost all computers that we have.
13:25And this is the bottleneck of moving a program to where you execute the program. and for neural networks that's basically move the weights in the inputs to the execution where you execute the program. That would be the TensorFlow. There are not many ways how we can go around this bottleneck. The only way is to store the memory locally and do the local computation there. And there's some processors that do something like that. For example, this Rebirth processor. So they don't have this von Neumann bottleneck in a major way on the ship. But then they need to also pipe data into that ship. And so the von Neumann bottleneck moves basically away from the ship to your storage or 3G network.
14:15And so you just move it away, but still the same bottleneck. You need to load the program, which might reside on disk or memory, through the network to the ship. And same physical problem. You just move a couple of variables around. That is sort of one part of the problem. We don't have architectures that can solve this problem. That's sort of the second part where my argument kicks in. And that is you need new technology to overcome bottlenecks. But once you have leveraged that technology, you need new technology to get over that. If you look at what we can do is we move from DRAM to HBM. So that's DRAM that's stacked.
15:00That's much faster. But you can only stack it that high because it's very difficult to manufacture and test in greatness. The yield is very low. And we were on actually in 2026. There's not enough HBM. You can't build the nice processes anymore because you run out. It's just too difficult to manufacture them. And with that, we have all these innovations. One of us, TensorFlow, big step up. Then we have an 8-bit precision, another step up. Then we have 4-bit precision with particular block-wide quantization, particular data types. From my research and other research, we know that's close to information, theoretically optimal in a practical sense.
15:38If you train on enough data, 4-bit precision is not up enough. You need actually 8-bit precision. So you can't go further. The hardware is maxed out. We have no new technology. We can make it easier to manufacture and a little bit cheaper, but not faster. And you have maxed out on the additional features. Sparsity could be something. People tried it for 50 years. I tried it present well. And so that might be the last thing, but forward precision is the end of quantization. And so that's the end of it. We have maxed out the features. We have maxed out the hardware. That's what we get. Okay, fascinating.
16:13All right, Dan, what is your perspective on all of this? I really appreciated Tim's post because I think one thing that I really appreciated is that there's some AGI talk that if you just kind of like trace the exponential, at some point you get, you know, the thing that will eat up the universe or whatever, which I always found a little bit odd to think that way. I appreciate the thing in terms of the actual physical constraints because, you know, like Tim said, these are physical systems with physical inputs and actually doing physical computation. I think my perspective was that if you look at where the systems are today and you look at the models that we've trained, we are just so far from even using the last generation of hardware as efficiently as possible.
16:57So not to mention all the new hardware that's being built out. So I think on the technical side, I'd say there are two major points I wanted to make in my post, which was one, if you look at the models that are kind of the really great ones that we know today. In my blog post, I mostly talked about open source models because they talk a little bit more about how they train the resources behind it. We don't have public figures behind how much OpenAI and Enthropec are using. But if you look at the DeepSeq model, for instance, this is one of the best open source models we have out there today. It was trained at the end of 2024 on last generation, kind of nerfed GPUs, H800s instead of H100s.
17:40The 800 is nerfed by all sorts of ways from NVIDIA to get around the export restrictions at the time. And they were trained with, let's call it like about 2 ,000 H800s, according to the report, for I think about a month. And when you compute how much, how long that took, when you see how much compute was actually available on the chip, you get something like a 20 % effective chip utilization or something like that. The number that the term of art is called MFU model flop utilization. But basically, that's a 20 % utilization number. Meanwhile, you know, in I think in the early earlier in the 2020s, we were seeing lots of training runs on older hardware, different model architectures that were easily achieving 50, 60 % MFU.
18:29So if you just take that number and then say, hey, maybe there's a way to get it out there. Since then, my good friend TreeDao has released a whole new set of kernels on how to train these models better. And you say, okay, there's a 3x there just from that one piece. Then the other thing to realize is that, so that is a model that is being used today in early 2026 as the base for some of the best open source or open source adjacent models out there. It would have started training the base model at least a year and a half ago. So let's call it mid 2024. Since then, we've started building out completely new clusters with the current generation of hardware.
19:09So on NVIDIA, these are the Blackwells. We started, you know, they're companies like Poolside. They're building out tens of thousands of B200, GB200 chips. You know, there's other folks like Reflection who are building out tens of thousands of B200 chips. So this is comparing, we have a new generation of hardware where even if you take the exact same precision as you had before, exact same everything, 2 to 3x faster compute, 10x larger clusters, plus maybe 3x lurking in terms of just pure optimization, that's 3 times 3 times maybe another 10, that's like another 90x of compute available. and that's not even looking at future build-ups that is literally clusters that you can point to today that people have started training on that you might hope that at the end of that you'll get much better models the point I really wanted to make was if you just look at it from those basic inputs you can look around you can squint a little bit you can see up to two orders of magnitude more compute available compared to the models that we are indexing on today now we can argue about, you know, is there, are there going to be diminishing returns in terms of scaling up?
20:20Are there going to be, do we expect the scaling curves to hold and all that? But you can just look around and see it and that's, you know, a hundred X more compute. So I think from the physical, just a pure compute perspective, there's a lot more available, a lot more that we're not doing. This is not even to mention a bunch of the great points that you mentioned, Tim. So these are eight bit training runs. We just started writing the papers about how to do a four bit training run properly. There's new things like the, on the GB200, you have 72 of these really connected really, really quickly. I don't think we've even seen the first pre-trained model come out of that yet.
21:01GBT 5.2, I think was the first time that you saw in one of opening eyes reports, hey, this was trained on H100, H200 and GB200, which to me suggests that was actually pre-trained on one of the really old clusters, maybe some fine tuning was done on the new GP200s. You make the point that not only is the hardware underutilized, but you also say that models themselves are a lagging indicator. Yeah, so the models that we see today that we can play with today have been pre-trained on clusters that were built out a year or two ago because you need enough time to get the cluster running, you need enough time to do the large pre-training run, and then you need enough time to really post-train it, fine-tune it, do all the RLHF and all that stuff.
21:46So the models that we have a snapshot today that at the beginning of the conversation, it's like, maybe it is AGI, maybe it isn't, are already trained on clusters that are a year and a half old. We've built out much larger clusters since then. We can expect, you know, you can expect that they're going to use them for pre-training. The models that we see today that we index on quality today are actually trained on pretty old hardware. and we've got new generations of hardware, more software choices we can make, not to mention architectural choices. So Tim, you were mentioning this thing about you need to move data and then compute on data.
22:21We've actually seen the transformer change in architecture a little bit slowly for a researcher, a little bit slowly for my taste, but you've seen the fundamental way we do the computation change. Even if you find another 1.5x or 2x there, Now you're talking, you know, 100, 150x more compute. So there's a lot more compute out there to train better, higher quality models. If I understand this whole discussion correctly, all of this is about pre-training, right? And whether we can train a bigger model with, you know, more data and more compute. But in conversations on this pod, a lot of the conversations have been about the importance of post-training and building AI systems with pre-training plus RL.
23:09Where does that fit? That's a great question. And I think another piece that we didn't, you know, I don't think either of our blogs particularly hit on. One way I like to think about it is that pre-training is like the general strength training that you do in the gym. You go lift heavy weights, you improve your strength, improve your general ability. and then post training is like the specific drills that you run to get a good at a particular task. So historically, the vast amount of compute has gone to pre-training. So just building models that are more generally capable of doing many things, have a lot of knowledge, get to a point where maybe they have more knowledge than your average person.
23:46I certainly don't know as much as ChatGPT, for instance. And then the post training is both how do you make it helpful. So, So, you know, chat GPT, you ask it to do something and then it actually listens to you and tries its best to do it. But I think the other thing that we've started to see increasingly in post-training is that you can start to post-train specific skills. So the model that's really good at helping you code uses a lot of the knowledge that you got from pre-training, but it's actually adapted to be particularly good for coding. Or the model that's really good for legal work, for instance, has a lot of the pre-training backbone.
24:22But then the post-training is really what gets it to that place where it's really useful. From a pure computational perspective, pre-training is usually much more compute expensive than post-training. Post-training, the work that you have to do, I think, I'm not a post-training expert, but the work ends up looking a lot more like how do you build a useful product? How do you get user feedback? How do you do things like that? Even then, there's a world where maybe the next generation of pre-trained models is a strong enough base that if you go tackle each vertical of the economy that you care about, you could actually post-train it to something quite useful.
25:03So I think that's a whole other computational aspect of it. So maybe we don't even need that 100x more compute that may be out there. maybe it's more a more traditional work of, you know, let's understand this problem deeply, let's understand how to train in almost the human sense. Like how would you take an intern and train this intern to do this specific task? How do you get this very powerful pre-trained model to do something really useful in this post-training sense? Is this concept of usefulness that you both mentioned where both of your point of views maybe converge? Some ways AGI is something, But what ultimately matters is where you land in terms of usefulness in the industry.
25:43And therefore, even though you, one, may not be able to reach that kind of ethereal definition of AGI that nobody really understands due to diminishing returns, in some ways doesn't matter because we still have like so much juice to squeeze that we have enough to go until we get to a place where this is truly useful, not just for coding, but for the rest of the economy. Yeah, the main conclusion of sort of my blog post was exactly that. I shouldn't pay too much attention to AGI, but more about thinking about how can we make it most useful? That might go beyond of how useful is a model. I mean, Dan mentioned post-training as a product.
26:22An important part we saw with computers is diffusion in the economy. That requires a very different mindset. The best mindset is build the best model and then everybody will use it. but that can help to really figure out how can you benefit the most people in the most pragmatic way. And I think that's sort of more of a Chinese mindset. And so that sort of mindset, so if I think of usefulness, one is model, the other is sort of the mindset. But I would agree, I think both Dan and I, I think most people would agree that if you have AI that does very impressive of things, like math Olympiad things and that sort of thing, but it can't do anything useful, is that AGI?
27:06And so models are already useful, so that scenario will not happen. But I think what we really want is very, very useful models. And I think we have that, and I think we can prove that. But I don't think we get to AGI by cert definitions, but we will see a significant an impact. Yeah, I think I would just add to that, Tim, you had this point about how much of the economy is physical, how much of it is knowledge work. And I think that the US-China contrast is really interesting there. And there's been these analyses, this book by Dan Wong going around about the manufacturing economy, the engineering economy versus the more orderly economy.
Read the full transcript
27:49and I think there's certainly a lot of great knowledge work to be done in the US. I think there's also, if you look at what the actual sectors of the economy are, so a large portion of it is healthcare, a large portion of it is education. Tech is certainly also a large portion that's kind of leading the stock market and driving the stock market. There's a lot of great people who are trying to use the new models to try to do things like develop new drugs or understand how to make a real impact in healthcare. Or if we can get robotics off the ground and do things like start taking, helping with some of the physical labor, maybe not necessarily building houses, but the day-to-day household labor that could be large, untapped portions of the economy.
28:36Those pieces are really great. You can almost start to see the first pieces towards it. But the self-driving analogy is really interesting to me because early on, I'll say early on in my PhD, I was quite skeptical about self-driving. So let's call this 2018, 2019. It felt like self-driving was always a year away or two years away. Or if you ask the experts, they say, oh, five years away. And then last year I rode in a Waymo. And today I just actually got access to Waymo on the highway. So now conceivably I could potentially sell my car. And I live in the Bay Area in California. I won't because I personally like driving.
29:14But there's a lot of the progress is funny in this way where it's kind of it's not there. It's not there. It's not there. And then one day something, a switch flips and then you're suddenly like, oh, not only is this thing pretty good, it's actually a lot better than the service that I'd get in an Uber or a taxi or something like that. That's a really exciting thing. If we see that happen, if we see that switch flip for like, you know, household cleaning or putting away the dishes or things like that. I think it would really be really exciting. It would change a lot of folks' perspectives. I'm not a roboticist myself, but I'm really watching that space with a lot of excitement.
29:48Dan, as a quick tangent, do you think we're evolving towards a multi-hardware, multi-chip kind of world based on what you see? You know, obviously there's been Grok and NVIDIA, there's Cerebrus, there's like a bunch of sort of special ASICs companies coming up from your kind of like low level in the stack, Vintage Point, what do you see? Yeah, that's a great question. So it's something that I spend quite a bit of time thinking about, more so I'd say on the lab side than necessarily on the industry side. Although of course we're paying close attention kind of on both sides. I think it's at a really exciting time where the NVIDIA chips are really strong, really reliable.
30:32there's a lot of software support around them that has built around. We're starting to see the same things happen, for example, on AMD chips with some of the research there. So on the lab side, we put out recently a library called HipKittens led by my great friend Simran Arora. And she was really looking at, okay, how can we take, what are the right software abstractions to program on these AMD GPUs? And it turns out they're not exactly the same as the NVIDIA GPU. So even two GPUs that have relatively similar specs, certainly compared to Grok or Cerebrus or Samba Nova or one of these other chips, even though they're relatively similar, they actually have pretty different software abstractions you need to use.
31:13And I think more people are getting excited by that and investing time and energy into that. We saw the Grok acquisition from NVIDIA. A lot of people are excited about TPUs today. I think Cerebrus and OpenAI just announced their partnership. So I think certainly it's going to be a wave of things coming forward that you're going to see a lot more. I'm sure NVIDIA will still do great and still grow beyond their$5 trillion company or whatever it is at the time of recording. But I think you're going to see a lot more diversity, especially around, I think, inference of the model. So training and inference are actually quite different computations.
31:55And as a result, you might actually want quite different chips to do it. On the inference side, you might want, for example, your models to live locally, on your phone, on your laptop. You know, my phone, my iPhone, which is a few years old at this point, is already more powerful than some of the GPUs that I had when I was starting my PhD. That growth of that hardware power is really exciting to see. And then you mentioned a second ago in reference to self-driving cars, that moment where things flip, switches, is turned on. Has that happened with agents already? You talked about software singularity.
32:30Are we at that moment for agents? Yeah, I think that so personally in my life, I'd say that moment was last June ish. So June 2025 was the moment that it really flipped for me to give some context here. So what I do in my day job at Together AI is we write a lot of these GPU kernels. I don't know how popular, but in the general ML zeitgeist, GPU kernels are thought of as kind of like the final boss of the thing that you learn how to program. They're very hard. They're very highly parallel. You don't write them like you have to write in C++, which is this old language that the old systems people used decades ago or whatever.
33:07They're not in Python, etc. When you're trying to hire for people who can write kernels, it's very hard. It's a very challenging skill set. It's certainly the tip of the spear in terms of the programming strength. And last June, we had this really interesting realization where we realized that cloud code, cursor agent, these agentic coding assistants were actually very good at writing these kernels. So there was one week where I think I wrote like three or four different features that usually would have taken me a week each. And I wrote all of those in a single day. And I was like, oh, my God, this thing is making me five times more productive as a kernel expert.
33:45I got my team on it. Now my team has all these really complex systems that they've built where they can write a whole feature that I think would have taken months of a whole team's time before. And this is kind of that final boss of programming challenge that was really challenging. So from our perspective, for coding, for this really technically challenging GPU kernel programming, it kind of crossed the Rubicon for us already. And I think, you know, I gave this talk a few months ago at Slush about what we're calling the software singularity, where we realized, hey, in terms of software engineering, even for these really niche skills, it's certainly better than the average programmer.
34:30programmer, it's at a place where it can accelerate the really expert programmer. So it's right now, as of today's recording, it's at a place where if I just let it on its own, it might not generate the right thing for you. But if you give an expert programmer this set of tools, they can go 10 times faster than they were able to go before. And I think that that's a really exciting place to be. And on the topic of agents, Tim, you just wrote another great blog post called Use Agents So be left behind. And part of what you talk about is coding agents versus agents for the rest of tasks. Where are we in that arc where agents are transitioning from being excellent at code to useful for the rest of our lives?
35:15That blog post flows also as a reaction to what I see is there are large productivity gains if you use coding agents for all kinds of tasks. tasks and as a professor you don't code that much you can actually code more easily which probably other professors previously would not do so easy now but yeah also for non-coding tasks is super useful and when i look at um the productivity gains that i have some others smaller like two or three sometimes it's like 10 times faster i do tasks 10 times faster the quality is not degraded sometimes the quality is um higher um an agent might not be as good as i am but the agent gets get tired.
35:55Agent doesn't make bad mistakes. And you see, you cognitively struggle with complicated information that you put together similar to booter kernels, but what Dan mentioned, all of that is working. And I mean, Matt, as you put it, is coding agents and agents for other stuff. But how I would see it is, it's just coding agents. Coding agents are general agents. Coding agents can write programs that solve other problems. And code is so general, if there's a digital problem, you can solve it for code. And coding agents make the thing so easy that now you can solve a variety of problems in a way that you couldn't solve before.
36:37And this angle makes you productive. I would say this is the main thing that has been changing. Coding agents allow you to take a problem in a way that you could think before. and it's at a pace that you could think before, you can paralyze a lot of different tasks. The agent doesn't get tired. You just keep going. The work's much easier. One bit I love in your post is that you're careful to separate the hype from reality at the beginning, but then quickly you land from your experience of experimenting with agents, for the Lystra in particular, you land at the conclusion that more than 90 % of code and text should be written by agents.
37:15You need to do so or you will be left behind. I think for many people that are engineers, that's already true. And there is sort of this thinking, oh, if you produce text or code, and everything's done by agents, it must be low quality, much too bad. But the key thing is you inspect the code, you inspect the text. You might make some slide edits. The 10 % that you do might make a big, big difference. Through this basically sort of editing, just reviewing of output, you kind of make it your own. AI-generated things are not less personal than things you've written on your own. I see that. If I write a grant proposal with the help of an agent, it's sort of alive.
38:03I can feel like it's exciting. Like this person reading it said, like, this is good research. I want to fund this. I think that's just the reality. Like if you just generate things and don't look at it and just say like, yep, that's good. That will not help you. But you can quickly review content. You can skim it. You can look at like, ah, this doesn't look right or I want it different and edit it and you're good to go. That will be the reality. And the skills that you need to work in that way, they are not fully developed for most people. They're also not fully developed for me. It's still sort of a phase of experimentation.
38:41Models move, frameworks move. And so you need to adapt. You need to learn. There's a lot to learn. But if you do it, payoffs are huge. And I think there was the thinking that software engineers will no longer exist. But I think people no longer believe that. It's like software engineers are so productive. That is essentially what you need to learn. If you use agents well, you can do so many things. And I think that's the core thing. If you don't know how to use agents well, you will be left behind. That will become a critical skill. Practically, how do I do that if I'm not a coder and I think about automating some parts of my job?
39:21What are some of your recommendations on how to approach that problem? Yeah, I mean, the best thing is just sort of being sort of very pragmatic. Just think about things and trying to code them. Particularly if you're not a coder, that's very difficult. and there's sort of this barrier with, say, like, I haven't coded before. Like, I don't know this. But if you interact with agents, they can just build stuff. And with minimal learning, I mean, they can also explain stuff. With minimal learning, you can get there, execute programs, build websites. Particularly if it's visual, you get quick feedback.
39:55It's not that difficult anymore. I mean, often I mentioned you need to inspect things, but if you build simple tools for yourself to make your life easier, Often you don't need to do that. The agents write good code. If you work in a company, you need to integrate it in a good code place, you probably should review it. But if you build a small program on your own to make your work more productive, that's easy. Just a defining example that might be relevant here. I put, for example, a tool that if I have a video where I talk, so I record videos of how I interact with agents, then there are certain phases where I just look at outputs and try to understand things.
40:34and their faces where I talk. So I just build a tool that recognizes the speech that when I'm saying, it's a timestamp, then it slices the video. So basically you have an entire video where I talk rather than sort of moments where nothing happens. And that's very easy to do. I built this like in 20 minutes. I think everybody can do it because I didn't look at the code, the agent that did it. And then I looked at the video and like, oh yeah, it's still right. If you get started with a feedback loop, you don't need to code. you just need to inspect the output that you can understand or learn how to execute a Python program or a bash style and you're there.
41:10How do you pick what you want to automate? Like, how do you, how do I think about automation in my life? Yeah. So, so I also talked about that in a blog post and it can be sort of a sort of a more intuitive thing and then a more nuanced thing. I think the more intuitive thing is like, just think about what could be useful and then can That even be like something more complex. You say, I want an Android app or an iPhone app that does this thing. And you initially might think that's complex, but then you throw a coordination and it works immediately. The world's noisier. It's like there's so many things that you can just do.
41:48And you can be very creative and say like, what I always wanted to have and it wasn't there. Nobody built this product. Can I do it now? And I think that mindset gives you useful things. That makes you more productive. but it also flexes your muscle and sometimes it doesn't work and then you understand like okay at aging cycle this or this is what i still need to learn to make both these kind of things and i think that is sort of the more intuitive perspective that's very useful and that quickly get you started on the path where you say like um i mean first there's excitement then it's sort of sober there's a sober reality but then you pick up again and it's like you realize okay if I do it like this, I get more and more productive day by day.
42:30That's the more intuitive part. The more nuanced part is the part that I learned in the automation industry. I worked three years in automation industry in Germany, automating factories. It's a very calculated approach where you look at how you work, you time each of these steps, and then you say, if I automate this exact step in such a way, what could be the payoff? How much time would I save? And then you calculate what is the productivity gain? And then you calculate how much time do I need to develop this automation? And if you do that, you can quickly realize that automating certain things will not make a difference.
43:12The blog posts are matching emails. It doesn't really work. And there might be other things. A big thing is always calendar invites. Nobody likes to create invites for a meeting. But then if you think about it, you're also very particular about meetings. Some days you want more meetings, you have a meeting day, or you say, I can put this in before lunch. An agent doesn't know that. And if you specify that to an agent, you could also just create the calendar invites, just meeting pipeline. And it doesn't increase particularly much. And so there are a lot of problems if you think about this nuanced way and can say, okay, I don't do that.
43:49Ah, this will help me. And then on your end, what have you learned or observed in terms of agents? What works, what currently doesn't work, but will work soon? How to manage them? I think there's two broad things I've noticed for agents. So the first is making the agents effective ends up being a lot like managing junior folks on your team or at a company. So, for example, the new intern who shows up on your team, you're not going to go to the intern and say, hey, go fix our revenue for the year, double our revenue for the year or something like that. Maybe you'll try that once, but you're unlikely to see the payoff from that.
44:31Instead, what you often do with junior folks is you say, hey, here's a first little task that you can do to get to know this complicated code base. And here are the things that you might run into because you've kind of done it before. And when you give the agents that context, give them that ability to look at those things, then they can usually figure things out. The other bit is that when you have a new person on your team, you maybe won't give them access to all the production credentials and all the production database and all those things, but you're going to give them enough tools to be productive.
45:03So sometimes there's this tension between, oh, I don't want my agent to go delete my everything in production. So I'm just going to have it be hamstrung and watch every little thing it does. Whereas if you did that with a person, you would never expect that person to be productive. That's another bit. You kind of want to think about the agents as, at least today, as maybe interns or more junior folks. The other really interesting thing that I've noticed, and I think when I think more the educational role of a professor and how do you prepare people for this future where agents are going to be such an important part of workflows, is that how do you train for that?
45:39And one of the things that I've noticed is that the more expertise somebody has, so whether that's in, for example, Tim with process automation or expertise in what I do in kernels and writing these very highly specialized programs, the more expertise you have, the more powerful the agent makes you. And that's because you can work at such a higher level of abstraction. You know what the important things are. You know how to set the direction. You know what the common pitfalls are, like what is easy, what's kind of hard, what you need to break up into multiple steps. One piece of conversation that was coming up for a while was like, are agents going to replace all software engineers or things like that?
46:17Or are they going to replace all junior people or something like that? I think where we are today, that's probably, you know, clearly not the case anymore. where, you know, if I have a tool to make my team 10x stronger, I'm not going to fire nine people on my team. I'm going to say, okay, go do this, become 100x more productive than you were before. That's one bit. But then also kind of the script for how you become an expert at something is probably pretty similar to the way that it was before. You're going to study things deeply. You're going to try to understand things a lot. You're going to want to do things yourself with with your, you know, hands on and really get things done.
46:55In this world, you know, the chat GPT can teach you a lot of things. I was personally, personally, I was trying to get chat GPT to teach me all the little ins and outs of how a car works. I don't know how effective it was so far, but, you know, there's, it's, it's a lot easier to, to learn things now than it was certainly even five, two, three years ago. So those are kind of some of the, the, the things I'd say. So, you know, you want to treat the agents as if you're, you are in that manager role, you want to help them get unstuck. You don't want to just say like, throw the agent out of the problem, walk away and never look at it again.
47:29But you also want to kind of figure out how to level yourself up so that you can be a better manager, have more domain expertise and really understand things in a deeper way. So the fact that you need to learn and be an expert and that doesn't change, I think that's very interesting and that makes a lot of sense. This question is like if you show up on the job as a young kernel engineer and that's your first day, typically they would be, okay, well, you do this simple task and that other simple task. And then by year two, you graduate to a more complex task. How does this sort of hands-on job training look like?
48:04Right. Yeah. Yeah, so we think about this a lot together because we're still hiring aggressively even in this world where the models are very good, the agents are very good. The way we think about it, so first I actually went, the professor and me went and actually recorded a bunch of lectures on how GPUs work. So I make everybody watch those and then I still give them a task from scratch which is, okay, go take this flash attention kernel and modify it to do some other thing. You can pick the extra feature. The nice thing about the agents is that you can dive into that higher impact role in a way that you weren't able to before.
48:39So it's really impactful, I think, when a junior IC goes to try to manage someone for the first time, because you're suddenly starting to think in much more precise terms. So, you know, the classic software engineer thing is, hey, the PM asked me to do this and wrote this super long doc with all these requirements. But then the minute that you try to go ask someone else to do something, you realize the specificity that you need when you want to address a feature or something like that. The nice thing about agents is you can almost start to shortcut that process where you can have the junior IC still be an IC, still do IC style contributions, but they can now act as that manager role and act as their own PM.
49:18Because when you're communicating with the agents, you need to be as precise about what you're trying to get done. So in some sense, I've seen with the junior folks who are joining my team. So these are folks fresh out of college or fresh out of a master's when they are really gung-ho about understanding and being able to use the AI agents. They're able to communicate so much better than in the olden days. They're able to level up their level of understanding a lot faster. And then, of course, they can do things and build tools at a speed that would have been really, really hard to do, you know, five, ten years ago.
49:56And maybe I'll add sort of the educational perspective there. This I think is sort of quite interesting. It's a little bit sort of contrasting. So what is also quite interesting is the educational perspective on agents. I talk quite a bit that basically use agents will be left behind. And that is also true for students. But just like as Dan said, you need the domain expertise. You need to have some knowledge to use agents well. What we're seeing is if we allow students to use agents, they are very productive. but sometimes they build solutions that look correct that are actually very bad or just wrong and they don't realize we're at this point where it's almost very difficult to learn both domain expertise and agent use that's a very difficult balance to achieve because we don't want to have students that don't understand things but we also want to have students that and basically can use agents.
50:56And so if they can't do that, they will not be effective in the workforce. What Dan said is you already have a pretty good person with strong background knowledge, and then that person can level up their equipment with agents. But what do you do if you have someone that just learns computer science? How much AC should they learn? How much work should you do without agents? And that's a very tricky balance, and we don't know how to solve it. If we let people use agents, they perform very poorly on basic knowledge. And if we let people just do the basic knowledge, they don't know how to use agents and they can't compete.
51:32So they can't do useful work in the workforce nowadays. Maybe their solution is to all sort of the basic knowledge first and then agents, but that's not what students do. Students have access to these AI tools. They will use them because it's easy. And so maybe the solution is just you need a way of thinking of working with information and knowledge that you don't understand. And develop, I mean, there's critical thinking. I think this goes beyond critical thinking. You basically need to know the unknown unknowns, things that you didn't consider and don't understand and that you didn't even think about.
52:11You need to have the ability to think more about that to really keep up with agents. Because I think in the future it's realistic that we work on problems that we don't understand, that agents understand, but we need to keep up in some way, that would be difficult. All right. So to switch tags as we get closer to the end of this conversation, what are you guys currently working on? What's top of mind for you at Allen Institute on the one hand, together on the other hand, whoever wants to take it first? We have actually sort of a very exciting project that will be released very soon in the next weeks.
52:50And so I worked quite a bit of efficiency. I've been switching my work basically to coding agents. And so we will have a major release of an open source coding agent that has a couple of key features. For one, training is 100 times cheaper. You need to generate synthetic data and you need to train on it. And so we have a method that's roughly 100 times cheaper, but it still gets state-of-the-art performance. And then we have another sort of major result, which you can almost see as the holy grail of open source models. And that is, we can take a private code base. Like if a company, you have this code base.
53:33Cloud doesn't know your code base, but you have the data. You can fine-tune a model on it. So what we have is you can just point our method to that repository. You don't need to have any sort of tests or need to understand how to generate the data. It's just automatic. You quickly generate the data. And then you have an agent that is as good as a frontier model, but you can have like a 32 billion agent. You can depart locally. You can have an army of specialized models for particular tasks, for particular code bases, and so forth. And yeah, I think that is a very powerful result. All of that is also packaged with a science of coding agents.
54:15There are a lot of confounding factors, a lot of hidden things in papers that are not mentioned. We sort of unearth them, build scaling laws, and show what does matter and what doesn't matter. And if you put all of that together, I think the sheet agents, very few GPUs, everybody can use it. And very easily, we unblock people by revealing all the sort of secrets. I think there will be a vast change in terms of how quickly we can progress in coding agents. Very cool. What about you, Dan? Yeah, so together, I think the major question that we're trying to answer today is we have all these powerful AI models that can do all these amazing things.
54:53They're very expensive to run today. So, you know, the question in all the public markets is like, is OpenAI going to be able to turn a profit ever? and what's really exciting about together is that we are kind of on the forefront of getting these models to use the hardware as good as it can. We talked a little bit about training early on in the podcast. At inference time, when you have the model, when it's already been trained, already been post-trained, the hardware utilization is like less than 5%. So it's at a place where there's so much more that we can do. There's so much more efficiency that we can push out of these and we're really excited about figuring out how to use the hardware the best way that we can, whether it's serving customers like Cursor or whatever the next great foundation model company is.
55:37But yeah, we're really excited about pushing that frontier and then really getting us to this future that if these models are going to be as impactful as you think, if they're going to be running everywhere, running your daily lives, running your toaster, well, we better make it the best possible toaster that we can. Do you want to talk about megakernels and Together Atlas for a minute? Together megakernels, Together Atlas, these are both projects along these lines. So let me dive into the megakernels first. So to understand this, the first thing when we say kernels is we usually mean we are going to write a specialized GPU program for a single operation in a model.
56:15You can think of a model as one of these trained models. It's like ABC, different operations in a row, and there will be hundreds of these. And the way that we've been writing kernels for the whole history of, let's say, call it NVIDIA hardware, is that you really specialize a single kernel for a single operation. With these mega kernels, we're doing something quite interesting, which is we can take the entire model, however many billions of parameters, and put it into a single GPU kernel. And with that, you can start to do a lot more fine-grained optimization than you were able to do before. It actually starts to make the NVIDIA GPU look a little bit more like a Cerebris chip or look a little bit more like a Samba Nova chip in terms of the optimization that you're able to do.
56:55And this is really critical at inference time. So we're able to see 2x, sometimes 3x speedups over even highly optimized inference engines. So we're working on bringing that to really work in production, bring it to fruition and use it across our whole stack. Together Atlas is another really great, interesting research project that we've done recently where we can get the model. There's this technique called speculative decoding that we use at inference, where basically we have a little model that is trying to guess what the big model is going to do. And because of the ways that we've designed language models, if the little model guesses correctly, you basically get those tokens for free.
57:34So if you do the speculative decoding right, you can get, again, 2x, 3x speedups over just running a vanilla model. With Together Atlas, we do one extra thing, which is we say, okay, we can get that little model and we can actually adapt it to your traffic. The longer you use this model, the more it learns the patterns of what you're asking, of what you're saying, and it can actually get faster over time. So all these things, we have efficiency in mind, inference efficiency in particular, and we're all kind of pushing towards making all these things faster. What are you excited about for 2026 at a reasonably granular level?
58:14What do you think happens? What do you think doesn't happen? Yeah, so I think they're both sort of split. I think a lot of things will be very boring and not much innovation. But then we're also surprised by a couple of things that we maybe don't see. And I think actually sort of in the frontier, we will be less surprised. I mean, it's no secret that we ran out of pre-training data. And as Dan said, these are sort of the muscles. Then you can sort of smooth over. And you smooth over with synthetic data. And that's how you build coding agents on lots and lots of different environments. You combine the data.
58:52we make some problems there but I think you already see a bit of machine returns I don't think coding agents will be that much better the user experience will probably improve but you see it that all these models get almost equally good I had my config like GLM 4.7 set up and I used it and I thought I was using Opus 4.5 and then realized oh wait I used a different model because they're quite quite similar. And so I think we see less progress there. Where I think we see more progress is actually the small models. If you train smaller models on more specialized data, they can do quite well. And the smaller models that you get, they're pretty powerful.
59:39A hundred billion parameter model, you can fit it pretty well, even sort of low-grade data center GPU, like an RTX 6 ,000 or$6 ,000. I think for a lot of companies, it will be very interesting. They don't need to rely on the frontier models. The small models might even be better because they're specialized. The big problem, and Anthropic CEO pointed it out, we have these powerful open-rate models, but nobody uses them because the deployment is so complex. And that is because once you go beyond eight GPUs, you first need the users to make efficient, but then also very complicated inference systems.
1:00:13There's no open source system that can do that at the moment. We have disaggregated entrance, separation across sequence lengths, and so forth. Perhaps we can build this. We can build this also for an 8GPU machine for smaller models. And then the efficiency that you see with 100 billion models will rival what frontier models have. So you will get the efficiency of small models. You get the flexibility of small models. Performance on the frontier will stagnate, but on the smaller level, we get more and powerful models still, because you can distill from these large models into these small models.
1:00:47And taken together, I think that will change things. I'm also really excited about small models. I think we're going to see a lot more capability out of them. I'll be watching the open source models pretty closely, I think. With things like JLM 4.7, you're starting to see it. You're starting to see the open source models rival some of at least our current best frontier models. I think we're going to see another big jump in open source capabilities this year. I'm really excited to see new hardware. So we're starting to hear a little bit about Rubin, the next generation of NVIDIA GPUs. I think we're starting to hear a bit about the AMD 400 series of GPUs.
1:01:24I'm really excited to see kind of what that next jump in hardware capabilities is, even as we haven't fully kind of used even the current generation of hardware. I'm excited to see what people do kind of with all the other modalities. I think last year, video generation models had a little bit of a moment with Sora 2, with Gemini, and Vio, I think they called it. Really excited to see what they can do. And yeah, really excited to just see, you know, what is that frontier of intelligence that you can get on your laptop or on your phone and how fast can you push it, how far can you push it. I think it's never been a more exciting time to work in AI.
1:02:02You both mentioned state space architectures earlier in the conversation. Do you think that's part of the near future that we sort of evolved to post-transformer architectures with a state-space, JEPA, world models, whatever direction? Is that something that you see in the near-term horizon and that you think is desirable? I think in a lot of places, they're already there. So some of the best audio models in the world are at least partially based on state-space models. I think NVIDIA released a bunch of really great hybrid models recently. I think NemoTron is what they called them. And so there's a lot of really great work kind of there already.
1:02:43I think we will see the architectures continue to evolve. So in some sense, the DeepSeq MLA compression takes some of those ideas. One of the Minimax models had sort of a linear attention idea. So I think you're going to see a lot more diversity in architectures. You're kind of already seeing it. But certainly, I think out of the Chinese labs where there isn't really an open AI of China, right? So there's an open AI or a Thropic or a Google Gemini that kind of brings all these centers of product and model and revenue and all those together. So I think you see a lot more risk taking out of the Chinese labs where you're trying to differentiate the next model, your next open source model.
1:03:25One way to do that is architecture. Another way to do that, of course, is just pure quality. So I think we're going to see a lot more explosion of different architectures. All right. Well, this has been a fascinating conversation. I really appreciate the time and insight and thoughts. Thank you so much. Really appreciate it. Yeah, thank you so much, Matt. Thanks so much for having us. Great to see you, Tim. This was a lot of fun. Hi, it's Matt Turk again. Thanks for listening to this episode of the MAD podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing if you haven't already or leaving a positive review or comment on whichever platform you're watching this or listening to this episode from.
1:04:01This really helps us build a podcast and get great guests. Thanks and see you on the next episode.
From the publisher
Will AGI happen soon - or are we running into a wall?
In this episode, I’m joined by Tim Dettmers (Assistant Professor at CMU; Research Scientist at the Allen Institute for AI) and Dan Fu (Assistant Professor at UC San Diego; VP of Kernels at Together AI) to unpack two opposing frameworks from their essays: “Why AGI Will Not Happen” versus “Yes, AGI Will Happen.” Tim argues progress is constrained by physical realities like memory movement and the von Neumann bottleneck; Dan argues we’re still leaving massive performance on the table through utilization, kernels, and systems—and that today’s models are lagging indicators of the newest hardware and clusters.
Then we get practical: agents and the “software singularity.” Dan says agents have already crossed a threshold even for “final boss” work like writing GPU kernels. Tim’s message is blunt: use agents or be left behind. Both emphasize that the leverage comes from how you use them—Dan compares it to managing interns: clear context, task decomposition, and domain judgment, not blind trust.
We close with what to watch in 2026: hardware diversification, the shift toward efficient, specialized small models, and architecture evolution beyond classic Transformers—including state-space approaches already showing up in real systems.
Sources:
Why AGI Will Not Happen - https://timdettmers.com/2025/12/10/why-agi-will-not-happen/
Use Agents or Be Left Behind? A Personal Guide to Automating Your Own Work - https://timdettmers.com/2026/01/13/use-agents-or-be-left-behind/
Yes, AGI Can Happen – A Computational Perspective - https://danfu.org/notes/agi/
The Allen Institute for Artificial Intelligence
Website - https://allenai.org
X/Twitter - https://x.com/allen_ai
Together AI
Website - https://www.together.ai
X/Twitter - https://x.com/togethercompute
Tim Dettmers
Blog - https://timdettmers.com
LinkedIn - https://www.linkedin.com/in/timdettmers/
X/Twitter - https://x.com/Tim_Dettmers
Dan Fu
Blog - https://danfu.org
LinkedIn - https://www.linkedin.com/in/danfu09/
X/Twitter - https://x.com/realDanFu
FIRSTMARK
Website - https://firstmark.com
X/Twitter - https://twitter.com/FirstMarkCap
Matt Turck (Managing Director)
Blog - https://mattturck.com
LinkedIn - https://www.linkedin.com/in/turck/
X/Twitter - https://twitter.com/mattturck
(00:00) - Intro
(01:06) – Two essays, two frameworks on AGI
(01:34) – Tim’s background: quantization, QLoRA, efficient deep learning
(02:25) – Dan’s background: FlashAttention, kernels, alternative architectures
(03:38) – Defining AGI: what does it mean in practice?
(08:20) – Tim’s case: computation is physical, diminishing returns, memory movement
(11:29) – “GPUs won’t improve meaningfully”: the core claim and why
(16:16) – Dan’s response: utilization headroom (MFU) + “models are lagging indicators”
(22:50) – Pre-training vs post-training (and why product feedback matters)
(25:30) – Convergence: usefulness + diffusion (where impact actually comes from)
(29:50) – Multi-hardware future: NVIDIA, AMD, TPUs, Cerebras, inference chips
(32:16) – Agents: did the “switch flip” yet?
(33:19) – Dan: agents crossed the threshold (kernels as the “final boss”)
(34:51) – Tim: “use agents or be left behind” + beyond coding
(36:58) – “90% of code and text should be written by agents” (how to do it responsibly)
(39:11) – Practical automation for non-coders: what to build and how to start
(43:52) – Dan: managing agents like junior teammates (tools, guardrails, leverage)
(48:14) – Education and training: learning in an agent world
(52:44) – What Tim is building next (open-source coding agent; private repo specialization)
(54:44) – What Dan is building next (inference efficiency, cost, performance)
(55:58) – Mega-kernels + Together Atlas (speculative decoding + adaptive speedups)
(58:19) – Predictions for 2026: small models, open-source, hardware, modalities
(1:02:02) – Beyond transformers: state-space and architecture diversity
(1:03:34) – Wrap
