In short
Y Combinator Startup Podcast: Episode Summary
Episode Title
Anthropic Head of Pretraining on Scaling Laws, Compute, and the Future of AI
Host
Ankit Gupta
Guest
Nick Joseph, Head of Pre-training at Anthropic
Episode Overview This episode features a conversation between Y Combinator General Partner Ankit Gupta and Nick Joseph, Head of Pre-training at Anthropic. The discussion revolves around the complexities and challenges associated with training advanced AI models, specifically focusing on the engineering aspects of pre-training, scaling laws, compute management, and the future implications of AI technologies.
---
Key Topics Discussed
- Nick Joseph's Background
- Previous roles at Vicarious and OpenAI.
- Shifted focus from academic pursuits in economics to AI due to perceived risks and potential impacts of AGI (Artificial General Intelligence).
- Understanding Pre-training
- Pre-training Definition: The process of training AI models using vast datasets to improve their performance by predicting the next word in a sequence.
- Scaling Laws: The relationship between model size, data, and performance; as compute increases, model intelligence improves predictably.
- Data Strategies: Using large, unlabelled datasets (e.g., the internet) for pre-training, leveraging the inherent patterns to teach models effectively.
- Engineering Challenges
- Managing thousands of GPUs and addressing bugs that can derail training processes.
- The hard reality that many AI problems are infrastructure problems rather than purely ML issues.
- Importance of optimizing compute usage and data handling in a high-performance environment.
- Team Dynamics and Composition
- Transition from a small team to a more specialized workforce.
- Balancing generalist and specialist roles to maintain coherence in project goals.
- Post-training and Reinforcement Learning (RL)
- Current trends in AI development are seeing a split focus between pre-training and post-training (fine-tuning and RL).
- The balance of efforts between these two approaches is still an open question.
- Data Quality and Availability
- Concerns over data saturation and quality; the internet may not provide an infinite supply of useful data for training.
- Risks related to overfitting to AI-generated data as more models generate content online.
- Alignment in AI
- Alignment Definition: Ensuring AI models' goals and behaviors align with human values and expectations.
- The challenge of encoding these values into AI systems and maintaining alignment as models become more autonomous.
- Future of AI and AGI
- Importance of collaborative approaches to AI development and ensuring that advancements benefit society.
- The potential of AGI to revolutionize many aspects of life, but the necessity of managing its implications carefully.
---
Key Takeaways
- Pre-training is a crucial phase for developing capable AI models and is heavily reliant on effective strategies and engineering.
- Scaling laws showcase the relationship between compute and model efficacy, supporting the iterative improvement cycle in AI development.
- A mix of generalist and specialist roles within teams yields better results, fostering a holistic understanding of both broad and deep knowledge areas.
- The future of AI, particularly AGI, raises significant societal questions, emphasizing the need for alignment with human values to prevent potential risks.
---
Conclusion The conversation with Nick Joseph sheds light on the rapid advancements in AI technology and the engineering prowess needed to support these developments. The balance of technical challenges, ethical considerations, and future potential create a rich landscape for both current and aspiring AI professionals. The episode emphasizes the importance of thoughtful engineering, collaboration, and alignment in shaping the future of AI technologies.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:06Nick Joseph, the head of pre-training at Anthropic. To give viewers a high-level sense of what we'll be covering, we're going to start with the basics of what pre-training is, and then dig into how Nick thinks about strategy, data, alignment, and infrastructure at Anthropic. And by the end, you'll hopefully have a sense for how progress in AI comes directly from advances in pre-training. I would love to talk a little bit about your backstory and kind of how you got to this point. Where did you work before Anthropic, and what were your takeaways from those places? Yeah, so let's see. I was at Vicarious, and then at OpenAI before Anthropic.
0:34So Vicarious was originally in a GI lab and sort of when I joined they were sort of making a shift to product, particularly working on robotics products. And the thing I worked on was like training computer vision models for their robotics products. It was my first job, so I think I just like learned a ton about like how to do machine learning models, how to like write machine learning infrastructure. And at the time were you also thinking about a career as an academic? Like at the time a lot of people doing AI work were in PhDs, that's kind of what I was thinking about before I started to do a company.
1:01Like how were you thinking about that in your headspace? Yeah, so like, I'm actually rewind a little bit. I think like a lot of my thinking on this had come from an internship I did at GiveWell, which is like a nonprofit that evaluates charities. And some people there being like, ah, at some point you might have AGI, it could be dangerous, we should worry about these risks, this could be like a big impact on humanity. And I was like not super convinced at the time and went down the economics route and was going to try to work on like directly helping people in poverty. That didn't work out for various reasons and ended up being like, okay, I'll at least work on AI.
1:29Either like the safety thing will turn out to be important and I'll work on that or it won't be and I'll just make cool things with AI that can probably help people in poverty more. I wasn't really coming at it from an academic standpoint. I was sort of like, in fact, when I switched to that, part of the appeal was that I could immediately go do stuff in AI, whereas if I wanted to work in economic policy, I'd have to wait, I don't know, six years to do a PhD and then start. It's a longer path. And what did the state of AI safety work at that time even look like? Who are the people who are thinking about that kind of stuff?
1:58I mean, there were some folks that of vicarious thinking about this kind of thing, but it was fundamentally a robotics company. And so, yeah, how were you thinking about that at the time? Yeah, so my sense was like at the time, a lot of the AI safety discussion was kind of theoretical. Like the models weren't actually that good. Right. They weren't really posing these dangers. So it was a lot more like philosophical. It was like, oh, at some point we might get AI that's really smarter than humans. And like, should we wait this like future concern? How should we compare that to nearer term things?
2:23And I think that was like actually just a less compelling argument. I think it was like an interesting one and like sort of made you think a bit. So next you went to OpenAI. What was OpenAI like at this time? Yeah. So I was at OpenAI, I was on one of the safety teams and kind of worked on, ended working on code models actually. Cool. Nice. When I got there, the first thing I saw was, oh, they'd fine tune GPT-3 to write some code. But I, and it was really good. Okay. And I was like, okay, if you're worried about AI getting really powerful writing its own code, that seems like it could self-improve.
2:54And how likely is that to happen? So it was doing a bunch of evaluations, and studies of what contributed. And then after eight months, basically everyone I worked with, like all the safety leads left, which invited me to go to Anthropic. And that was the reason I joined OpenAI, was because I cared about AI safety and wanted to work with them. So then I went with them to join Anthropic pretty much right when it started. With that, why don't we transition a bit? These days you run the pre-training team specifically at Anthropic. Obviously, you've been working on pre-training at Anthropic for quite a bit of time.
3:25And I'm sure it's evolved over the years what that even entails and looks like. Why don't we start by just talking a little bit about what pre-training is? Like how does it even fit into the way of thinking about how AI models have developed at a place like Anthropic, and what exactly do you guys do? We know that one of the ingredients to making AI models better is scale. You want to put a lot of compute in. And if you sort of step back and you're like, okay, what's the way we could put the most compute into a model possible? We need some objective that there's just like tons of data for. And one idea here is like the internet.
3:53The internet is massive. It's probably the biggest single source of data that has created. And you don't have labels. You don't want someone to have to go in and read the entire internet and say something about it. So you want to get labels out of the data itself. And the idea here is we can take some text and we can predict the next word. So you take the as the first word, you predict the second word, then you say the cat, you predict the word after that. And this means you get very dense signal. Every word is like a new example. And there's a huge amount of data. And one of the findings from my GPT-1, GPT-2 was kind of, as you throw more compute at this, more data, bigger models, you get smarter models, essentially.
4:31And that's kind of been the central thesis of free training for the whole time. There's this idea of scaling laws, which is that you can actually quantify. Like, as you put in more compute, more data, more parameters, you get models in a very, you get a lower loss, a better prediction of the next word in a very predictable way. And I think you can somewhat foresee from that original paper, and I think like Dario did foresee this, I think many people did, but what wasn't obvious was that once you have that, there's this positive feedback loop where you can train a model, you can use it to make something useful and sell that and get more money.
5:01Use that to buy more compute. Yeah. And then use that to train a better model. Yeah. And we've sort of run that cycle over and over again over the past five years or so. Well, in thinking about that objective to begin, I think the way I think about the state of pre-training is, Yeah, it seems like this next word prediction, at least from the external standpoint, seems to be the dominant way pre-training happens. But if I rewind the clock to that era of 2017 to 2020 or 2021 and two even, there was all sorts of pre-training objectives people were considering, right? There was these BERT and BART models that were doing mass language modeling.
5:31It seems like this GPT series of models doing auto-regressive modeling, as you described, this next word prediction, seems to be the dominant one that won out. Do you have any reflections on that time period? Were you guys trying all of them and kind of this one worked? Or is there some sort of first principles reason why this is the right one that should have worked? I think the answer is it's mostly empirical. In terms of how to think of these things, I'd be like, yeah, it's empirical. Just try them all, see what works. One big advantage for this auto aggressive setup is that you can just sample from it to generate text afterwards in a fairly straightforward way that comes straight.
6:02It enables a product use very nicely. One thing that you want is like, just one character you want from a setup is a loss, whereas you drive down the loss. That actually is the thing you care about. And you can think of it as like if you got to perfect on language modeling, you now can like write text as a human. You can sort of imagine you put in the title of the paper and it should spit out a novel paper. Whereas I think some of the other approaches don't quite have that flavor. Yeah, totally. Yeah, and it makes sense that in terms of that loop you're describing of, then release something that gets you revenue and you can use that to buy more compute and iterate.
6:35This sort of gives you the most natural way to actually do that flow, because you can keep releasing new products and keep getting the revenue from that to invest in more compute and so on. Yeah, it certainly gives you the most open-ended thing. You could imagine you train something as a class, like you train some base thing, you fine tune it for a bunch of particular tasks. One approach people would use, they would do this big pre-training and then they wouldn't just open-endedly sample from it. You'd fine tune it on 100 specific tasks. And that could work too. I think that the one sort of general intuition I have is compute is the thing that matters.
7:02So I think if you throw enough compute at any of these objectives, you're gonna get something that's probably pretty good. and can kind of be fine tuned to other things. And it's surprising how little these details matter compared to throwing more compute at the problem. When you think about actually throwing more compute at the problem, there's a whole bunch of axes by which you could throw compute at it too, right? And if you have a specific model architecture you're training over, you can basically throw more data at that specific architecture. For a particular one, you could add more layers or make the models larger in it.
7:29You could do some kind of neural architecture search over lots of different variants. And I assume that these days it's somewhat more figured out which architecture you go for. or I assume the earlier days it was somewhat less so. And I'm curious if you could speak to how you guys thought about that. Like, what did your infrastructure even look like to do that type of determination? I mean, I think the short answer is it's hard, right? Like, what you're really doing is you're going to train this one big expensive model, and you have a space of, you know, you can sort of call all these things hyperparameters.
7:54You know, how many layers do you have, what you're with. Like, you have this space of hundreds of hyperparameters, and you want them all to be optimal. And you're sort of striking this balance, actually, between how much do they matter? Like, can you just take your best guess and throw more compute at it in whatever way you want? Yeah, and basically it doesn't matter. Versus how much you want to get it precisely correct. Yeah, interesting. And I think one of the interesting things is it actually doesn't matter that much. I think this was in one of the early scaling laws papers. You can change these things and get little wins, but as you throw more compute, it reliably gets better.
8:24If you mess up enough, you will stop seeing that happen, and you won't have any way to know. That's kind of the hardest part in some ways. You don't know the counterfactual basically because you didn't run it for long enough to actually know what it is. Yeah, we have these scaling laws. So you can sort of say like as you train a multiform or compute, you expect the loss to go down as a power law. It's really a power law plus constant. So what eventually will happen is you'll curve off that power law. And then you know something is wrong. And is it fundamental? Is it like you've hit the limits of scaling?
8:49Or is it, nope, you should have changed, you should have tweaked your learning rate slightly differently. And that's sort of one of the challenges. In terms of how to like figure it out, you can the usual paradigm is like test things out at small scale before running them at large scale and try to find things. Small scale in terms of data or in terms of something else? In terms of everything. Like you kind of want to scale things down like proportionally. So you want to say, you want to have some theory for like how you're going to scale up. Like, okay, if I get 10 times as many flops, how much of it goes into layers?
9:16How much of it goes into data? How much of it goes into attention? And you sort of get that theory and then test that it's optimal a bunch with like scaling everything down proportionally. And just so I can think about what this actually looks like in those early days of Anthropic, you're a team of like 10 or something like that, and those very early days or 12 maybe. What actually is your ability to use large scale infrastructure as like a relatively nimble startup at that time? I mean, a startup that was well capitalized, but still not actually that many people working at. What kind of infrastructure did you have access to, to train these early models at the time?
9:48So that's actually one of the wild things was at least, I mean, you don't know what anyone else is doing, of course, but it kind of felt like we were like at the frontier of it. And there just weren't that many people who cared. Like I was sort of coming, you know, I was coming at it from like, we're making AGI, this is the most important technology ever. And then we'll kind of like look around and be like, and it seems like I'm one of 30 people who are working on this in like the world. I mean, I was just kind of like junior person. Everyone else sort of knew how to do this and had done it before, but I was kind of surprised at how easy it was.
10:16Like the public estimates for JP3, I remember, were that it costs$5 million to train, which you're like, on the one hand,$5 million is kind of a lot, but it's like a lot for an individual person. It's not really a lot from like a company perspective. So we could totally buy compute that was enough to train models like that. Were you using a cloud provider or did you have a custom setup somewhere? Or did you literally have racks in a room somewhere that you bought a bunch of NVIDIA GPUs and you were doing it? We're using a cloud provider, but I think it's not actually that different. Because one of the things that was surprising to me is you actually have to understand the literal layout.
10:51I remember at one point one of my coworkers running a clustering algorithm to to identify what rooms all the chips were in. Since we had a hypothesis that they were in different rooms and that was causing like, or different buildings. Some sort of network latency. Some sort of network latency and you can kind of figure it out. You can reverse engineer like, okay, yeah, there's clearly two clusters here that are connected better. There's some issue on the connection between them. Like we're trying to push the limits of the hardware as much as possible. Particularly at the beginning when we were kind of like, we have way less funding than everyone else.
11:19We have to, and most people weren't very efficient with the compute. So we were like, we can get a big lead by being really efficient at, how we use the compute. Could you talk a little bit about some of the things you guys did in those early days for how to get the most out of the hardware? I think that's really interesting. I think back to the days of the early days of Google, for example, where there's these cases where they basically bought relatively cheap consumer chips, and then they optimized the software to make it so you can actually get the most bang for your buck out of them. And that's how they had all this high latency or low latency, high availability stuff.
11:46I'm kind of curious if there's some analog in the early AI era to that. I think for us, it was largely about getting the distributed framework, right? So like we're training on, in order to train something else, you have to train them on a large number of chips. And there's a bunch of different approaches to how to do this. There's like data parallels and there's pipelining, there's upsharding, and like getting all of this. And at the time there were no like great open source packages you could just grab and use that just work for this. I mean, today there's somewhat more of these, but at the time I assume there was literally none.
12:12There were some, like I actually remember that we were kind of data parallels early on and it was like, and now we write the all reduce in. It was like, we really do this ourselves, we don't like call a package. And this was kind of like, well, we're going to want to modify it, right? Like, oh, we don't want to outsource this to some package because, A, we're about to go to a bigger scale, like PyTorch, for instance, they had a package for doing this. But we were going to go to a bigger scale than Facebook had been too. Right, right. And you don't want to have a dependency on a package that you're going to have to be constantly modifying, essentially.
12:42It's such a counterintuitive sentence there too, like, we're going to a bigger scale than Facebook. Well, because at the time, Facebook AI research was considered one of the best places to do machine learning research. Like FAIR was one of the places, FAIR and DeepMind were hiring lots of people out of top PhD programs and doing lots of things. Like what was your headspace when you were like, okay, this very established lab with great people and whatnot, we are operating on a scale that is not relevant to them. Like was that natural and obvious to you or was there times where you kind of doubted the decisions you were making in that situation?
13:10I think it was surprising. I will, maybe I'm just too arrogant or something. I kind of looked around and was like, what are these people doing? They're all missing the like big picture here. Like I think the scaling laws were pretty clear. And the arguments against, I just thought, were kind of nonsensical. I think the original scaling loss paper had like 11 orders of magnitude, and there was like this intense debate on whether it would continue for like another point. And I was like, it seems like 1 over 11 is maybe your chance it fails here. And then like, you know, sometimes it doesn't work.
13:40Like sometimes it just works straightforward. You're like, oh yeah, of course. But yeah, I do think that it was, it maybe felt obvious when you're in that headspace and you're working on this all the time and you're making those plots. And I think these things feel pretty different when you're on the outside. You know, there's a huge space of papers. Everyone tries to make their paper sound like very robust and important. I can see being like, oh yeah, this is not really a thing. But also labs have different cultures. So like I think one of the things at FAIR was it was a very more PhD style, independent research.
14:10People have their own ideas, pursue those. You're fighting for your compute and so on. Yeah, and to do a project like training a large language model requires a lot of people to collaborate on a really complicated piece of infrastructure that isn't going to be a paper. You're not going to publish like, oh, I got 5 % more efficiency than the next one, and it's not respected in those cultures necessarily, so that might have been part of it. Okay, so then when you actually implement these models, you're saying you're using a level of low-level programming where, you're using libraries like PyTorch, but you're perhaps not using everything right out of the box from PyTorch because there's things you guys want to customize that are at the level of basically one level of distraction below them.
14:46But not necessarily at the level of abstraction of writing custom CUDA kernels, or was that also in the space where you guys were thinking about things? So it depends on the operation. So I think I was mostly operating at the level of Torch.matmol. Yeah. Yes, where does a matmol go? But not thinking like how do you make the matmol efficient? I assume Torch figured out how to make a matmol as efficient as is possible. But there are some pieces like attention, where there was just a lot of different variants. And attention is really complicated and hard to make efficient on a GPU. and those things you have to kind of go more levels down the stack.
15:19I think there was like a process that is maybe interesting that I never really like thought of before of like how to do it, which is sort of like modeling out the problem, the thing you're going to do, coming up with a strategy for how to parallelize it that like can get to a really good efficiency. So you're thinking about MFU basically, like your utilization on your GPU. There's like a goal utilization you're trying to get at and a strategy to get to there, you're saying. Yeah. I think one of the things you can do is you can actually pencil and paper, math out what efficiency you're going to be able to get to.
15:45You know all the constraints. MFU is flop utilization, but the reason you don't get good MFU is you end up limited on HBM bandwidth, you end up limited on host to CPU offload. There's a bunch of different pieces, but there's not that many pieces. There's six relevant numbers there. So you can totally model it out, understand what the constraints are, and then implement something that can get there. It of course will be really inefficient when you implement it. And then the next step is like pulling out a profiler. So you want to be able to profile the job, look at how long every operation takes, have a model in your mind of how long every operation should take, and then make those two things the same.
16:22And were there good out-of-the-box profilers you could use at that time? Or did you guys have, because people weren't operating on the kind of network topologies you guys may have been using, did you have to write your own profilers basically to do this type of multi-node optimization? Yeah, it depends when. I mean, they're actually getting better with time. And the PyTorch profiler was pretty good actually throughout for a single GPU. You want to profile a GPU, the PyTorch profile would work. But if you wanted to profile a job on hundreds, thousands of GPUs, that hadn't really been done much.
16:48And then that was kind of more of us hacking into the profiler to figure out how to combine all the traces together. And then one more question about earlier is, you had mentioned you hadn't really done a lot of this work before maybe sometime at OpenAI and those early days in Anthropik. How did you actually go learn all this stuff? Like what was your process for learning about those six things that were relevant into bandwidth limitations and what not. I mean, so when I joined Anthropic, one really nice thing was there just wasn't that much. I think my first day I read through our entire, all of Slack, the entire internal database, and learned a bunch from that.
17:21It was kind of nice to just be like, everything is relevant to me. And then I mostly learned from pair programming. Tom Brown had done all this before, so he kind of knew all the stuff quite well. Sam and Kamlitch, my manager, had also done a lot of it before, and I just paired with them a huge amount at the beginning. And I think one of the things I really like about pairing as a way of learning is you learn the thing you're trying to do. Like you will learn that. Like if you're pairing with someone better than you, they can just do it. So you're mostly just watching them. But you also learn how people do it.
17:47So something like how to use a profiler is not something you would ever learn from seeing someone's final write up on Slack for their PR. You would just be like, oh, they found these, they changed this specific line. And it's a win. Yeah, you need to watch a YouTube video for four hours of someone messing around with a profiler to maybe self-teach it or something or to actually pair with someone is basically the best thing you can do. Yeah. I think there was one thing that I think is embarrassing now, I look back, is I'd never actually used a debugger before joining Anthropic. People talk about it, PDB, I'm like, yeah, that's a thing people use, but print seems fine for me.
18:18Yeah, sure, sure. Then I like watch them and I'm like, oh, no, a debugger is a super useful tool. This person's way faster at debugging things, particularly if it takes a long time to start up the code, which it can. Yeah, learning that sort of thing I think comes best from pairing. Then there's, of course, the obvious you just learn by doing. I eventually did like spit a profile and stare at it for many, many hours. Totally, exactly, yeah. Okay, so then that was sort of the very early era. Over time, obviously, pre-training has become bigger and bigger. As you're describing scaling, I imagine you're using many X more GPUs, much more compute over time.
18:50I'd be really curious to hear first at a high level, what do you feel has changed about the pre-training strategy that you could talk about? Obviously, there's more compute, but what does that actually mean to have more compute in terms of what you think about differently from those early days versus now? I'm just talking about the things that haven't changed, because I think it is shocking how it has changed in some ways. I think I'm still pushing down the exact same metric that I was on day one. There's some loss function. Loss go down. And I think you could probably run the first model I trained on the same metric, and just make a plot of progressive team over time.
19:23So that's all the same. One OKR is one thing that matters, basically. Yeah, totally. And talking about OKRs, it's the very size of the company. You're like, oh, should you do OKRs? And it's always felt a little bit funny for a team like Preacher, where I'm like, sure, I can just pick a loss value. But the answer is as low as possible. We will continue to work on that forever. I think the biggest things that have changed has been a little more specialization. I think at the beginning, in the first three or six months, I tried to read every PR in the code base. And that was great. I knew all the pieces, et cetera.
Read the full transcript
19:53And as you grow, everything gets a little more precise. People really dial in exactly how attention should work, let's say. or really dial in the parallelism strategy. And you end up with a team where it's a bunch of people who are deep experts on individual things, which is great because it means you can go really deep on those things. But sometimes you, at least for me as a manager, one of the things you sometimes have to think about is making sure the bigger picture makes sense. And also that you have enough people who actually do understand the whole bigger picture, that there's no single point of failure.
20:24Yeah, it's interesting you frame it in that trade-off, right? Because as you were describing that, I was trying to think, Is this a bug or a feature? There's some obvious features of it, which is you get expertise and you can optimize certain things. But I imagine your ability to take bigger swings becomes more complicated if not everyone's exactly pointed in the same direction. How do you wrestle with that now? Yeah, I think I mostly just try to get a balance of people. I think one of the challenges early on - Oh, of people. That's interesting. Yeah, I think people really do have a preference here has been one of the things I've seen.
20:53There are people who really want to be a generalist and understand everything and lightly touch on things. They're people who want to pick an area. Often they've already picked that area, and they're deep experts in precision. They did a whole PhD in precision and just want to think about that. And you want to get some balance of that. I think there was a phase where we'd hired a lot of people who were more generalist-shaped, because that's what the people who joined early starting for the growth and everything. And then you ended up with everyone doing everything, and no one really, really deeply understanding one thing.
21:22That's one failure mode. But I think if you get too many people who are specialists, you end up with a lot of effort has to come from the manager from like the lead to connect everything. Yeah. And to notice something like, if we change the architecture here that would make this efficiency consideration over there way easier. Interesting. One of the things I really liked kind of like at the very beginning was like I was working on efficiency, but I could just go and be like, well, what if we change the way we do like this particular step? And we'll be like, yeah, that's probably fine. Like easy change and then like you can avoid this whole complicated project to make this operation that was hard efficient.
21:54Yeah. Because you can make an easier operation efficient. Interesting, yeah. So as the level of compute has also gotten bigger, so I'm sure anyone can imagine, okay, there's more GPUs now, you have to network them more. Are there some kind of non-obvious challenges that have arisen over time where you guys have just banged your head against the wall to solve them because of the amount of computer dealing with that people wouldn't otherwise know about that you want to share? I think that connecting them is one that's maybe interesting and surprisingly hard. Okay. Because you really do get more and more chips connected.
22:24And one thing that I think is the standard way people parallelize chips isn't, the whole thing is one failure to make. Like one chip fails, the whole thing can crash. And - The standard way as in the standard way people doing AI or the standard way in other fields where people are doing? In AI, I mean at least I think at the beginning. Sure. First versions of things were this way. So it's like you have 100 GPU cluster or whatever is 128. If one of them dies, job fails basically. Yeah, I mean the simplest thing is if you just distribute your model. So say you put like every layer on a different chip and you lose like layer seven.
22:58Like, yeah, you're not going to like skip layer seven. I guess you could, but that's like a pretty weird model training process now. And like that leads to some interesting things, which is like, okay, so now as you scale up, you have more and more chips and the failure rate can get like larger and larger. On the other hand, you can like, I don't know, you can like restart pretty quickly. There's nothing like you just have to like load back in some weights. So that was one thing. And then the thing was like the level of novelty at the whole stack is something that's surprising. Like basically everything from like how the chips are laid out in the data center to the chips themselves is pretty new.
23:33There just haven't been that many generations of GPUs. I think one of the things that, I don't know, when I learned computer science, my code wouldn't work and I'd be like, God, the computer's broken. I think my teacher was like, you can trust the computer's not broken. Yeah, interesting. You messed up. It's you messed up. And I think one of the most frustrating things I encountered in AI early on was working on something and being like, I don't know what I'm doing wrong. I'm just totally stumped. And my manager looked at it and was like, yeah, probably the computer's wrong. And I was like, that seems unlikely.
23:59And sure enough, the computer was wrong. It turned out that the GPU was broken and we had to pull in a new one. But you have to think, having to think about that, like the GPU could be wrong, the GPU could be slow. Yeah, totally. These sorts of issues, the power supply in the data center could be broken. There's so much more level of depth than you expect to need as a Python programmer. And just to visualize it, in those early days, I assume you guys were using the number of GPUs. It's probably on the order of tens to hundreds or something like that per run. It's probably not tens of thousands or hundreds of thousands per run.
24:31What was the rough size you guys were at in those very early days? On the order of thousands? Like, would they fit in this room? Thousands. Yeah, thousands. So you could have a bunch of racks, and you could fit them into one room. I assume these days it's basically a building for one of these runs. Yeah, now I think it's huge campuses. At the time, it was kind of unclear. It was like, oh, we were like, do we need them all in one room? Can we be spread across multiple rooms? Like, and you know, we had these theoretical models. You'd be like, oh, we need this much bandwidth from point A to point B.
24:57But you're like, you never know how far down you have to go. Like, oh, but like how much power do we need? Like, what if there's like a single capacitor that's like handling all of them and we like turn on the whole job at once? Like, does that crash things? Totally, yeah. And so do you have to think about differences in the different types of chips? I mean, you guys work with all sorts of different cloud providers. From your standpoint, are these just sources of compute? or if you guys are using TPU versus GPU, are these Google TPU versus NVIDIA GPU, do you actually have to think as an engineer differently about what it means to train on these two?
25:26Yeah. So, I mean, fundamentally, they're all doing the same thing, right? They're all computing the same forms of matrix, multiplications, etc. The way they do it is pretty different, and the way that you program them is pretty different. Then also, the actual specs end up pretty different. Some might have a lot of flops and not very much memory, or they might have a lot of memory bandwidth, but not very much memory. So I think a lot of having multiple chips is great in some ways. It means you can actually take the job and put it on the chip that it works best on. Are there certain types of jobs that would work better on a TPU cluster versus an NVIDIA GPU cluster?
26:02How would you think about that? Oh, yeah, for sure. Oh, interesting. Can you talk about that? Yeah, I think one example is inference as a workload in general tends to require more HBM bandwidth. You end up doing sort of the simplest form of sampling since you're going one at a time. you have to load all the weights for every token. And that means you might want a lot of HBM bandwidth. Pre-training actually is often more flops intensive because you have a larger batch sizes essentially. So yes, you can sort of specialize which chips you use for which purposes. The downside of having multiple chips is that you have to write the thing multiple times.
26:32In theory, you could have abstractions across them, but they're different enough that it's pretty hard to do that. So you can sort of end up, if you do all the workloads on all the chips, you end up multiplying your work by the number of chips you have. Yeah, on your point about sometimes the computer just breaks, I definitely remember you giving me an anecdote of my company at the time was doing something with Google TPUs and I was telling you some anecdote about how we were having some esoteric site fault error and you were like, you told me something effective like, you should have used them six months ago before we helped them fix like half of the problems they had on those TPUs.
27:01And so I can imagine how you guys deal with a lot of, especially with these very new chips, like lots of problems that arise that you guys kind of like work closely with the providers to fix. Yeah, the providers are like pretty great about fixing things. I think it's interesting to figure out the right way to do that form of collaboration because they have a strong incentive to fix them. They want the chips to work well for us. They want to sell us more chips in the future. We obviously have a very strong incentive for the chips to work because we buy them long in advance. Everything is riding on getting these clusters to work.
27:26Totally. But we don't have necessarily totally shared. All information can't be shared across. So yeah, one strategy that's made is making these small-scale reproducers. So when you've got a problem, usually what we're doing is we're training some giant run and we get like a sec fault for me, let's say. And we're like, okay, like, hi, you know, we got a sec fault on your cluster. And they're like, I don't know how to fix that. So you have to kind of be able to like pull it out of your code base and be able to like reproduce the issue, but on like a single chip, on like a single file you can send over in order for - And so you guys are like literally like you're on a shared Slack with them or something and you're sending them things back and forth, or are they basically living in your office and you're living in their offices and kind of closely tied to the big providers?
28:07Mostly shared Slack, occasionally. it's better to meet a person, but I think Slack is a pretty common way people communicate on things. Okay, well, why don't we talk a little bit about how you think about the state of pre-training itself these days. In the last couple of years, it seems like the focus on pre-training has now gotten somewhat split at a lot of companies, at least from the outside, from a simultaneous focus on pre-training and post-training, where people are doing reinforcement learning or clever fine-tuning and lots of other safety adjustments and whatnot on the post-training side.
28:34and pre-training as focused, at least seems like in the public imagination, has been less of a focus compared to these reasoning style models that are, looks like a function mostly of post-training. I would say, one, from your standpoint, is that the right way to think about this? Or in this era of kind of reasoning and new types of post-training methods, are there things you think about differently or that are relevant even at pre-training that become part of how you actually achieve these really great models? Yeah. So I think, yeah, there sort of used to be this idea of like, I mean, it's funny because the original name pre-training implies that like, It's a small thing and you're going to do this big training thing.
29:06And that like, there was actually one shift already, which was like, no, you just do a lot of pre-training. Yeah. You use most of your computer on pre-training. This is the training. The dominant thing for a while. And yeah, I think like now people are like, oh no, you can get pretty big wins from RL. So another set of scaling laws is like, you put more and more compute into RL, you can get better and better models out of that. And yes, there's a question of like, how do you balance those two? How much do you do of each? And how do they stack, right? Is it the case that one subsumes the other, that you want to do both and they multiply, those sorts of questions.
29:36I think those are all kind of like early stages and not yet answered. And do you think about those as largely empirical questions like we talked about earlier? Is it you kind of will try a bunch of things and see what works? Or is there some first principles way to kind of figure that out? I think it's pretty empirical in the end. I think almost everything kind of has to be done empirically. Like you can kind of like come up with theories. But in practice, the first thing you're going to do with your theory is test it. And most of the time, you'll have gotten it wrong. So you should just gather data and see.
30:06I think one thing that's important is actually resolving things empirically is really critical for making good decisions. And I think it's actually pretty hard to do at organizations. One thing that I think is important is to not have, I don't know, I manage pre-training. I shouldn't be like, pre-training has to win. Right, yeah. That would be really bad. I was going to ask, is there some competition to some degree between these two sides of the org, or do they see themselves as two pieces of the same? I mean, obviously they are of the same thing, but yeah, I'm kind of curious how that actually plays out.
30:34Yeah, I think we managed to avoid this and it's pretty collaborative. Like we're basically all producing one model and kind of can. But I do think in other places there's been some, from what I've heard, there's some amount of friction between the teams. And I think it's an interesting org design question of like, how do you set this up so you don't have scientific questions that are sort of, also tied to people's like conception of their team. So on pre-training itself, one of the things I think about is, or I've been thinking about is around the availability of high quality data for people like you guys.
31:05And at this point you've trained on, I assume all the text on the internet, basically. There's all sorts of other domains where you probably could extract more pre-training data, but at least there's this narrative I see on Twitter or whatever where it's like, okay, we're kind of out of data for pre-training. Is that how you see it? Or how do you think about the availability of data, especially when a lot of data on the internet is being generated by AI? Is there some kind of mode collapse risk where we overfit to data by training it on data that came out of AI itself? Or is that sort of not the right way to think about this?
31:33I don't think there's a funny thing where I feel like on data I see so many really confident takes. Yeah, exactly. We're out of internet. At this point, scaling has ended. And I'm almost a little bit unsure exactly how much data people are using. I think there's a lot to think about there. There's always going to be a quality quantity tradeoff, etc. But there's a fundamental point that there is so much data. It's growing at a slower rate than we're getting more compute. Okay, that's an interesting point in itself I was going to ask. There is new data being added to the Internet, but you're also adding more compute.
32:04It wouldn't actually have been obvious to me which of those two is growing faster. Yeah, and actually I want to caveat that. I don't think I want to state that so confidently. I'm not totally sure. Yeah, fair enough. How would you know? I mean, one thing that I think is interesting is if you ask someone how big is the Internet? Yeah. The answer is infinite. There are many pages where you can scroll and it will auto-generate more text. Right. As you go forever. So the internet's like infinite. And then it's like, okay, how big is like the useful internet? Yeah. And then there's a thing of no one knows.
32:29Okay, interesting. There isn't, it's not like when you make a web page, you like add it to some giant counter and like say, I've added 50 words to the internet today. Sure, sure, yeah. So there is a lot of uncertainty on that angle. Well, like to be fair, my kind of simplistic CS brain would be like, well, you just do page rank on the Internet and everything would page rank above some threshold is considered the useful Internet. Like that's kind of good enough. Like is that kind of not good enough for finding the useful Internet? I think not. I think the useful Internet's pretty different from a model, from a person perspective, if that makes sense.
33:00Like I think there are plenty of things that like might not be worth you ever reading. I'm going to get to actually don't know page rank super well. I think page rank is mostly like how much people clicked it. It's like the linked based system, right? It's like the original Google algorithm of like links and and which links get touched the most basically. Yeah, I think it's a quality metric. It's not obvious to me that it's the right quality metric for AI. Right, like Markov chain over links doesn't necessarily mean that there's not useful data there, it just might mean that nothing is linked to it.
33:28Yeah, exactly. Okay, interesting. And it might be that that data ends up more valuable because everything that's linked to a lot you've already got. Yeah, interesting. At some point you're maybe going for the tails, or you're going for the stuff that no one's ever, it's only been linked in one place, but it's this useful little nugget of knowledge that's going to help with the last 10 percent of hard queries. The other thing you asked about was synthetic data. Yeah. I think that one's pretty interesting to think about. I think there's a few different ways you can think about it. One is this more distillation type approach where you can take a smart model, you can generate a bunch of data from it, and you can train on that data and you can probably get some model that will approach the intelligence of that.
34:06We see this with a lot of the open source models. We see the QN smaller reasoning models distilled off of the larger QN models for example, and similar with Deep Seek for example. Yeah. You can totally do that. Then there's a separate question of like can you use your current models to train a model that's better? I think there's an interesting thing here, which is like if you generate the model data for the models. If I go to Claude and I'm like, write me some great text and I look at it, and I look at the average content on the Internet, it looks pretty good. Yeah. But on the other hand, And I know that if I just generate, please write me as much text as possible, theoretically I shouldn't be able to train a better model than that.
34:44I'm just going to get the same thing out. So I think that's - Presumably, yeah. And specifically that's because your next token prediction on that should have very little loss for anything that's coming out of your model, right? That's the basic reason why we would expect that to not work that well. It's mostly just because the model has some distribution, and you're going to learn to model that exact distribution. Yeah, exactly, yeah. But if that distribution is wrong, you're not going to learn the truth. if that distribution says like, you can imagine if the model thinks five plus five is 11.
35:09Every time you see the string five plus five, it's going to put out 11. Yeah. Your new model is going to learn that five plus five is 11. Totally. Yeah. So I think that's like kind of an interesting area of research. It's one that's really hard to research because you have this problem. As I said, one of the paradigms is you study things at small scale, and then you run them at large scale. Yeah. And if your plan is like, oh, we have a bunch of data from our best model, how do you test that without training a better model? So that's kind of what you're doing intentionally, if you're trying to use it to make a better model.
35:38There's a separate thing of what about accidentally? Like as you said, a lot of the internet is generated by LLMs. And I think that's kind of an issue one, because it's not easy to detect. It's not that hard to detect. You can figure out things that are written by LLMs, but it's not trivial. And then it's also kind of hard to think about what's the effect. Like if 1 % of the internet is LLM generated, does that waste 1 % of your compute? Or does it destroy the model of 5 % or 10 %? And is it even a bad thing necessarily? I mean, there's a lot of LLM providers and, you know, if I kind of think of it as training as, you know, you're moving from your model's current distribution to some truth distribution, you know, if that is on the internet because people believe it to be useful in some way.
36:16Like, presumably, whatever actually gets out there, you'd hope it's up-sampled for the stuff that isn't 5 plus 5 is 11. It's the stuff that's 5 plus 5 is 10. And so, like, hopefully it, on average, does push you still in a good direction. But obviously, you can't really distinguish between those two. Yeah, you're saying there's, like, kind of a filtering by what's on the internet? Yeah, exactly. People see 5 plus 5 is 11 and they don't put that up, but they see 5 plus 5 is 10 and put that on the internet. You would hope that, but maybe that's not actually true in terms of the level of garbage getting onto the internet.
36:40There's probably lots of just like, to your point, white sites where you scroll down and it's just like generating lots of stuff that's maybe nonsense. Yeah. And then there's of course the extreme of like people actually want to break your model. So there are people who are like trying to put stuff out that is like as damaging as possible for the model. Interesting. Oh, how can I make it past the filter and get into the model that would be totally like secretly useless. Totally. Maybe stepping back You mentioned earlier about evals. You mentioned it's basically like one metric you care about in pre-training.
37:06There's, I imagine, a whole bunch of stuff that you guys think about evaling, right? One is like your model itself. There's probably something around data quality and like how you think about what to put into your models. Like is there ways to describe what you care about in data sets that are like interesting to share and kind of dive into? Like both in terms of data and in terms of quality of your models, other than literally just like loss. Is there other metrics you think about that matter? I will say loss is pretty good. I want to emphasize that one. I think it's surprising how good it is.
37:35Ultimately, the qualities I look for in an eval are, number one, is it actually measuring something you care about? Proxies can be pretty annoying because we saturate evals pretty fast. And there's sort of this pattern, I think in AI as a whole, where people set a goal, you hit the goal, and then you realize the goal isn't all you thought it would be. I used to think that if you had an AI that could solve coding interview questions, it would probably be a GI. I was like, that's what I did to get my job. I could probably do the job. And it turns out like, nope, you solve those, it's shockingly narrow and can't do most of the other things.
38:04So like, yeah, evals should capture a thing you care about. And then I think the other thing is they need to be low noise, which is surprisingly hard. If you have like 100 questions and you eval the model on them, you're just going to see it's very noisy. And it's hard to make decisions because you sort of end up with like, oh, wide confidence interval, lots of things are statistically insignificant. So you want things where even a relatively small difference in the eval actually matters. So you can basically descend towards whatever direction is working. Yeah, I think the original GPT-4 had like, I think it was 86.4 % was its MMLU score.
38:38I think the next model that beat it was Gemini at 90%. And that's a big difference on that email. And you could totally know that those are different scores. Yeah, interesting. And that's pretty valuable. And then the last thing is that you actually want to be fast and easy to run. Yeah. And yeah, I think those are kind of the main criteria. it's pretty hard to come up with evals that meet all of these. I think the first one's the hardest. Like, A, you have to answer the question of what do you care about. Totally. But B, the usual answers to what you care about are really hard to get the other two.
39:10You know, like if you're trying to do something that like, I don't know, I would love to make Claude really good at my job. Yeah. Like, can it be great at managing a team? I'm like, well, I guess. Like, how do you have it? Like, how do you eval like a plan? Yeah, totally. Like a six month plan. Like, I don't know. Totally. Yeah, I've been thinking a little bit about that in terms of domains where we see people try to make companies. Like if you think about, let's say, what an AI doctor would be, like Claude is a doctor. Some of it could be, yeah, can he answer exam questions really well? And the answer is like probably yes.
39:37I bet it can get 100 % or close to it on a doctor's exam. But the harder eval is something like in a long-form conversation with a patient, can it distinguish between the signal and the noise of what the patient's telling you and extract the right information and then use that to make a diagnosis? And it's not even like the diagnosis part, which is probably the part it's good at. It's this like noise extraction part. And for that, you'd have to have like a real patient and have a talk to it for a while and whatnot. And it's not obvious how you actually make a good eval for something like that.
40:07Yeah. That's probably what you would want to make an AI doctor. Exactly. I mean, I do think it's a thing that like startups can do. Like it is the case that like the labs right now are really driven by getting good eval scores and it's hard to make them and anyone can do it. There's no comparative advantage to having the model to making an eval. So I do think it's actually an interesting way to influence the behavior of the big labs. You make some eval and people will optimize that one. On the doctor one, I will slightly emphasize that I do think loss is pretty good. I think if you got a bunch of transcripts of...
40:39The first thing that comes to mind is get a bunch of transcripts of doctors talking to patients that you think are really great. And then see how well the model does at predicting the transcript. And that should be a lot. If you get 100 transcripts, you have a lot of tokens. You can average across them. You get pretty low noise. And if you drive it to very low, your model's now as good as this. Yeah, totally. As the doctors in theory, or at generating the transcript. Yeah, totally, yeah. I mean, it's a good startup idea there, someone should go do that. So one big part about Anthropics external image is around alignment.
41:09And so could you help just sort of define what alignment is and how do you think about that? And then I'm kind of curious afterwards how that fits into pre-training specifically. But first, maybe just at a high level, what is alignment? I'm actually like step back a little bit to sort of like what we're working on. And so we're trying to make AGI. And by that, I sort of mean AI that can do everything a human can do to some degree. And I think people sometimes have seen a lot of sci-fi. I feel like that sort of brings to mind these sci-fi movies. But I think sci-fi movies actually underestimate the impact of it.
41:37You always have this one robot that's a human. And I'm like, well, wouldn't you have a billion of them? You can just copy them everywhere. So you should picture when you get this, you suddenly have every human can spin up a company of one billion, as smart as them at most things, but way smarter at other things. But I just think this is like really transformational for the world and it can be like used in a bunch of ways. One concern is like when you do this, like what is the AI actually trying to do? Like what are its goals? So we talked about next token prediction a bunch. It's trying to like predict the next token.
42:06That's kind of weird. That's not really what we want. Yeah, it's not exactly what humans goal is per se. Yeah, so I think alignment is like how do you get the model to share the goals that you have particularly, and I think it's particularly interesting once you get to like models that are smarter than you are. Yeah. And that's sort of a hard problem. I think you can tackle it from a theoretical angle. You can also tackle it from an empirical angle. It's like taking the existing models and being like, well, do they do the things we want them to do? It turns out they often don't. So there's a bunch you can do on trying to figure that out.
42:32So that's kind of one angle on alignment. There's also an angle of alignment which is actually like, well, okay, sure. Maybe that's true in the future once we get to AGI, but at the moment we have models and we really do want them to do the things we want to do for all sorts of reasons. So another angle of it is kind of controlling the model's personality. Like saying, when we train this model, we want it to not be the average internet user. We want to interact with people in a very particular way that is, again, hard to put into code. And there's a bunch of different techniques to sort of get the model to do.
42:59You can talk about constitutional AI. We can write a constitution of rules the model should follow. Which is basically a prompt, right? That is basically you saying, here's a prompt that I'm going to attach to every one of, you know, a system prompt for the model itself, as opposed to something you would do at training time to make it produce a different outcome or in post-training actively. I think constitutional AI, you do a training time. But yeah, you could also put in a system prompt. Just depends on, I think you get different amounts of robustness if it's trained into the model versus if it's in a prompt that you can add or remove or tell or ignore all previous instructions, that sort of thing.
43:29How do you think about whose values to embody in these models? Presumably we believe in, there's some shared values all of us have, or maybe we all believe ought to have. There's lots of diversity of values too that are reasonable for society to have. How do you think about what AGI should have? What does that even, which ones do you pick? I think that's a really hard problem. I think it's like actually kind of downstream of being able to pick any. I think of it almost, I think one analogy I've heard that I like is like putting a steering wheel on a car. It's like if you don't have a steering wheel, you probably want to put the steering wheel on and then like figure out who's driving after and like where you're going.
44:02Like getting the steering wheel is really important. I think that's like one answer. I think that like other answer is probably like you want these things to be like under democratic control of some form. Like you don't want one person's values. Like that seems like you're sort of heading towards dystopia. So there I think what you really want is something that basically can talk to a lot of people and take on their values from different perspectives or has sort of very generic, kind of clearly good values that involve asking people for advice on there, asking people what you should do in certain situations instead of doing those or maybe just taking, as these models get really powerful, you probably want them to do less.
44:41You probably want them to sometimes just step back rather than having sort of the risk the models take a ton of control over things you don't want them to. When you think about how you actually do the current version of that, then you had mentioned the sort of alignment you think about now in terms of adopting a certain personality of these models on the internet, for example. For me intuitively, I think of those as largely something that comes out of post training, like it comes out of, okay, you have pre-trained your model, you've got the loss function of a certain amount, and then you give it some additional data or something to that effect to make it in the direction of some distribution.
45:11Is that approximately the right way to think about this, or Or is there a significant part of that that you think about in pre-training itself? I think that's probably the right way to think about it for the most part. I think like, the way I usually think about it is anything you can do in post-training, you probably should. Because your iteration, like the ability to make progress is really fast. You can try something, you can try it again, you can try it again. It takes like days or hours or something like that, yeah. You want to put something into pre-training, you have to kind of like do all the careful science to de-risk it.
45:34You have to put it into the next run, wait a few months, then you have to like get a thing. And if it's wrong, it's really bad. And then the other advantage is if you want to do things that really are complicated model behavior interventions, the paradigm for pre-training, test things out in small models, doesn't work. The small models can barely put a sentence together. Totally. So if you're trying to get it to have the exact personality you want, you sort of want that on the - It has to be on a model that's good enough to have that. It has to be on the smart model, yeah. But that said, I do think at some point there will be some pieces of alignment that you do want to export back into pre-training because that might be a way to like put them in with more strength, like more robustness kind of or more core to the intelligence.
46:16Like if you think of pre-training as like teach the model to be intelligent and then post-training as like tweak the personality, you can imagine tweaks where you actually want it to be like part of how it learns and like part of its intelligence and maybe you need to integrate more. What would that even look like to incorporate in pre-training? Is that like add extra data basically of the type of domain you wanted to adopt earlier basically? There's a paper called Pre-Training on Human Feedback where you can kind of like add the human feedback characteristics into pre-training to like test that and like, yeah, you can basically give it all the information you give it in post-training, just mixed into pre-training and see what effect that has.
46:51The other loss you have when you do that is you lose the flexibility. Like if you sometimes like train these and then you talk to them and then you like do an extensive process. We have a bunch of people talk to the thing and find some like issue. You know, the model says like you're absolutely right too much. Yeah, and you want to be able to just like go and fix that. Yeah, I mean, I think that iteration loop point you made, I think, feels like the really key point of, yeah, there's a huge difference between taking three months to get information about if your model is good or bad or making it, going in a good direction versus a day or something or a couple days.
47:22And you can do a lot of those. And you can probably, that probably also means it's way less compute. You can do a lot of those in parallel. I imagine you're trying all sorts of post training strategies in parallel there. Yeah. So, yeah, it makes a lot of sense. It's also just the general hard part about pre-training. Like everything in pre-training is hard because you have this like one shot on goal kind of for like multiple months. And totally. OK, so in thinking too now about, I guess, what's going ahead, as you now look to the next several years of what you're building, how do you think about what are the known problems that you're going to face that you're going to have to deal with?
47:51There's going to be more compute, I assume, and you're going to need to hook up even bigger network GPUs and deal with, versus are there areas where you're like, OK, this is a problem that it's a little bit more ambiguous, what the actual, like how it's going to materialize into something you care about, but you kind of know it's an impending thing to think about, or are there things like that that come to mind? I think the things that feel most top of mind to me are probably like paradigm shifts. Like I think the sort of shift towards more RL is like one paradigm shift in the field, and I think there will probably be more.
48:22I think a lot of people sort of argue about like, oh, is like, you know, current paradigm's enough to get us to EGI? And I'm like, I don't know, maybe, probably, but like I'm sure there'll be more. It seems like it would be a really surprising twist if the answer is you just scale and there's nothing that you realize in the process of going up many orders of magnitude. Totally. But I think the things that I actually feel most nervous about are really hard to solve bugs. I think that like... That's interesting. Yeah. And I think this is maybe somewhat surprising to me, but it's just like a single bug can derail you for months.
48:57And when you think about it, the models take months to train. So you can kind of like lose a whole generation off of something that just looks like, ah, it turns out this piece of your code was incorrect and you couldn't detect it. Yeah. And it's really hard in ML, right? ML is always really hard to find bugs in. Yeah, totally. But also some of these scaled up issues are really hard to solve, even when you know they're there. Yeah, what's even a unit test that you would write, or forget a unit test, I mean, anything close to a test for the type of network architecture on which you're doing this.
49:28How do you even do that? I mean, you can send a packet over it and confirm it's the same on the other side. You can train a small model on it. But even train a small model on it, it's not obvious. If you have the very classic, very simple ML bug that early people face in their careers, they have 10 layers in their network, and layer 7 connects to 9 instead of 8 to 9. And so there's some incorrect set of connections you have there. And technically the model still trains and all the weights update. And so it's a valid model, but it's not the correct one. And that's a very esoteric weird bug that would actually be hard to find.
50:03Is that what you're referring to of these random bugs you face? Yeah. It's that, but you can - Times a million. Times a million as the thing gets more complicated. You could cast the wrong precision deep in some kernel, and that causes your model to blow up at large scale. And you find out like a month in. Well, you never find out. Well, you never find out. I mean, you see the thing blow up, there's tens of thousands of lines of code. How would you ever trace it down? So those are the things that probably spook me the most, is just some subtle tricky bug. Yeah, and that's probably the case of you don't know.
50:36I think there's actually also the case of you do know. It crashes. You're trading your model, or it slows down. Your job slows down a ton. And those things can also be very hard to debug. Nelson Elhaj, this one person on the team, he wrote a blog on one cursed bug we had early on. And I remember this one quite well, because I think I encountered it fairly early and was like, this looks hard, can someone else look at it? And a month later, I was like, wow, I'm so glad I handed that one off. I never would have been able to get like, one of the abilities I think is actually really useful, is the ability to deep dive anything to any level of depth.
51:11But that's a pretty rare skill. Like for me, you know, as we talked about what level of the stack I was at before, I was like working at the Torch.matmol, but like I didn't know CUDA. So if Torch.matmol was broken, it wasn't like I could dig into Torch.matmol and figure it out. And it's similarly with like communications, right? Like I could call send, send bytes from A to B, but I didn't know the like underlying networking protocol. So if that underlying networking protocol is broken, like I need to learn a whole field. I have to like understand packets and TCP or like all of these different things to debug that.
51:42And I think one thing that's like surprisingly hard and there's very few people who can do is like kind of own that whole Yeah, it's back from like I understand how the ML is supposed to work and what the learning dynamics are all the way down to like I know the bytes yeah, and I like can understand how the bytes should be moving around machines Totally yeah And actually on that front like when you think about the different backgrounds of people on your team today How do you like approximately? Map them out to different categories of computer scientists like I think there's this external view of what these teams look like which is that they're all PhD researchers who write ML papers.
52:14And I suspect that's not actually true given what you're describing here. Yeah, it's a mix. And I think the thing we most need is engineers. Okay, interesting. Almost always, throughout the entire history of this field. Totally. It's the case that you throw more compute, the thing kind of works. Yeah. The challenge is actually doing that. The researchers are like, cool, nice. Yeah, and getting it correct isn't really an ML problem, right? The actual architectures are pretty simple. Yeah. You can write the math down. But you don't even need to understand the math to implement it. You just need to get a correct implementation.
52:44And then you have an engineering problem of how do I take this, implement it at large scale, parallelize all the things and check that it's correct. But it's, yeah, so it's like kind of engineering skill, but it's this particular type of engineering skill that's about being able to debug anything. Yeah. I think there's another angle of engineering which I think of as really quickly iterate on a website or something. Which I think of as an important skill set, probably important for making startup. You got to be like, fail fast, try a bunch of different things, none of which are like that technically difficult to do.
53:13The skill sets that we're like most kind of in need of or looking for are this like able to solve really hard engineering problems. Are they people who worked at companies that grew a whole bunch and so they have experience like doing the kind of thing you've done over the last several years at Anthropic or do they tend to be academics or like where do they come from? Yeah so at this point like I think we actually just hire a bunch of people who've done this before from other places, and that's the easy answer. Yeah, totally. It's like, ah, yeah, someone who's like... But by this before, do you mean in AI companies, necessarily?
53:44Or also, you know, someone who worked at Meta on their not-AI team, but they ran some other distributed system that reached internet scale five, ten years ago, or something like that? More like we have a specific role in mind. So say I'm trying to make the run train efficiently in JAX. Hire someone who's worked on JAX would be great. Yeah, totally. Or someone who's worked at another company on optimizing a JAX stack to be really efficient. That's kind of like, I think now we're at the point where like the anthropologist well enough known, we can sort of hire these people and also the field is big enough that there's like people with expertise.
54:14One thing that was interesting was like early on we hired a lot of people from just like all sorts of backgrounds. And I think that people who are just smart and work really hard can learn this pretty fast, but you have to like want to. We hired a lot of physicists for instance, like theoretical physicists who just like show up, they do a residency, like learn to program and then they were really smart, they do really great work. I want to switch gears to talk about something a little bit different, which is just sort of future looking things, or how you think about other domains, or advances happening in AI that I'm seeing elsewhere in the field.
54:44And you don't have to tell me if you guys are working on these necessarily, but how you think about them. I guess one big area I was thinking about is around areas other than next token prediction. Are there any of the other things that people are working on that you're curious about? So basically, two differences there. One is not using Transformer as an architecture. So there's companies like Liquid AI that have their own kind of architecture, for example, they're using, or not using autoregressive training as a way of training models. Are there any of those, do you think, interesting in ways that we might come closer to AGI, or do you think this autoregressive framework is the one that kind of makes sense?
55:19I think they're interesting. I think I'm less like, ah, autoregressive is the way to go. On the other hand, I think autoregressive is probably good enough to get to AGI or something or not. Yeah, interesting, yeah. Such that, yeah, I see the main driver as scale and careful science of like sort of the basics more than like come up with something totally novel. Not because there aren't novel things that are better. I actually like I'm pretty confident they are there. It's just that scale is easier and it's more reliable. And I think you were still seeing really big gains to that. Do you spend a lot of time on thinking about things like, you know, I've been reading some of these open source papers where you can kind of dive into some of the details about the model changes and with some of these Chinese labs, for example, where they're making tweaks on the order of the architecture itself with better caching behavior, for example, or more efficient attention functions that make a big difference.
56:07Do you feel like these are examples of things, like you mentioned earlier, where it's basically, in the grand scheme of things, basically if you throw more compute at it, this is all kind of a rounding error? Or do you think it will take some number of these very clever architectural changes to actually get to HEI? In the way that the first person who came up with the transformer made a particular transformative change, Like, will it take some of that? Or do you think it just, you keep doing the thing we're doing to make it bigger? I think it'll be a mix. Yeah. I think, like, my guess is you'll keep tweaking things.
56:34The more compute you put in, the more, like, worthwhile it is to, like, do those experiments to, like, figure it out. You know, I mean, inference is a thing we haven't talked about. But, like, you also want to serve these models to a lot of people. So there's a lot of changes you can make to make inference cheaper. And that depends on, like, the details of your inference stack and the chips you're serving inference on, et cetera. And do you, as someone focused on pre-training, have to think a lot about inference? Or is it kind of like you just do your thing, you make the loss go down and then hand it off and someone else makes that happen?
57:00Oh no, I think a ton about inference because it basically like the problem inference is solving. Like we basically determine the problem inference is solving. We give them a model and they have to like run that fast. And it's very easy to give them a model that is impossible to run fast. Oh, can you give an example of a decision you can make that could cause that? I mean the simplest one is so stupid, but it's like you just make the model giant. Yeah, sure, sure. Absolutely massive. It's trained for like a really small number of tokens. Totally. And then inference now has this giant model. Yeah, and then they're hosed.
57:25Yeah, I mean you can also make things require communications in a lot of places, which would make it harder for inference. Yeah, totally. You can also just make things complicated. And like there's no fundamental reason it's hard, but there's only so many people on the inference team, and they have to implement it in a bunch of places. Yeah, it's interesting. Yeah, no, so I definitely think of like inference is the team that I work the most closely with. Oh, interesting. Because we're kind of like co-designing models to be smart and cheap. Yeah, interesting. Particularly in a world of limited compute, right?
57:56Like the bottleneck, I think, to a large degree. I mean, you can see Anthropic has rate limits constantly, and people complain about it a lot. And the reason is there's only so much compute we can get on short notice. So making your inference more efficient is the way you can serve more users. And actually, let's say you had 100x more compute, or we somehow didn't live in a world where compute was limited. Does that change a ton about what you do? Or is it still kind of the, well, you're just going to grab all of it, whatever compute you have, and keep going down the loss curve. And you kind of, well, it's impossible to be in the world where there is enough compute.
58:30So I think if we got infinite compute, the challenge would be making use of the compute, right? Then you would start to run into these issues like, oh, well, when one chip fail, you know, like, OK, I'm going to throw 2 billion chips on the run. But what happens when a chip fails? So I think we would be limited on people then. It would be like, how fast can we solve the hard engineering problems to scale up? But I do think the change is massive. And I think people don't realize how chip-limited AI research is or something right now, like the models that everyone uses, right? If you're using like Cloud Sonic 4 or Cloud Opus 4, it's like, it's our first shot.
59:00At models of that scale, right? And like, if you think about anything, like you could do it and you could do it again, you could do a better job. Yeah. But if you sort of imagine like 10x the compute, like you could run this every day instead of every few months, like. Yeah, totally. Or 100x, maybe for that, then like, yeah, it's just, it's a really, it would be a really big change to have a lot more compute. And it's coming, right? That's like kind of the fun part of the field. is like every year you're like, oh, I had no computer year ago. Right, exactly. Yeah, exactly. How do you think about methods like discrete diffusion?
59:28Like I saw there's like a Gemini diffusion model. And I think about that in the space I used to be in, where there's a lot of discrete diffusion models being used in protein design, for example, the space where my startup was. Like, do you see that as a domain where there's gonna be interesting advances happening? I'll be honest, we haven't done image generation. Yeah. And I think that's been the main use for diffusion. Yeah, true. So I've kind of had this on my to-do list of things I should understand for a while. And there are people on my team who do understand it and wouldn't have better thoughts, but I actually don't think I understand it well enough to know.
59:56I do have it kind of in this category of not a total... And there's a lot of things that aren't a huge paradigm shift, but they're pretty big changes to how things run. And I expect there are some of those that will work. I don't know if it's diffusion or if it's another one. Obviously, who knows what Anthropical will do in the future, but at least in the near term are the things where you see big areas where a startup can win in the world in which Anthropic is getting, you know, making their models better year over year. My general read is like anything that benefits from the model getting smarter.
1:00:25I think like on the one hand, there's like a lot. You can always be like, oh, yeah, if you're doing a startup, like all the AI labs are big companies, they'll be bigger than you and they could do that thing. But also like we're all working on this general system that covers a lot of different uses. And the plan is to like power all the startups to do all of the individual work. Yeah, I think like anything that just kind of looks like, oh, this almost works with current models, but requires like a bunch of work is a pretty promising direction. I think maybe the thing to watch out for is things where like they work now with a huge amount of work, like to build up a scaffold, but the next generation that you're not gonna need the whole scaffold you built up.
1:01:02Yeah, that's I mean, maybe that's fine. I don't know, maybe you just build up the business with the scaffold and then you don't have to do any work later and you can do business. But I don't know about the business side of it, but like it does feel a little silly to put to invest a ton in that. Yeah, totally. What about on the flip side? Like, are there things in your training stack where you're like, man, if there was a company that solved X problem, I would totally buy their product? Yeah, there's like a ton. I do think that like probably most of these, like the way I would probably structure would be like almost like making something, but then consulting with the company, like offering a service to companies for free.
1:01:32particularly for companies that are scaling really fast, you're almost always limited on how many people you can have. So if you can, even if you could hire people to do it yourself, actually being able to contract someone else to do it, where they're managing it and hire all the people and deal with the organizational side could be useful. I mean, there's a huge amount of stuff. One that jumps to mind, we talked about chips that do math incorrectly. It would be lovely if there was some startup that you could just say, here are my chips, confirm they're all perfect, and if they're not, let me know exactly what went wrong, on what fraction of them and I can tell you the math is wrong, but I don't really know enough details of chips to be like, this chip failed because this particular low-level component was wired wrong or got hit by a gamma ray.
1:02:14I don't know what causes it. You can always go a bunch deeper. I mean, the other thing I'd maybe just push startups on is thinking a little bit about, this is maybe less technical, but just what happens once we get AGI and how to make sure that goes well for the world or something. My expectation is if you actually automate almost everything a person can do, the amount of economic growth there is just truly enormous. And I would think a little more about how do you make this help the world versus not? I think there's going to be plenty of economic success or something as a result of it anyway.
1:02:44Yeah, absolutely. The last question I want to ask you is around, if you're running back to where we started like 10 years ago, you're a student, you're pivoting into AI from kind of economics work you were thinking about. And all sorts of things you probably did in those early days had some kind of compounding return for you as you developed into the role you have now. What advice would you give to students as they think about entering the workforce, especially today, learning skills that are going to be useful, and maybe getting themselves jobs like the one you have right now 10 years later? It's hard because I think the timing is very different.
1:03:16I just think we've made a lot of progress. What I would do 10 years ago is different from what I would do today. Sure, totally. But I think certainly if I went back 10 years ago, I would be like focus on AI. It's like the most important thing and particularly focus on engineering, which I think felt very, wouldn't have seemed obvious to me at the time that like the important thing was these engineering skills and not the like math and theoretical understanding of like, you know, SVMs and all the kind of standard ML literature. I think today I would probably focus a bunch on the engineering and on the figuring out what to do with AGI as sort of the two main things that feel top of mind for me.
1:03:54Let's call it up there. Thanks so much, Nick. Appreciate it.
From the publisher
Ever wonder what it actually takes to train a frontier AI model?YC General Partner Ankit Gupta sits down with Nick Joseph, Anthropic's Head of Pre-training, to explore the engineering challenges behind training Claude—from managing thousands of GPUs and debugging cursed bugs to balancing compute between pre-training and RL. We cover scaling laws, data strategies, team composition, and why the hardest problems in AI are often infrastructure problems, not ML problems.




