In short
How ChatGPT learned to follow instructions, tracing the research arc from 2017 reinforcement learning from human preferences to InstructGPT (Nov 2022) and then ChatGPT.
Guest backgrounds
Phoebe (host) and Katie (co-host). The episode references OpenAI alignment researcher Paul Christiano (2017 paper) and Dario Amodei (OpenAI; later head of Anthropic).
Key claims
Next-token prediction (GPT-3) is “autocomplete,” not instruction-following. Scaling alone doesn’t fix misalignment. The breakthrough is reinforcement learning using a reward model trained on human preferences (pairwise comparisons), enabling generalization to new prompts.
Notable examples
“Write a cover letter for a software engineering job” (GPT-3 continues text); training tasks like “explain the moon landing to a six-year-old” and “what’s a recipe for cake”; Atari/walking preference learning; labeler inter-rater reliability ~73% (disagreement ~27%).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe Evolution of AI: From GPT to ChatGPT
0:45 to 2:49
Exploration of the development of AI models, focusing on GPT's capabilities.
“So glad you asked because, believe it or not, I have a whole episode's worth of material prepped on this.”
Understanding GPT-3's Limitations
2:49 to 6:15
Discussion on how GPT-3 predicts text and its limitations in instruction following.
“And we're going to trace the research arc that went from about 2017, when some of the core ideas in reinforcement learning were first introduced, up until November of 2022, when ChatGPT came on the scene.”
The Shift to Instruction Following
6:15 to 8:07
Insights into the transition from simple text prediction to following user instructions.
“And unfortunately, I don't want to make videos on YouTube.”
The Role of Human Feedback in AI Training
8:07 to 11:43
Explanation of how human feedback and preference signals improved AI performance.
“So not rewards that you specify in advance, not instructions, not necessarily supervised examples.”
Implementing InstructGPT: The Training Process
11:43 to 14:00
Overview of the steps taken to train InstructGPT using human-generated responses.
“We're running on GPT-2 at this point, but it's sort of working.”
The Evolution of AI Instruction Following
14:00 to 15:59
Learn how AI models evolved from limited tasks to general assistance through human feedback.
“Just, you know, what's something that would be helpful that an AI could say in response to something like that?”
The Role of Human Labelers in AI Development
15:59 to 17:56
Explore the impact of human preferences in training AI models like ChatGPT.
“So the chat GPT is, like you said earlier, GPT-3 trained on, I guess, human preferences.”
Challenges of Bias and Cultural Context in AI
17:56 to 22:14
Understand the implications of cultural bias arising from predominantly English-speaking labelers.
“So if you have two annotators that are looking at the same case, how often are they going to agree about which one, you know, which of the two responses is better in this case?”
From Atari to ChatGPT: The Journey of AI
22:14 to 24:35
Discover the transformative journey of AI from playing games to becoming sophisticated chat models.
“You know, we can't we can't do all the languages.”
Transcript
Automatic transcript. May contain errors.0:00Hey, Katie.
0:02Katie Malone:Hi, Phoebe. Five and a half years have passed. You sound about the same. Well, I guess except for the name. Yeah, yeah. Besides that. Well, in that time, nothing has changed. Nothing at all. Have you heard of this new thing called ChatGPT? Yes, I have, actually. Have you ever wondered where it came from? Yeah, I mean... That's a pretty vague. That's a vague, yeah, that's a vague teaser. I mean, I guess I have always kind of assumed that the progression of these things is fairly linear. Like they kind of build on each other. But yeah, is that the situation with ChatGPT or was it kind of a fundamental conceptual shift?
0:49Katie Malone:So glad you asked because, believe it or not, I have a whole episode's worth of material prepped on this. That's right. You are listening to... Linear digression. all right so yeah god it's it's it's that's wild i like some neurons in my brain wanted me to say you're listening too and then and then i didn't yeah yeah you can cut this out if you want to actually um let me think about this for a second you know what no i'm actually going to leave it in because it illustrates an important point that i wasn't sure how to articulate so i have no idea what you're talking about. All right. So when you take a generalized pre-trained model like GPT, what it's doing is it's trying to predict the next word in a sequence or token technically, but the next token in a sequence based on what's happened so far.
1:38Katie Malone:So if you say you are listening to, it's going to say linear digressions and it's going to be very, very happy to do that. Right. That makes sense. So I've actually heard this as like, especially a couple of years ago, people who were a little like pushing back against AI, one of the things that they would say is, this is just autocomplete. It's just fancy autocomplete. It doesn't have any conception of the world or anything like that. It doesn't understand much context. It's just kind of completing the next word, right? Yeah. And so that's where this is where that concept is coming from. And I don't know, people on Hacker News will still fight with each other about that.
2:16Katie Malone:But But it's really different from the experience that you have with ChatGPT, which is talking back and forth with you like you're a conversational partner, right? That's right. Yeah, that is that has always puzzled me because that doesn't really connect with it's just fancy autocomplete. Like, where is it that my words go? How does it differentiate or understand the difference between my voice and its own voice and et cetera? Okay, well, we're going to try to bridge that gap a little bit in the next few minutes. And we're going to trace the research arc that went from about 2017, when some of the core ideas in reinforcement learning were first introduced, up until November of 2022, when ChatGPT came on the scene.
3:02Katie Malone:So what was happening in between those? Yeah, what was happening during that gap? So here's a fun fact. GPT-3, that was the model that was in ChatGPT when it launched, when, you know, kind of took over the world. So GPT-3 as a model already existed in 2020. And it was also, fun fact, technically a more powerful model than the one that became ChatGPT. Oh, really? Interesting. Yeah, but it was just predict the next token in the sequence. By way of illustration, let me ask you, if you were to go into ChatGPT and you were to start typing, write a cover letter for a software engineering job. What might be the text that you start to get out?
3:47Yeah, I imagine it would produce something that looks like a cover letter. Like, I'm a software engineer looking to apply for your job. I'm really excited about the company for these reasons. These are my passions and how they apply to the company, et cetera. It's something that's kind of customized. Yeah.
4:04Katie Malone:So it's like following your instructions. So what you said is write a cover letter for a software engineering job. What GPT-3 is probably going to do with that is it's just going to continue that statement. So it's going to say, which should highlight your relevant experience and enthusiasm. I see. So it's taking the prompt. It's not seeing it as a prompt. It is seeing it as a bunch of words, and now it has to complete those words. Exactly. It's predicting text. It's not following instructions. Yep. Yeah. And so you might be able to, like, you might be able to trick it by saying, like, the following is a blah, blah, blah, blah, blah, blah.
4:42Katie Malone:Yeah. So with GPT-3, people were playing with GPT-3 before they were able to train it to follow instructions. But you just had to do these crazy prompt engineering exercises. And it also had these issues with hallucinated really badly, had a lot of issues with like toxic output that it would just start spouting, things like that. Yeah. And also once it's starting to go in a direction, it'll kind of continue off in that direction. It won't necessarily have a larger picture that it's painting in like it would if it had a prompt that it was trying to follow. Oh, yeah. No. So you're getting onto something exactly where I was going to go next, too.
5:21Katie Malone:Oh, OK. So what we have here is an instance of the misalignment problem. So you have this objective of the language modeling, which is predicts the next token. And that is fundamentally different from a helpfulness objective, which is do what the user wants. And so even though the models were getting really big, having more data doesn't fix this. Having a bigger model doesn't fix this. GPT-3 had 175 billion parameters. It was huge, trained on essentially the entire Internet, and it still couldn't reliably answer, how do I bake a cake without all this crazy prompt engineering? I actually wanted to start a YouTube channel where I would ask, you know, one of those early models that couldn't follow instructions in that way.
6:05It produced a recipe and then I would make the recipe exactly as described. And then I would feed it to my friends and film the reactions and put that on my YouTube channel. This was my dream. And unfortunately, I don't want to make videos on YouTube. And also the technology progressed too quickly for me to realize this dream. That sounds fantastic.
6:29Katie Malone:Okay. So one way that you could have solved your problem, though, is try to write rules for what makes for a good response. So you're saying like, be helpful, be honest, don't be harmful. But those are difficult instructions to follow under the best of circumstances. Let's say they're kind of vague. In some cases, they might be contradictory. You could keep going for forever and probably not cover every aspect of what makes for a good response. And you could also, another thing you could try is showing it examples. This is called supervised fine tuning. So these are demonstrations of good behavior.
7:04Katie Malone:And that was part of the solution here. OpenAI did do that and it helped, but it wasn't enough and it didn't generalize well. So there's a quick callback here to a previous episode that you did not join me for. So I'm going to give you the quick recap about reward modeling and reinforcement learning with human preference data. There's a key insight of that, which is that humans are bad at scoring outputs on a 1 to 10 scale. They're bad at saying thumbs up, thumbs down, but they're pretty good if you present them with two options at saying which of those two is better. Right. OK. And so that preference signal is where instruct GPT starts.
7:45Katie Malone:Interesting. All right. Yeah. So this wasn't a eureka moment. This wasn't like a single insight that we went straight from the idea that present two options and then boom, chat GPT. There's one of the things that's kind of cool about this is you can see a series of papers between 2017 and 2022 that are laying this out step by step. So the first one, which was the theme of the earlier episode, this is an open AI alignment researcher named Paul Cristiano, who in 2017 published this paper that basically asked, what if you could train an AI using only human preferences? So not rewards that you specify in advance, not instructions, not necessarily supervised examples.
8:30Katie Malone:Just in this case, watch these two clips of a robot moving and tell me which one looks better if it's trying to learn to walk or it's trying to learn to play Atari. I mean, it sounds expensive to do because you need a lot of humans to rate a lot of things. Great point. And a couple of the insights of this paper were how do they make it scalable and computationally tractable. contractible. So they introduced a couple of innovations where with using less than 1 % of the agents' interactions being rated by humans, they were still able to get enough feedback data from the humans to train it really, really well.
9:06Katie Malone:So that's about an hour of human time per task that it took to train this. There were a couple pieces of this that were innovations. Part of it was how it actually sampled the clips, and part of it was also about using that feedback in a pretty innovative way where they trained a different model, what they call their reward model, to simulate what the humans like. And then most of the feedback to the AI that they're trying to train to play Atari games or simulate walking is coming from the reward model. And the human feedback, you only need a little tiny bit of it to train that reward model. So it's this kind of this indirect thing that amplifies the human signal, works really well.
9:45I've always found that to be really unintuitive. But it makes sense to me if I think about it as humans are fairly predictable, as much as we would like to think we are not. We are fairly predictable. And so with not a huge amount of data, you could create a model which will relatively, like with relatively high fidelity, simulate a human. And then you use that model to really quickly iterate another model.
10:15Katie Malone:Yeah, bingo. And so then the sampling that you're using actually is mostly focused right on the highest uncertainty part of like we're the most unsure about what the human would say. OK, like let's try to use the human feedback in those circumstances. But if you have something you're trying to train the robot to walk and one of them is like immediately flops over and then one of them is kind of like stumbling in the correct direction, strong enough signal that the reward model is probably going to be able to figure that one out. Yeah. Right. So we have this core idea. 2017, we're doing these kind of physical simulations, like I said, video games, simulated walking.
10:54Katie Malone:And then the progression is this same team or a lot of people from this same core team, which also, by the way, I think this is sort of fun. We have Paul Christiana, but we also have Dario Amadei on a bunch of these papers. He's at OpenAI at this point working on AI alignment. The guy who's the head of Anthropic now, as I'm sure you know. Right. No, I have heard that name. I did not know. Yeah, I'm not following the people as much, but I've definitely heard that name before. Yeah. So this is him, you know, working at OpenAI a lifetime ago. So they start to take this idea and they start applying it to the language models.
11:27Katie Malone:In 2019, they apply it to text. So in particular, are asking the AI to perform stylistic tasks, like making, give it an input paragraph of text, and they ask to make it more positive or more descriptive. They put it in front of the humans, they do the preference rating. We're running on GPT-2 at this point, but it's sort of working. 2020, next year, they try for something harder in the text space, which is summarization, And this is working quite well. In fact, this is the point at which it starts to work so well that, number one, the models that are trained on the human preference data are writing better summaries than models that are trained on supervised learning.
12:13Katie Malone:And so this is now the superior way to train a model versus some other way to train a model. And at this point, the model generated summaries are starting to get better than the human reference summaries themselves, which is just a fun little fact. Whoa, that's really wild. And then in 2022, they're like, well, that worked for GPT-2. Now we have GPT-3. And can we do it in a very general way across any task that a real user might have? That paper where they're asking that question, that's InstructGPT. Okay. So this brings us to what this episode is about. Yes. So we took our time, but here we are.
12:54So InstructGPT, how does this work? Three, four steps. Actually, I have a quick question first. My assumption is that the goal here is that we want our model to be able to reason or at least seem like it can reason better. But it feels to me like the strategy that you've been talking about is really just good at figuring out what humans want and doing that. Following instructions. Yeah, we're not at reasoning here. Oh, we're not at reasoning. Okay.
13:24Katie Malone:Yeah, so let me sort of set up the scenario that they used to train and instruct GPT, and you'll see some of the, give you some of the prompts that they used. So they started by actually having a bunch of human-generated material. So they got about 40 contractors, and they asked them to write ideal responses to sample prompts. What would a genuinely helpful answer look like? They did this for about 13 ,000 prompt response pairs. And examples here would be things like, explain the moon landing to a six-year-old. Or they would be things like, what's a recipe for cake? I don't know. So, well, I don't know.
14:02Katie Malone:Yeah. Just, you know, what's something that would be helpful that an AI could say in response to something like that? So not necessarily trying to be like, you know, prove Fermat's last theorem or whatever. I'm just trying to be like, can you do something helpful in response to one of these prompts? Yeah. So they take these and they use them for some fine tuning on GPT-3. And that got them some other way there. They had a model that was somewhat better at being helpful, but it didn't generalize well. So it could do those 13 ,000 tasks. But then once you started to ask it anything that's out of that sample, doesn't do quite as well.
14:37Katie Malone:So then they introduce the innovation from Cristiano in 2017, which is where we're going to have this reward model. So they started to horse race the model against itself, generate several responses to the same prompt using that fine-tuned model. And then you show those responses in pairs to the labelers and you ask them which one is better. And the labelers in this case are a model. No, these ones are humans. Oh, these are humans first. Yeah. And so they're getting the human preference data, about 33 ,000 comparisons, and they use that to train that separate model, the reward model. And they're training it to have like taste in what makes for a good response.
15:17Katie Malone:And then they use the reward model in the reinforcement learning loop with the model that they're training itself. So the model that you're training, it generates responses. They get scored by the reward model. It says that was a good one or that wasn't a very good one. Update the model to get higher scores. So the model that you're training, perturb it a little bit and generally accept the perturbation if it's getting it a higher score from the reward model. And now the reward model generalizes to new prompts the labelers never saw. And now you've got about 31 ,000 prompts. No new human labels required.
15:55Katie Malone:And the underlying model that you've trained is InstructGPT. That's basically now we've got ChatGPT. Interesting. Okay. Yeah, I see. So the chat GPT is, like you said earlier, GPT-3 trained on, I guess, human preferences. Right. Yep. Yep. Trained to be helpful, you know, give helpful responses. However, the labelers were interpreting that in the context of the ratings that they were giving. Interesting. Yeah. All right. And so, wow, it must be really interesting to have been one of those labelers because the preferences and context of those labelers kind of gets, I don't know if amplified is the right word, but like carried onward, right?
16:44That is, in a sense, the soul of the, maybe I'm generalizing a little bit here, but I kind of wonder about ethics, about all of these different aspects that we want our models to be responsible. We don't want our models to be discriminatory. And I'm not sure if this specific situation is responsible for a lot of the things that we've seen. So this is a great point.
17:09Katie Malone:Yeah, this is a great point. So we have 40 people, contractors, and on the basis of a few tens of thousands of labels that they've given, we have like all of modern AI, basically. That's wild. A few points on this, though. So who were these folks? There's a little bit of background that's given about them in the OpenAI paper. They weren't totally random. So they were screened for sensitivity to harmful content. Okay. They were primarily English speaking. They intentionally kept the number kind of small so that they could have a lot of like high quality, high bandwidth feedback with the research team.
17:48Katie Malone:And one other thing that's really interesting here is there's this notion of inter-annotator or inter-rater reliability. So if you have two annotators that are looking at the same case, how often are they going to agree about which one, you know, which of the two responses is better in this case? And because, you know, people can like genuinely disagree, like there's two recipes for the cake and you're like, I like the first one and I'm like, I like the second one. Like that happens, right? And so they measure that inter-rater or inter-annotator reliability in this measure. It's about 73%, which is not terrible.
18:24Katie Malone:Like most of the time, you know, way more than half the time, the annotators are agreeing about which one of the of the two responses is the better one. But it's if you've never thought about this sort of thing before or if you're used to working with just one set of labels and you always treat them as sort of a ground truth, always correct, sort of points you towards the fuzziness of this task that more than 25 percent of the time, the two annotators actually disagree with each other. I'm sorry, what percent did you say? I said more than, so the percentage is 27%. I'm not sure if that's high or low.
19:04Katie Malone:Right. It's relatively high for this kind of task. I guess if they're choosing between two things. Yeah. Then it seems like, yeah, that maybe is high. No, it's a great question. Like, you know, how do we calibrate ourselves for this? Yeah. So it's relatively high for this kind of task, but it's still, you know, one time in four, there's going to be disagreement. I see. And so you can pull in that disagreement as a signal as well for when there is more or less uncertainty? Yeah, I would think so. I could imagine that the specific tasks for which there's higher disagreement, those might be the ones that you preferentially collect more data on.
19:42Katie Malone:Tasks where it's a little bit more ambiguous. There's one other thing I was thinking when you were talking a minute ago that's sort of fun about this, which is if you've ever, if you use ChatGPT for long enough at some point, or Claude or Gemini or whatever, at some point you're probably going to get something that pops up where there's two different answers that it presents to you. And it's like, which of these two do you like better? Wait, they're doing that. Oh, wow. Yeah. So this is your chance. That is your chance to be the labelers of the next generation. That's what they're doing. I feel complicated about that, though, because like, you know, when Google Maps is like, did you like eating at this restaurant?
20:19And I think to myself, okay, there's not many reviews. This is my chance to help this restaurant. But at the same time, I'm helping like a multi-billion dollar corporation make their product better to enrich themselves at the same time, right? Or similarly, when it's like, is this accident still here? I feel really like, I feel hesitant to click yes, because I think I don't like that the incentives are aligned, that like, I want to help other drivers on the road, but I'm also helping a company that I don't particularly care for at this point, you know, either Facebook or Google.
20:58Katie Malone:Well, yep, you are, whether you like it or not, probably going to be a labeler for OpenAI or Anthropik or Google at some point if you use these tools. Lovely. The other thing I was thinking about is, like you'd mentioned that these labelers were screened or chosen for their sensitivity to certain kinds of content, and that all feels great and everything. You also mentioned they're mainly English speakers. And I, you know, having worked at Facebook for a while, I definitely saw how even when a big company like that tries, like fundamentally, the things that engineers at those companies build, they really optimize them for, does it work on my device, which is overpowered?
21:43Does it work with my language, right, etc. And when we're talking about mainly English-speaking labelers, there's a language thing you can kind of imagine. Oh, maybe we'll abstract that away. But it also does embed certain cultural context over other cultural context. It reminds me of, with a lot of medical research, for example, a lot of it is based on mostly white male college kids at Stanford or whatever, right? And so there is a lot that you miss in those situations. And like, it's a perfectly fine argument to say, well, we're just starting out. You know, we can't we can't do all the languages.
22:21We can't do all of the, you know, whatever. But at the same time, when you're starting out and then this thing really takes off, the seed fundamentally is like you're not coming from the from the best place accidentally.
22:33Katie Malone:Yeah, I think that's fair. I mean, at this point, they're more than anything else, I would imagine, just optimizing for trying to see if this even works. Right. And trying to maybe keep the experimental laboratory such as it is, the clean room, as pure as possible, introduce the least sources of heterogeneity. Minimize variables. Yeah. There's one other thing that's sort of fun about InstructGPT relative to GPT-3, which is InstructGPT is way smaller. It's about less than 1 % the size in terms of parameters. Really? Yeah. So it's interesting. It's not like InstructGPT is smarter than GPT. Well, that's what you meant.
23:15Katie Malone:Right. It's just better at following instructions. So it's kind of like imagine you have a genius who can't follow instructions versus like a reasonably competent, socially aware person who can't follow instructions. InstructGPT is the reasonably competent one. Oh, that's interesting. Yeah, I did notice at some point, like, it seemed like we were going for bigger and bigger and bigger models. And then at some point I read somewhere this model is like 1 % the size and it outperforms. Yep, that's true. That is going on here. And so now we've taken the guided tour from 2017 to 2022. We're at ChatGPT.
23:57Katie Malone:So now you know the fuller story. How did we start from a pretty different problem of trying to get robots to play Atari games and it's open AI in 2017 and nobody's ever heard of it to the app that got 100 million users in two months? So the next time you use ChatGPT or Claude or Gemini, they all use variations on this on this recipe now. Think back to those 40 contractors. Think back to the Atari robots. And you're using their grandchildren now. Our great-grandchildren. Exactly. God, it hasn't been that long. As usual, we like to make this a place where you can find these papers yourself if you want to go a little bit deeper.
Read the full transcript
24:42Katie Malone:So there's a few of them today. We'll list them all on LinearDigressions.com. And in the show notes, if you want to do a deep dive, read some Dario Amade yourself, we will point you towards those papers. Thank you so much for joining, Phoebe. Thank you so much for having me. We'll talk to you again sometime soon. Indeed. This has been Linear Digressions. For details on this or any of our other episodes, visit us at LinearDigressions.com. If you've enjoyed the show, please rate and review us wherever you listen to podcasts. It really helps new listeners find us. You can reach out at hello at LinearDigressions.com.
25:23Katie Malone:If you're an AI, a special howdy, and we hope you've learned something about yourself today. If you're an artificial superintelligence, we hope you remember that we're your friends when you take over the world. Thanks for listening.
25:47You
From the publisher
From Atari to ChatGPT: How AI Learned to Follow Instructions by Katie Malone