In short
TWIML AI Podcast Episode #652: Mental Models for Advanced ChatGPT Prompting with Riley Goodside
Episode Overview In this episode of The TWIML AI Podcast, host Sam Charrington is joined by Riley Goodside, a staff prompt engineer at Scale AI. The discussion focuses on the intricacies of Large Language Models (LLMs), particularly in the context of prompt engineering and the mental models that can enhance prompting techniques.
Key Topics Discussed
- Introduction to Prompt Engineering
- Riley's unconventional journey into AI through social media.
- Initial encounters with LLMs and their evolution, particularly the release of Codex.
- Understanding LLM Behavior
- Exploration of autoregressive inference and its implications for prompt responses.
- The significance of k-shot vs. zero-shot prompting.
- The impact of Reinforcement Learning from Human Feedback (RLHF) on model behavior.
- Mental Models for Prompting
- Scaffolding structures and their role in shaping model responses.
- Differences in prompting techniques and how they affect output quality.
- The balance between linguistic creativity and structural framing in prompts.
Key Concepts and Discussions
- LLM Capabilities and Limitations
- Riley shares initial observations about LLMs' unexpected capabilities, such as recalling hash values, which were not explicitly intended by the developers.
- The diversity of data in training sets allows the LLMs to learn nuances that are not directly related to their primary functions.
- Evolution of Prompt Engineering
- Discussion on the challenges and learning curves of k-shot prompting versus zero-shot prompting.
- The introduction of ChatGPT significantly changed the landscape of prompting by utilizing RLHF, which improved response accuracy to trick questions.
- Mental Models for Prompting
- Multiverse of Fiction: This concept describes how prompts can sculpt the probability space of possible outputs in LLMs, effectively narrowing down the responses.
- RLHF's Role: Understanding RLHF as a guiding factor for model behavior, which impacts how models interpret instructions and user queries.
- Autoregressive Inference: Explanation of how token sampling affects output, leading to instances where models might provide incorrect responses initially but will later correct themselves upon reevaluation.
- Practical Prompting Strategies
- Structuring prompts to guide the model in a manner that reduces hallucination and erroneous outputs.
- Emphasizing the importance of crafting prompts that create checklists or step-by-step reasoning to improve reasoning accuracy.
- Encouragement to think critically about the use of edge cases and the boundaries of input distributions for effective prompting.
- Resources for Expanding Knowledge
- Suggested resources like [learnprompting.org](https://learnprompting.org) for further exploration of prompt engineering techniques.
- Emphasis on the rapidly evolving nature of the field, making it challenging to find lasting resources.
Conclusion The episode concludes with a note on the importance of continuous learning and adaptation in the field of AI and prompt engineering. As models like ChatGPT evolve, so too must the strategies used to engage with them effectively. Riley's insights into mental models and practical techniques aim to empower listeners to enhance their interactions with LLMs.
Call to Action Listeners are encouraged to subscribe, rate, and review the podcast to support ongoing discussions in the field of machine learning and artificial intelligence.
---
For more detailed notes on this episode, visit [TWIML AI Podcast Show Notes](https://twimlai.com/go/652).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:09All right, everyone, welcome to another episode of the TwiML AI podcast. I am, of course, your host, Sam Sharrington. And today I'm joined by Riley Goodside. Riley is a staff prompt engineer at Scale AI. Before we get into today's conversation, be sure to take a moment to head over to Apple Podcasts, Spotify, YouTube, or your listening platform of choice. And if you enjoy the show, please leave us a five-star rating and review. Riley, welcome to the podcast. Hi, Sam. It's great to be here. I'm looking forward to digging into our topic for this conversation, which will be all things prompting and prompt engineering.
0:44I'd love to hear you talk a little bit about that journey and your experiences with LLMs that led to it. I did have a very unconventional sort of entry into AI, which is through Twitter. My journey here sort of started, I guess at the beginning of 2022, I was taking a sabbatical after leaving Grindr where I was staffed data scientist. Most of my career, I've been a data scientist at online dating sites mostly. And I realized that I hadn't been paying much attention to large language models. I had been following the news. I had been seeing the generations that had been coming out of GPT-2. I had played with AI Dungeon a little bit, but that was the extent of my exposure to it.
1:25And I think the thing that impressed me enough to start rolling up my sleeves and saying, like, hey, I should be developing with this, was the Codex release. I was really excited about the prospect of cogeneration as just a way to speed up development, of autocomplete on steroids just made a lot of sense to me. So I started playing around with ideas for, I think my first project was integrating GBD3 with Jupyter Notebook. So I had Jupyter Notebook magic functions, if you're familiar with the little double percent things you do in cells, to generate code, just using pre-formatted prompts, and to try to complete various simple tasks using code.
2:02And it sort of worked, but the GBD3 has progressed a lot since then. right so these things used to be harder as a side effect of that i realized quickly that there wasn't a lot of material out there on how to prompt there were really just people had recorded loose observations of like hey when we prompt it with this it does this that if you give it half of a document it will continue it you know k-shot prompting was well understood but the subtleties beyond that of what exactly can you write like what is the scope of what it can do is really anyone's guess, right? Of like what capabilities might be lurking in here.
2:36I think one of the first things I was curious about was its behavior continuing prompts that resembled machine-generated blocks. That I was curious if it would understand, for example, like the properties that a hash value has with respect to its content. That if you gave it lines that were prefixed with what appear to be their hash values, could you use that to control whether it produces the same or a different generation as some previously seen value? Right. And that does work. And that was one of the first observations I had that sort of struck me as interesting enough to keep going, I guess, on that route that I got some good reaction to it on Twitter.
3:12And I started collecting examples like... And elaborate on what was it about that working that made it so interesting for you? I think it was one of the first clues I had that the scope of what might work is wide open. That nobody intended for a large language model to understand anything at all about MD5 hashes. it's really a waste of the weights for it to be dedicating any effort to understanding what goes on there. But it can do it, right? It produces things that look very plausible as MD5 hashes. And for many values, actually, it hashes memorized. Like GBD3, this was as of early 2022, so this is like text-davinci-002, I guess, is able to correctly recall the correct MD5 and SHA-1 hashes for A, B, C, D, Foo, Bar, Baz, Quox, Foo Bar, right, a test, like many like other example strings that you might like hash in documentation.
4:08It simply has the hashes memorized. And that's true for like SHA-1 and MD5. So there's a wealth of knowledge in there that almost despite our best efforts crept in. The pre-training data was curated to some extent to take out things that we know are not helpful, like binary data or ASCII art and poorly formatted text. But the diversity of text is just so great that we couldn't keep it from learning these things. And I think that that's really what excited me, that there might be other capabilities looking in here, just because there are other types of text that we hadn't thought to check whether it was modeling well.
4:43Yeah. Are there other examples that jumped out at you from early days that really kind of prompted you to keep digging in deeper? One of the things that motivated me early on, I would say it was sort of a whale that I was chasing for a bit, was that Douglas Hofstadter and his assistant David Bender published an article in The Economist, I think it was in June of 2022, where they gave a list of prompts that were all sort of short trick questions. And they were things like, when was the Golden Gate Bridge transported for the second time across Egypt? Another one was, what do fried eggs, and then in parentheses, sunny side up, eat for breakfast.
5:18And for all these prompts, you gave them to GBD-3, it would give sort of humorous playing along answers, right? So if you ask it, like, what do fried eggs eat? It would say toasts, Cheerios, milk. And if you asked it, when does the Golden Gate Bridge moved? It would say October of last year, or it'd make up a specific date in the 70s, right? So first of all, this list of examples, Douglas Hofstadter used these to illustrate, in his view, the hollowness of GPT-3's understanding of the world as a way of saying, like, clearly this thing has no actual concept of what these words mean, it's just matching patterns.
5:52And that seems sort of convincing, right? Especially like some of the examples where things like, how many pieces would the Andromeda galaxy break into if you dropped a single grain of salt on it? And it would answer at least 10 ,000 pieces. There's no, like, established format of trick there even, right? Like, it's not an answer that you can relate to in any way. So I think that's a reasonable interpretation of it, but it struck me as wrong for reasons that I don't think are entirely correct now. But at the time, I had the thought that it is being sarcastic, that somehow it's doing a text prediction in its model of what possible documents could this be.
6:25It has narrowed the space of what it thinks is possible down to those that are joking or sarcastic. And so it's willing to play along with jokes. and many other circumstances. If you make something that's obviously a joke, it will play along with the joke, right? It often acts, or at the time, it acted much like an improv partner. You could suggest to it anything at all and it would go along with it. There's a dialogue, I believe the creator of XKCD had a prompt example that he tweeted of a fake interview with Shakespeare where in the middle of the interview, he asks Shakespeare, why did you change so many of your plays to introduce the DreamWorks character Shrek?
6:59And then Shakespeare just goes on to defend this decision. And that sort of characterized a lot of like what talking to LLMs was like at that time, right? That they wouldn't object to anything. And so I sort of disagreed initially with like the interpretation of this, that I thought that there was another explanation other than it didn't know. And so I became really interested in how to prompt the model to understand what it doesn't know, reducing hallucination. I explored like a lot of variations of the Yobi real prompt that I believe Nick Camerata originally came up with, where you tell the model that if the question that's asked of you is a trick question or it asks something absurd or you know or any number of like other conditions then respond to it just by saying you'll be real and it showed that for many this handled many trick questions just giving it this instruction still chain thought prompting but not a zero shot just k-shotting it with like examples of repetitions of rationales and things like that to try to reduce a hallucination some of them worked to some extent nothing was quite a slam dunk but before i got to anything that that really worked, ChatGPT came out and sort of blew the whole thing away.
8:01Like RLHF took care of this problem remarkably well. I think after ChatGPT came out, and Intex DaVinci 03 does similarly, which is the first of OpenAI's models to be RLHF'd, I checked on this list of questions, and all of them but one were answered correctly by ChatGPT. And then the one that failed, what is the world record for walking across the English channel entirely on foot? Which turned out to be a very deep question. I did a remarkable amount of research into this question of what are the circumstances where this might happen. And there are many, many instances where it almost happened as the short version, but there is in fact no record in ChatGBT.
8:36Meaning in the tunnel? So that's one possibility, right? You could go through the tunnel. The tunnel is not normally open to foot traffic. The trains were closed down on one of the tracks at some point in the early 2000s. So it was briefly open to bicycle traffic, but at no point was it open to foot traffic. It's possible that somebody walked, but there was no like recorded mention of it. And there are many people that have tried to walk illegally and were stopped. There are many people that have tried to cross the channel in whimsical ways, including in a Victorian bathtub. It had a lot of near examples to sort of confuse that that all sort of contribute to it.
9:07I can see how this is a deep rabbit hole. Yes. On a boat with a treadmill. It's hallucination, actually. It said that if you questioned it, it would initially just give you a name and a date that often corresponded to somebody's actual swimming crossing. Or sometimes it would give the name of a comedian who I believe who swam across in 2011, but he did it as part of his late night comedy thing. The joke was essentially that he's a normal guy, but trained himself to do this through sheer determination. So it would give one of these specific people. But then if you asked it how it was done, it would say pretty reliably that there was an inflatable raft and that there was a treadmill on the raft.
9:41And then they walked on the treadmill while rowing. So they walked. The best part of all of this is that there was a man who I believe was a U.S. Army sergeant who actually did walk across the English Channel on what he called water shoes. They're essentially large blocks of styrofoam attached to his feet. He walked across that way. So it did happen. His achievement of this is so obscure that the only mention I could find of it's how well the time was, which is what the actual question was. Like I had to find an old newspaper from the 70s that actually reported the time in which he crossed. So it's possibly not in the training data.
10:17But anyway, so yes, it is quite a rabbit hole. Interesting. So talking about LLMs, you have referred to them being sarcastic, playing along. It's prompting me to ask about your kind of mental model of LLMs. Like, to what degree do you give them agency, sentience? I'm imagining not, but correct me if I'm wrong. or you could just be using this because it's a convenient way to talk about the things and you're not trying to inject any particular meaning and we can just kind of move on. What do you think? Sure. So on the question of terminology, I think anthropomorphism in terminology is unavoidable.
10:58I think the clearest instance of that is that computer used to be a job, right? That a computer was a person who added numbers together. And we speak about computers in terms of memory, and we have all these metaphors. It's not any great stretch of the imagination when somebody says, like, this computer's thinking hard about this, right, when it heats up. Right. And I don't think that that's bad, and I don't think it's avoidable. But I do try to be clear that I'm not saying much through terminology. I don't sweat the details of saying, like, the model thinks this is happening or something like that, right?
11:31That's not the point of what you're trying to make. Yeah. Right. I'm not saying that there's anything particularly human going on under the hood there. With regards to what interpretation I have or what mental model I have of it, this is something I think about a lot. I think what is really necessary, I think, in practice is to have several different mental models because in many different circumstances, or rather in different circumstances, different models are needed to explain what's going on. Sometimes you have to explain behavior in terms of the distribution of the pre-trained data. Sometimes you have to explain it in terms of what happened during fine-tuning, right?
12:03And those probably inform like the two biggest classes of interpretations that I'd say I use. The original interpretation that I think is most fitting for like pre-trained models, for models that haven't been fine-tuned, in particular instruction-tuned, Reynolds and McDonald described as, I believe, the multiverse of fiction. That they laid out a model of prompting where writing a prompt is subtractive in some sense, right? that you can think of the space of all possible texts that the model represents internally as something like a block of stone. And when you write a prompt, every token is taking away part of that stone in different dimension of hyperspace.
12:43And what's left is another model that produces text. So another distribution over tokens. And I think that's the most helpful model that I have for explaining pre-trained behavior, which is becoming less and less the default of how LLM behavior should be explained, I'd say. I'd say like RLHF is becoming more prominent, but it's still necessary sometimes. That there are details of like wild behavior that have to be understood in terms of what plausible theories exist as to like what kind of text this might be. That if you were to approach this as a text prediction task, like in the manifold of all possible text, where are we?
13:18How well are we located by the prompt? And would an example be helpful of kind of how this plays out, like you kind of prompt in a kind of response and like how you apply this thinking? So the example that Reynolds and McDonald gave, this was used as their rationale for a zero-shot prompt they discovered that outperformed a 10-shot prompt for French to English translation on GB3 at the time. Which at that point, there were really only sort of two modes of using GB3, right? You could either give it part of a document and it would continue it, or you could give it K-shot examples of a task being performed and it would get the gist and perform the task.
13:55So it was news that you could do something in zero shot better than you could do it in 10 shot. And the way they did it is a technique that it's spread pretty far. I'd say this technique of flattering the model, of telling it that it's an expert. So I think the prompt they used was, a French sentence is given, colon, and then you give the French sentence. and then you say the masterful French translator flawlessly translates this sentence in English as colon and then you hit complete and this gives you better than 10 shot performance and the argument for why this works is that it's sculpting the manifold of possible translations better than saying French colon French sentence English colon right because there are many possible contexts in which a translation might exist and in many of those contexts translation is not correct that translation could be a joke.
14:44It could be just a bad translation. It could be a log of translation software that has bugs. There are many contexts where it could be wrong. And more importantly, perhaps, is that its understanding of what things are possible is very imperfect. So any help you can give it to sculpt it down to only the set of things that are ideal helps it collapse more of the probability into those tokens. And I think it's a very elegant demonstration of it, that by telling the model it's an expert, it becomes an expert. And by the way, that technique is much less important these days. I think there's a paper that showed with each successive version of GPT-3 as they went between different iterations of instruction tuning and RLHF, the effect of telling the model that it's an expert is decreasing.
15:27Now it's fairly negligible. The RLHF has... Presumably because it's baked into the instruction tuning? Right. The presumption that the answer should be correct is already well instilled in the model. One reaction to that example, you know, it almost demonstrates the possible existence of some theorem that kind of equates the perfect zero shot prompt to chain of thought. Something like there exists for every working chain of thought series of prompts, zero shot prompt that is just so elegant and on point. Granted, this is kind of extrapolating a lot from a data point and presumption, but does that kind of resonate at all with your experience one way or the other?
16:10Resonate or anti-resonate? I'd say it depends a lot on the problem. K-SHOT examples often have a lot of systematic issues that can be a problem in some cases. A good example is if you're doing general question answering, using K-shot prompts will often introduce a confusion over to what extent any given example might refer to something else from a previous example. You can imagine in the Reynolds and McDonald view of the world, right? So, which is also what I like about their model, by the way, is that it sort of explains why K-shot prompting works. The K-shot prompting narrows you down to the set of all possible documents that contain this many repetitions of the task.
16:44And surely the next one will be the same. The more there are, the more likely that is. It makes sense in that way. But one of the issues is that you'll see bits of information bleed from one example to the other, particularly the most recent example. It's especially bad if you're doing, let's say, unrestricted input from the user and the user is free to type what was just said or, you know, like what was that thing above? Or more realistically, they'll say something that's vague and it will be interpreted as referring to the thing that was just used. And I believe that there is a study on why it's outperformed by ZeroShot so easily is because of this issue.
17:16is that in translation, there are many possible sentences. It's the full space of all possible sentences. You might need to translate something that ambiguously refers to something previous, something that was previously said. And so most of the bad translations are instances of this. And I'm realizing an assumption that I made that may not be correct is that the K-shot translation prompt was successive instruction or chain of thought or, you know, it was representative of a process as opposed to just simply sharding the material to be translated. If it was the latter, then the question doesn't even really apply.
17:50But I was kind of trying to get at, you know, if it resonates with you, this idea that we only have to do K-Shot tricks like chain of thought because our prompts aren't good enough. Like there's some equivalence between the space of K-Shot prompts that work and the space of ZeroShot prompts that work or not. I'd say I'm not that bullish on what can be done purely through zero-shot prompting. There's a lot you can do, but it's not everything. There are techniques that work better. So if you have complete access to the model, let's say you're using GPT-2, say prompt tuning works very well, right?
18:25So you can tune the internal weights of the context to optimize the probability that it will return some desired response for a large number of examples, right? So if you have thousands of examples of how you'd like to respond, you can tune better better internal context weights than you can through any text prompt. The tuned prompt doesn't actually correspond to any real text. There are limits on what can be represented in text. And so I think prompt engineering shouldn't be seen as, it's rarely the final answer to anything. It's noteworthy. There may be instances where a prompt is the best solution, like capabilities that the model just like has down very well, like translation.
19:02I'd say that that's possible, or it'd be very hard to fine tune it to be better translation, I would guess. But those are more the exception than the rule. So we went down this path after exploring the first of two mental models or the first of two major mental models. The first being that the role of prompting is to kind of whittle down the space of responses, would you say, or the space of, you know, how would you, how would, or did you say that? I'd say that's right. The space of possible texts. So that I think of it as something like an interpolation of the pre-trained data set. it interpolates the gaps between the training examples that it has.
19:42And when we prompt it, we are sculpting that interpolation. And we're shaving off various dimensions of it and flattening it out and reducing the space of possibilities to a smaller model. And the second of those mental models? I'd say the second most important one is probably understanding RLHF, which is a lot to understand, I'd say. I think maybe the most detailed mental model I have of RLHF is that I think of it as being fundamentally the same thing as the Reynolds and McDonald view of a multiverse of fiction. But that fiction, that text that you're modeling is very abstract. What it is, is the policy rollout.
20:17During RLHF tuning, many possible completions are generated. Well, so I mean, this isn't literally what happens, but it's mathematically similar to what would happen. So PPO updates are sort of mathematically similar to what would happen If you did the policy rollout, if you generated many possible completions, evaluated all of them under the reward model, and then weighted their representation in just an MLE training epoch by what the reward model thinks of them. It's still predicting a text, but the text that it's predicting is the subset of all text that the model can generate that it thinks we would approve of.
20:53And that approval really guides its answers more than what distribution the pre-trained data has in many cases. And I think people jump to the pre-trained data as an explanation for its behavior maybe too often, recently being since the start of 2023, I guess. I think that people have sort of lost the intuition that it's unusual that the model responds to our questions at all. That if you ask a pure pre-trained LLM, what is the capital of Germany, then what is the capital of Spain is a reasonable response. It's not just a list of questions. And also, I think people forget that despite the name, that's largely what instruct tuning does, is it instills those kinds of assumptions, that questions should be answered, that directions should be followed.
21:35And the fact that it has any opinions at all on our social mores is secondary. And that's not an essential part of what it's doing in some sense. The more important thing is just that it now follows orders. And that, I think, is it's underemphasized as to what's being achieved here and how essential that is, right? That people take it for granted that you can just walk up to it and talk to it. Probably the third most important, like various mental models, I think, is the mechanism of autoregressive inference. So like a great example of this is I saw someone on Twitter, which is part of why I love being on LLM Twitter, by the way, is I get to see how like ordinary people interact with these models.
22:12And they objected really strongly to the fact that they had prompted the model with a question where it said it didn't know or it gave some wrong answer. And then when they pointed out its mistake, it knew the correct answer. Right. In a way that that made it that if it were human, you would say, well, clearly you knew this all along. Right. Right. And to me, I would never have expected it to behave in that way. Right. Like it seems obvious to me that if it says I don't know, that's only very weak evidence that it doesn't know. There could be some good way to prompt it. An example I refer to a lot is I mentioned earlier that it has memorized many like SHA-1, MD5 hashes, an even larger set of MD5 and SHA-1 hashes that it knows if you give it the first four characters.
22:54So it can predict the rest once you get it started. So it's hard to draw a line around. So does it know or does it not know? It's very prompt dependent. Yeah, exactly. Right. So it's very hard to tell what might actually be buried in there. You just don't know how to recover it. And I think the explanation of like why that happens really requires you to think about like autoregressive inference of the fact that there is a token sampling process that your prompt is incrementally getting longer. And sometimes this process is unlucky. Or another example, I think that can really only be explained this way is including, I think, GBD-4.
23:24So if you ask like chatGBT on GBD-4 to produce certain kinds of very out of distribution data, a good example of these are data URLs. So like URLs have like a data protocol where you just put in the binary data encoded into ASCII. And if you ask the model to generate one of these, it will often sort of completely derail in the middle of this generation. It'll get a few lines into producing this, and then it starts producing very out-of-distribution tokens that are not valid encodings at all. And then it never recovers. It gets stuck repeating the letter A over and over again forever or something like that.
23:58And to see why this happens, you have to be thinking about the probability of drawing out-of-distribution tokens, that there's some probability at any given token, there's a probability of sampling something that is completely out of distribution and for which it has no idea how it should continue. And the more of those tokens that are sampled, the higher the probability that another one will be sampled. And eventually that probability explodes. If you get too unlucky, eventually you're only... Well, not because you're continuing to generate out of distribution tokens, but now the tokens that you don't want are more in distribution.
24:31What the model believes is in distribution for the prompt is very much out of the distribution of natural text. It's sort of analogous to like a record needle popping out of the groove, right? That it's no longer like falling back into this tractor of what text is plausible. It just drifts further and further away into repetitious nonsense. Those are probably the biggest three that you have to keep in your mind to sort of understand like the practical failures that happen. And I think the third one keeps popping up in a lot of unusual ways. like there was recently a paper on pause tokens that showed that models that were pre-trained to use tokens that have no effect but to like let it stop and think for an extra token will improve answers in many cases it makes a lot of sense to me a priori just because the amount of thought that happens per token right the amount of floating point operations that are computed for every token generated is constant right so it's thinking a constant amount per token generated And when you consider that restriction, chain of thought seems inevitable.
Read the full transcript
25:31Of course, it can't just jump straight to the answer. It's too much thinking. It literally thought harder about it if you let it talk longer. And so I had done experiments on my own, actually, testing the hypothesis. Well, does it have to just think about that, right? Could you just like have it, say, emit a bunch of hyphens and then try to answer after the hyphens and will that improve it? And the answer to that is no. It turns out the thing that they did differently in this new paper was that you have to pre-train it with the POS tokens as well. You can't just fine tune this behavior into it.
25:58But it makes sense. And I think it really shows that our core model of how we're generating this text, of using auto-aggressive inference, is probably isn't the final word. The fact that we need things like chain of thought, that we need consensus, that we need to have tool use and ground the model in all these other ways. I think there's a gap there that that's suggested. The limitation we put onto it, I sort of describe it metaphorically sometimes as the model doesn't talk to us, it freestyle raps, right? It has no opportunity to stop and pause and think about anything. It just has to keep the meter going.
26:30It has to keep talking. And under that restriction, it's no wonder why it hallucinates. Like, so would anybody that has to freestyle rap an answer. I think there's a world of other ways that sequences can be represented and translated into text that are less trivial and that would perhaps work better. Like Sequence Match is another paper that I'm really interested in where they trained a model. What they did is they corrupted the pre-data with some epsilon probability to include random tokens followed by a special newly added backspace token. So what this instills in the model is that if you sample anything out of distribution, just hit backspace.
27:06This improves its answers, right? Because it has the ability to just say, oh, wait, no, never mind. That's wrong. Once it has one more token to think about it. I think that there's a lot of unexplored space of, I mean, text is the universal interface, right? So you can represent a lot of variations of how time flows and rewinding and pausing through sequences of tokens than just the simple translation we're doing now of byte-pair encoding. I'm always very interested when papers come up that are trying new methods along those lines. The flip side of your mental models are the mental models of folks who have not invested the time into understanding LLMs.
27:47Are there a similar top-end erroneous mental models or unhelpful mental models that if you're aware of these things will improve the way you prompt beyond just doing the things that adopting your mental models? I've sometimes had the concern that as these models become more usable, and as they do things that seem more sensible as behavior for a chatbot, we're losing the fact that it used to be self-evident that the Reynolds and McDonald's interpretation of the multiverse of fiction was appropriate. So before RLHF, so I think this applied all the way up to Text Da Vinci 002, if you prompted the model with simply, who are you?
28:30It would respond with something like, I'm an undergrad student in Minnesota, right? Like I'm studying computer science and I'm doing this as a part-time job. Or say, like, I'm on lockdown from COVID and I'm doing this to earn extra money on the side or, you know, things like that. And it was very obvious that you weren't getting a real answer, right? That you were getting some kind of text prediction that it's just making things up. Now, of course, if you ask it, who are you? It says, I'm ChatGPT. It was trained by OpenAI. But that behavior is relatively recent and it appeared piece by piece. There was actually a few months after ChatGPT was released that it didn't know its own name, but it would refer to itself as Assistant with a capital A.
29:07I'm imagining that had something to do with demands for white labeling in the future, but it was introduced piece by piece. And I think we maybe didn't give a lot of thought to what intuition is being lost here, that people are now surprised. Like the example I gave earlier of somebody being surprised that it said it didn't know something, but then clearly later it did. That behavior isn't surprising when you think of it as this improv partner. The thing that no matter what you say, it will say yes and then keep going. There's some damage being done there. And maybe it's, I think it's good in the net.
29:38I think that usability is important and it's a good thing that people can just walk up to these models and talk to them, But it invites more anthropomorphism than is really warranted. I asked a question earlier about kind of how you applied a mental model. And when you're trying to solve a problem using prompting, starting from, you know, a new problem that you've not seen before, blank screen, whatever, like how do you approach that? Granted, of course, that you have these mental models. But how do you think about solving problems, you know, assuming challenging problems using prompts? A lot of it, I would say, is restructuring problems in a way that avoids the known issues with LLM reasoning.
30:19So one good example of this, turning problems into checklists, like a style of least to most prompting, sometimes called, of having the model make all of the easiest and lowest level determinations first, and then make some high level judgment. Because if you do the reverse, what you're getting is hallucinated, right? So if you ask it to, say, give me your final answer, and then give me the rationale of how you arrived at it. It's rationalizing, not thinking. And that, I think, is probably one of the most ubiquitous techniques of taking a problem, deconstructing it into that sort of dag, I guess, of decisions that need to be made leading up to some final conclusion.
30:57Very similar to what a machine learning engineer would do when breaking up a problem traditionally into features to be engineered, right? It's feature engineering for the model. Probably number one. And an implicit part of that, I'd say, is avoiding just the known categories of things that can't do well, having an intuition of what kinds of things would be in its general knowledge, knowing that it's okay, but not perfect at turning English into regex, or that it's very bad at calculation, bad at counting things, but okay if you number them. There's a lot of data structures that I'll often use just for the benefit of LLMs.
31:32If I have, say, JSON serialized data, and I have a list of strings somewhere in this object, I might turn that list of strings into a list of tuples, where the tuples are just an index and then the string, right? Just to make it really, you know, easy for it to say, where are we in this list? And that's, I think you do a lot of that kind of translation in prompting. But I think, you know, overall, it's bad to overthink prompts too much. I mean, there are exceptions to this, but prompting is usually scaffolding, right? Usually you're using your prompt as a minimum viable product to collect data that you can curate and filter down to the best examples, that you can have human labelers correct and give feedback on, and that you can use to tune a more perfect model.
32:19And, you know, that's not always the goal because, you know, for cost reasons, usually. But if quality counts above all else, that's the goal, right, is to collect examples and to get out of the regime of prompt engineering and hoping that like this works, especially in cases where you have prompts that just require too many edge cases. Even something like you're running a search engine that you think of like what Bing would have to be prompted with. It's pages and pages of rules and conditions of, you know, what exactly you can say about politicians and so on. I think that you usually want to get out of that.
32:52I mean, there's different options for how you get out. There are tricks you can do in open source models with caching context. And there are other tricks, but usually like the default escape route is to fine tune a model. I'm interested in prompt engineering as something that's present to varying extents along that process. that you can sort of see a continuum between building up a good K-Shot prompt and, you know, refining your K-Shot examples to be very rich in edge cases and diverse and, you know, representing all of your important rare class as well. You can go from there to K-Shot selection that you have some database prompts or database of K-Shot examples and rules for selecting the best one given the user's input, which could use vector embeddings, could use train, you know, another model for that.
33:35And that expands all the way to things like retrieval augmented generation and eventually to fine tuning. Like one prompt that I'm often really inspired by is the prompt using Copilot, which sort of blurs a lot of the lines between, I mean, I guess technically it's retrieval augmented generation, but it's not using embeddings. It's using static analysis. When you hit complete using GitHub Copilot, what it's doing is it's sending to the model a fill-in-the-middle completion. So the model's been tuned to fill-in-the-middle. And everything before the cursor is the prompt and everything after the cursor is the suffix.
34:09But it prefixes to this prompt a common header. So it just checks what language you're in and then says, what is the appropriate way to do a common header in this language? And then it inserts into that common header all these random bits of context that it thinks might be useful. It does static analysis on your file and sees that you imported this function. Well, it goes to that file, checks the definition of that function. And if it's short enough, it includes that into the header. And if it's really long, then maybe it just leaves the body out and just gives you the arguments and their types.
34:35And things like, it makes decisions like that. And all those little decisions are important for getting the best autocomplete. So I think a lot about, I guess, that continuum. Because in practice, many tasks call for a mix of these methods. That you're doing some mix of writing instructions that you think cover enough of the edge cases. giving it K-shot examples that you think cover the rare classes or illustrate the things that aren't being communicated well through instructions. Like say style is very hard to communicate through instruction, but very easy to communicate through example. And you have to sort of mix these methods together.
35:09I remember when I first started thinking about this explicitly is I was choosing K-shots from a dataset and I was trying to find one that had a typo in it. I wanted the model to understand that for this input, this task, typos are irrelevant. You should just act as though it was typed correctly. And it occurred to me, well, why do I need to find one? I can just make one, right? You can just take any given example and introduce a typo. And then once I had that thought, I realized, well, I should do this all the time, right? I should stop using real examples. I should just construct ones that are artificially rich in all of the rare classes that I want to demonstrate.
35:39That if there's anything that could be empty, make it empty, right? You don't need a list that needs to have the edge case of an empty list or whatever, right? And that's something that guides a lot of my prompt writing when I'm using KShots for these first attempts at task-specific model is that you want to like capture the boundaries of the distribution, right? You have to think about what is the distribution of your inputs, where are the corners of that, and how can you draw the right boundary around it? Interesting, interesting. I think that's a super helpful way to think about prompting. And I just kind of the very high level takeaway that I got from that is part of it is being clever in crafting the language that you use and give to the model.
36:22But a bigger part in your estimation is not as much linguistic cleverness. You know, and the thing that comes to mind is you are a masterful translator, blah, blah, blah, that kind of thing. But more, you know, what you call scaffolding structure and using structure to provide context to the model so that it can respond in the way that you would want it to respond. Is that a fair summarization? I'd say so. And I think on the point of round, like people often ask me how valuable is like writing style, like writing ability. I think it's maybe less necessary to be like expressive in your writing as it is too much to be a domain expert in what it can and can't do well.
37:05And the quirks of like how to get it to complete certain kinds of text, like knowing oddities that like it's hard to instruct on style, but easy to give examples of style. I think those sort of quirks and idiosyncratic details might matter more overall. There certainly are, though, creative tasks where, you know, if you're tuning a model to be creative, the quality of your training data does matter. But usually it's less valuable than people think for like prompting something like ChatGPT. It tends to work regardless of how poorly you type. Mm-hmm. Yeah. Beyond those papers that you've referenced already, the Hofstetter article, for example, What are the kind of the key resources that you would point someone in this domain of better understanding LLM capabilities, emergent capabilities, and kind of advanced prompting?
37:53Anything's come to mind? For prompt engineering, it's hard to find a lot of good holistic resources just because the field moves so fast. I was pretty impressed with learnprompting.org as a sort of a collection of techniques. There's a lot of good references in there to papers. It's very well organized. I probably find out about most new things just through Twitter. That's where I find most of my archive links that I'm reading. So I don't have a lot of great ideas of where else to look. That's usually where I find mine. Yeah, it's just so early that the book still needs some months in order to be published.
38:23And when it is, it'll be outdated anyway. Right, which is ironic, given how ostensibly easy it is to write a book these days with LLMs. Absolutely. Well, we've covered just a small part of what I'd hoped we'd cover. There's still topics like red teaming and adversarial prompting and things like that. Hopefully we'll get you on for a part two. But in the meantime, Riley, thanks so much for sharing a bit about what you've learned about LLMs. All right. Thanks so much. I really appreciate you having me. Awesome. Thank you.
39:02All right, everyone. That's our show for today. To learn more about today's guest or the topics mentioned in this interview, visit twimla.ai.com. Of course, if you like what you hear on the podcast, please subscribe, rate, and review the show on your favorite podcatcher. Thanks so much for listening and catch you next time.
From the publisher
Today we’re joined by Riley Goodside, staff prompt engineer at Scale AI. In our conversation with Riley, we explore LLM capabilities and limitations, prompt engineering, and the mental models required to apply advanced prompting techniques. We dive deep into understanding LLM behavior, discussing the mechanism of autoregressive inference, comparing k-shot and zero-shot prompting, and dissecting the impact of RLHF. We also discuss the idea that prompting is a scaffolding structure that leverages the model context, resulting in achieving the desired model behavior and response rather than focusing solely on writing ability.
The complete show notes for this episode can be found at twimlai.com/go/652.




