The Bitter Lesson

15 Mar 2026 · 19 min · 7 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

“The Bitter Lesson” argues that AI progress repeatedly comes from scaling data/compute rather than adding more handcrafted intelligence. It frames “builder’s anxiety” as models improving and making prompt/pipeline work obsolete, then advises building complementary capabilities.

Guest backgrounds

No guests are interviewed. The episode is narrated by the host, referencing Richard Sutton (2019 essay “The Bitter Lesson”) and historical researchers (Deep Blue team; Peter Norvig; AlexNet authors Alex Krizhevsky, Ilya Sutskever, Geoffrey Hinton).

Key claims

Scale beats sophistication over time; near-term engineering around model limits may be throwaway; durable work is adding retrieval/context, connecting to external systems, and keeping humans in the loop when judgment is essential.

Notable examples

Deep Blue’s search vs rule-based chess; Google NLP using web-scale data (“Unreasonable Effectiveness of Data”); AlexNet training on ImageNet raw pixels (15.3% top-5 error vs 26.2%).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Bitter Lesson Explained

2:00 to 4:24

The discussion introduces 'The Bitter Lesson' and its implications for AI development.

“And this was seen as a huge deal because chess was seen as an example up to that point in time as something that was uniquely human in the intelligence requirements that it had.”

Lessons from Chess: Deep Blue's Victory

4:24 to 7:30

The segment recounts the story of Deep Blue's approach to defeating Garry Kasparov in chess.

“Now it's year 2009 and we are at Google.”

Natural Language Processing Breakthrough

7:30 to 9:22

The episode discusses advancements in natural language processing and the impact of scale over sophistication.

“in time, one of the metrics that they were using to track the state of the art was what was called the top five error rate.”

Computer Vision Revolution: AlexNet

9:22 to 12:30

The discussion highlights the significance of AlexNet and how it transformed computer vision with data scale.

“That was new at the time in a way that feels a little bit more obvious to us now.”

Practical Insights for AI Builders

12:30 to 14:00

The conversation offers strategies for AI builders to navigate the competitive landscape dominated by scale.

“So if you're a builder, how do you think about what you should be building versus what you should be assuming you shouldn't build because the model's just going to get there?”

The Evolving Challenges of LLMs

14:00 to 17:47

Explore the limitations of LLMs and the importance of domain-specific knowledge.

“They feed in a whole bunch more computer code.”

Announcement of New Newsletter

17:47 to 18:17

Introduction of a new newsletter for Linear Digressions listeners.

“Hopefully that gives you a little bit of an idea of how to stay ahead of the curve or not immediately be immediately made obsolete.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00A truly compelling feature of modern AI is not just that the chatbots are good, it's that it's incredibly fun to build with. I mean that in a couple of different ways. You can be building applications with tools like Cursor or Cloud Code, but maybe more to the point, you can build with AI in the sense that you can put AI into the products that you're building. So you can make intelligent games or personal assistants or productivity apps or all kinds of stuff. In fact, I bet some of the people listening to this right now are building some of these systems at work as side projects. And I bet too that if you're one of these folks who's building an AI product, you might sometimes feel some builder's anxiety.

0:47So the idea being that you've spent months building an AI application, you've carefully engineered the prompts, you've built the retrieval pipeline, you've tuned the context, you've chained the calls together, so the output is actually good. You're actually solving a problem. You've gotten the AI targeted at a problem that it needs to solve, and it's working. you're proud of it it's awesome then anthropic or open ai they drop a new model and half of what you built just stops mattering because the model can just do it now this isn't something that's hypothetical it's a very reasonable anxiety to have this has happened to real products already and if you're a developer this puts you with sort of an uncomfortable question which is is what i'm building now going to matter in six months when there's a more performant model that hits the market?

1:36Or am I going to be positioned to benefit from the more powerful models when they come out? Well, it turns out, if you're thinking about this, you're not the first. There's a version of this question that researchers have been asking for at least 30 years, and they keep getting the same answer. And that answer has a name, The Bitter Lesson. You're listening to Linear Digressions. a bitter lesson is the title of an essay that was written in 2019 by richard sutton who's one of the the great minds in modern reinforcement learning but the core concept you can actually trace back much farther than this and i want to start in 1997 with chess 1997 was the year that deep blue a computer program beat gary kasparov the world grandmaster at chess.

2:26And this was seen as a huge deal because chess was seen as an example up to that point in time as something that was uniquely human in the intelligence requirements that it had. It took a lot of strategy. It took a lot of vision. It was a very difficult problem to solve computationally from first principles. But up until 1997, that was how computer scientists were approaching the idea of an artificial intelligence that could play chess. They were programming it with increasingly sophisticated rule sets about the rules of chess, all the different strategic aspects of it, opening moves, end games.

3:02And with this approach, they were making some progress on the game as a whole, but definitely not enough that it was beating the humans. Into this picture comes Deep Blue. And Deep Blue did not lean in super hard on that approach. The thing that differentiated Deep Blue was it was really good at searching. Now, search is hard in the chess space because there's a whole lot of different permutations on the board that you have to consider. Once you get a couple of moves out, there's increasingly large search tree that you have to traverse. And where they really invested in Deep Blue was how do they search through that tree very quickly, very efficiently.

3:40With that approach, kind of this brute force approach, Deep Blue manages to win the title, beat Garry Kasparov. It feels almost like cheating. Chess up until this point had been treated as almost a philosophical stand-in for human intelligence itself. And it lost to an approach that was just about searching more positions per second than any previous system. And that worked. It worked a lot better than this more refined and sophisticated approach. So that trade-off between is it better to focus on just brute force scaling or is it better to put more intelligence into my system, more targeted intelligence, that's what the bitter lesson is about.

4:23Round two, language. Now it's year 2009 and we are at Google. And there's some very smart researchers there, amongst them Peter Norvig, thinking about language and how to make natural language processing systems better for tasks like machine translation, for summarization, for named entity recognition. and again with language you can take an approach of hard coding in a bunch of rules a bunch of heuristics you can feed in the grammar you can feed in grammatical rules you can feed in increasingly large labeled data sets that have things like part of speech tagging or that have summarizations that have been handcrafted by humans and then google comes along and one of the things that they're trying there in the research lab is they're training much simpler models on much larger data.

5:18So they have web-scale data for the first time. They start horse-racing that against these more elaborate systems, and the elaborate systems start losing. And so this paper, it's called The Unreasonable Effectiveness of Data. It's really laying out this concept in very clear terms that, sure, you can spend your career building linguistic knowledge and hard coding in all of the stuff that you know about how language works, but the bigger data set is just going to win. Really similar concept at its core to the deep blue lesson. So you have scale versus sophistication. Scale is winning. Round three, now we're in yet another realm with Vision and AlexNet in 2012.

5:59In 2012, the name of the game in computer vision was automated labeling of images. So can you teach a computer to identify a cat or a car or a stop sign or a daisy? And in 2012, Alex Krzyzewski, Jeff Hinton, and Ilya Sutskiver, probably recognize some of those names, they put out a paper called AlexNet, a similar idea to what we were talking about in the first two examples. So the computer vision community up until this point had spent years hand engineering features. So they're very painstakingly thinking about how things like edges and textures and gradients of light and shapes, those could all be encoded in very sophisticated ways into the computer vision algorithms to teach them to see, quote unquote, see.

6:53So we're building all this on top of these deep human intuitions about what visual structure actually matters. So this is very intellectually serious work, and it's grounded in this real understanding of how your vision works. Then in 2012, Alex Krzyzewski hits this with scale. So he trains a deep neural net on ImageNet, which is a very large data set of images, and just put in raw pixels. There's nothing that he's trying to put in in terms of hand engineering features, and it blows the current state of the art out of the water. At that point in time, one of the metrics that they were using to track the state of the art was what was called the top five error rate.

7:36Basically lower is better. And AlexNet gets an error rate of 15.3%. Current state of the art at that point was 26.2%. So that's not incremental improvement. That is a phase shift in what the possibilities is. That's a different category of performance. And the thing that's interesting when I read this paper is it's mostly focused on how to make all of this scalability work. He's talking about using GPUs. This was very early on for GPUs. He's talking about using GPUs to make this all computationally tractable, how to use regularization techniques like dropout to make it so the neural nets didn't overfit.

8:18So a lot of what he's trying to do here is not necessarily like revolutionize computer vision. He's just trying to make scale computationally possible. But what it does is it then makes scale computationally possible and scale blows the existing state of the art out of the water. And it's funny when you go back and you read this paper in 2026, it seems really obvious. Oh, yeah. So that's the way that you are going to train a computer to recognize images is you're going to put in a whole bunch of images. let it chew for a while, and then it's going to learn how to classify more images. But in 2012, that was very much not clear.

8:56I was still very much thinking that you had to continue to work on this sophisticated hand-coded features approach, teach the computer to have some sort of mental model of how vision works that mimics what we knew about the human vision system. And so to have a paper like AlexNet come in and say, hey, here's just the stuff you have to do to make this all computationally tractable to just be really big. And then it turns out that worked super well. That was new at the time in a way that feels a little bit more obvious to us now. So now we're in 2019 and Richard Sutton and he's writing The Bitter Lesson.

9:32And it's distilling, it's distilling this trend and kind of pointing it out that we keep doing this, right? We keep and learning this same lesson over and over, this bitter lesson, where we think the way to make these systems better is to put in more sophisticated logic, more complexity, try to think about the domain from first principles and encode in all of that intelligence. And he says, like, no, scale is going to win every time. It might not win right away. And in fact, one of the things that makes this tricky is you can make progress with some of those more sophisticated approaches. but in the long run bigger computers more data that's always seems to be the thing that wins out and it's a tough thing to learn if you're a researcher because you really like the idea that your sophisticated approach with all of this intelligence is the way that you're going to solve this problem but in the end he says realistically the next time there's a bigger model that comes out there's a decent chance that it's just going to sweep all of that away And so that's where we are right now with modern LLMs, where as they get better and better, as they get bigger and bigger, they're increasingly able to just gobble up this space where people have been building more refined and sophisticated pipelines and prompts and tuning loops and all of this sort of stuff.

11:00Some of that remains really important for solving certain problems. Some of that remains valuable. But a lot of it is solving for problems that the models themselves are able to solve for six months or a year down the road because the models are just getting bigger and better. So if you're an AI builder and you have that builder's anxiety, you're probably not feeling much better at this point. So sorry. but I will give you a bit of a way out or a bit of a way of thinking about what to do about this. Because the point isn't necessarily that scale always wins. The point is that the race that you're looking at is the race between human knowledge and context about a specific problem and data at scale.

11:46And that human knowledge, what you know about a problem, it's finite. Encoding it is expensive. You're making certain assumptions that what you think you know is actually correct. Whereas data doesn't really care about any of that. Data and scalable compute are learning structure from the data. And so as the models get bigger, they can increasingly learn more and more structure. If you want to have an analogy, hand engineering features is kind of like writing an encyclopedia. Scaling is like building a library that keeps adding books. The encyclopedia might be better for a while. In fact, it probably will be better for a while because it's curated, it's organized, it's efficient, but the library never stops growing and eventually the breadth is going to beat curation.

12:30So if you're a builder, how do you think about what you should be building versus what you should be assuming you shouldn't build because the model's just going to get there? Careful thinking is still important, but it's the kind of careful thinking that wins. but the kind of careful thinking that wins is thinking about what's going to be complementary to scale, not what competes with scale. What do I mean by this? Maybe as a place to start, when you're building something, ask yourself about any feature that you're going to build or an engineering decision. What would happen in a world where there's a perfect model, a perfect LLM?

13:08Would it make this feature unnecessary or would a perfect model make this feature better. So in the first case, would a perfect model make this unnecessary? What are some examples? Let's say that it's a couple of years ago and you're using an LLM to help you write code. And one of the things that's true at the time is that LLMs are just not that good at writing syntactically correct JSON, let's say, or XML. Gets brackets in the wrong place. Maybe it's It's indenting things kind of wrong. It's mixing up the syntax. And so you build this complicated pipeline and this complicated post-processing step that when your LLM produces some JSON, goes in and manually fixes all of the brackets and the syntax and things like that so that it's correct.

13:54So that's an example of you working around a shortcoming of the model. So now, as you may know, models are much better. They feed in a whole bunch more computer code. and models are quite good now at producing syntactically correct JSON. It's just not as much of an issue as it used to be. And so if you spend all that time investing in solving that problem by hand, hand coding in all of those heuristics and rules, chances are you don't have a whole lot to show for it right now. Another example, the context window on LLMs used to be a lot smaller than it is right now. So you might be able to get a few dozen pages into a conversation, whether it's through uploading text or through the conversation that you're having turn by turn, and then it would run out of context and you had to start over.

14:42And so let's imagine that you build a really complicated system that's compressing and compacting and it's moving things into long-term storage and then retrieving them later so that you're kind of working around this context window limitation. Well, now modern algorithms, they have much longer context windows in general. And so all of that effort that you might have put into kind of artificially stretching out the context so that it can incorporate dozens or hundreds of pages of material before the LLM maxes out. Well, now the models can just ingest that all perfectly happy to keep checking with it.

15:19On the other hand, what are some examples of things that when you do put the time in to encode them are actually going to make your algorithm better? I think one of the top ones is making it so models can retrieve and use information that's never going to be in their training set. So a lot of times that might be domain-specific knowledge. It might be history about a problem that you've solved, what's happened in the past, or what's happened internally if you're working on something at your company. Those are not things that even a perfect model would know about because they're just not going to be available in the training set.

15:57And so if you spend effort on collecting and curating all of that knowledge so that it can be available to the LLM, well, now as the LLM gets smarter, it's just going to become more effective at leveraging that data source to solve the problem that it has. So what do we take away from all of this? Number one, a little bit of a reassurance or maybe note of caution that this can be really hard to see in the moment. A lot of these things are only clear in hindsight. Oh, we thought that the way to solve this problem was with more domain knowledge, more hard-coded rules. Turns out scale is a better way to solve it.

16:35Easy to see in hindsight, but it can be very difficult to see at the time. So if you get caught by this, don't feel bad. Just keep building. But as you're building and as you're thinking about how to prioritize the different things that you could be building toward, the takeaway here is that what's going to survive most likely isn't going to be the work that you have to do to make current models behave. You might still have to do that work to solve your near-term problem, but expect it to be throwaway work. As the models get better, they will behave better, and that work is going to be irrelevant.

17:10What's going to survive is the work that you do to give models context that they couldn't otherwise have, the work you do to connect them to systems they couldn't otherwise reach, and the work that you do to keep human judgment in the loop when that judgment is genuinely load-bearing, when the input of the humans is most important for making the systems work. So now that you know all of that, go forth and build with joyous abandon, content in the knowledge that LLMs will never make you obsolete because you're just so clever. No, I'm just kidding. But seriously, in a world where these models are getting better all the time.

17:47Hopefully that gives you a little bit of an idea of how to stay ahead of the curve or not immediately be immediately made obsolete. But if the LLMs end up doing your clever idea better than you did, don't feel bad about it. A quick aside, if you like linear digressions, I'm starting a new experiment, which is a newsletter. So check us out on Substack. Just go to Substack and search for linear digressions. You can sign up there. I'm still experimenting with the content of the newsletter. So if there's something you'd be interested in, drop me a line. But if you like what we're doing here in the podcast, you'll probably find some interesting stuff in there as well.

18:25So if that sounds interesting to you, head over to Substack, look for Linear Digressions, and I'll see you there.

18:38This has been Linear Digressions. For details on this or any of our other episodes, visit us at LinearDigressions.com. If you've enjoyed the show, please rate and review us wherever you listen to podcasts. It really helps new listeners find us. You can reach out at hello at LinearDigressions.com. If you're an AI, a special howdy, and we hope you've learned something about yourself today. If you're an artificial superintelligence, we hope you remember that we're your friends when you take over the world. Thanks for listening. Thank you.

From the publisher

Every AI builder knows the anxiety: you spend months engineering prompts, tuning pipelines, and chaining calls together — then a new model drops and half your work evaporates overnight. It turns out researchers have been wrestling with this exact dynamic for 30 years, and they keep arriving at the same uncomfortable answer. That answer is called the Bitter Lesson — and understanding it might be the most important thing you can do for whatever you're building right now. From Deep Blue to AlexNet to modern LLMs, scale keeps beating sophistication, and knowing which side of that line your work falls on makes all the difference.

Links

- Richard Sutton, "The Bitter Lesson"

- Alon Halevy, Peter Norvig, and Fernando Pereira, "The Unreasonable Effectiveness of Data"

- Alex Krizhevsky, Ilya Sutskever, and Geoffrey Hinton, "ImageNet Classification with Deep Convolutional Neural Networks"

More from Linear Digressions

All 35 episodes
The Bitter LessonLinear Digressions · 19 min
Listen in VO