Self-Adapting Language Models

12 Oct 2025 · 17 min · 8 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Self-adapting language models (SEAL) that improve by generating “self-edits” (natural-language instructions) to create their own training data and update their weights, using meta-learning.

Guest backgrounds

No guests are identified in the transcript; it’s a host-style “Deep Dive” conversation.

Key claims

Standard fine-tuning needs lots of perfect task data and doesn’t “assimilate” new knowledge well; SEAL uses an outer RL loop to learn effective adaptation strategies and an inner supervised fine-tuning loop to apply them.

Notable examples

On SQuAD-style reading comprehension, baseline accuracy was 32.7%, raw-text fine-tuning 33.5%, synthetic data from the base model 39.7%, GPT-4.1 synthetic data 46.3%, and SEAL (smaller model generating its own RL-guided synthetic data) 47.0%. On ARCAGI-like visual reasoning with few examples, in-context learning scored 0%, test-time training with base self-edits 20%, and SEAL-configured adaptation 72.5%. Limitations: catastrophic forgetting and high compute cost (30–45s per self-edit evaluation; ~6 hours per batch), with proxy rewards reducing RL evaluation time to ~5 minutes.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Challenges with Fine-Tuning AI

0:29 to 2:07

Exploration of the limitations of traditional fine-tuning in AI models.

“And the source material we looked at uses a great analogy for this.”

Introducing Self-Adapting LLMs

2:07 to 2:37

Discussion on the SEAL framework for self-adapting language models.

“It's a framework designed to tackle this exact problem, this AI stasis.”

Mechanics of Self-Edits in AI

2:37 to 4:25

Detailed explanation of how self-edits enable AI models to learn more effectively.

“So our mission today, let's unpack how SEAL pulls this off, this self-directed learning.”

Applications of SEAL: Factual Knowledge

4:25 to 8:11

Demonstration of SEAL's effectiveness in improving factual knowledge absorption in models.

“There are these two nested loops working together.”

Applications of SEAL: Few-Shot Learning

8:11 to 11:55

Analysis of SEAL's role in enabling models to learn from few examples without human intervention.

“It really speaks to the power of optimizing the synthetic data for the specific learning process.”

Limitations of Self-Adapting Models

11:55 to 13:43

Discussion of the challenges and limitations faced by self-adapting language models.

“It really does sound like a step towards models that can genuinely adapt and learn on their own.”

The Future of Synthetic Data in AI

13:43 to 14:01

Exploration of the implications of synthetic data generation for the future of AI.

“However, they did explore a potential workaround.”

The Future of Self-Adapting Language Models

14:01 to 16:28

Explore how self-adapting models may revolutionize AI learning and data generation.

“Things a human could quickly assess or maybe another model.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:28Welcome back to the Deep Dive. pick up some new facts or learn a new way of reasoning or just get better at a specific task well the old way fine-tuning it often just doesn't really move the needle much and it's not like the model isn't smart enough it's more about the process right fine-tuning kind of assumes you have tons and tons of perfect task specific data just lying around which in the real world you almost never do you might get what the single new article or maybe three examples of how to do something new right and that handful of raw data it's just not enough to effectively nudge, you know, the billions of connections inside the model.

1:02The weights, precisely. And the source material we looked at uses a great analogy for this. Think about human learning. Okay. Like a student cramming for a big exam. They don't just passively read the textbook cover to cover and hope it sticks. No, definitely not. You rewrite notes, draw diagrams, connect ideas. You assimilate. You restructure the information. You make it your own. You might augment it with other things you know. That act of processing, that's the key for making learning last in humans. But standard LLM fine-tuning, it completely skips that step. You're just feeding it the data as is, like pouring water onto a rock.

1:39And hoping some of it seeps in. Yeah. And if that raw data isn't formatted in just the right way for how these models learn, you know, gradient-based learning, then the learning is super inefficient. The model doesn't inherently know how to turn a simple paragraph into the best internal notes for remembering it later. Okay, so what we need is an LLM that can figure out its own study plan, like design its own best way to learn new things. You got it. And that brings us right to the heart of today's deep dive, self-adapting LLMs. They call it SEAL. SEAL. Okay. It's a framework designed to tackle this exact problem, this AI stasis.

2:13It lets the model generate its own instructions for learning. They call these self-edits. Self-edits. So natural language instructions that tell the model how to create its own training data or how to run the learning process itself. Exactly that. It's like the AI learning to be its own teacher, figuring out the best curriculum and the best teaching method for itself based on the new info it gets. Okay, this sounds fascinating. So our mission today, let's unpack how SEAL pulls this off, this self-directed learning. It uses something called meta-learning, right? And then we need to look at the results because apparently they're quite impressive on some key challenges.

2:50Yeah, two big ones. Getting new facts into the model and handling tasks with very few examples. All right, let's start with the core mechanism, these self-edits. What exactly are they? If it's not just outputting an answer, what does this strategy look like? So a self-edit is basically the model's own generated text that lays out how its internal structure, its weights should change. It's like a recipe for self-improvement. A recipe. Okay. Okay. What kind of ingredients can it specify? Oh, it's quite flexible. It can tell the system to generate synthetic training examples based on, say, a new piece of text.

3:24So making up its own practice questions. Kind of, yeah. Or it could specify the nitty-gritty details of the training itself, like the learning rate or how many times to go over the hyperparameters. It could even tell the system to use external tools, maybe for doing complex data augmentation, like manipulating images or something. Okay, but hang on. Adding this whole extra layer. You mentioned reinforcement learning earlier. Doesn't that make things way more complicated, like computationally expensive? Is this even practical? That's a really critical point. And yes, it is more complex. The core idea here is meta-learning.

3:59Learning to learn. Exactly. Learning the strategy for learning. Think of it like an athlete. They don't just train their muscles. A really smart athlete or their coach figures out the best possible training plan based on what's worked before. That's meta-learning. Right. So SEAL isn't just training the LLM to give good answers now. It's training the LLM's ability to generate strategies that lead to a better updated model later on. Precisely. It's optimizing the adaptation process itself. And the way it's built reflects this. There are these two nested loops working together. Figure 1 in the source really helps visualize this.

4:33Okay, walk us through these loops, an optimization within an optimization, you said. Yeah. So the outer loop uses reinforcement learning, RL. This is where the strategy gets learned, refined. The model generates a potential self-edit. That's the plan, the synthetic data, the instructions. Then that edit gets applied in the inner loop, which we'll get to. And then the system checks how well the newly updated model performs on the actual task, that performance score. That's the reward signal for the outer loop. Ah, okay. So the model tries out a study guide, metaphorically speaking. Takes the test, and the score it gets tells it whether that study guide was any good for learning.

5:09Exactly that. The RL feedback basically says, hey, that strategy worked well, try more like that, or nope, that didn't help, try something different. Okay, so that's the outer strategy loop. What about the inner loop? The inner loop is more familiar. It uses standard supervised fine-tuning, SFT. This loop takes the winning self-edit, the one generated and selected by the outer loop, and actually does the work. It executes the plan. Right. It uses the synthetic data, follows the parameter instructions, and runs the standard gradient descent process to actually update the model's weights. This is what makes the change stick.

5:43It leads to those persistent updates in the model's knowledge. That's a really clever setup. Use RL to figure out how to change effectively, and then use SFT to actually make that change permanent. Okay, let's see this in action. Application number one, getting new factual knowledge into the model. Yeah, this is a huge one, a constant headache for LLMs. How do you get them to really absorb new information, like from a news article or a research paper? So they can recall it later without needing the original text right in front of them in the prompt. Exactly. You want real lasting memory, not just look up from the context window.

6:19So how did they test this? How do you measure if the knowledge is really sunk in? They used passages from a standard data set called SQUAD. It's a tough reading comprehension benchmark. They took a model, QN 2.5 7B in this case, and here's the crucial part. Instead of just training it on the raw squad passage. Which we already know doesn't work very well. Right. Instead, they fine-tuned it only on the synthetic data generated by the SEAL framework. So SEAL read the passage and created its own implications, questions, facts based on it. That was the training material. Okay. And then they quizzed the updated model on the passage, but without giving it the passage again.

6:58Correct. Just asking questions based on the text it was supposed to assimilated. And this is where, as you said, it gets really interesting. The results tell a story. Let's hear the number. Okay. Baseline model, no adaptation at all. 32.7 % accuracy on the questions. Pretty poor. Then, just training on the raw passage text. Accuracy barely budged. 33.5%. Confirms that raw data isn't easily assimilated. Right. So the self-editing is needed. What happened when they used synthetic data? Well, first they tried synthetic data generated by the base Quinn model itself without the seal optimization. That got them up to 39.7%.

7:35Better, but not amazing. Okay. What about using a really powerful model like GPT-4 to generate the synthetic data? That should be good, right? You'd think so, and it was definitely better. Using synthetic data generated by GPT-4.1 pushed the accuracy up quite a bit to 46.3%. So state-of-the-art generation helps create better study material. Makes sense. But where did SEAL land? This is the kicker. SEAL, where the smaller Quinn model generated its own synthetic data, guided by that RL outer loop, achieved 47.0 % accuracy. Wow. Okay, let's pause there. That's huge. The 7 billion parameter Quinn model using SEAL self-adaptation created training data for itself that was more effective than the synthetic data generated by the much, much larger GPT 4.1.

8:18Exactly. It outperformed the giant. It really speaks to the power of optimizing the synthetic data for the specific learning process. So why? What made the SEAL-generated data better? Was it just more data? No, not necessarily more, but better structured for learning. The qualitative analysis they showed in Figure 5 suggests the RL training pushed the model to generate edits that were more digestible. Digestible? Yeah, like breaking information down into smaller atomic facts or generating really well-formed question-answer pairs, things that are just easier for the gradient descent process in the inner loop to actually learn from effectively.

8:55GBC 4.1 might generate a lot of text, but SEAL generated highly learnable text. Okay, that makes sense. Optimized for learning, not just for volume or fluency. So that first application was about teaching the LLM to remember things better. What about the second one? Autonomous configuration for few-shot learning. This sounds like teaching it how to tackle new kinds of problems. Precisely. This addresses the challenge where a model needs to figure out how to do a completely new task when it only gets shown a tiny number of examples, maybe just two or three. And critically, it has to do this without a human stepping in to tweak the learning strategy.

9:30Right. The model has to configure itself. How did they set up this test? They used a benchmark called ARCAGI, or a simplified version of it. It's all about abstract visual reasoning, looking at grids of colored squares, and figuring out the pattern or transformation rule from just a couple of before and after examples. Ah, those are tricky. Requires inferring rules. Very tricky. So they used a LAMA 3 model here, a smaller one, 1 billion parameters. And its job was to generate a self-edit that was essentially a complete configuration file for its own training process on this new ARC task. A config file.

10:06What was it configuring? Everything. It had to choose which data augmentation techniques to apply to the few examples it had, like should it rotate them, flip them, resize them, maybe chain a few augmentations together. And at the same time, it had to pick the best optimization settings, the learning rate, the number of training cycles, even the specific type of loss function to use. It was literally designing its own adaptation pipeline on the fly for each new ARC problem. Wow. Okay, so how much difference did this self-configuration make? What are the results here? The results are, well, they're pretty stark.

10:38Let's start with the baseline. Standard in-context learning. That's just showing the model, the examples, and the prompt and asking it to solve the new one. How did that do? 0 % success rate. Couldn't solve any of them reliably. Okay. Not surprising for ARC. What about adding some training, but without the smart seal configuration? They tried something called test time training, TTT, where the model does some quick learning on the examples. If they combine TTT with a self-edit generated by the base model without the RL optimization, the success rate got to 20%. Better, but still not great. Not great.

11:13But then they applied SEAL. They trained the model using that outer RL loop to learn how to generate the best config file for each type of ARC task. Yeah, success rate. Jumped to 72.5%. Whoa, from 0 % or 20 % up to nearly three quarters. That's not just improvement. That's transformative for this kind of task. It really is. It shows the model learned to look at the visual features of a new, unseen ARC problem and correctly predict the most effective adaptation strategy. Like if the example transformation involved symmetry, the SEAL train policy learned to include flipping or rotation augmentations in its self-generated config file.

11:52That made the limited training data much more powerful. That's amazing. It really does sound like a step towards models that can genuinely adapt and learn on their own. Yeah. But we need to be realistic. This sounds almost too good to be true. What are the catches, the limitations? Yeah, absolutely. The researchers are very clear about the limitations, and there are two big ones right now. First, the classic problem in continual learning, catastrophic forgetting. Ah, right. The model learned something new, but in the process, it forgets what it learned before. Exactly. If you update the model sequentially, teach it about passage A, then passage B, then passage C, its performance on passage A starts to drop off as you add more edits, the new learning overwrites the old.

12:33And SEAL doesn't magically solve this. Not yet, no. The source shows this clearly in figure six. Performance degrades over sequential edits. They suggest future work could add a penalty to the reward function, basically, rewarding the model for learning the new stuff without messing up the old stuff. But that's not implemented yet. Okay, so catastrophic forgetting is still a hurdle, what's the second major limitation? This one's more practical, but just as important. Computational cost. That outer RL loop, the one that figures out the best strategy, it's expensive. That was expensive. Well, to calculate the reward for just one potential self-edit, you have to actually run the inner loop.

13:12You have to fine-tune the model with that edit and then evaluate its performance. Woof. Yeah. They reported it took about 30 to 45 seconds for each self-edit evaluation on their hardware. and the outer loop needs to try many edits. So training this thing must take forever. Pretty much. For one of the experiments, optimizing just a single batch of updates took around six hours. Six hours per batch. Okay. Okay, that definitely limits who can actually do this kind of research or deploy it right now. Only places with massive computing resources. For sure. It's a significant bottleneck. However, they did explore a potential workaround.

13:47Oh. They tried using a proxy reward. Instead of doing the full fine-tune and evaluate cycle, they estimated the quality of a self-edit based on simpler heuristics like how clear is it? How detailed? Does it cover the key info? Things a human could quickly assess or maybe another model. And did that help with the speed? Dramatically. Using this proxy reward cut the evaluation time down from 30, 45 seconds to just around 5 minutes total for the RL part of a batch, apparently, while still getting strong results. So that shows there might be engineering pathways to make this more feasible. Okay, that's promising.

14:22So let's zoom out. Connect SEAL to the bigger picture. We keep hearing about this data wall coming for AI. Yeah, the projections are pretty stark. Some sources estimate that we'll basically run out of high-quality, publicly available human text to train the biggest LLMs on by. Maybe 2028. Could be sooner. We're exhausting the Internet's worth of text. Pretty much. And if that happens, if the firehose of human data dries up, then future AI progress might completely depend on models being able to generate their own useful training data. Synthetic data becomes essential, not just nice to have. And SEAL seems like a way to make that synthetic data generation much smarter.

14:58Yeah. Not just generating random text, but generating text specifically optimized for learning. Exactly. It avoids relying on fixed rules or heuristics for generating data. It lets the model learn what kind of synthetic data is actually useful for its own improvement. So if SEAL or something like it becomes standard, what does that future look like? It seems like it moves LLMs away from being static databases. It really does. They become continuous, autonomous learners. You could imagine a future LLM reading, say, a new scientific paper, generating its own explanations, figuring out the implications, maybe creating practice problems for itself.

15:36and then actually distilling those insights back into its core parameters all on its own. A cycle self-improvement. Yeah, this loop of self-expression, self-editing, self-refinement. It means models could potentially keep getting better, especially on niche or complex topics, even if there's no new human data coming in for that specific area. That really opens the door to AI systems that are much more agentic, dynamically acquiring and updating their knowledge as they operate in the world. Could a system like this even decide when and how much to update itself? Like make a tiny tweak mid-task versus doing a major knowledge consolidation afterwards.

16:14That's the logical extension, isn't it? Moving from systems that just respond to systems that actively manage their own learning, their own internal state. Adapting based on literally everything they experience. That is a powerful thought to end on. The future of AI might not just be about what data we feed it, but about how well it learns to feed itself. Well, that's all the time we have for this deep dive into the potentially less static, more self-adapting future of AI. Thanks for joining us. Thanks for having me. We'll see you next time on The Deep Dive.

From the publisher

This paper introduces Self-Adapting Large Language Models (SEAL), a novel framework that enables LLMs to autonomously improve by generating their own training data and finetuning instructions, termed "self-edits." This adaptation process is driven by a reinforcement learning (RL) loop that rewards the model for generating self-edits that subsequently improve its performance on downstream tasks, contrasting with static models that learn from data "as-is." The authors demonstrate SEAL's effectiveness in two key domains: knowledge incorporation, where it generates synthetic data to efficiently integrate new facts, and few-shot learning, where it autonomously configures optimal data augmentations and training hyperparameters. Although promising, the work notes limitations regarding computational overhead and susceptibility to catastrophic forgetting during continuous adaptation.


More from Best AI papers explained

All 475 episodes
Self-Adapting Language ModelsBest AI papers explained · 17 min
Listen in VO