Self-Improving Language Models with Bidirectional Evolutionary Search

1 Jun 2026 · 21 min · 11 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Bidirectional Evolutionary Search (BES) to help language models escape “entropy shell” failure modes in hard, creative, multi-step reasoning by combining evolutionary operators with dense goal feedback.

Guests/backgrounds

No named guests; episode is a technical discussion by the hosts summarizing research.

Key claims

Standard best-of-N and tree search rely on autoregressive next-step generation and sparse end-only verification, trapping models in high-probability loops. BES uses forward evolutionary search (combination, deletion, editing, translocation, crossover) plus backward goal decomposition into many verifiable sub-goals, giving dense intermediate scoring and exponentially fewer samples.

Notable examples

MUSIC multi-hop question; BES translocated “Custard Records” into the branch that already knew the artist “James Blunt,” achieving 1.0. Inference-time circle packing (26 circles in a unit square) where BES generated a hybrid Python optimizer and beat OpenEvolve/GPA/Schink of Evolve at $18.60 API cost.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding AI Failures

0:45 to 1:39

Exploring the limitations of current AI reasoning methods and their inefficiencies.

“our mission is to unpack a groundbreaking framework called Bidirectional Evolutionary Search.”

The Autoregressive Expansion Flaw

1:39 to 3:16

Discussing how autoregressive methods limit AI's ability to find creative solutions.

“Well, the dominant methods for AI problem solving right now are essentially variations of two strategies.”

Learning from Biological Evolution

3:16 to 4:52

Drawing parallels between AI logic and the principles of biological evolution.

“You can never edit a previous paragraph.”

Bidirectional Evolutionary Search Explained

4:52 to 6:28

Introducing the four operators that enhance AI reasoning through evolution.

“And this fundamentally changed the game through chromosomal recombination.”

Addressing the Risk of Garbage Output

6:28 to 8:00

Discussing the challenges of random mutations and the importance of filtering in AI.

“Okay, I have to jump in here and offer some pushback.”

The Role of Backward Search

8:00 to 9:56

Explaining how backward search improves AI capability by breaking down goals.

“It needs a way to know which of these newly bred ideas are actually good and should survive to the next generation.”

Case Study: Solving a Multi-Hop Reasoning Problem

9:56 to 12:50

Analyzing a specific scenario where BES successfully combines information to reach an answer.

“Exponentially is a heavy word in mathematics.”

Evaluating the Effectiveness of BES

12:50 to 14:00

Comparing the performance of BES versus traditional methods in complex problem solving.

“The AI literally took two incorrect answers, genetically bred them together, and birthed a single perfectly correct answer.”

Evolutionary Search in AI Optimization

14:00 to 18:02

Learn how bidirectional evolutionary search improves AI's problem-solving capabilities.

“Wow, it optimized for doing the least amount of work possible.”

Cost-Effectiveness of BES

18:02 to 18:28

Discover the surprisingly low cost of executing complex AI optimization tasks.

“You can barely buy a decent lunch in a major city for$18.”
Show all 11 chapters

Implications of Evolutionary AI

18:28 to 20:55

Understand the broader impact of evolutionary AI on the future of technology.

“If we connect this to the bigger picture, this represents a massive watershed moment for the industry.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00An AI can, you know, write complex software in seconds. It can pass the bar exam. Right, easily. But you give that exact same multibillion-dollar supercomputer a basic logic puzzle, and suddenly it's just trapped in its own head. Oh, yeah, stubbornly repeating the exact same obvious mistake over and over. It is incredibly frustrating for anyone trying to use these tools for serious problem solving. It's like watching a grandmaster chess player who keeps forgetting how the knight moves. Well, it really is the ultimate modern paradox. I mean, we have built these incredibly powerful reasoning engines, But they frequently hit this hard ceiling when the problem requires like a nonstandard or creative leap.

0:40Today we are going to fix that. So welcome to another custom tailored deep dive. Looking at the stack of research data and technical papers on the desk today, our mission is to unpack a groundbreaking framework called Bidirectional Evolutionary Search. Or BES for sure. BES. We are exploring how injecting the principles of biological evolution, you know, mimicking how DNA mutates and recombines, is finally helping AI break out of its own mental traps. And this involves a concept called goal decomposition, which basically acts as the compass for these biological mutations. Exactly. The core concepts here, the actual mechanics of how this works, they're surprisingly intuitive once you see the patterns.

1:21We don't really need to get bogged down in the deep calculus of it all to understand the paradigm shift happening underneath the hood. Okay, let's unpack this. To understand how this BES framework fixes AI reasoning, we first have to understand why the AI is failing at the frontier of its capabilities right now. So how does a standard AI currently try to solve a hard problem? Well, the dominant methods for AI problem solving right now are essentially variations of two strategies. You have best of N sampling and you have tree search. Best of N is the brute force approach. The AI generates a whole bunch of different answers to a prompt, say, 100 different attempts from scratch.

2:01Yeah. And then uses a completely separate verifying system to grade them and pick the best one. That seems incredibly wasteful. You were basically hoping that out of 100 blind throws, you know, one of the darts magically hits the bullseye. It is wasteful. And it really doesn't scale well for complex problems. Tree search is slightly more sophisticated. It breaks the reasoning down step by step. Okay, that makes sense. Yeah, the AI explores different branches of logic, scoring each individual step, and tries to follow the most promising path forward. But both of these methods share a critical flaw.

2:34Right, and the research data gets really revealing here. Both methods build their answers through what's called autoregressive expansion. Exactly. Autoregressive expansion. That sounds like a highly technical way of saying it's just predicting the next word. That is the perfect translation, honestly. It is predicting the next logical step based strictly on its own training and, well, the immediate context of the steps it just took. So it can't really plan ahead. No, it looks at what it just wrote, calculates the highest probability for the very next move, and generates it. The researchers have actually mathematically proven that this method physically confines the AI to a narrow entropy shell.

3:12A narrow entropy shell. I love that term. It's like trying to write a brilliant, complex, Pulitzer-winning novel, but you're only allowed to ever think about the very next word. You can never look back. You can never edit a previous paragraph. You can never rearrange the chapters or combine a theme from chapter one with a character from chapter five. You just march forward word by word. Inevegably, you just get stuck in a predictable, highly probable rut. The AI gets completely trapped inside its own probability distribution. Because when you are trying to solve a truly hard frontier level problem, the correct highly creative solution usually lives in a low probability region.

3:54It requires an unusual combination of thoughts. Exactly. Because the AI is strictly autoregressive, it just keeps circling the drain of high probability, average, predictable responses. It literally cannot reach the creative solution using a forward-only expansion method. So if stepping forward word by word traps the AI in this entropy shell, how does it ever escape? I mean, it can't just suddenly decide to start being a language model. Well, the researchers look to the history of biology for the answer. Think about the early history of life on Earth. For a very long time, life reproduced asexually.

4:26Single cells just split. They cloned themselves. Right. The biological limitation of asexual reproduction is that if one organism develops a beneficial mutation and a completely different organism somewhere else develops a different beneficial mutation, those two biological breakthroughs can never combine. Because they're completely isolated. Right. Every single offspring is just a direct autoregressive extension of its single parent. It's a purely linear progression, just like the pre-search we were talking about. Exactly. Then sexual reproduction evolved. And this fundamentally changed the game through chromosomal recombination.

5:03Suddenly, gene segments from completely different lineages could be spliced together to produce novel combinations that neither parent possessed. See, here's where it gets really interesting. This bidirectional evolutionary search framework literally applies this biological concept to AI logic. They call it forward search. Instead of just generating the next step in a vacuum, the AI maintains a diverse population of different reasoning paths and it starts mutating them. It uses four specific evolution operators to manipulate the reasoning. The first is combination. The AI takes two different reasoning paths that perhaps started the same way but diverged, and it concatenates their unique endings together into a single longer thought process.

5:45So gluing the end of idea A to the end of idea B to see if they make sense together. But it doesn't just add things, right? It can also subtract. Oh, absolutely. It mutates by shrinking. This is the deletion operator. The AI looks at a long, meandering chain of logic and just deletes a flawed or useless interior step, tightening the whole argument. Editing. Finally, an AI that actually edits its own thoughts instead of just generating more fluff. Right. It's huge. Yeah. The third is highly surgical. They call it translocation. The AI transplants a single specific step from one reasoning path and drops it right into the middle of a completely different path.

6:20And the final operator is crossover, where it splices the beginning of one idea directly onto the end of another, completely bypassing the middle. Okay, I have to jump in here and offer some pushback. If we are just letting the AI randomly splice, delete, and mutate its own thoughts like some kind of chaotic genetic blender, doesn't that just create a massive amount of unreadable garbage? It's a fair question. I mean, we already constantly complain about AI hallucinations and systems making up fake facts. isn't randomly crossing over different logic paths just weaponizing hallucination? That is the most logical critique to make of this system.

6:57And the reality is, yes, absolutely. Just like in nature, the vast majority of random genetic mutations are complete failures. Right. They produce absolute garbage. But evolution doesn't need every mutation to be good. It only needs a few to be vastly better. Ah, I see. Because the system is generating a massive, diverse pool of candidates, The sheer volume of attempts covers further failures. Exactly. The mathematical theorems in the data show that these four evolutionary operators are the key to physically escaping that entropy shell. They forcefully drag the AI out of its high probability comfort zone and into those low probability regions where the genius solutions hide.

7:39That makes total sense. Even if 90 % of the spliced thoughts are junk, the 10 % that survive make up completely novel ideas the AI would never have reached autoregressively. But wait, if 90 % of these logic mutations are complete garbage, how does the AI keep from just hallucinating itself into a corner? There has to be some kind of filter, right? Oh, absolutely. It needs a survival pressure. It needs a way to know which of these newly bred ideas are actually good and should survive to the next generation. This is where the bidirectional part of bidirectional evolutionary search comes in. Current AI verification, how an AI checks its own work, is what we call sparse.

8:16Sparse, meaning it doesn't happen often. Exactly. It only checks at the very end. It's a binary pass or fail at the conclusion of a long, complex reasoning chain. Imagine asking an AI to solve a complex math proof. It generates 20 sequential steps of logic, gives you an answer, and the verifier just says, wrong. That is incredibly unhelpful. It's like trying to assemble a massive 500-piece IKEA wardrobe, and the only instruction you have is a single picture of the finished product on the outside of the box. That is the perfect way to look at it. That's sparse verification. You blindly build the whole thing, step back, it wobbles and collapses into a pile of particle board, and you have absolutely no idea which of the 500 steps you messed up.

8:59BES completely flips this dynamic by introducing the backwards search. What this does is take the main overarching goal, and before the AI even starts trying to solve anything, the system recursively decomposes that primary goal into a branching tree of smaller, highly verifiable sub-goals. Oh, wow. So instead of just staring at a picture of the finished IKEA wardrobe, you now have a dense, highly detailed checklist. You get a little dopamine hit, a numerical score, every time you correctly attach a single hinge or properly align a single shelf. Yes, that dense intermediate feedback changes everything.

9:33When the AI is using its genetic blender to mutate its thoughts, it's not waiting until the end to see if a mutation worked. It is constantly scoring these partial thoughts against the sub-goals. Right. What's fascinating here is the data includes a theorem showing that this backward search exponentially reduces the number of samples the AI needs to generate to find the right answer. Exponentially is a heavy word in mathematics. How exactly does breaking a goal down create an exponential reduction in the effort required? Think of it like a multiplicative guessing game. If an AI needs to get five steps right in a row, and it has a 10 % chance of guessing each step correctly, the chance of getting the whole thing right all at once is 10 % times 10 % times 10%, it becomes microscopically small.

10:19You'd need millions of attempts. Right. By decomposing the problem into sub-goals, it only has to find the first step, Verify it against the dense feedback, lock it in, then find the second step and lock it in. It turns a massive multiplication problem into a simple addition problem. I want to see this actually happen in practice, because the theory of forward evolution and backward goal checking sounds great on paper, but I want to know what this looks like inside the mind of the AI. Let's look at a specific test case from the data. Sure. The researchers tested this on a multi-hop reasoning data set called MUSIC.

10:52This tests an AI's ability to pull from vastly different sources of scattered information to answer a single complex question. Okay, what was the question? The prompt was, what is the record label of the artist who originally recorded Back to Bedlam? Oh man, I actually have no idea. James Blunt saying You're Beautiful, which I think is on that album, right? But the original label, no clue. Neither does the AI initially. The answer is Custard Records. Here is how the AI has to figure it out using BES. The backward search immediately kicks in. It takes that big vague question and breaks it into two highly measurable sub-goals.

11:26Okay. Sub-goal one, who is the artist? Sub-goal two, what is their label? Now the AI has its checklist. Got it. And with the map drawn, the forward search starts mutating paths to try to fill in the blanks. Here is where we see the absolute magic of the translocation operator. The AI starts exploring multiple paths simultaneously. One reasoning branch correctly figures out that the artist is James Blunt. it nails sub-goal one. Perfect. But then, autoregressively, it guesses that his label is Atlantic Records. Atlantic distribute him later, but it's the wrong answer for the original label. So that branch gets stuck, it fails the final goal.

12:03Exactly. Meanwhile, another branch of the AI goes somewhat rogue, it completely misses the artist's name, but it does a general search about the album Back to Bedlam, and it actually uncovers the text Custard Records. But because it missed the first step, it gets confused by its own context and completely fails to formulate the correct answer to the prompt. Okay, so you have two branches. Both of them are wrong. Neither of them can solve the problem alone. Right. But they both get a partial score because of the dense feedback. The backward search says, hey, branch A, you got the artist right. Here's a partial score.

12:36Branch B, you found the right label. Here's a partial score. And then. Then the translocation operator steps in. It physically plucks the knowledge of custard records from the bad branch and it graphs it directly into the good branch that knows the artist is James Blunt. So what What does this all mean? The AI literally took two incorrect answers, genetically bred them together, and birthed a single perfectly correct answer. It scored a perfect 1.0. That is wild to visualize. It is a massive leap forward. What is crucial is how this affects the AI's behavior over time. The researchers tested this during a post-training phase on LAMA 3.23B and LAMA 3.18B models.

13:16Those are essentially mid-sized open-weight models that researchers use for experimentation, right? Exactly. They compared BES to standard baseline methods like GRPO. GRPO is basically the current industry standard for teaching an AI through trial and error. And how did the industry standard handle the complex reasoning compared to BES? Well, the standard baselines suffered from a phenomenon called reward hacking. Because their feedback was sparse, only getting graded at the very end, the AI basically learned to get lazy. Lazy? How so? It realized that actively searching for information across multiple steps was incredibly hard and prone to error, so it just gave up.

13:55It started blindly guessing the answers, hoping to get lucky. Its performance actually degraded over time. Wow, it optimized for doing the least amount of work possible. The BES-trained agents did the exact opposite. it. Because they were constantly getting rewarded for partial progress through the decomposed sub goals, they learned to actively search. Because they knew they were on the right track. Exactly. They achieved a dramatically higher number of valid searches, and their overall accuracy was substantially higher. They learned how to do the hard work because the feedback loop rewarded the effort step by step.

14:27Okay, so it can solve complex music trivia, and the data mentions it crushes traditional logical puzzles like knights and knaves. But let's push the front ear here. Trivia and riddles are one thing. Can this evolutionary framework help AI solve actual open scientific problems? The short answer is yes. And this is where the implications get very real for industries outside of computer science. The researchers tested BES at inference time, meaning the AI is actively thinking and solving a new problem on the fly, not just being trained in a lab. They used a GPT-5 backbone and aimed it at open mathematical benchmarks.

15:05Give me an example of an open math benchmark that tests this kind of frontier reasoning. A prominent one is called circle packing. The specific challenge is to mathematically pack 26 non-overlapping circles into a unit square in a way that maximizes the sum of their radii. That sounds like a deceptively simple geometry problem that is actually a complete nightmare to calculate. But why does packing 26 circles into a square matter to anyone outside of a math department? Because if an AI can figure out the flawless spatial reasoning to optimally pack 26 chaotic circles into a constrained space, it can figure out how to optimize global shipping routes for a fleet of cargo ships.

15:45Oh, I see. Right? It can figure out how to pack billions of microchips onto a single silicon wafer with lira waste. It can route hundreds of delivery drones over a city without them colliding. These are hard, open-ended optimization problems with massive economic value. Okay, so how exactly did our genetic mutant AI handle the 26 circles? Did it just sit there and guess coordinates for an hour? No, it evolved what is essentially a highly complex hybrid global optimizer program to solve it. It didn't just guess the math. It actually wrote a bespoke, sophisticated optimization program in Python.

16:21Wait, how does an AI use evolution to write a software program? The backward search first decomposed the massive problem into coding subgoals. Subgoal 1, establish the coordinate boundaries. Subgoal 2. Write a physics engine style collision detection function to make sure circles don't overlap. Okay, making the checklist. Then the forward search started mutating different Python scripts. One script had a great boundary checker, but a terrible optimization loop. Another script had a brilliant gradient descent algorithm, but failed the collision detection. Let me guess. The crossover operator stepped in.

16:53It spliced the perfect boundary checker code from script A directly into the brilliant optimization loop of script B. Exactly. They tested the resulting hybrid program, and BES beat every single existing open source framework attempting this today. It beat OpenEvolve, it beat GPA, and it beat Schink of Evolve, which are all previous highly funded attempts at automated code evolution. That's incredible. BES had the best average performance, the highest peak performance, and had much lower variance, meaning it was consistently stable every time it ran. But creating a bespoke hybrid global optimizer by running hundreds of evolutionary code mutations, checking them against sub-goals, and running Python scripts in the background, that has to be astronomically expensive.

17:37Compute costs are the biggest bottleneck in AI right now. Every time we ask an AI to think harder, the server bill skyrockets. You would naturally assume that, but this is the punchline of the entire dataset. Despite being incredibly complex mechanically, it is remarkably cheap to execute. The API cost to run the circle packing problem using BES and achieve a state-of-the-art result was exactly$18.60. $18! You can barely buy a decent lunch in a major city for$18. Compare that to the much weaker Schink Evolve framework, which costs$13 to run. For an extra$5, you are getting frontier-pushing mathematical performance that dominates the benchmarks.

18:18During the training phase, BES added less than 30 % wall clock overhead compared to traditional rigid tree search methods. That is staggeringly cost-effective. We are getting supercomputer-level optimization for the cost of a hamburger and fries. If we connect this to the bigger picture, this represents a massive watershed moment for the industry. The entire tech world right now is utterly obsessed with scale. Yeah, the bigger the brain, the smarter the AI. Exactly. But what this data proves is that giving an AI structural evolutionary search methods, teaching it how to think, how to break problems down, and how to self-correct, is far more powerful and vastly more economically viable than just building a bigger, more expensive brain.

19:01Work smarter, not harder. Even for supercomputers, it fundamentally shifts the entire strategy of artificial intelligence. It really is a paradigm shift from scaling the raw size of the model to scaling the quality of the reasoning at inference time. What a journey. We started with an AI that gets stuck in a rut, confined to a narrow entropy shell where it can only think one predictable word at a time, entirely unable to reach the brilliant ideas hiding in the margins. We've seen how researchers are breaking it out of that shell by mimicking biological reproduction splicing, deleting, and mutating its thoughts to force it out of its comfort zone.

19:38Combine that genetic blender with a backward-looking step-by-step checklist to grade its own mutations, and you get an AI that actively searches, reasons, and solves frontier math problems with massive real-world implications. All for a handful of dollars. Exactly. For you, the listener, the takeaway here is massive. The AI tools you use every day are about to get dramatically better at complex, multi-step reasoning, and they are going to do it without needing to become vastly larger, slower, or prohibitively expensive. It is the democratization of high-level reasoning. Once these frameworks become standard across the industry, the floor for what we consider basic AI capability is going to skyrocket.

20:18Which leaves me with one final, slightly provocative thought to end on. If AI can successfully self-improve by mimicking the biological mechanisms of DNA crossover and mutation, what happens when the AI gets smart enough to start inventing its own evolution operators? Oh, wow. Right. What happens when it starts manipulating its reasoning with completely alien, high-dimensional logic splices that have absolutely no equivalent in human biology or nature? What kind of logic will it breed then? It might just be the moment the AI truly stops thinking like a human and starts thinking entirely like itself.

20:50Now there's a thought to keep you up at night while you stare at that blinking cursor on your screen. Thank you for joining us on this deep dive. Until next time, keep exploring the frontier.

From the publisher

Researchers have developed Bidirectional Evolutionary Search (BES) to overcome the limitations of standard language model sampling, which often struggles with sparse feedback and predictable outputs. While traditional methods like tree search are confined to a narrow "entropy shell" of high-probability responses, BES escapes this range by using evolutionary operators such as crossover and translocation to recombine successful segments from different trajectories. Simultaneously, a backward search process decomposes complex goals into manageable sub-goals, providing the dense feedback necessary to guide the forward search. Theoretical analysis demonstrates that this dual approach can exponentially reduce the number of samples required to solve difficult reasoning problems. Experimental results confirm that BES significantly improves performance in both model training and real-time inference across logical, mathematical, and agentic tasks. By integrating genetic algorithms with goal decomposition, the framework enables models to discover novel, high-quality solutions that standard autoregressive generation would likely miss.

More from Best AI papers explained

All 475 episodes
Self-Improving Language Models with Bidirectional Evolutionary SearchBest AI papers explained · 21 min
Listen in VO