GEPA: Generative Feedback for AI System Optimization

29 Jul 2025 · 15 min · 7 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

GEPA (Genetic Pareto) is a prompt optimizer for LLM system optimization that targets sample inefficiency in reinforcement-learning-style prompt tuning (e.g., GRPO/GRPO-like RLVR), which can require tens to hundreds of thousands of rollouts.

Key claims

GEPA uses reflective prompt evolution—natural-language reflection on full “trajectories” (reasoning steps, tool calls/outputs, and external diagnostic feedback like compiler errors)—to iteratively mutate prompts. It selects candidates via a Pareto frontier (best strategy per instance) to avoid local optima.

Notable examples

IFBench instruction-following—optimal with 678 rollouts vs GRPO’s 24,000; 79 training rollouts to match GRPO best. Compared to MIP-AV2, GEPA improves up to 11.1% (GPT-4.1 Mini) and 10.3% (Qwen38B) and produces prompts up to 9.2x shorter. It also optimizes code at inference time (kernel writing, AMD XDNA2 via NPEvil, NVIDIA V100 CUDA via KernelBench), reaching up to 70% vector utilization using compiler-error feedback.

Guests

none mentioned.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Challenges of Traditional Optimization

0:45 to 2:14

Discussing the inefficiencies and costs of traditional methods like reinforcement learning.

“They need just an enormous number of rollouts.”

Introducing GEPA: A New Approach

2:14 to 4:24

Exploring GEPA, a prompt optimizer that improves efficiency in AI learning.

“That includes algorithms like group relative policy optimization or GRPO.”

Core Principles of GEPA

4:24 to 8:01

Delving into GEPA's principles including genetic prompt evolution and natural language reflection.

“GPA iteratively proposes better and better candidate prompts.”

Performance Comparison with Traditional Methods

8:01 to 11:10

Comparing GEPA's performance with GRPO and MIP-AV2 in terms of efficiency and effectiveness.

“If you look only at the rollouts specifically used for training, the ones where it's actually learning, not just validated, it's even more dramatic.”

Real-time Problem Solving with GEPA

11:10 to 12:41

How GEPA adapts for real-time programming tasks and optimization.

“But the core GPA itself is already delivering huge gains.”

Implications for the Future of AI

12:41 to 14:00

Discussing the broader impacts of GEPA on AI learning and optimization workflows.

“It's like having an AI that can learn from expert coaching and compiler feedback in real time to write better code.”

Exploring Future AI Learning Approaches

14:00 to 15:07

Discover how GPO's linguistic insights could merge with weight-based learning.

“and this really cool application in generating highly optimized code.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Welcome to the Deep Dive, your shortcut to being truly well informed. Today we're jumping into something pretty exciting, the really fast-moving world of large language models, you know, LLMs. And specifically, we're tackling a big challenge. How do we make these incredibly powerful AI systems, well, smarter and more efficient without, you know, breaking the bank or taking forever? Because traditionally, getting these LLMs optimized, it's relied a lot on methods like reinforcement learning. And they work, definitely. But they often come with this huge cost, time, computation resources that, let's face it, aren't infinite for most of us.

0:39Okay, let's dig into this. Right. And what's really fascinating, and honestly, a major bottleneck is how these traditional methods learn. They need just an enormous number of rollouts. Think of a rollout as like one attempt, one trial and error cycle for the LLM to try something and get feedback. Okay. And we're talking tens of thousands of these attempts, sometimes even into the hundreds of thousands, just to learn a new task. Wow. That's a little... It is. Imagine if every time an LLM needed to learn something new, it had to, I don't know, run a marathon just to figure out how to walk properly.

1:08Yeah. It's that kind of scale. And this sample inefficiency, it becomes a real problem for you, the user, pretty quickly. Because lots of LLM applications, they involve calling expensive external tools, right? Right. Like APIs or databases. Exactly. Or maybe you just have a limited inference budget. Yeah. You literally can't afford to call the LLM thousands and thousands of times for every little learning step. Or sometimes you just can't fine-tune the weights of the biggest, best models. They're too large or maybe they're proprietary. So that constant costly grind is a huge hurdle. Yeah, that really paints a picture of the problem.

1:41Effective maybe, but just brutally inefficient for so many situations you actually encounter. But what if there was another way? A way to kind of bypass this bottleneck using language itself. So our mission for this deep dive is to explore something called GPA. That's G-E-P-A, genetic Pareto. It's a new prompt optimizer. And we're going to see how it's really challenging the status quo, offering, well, a faster, maybe smarter way to get AI systems performing at their best, often with some pretty surprising efficiency gains. Okay, so let's dive deeper into that costly grind we mentioned. A really popular way to adapt LLMs now is reinforcement learning with verifiable rewards, RLVR.

2:19That includes algorithms like group relative policy optimization or GRPO. That's right. And these methods, they are, like you said, quite expensive. The core idea is they boil everything down to just one number at the end of a rollout, a scalar reward, success or failure. Just a thumbs up or thumbs down, basically. Kind of, yeah. A score. And then they try to work backwards from just that single score to figure out, OK, how do we improve the LLM strategy, its policy for next time? But the crucial detail, the bottleneck we talked about, is that this usually needs tens of thousands of rollouts. For full training, it can easily be hundreds of thousands.

2:55So, for instance, the research we're looking at today, they benchmarked GRPO using 24 ,000 rollouts for their experiments. But other papers using GRPO, they often mention needing up to hundreds of thousands of rollouts. That's the core inefficiency GPA is designed to tackle. It's like running that marathon just to tweak your walking stride a tiny bit. Okay, so that's the old way. Powerful in some respects, but comes with this massive overhead. Now, let's bring in the game changer, GEPA, genetic Pareto. You said the real ingenuity here is how it uses language itself as a richer learning medium, not just that single score.

3:29Exactly. That's the heart of it. GPA uses what's called reflective prompt evolution. So instead of just that sparse final number, GPA actually looks at the whole journey. It samples these system level trajectories. You know, the LLM's internal reasoning steps, the calls it makes to external tools, the outputs it gets back from those tools. The whole chain of thought and action. Precisely. And then it literally reflects on them in natural language. Yeah. It tries to diagnose problems, figure out what went wrong or right, and then proposes and tests, updates to the prompt. It even does this cool thing where it combines complementary lessons from the Pareto frontier of its own attempts, which is, well, a pretty sophisticated way to learn from multiple good ideas at once.

4:12Okay. That sounds really different. How does it actually do that? You mentioned three principles. Yeah. There are three core ideas driving GEP. First is genetic prompt evolution. Sounds fancy, but the idea is kind of intuitive. GPA iteratively proposes better and better candidate prompts. How? By modifying existing ones, maybe tweaking a word, combining parts of two successful prompts. They call it mutation or crockover, borrowing from biology. Like evolution for prompts. Exactly. And it's constantly informed by new rollouts, the results of its latest attempts. So it accumulates lessons along the genetic tree, building on what worked before.

4:50Okay, that makes sense. What's the second principle? Second is natural language reflection. This is, I think, the really key innovation. TPA uses the natural language traces that get generated when a system runs. And you mentioned compound AI systems earlier. What are those exactly? Right. Good question. Basically, think of modular systems. They're made up of different parts. LLM calls, maybe some external tools like a calculator or a search engine, and some logic controlling the flow. Frameworks like React or DSPY often build these kinds of systems. Got it. So GPA looks inside those systems. Yes.

5:21It looks at the instructions given to the LLM, its chain of thought reasoning, the tool calls, the tool outputs, everything. And crucially, it can also use feedback from outside the system, like evaluation traces. Think about, say, compiler error messages if the LLM is writing code. That's highly diagnostic information. It tells GPA exactly what went wrong. Ah, okay. So it's getting much richer feedback than just good or bad. Way richer. This detailed language-based feedback allows for really large and effective updates to how the whole system behaves. It's not just tiny nudges. Okay, and the third principle, Pareto.

5:58Pareto-based candidate selection. Yeah. Yes, this is super smart. Contrast it with the naive strategy, right, which is just pick the single prompt that got the highest score overall. Which sounds reasonable. It does, but it often gets you stuck in a local optimum. You find a good solution, maybe quickly, but you miss out on potentially much better ones because you stopped exploring too soon. Right. You optimize for one thing and miss the bigger picture. Exactly. So GPA does something different. For every single training example, every individual task is trying to solve. It identifies the highest score achieved across all the candidate prompts it tried.

6:33This collection of best case strategies for each specific instance forms the Pareto frontier. Think of it like building a dream team. You don't just pick the single best player overall, you pick the best player for each position, each specific need. Okay, so it's keeping track of diverse strengths. Precisely. It filters the pool down to these varied winning strategies. This helps GPA escape local optima, find those truly great solutions, without expanding the search excessively. It's a clever way to balance exploring new ideas with exploiting what's already working well. That's a really elegant set of ideas, genetic evolution, reflecting in language, and this Pareto selection.

7:11Okay, so you've got these sophisticated learning mechanisms. The big question then is, does it actually work? What kind of performance leap are we talking about? What are these unbeatable results we teased? Well, the results are, frankly, pretty striking. Let's look at the head-to-head comparisons. First, GPA versus GRPO, that reinforcement learning method we discussed. GPA is just highly sample efficient. On average, it outperforms GRPO by 10%, and in some cases by up to 20%. But here's the kicker. It does this while using up to 35x fewer rollouts. 35 times fewer? That's incredible. Really, it's to give you a concrete example.

7:48On a benchmark called IFBench, which tests instruction following, GPA hit the optimal performance level using only 678 rollouts. Okay. Compare that to GRPO, which needed the full 24 ,000 rollouts they allocated for the experiment. So yeah, 35 times fewer. If you look only at the rollouts specifically used for training, the ones where it's actually learning, not just validated, it's even more dramatic. For that same IF bench task, GPA needed just 79 training rollouts to match GRPO's best scores. 79 versus potentially tens of thousands. Exactly. So what does this mean for you, the user? Massive savings.

8:21Less compute time. Lower costs. Faster optimization. It really changes the economics of tuning these models. Okay, that's GRPO. What about other state-of-the-art methods? Right, so they also compared GPIA to MIP-AV2. This was considered the previous state-of-the-art prompt optimizer. A key difference is that MIP-AV2 tried to optimize both the instructions and the few shot examples you give the LLM. Okay, tuning two things at once. Yeah, but GPIA, which focuses only on the instructions using its reflective process, consistently outperforms MIP-AV2 in all settings. The performance margins were significant, too, as high as 11.1 % for GPT-4.1 Mini and 10.3 % for QEN38B.

9:03Across the board, GPIO more than doubles the aggregate gains over baseline seen with MIPF2. Wow. So optimizing just the instructions but doing it smarter with reflection was actually better. It seems so, yes. It suggests that GPIO's way of learning from language traces is really powerful, maybe even more powerful than trying to tune both instructions and examples simultaneously without that deep reflection. And there's another benefit, almost unexpected but really valuable. The prompts GPA generates are shorter, much shorter. Shorter prompts, why does that matter? Well, think about cost and speed again.

9:35GPA's instructions, Evolve Through Reflection, turned out to be up to 9.2x shorter than the prompts generated by MIPOv2. Shorter prompts mean they're computationally cheaper. For anyone using LLMs via APIs, you know you're usually charged per token, including the input tokens in your prompt. Right. So shorter prompts directly reduce your costs. Exactly. But it also decreases latency. The model responds faster and just generally improves the overall efficiency of LLM serving systems. It's a win-win-win. Better performance, lower cost, faster response. That's amazing. Better results, and it's cheaper and faster to run.

10:10You can actually see this qualitatively too, right? There's a figure in the paper. Yeah, figure two. It's a great illustration. It shows how GPA started with a really simple kind of generic seed prompt for a multi-hop question answering task. And through its evolutionary reflation process, it transformed that into this highly detailed, super effective prompt. It doesn't just tell the LLM to answer the question. It guides its reasoning. It specifically instructs the LLM to think about what information is missing and to target missing but logically linked documents in its search rather than just, say, paraphrasing what it already knows or restating the question.

10:48It teaches it a much more sophisticated search strategy. That's a concrete example of how that reflection leads to better instructions. Precisely. And just as a quick note, they did experiment with a variant called GPA plus merge, which tries to combine different successful strategies later in the process. It showed potential for large gains, maybe up to 5 % more improvement. Interesting. Yeah, though they mentioned that figuring out the best time to apply this kind of crossover strategy is something for future research. But the core GPA itself is already delivering huge gains. This really sounds like it's more than just tweaking prompts better.

11:22It feels like it opens up new possibilities. You mentioned using it for inference time search, especially for code. Yes, this is a really exciting direction. The idea is you can use GPA not just for general prompt tuning, but to essentially overfit to a specific set of problems you need solved right now at inference time. So GBA keeps iteratively proposing better solutions to every problem in the set until it finds really good ones. So like real-time problem-solving improvement. Kind of, yeah. The paper gives concrete examples. Using GPA for writing kernels, that's low-level code, for AMD's recently introduced XDNA2 architecture using a benchmark called NPEvil, and also for generating CUD code for NVIDIA V100 GPUs using another benchmark, KernelBench.

12:06These are complex, performance-critical coding tasks. And how does that work? How does GPA optimize code? Well, this is where that feedback engineering comes in strong. You can inject domain expertise. For kernel writing, you can provide feedback based on kernel development expertise, or crucially, feed it compiler error messages. Ah, like we discussed with reflection. Direct diagnostic feedback. Exactly. That feedback gets dynamically injected into GPA's optimization loop. And the results were pretty impressive. For those NPU kernels, they achieved up to 70 % vector utilization, which is a measure of code efficiency, and significantly higher than other methods managed.

12:46It's like having an AI that can learn from expert coaching and compiler feedback in real time to write better code. So connecting all these dots, GPA isn't just a tool to make existing LM workflows slightly better. It seems to be about enabling them to learn much more efficiently, adapt faster, maybe tackle problems they couldn't before, especially when, you know, money or data is tight. I think that's right. It feels like a genuine step towards more, I don't know, human-like learning for these systems. more adaptive, more efficient, leveraging language in a deeper way. Okay, so let's try and wrap this up.

13:18Today we've taken a deep dive into GPS, this novel prompt optimizer that's really shaking things up for large language models. Its power seems to come from this unique combo, using natural language reflection to understand mistakes, genetic evolution to improve prompts, and that smart Pareto selection to balance exploration and exploitation. Right, and the outcome is pretty clear. Unprecedented sample efficiency. Remember that 35x figure and significant performance gains over traditional RL methods like GRPO and even over prior state-of-the-art prompt optimizers like MIP-Rof2, plus the bonus of shorter, cheaper prompts.

13:56And the practical value seems huge. We talked about improving complex reasoning, instruction following, even things like privacy-aware delegation, and this really cool application in generating highly optimized code. Yeah, it offers a genuinely practical path forward for optimizing complex AI workflows. especially as you said in those common real-world scenarios where data or budget or time are constrained. So reflecting on all this, what does it mean for the future of AI? It feels like we're pushing boundaries here. Perhaps the most intriguing question this deep dive leaves us with is about that fuzzy boundary between prompt-based and weight-based learning.

14:31If GPO can achieve so much just by learning lessons in language through prompts. That's a great point. Could we see a future where these methods start to merge? where GPO's way of generating linguistic insights is used to, say, guide the rollouts in traditional reinforcement learning or somehow inform weight space adaptation more directly. That's a fascinating thought. Could you unify these approaches, using language-based reflection to make weight-based learning more efficient, or vice versa? Exactly. It really makes you wonder about how AI learning itself is evolving, moving beyond just brute force computation towards something maybe a bit more reflective.

15:07It certainly leaves us with a lot to think about.

From the publisher

This paper introduces GEPA (Genetic-Pareto), a novel prompt optimizer for large language models (LLMs) that significantly outperforms traditional reinforcement learning (RL) methods like GRPO and other prompt optimizers such as MIPROv2. GEPA achieves this by leveraging natural language reflection from system-level trajectories and a Pareto-based multi-objective evolutionary search, allowing it to learn from significantly fewer "rollouts" or trials. The research demonstrates GEPA's superior sample efficiency and robust generalization across various tasks, including question answering and fact extraction, while also producing shorter, more computationally efficient prompts. This approach offers a practical solution for optimizing complex AI systems in data or budget-constrained environments.

More from Best AI papers explained

All 475 episodes
GEPA: Generative Feedback for AI System OptimizationBest AI papers explained · 15 min
Listen in VO