In short
EGROL (Evolution Guided General Optimization via Low-Rank Learning) removes the memory/compute wall that previously prevented evolution strategies (ES) from scaling to billion-parameter models, using low-rank perturbations and enabling integer-only, hardware-efficient training.
Guests
No guest names or backgrounds are provided in the transcript.
Key claims
ES is gradient-free and robust to noisy/binary rewards; naive ES is too costly because it needs full-rank perturbation matrices per population member; EGROL replaces full-rank perturbations with low-rank A·Bᵀ, yielding ~100x throughput and fast convergence (~1/R).
Notable examples
Jumanji Snake (40.68x faster than OpenES), RWKV7 fine-tuning on countdown/GSM8K (1.5B: 35% vs GRPO 23%; 124 parallel generations per GPU vs 32), and EGG (Evolved Generative GRU) integer-only NT8 weights/NT32 activations, MinGRU recurrence, and no activation functions (nonlinearity via saturated addition) trained with population size 262,144.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding Evolution Strategies
0:45 to 2:26
Exploring the benefits and mechanics of evolution strategies in AI.
“that stopped evolution strategies from ever touching billion parameter models.”
Challenges of Naive Evolution Strategies
2:26 to 3:53
Discussing the limitations of traditional evolution strategies in scaling.
“You ship out parameters, get a single number back, and aggregate.”
The Breakthrough of EGROL
3:53 to 5:48
Introduction of EGROL and how it innovates low-rank learning for optimization.
“And this is precisely why eGrawl is such a breakthrough.”
Performance Gains of EGROL
5:48 to 7:42
Analyzing the speed improvements and efficiency of EGROL in reinforcement learning.
“Eggroll achieves a hundredfold increase in training throughput for billion-parameter models at scale.”
EGROL in Real-World Tasks
7:42 to 9:36
Examining EGROL's performance on various tasks against traditional methods.
“Is OpenES just that slow, or is EGRL really that good?”
Innovative AI Architecture with EGROL
9:36 to 12:16
How EGROL enables unconventional AI architectures through pure integer training.
“They were proving that ES can design and train a model based purely on hardware efficiency, totally ignoring the limits of traditional gradient methods.”
The Future of Optimization in AI
12:16 to 14:01
The implications of EGROL for future AI architectures and hybrid systems.
“That number, that population size, is literally two orders of magnitude beyond what was possible in previous ES work.”
Exploring Non-Differentiable Symbolic Modules
14:01 to 14:18
Learn how Evolution Strategies can enhance training for hybrid AI architectures.
“So how might an ES algorithm like eCrow enable the training of these massive hybrid AI architectures, where the human design part and the learned part are finally optimized together seamlessly?”
Transcript
Automatic transcript. May contain errors.0:28Welcome back to the Deep Dive. I have to say a delightful name, Evolution Guided General Optimization via Low-Rank Learning or EGROL. EGROL, yeah, it's a great name. So our mission today is to really understand how EGROL doesn't just, you know, improve on an old method. It seems to completely demolish the scaling wall that stopped evolution strategies from ever touching billion parameter models. We're going to explore how it does that and then get into its most, frankly, astonishing feat, training a functional large language model using nothing but pure integer math. It is a genuine paradigm shift.
1:01And to start, evolution strategies, or ES, are a class of what we call black box optimization methods. They're inspired by natural selection. You know, you have population of solutions, you test their fitness, you iterate. The key benefit, the thing that makes them so attractive right now, is that they're gradient free. And being gradient free, that's everything, that lack of reliance on calculus. Exactly. It means you don't need differentiability. And that just opens the door to optimizing systems where gradients are messy or even mathematically impossible. Okay, so give us an example. Well, think about fine-tuning an LLM where your only reward is a binary outcome.
1:39Did it generate the correct answer or not? ES can handle that noisy outcome-only reward perfectly. So ES is the solution when you only care about the final result, the destination, and you don't need that perfectly smooth, differentiable roadmap that backpropagation absolutely demands. That's a great way to put it. It's incredibly liberating. And beyond that flexibility, ES is just inherently more robust. Or stable, you mean. Yeah, because you're sampling fitness across a whole population, that exploratory step naturally smooths out noisy or ill-conditioned optimization landscapes. The search is just far more stable.
2:14And critically, it is built for the era of hyper-parallelization. Every member of your population evaluates its fitness on its own. So you just send out the work and wait for the scores to come back. Right. You ship out parameters, get a single number back, and aggregate. This minimal communication means you get near-linear speedups on massive clusters. Something standard gradient descent really struggles with because it has to sync these enormous gradient tensors all the time. Okay, but hold on. Before we get to the solution, let's really unpack the problem. If ES is so flexible and so parallelizable, why isn't everyone already using it for these massive models?
2:51It all comes down to cost. In the multi-billion parameter world we live in now, you're dealing with matrices that have billions of weights. Let's just call that big matrix WN. A naive evolution strategy requires you to generate a unique, full-rank matrix perturbation. We'll call that E for every single member of your population. And just to be clear, when you say full rank perturbation, you mean E is basically a random matrix, the exact same size as the entire model. So if your model is 10 billion parameters. Your perturbation matrix is also 10 billion random values. Wow. And if your population is, say, 1 ,000, you suddenly need to store 1 ,000 copies of that 10 billion parameter matrix just for the exploration step.
3:32That sounds completely impossible. It leads to two paralyzing costs. Yeah. One is massive memory replication, just constantly moving and storing these huge tensors. And two is the computational overhead. Every forward pass involves matrix multiplications that scale with your population size. It just made naive ES completely off the table in the billion parameter regime. And this is precisely why eGrawl is such a breakthrough. It comes in with this extremely elegant solution from low-rank learning. The insight is it's actually pretty simple. Instead of sampling that enormous full-rank perturbation E.
4:0610 billion parameter one. Right. E-Growl generates the perturbation using two much, much smaller matrices, A and B, so that the perturbation E is just their product. E equals A times B transpose. Okay, so you've taken one giant object, E, and replaced it with two tiny objects, A and B. It sounds simple conceptually, but what does that do in practice? Well, think of it this way. Creating the full rank matrix E is like trying to map an entire continent by surveying every single square inch. Low rank learning is like capturing the essential features, the mountains, the big valleys, with just a quick sketch.
4:43A sketch defined by A and B. Exactly. And that sketch, where the rank R is tiny compared to the matrix dimensions, is enough to guide the search effectively. You know, if this sounds familiar to anyone listening, it should. It's conceptually very similar to LoRa Low Rank adaptation, which is used all the time in gradient-based fine-tuning. It is, yeah. They've just cleverly adapted that concept to the random non-differentiable exploration phase of evolution strategies. And when you look at the numbers, the gains are just staggering. The math is complex, but let's just translate the practical savings.
5:16Your auxiliary storage per layer, so the extra memory you need for the perturbation, used to be proportional to the total number of weights. Now it's proportional to the rank r times the sum of the dimensions. That decoupling is the whole game. The complexity used to be driven by the total number of weights, Now it's driven almost entirely by that tiny rank car. So it's not just a theoretical win. Not at all. This is what lets them fit a billion-parameter model on hardware that could previously only handle a fraction of that. The memory footprint just gets slashed. The paper puts a number on it.
5:49Eggroll achieves a hundredfold increase in training throughput for billion-parameter models at scale. A hundredfold. It's revolutionary. It brings the speed of this black box method up to nearly pure batch inference speeds. I think they hit 91 % normalized speed. And what's really fascinating here is the theoretical confidence they have that this isn't, you know, sacrificing performance. You might assume that by using such a tiny rank R, you're throwing away all the good stuff. That was my first thought. But their analysis proves that the low rank update converges very quickly to the full rank version at a rate of order 1 over R.
6:25Wait, so does that mean you could use a rank of 1, just R equals 1? Does that mean you're throwing away almost all the complexity, all the rich paths that a full-rank matrix would offer? Why doesn't that cripple the performance? That's the magic of that fast convergence rate. It suggests that even in these extremely low-rank regimes, the main direction of the gradient or the estimate of the gradient in ES is still well captured. The optimization landscape, especially near the good spots, often has a much lower effective dimensionality than the parameter count would suggest. Oh, I see. The low-rank perturbation captures the most important directions of movement, so you get the benefits of the full strategy without the insane computational cost.
7:05Speaking of insane speed games, let's look at how this actually played out in the real world, starting with standard reinforcement learning tasks. So the research pitted EGROL with its efficient low-rank updates against OpenES, which is the traditional full-rank implementation. The brute force method, yeah. And the performance was competitive, which is the first big checkmark. EGRL matched or beat OpenES in 14 out of 16 environments. But the real story is the speed. A 40.68 times faster training time on the Jumanji Snake task. I mean, that is absolutely mind-boggling for a black box method that's supposed to be heavy.
7:42Is OpenES just that slow, or is EGRL really that good? It's a bit of both. OpenES is fundamentally bottlenecked by those memory costs we talked about. EGRL just solves that bottleneck. When you can scale up your population size without immediately hitting a memory wall, you naturally speed up exploration. You get 40 times faster because you can run 40 times as many optimization steps in the same amount of clock time. And that speed advantage carries over into LLMs too. Directly. They applied EDROL to fine-tune recurrent language models. Okay. Specifically, a version of the RWKV7 architecture on really complex reasoning tasks, things like countdown and GSM 8K.
8:19And how did it do against, you know, established gradient-based methods? Very, very well. On the countdown task, with a 1.5 billion parameter model, EGROL converged to 35 % validation accuracy. And for comparison. It significantly surpassed a competing state-of-the-art method called GRPO Group Relative Advantage Optimization, which only hit 23%. And GRPO is a sophisticated gradient-based algorithm designed for these tasks, which makes the result even more impressive. Wow. And if we connect this back to the bigger picture, the reason for this superior performance, it comes back to scale. The population size again.
8:53Exactly. During fine-tuning, eGroll allowed for 124 parallel generations per GPU. GRPO, because it has to deal with these huge gradient tensors, only allowed 32. Being able to scale your population by two orders of magnitude is fundamentally what ensures robust exploration and stops the optimizer from getting stuck on these really complex LLM tasks. Okay, this brings us to what I think is the most radical application, where EGROL wasn't just speeding things up, but was actually enabling entirely new, unconventional AI architectures. Yes. Here's where it gets really interesting for me. The researchers used EGROL to pre-train this highly specialized architecture they called EGG, the Evolved Generative GRU.
9:34This was a proof of concept. They were proving that ES can design and train a model based purely on hardware efficiency, totally ignoring the limits of traditional gradient methods. And the EGG model had three really radical design constraints that standard backprop would just, it would struggle or flat out fail to optimize. Yeah. And every single choice was made for pure performance and energy efficiency. Okay. What was the first constraint? First one, pure integer training. So normally training needs 32-bit floating point numbers to calculate gradients. The EGG model keeps all its weights as NT8 and its activations as NT32.
10:09It never, ever casts a floating point. And why is that so critical for performance? Because innate matrix multiplication is massively faster and way more energy efficient than floating point math on modern chips like the H100s. But if your weights are integers, you can't calculate a smooth derivative for backprop. It just breaks the math. But EGRL doesn't care. EGRL is black box. It doesn't care. It sends integer weights, gets a fitness score back, and updates. That immediately cuts off the traditional path. What was the second constraint? Second, they use a nonlinear recurrent neural network, a MinGRU variant.
10:45RNNs are great for tracking state and sequences, but for backprop, they're famous for stability problems, vanishing, or explosion gradients. Right. Since ES estimates the gradient from population fitness instead of the chain rule, it's just inherently robust to those internal stability nightmares. So if your architecture is prone to chaos with backprop, ES might be your only viable path. And the third constraint, this one was the most unusual. The third was the removal of all traditional activation functions. No real U, no sigmoid, no tan. They just designed the model without them. So where does the non-linearity come from?
11:17It comes implicitly from something called saturated addition. It's basically the clipping that's inherent to the NT data type. When you add two NT8 numbers and the result is too big, the system just clips it at the maximum value. That clipping effect, which is a sharp, non-differentiable break, is what gives the network its complexity. It's using the physics of the hardware data format as its activation function. Exactly. It's a non-differentiable effect that only a black box optimizer like IGJ could learn to exploit. This is the whole thesis. If you design the most hardware-efficient model possible, you often end up with non-differentiable parts.
11:55And IGEL gives you a path to actually train it. And the stability of training IGJ also highlighted how necessary this whole approach is. The researchers found that maximizing the number of unique perturbations, the population size, was absolutely critical. So small populations didn't work. They were highly unstable. Zero order methods, which are like a population size, one just fell apart. And by using EGRL to solve that memory bottleneck, they could stabilize this integer training by scaling the population up to, what was it, an astronomical 262 ,144. That number, that population size, is literally two orders of magnitude beyond what was possible in previous ES work.
12:35I mean, just think about the exploration space they unlock. Yeah. It proves that for this kind of de novo pre-training of radical architectures, you need that hyperscale population to navigate the vast non-differentiable search space. So what does this all mean for you watching the cutting edge of AI? EGRL has, it seems, fundamentally solved the scaling challenge for black box optimization. It has. It's made evolution strategies not just viable, but efficient for billion-parameter models by using these elegant low-rank perturbations. It means we now have a tool that is no longer beholden to the strict constraints of continuous calculus.
13:10We're free to design and train these highly specialized energy-efficient models, like the integer-only EGG model, that gradient methods simply could not touch. We're no longer limited to designing architectures based on how smooth the loss surface is. We can prioritize speed, energy, performance. The optimization tail is no longer wagging the architecture dog. Exactly. And that leads us to a really profound final thought here. The true lasting significance of eGroll is the permission it gives us to connect differentiable neural components with non-differentiable symbolic components. A bridge between two worlds.
13:45A bridge, yes. Think about the next generation of AI, large-scale, end-to-end mirror symbolic systems. These systems might need to interface directly with symbolic modules for things like high-precision calculations or talking to external memory systems. Those symbolic modules are almost always non-differentiable. So how might an ES algorithm like eCrow enable the training of these massive hybrid AI architectures, where the human design part and the learned part are finally optimized together seamlessly? When training stops needing a derivative, the design space, it truly explodes.
From the publisher
This paper introduces Evolution Guided General Optimization via Low-rank Learning (EGGROLL), a novel algorithm that enhances the scalability of **Evolution Strategies (ES)** for optimizing neural networks with billions of parameters. ES is an optimization method that bypasses the need for gradient backpropagation, offering advantages like handling non-differentiable objectives and superior parallelization potential. EGGROLL overcomes the memory and computational bottlenecks of traditional ES by substituting expensive full-rank parameter perturbations with **efficient low-rank matrix perturbations**, significantly increasing training throughput, as demonstrated by a **hundredfold speedup** for large models. The architecture is tested across various domains, including reinforcement learning, large language model (LLM) fine-tuning for reasoning tasks, and the stable pretraining of a custom **purely integer-based recurrent language model (EGG)**, showcasing the method's efficiency and flexibility. A **theoretical analysis** also confirms that the low-rank update rapidly converges to the full-rank ES update.




