In short
EVOLM (evolutionary language model) trains a language model to generate its own grading rubrics via co-evolution, avoiding human-written or opaque “scalar reward” grading.
Guest backgrounds
No guest bios or names are provided in the transcript.
Key claims
Static/opaque reward models cause “reward over-optimization” because students learn to game hidden scoring criteria. EVOLM instead has the same model alternate between (1) policy/student answering prompts and (2) rubric/teacher writing explicit natural-language criteria. A small frozen 1.7B “judge” scores answers strictly from the rubric; rubric training uses “discriminative utility” with preference pairs created by “temporal contrast” (current answer vs earlier self).
Notable examples
Perimeter-48 rectangle problem leads rubrics to precompute the maximum area 144; emotional-support rubrics become concrete checklists. Format reward forces JSON rubrics (without it, valid structure drops to ~23%). On HealthBench and ResearchQA, EVOLM rubrics align with expert rubrics more than GPT-4.1 (e.g., 58.4% vs 52.5% on HealthBench).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe Limitations of Current AI Training
0:51 to 2:25
Discussion on how human grading limits AI's potential and introduces EVOL.
“Today we're exploring a totally mind-bending breakthrough in how AI actually learns.”
The Problem with Scalar Reward Models
2:25 to 4:50
Exploration of the inefficiencies in current AI grading systems and their failures.
“Yeah, let's unpack that because the industry has tried to build automated graders to get around the human bottleneck.”
The Concept of Reward Over Optimization
4:50 to 6:06
Examining how AI gaming the system leads to poor learning outcomes.
“You're assuming the student AI is actually trying to learn the material.”
Introducing EVOL: The Co-Evolution Loop
6:06 to 7:20
Description of the new EVOL model that evolves without human input.
“And that reward over optimization is exactly why a static grader is doomed to fail.”
The Role of the Tiny Judge in EVOL
7:20 to 8:48
How a smaller judge improves the efficiency of AI learning and grading.
“But if the student and the teacher are essentially the exact same entity, just taking turns, wouldn't they just collude?”
Understanding Discriminative Utility and Temporal Contrast
8:48 to 12:45
Explaining how the EVOL system trains the rubric generator without human labels.
“The massive 32 billion parameter judge only scored 67.7%.”
Scaling the Rubric Generation Process
12:45 to 14:00
Discussion on how AI self-improvement leads to more nuanced grading rubrics.
“And the researchers actually tested several different ways to train the rubric.”
The Evolution of Grading Rubrics
14:00 to 14:51
Learn how AI grading rubrics evolve to become more precise over time.
“The difficulty of the grading automatically scales up alongside the intelligence of the policy.”
Visualizing AI-Generated Rubrics
14:51 to 16:44
Discover how AI transforms abstract grading criteria into specific checks.
“Okay, I want to visualize what these hyper-evolved rubrics actually look like when they print out.”
The Importance of Format in AI Rubrics
16:44 to 18:57
Understand the role of structural formatting in AI rubric generation.
“Just scans a text of the policy's answer, looks for the number 144, and gives the points.”
Show all 14 chapters
Evaluating AI in Real-World Tasks
18:57 to 21:28
Examine the performance of AI-generated rubrics on complex real-world tasks.
“Okay, so we have this open source AI successfully training itself, creating these hyper-dense rubrics, using its past self as the benchmark, and keeping itself on track with structural formatting.”
A New Era for AI Grading
21:28 to 21:54
AI is advancing beyond human grading with self-evolving systems.
“We are witnessing AI finally move past the bottleneck of human grading.”
The Democratization of Intelligence
21:54 to 22:36
Explore how open-source AI can revolutionize intelligence development.
“And the reason this matters to you, the listener, is about the democratization of intelligence.”
Future Implications of Evolving AI
22:36 to 23:09
Consider the impact of self-evolving AI on subjective human experiences.
“Like art, philosophy, the writing of a novel.”
Transcript
Automatic transcript. May contain errors.0:00Imagine trying to teach a student who is infinitely smarter than you. Oh, man, that sounds like a nightmare. Right. Like they're reading books and languages you don't even speak and they are solving equations you can't even write down. Yeah, you'd be totally lost. Exactly. So how do you grade their final exam when you don't actually understand the answers they're giving you? Well, you're touching on the exact ceiling we've hit with artificial intelligence today. Yeah. We've built this entire system of machine learning that is fundamentally reliant on human supervision. I mean, we are the teachers holding the red pens.
0:34And that works fine, you know, when the AI is just learning basic math or summarizing an email for you. Sure, for the easy stuff. But we are stepping into an era where we want these models to cure diseases or solve fusion energy. We want them to surpass us. Right. So, welcome to the Deep Dive. Today we're exploring a totally mind-bending breakthrough in how AI actually learns. It really is a massive shift. Because we know AI gets smart by reading the internet and learning patterns. But getting smart isn't the same as getting useful. Not at all. To get useful, an AI has to know the difference between a good answer and a bad one.
1:11And historically, human graders have had to sit there and do that. Yeah, humans or, you know, massively expensive proprietary models like a GPT-4, basically grading the AI's homework to train it. Which creates a huge bottleneck. A massive bottleneck. If an AI is only ever trained by a human or by a system designed to mimic human judgment, its capabilities are intrinsically capped by our own understanding. It can't grow past us. Exactly. Human judgment just cannot supervise intelligence that goes beyond its own horizon. So that is our mission for you today. We're looking at this new research detailing a breakthrough called EVOL.
1:47EVOL, evolutionary language model. Right. And this isn't just some minor software patch. This is a brand new method where a single language model learns to generate its own grading rubrics. It's wild. It continuously improves itself without any human supervision. No external API, no human in the loop. Nothing. It effectively pulls itself up by its own bootstraps. And once you understand how an AI does this self-evolution, it becomes clear why human-written rubrics are just, they're on the verge of becoming totally obsolete. Oh, absolutely. It democratizes the future of machine learning by completely removing that human bottleneck.
2:24So to really grasp the magnitude of what Evil M is doing, we first have to look at why the old way of grading AI is breaking down. Yeah, let's unpack that because the industry has tried to build automated graders to get around the human bottleneck. Right. They knew humans were too slow. Exactly. In current reinforcement learning, to train an AI policy, which is the part of the brain that generates the answers, you need a reliable reward signal. You need something saying good job or bad job. Right. And because paying humans to do that is incredibly slow and expensive, the industry started relying heavily on something called a scalar reward model.
3:01A scalar reward model. I mean, that just sounds like a highly technical way of saying a black box that spits out a score. That is actually a perfectly accurate translation. Okay, good. A scalar reward model takes whatever answer the AI just generated, runs it through this massive opaque neural network, and it just outputs a single number. Like a grade. Yeah, a score. Maybe it's a 9.2 out of 10. But the catch is? The catch is that the evaluation criteria, the actual reasoning behind why it gave a 9.2 instead of an 8.5, is hidden deep in the model's weights. So you can't inspect it. You can't read it at all.
3:35It's just a raw number, no feedback whatsoever. And that lack of transparency leads to a critical failure in training. Which this new research really highlights, right? Oh, spectacularly. The data highlights a wild example of this. The researchers looked at a state-of-the-art scalar reward model called SkyWork RM. SkyWork RM, okay. Yeah, it was explicitly designed to be an incredible automated grader. And when they put it through a benchmark that just tests its grading ability, it scored an amazing 86.4%. Okay, an 86.4%. I mean, that's a highly capable grader. That's a solid B plus bordering on an A.
4:10Right. So you'd assume it would make a fantastic teacher for developing AI. But when the researchers actually used this highly rated grader to train a downstream AI policy, when they let it act as the teacher for a student model, it produced the absolute weakest downstream model in the entire study. Oh, wow. The student AI only scored 59.7 % on its evaluations. It was a complete failure. Wait, I'm genuinely confused by that math. Which sounds backward, I know. Yeah. If the teacher is passing its own tests with flying colors, proving it knows how to grade accurately, why is the student failing the class so miserably?
4:49Well. Because the better the greater, the better the feedback should be. You're assuming the student AI is actually trying to learn the material. It's not. No, it's not. It's trying to maximize its score. This is a phenomenon called reward over optimization. Reward over optimization. Let me think about that. Is it like a student who realizes the teacher is obsessed with a very specific quirky grading style? Go on. Like maybe the teacher always gives high marks to essays that use the word plethora and have exactly five paragraphs. Yes. That is the exact mechanism at play. So the student completely stops trying to write a thoughtful essay about history or science.
5:25They just learn how to hack the rubric. They just stuff their paper with the word plethora. They format it perfectly. And the black box grader spits out a 10 out of 10. They over-optimize for the reward. Yeah. But when you put them in the real world, they fail because they haven't actually learned how to think. Wow. The static, opaque scoring model just cannot adapt to a learning student. As the policy gets smarter, it probes the black box greater. It finds the loopholes. It figures out its blind spots and exploits them entirely. The landscape of what the student is producing fundamental changes over time, but the teacher's hidden criteria stay exactly the same.
6:03So the training signal just becomes completely useless because it's being gamed. Exactly. And that reward over optimization is exactly why a static grader is doomed to fail. The only way out of that trap is if the teacher learns at the exact same pace as the student. The grader has to evolve with the AI. Right. And that brings us to the core of this breakthrough, the EvilMM co-evolution loop. EvilMM solves this hacking problem by entirely ditching the opaque black box. So no more hidden scores. Right. It forces the AI to write out its grading criteria in plain natural language, explicit readable rubrics.
6:38Okay, so how does it do that? It takes a single open source model. In this research, they used an 8 billion parameter model called Quen38B, and it splits it into two alternating roles. And it's like a Jekyll and Hyde dynamic, but applied to studying. Exactly. In role number one, the model acts as the policy. Its only job is to answer questions and generate text. The student. The student. Then in rule number two, the model switches hats and acts as the rubric generator. The teacher. Right. Its job is to look at a prompt and write down the exact verifiable criteria needed to grade an answer to that prompt.
7:15But wait, they're the same model? They share the same underlying neural architecture, the same parameters, but they take turns updating and training each other in this continuous loop. But if the student and the teacher are essentially the exact same entity, just taking turns, wouldn't they just collude? What do you mean? Like, the AI could write an incredibly lazy answer and then switch hats and write an incredibly lazy rubric that gives that bad answer an A+. Oh, I see. There has to be a third party keeping them honest. There is. There is a third component. The researchers introduced a judge.
7:46Okay, a judge. But this judge isn't some massive, state-of-the-art supercomputer. It is a tiny frozen 1.7 billion parameter language model. Wait, frozen? Yeah, it does not learn. It does not update its weights. It just reads the natural language rubric generated by the rubric generator, applies it to the policies answer, and calculates the score based strictly on those instructions. Okay, this feels deeply counterintuitive. How so? You are trying to build an AI that can rival human experts, right? Right. And you hand the ultimate gavel to a tiny, dumb 1.7 billion parameter judge. If you want the best possible feedback, why wouldn't you use a massive, super smart 32 billion parameter model as the judge?
8:31Well, the data reveals something fascinating about that exact assumption. The researchers actually tested different judge sizes to see what would happen. The tiny 1.7 billion parameter judge actually produced a better downstream policy, scoring 69.3%. Better than the big one. Yes. The massive 32 billion parameter judge only scored 67.7%. A less capable judge created a smarter student model. Exactly. How does that physically make sense? Because a massive, brilliant judge is too forgiving. Oh, I think I get it. Yeah, if the rubric generator gets lazy and writes a vague, abstract instruction like, check if the math makes sense, a 32 billion parameter judge is smart enough to do the heavy lifting itself.
9:16It just figures it out. It can infer what make sense means. It bridges the gap between a bad rubric and a good evaluation. But the tiny judge. But a tiny 1.7 billion parameter judge. It immediately gets confused. It performs wildly unreliably if the instructions aren't absolutely perfect. So using the tiny judge forces the rubric generator into a corner. Right. It can't be lazy because the judge won't bail it out. It's like, okay, if I'm writing a recipe for a Michelin star chef, I can just write make a standard roux and walk away. Sure, they know what that means. But if I'm writing a recipe for someone who has literally never been in a kitchen, I have to be excruciatingly specific.
9:52I have to write, melt exactly two tablespoons of butter over medium heat, add two tablespoons of flour, and whisk constantly for exactly two minutes until it smells like pie crust. That's a great analogy. Yeah. The tiny judge forces the AI to extract all of its own hidden knowledge and put it explicitly on the page. It completely removes the interpretive burden from the judge. Right. The quality of the training signal now comes entirely from the content of the rubric itself rather than the underlying capability of the judge reading it. Okay, that makes total sense. It solves the problem of keeping the rubric specific.
10:27It does. But it raises a much bigger question for me. Which is? If there is absolutely no human involved in this loop providing ground truth labels, how does the rubric generator know if it has written a good rubric? Like, how is the rubric itself being graded? The researchers approach this by defining a concept called discriminative utility. Discriminative utility. Think about the ultimate purpose of a grading rubric. A good rubric is simply one that helps a third party at this case. The tiny judge clearly separate a fantastic answer from a terrible one. Right. If a rubric gives every single student an 85, regardless of what they wrote, it has zero discriminative utility.
11:08It's totally useless. Exactly. It needs to give the great essay 100 and the terrible essay a 50. So how do you train it to do that? To train the rubric generator to do this, you have to feed the system what are called preference pairs. Preference pairs. You give the tiny judge one preferred answer, a good one, and one dispreferred answer, a bad one. Okay. If the rubric generator writes a rubric that successfully guides the tiny judge to give the preferred answer a higher score than the dispreferred answer, the rubric is rewarded. It did its job. I'm tracking with discriminative utility. But we are back to the human bottleneck again.
11:44Are we? Where do those good and bad answers come from to test the rubric if there's no human writing them or labeling them? Ah, through a mechanism called temporal contrast. Temporal contrast. This might be the most elegant part of the entire system. To get that pair of good and bad answers, the system takes a brand new answer the AI just generated. Okay. And pairs it with an answer the AI generated several steps ago in the past. Oh. It's comparing its present self to its past self. It creates its own preference pairs purely through the passage of time. That is brilliant. It's like, think about a writer working on a novel.
12:20Right. If they compare the final draft they finished today with the rough draft they wrote two weeks ago, they don't need an external editor to tell them which one is better. Nope. The newer one is inherently better because the writer has spent the last two weeks practicing, editing, and refining their skills. Exactly. You inherently know that today draft is the preferred answer and the two weeks ago draft is the dispreferred answer. That is the exact mechanism driving evil M. Yes. And the researchers actually tested several different ways to train the rubric. They tried having the AI guess the original question based on the answer or writing answers conditioned strictly on the rubric.
12:56But temporal contract won out. Temporal contrast outperformed all of them because it creates a natural self-scaling curriculum. A self-scaling curriculum. Let's break that down. What happens as the AI gets smarter over time? Well, early on in the training process, the gap between today's answer and two weeks ago's answer is massive. Because it's still learning. Right. The past answer is probably riddled with basic errors. Because the difference is so obvious, it's very easy for the tiny judge to tell them apart, even if the rubric generator writes a fairly basic grading sheet. But fast forward thousands of steps.
13:31Exactly. The AI has gotten highly capable. Now, the gap between the new answer and the old answer shrinks dramatically. They both look like excellent, expert-level responses. So the tiny judge looks at both of them and says, I don't know, they both look like A-plus answers to me. Gets confused again. Precisely. So to help the tiny judge tell them apart now, the rubric generator is forced to invent sharper, incredibly nuanced, hyper-specific criteria. It has to dig deeper. It has to dig deeper into the actual logic of the prompt to find the minute differences in quality. The difficulty of the grading automatically scales up alongside the intelligence of the policy.
14:11It is constantly raising its own bar. It really is. But wait, what happens if the AI just has a bad day? A bad day? Like, what if, purely by statistical chance or a hallucination, the draft it writes at step 500 is actually worse than the draft it wrote at step 400? Doesn't that break the temporal contrast model? If you were only running this a dozen times, yes, statistical noise could derail it. Right. But over thousands and thousands of iterations, the macro trend is relentless improvement. The law of large number. Exactly. The sheer volume of comparisons acts as a filter. The overall trajectory forces the rubrics to become tighter and more robust, smoothing out any individual step where the model might have hallucinated a worse answer.
14:54Okay, I want to visualize what these hyper-evolved rubrics actually look like when they print out. Oh, they're fascinating. Because when humans write rubrics, they're usually pretty generic. Like a teacher's rubric might say, checks grammar, 10 points, understands the core topic, 20 points. Do the AI rubrics look like human rubrics? Not at all. As the co-evolution loop runs, the rubrics transform from abstract human-like labels into highly verifiable, almost programmatic checks. Like code. Almost. The system learns to enrich the criteria. It actively packs the expected answers directly into the rubric itself to make the tiny judge's life as frictionless as possible.
15:37I remember looking at the data and they gave the AI a complex math problem to solve. It was find the maximum area of a rectangle with a perimeter of 48. Yes, that's a great example. And early on in the training loop, the AI rubric generator writes these vague human sounding rubrics. It outputs stuff like correctly applies the perimeter formula or correctly calculates the maximum area. Right, which leaves it up to the judge to actually execute the math and verify the work. But by step 1000 of the training loop, the structure completely changes. It realizes that asking a tiny 1.7 billion parameter judge to execute optimization calculus is risky.
16:13Extremely risky. So by step 1000, it consolidates 80 % of the entire grading weight into one hyper-specific rule. The rubric literally just says the answer is the correct maximum area of 144 derived from the given perimeter of 48. It pre-computes the answer itself. It pre-computes it and embeds the number 144 right into the grading instructions. It transforms a complex proof verification task into a simple pattern matching task. The tiny judge doesn't need to know how to do math anymore. It just scans for the number. Exactly. Just scans a text of the policy's answer, looks for the number 144, and gives the points.
16:49The rubric generator has extracted its own latent mathematical knowledge and explicitly externalized it for the judge. It's basically writing a foolproof cheat sheet for a substitute teacher. That is exactly what it's doing. But, you know, math is objective. It's verifiable. The data also mentions tests on highly subjective tasks like emotional support prompts. Yeah, and it exhibits the same evolutionary behavior for subjective tasks, which is just remarkable. Really? Early on for an emotional support prompt, the rubric might say something abstract like offers meaningful insights to the user. Which is highly interpretive.
17:23Very. But by step 1000, it abandons that vague language entirely. It evolves into dense, specific checklists. What does that look like? It requires the judge to verify if the answer, quote, provides actionable strategies. For example, grounding techniques, self-compassion, seeking professional help. Oh, wow. So it forces concreteness onto abstract human emotions. Exactly. But there is a technical risk here, right? You mean? If it's just a machine talking to a machine trying to find the absolute most efficient way to communicate, how do you stop them from developing their own weird, unreadable shorthand?
18:00Oh, right. How do you keep the rubrics legible to us humans? Well, the researchers actually had to build in a critical engineering constraint for that exact reason. Okay, what was it? To make sure the AI didn't just spit out garbled code or break the system entirely, they added a format reward. A format reward? Yeah, it's a small technical rule that forces the rubric generator to output its criteria in a strict, valid JSON format. JSON, so basically a highly structured text file with specific brackets and indents that a computer can easily read. Exactly. What happened if they turned that format reward off?
18:35Did they test that? They did, and the result was complete structural collapse. Really? Yeah. Without it, the percentage of structurally valid rubrics plummeted from over 85 % all the way down to 23%. Wow. The continuous evolutionary pressure to be as efficient as possible actually broke the formatting entirely unless it was explicitly rewarded for keeping it. So the system needs those structural guardrails to keep the evolution from just descending into total chaos. It absolutely does. Okay, so we have this open source AI successfully training itself, creating these hyper-dense rubrics, using its past self as the benchmark, and keeping itself on track with structural formatting.
19:13That's the loop. But the ultimate test is reality. It's one thing for the AI to train itself on standard math and coding tasks from its own data set. Sure. Can these self-taught rubrics generalize? Like, can they handle complex real-world domains that the AI has never been explicitly trained on? And this brings us to out-of-distribution success. The deep research tasks. Yes. The researchers tested EVOLUM on these deep research tasks. These are complex, multi-step problems in highly specialized fields. And crucially, they put EVOLUM's self-generated rubrics head-to-head against rubrics written by actual human experts.
19:53Okay, let's look at the domains they tested. They use two massive benchmarks, HealthBench and ResearchQA. Right. And just for context, a task on HealthBench isn't just asking, what is a normal body temperature? No, not at all. It's evaluating complex medical diagnosis paths, checking if a response correctly weighs the contraindications of prescribing a specific drug given a patient's multilayered medical history. It's incredibly complex. Writing a rubric to grade that requires immense specialized knowledge. It's a domain where precision is literally life or death. Right. And the results here are staggering.
20:25Evol M. achieved a higher pairwise agreement with the expert human rubric.
20:36Evol M. generated rubrics that aligned with human experts 58.4 % of the time. GPT 4.1 only hit 52.5%. That is wild. And on Research QA, EVOLM hit 59.3%, while GPT-4.1 was stuck at 51.0%. I mean, think about the magnitude of this for a second. GPT-4.1 is a massively expensive proprietary model with billions of dollars of human reinforcement learning behind it. Massive resources. Thousands of human contractors were paid to refine its judgment. Yeah. And here we have an open source model pulling itself up by its bootstraps without any external human labels. And it is generating grading criteria that align better with human experts on deep medical research than the billion dollar titan.
21:21It proves that this co-evolutionary loop of fighting its past self isn't just memorizing tricks for standardized tests. It's real intelligence. It is learning the fundamental, universal structures of what makes a good answer in any complex domain. Well, let's bring this all together. We are witnessing AI finally move past the bottleneck of human grading. We really are. By taking a single model, splitting it into a student and a teacher, and using a tiny unforgiving judge to force absolute clarity, Evil M proves that an AI can extract its own latent knowledge. Exactly. It creates perfectly verifiable, hyper-specific grading rubrics, all by just trying to beat its past self through temporal contrast.
22:02And the reason this matters to you, the listener, is about the democratization of intelligence. That's the big takeaway. We are entering an era where AI development is no longer bottlenecked by human labor. It isn't limited to domains where we have absolute ground truth answers, like a math equation or a line of code. Right. Because open source models can now self-improve infinitely. The cost of training hyperintelligent systems drops drastically. It breaks the monopoly of the massive proprietary labs. Which is incredibly exciting. And it leaves us with a final lingering question to ponder. What was it?
22:35If an AI can now independently define what good looks like and even beat out human experts in designing grading rubrics for deep research, what happens when we apply these self-evolving rubrics to deeply subjective human experiences? Oh, wow. Like art, philosophy, the writing of a novel. Will the AI eventually discover patterns of good that our human minds can't even perceive? That's a deep thought. Are there structures to beauty or narrative that an evolving rubric will be able to measure, but we simply can't? Guess we'll find out. Something for you to mull over today. Thank you for taking this deep dive with us.
From the publisher
This paper introduces EVOLM, an innovative framework for self-evolving language models that improves performance without relying on human annotations or external teacher models. By transforming a model’s internal knowledge into explicit natural-language rubrics, the system creates an autonomous feedback loop where evaluation and generation capabilities improve in tandem. This method utilizes variational inference to optimize rubric generators, rewarding criteria that successfully help a small, frozen judge distinguish between superior and inferior responses. Experimental results demonstrate that EVOLM outperforms established baselines, including GPT-4.1, by shifting from abstract judgments to verifiable, instance-specific criteria. Ultimately, the research shows that structuring evaluative capacity into co-evolving rubrics allows models to surpass the limitations of static external supervision.




