Learning to Discover at Test Time

23 Jan 2026 · 16 min · 10 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Test-time training for “discovery” (TTT Discover) where an AI updates its weights during problem solving instead of only searching with a frozen model.

Guests

The transcript doesn’t name the guests; it’s a two-person discussion (an interviewer/host and a guest co-speaker) about a Stanford/NVIDIA/UC San Diego paper.

Key claims

Standard AI “guesses” at test time (frozen brain) and can’t learn from failures; TTT Discover performs reinforcement-learning updates during the test, optimizing for maximum reward outliers via an entropic objective and PUCT with max-reward state reuse.

Notable examples

Erdos minimum overlap problem (0.380876 vs AlphaVolve’s 0.380924) using a chaotic 600-piece asymmetric step function; GPU kernel optimization for NVIDIA A100/A100 competitions (TrimMol, MLA decode), achieving ~50% faster kernels than top human submissions; “decompose and distill” (fuse ops, FP16, delegate matmul to cuBLAS/Cubless). Cost claim: about $500 per problem using GPT-OS 120B and Tinker API.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Student Exam Analogy

0:46 to 2:16

Exploring the connection between student learning during tests and AI learning.

“You are not just retrieving what you studied.”

The Frozen Brain Problem in AI

2:17 to 4:01

Discussing how traditional AI struggles with real-time learning from mistakes.

“That sounds a little bit like semantics, but I have a feeling it's a huge technical shift.”

Introducing TTT Discover

4:02 to 5:50

Overview of the new approach called Test Time Training to Discover and its implications.

“And this is crucial because of something called out-of-distribution problems.”

Shift from Searching to Training

5:51 to 7:42

Understanding the difference between traditional AI approaches and TTT Discover's methodology.

“you don't care if your average training run is just okay.”

Revolutionizing Scientific Discovery

7:43 to 10:09

Explaining how TTT Discover changes the rules for AI in scientific research.

“TTT Discover looks at the maximum reward anyone has ever gotten from that path.”

Practical Applications of TTT Discover

10:10 to 13:30

Examining the real-world applications of TTT Discover in solving complex problems.

“It found a solution that we essentially refused to look for because it was ugly.”

The Cost of Discovery

13:31 to 14:00

Discussing the economic implications of TTT Discover and its accessibility.

“You could fund a scientific breakthrough with a bake sale.”

Challenging the Scaling Laws

14:00 to 14:31

The discussion challenges the notion that larger AI models are always better.

“The dominant narrative lately has been scaling laws.”

From Retrieval to Discovery

14:31 to 15:46

Explores the shift from AI as a retrieval engine to AI that can create new knowledge.

“We've stopped caring about the average safe answer and started hunting for the max reward outlier.”

AI's Potential and Perils

15:46 to 16:00

Discusses the thrilling yet terrifying prospects of AI solving problems beyond human understanding.

“If the AI can find asymmetric math solutions that humans overlook because of our biases, it might start solving problems in ways we can't even intuitively understand.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00I want you to picture a scenario. It is late. You are a student, maybe back in college, and you are staring at a final exam that is just a beast. It is way harder than anything you saw in the textbook. Oh, yeah. The classic out of distribution nightmare. We have all been there. Exactly. You look at the first problem. Your mind goes blank. You write down a formula, realize it's wrong, scratch it out. You try a different approach, hit another dead end. You are, you know, panicking. But then, if you are lucky, something clicks. You stop just throwing random guesses at the page. You stop panicking.

0:37You look at why your first attempt failed. You look at the error. You adjust your mental model. And you actually learn something new right there in the quiet of the exam hall. You get smarter during the test. Yes. You are not just retrieving what you studied. You are adapting. You're literally updating your own neural weights in real time to match the challenge in front of you. But here's the thing. until very, very recently, artificial intelligence, even the really impressive stuff, it didn't really do that. When an AI hits a hard problem, it usually acts like that panicked student just guessing.

1:10It just keeps rolling the dice based on what it already knows from its training months ago. It doesn't learn from its mistakes during the problem-solving process. That's the frozen brain problem. Once the training is done, the learning just stops. It's strictly retrieving, not reasoning, or at least not reasoning in a way that lets it improve from its own mistakes in the moment. But today, we are unpacking a paper that completely flips that script. It's called Learning to Discover at Test Time. And it's coming out of a huge collaboration between Stanford, NVIDIA, UC San Diego, and a few others.

1:44And they are introducing something called TTT Discover. Which stands for Test Time Training to Discover. And I've got to tell you, the results in this paper are absolutely wild. We are talking about an AI that solved a 60-year-old math problem that has stumped humans for decades, and maybe even more impressively. It optimized computer code to be faster than the best human engineers in the world. And it did it by fundamentally changing how it thinks about a problem. We are moving from a model that just searches for an answer to a model that trains itself to find the answer. So let's unpack this.

2:17Moving from searching to training. That sounds a little bit like semantics, but I have a feeling it's a huge technical shift. What is the actual difference? Well, to understand the shift, we have to look at the status quo. Before this paper, if you wanted an AI to solve a really hard problem, like proving a theorem, you used what we call test time search. You might know it from methods like AlphaVolv. Is this kind of like how I might ask ChatGPT to generate 10 different headlines and I just pick the best one? It's a similar concept, yeah, but a bit more structured. Imagine you have a large language model.

2:52Its neural weights, the actual connections that make up its intelligence, are frozen. They were set months ago during its initial training. When you give this model a hard problem, you prompt it to generate a solution. If it fails, maybe you have a system that prompts it to try again or try a different angle. So it's iterating on the output, but the model itself isn't changing. Precisely. It creates a buffer of attempts. It might generate thousands of guesses, but the model have amnesia. Attempt number 1000 is generated with the exact same baseline intelligence as attempt number one. It doesn't look at the previous 999 failures and say, oh, I see a pattern here.

3:31Let me rewire my understanding to fix that. It's just brute force. It's like trying to pick a lock by just jamming random keys in, hoping one fits instead of actually studying the tumblers inside. That is a perfect analogy. Now, TT Discover test time training takes a completely different approach. It says, why are we throwing away all that experience? Right. When the model makes an attempt and fails or gets a partial success, TTT Discover uses that experience to actually perform reinforcement learning updates on the model's weights while it is taking the test. So it is rewriting its own brain wiring in real time, specifically for the problem it's trying to solve.

4:07Yes. And this is crucial because of something called out-of-distribution problems. discovery of problems like proving a new mathematical theorem or inventing a new algorithm, they require ideas that just aren't in the training data. Because they're new to humanity. They're new to humanity, so you can't just retrieve the answer because the answer doesn't exist yet. You have to invent it. And by letting it train on its own attempts. Correct. By allowing the model to train on its own attempts during the test, those attempts become a new, highly specific data set. The model essentially builds a little mini curriculum for itself to master this one specific, incredibly hard challenge.

4:44It becomes a specialist in that one problem for a few hours. That is fascinating. It's not just trying harder. It's getting smarter for this one problem. But this brings up a question I had while reading. We already have reinforcement learning. We have PPO. We have all these acronyms used to train the base models. Why couldn't they just use standard RL? Why did they need this new discover method? This is where we get into the philosophy of the paper, which I find really compelling. It challenges how we think about success in AI. Standard reinforcement learning. Think of the algorithms used to train a robot to walk or a car to drive.

5:19It's designed to find a policy that works well on average. Right. You want the self-driving car to be safe 100 % of the time. You want high reliability. Exactly. You want a robust policy. You maximize the expected reward. But in scientific discovery, the rules are completely different. You don't care about the average. You don't care if the model fails 999 times. You only care about the one time it succeeds. It's the difference between a daily commute and breaking a world record. Perfect analogy. If you are Usain Bolt trying to break the 100-meter world record, you don't care if your average training run is just okay.

5:54You don't care if you trip and fall 500 times. You are hunting for that one perfect lightning-fast run that breaks the record. And once you have the discovery, the policy, or the brain that found it, it doesn't matter anymore. It doesn't matter. You have the answer. So TTT Discover is effectively hunting for outliers. It wants the spike in the graph, not the flat line. How does it technically force the model to stop playing it safe? They introduced two main technical mechanisms to drive this one perfect solution philosophy. The first is what they call the entropic objective. That sounds like something out of a sci-fi movie.

6:29What does it actually do? In simple terms, standard RL optimizes for the expected reward, the average. The entropic objective changes the math to exponentially favor the maximum reward. It tells the model, I don't care about steady progress. I want you to take risks. I want you to find the spike. So it incentivizes the model to swing for the fences. It rewards the jackpot, not the steady salary. Yes. If a certain path has a 99 % chance of failure, but a 1 % chance of a massive breakthrough, a standard model avoids it because the average outcome is bad. The entropic objective tells T.T. Discover to dive right in.

7:06Okay, so that's the first mechanism. And the second? The second is about how it explores. They use a technique called P-U-C-T predictor plus upper confidence bound applied to trees. It's basically a way of reusing past states. State reuse. That's like saving your game before a boss fight so you don't have to replay the whole level. That's a great way to put it. But here is the twist. In traditional systems like AlphaZero, which mastered chess and Go, the system decides which state to revisit based on the mean value of that path. It asks, on average, is this a winning position? But TTT Discover doesn't care about the average.

7:42Exactly. TTT Discover looks at the maximum reward anyone has ever gotten from that path. It prioritizes paths that have produced one spark of genius, even if the rest of the path was absolute garbage. So if I went down a rabbit hole 10 times, 9 times it was a dead end, But one time I found a gold coin, TTT Discover says that's a gold coin path. Whereas AlphaZero says that's a 10 % success path. Stay away. Correct. This allows the model to extend its horizon. Instead of getting stuck in short loops, it builds complex multi-step solutions over time because it is constantly lured forward by those glimmers of maximum reward.

8:19Okay, so we have a model that learns in real time and is just obsessed with finding one perfect outlier. Let's see what this actually looks like in practice. The paper dives into a math problem called the Erdos minimum overlap problem. This is a classic. It was posed by the legendary mathematician Paul Erdos back in 1955. And for those of us who aren't mathematicians, what's the goal here? Simplistically, you are trying to partition or, you know, cut up a set of numbers to minimize the overlap of their differences. It's a combinatorial number theory problem. For decades, mathematicians have been trying to find the step function that certifies the lowest possible overlap constant.

8:58And the previous state of the art was set by AlphaVolv, the search-based method we talked about earlier. Right. AlphaVolv found a solution with an upper bound of about 0.380924, and it did this by finding a symmetric construction. Symmetric meaning it looks neat, balanced. Yes. Humans and the AI models trained on human data have a strong bias towards symmetry. We assume the perfect answer must be beautiful and balanced. AlphaVolve found a 95-piece symmetric step function. And then TTT Discover comes along. And TTT Discover blew it out of the water. It found a bound of 0.380876. That sounds like a small difference in numbers, but I'm guessing it's not.

9:40In this field, that is a massive leap. But the fascinating part is what it found. It didn't find a neat symmetric pattern. It found a chaotic 600-piece asymmetric step function. So the AI realized that our obsession with symmetry was actually holding us back. Exactly. It broke the human bias. Because it was learning a test time and because it only cared about the max reward, it was willing to venture into this messy asymmetric territory that a human or a frozen model would likely ignore as wrong. That is the aha moment right there. It found a solution that we essentially refused to look for because it was ugly.

10:13It did. and it proves that discovery requires breaking away from the distribution of what we already know. Math is great, but let's pivot to something that runs the modern world, code, specifically GPU kernels. This is where the rubber meets the road. If you are using ChatGPT or Claude or any AI, you are running on GPUs. And those GPUs run on kernels, low-level code that manages how data moves and multiplies. And optimizing these kernels is basically a dark art. There are entire competitions for it. There are. The paper looked at a competition for the NVIDIA H100 and A100 GPUs. Specifically, they targeted TrimMol, triangular matrix multiplication, which is a core component of AlphaFold.

10:56So making this faster means we can do biology faster. Precisely. And they also looked at MLA decode, which is used in DeepSeq. The competition is to write code that runs as fast as possible. And how did TT2 Discover do? It beat the best humans. On the A100 GPU leaderboard, the kernel written by TT2 Discover was 50 % faster than the top human submissions. 50 % faster. That is not a marginal gain. In optimization, people fight for 2 % or 3%. That is a generational leap. How did he just find some clever math trick? It did something more impressive. It understood the hardware. This brings us to a concept the researchers dubbed decompose and distill.

11:33Okay, break that down for us. What does decompose and distill actually look like? So the AI first decomposed the problem. It analyzed the task and realized that for this specific operation, the bottleneck wasn't the math. It wasn't compute bound. It was memory bound. It was spending too much time moving data back and forth. Exactly. It's like having a Ferrari engine but being stuck in traffic. The chip was waiting for data. So the AI decided to distill the operations. It fused them. It took the input layer norm, the sigmoid gating, and the output layer norm and fused them into single kernels to stop the data from bouncing around.

12:09It's like doing all your grocery shopping in one trip instead of driving back home after buying the milk, then going back for eggs. That's a great analogy. And it didn't stop there. It converted inputs to FP16 mixed precision to save bandwidth, and it delegated the heavy matrix multiplication to Cubless, which is a library optimized for the hardware's tensor cores. So it knew what it was good at, and it knew what the hardware library was good at, and it acted like a chief engineer delegating tasks. It demonstrated a profound understanding of hardware constraints, memory I.O. versus compute. And remember, this wasn't a coding model trained specifically on kernel optimization.

12:46It learned to do this during the test time training. It poked the hardware, saw what worked, updated its weights, and iterated until it found this decompose and distill strategy. That leads me to the big question. When we hear about these massive breakthroughs, alpha, kalfa, fold, we usually assume it required a supercomputer and millions of dollars. Who actually built this? That is the kicker. This wasn't GPT-5 or Gemini Ultra. This was done using GPT-OS 120B. An open model. An open weights model. And because they used this TTT method where they explore and train on this specific problem, the cost was incredibly low.

13:23How low are we talking? They used the Tinker API for the training runs. The cost to find these state-of-the-art solutions beating 60 years of math and the best NVIDIA hackers was about$500 per problem. $500? You could fund a scientific breakthrough with a bake sale. It completely democratizes discovery. It suggests that the barrier to scientific breakthrough isn't necessarily having the smartest base model or the biggest cluster. It's about having a model that can think, learn, and iterate on the specific problem you put in front of it. It's the difference between hiring a genius who knows everything but can't learn versus hiring a smart, hardworking person who is given the time and resources to master a specific task.

14:05Exactly. The dominant narrative lately has been scaling laws. Just make the model bigger. This paper challenges that. It says maybe we don't need bigger brains. We need brains that can concentrate better. And the decompose and distill success proves that this method can handle complex, multifaceted engineering challenges, not just pure math. Right. It can plan. It can understand systems. It's a very versatile type of intelligence. Right. So let's wrap this up. We've moved from guessing to learning. We've stopped caring about the average safe answer and started hunting for the max reward outlier.

14:39And we've proved that open models can beat humans at their own specialized games for the cost of a gaming console. It is a fundamental shift. We are proving that search, or rather, discovery scales. What does this mean for the future? I want to leave you with this thought. Right now, we treat AI largely as a retrieval engine. We ask it questions, and we expect it to give us answers based on what humanity already knows. summarize this, translate that. But TT Discover shows us the path to something else. If an open model can beat human experts and closed frontier models simply by thinking and learning on a specific test problem for a few hours, what happens when we apply this to problems that currently have no solution?

15:21Problems where there is no training data because no one knows the answer. Exactly. We aren't talking about optimizing a kernel anymore. We are talking about protein folding for specific rare diseases. We're talking about finding the magnetic configuration to stabilize fusion energy. We are moving from AI that retrieves knowledge to AI that creates new knowledge. The transition from student to inventor. And perhaps even beyond inventor. If the AI can find asymmetric math solutions that humans overlook because of our biases, it might start solving problems in ways we can't even intuitively understand.

15:56We might just have to trust the results. That is both a thrilling and slightly terrifying thought. Yeah. The idea that the answers to our biggest problems are hiding in the asymmetric outliers that we're too afraid to look for. But now we have a machine that isn't afraid to look there. Well, on that mind-bending note, I think we will leave it there. Thank you so much for joining us on this deep dive into TTT Discover. It's been fascinating. It's a pleasure. And to everyone listening, keep learning and maybe let your models learn a little bit too. We'll catch you on the next deep dive.

From the publisher

This paper introduces TTT-Discover, an innovative system designed to solve complex science and engineering problems through test-time training. Unlike traditional static models, this approach enables an open-source AI to continuously learn and refine its policy while actively seeking solutions for a specific task. By utilizing an entropic objective and adaptive reinforcement learning, the system successfully established new state-of-the-art results in mathematics, GPU kernel engineering, and biology. The researchers demonstrate that this method can outperform elite human experts and powerful closed-frontier models at a fraction of the typical computational cost. This framework effectively transforms the problem-solving process into an iterative search and learning environment where the model improves itself until a breakthrough is reached. Notably, the paper details successful applications in Erdős’ minimum overlap problem and high-performance algorithm design.

More from Best AI papers explained

All 475 episodes
Learning to Discover at Test TimeBest AI papers explained · 16 min
Listen in VO