In short
The episode argues that today’s transformer LLMs mainly do next-token “parroting” due to attention’s effectively infinite lookback memory, which encourages shortcut learning (“epicycles”) instead of compact world models. It presents Next-Latent Prediction (NextLat): train a transformer to predict its own next hidden states via a small 3-layer latent dynamics side model, producing compressed “belief states” that support longer planning and better reasoning.
Notable examples
Manhattan taxi text—standard models get 100% next-turn accuracy but have a messy internal “map,” while NextLat yields a grid-like representation with lower effective latent rank (160.1 to 52.7). Countdown math (game of 24-like)—reduces “regretful compromise” (42.3% to 54.8% valid). PathStar graph routing—near-100% solve rate on large graphs. A5 anomaly—tiny side model trained on 12 tokens solves 36-token sequences (>95%) when the main transformer fails (0%). Variable-length self-speculative decoding—up to 3.3x faster generation (1.3B model).
Guests
No specific guest names or backgrounds are provided; it’s a two-host discussion.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VONext Latent Prediction Overview
0:45 to 2:16
Discover the new paradigm of next latent prediction in AI architecture.
“And we tend to anthropomorphize these models.”
The Flaw in Standard AI Models
2:16 to 3:41
Understand the limitations of current AI models based on transformer architecture.
“It is the difference between a parlor trick and a true reasoning engine.”
Ptolemy's Geocentric Model as an Analogy
3:41 to 5:50
Learn how Ptolemy's model illustrates the flaws in standard AI logic.
“Why do the hard work of building a cohesive mental model of a subject when you can just look up the exact word you need whenever you want?”
Introducing NextLatentPrediction Mechanics
5:50 to 8:00
Explore the mechanics of Nexlat and its potential to improve AI learning.
“Standard AI, with its infinite lookback memory, completely lacks the pressure to form a simple compact explanation of the world.”
Benefits of Nexlat's Approach
8:00 to 9:16
Understand how Nexlat combines the efficiency of transformers with RNN capabilities.
“Because I thought the whole industry abandoned RNNs years ago.”
Testing AI with Manhattan Taxi Rides
9:16 to 11:10
Learn about an experiment testing AI's spatial reasoning through taxi rides.
“But it's one thing to say this creates a world model or a belief state in abstract theory.”
Countdown Math Task and AI Limitations
11:10 to 12:50
Discover how standard AI struggles with mathematical reasoning tasks.
“Nexlat builds an internal mathematical map that closely mirrors the real Manhattan grid.”
Nexlat's Impact on Abstract Reasoning
12:50 to 14:01
See how Nexlat improves AI performance on abstract reasoning tasks.
“The AI is given a list of random numbers and a target number, and it has to use basic arithmetic like addition, subtraction, multiplication, division to combine the random numbers and hit the target exactly.”
Nexlat's Future Predictions and Consistency
14:01 to 14:31
Learn how Nexlat enhances AI's predictive capabilities for future states.
“But because Nexlat forces the model to predict its future latent states, its future thoughts, it enforces global planning.”
Clever Hans and AI Learning Shortcuts
14:31 to 15:00
Discover the Clever Hans phenomenon and its implications for AI learning.
“Yeah, Clever Hans was a horse that appeared to do math but was actually just reading the subtle body language of his trainer to know when to stop tapping his hoof.”
Show all 17 chapters
PathStar Graph Task Performance
15:00 to 15:28
Explore Nexlat's performance on complex logic tasks compared to standard models.
“But Nexlat maintains a near 100 % solve rate, even on the largest, most complex graph sizes.”
The A5 Anomaly Explained
15:28 to 16:18
Understand the A5 anomaly and its challenges for traditional transformers.
“The A5 word problem is a highly complex state tracking task.”
Training NEXLAT on A5 Problem
16:18 to 16:43
Learn about the training process and unique success of the NEXLAT model.
“By definition, a shallow transformer fundamentally cannot solve the A5 problem for long sequences.”
Tiny Models Surpassing Larger Ones
16:43 to 18:20
Explore how tiny models can outperform larger transformer networks.
“So researchers trained a NEXLAT model on short sequences of 12 tokens for this A5 problem.”
Benefits of Latent Prediction in AI
18:20 to 20:21
Discover how latent prediction enhances speed and efficiency in AI generation.
“It really challenges our assumptions about how we need to scale AI moving forward.”
Transforming AI Understanding and Efficiency
20:21 to 21:28
Learn how Nexlat shifts AI from memorization to understanding and planning.
“Nexlat accelerated text generation by up to 3.3 times compared to standard generation.”
Future of AI: Tiny Minds and Ecosystems
21:28 to 22:30
Contemplate the future of AI as potentially small, efficient models evolve.
“Think back to that A5 anomaly we unpacked, where the tiny auxiliary inner network learned logical rules that completely exceeded the cognitive limits of its massive parent network.”
Transcript
Automatic transcript. May contain errors.0:00You know, there is this really fascinating contradiction at the heart of modern AI right now. and you've probably experienced it firsthand. Oh, absolutely. Everyone has by this point. Right, because if you sit down right now and ask a top-tier AI to, I don't know, write a beautiful, deeply emotional Shakespearean sonnet about a toaster. It'll do it instantly. Yeah, in like three seconds flat. Yeah. The rhymes will be perfect. The meter is flawless. But if you ask that exact same AI to navigate a really simple logic puzzle or figure out the physical layout of a room, it just, it often completely fails.
0:36It hits a brick wall. It creates this wild illusion of, like, profound brilliance, which is then followed immediately by a display of total incompetence. Exactly. And we tend to anthropomorphize these models. You know, we think of them as these thinking engines. But fundamentally, historically, they really haven't been doing any thinking at all. And the reason why is that standard AI architecture is essentially playing this giant, highly sophisticated game of, guess the next word. It's a parrot. A very brilliant parrot. A brilliant parrot with an enormous vocabulary, sure, but a parrot nonetheless.
1:08It doesn't actually understand the physical or logical world it's talking about. But today, our mission for this deep dive is to unpack a groundbreaking shift in AI architecture. Well, it changes all of that. We are exploring a new paradigm called next latent prediction, or next lat. Yeah, Nexlat. And it represents a massive fundamental shift in how we actually build these systems. Because instead of training an AI to just predict the next word in a sequence, Nexlat trains the AI to predict its own internal thoughts. Its own future states. Exactly, its own future internal states. And the downstream effects of making that one specific change are just incredible.
1:49We're talking about AI that suddenly gains genuine long-term planning skills, AI that stops making these bizarre unforced logic errors, and surprisingly, AI that can actually generate answers up to three times faster. Which is huge for the industry. Right. So if you use these tools or you just want to understand where the technology is heading, you really need to understand this shift. It reveals the exact boundary between an AI that is just parroting data and an AI that actually comprehends the tasks you are asking it to do. It is the difference between a parlor trick and a true reasoning engine.
2:21To really grasp why NextLet is such a breakthrough, we first have to dissect the fatal flaw in the current gold standard of AI. Because almost every major language model you interact with today uses the transformer architecture. Running on next token prediction. Right, next token prediction. And the core issue there lies in how a standard transformer handles memory. So let's talk about that memory. Transformers have this massive memory capacity that grows with the sequence of text. And they use a mechanism called self-attention. Let's break that down for a second. Because self-attention basically means the AI has the ability to look backward at any previous word in its context window at any given time.
3:01At any time, yeah. It has this ad hoc, perfectly accessible, infinite archive of everything that has been said so far in the prompt. Right. So think about the implications of that for a learning system. Well, if you have perfect instantaneous recall of every single data point, you never actually have to understand the relationships between those points. It's like taking a test, but it's completely open book. And not just open book, you have a photographic memory of a textbook. Exactly. You don't need to learn the underlying concepts of physics or history. You just need to know exactly which page to flip to when a specific question is asked.
3:34So that open book nature basically removes all the friction required for true learning. Because it has that perfect infinite look back memory, the AI has absolutely no mathematical incentive to compress all that chaotic history into a set of consistent rules. Why do the hard work of building a cohesive mental model of a subject when you can just look up the exact word you need whenever you want? Right. And worse than just being lazy, it starts memorizing complex task specific shortcuts. Yeah, it finds superficial patterns that just happen to predict the next word in the training data, but it completely fails to build a generalizable structure of how the world actually works.
4:14Which we see all the time. Like, an AI might correctly guess the next word in a math problem because it's seen a similar string of numbers before, not because it actually understands addition. Yes, exactly. You know, the more I looked at this, the more it reminded me of Ptolemy's geocentric model of the universe. Oh, that's a great analogy. Right. Back in ancient Greece, Ptolemy built this incredibly complex mathematical system to predict where the stars and planets would be in the sky. And it was all based on the completely false assumption that the Earth was the center of the universe. But mathematically, it worked.
4:48It worked. His math accurately predicted the night sky from our perspective on Earth. But the underlying model was convoluted and totally wrong. He had to invent epicycles, these crazy little loops within loops that the planet supposedly traveled on just to make the math match his observations. Exactly. To make the geocentric math work, his model implied that the moon would sometimes come twice as close to the Earth as at other times, which we know would cause catastrophic tidal waves. It's completely absurd. Right. The output looked correct to a casual observer, but the internal logic was fundamentally broken.
5:22And the parallel to standard NextToken AI is just striking. It really is. Ptolemya's model lacked a simple, generalizable underlying structure. Copernicus later provided that structure simply by putting the sun in the center, which eliminated the need for all those crazy epicycles. Right. In learning theory, there is this core principle that simpler explanations of observations generalize better. It's basically Aukamara's razor applied to machine learning. Standard AI, with its infinite lookback memory, completely lacks the pressure to form a simple compact explanation of the world. It just builds Ptolemy's chaotic epicycles over and over again.
6:00Exactly. So if giving an AI this massive uncompressed memory leads to lazy shortcut learning in epicycles, how do we force it to build a true compact model of the world? Well, that is exactly where NextLatentPrediction comes in. We change the fundamental goal of the training process. Normally, you just train the model to output the next token, the next word. But in NextLayer, we add a very lightweight secondary model inside the architecture. Okay, what is that? It's called a latent dynamics model, and it's really just a simple three-layer neural network. Three layers, so that's tiny compared to the main multi-billion parameter network.
6:36Oh, it's a fraction of the size. But its job is crucial. It forces the main transformer to predict its own next hidden state based on its current state and the next action or token. We should probably define hidden state here for a second. Yeah, good call. When we say hidden state, we mean the actual mathematical configuration of the AI's internal network at a specific moment in time. So instead of just predicting the external word like Apple, the AI is now being forced to predict the mathematical shape of its own brain exactly one millisecond before it says the word Apple. Right. And by forcing the model to predict its own internal transitions, we're forcing it to build what we call belief states.
7:15Belief states. Yeah, this is a concept borrowed from reinforcement learning. A belief state is a perfectly compressed summary of the past. It is a sufficient statistic, meaning it contains absolutely everything you need to know about the past to accurately predict the future. With no wasted space. None. No wasted space, no unnecessary details. NextLat mathematically forces the Transformers' messy hidden states to converge into these perfectly compact belief states. Wait, hold on. Let me push back on this mechanic for a second. Sure, go ahead. If we are forcing the AI to update its internal state step by step, carrying a single compressed memory forward through time, isn't that just turning it into a recurrent neural network?
8:00An RNN? It sounds a lot like it, yeah. Because I thought the whole industry abandoned RNNs years ago. They process data sequentially, one step after another, which makes them incredibly slow to train compared to transformers. Are we just taking a massive step backward in efficiency here? So that is actually the genius of the Nexlat design. We are getting the benefits of an RNN without the sequential bottleneck. How is that possible? Because Nexlat retains the massive parallel training efficiency of a standard transformer. You still train all the tokens across the whole sequence at the exact same time, just like normal.
8:32But we artificially inject a recurrent inductive bias using that tiny three-layer sidekick model. Okay, let's unpack inductive bias for a second. That basically means the mathematical pressure or the rules of the game applied to the AI during training, right? Precisely. We aren't changing the physical transformer architecture to make it sequential. We are just changing the penalty and reward math applied to it during that parallel training. The AI is still reading the whole book at once, essentially. But it only gets a good grade if step one logically flows into step two and step two logically flows into step three.
9:08So it enforces the discipline of step-by-step compression, but without losing the processing speed of a transformer. Okay, that makes sense. We're giving the AI the organizational skills of an RNN, but keeping the blazing speed of a transformer. But it's one thing to say this creates a world model or a belief state in abstract theory. How do we prove the AI actually understands the space it's navigating? Well, there was a brilliant experiment done using taxi rides in Manhattan that visualizes this perfectly. Right. This is such a cool test. It really is. It's one of the best stress tests for AI spatial reasoning.
9:44In this experiment, AI models were trained on 91 million sequences, which is about 4.7 billion tokens, of purely random taxi rides. And to be clear, this wasn't a map. It was just text. Right. Just lists of intersections like pick up at point A, turn left, go straight, turn right, drop off at point B. Just a massive text file of turn-by-turn directions, no visual data whatsoever. None at all. Now, when you train a standard NEXT token transformer on this data, it performs beautifully on the surface. It gets a perfect 100 % accuracy on predicting the next legal turn. So if you asked it what to do at Fifth Avenue and 42nd Street, it will never give you an illegal move.
10:20Exactly. Which makes it look like it completely solved the problem and deeply understands New York City. You would absolutely think so. But researchers took the internal math of that standard AI, its hidden states, and mathematically visualized it as a 2D map to see what the AI's internal representation of the city actually looked like. And it is a complete disaster. Oh, it's a mess of impossible orientations. You have roads flying over other roads, intersections that don't connect in physical space, parallel streets crossing each other. So the AI has 100 % accuracy, not because it understands the grid of New York, but because it just memorized 91 million local shortcuts.
10:59Exactly. The S. Ptolemy is universe again. The output looks right, it tells you to turn left, but the internal logic requires these impossible epicycles to make the math actually work. But when you apply Nexlat to the exact same text data, the result is stunning. Nexlat builds an internal mathematical map that closely mirrors the real Manhattan grid. The streets are parallel, the intersections connect logically. The errors are incredibly sparse and localized. And what's wild is that we can actually measure this difference mathematically without even looking at a visual map. Right, using a metric called effective latent rank.
11:34Effective latent rank. Effective latent rank is essentially a measure of dimensionality. Think of it as how many independent coordinates or dimensions the AI needs to describe its world. So like if you want to describe where a box is in a room, you only need three dimensions, length, width, and height. Exactly. And if a computer system is using 160 dimensions to describe that same box, It is horribly bloated and doesn't actually understand 3D space. It's just overcomplicating the universe. Heavily overcomplicating it. And the standard GPT model had a massive effective latent rank of 160.1 on the Manhattan data.
12:09It was wildly inefficient. NextLat, however, had a rank of 52.7. Wow. It's basically a third of the size. Yeah. It successfully compressed all that chaotic random text data of taxi turns into a highly efficient, structurally accurate representation of the city. By forcing it to predict its own internal states, Nexlat forced the AI to build a real functioning map in its mind. Okay, so building a physical map of New York makes perfect sense. But what about abstract spaces? Like if it can plan a route through Manhattan without roads overlapping, can it plan a route through a complex math equation without hallucinating?
12:46That brings us to the countdown math task. And this is a brutal test for abstract reasoning. It's essentially the game of 24. Right. The AI is given a list of random numbers and a target number, and it has to use basic arithmetic like addition, subtraction, multiplication, division to combine the random numbers and hit the target exactly. And standard NextToken models suffer heavily on this task from a phenomenon called the regretful compromise. Yes. You know what that sounds exactly like? A bad improv comedian. A bad improv comedian. How so? Well, if you've ever watched Bad Improv, a comedian will start a scene, they'll paint themselves into a narrative corner, suddenly realize they're completely trapped, and to end the scene, they just shout, and then aliens abducted them.
13:29Oh, man. That is the perfect analogy for what standard AI does with math. Because the standard AI does the exact same thing. It starts doing calculations step by step. It gets to the very end of the equation, realizes it is nowhere near the target number, and because it can't go backward to fix its early mistakes, it just hallucinates an invalid final equation. It essentially says 14 plus 2 equals 85, just to pretend it succeeded and end the scene. Exactly. It completely lacks look-ahead planning. It cannot see the consequences of its early steps. But because Nexlat forces the model to predict its future latent states, its future thoughts, it enforces global planning.
14:08It literally has to think ahead to maintain consistency in its own internal math. And the numbers on this are wild. Nexlat drastically reduces that regretful compromise error. On the countdown task, the validity of the final equation generated by the AI jumps from 42.3 % in standard models all the way to 54.8 % with Nexlat. It avoids what we call the Clever Hans cheat. Clever Hans, the horse, right? Yeah, Clever Hans was a horse that appeared to do math but was actually just reading the subtle body language of his trainer to know when to stop tapping his hoof. Standard AI does the same thing. It exploits local short-term shortcuts instead of learning the actual rules.
14:47And we see Nexlat's superiority validated again in the PathStar graph task. This is a complex logic puzzle involving routing through a star-shaped graph. Standard models completely fail as the graphs get larger. But Nexlat maintains a near 100 % solve rate, even on the largest, most complex graph sizes. It is demonstrably better at holding a long-term plan in its head. Yes, exactly. Okay, so we've seen it map a city. We've seen it plan ahead in math. But there is a fascinating quirk about this architecture that actually seems to defy the theoretical mathematical limits of what transformers are supposed to be able to do.
15:24Ah, yes. This is the A5 anomaly. This anomaly is the part that literally keeps computer scientists up at night. The A5 word problem is a highly complex state tracking task. It involves tracking even permutations of five elements. Let's ground that. What does that actually look like for the computer? Okay, think of it like a complex cup and ball magic trick. If a magician is swapping five cups around, you cannot process the final location of the ball just by looking at the whole scene at once. Right, you'd have to watch the movements. You have to watch every single move, sequentially, one after another, to know where the ball ends up.
15:59The problem belongs to a complexity class that fundamentally requires deep, step-by-step, sequential tracking. And a standard, shallow transformer processes everything in parallel, all at once. It physically doesn't have the step-by-step architectural depth to follow the cup and ball trick. Exactly. It is mathematically restricted. By definition, a shallow transformer fundamentally cannot solve the A5 problem for long sequences. It just doesn't have the architecture to do the computation. Wait, if the architecture itself is fundamentally blocked from doing this, that means giving it more training data won't help, right?
16:37You could feed it a trillion more examples and it still would fail. It will always fail. It hits a mathematical ceiling. So researchers trained a NEXLAT model on short sequences of 12 tokens for this A5 problem. Remember, NEXLAT is two things. It's the big transformer and the tiny three-layer recurrent latent dynamics model riding alongside it. Okay, so they train it on 12 tokens. Right. And then when they tested the system on much longer sequences, like 36 tokens, the big transformer failed completely, as expected. It hit its ceiling and dropped to zero accuracy. Or the twist. The twist. The lightweight three-layer sidekick model, the late Dynamics model, when tested on its own on those 36 token sequences, succeeded with over 95 % accuracy.
17:20I am still baffled by this mechanism. Yeah. How is it possible that the tiny sidekick model surpasses the massive master network that is literally providing its training data? It's crazy, right? How does the student become a master of something the teacher physically cannot do? Well, if you look at the mechanics of the training process, it synthesizes beautifully. The big transformer acts as an incredibly efficient, highly parallel lens. Within its short 12-token window, it creates a pristine learning environment. It organizes the data. Yes. It organizes the chaotic data so cleanly that the tiny recurrent model can observe the pure, generalized, sequential rules of the universe.
18:02The sidekick isn't limited by the transformer's architectural ceiling. It just uses the transformer to perfectly focus the data. It completely bypasses the architectural limits of the host network. So the big brain can't do the math, but it organizes the room so perfectly that the tiny brain can figure out the rules and then keep applying them forever. It really challenges our assumptions about how we need to scale AI moving forward. We've assumed bigger is always better. This proves that's not necessarily true. That's wild. And beyond making AI smarter, more logical, and theoretically fascinating, this latent prediction trick actually has a massive practical benefit for everyday use.
18:40Huge benefit. It makes the AI significantly faster. Generation speed is a huge bottleneck right now. Standard autoregressive generation, you know, spitting out text one word at a time is painfully slow. Very slow. Every time the AI wants to output a single word, it has to fire up and run its entire massive neural network. Right. And to speed this up, engineers currently use methods called multi-token prediction. It tries to guess a fixed number of words ahead. But those methods operate in the token space, the actual dictionary of external words. Yeah. They are rigid. If you train a model to guess two words ahead, it can only ever guess exactly two words ahead.
19:17Trying to predict multiple specific words at once often degrades the overall quality of the underlying logic. It just gets confused. But Nexlat has a massive advantage here because it doesn't guess words. It guesses thoughts. It operates in the latent space. We call this variable length self-speculative decoding. Because Nexlat has that latent dynamics model, it can run recursively inside its own internal math. Without firing up the big brain. Exactly. It can take its current state and just use the tiny three-layer sidekick to ask, what's my next state? And the state after that, and the state after that.
19:50It can spin out a whole chain of future thoughts without ever turning on the heavy, slow main transformer. It's just quietly daydreaming the next 10 steps in a fraction of a millisecond before it bothers to translate those thoughts into actual words. It adaptively drafts a flexible number of tokens ahead based on how confident it is. In testing, it was comfortably daydreaming up to 10 tokens out into the future. So what is the tangible benefit when you're sitting at your computer using one of these tools? It translates to a massive savings in time and compute power. When researchers tested a 1.3 billion parameter model trained on 100 billion tokens, Nexlat accelerated text generation by up to 3.3 times compared to standard generation.
20:32More than three times as fast. Three times faster with zero loss in accuracy. In fact, it preserved the quality of the logic better than any of the traditional multi-token guessing methods. Wow. Let's bring this all together. Today we looked at how the AI industry has essentially been building Ptolemy's universe. Models that memorize complex shortcuts and epicycles because they have unlimited memory and no pressure to actually understand anything. But by introducing Nexlat, by simply forcing the AI to predict its own internal states rather than just the next external word, we force a shift from lazy memorization to compact belief states.
21:09It builds a true, accurate map of the environment. It plans ahead to solve math problems without painting itself into a regretful corner with bad improv. And because it can strain together its own thoughts in the latent space, it generates text over three times faster. It is a fundamentally more elegant way to build an intelligence. Which brings me to a final thought for you to chew on today. Think back to that A5 anomaly we unpacked, where the tiny auxiliary inner network learned logical rules that completely exceeded the cognitive limits of its massive parent network. Yeah, the student surpassing the master.
21:42Exactly. So if a tiny sidekick can become smarter than the master, what happens if we let those tiny inner world models detach entirely? That is the exact question pushing the boundaries of the field right now. Could the future of AI not be about building bigger, massive, multi-trillion parameter transformers that take up entire data centers? What if instead the future is cultivating ecosystems of these hyper-efficient, highly logical, tiny latent minds that just talk to each other? The massive transformer just becomes the incubator. The incubator for true reasoning. Thank you for joining us on this deep dive.
22:19As you go about your week and interact with these tools, keep questioning how the systems you use actually think. because as we've seen today, sometimes the biggest breakthroughs come not from giving an AI more data, but from forcing it to understand the data it already has. Until next time.
From the publisher
This paper introduces Next-Latent Prediction (NextLat), a novel training framework designed to help Transformer models learn more compact and generalizable internal world models. Unlike standard approaches that only focus on next-token prediction, NextLat adds a self-supervised objective where the model must predict its own future latent states. This method encourages the formation of belief states, which are efficient summaries of past information that improve the model’s ability to reason, plan, and generalize. Theoretically, this injects a recurrent inductive bias into the architecture without sacrificing the parallel training efficiency or speed of the original Transformer. Empirically, NextLat demonstrates superior performance in world modeling and long-horizon reasoning compared to traditional baselines. Furthermore, the learned latent dynamics enable variable-length self-speculative decoding, which can accelerate inference speeds by over three times.




