In short
Pathway’s post-transformer “BDH” architecture is tested on the Sudoku Extreme benchmark (~250,000 hardest Sudoku puzzles) and reportedly crushes it, exposing a weakness in transformer LLMs on constraint satisfaction/search.
Guests
Adrian Kosovsky, Pathway co-founder and Chief Scientific Officer; previously discussed transformer reasoning limits (attention head dimension ~1,000 ceiling).
Key claims
BDH solves Sudoku Extreme with 97.4% accuracy; leading LLMs (O3 Mini, DeepSeek R1, Claude 3.7 Sonnet) score effectively 0% on the same task. BDH reasons in a larger latent reasoning space without chain-of-thought text, using sparse positive activations (~5% neurons firing), a state-based model (no full attention over the sequence), and continual learning (reach advanced beginner in ~20 minutes of a new game).
Notable examples
Sudoku as a proxy for backtracking constraint satisfaction needed in medicine, law, and operations planning.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOSudoku as an AI Benchmark
0:45 to 2:00
Explains why Sudoku serves as a challenging test for AI models.
“Extreme that consists of roughly 250 ,000 of the hardest Sudoku puzzles available, and the results are striking.”
Limitations of Current Transformers
2:00 to 3:55
Discusses the inefficiencies of transformer architecture in solving Sudoku.
“Large language models turn every problem into text, and then solve it by predicting the next token one step at a time.”
BDH Architecture Overview
3:55 to 6:00
Introduces the Baby Dragon Hatchling architecture and its internal reasoning capabilities.
“The analogy Pathway uses is a chess grandmaster playing 20 simultaneous games with her eyes closed.”
Technical Innovations of BDH
7:10 to 8:40
Describes the technical features that differentiate BDH from transformers.
“And now here's a detail that should get the attention of anyone thinking about the economics of AI.”
Implications of BDH's Success
8:40 to 10:00
Examines the broader implications of BDH's performance in AI architecture.
“When Adrian was on the show, we discussed how the architecture has been demonstrated at about a billion parameter scale, comparable to GPT-2 from now many years ago.”
Transcript
Automatic transcript. May contain errors.0:00This is episode number 978 on the post-transformer architecture crushing Sudoku Extreme.
0:08Welcome back to the Super Data Science Podcast. I am your host, Jon Krohn. Today's topic is a puzzle.
0:15Jon Krohn:Literally, we're going to be talking about Sudoku and why a game that millions of people around the world knock out every morning over their coffee is exposing a fundamental weakness in the most powerful AI models on the planet. So here's what happened. Pathway, a company I've had on the show before back in episode number 929 most recently. In that episode, I had their co-founder and chief scientific officer Adrian Kosovsky in the studio to talk about their post-transformer architecture called BDH. Pathway recently published a research article alongside a benchmark called Sudoku Extreme that consists of roughly 250 ,000 of the hardest Sudoku puzzles available, and the results are striking.
0:54Jon Krohn:Pathway's BDH architecture solved these extreme Sudoku puzzles with 97.4 % accuracy. The leading large language models, and we're talking O3 Mini, DeepSeq R1, Claude 3.7 Sonnet, scored effectively 0%, not low, 0, on a task that any reasonably practiced human can do with a pencil and some patience. Now, before we dig into why this matters and and what BDH is doing differently, let's quickly recap what makes Sudoku such an interesting test for AI. Sudoku is a constraint satisfaction problem. In a 9x9 grid, every move has to satisfy multiple rules simultaneously. The numbers 1 through 9 must appear exactly once in each row, once in each column, and once in each 3x3 box.
1:40Jon Krohn:A completed grid is trivially easy to verify, but producing that grid from a partially filled board requires searching through interacting possibilities without violating the rules. It's that combination of search, constraint, and backtracking that makes Sudoku a clean test of whether a system can reason under constraints rather than merely describe them. And this is exactly where today's transformer-based LLMs start to struggle. Here's the core issue. Large language models turn every problem into text, and then solve it by predicting the next token one step at a time. That works brilliantly when language is the right medium for a task—writing an email, summarizing a report, generating code—but But Sudoku doesn't live in language.
2:19Jon Krohn:Forcing it into a chain of text is painfully inefficient because the transformer architecture processes information token by token with a limited internal state at each step. The data that represent the model's thinking, its latent space, are constrained to roughly a thousand floating-point values per token, and critically, each decision gets locked in as text is generated. Transformers can't hold multiple candidate strategies in parallel, they don't have the ability to step back and reconsider earlier moves without verbalizing every intermediate thought. Now, if you've been a listener for a while, this constraint on the transformer's internal representation might ring a bell.
2:53Jon Krohn:When Adrian Kosovsky was on the show, he made this exact point. The size of the attention head vector dimension in a transformer has essentially stopped scaling, even as models get larger and larger in terms of parameters. It's converged to around a thousand dimensions, which means all concepts that the transformer works with have to be mapped into a vector space of about that size. Adrian described this as a fundamental ceiling on the transformer's capacity for nuanced reasoning, and during his interview on the show last year, I found that argument compelling. The Sudoku benchmark data now provide concrete, quantitative evidence for it.
3:27Jon Krohn:So what is Pathways BDH architecture doing differently that allows it to crush this Sudoku benchmark? BDH, which by the way stands for Baby Dragon Hatchling, I probably should have mentioned that earlier in this episode, is a playful name for the first model in Pathway's baby dragon family. And it's what Pathway describes as a native reasoning model. BDH maintains a much larger internal reasoning space, what they call a latent reasoning space, that isn't constrained to verbalizing every thought as text. The analogy Pathway uses is a chess grandmaster playing 20 simultaneous games with her eyes closed.
4:01Jon Krohn:She's not whispering each move to herself in words, she's internalized the patterns and can navigate the search space seamlessly. BDH is designed to enable that kind of internalized reasoning in a machine. On this podcast, I'm always going on about how Claude Code is mind-blowing, but now Claude Cowork is making my jaw drop as well. For example, I recently wanted to quantify how healthy my sales pipeline is for my AI consulting business. I simply asked Claude to estimate my sales for the coming quarter, and it brought info from relevant Google Sheets and my Gmail to create a professional spreadsheet of clients with estimated revenue for each one.
4:34Whoa, this might have taken me a day. Instead, it was done flawlessly with Claude Cowork in minutes. Claude is the AI for minds that don't stop at good enough. It's the collaborator that actually understands your entire workflow and thinks with you. Whether you're debugging code at midnight or strategizing your next business move, Claude extends your thinking to tackle the problems that matter. Ah, and you'll appreciate that I can ask Cowork to show me data, such as my sales spreadsheet, and it provides an interactive chart right in the conversation. For problems worth solving, get started with Claude at Claude.ai slash superdata.
5:06That's Claude.ai slash superdata. And check out Claude Pro, which includes access to all of the features mentioned in today's episode. Claude.ai slash superdata.
5:17Jon Krohn:There are a few key technical ingredients here that are worth understanding. First, BDH uses what are called sparse positive activations, meaning that at any given time, only about 5 % of the artificial neurons in the network are firing. This is radically different from a transformer where you're flowing information through essentially all the neurons, dense activation, on every single input. Mixture of experts models are, yeah, somewhere in between, but we're not covering that in this episode. Anyway, as Adrian explained when he was on the show, this sparse activation is far more biologically plausible.
5:49Jon Krohn:It's much closer to how a human brain actually works. We have around 80 to 100 billion neurons and roughly 100 trillion synaptic connections, but only a tiny fraction are active at any given moment. If our brains were densely activated the way a transformer is, we wouldn't have enough energy to power them. Second, like the best-known post-transformer architecture Mamba, BDH is a state-based model, meaning it doesn't rely on the standard transformer attention mechanism that looks back through your entire input sequence to find relevant context. Instead, it maintains and updates an internal state, somewhat analogous to how biological neurons continuously update their synaptic connections based on what they're processing.
6:26Jon Krohn:This is closely related to Hebbian learning, the foundational neuroscience principle that neurons which fire together wire together. That's how we learn, how biological animals learn. And Adrian discussed this at length back in episode number 929, and he explains how BDH draws inspiration from these biological mechanisms. And third, and this is particularly relevant to the Sudoku results, BDH achieves what Pathway calls continual learning. It learns from every interaction and internalizes that learning over time. According to Pathway, BDH can pick up the rules of a new game and reach an advanced beginner level in as little as 20 minutes, then improve through repeated play.
7:03Jon Krohn:This is a far cry from a transformer which has a fixed set of weights after training and relies on in-context learning or chain of thought prompting to tackle novel problems. And now here's a detail that should get the attention of anyone thinking about the economics of AI. BDH achieves its 97.4 % accuracy on Sudoku Extreme at materially lower cost than the leading LLMs achieve their near-zero scores. Because BDH reasons in its internal latent space rather than generating long chains of text, it doesn't burn GPU cycles verbalizing every intermediate step. Pathway reports that the cost is roughly 10 times lower compared to O3 Mini, DeepSeek R1, and Sonnet 3.7, with no chain of thought required at all.
7:45Jon Krohn:Now, you might be thinking, okay, Sudoku is interesting, but is this just a parlor trick? I think not, and here's why. The ability to solve Sudoku is really a proxy for the ability to navigate constraint satisfaction problems more broadly, holding multiple possibilities in parallel, backtracking when needed, and converging on solutions that satisfy all rules simultaneously. These are precisely the skills needed for countless real-world challenges in medicine, law, operations, planning, and many other spaces. domains where you're balancing competing constraints under uncertainty. A system that can reason through these spaces natively, rather than forcing everything into a text-based chain of thought, could eventually do more than summarize information.
8:25Jon Krohn:It could help generate strategy. Pathway, indeed, calls this generative strategy, looking at a problem, understanding the constraints, and creatively proposing what should be done, rather than merely remembering what has been done before. That's an exciting frontier. Now, I do want to be balanced here. BDH is still early. When Adrian was on the show, we discussed how the architecture has been demonstrated at about a billion parameter scale, comparable to GPT-2 from now many years ago. And Pathway hasn't yet released a massive frontier scale model. Adrian was clear that there's nothing stopping them from scaling much larger, but their current focus is on entering the reasoning model space, where this architecture's advantages are most pronounced.
9:05Jon Krohn:That said, I find the Sudoku extreme results that I've covered in today's episode compelling, as evidence that the transformers limitations are real, and that alternative architectures can address those limitations. The data are clear. 0 % from transformer-based architectures versus 97.4 % from BDH is not a marginal difference, it's a categorical one. And if you combine this with the theoretical arguments Adrian laid out on the show previously about sparse positive activations, about the biological plausibility of BDH, about its potential for lifelong learning and reasoning over long time horizons, the picture emerges that this is an architecture that could meaningfully push AI capabilities beyond what transformers alone can achieve.
9:48Jon Krohn:All right, we've got links in the show notes to Pathways' full research article on the Sudoku Extreme Benchmark, as well as to the previous BDH paper on Archive, so you can dig into the technical details to your heart's content. The transformer has been the undisputed king of AI architectures for the better part of a decade, and it's exciting to now see credible challengers emerge that are fundamentally rethinking how machines reason. All right, that's it. If you enjoyed today's episode or know someone who might consider sharing this episode with them, leave a review of the show on your favorite podcasting platform or on YouTube, tag me in a LinkedIn post with your thoughts.
10:22Jon Krohn:And if you haven't already, be sure to subscribe to the show. Most importantly, I just hope you'll keep on listening. Until next time, keep on rocking it out there and I'm looking forward to enjoying another round of the Super Data Science Podcast with you very soon. Thank you.
From the publisher
A game millions of people solve over morning coffee is exposing a fundamental weakness in today’s most powerful AI models. In this Five-Minute Friday, Jon Krohn breaks down Pathway’s new Sudoku Extreme benchmark, roughly 250,000 of the hardest Sudoku puzzles available and why leading LLMs like o3-mini, DeepSeek-R1, and Claude 3.7 Sonnet scored effectively zero percent, while Pathway’s post-transformer BDH architecture achieved 97.4% accuracy at a fraction of the cost. Listen to the episode to find out what BDH is doing differently, why Sudoku performance matters far beyond puzzles, and what this means for the future of AI reasoning.
Additional materials: www.superdatascience.com/978
Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information.




