In short
Neural computers (NC/CNC) that replace the classic hardware+OS stack with an AI model as the “runtime,” hallucinating a working command-line or GUI interface from learned screen behavior and user I/O traces.
Guests
No guest names or bios appear in the transcript; it’s a single conversational host format (“Thanks for having me,” “Today we are looking…”).
Key claims
Video-generation models can be trained to render UI and respond to input via latent “runtime state” (H_t) plus action encoders; however they lack symbolic reasoning and drift over time.
Notable examples
WAN 2.1-based prototypes: C-Elegend (command line) and GWorld (GUI). Text resolution issue: tiny 6-pixel fonts fail; 13-pixel terminal fonts work. Control: raw keystroke multi-hot vs meta-action intents; cursor handled via separate cursor layer + mask. Arithmetic probe: 10+15 scored 0% on base 2.1, 4% on NC prototype (Sora 2 scored 71%). Four CNC pillars: Turing-complete, universally programmable, machine-native semantics, behavior consistency (avoid “Excel becomes calendar” drift).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe Concept of Neural Computers
0:45 to 1:40
Discussion on how AI models are evolving to redefine computing environments.
“They are, and that classic hardware stack has been the foundation of digital life for decades.”
Comparing Classical and Neural Computers
1:40 to 2:32
Exploration of the differences between classical computers and neural computers, using metaphors.
“But agents are basically just clicking real buttons on a real operating system.”
Prototypes of Neural Interfaces
2:32 to 3:47
Overview of prototypes built to simulate command line and graphical interfaces using AI.
“the processor, and the interface all at once.”
Data Sets for Training AI Models
3:47 to 4:52
Insight into the datasets used to train AI models on operating system physics.
“They trained it on two massive, distinct datasets.”
Challenges in Rendering Text and Control
4:52 to 5:35
Discussion on the difficulties faced in rendering text and managing user input.
“It is, but it still runs into hard physical limits.”
Input Methods for AI Models
5:35 to 6:49
Explanation of different methods for input processing in neural computers.
“Right, so if the keyboard is just a stream of characters, that's somewhat manageable.”
The Complexities of Action Encoding
6:49 to 8:06
Delving into action encoding techniques used for AI to interpret user actions.
“That sounds like it would completely overwhelm the model over a long session.”
Limitations of Neural Computers
8:06 to 9:21
Highlighting the limitations of neural computers in performing logical tasks like arithmetic.
“The action is just absorbed into the ongoing physics of the system.”
Roadmap to Complete Neural Computers
9:21 to 10:39
Discussing the essential pillars required to achieve a fully functional neural computer.
“You type 10 plus 15, hit enter, and the computer hallucinates the number 30.”
Implications of Human Interaction
10:39 to 12:15
Exploring the consequences of training AI on human behavioral data.
“So with all these hurdles, the math failures, the memory drifting, why keep trying?”
Transcript
Automatic transcript. May contain errors.0:00Imagine opening your laptop, right? You type 10 plus 15 into the command line, hit enter, and your computer confidently spits out 30. Which is just wild. It is. And not because the processor made some weird math error, but because your operating system isn't an operating system at all. Well, it's a video generator hallucinating what a computer screen looks like. Yeah. Welcome to this deep dive. Thanks for having me. Today we are looking at a mind-bending paradigm shift in how we even define computing. Usually when we talk about the device you're using right now, there's this expectation of structure.
0:38Right, a comforting hardware stack. Yeah. You have the processor doing the math, the RAM remembering the state of the machine, and the screen and keyboard for input and output. They're distinct physical modules. They are, and that classic hardware stack has been the foundation of digital life for decades. But recent AI development is pushing us into a space where those boundaries are completely dissolving. Like completely vanishing. Exactly. We're moving towards something called a neural computer or an NC. And the ultimate goal of this field is the completely neural computer, the CNC. Which sounds like science fiction.
1:11It really does. We're exploring how AI models are evolving way beyond just, you know, generating text or making deep fake videos. They are literally attempting to become the computer's entire runtime environment. Okay, let's unpack this because it is a massive conceptual leap. To understand how radical this is, we have to map out the current landscape. Classical computers run explicit programs, right? Right, governed by a rigid operating system. And then we have AI agents. But agents are basically just clicking real buttons on a real operating system. They act through that external software stack.
1:48Yes, and world models just predict how an environment changes. But neural computers propose something entirely different. In a neural computer, the model is the running computer. So think of a classical computer like a traditional restaurant kitchen, right? You have a separate prep station. You've got the grill and the plating area. Everything is separate. That's a great way to look at it. But a neural computer is more like a sci-fi microscopic organism that just ingests raw ingredients and instantly morphs them into a finished meal totally internally. That is exactly it. In technical terms, that organism's internal state is what we call a latent runtime state.
2:26In the math driving these models, they denote this as a variable, H sub T. This single variable, this massive matrix of numbers, acts as the working memory, the processor, and the interface all at once. That's just crazy to think about. It holds the entire universe of the computer for that specific millisecond. This isn't just a smarter layer on top of Windows or Mac OS. It's an attempt to replace the operating system itself with a neural network that just updates and renders frames based on inputs. Here's where it gets really interesting. If a neural network is going to act as your entire computer, it doesn't have a graphics card drawing windows and fonts.
3:03It has to render an interface entirely from scratch. Right, and to test this, researchers built prototypes on top of a state-of-the-art video generation model called WAN 2.1. But not for making movies. No. Instead of generating cinematic video, they trained it to generate a functioning computer interface. They created two main prototypes. One is C-Elegend, which models a command line interface. The classic black screen with scrolling text. Exactly. And the other is GWorld, which models a graphical user interface, like your standard desktop. But to build a GUI from scratch, I mean, you can't just feed an AI a textbook on coding.
3:40You'd have to feed it thousands of hours of people just staring at and clicking on their actual screens, right? Which is precisely what they did. Let's look at the command line prototype CLagin. They trained it on two massive, distinct datasets. The first is a general dataset using open-ended Asinema terminal recordings. And for you listening, Asinema recordings aren't video files. They are lightweight text recordings of terminal sessions. So this general data set is incredibly messy. It's very human. Yeah, it's full of typos, weird pauses where somebody's trying to remember a command, hitting backspace, getting an error.
4:12It's raw workflow. Then you have the second data set, which they called Clean. This used deterministic VS scripts. Like perfectly paced replays. Exactly. They are dockerized, highly controlled programmatic executions of things like package installations, no wasted frames, and for the desktop prototype, GYWorld, they captured massive amounts of desktop RGB video synchronized with mouse and keyboard logs. So because of all that data, the AI has basically learned the physics of an operating system. When you watch the output, it knows how text is supposed to scroll up and disappear. It knows how a long prompt wraps around.
4:50Right. It is hallucinating the physical laws of a computer monitor. It is, but it still runs into hard physical limits. A critical technical detail here involves text resolution. The video model relies on a VAE, a variational autoencoder. Which compresses the image to save processing power. Yes, and then blows it back up. The researchers found that the VAE really struggles to reconstruct tiny 6-pixel fonts. So the tiny text just turns to mush. It becomes completely illegible. You get localized blurring even if the rest of the screen looks pristine. but, and this is key, the model perfectly handles the standard 13-pixel terminal font.
5:26It's a brilliant example of how practical design choices, like locking your training data to a 13-pixel font size, are critical for AI training. If the user can't read the text, the illusion breaks down. Right, so if the keyboard is just a stream of characters, that's somewhat manageable. But watching a video of a computer is not the same as using one. How does it know what to do when you actually type or move the mouse? That brings us to the core challenge. Control. The model has to learn short horizon control. If you hover over a menu, it needs to highlight. If you click, it opens. Exactly. To bridge the gap between your physical hands and the hallucinated screen, they designed specific action encoders.
6:07Let's break that down because they tried two different approaches for keyboard input. The first is the raw action encoder. This method feeds every single solitary keystroke into the model as an isolated event using a 169-dimension multi-hot stream. Wait, a 169-dimension multi-hot stream? I know, it sounds complex. For the listener, just imagine a giant piano with 169 keys where multiple keys can be pressed at the exact same millisecond, like holding shift, control, and C together. That's a perfect analogy. So if you type Ls-L, it sees the L, the S, the space as isolated piano chords. And it has to infer what that sequence means purely from that fragmented stream while keeping the video stable.
6:52That sounds like it would completely overwhelm the model over a long session. Wouldn't it make more sense to bundle those strokes into a single command? Yes, and that leads directly to their second approach, the meta-action encoder. Instead of a barrage of key presses, it bundles them into structured intents. Like treating Ls-L as one specific tool. Precisely. It's much easier for the model to process. Okay, so that's typing. But a mouse cursor is a totally different problem. You're dragging a little white arrow over a hallucinated image, which should technically smear the pixels behind it, right?
7:22You hit on a major architectural hurdle. If you ask a standard video model to move a cursor across a complex background, it often corrupts the desktop image underneath. Oh, because it gets confused about what's foreground and what's background? Exactly. To fix this, they introduced explicit visual cursor supervision. They isolated the cursor onto its own layer, feeding the model the cursor's visual foreground alongside a mask. A mask. So they tell the model, basically, only update these specific pixels shaped like an arrow and ignore the background. That's the mechanism. What's fascinating here is the architectural elegance.
7:59These action signals aren't explicit tokens sitting in a traditional transformer network. Wait, really? Where do they go? They are implicitly injected directly into that latent space, the H sub T we discussed earlier. The action is just absorbed into the ongoing physics of the system. So if it's absorbing all this, moving the masked cursor, rendering scrolling text, does this mean the AI actually understands the tools it's using? Or is it just acting like a very convincing parrot? That is the million dollar question. And it brings us directly to the system's absolute biggest weakness, symbolic reasoning.
8:33Right. What happens when it has to do hard logic? Well, researchers ran something called an arithmetic probe. They tested the command line prototype with 100 basic math problems in the hallucinated terminal, things as simple as 10 plus 15. The results illustrate a massive limitation. The base 2.1 model scored an abysmal 0%. Wow. And the neural computer prototype, it scored 4%. Wait, a plastic solar-powered calculator from 1975 gets 100 % on this test, but this cutting-edge neural computer gets 4%. It's staggering, isn't it? Now, to be fair, they did test a different model, Sora 2, which scored 71%.
9:13So there are vast differences in system-level advantages, but for this specific prototype, it failed completely. It's wild to imagine. You type 10 plus 15, hit enter, and the computer hallucinates the number 30. Because it's not doing math, it just thinks a 30 looks plausible there. This raises an important question about the fundamental nature of these models. Current video models are probabilistic pattern matchers, not symbolic calculators. They lack native symbolic reasoning. Right. They see a visual canvas of character patterns, not rigid logic. So, if it completely fails at basic math, how do we ever get to a completely neural computer that we can actually rely on?
9:52There is a clear roadmap to get from these early prototypes to a true CNC. They define four essential pillars. First, it must be Turing-complete. Second, universally programmable. Meaning you can install routines and call them later. Exactly. Third, machine-native semantics. It needs its own optimized way of operating. And the fourth pillar is behavior consistency. Behavior consistency. Let's ground this for the listener. Imagine you leave your computer running an Excel spreadsheet overnight. You wake up and your financial spreadsheet has slowly drifted into looking like a calendar app. Right.
10:26Catastrophic forgetting. It's literally forgetting what it is. A true CNC must not change its function during ordinary use unless it is explicitly updated. Right now, these models drift. Over long horizons, the rules of their universe start to warp. So with all these hurdles, the math failures, the memory drifting, why keep trying? Why not just stick to the classic hardware stack we all know? Because of the data advantage. High-quality, human-written code is hard to produce. But think about how we interact with screens. Every click, every scroll, every keystroke. It's an I.O. trace. We are constantly generating training data just by doing our jobs.
11:03If we connect this to the bigger picture, this data asymmetry changes everything. We don't need to write code to teach the computer. The sheer volume of our everyday digital behavior will train the runtime environments of the future. So what does this all mean? Imagine a near future where your computer isn't built on a static Apple or Microsoft operating system. It's a fluid, learned environment that literally adapts its interface to how you specifically work. It becomes a highly adaptable space that reshapes itself around you. We have taken quite a deep dive today. We started with a clunky hardware stack and explored how this field is smashing that into a unified neural runtime state.
11:40Yes, hallucinating the physical behavior of a UI. Exactly, down to the scrolling text to mask cursors. But we also grounded it in reality. These models are probabilistic painters, not symbolic calculators. Keeping them consistent is a massive challenge. But the path is being paved with our own daily interactions. And that leaves me with a final thought for you to mull over today. If the ultimate, completely neural computer is trained entirely on human I.O. traces, our actual chaotic clicks and rapid keystrokes, will it eventually inherit our digital bad habits? That's a scary thought. Right. If humanity's data is full of people getting distracted, opening 20 new tabs they don't need, and endlessly procrastinating, might your future perfect neural computer start exhibiting the exact same behavior?
12:29Just because that's the learned physics of how humans use screens. A computer that procrastinates because it assumes that's how it's supposed to function. Something to think about the next time you're clicking aimlessly around your desktop. So thank you so much for joining us on this deep dive. Look at your screens a little differently today, and we will catch you next time.
From the publisher
Researchers have introduced Neural Computers (NCs), a transformative computing paradigm that merges memory, processing, and input/output into a single learned runtime state. Unlike traditional hardware that executes rigid code, these systems use neural networks to internalize the functions of a running computer. Current prototypes utilize video models to simulate interactive command-line and desktop environments based on user instructions and actions. While these early versions excel at visual rendering and short-term interface control, they still struggle with complex symbolic reasoning and long-term stability. The ultimate vision is the Completely Neural Computer (CNC), a general-purpose machine capable of durable capability reuse and explicit reprogramming. By shifting executable state from external software to the model's own latent dynamics, this approach seeks to move beyond the limitations of current AI agents and world models.




