NextFlow: Unified Sequential Modeling Activates Multimodal Understanding and Generation

10 Jan 2026 · 15 min · 8 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

NextFlow, a unified sequential (decoder-only transformer) model that combines text/image understanding and generation, aiming to match diffusion-model visual quality while retaining LLM-like reasoning and instruction following.

Guest backgrounds

No guests are named; the episode is hosted by the “Deep Dive” hosts (no external guest bios provided).

Key claims

Prior unified multimodal models fused LLMs and diffusion inefficiently; NextFlow generates pixels via variable-resolution autoregressive “NextScale prediction” (not raster scan), producing 1024x1024 images in ~5 seconds with ~6x fewer FLOPs than diffusion. Uses multiscale 3D rotary positional embeddings (adds scale “z” coordinate) for resolution-invariant training; fixes coarse/fine data imbalance with scale-aware loss reweighting (alpha=0.9) and mitigates exposure bias via self-correction/residual features. Optional diffusion decoder can restore ultra-high-frequency detail at the cost of possible structural drift.

Notable examples

Inkedit image editing score 4.49 (e.g., in-painting bun cake with vanilla cream; view edits like bird’s-eye perspective). Omnicontext single-subject consistency score 9.22 (subject-driven generation). Interleaved text-image narrative generation from video-text sequences (e.g., step-by-step cooking; Ugly Duckling). “China’s national treasure” example: reasons to generate a giant panda instead of a red panda; improves WiseWorld Knowledge Benchmark 0.60→0.70. Mentions in-context learning (e.g., photo-to-oil-paint style transfer).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The State of AI: LLMs vs Diffusion Models

0:45 to 1:54

Exploring the strengths and weaknesses of LLMs and diffusion models.

“You know, traditional multimodal LLMs were fantastic at perception.”

Introducing NextFlow: A Unified Approach

1:54 to 4:00

Introduction to NextFlow and its significance in generative AI.

“Let's really unpack this because the core breakthrough here, it isn't just about making the model bigger.”

NextScale Prediction: Breaking Bottlenecks

4:00 to 6:28

How NextFlow's NextScale prediction improves image generation efficiency.

“That fundamentally changes the cost equation of deploying these things.”

Training Challenges and Solutions

6:28 to 8:13

Addressing data imbalance and error accumulation in model training.

“It ensures a stable foundation no matter what the final resolution is.”

The Role of Optional Diffusion Decoder

8:13 to 10:43

Understanding the trade-offs of using a diffusion decoder in NextFlow.

“Because of the inherent tradeoff in the VQ tokenizer they use.”

Multimodal Mastery and Application

10:43 to 12:39

Exploring NextFlow's capabilities in complex multimodal tasks.

“It learned from video, which taught it to model narrative progression.”

Future of AI with NextFlow

12:39 to 13:56

Discussing the implications of NextFlow's architecture for AGI.

“The model can see a few examples and infer a stylistic pattern.”

Exploring Autoregressive Models in AI Training

14:01 to 14:50

Learn how autoregressive models can enhance AI's understanding of both text and images.

“Because the model is fully autoregressive across both text and images, all the standard reinforcement learning techniques we use to align LLMs with human values, they can be directly applied here.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Welcome back to the Deep Dive, where we take the fire hose of information you share with us and distill it into genuine, usable knowledge. Today, we are plunging into the absolute frontier of generative AI. We really are. For the last couple of years, it's felt like the world of cutting-edge AI was split into these two separate powerful kingdoms. You had your LLMs, the large language models, who are just masters of text, logic, deep reasoning. And then over in this other lane, you had the incredible visual engines, the diffusion models. Right. The ones making those stunningly photorealistic images.

0:36Exactly. But they operated separately and that created this structural inefficiency or what researchers call friction. And that friction really just meant compromise. You know, traditional multimodal LLMs were fantastic at perception. They could look at an image and tell you what it was. Sure. But ask them to create a photorealistic image. The visual quality just wasn't there. Meanwhile, you have diffusion models, visually flawless. Yeah. But they have none of that deep logical reasoning. So a brain without hands and hands without a brain. That's a great way to put it. And that's where today's topic, NextFlow, enters the arena.

1:10It's a really radical proposal for a single, unified, sequential modeling approach. And this isn't some small proof of concept. We're talking about a 7 billion parameter architecture. It's aiming to perceive, reason, and create across text and images in one seamless system. So our mission today is to unpack how this one system manages to fuse those superpowers, how it gets that high quality visual output, but also keeps that deep logical reasoning. Okay. And the real significance for you, the listener, is that this isn't one of those sticky tape solutions we've seen before where they kind of bolt two different models together.

1:48Right. This is one transformer handling everything sequentially from the prompt all the way to the final pixel. Okay. Let's really unpack this because the core breakthrough here, it isn't just about making the model bigger. It's about fundamentally changing the generation process itself. It is. It's about solving a massive efficiency headache that just plagued all the earlier unified models. So if you look at those previous pure autoregressive models, things like EMU3, they generated visuals using what they called raster scan next token prediction. Which sounds very technical, but you can just think of it like a giant printer head scanning across a page.

2:23Okay. It decides on one tiny patch, one pixel at a time, left to right, top to bottom. So one by one. One by one. And the problem is that the number of decisions, the sequence length, it grows quadratically with the resolution. It gets exponentially harder. Which means what in practice? In practice, it meant generating a single high res, a 1024 by 1024 image could take over 10 minutes. It was academically interesting, but commercially, it was a dead end. 10 minutes. That's like waiting for dial-up internet to load a GIF. So how does NextFlow break that bottleneck? They introduce something they call NextScale prediction.

3:00It's part of a variable resolution autoregressive framework, or VAR. NextScale. So instead of scanning linearly, the model predicts the visual tokens scale by scale. Imagine you're a sculptor. You don't start by carving the tiny details on an ear, right? Yeah. No, you block out the whole figure first. Exactly. You start course. You lay down the low-resolution structural blueprint of the whole thing. and then you move to finer and finer resolutions to add the detail. So it's foundation first, details last. It's prioritizing the global composition of the image over the local little bits, at least early on.

3:33Precisely. And that dynamic, scale-aware processing is the engineering genius here. It lets them train directly at native resolutions, and it translates directly to enormous speed increases. And what are we talking about speed-wise? We're talking about NextFlow generating an image in about five seconds. And get this, it requires six times fewer FLOPs, that's floating point operations, than the big diffusion models at that same 1024 resolution. Six times fewer FLOPs. That's not just a technical footnote. That fundamentally changes the cost equation of deploying these things. It absolutely does.

4:07If you're running this on an API, that kind of efficiency is revolutionary. Right. But it all relies on how they structure that information. You mean the text tokens mixed with all the visual patches at different scales? Right. And for that, they use something called multiscale 3D ropey, or rotary positional embedding. I love that name. What's the intuition there? Why a third dimension? So think of normal ropey as a detailed 2D map. It tells the model where a token is, its X and Y coordinates. Okay. But NextFlow has to process a sequence that mixes text with visual tokens from different resolutions, the big picture and the fine detail.

4:44If it only had X and Y, it would get completely lost when it tried to zoom in. Ah, I see. So they add a z-axis, the scale dimension. For a text token at some position T, they just give it a simple coordinate like TTT. But for a vision token, they explicitly encode its spatial location, its PX pi, and its resolution scale. So that z-axis, the scale information, is the key. And I'm guessing that helps stabilize the training process. Yes. By mapping all these positions to a fixed normalized range, they achieve what's called resolution invariant training. The model learns the concept of scale, so it doesn't have to guess when you ask it to generate at a higher resolution than it was trained on.

5:24It just works. Which is a huge problem in other models. A huge problem. But, you know, architectural efficiency is only half the battle. When they started training this thing, they ran into deep structural instability. What was happening? As they scaled up to higher resolutions, they started seeing these artifacts. And the whole structure of the image would degrade. And the reason, as the material points out, is actually beautifully simple. It's a data imbalance. It is. The earliest scales, the ones that define the whole global layout, the foundation of the image, they just contain way fewer tokens than the later fine-grained scales.

5:58Right. So if you weigh all the errors equally. The model just naturally prioritizes getting the millions of fine details right over the handful of foundational structural tokens. It was basically painting the trim perfectly while the house was collapsing. Okay. And this is where it gets really clever. They fixed this structural failure using scale-aware loss reweighting. Yes. They introduced a hyperparameter, alpha equals 0.9, which, in simple terms, dramatically increases the penalty for getting those coarse scale tokens wrong. It forces the model to prioritize getting the global structure right first.

6:32Exactly. It ensures a stable foundation no matter what the final resolution is. But wait, if you crank up the importance of the coarse tokens that much, aren't you risking the model just ignoring the fine details, making things muddy? It's a delicate balance for sure. But the assumption is that once that structural integrity is locked in, the sheer volume of fine-grained tokens at the higher scales will naturally push the model toward high fidelity anyway. I see. The other big training challenge they had to solve is, well, it's the classic flaw of any autoregressive model. Error accumulation. Ah, right.

7:09Exposure bias. In training, it only ever sees the perfect next step. Precisely. So during inference, if it makes one little mistake, it then conditions its next step on that mistake and the error just snowballs. It's like a pilot who only ever trains in a perfect simulator. The second they hit real turbulence, they're in trouble. So what's the trick here? How do you make it robust against its own mistakes? Is there some kind of self-correction built in? Yes, exactly. They use a self-correction with residual features mechanism during training. So instead of always feeding it the one perfect ground truth token for the next step, they force it to sometimes sample from a distribution of, say, the top few most likely tokens.

7:49So you're intentionally giving it a slightly suboptimal choice sometimes. You are. You're forcing it to practice correcting these small, plausible errors. It learns how to recover and condition its next steps based on a real-world, potentially flawed history. It just makes it much more robust. That's brilliant. It's like training it to navigate imperfect sequences. Okay, so briefly, we should mention the optional diffusion decoder. Why bolt on a diffusion module if the main AR model is so fast? Because of the inherent tradeoff in the VQ tokenizer they use. That process of turning an image into discrete tokens, it's what enables the speed, but it inevitably loses some of the ultra high frequency details.

8:29Like what? Think of tiny testers on skin or small faces or really precise text in the background of an image. The optional diffusion decoder is just there to add that last bit of hyperrealism back in. But there is a cost to that, right? A trade-off. Definitely a trade-off. The diffusion process, it's a bit random. And while it makes the image look stunning, it can sometimes alter the structure just slightly. So for tasks that need strict spatial consistency, like very detailed local editing, you might actually skip it and stick with the core AR model. I see. Speed and precision versus that last ounce of photorealism.

9:03And this unified architecture is what lets NextFlow achieve this really remarkable multimodal mastery, especially in these challenging tasks that combine high fidelity output with following very precise instructions. Like image editing. Exactly, like image editing. On the Inkedit benchmark, the fine-tuned version got the highest overall score of 4.49. It just outperformed really strong diffusion baselines. And we're not just talking about, you know, adjusting contrast. We're talking about local edits that need a deep semantic understanding of the world. Precisely. For example, a local edit like D, in-painting the hollow center of a bun cake with vanilla cream filling.

9:39The model has to understand the geometry of the cake, the texture, the viscosity of the cream. Or changing the camera perspective. That's a view edit, like switching to a bird's eye view of a house. That's a fundamental change in geometry, lighting, scale. Right. And it also shines in what's called subject-driven generation. Which is when you take a specific object or pattern. Yes. And it's a really valuable task. You take a specific subject, say, the unique scrollwork design from a black metal bed frame, and you place it onto a totally different object, like a garden gate, while preserving its identity.

10:14And it does that well. Very well. It got a 9.22 on the Omnicontext single-subject benchmark for consistency, which is highly competitive. It's really fascinating, though, is how it handles narratives. It's not just one prompt in, one image out. That sequential nature means it's really built for interleaved generation. Yeah, the training pipeline for this was really smart. They used massive video text data sets and broke the clips down into sequences of image text frames. So it learned from video. It learned from video, which taught it to model narrative progression. It understands the concept of what happens next.

10:49And that allows the model to produce text descriptions and then the corresponding image in an alternating sequence. It's perfect for complex instructions. Think about generating step-by-step cooking instructions. You get the text, blend the ingredients, followed by an image of the mixer actually running. Or for storytelling. You could have it generate images frame by frame for a story like The Ugly Duckling, illustrating the transformation across several distinct steps. So what does this all mean for the big picture? This unified structure means the model can finally leverage those innate LLM abilities, that deep language understanding and planning, for purely visual creation.

11:29And this is the critical convergence point. It's chain of thought reasoning or co-T. Right. Because NextFlow processes everything sequentially. You can fine tune it to articulate a reasoning trace. It makes a little mental note to itself before it starts generating the visual token. So instead of just blindly acting on the prompt, it acts on its own internal plan for the prompt. Exactly. And that lets it fix semantic problems or ambiguities in the prompt itself. We have a great example of this. If you ask a baseline model for China's national treasure, it might just spit out a red panda. Right, because that phrase is probably in a training caption somewhere.

12:05It's technically associated, but contextually, it's wrong. The wrong panda. The wrong panda. Yeah. But by using chain of God, NextFlow reasons through the cultural context first. It literally says to itself, the universally recognized national symbol is the giant panda. And then it generates the correct image. And that made a real difference in testing. A huge difference. It boosted its performance on the WiseWorld Knowledge Benchmark from a score of 0.60 up to 0.70. That's a massive leap just from forcing the model to articulate its plan. And finally, that training on interleaved data also enables really robust in-context learning, or ICL.

12:46Absolutely. The model can see a few examples and infer a stylistic pattern. So if you show it a couple of examples of turning a photo into an oil painting, you can then give it a new photo. And it will just apply that same oil painting style without you having to explicitly describe it. Exactly. It recognizes and generalizes the transformation. Zero shot. So looking back across everything we've talked about, the core conclusion really seems to be that NextFlow is this powerful proof of concept. It is. It shows that a single decoder-only transformer can now efficiently rival the visual quality of the best diffusion models.

13:19While keeping the reasoning power and the instruction following that we really only associate with huge LLMs. And that inference efficiency is the game changer. An image in five seconds with six times fewer FLOPs. That makes these complex multimodal systems truly practical outside of a research lab. But we should quickly restate that tradeoff. The discrete nature of that VQ tokenizer, it still imposes a bit of an information bottleneck. Right, which is why you sometimes still need that optional diffusion decoder for maximum hyperrealism. That quest for perfect pixel fidelity with no secondary module, that's still the long-term challenge.

13:55And that points directly to the future, right? Native multimodal chain of thought and unified reinforcement learning. Yes. Because the model is fully autoregressive across both text and images, all the standard reinforcement learning techniques we use to align LLMs with human values, they can be directly applied here. Which means we might finally be training AI systems that truly self-improve across both language and vision at the same time, teaching them to align their visual output with really complex, nuanced human objectives. The researchers themselves call this the key step toward an AGI that can think with images.

14:32Okay, so here's a final provocative thought for you. If the ultimate goal is an AGI that can think with images, reasoning through intermediate visual generation to self-improve its own understanding and creation, how quickly will the visual world be unlocked for intelligent analysis and application compared to the purely text-based systems we've relied on for so long? Thank you for joining us for the Deep Dive.

From the publisher

This paper introduces NextFlow, an advanced autoregressive model designed for high-quality image generation and editing. It utilizes a decoder-only Transformer architecture and a multi-scale training approach to enhance visual fidelity and reconstruction accuracy. To support this technology, the authors present EditCanvas, a comprehensive benchmark containing over 5,000 human-verified samples across 57 distinct tasks. This dataset evaluates diverse capabilities, ranging from traditional image modifications like lighting and object removal to subject-driven generation. The research also details infrastructure optimizations, such as workload balancing and reinforcement learning techniques, which significantly improve training efficiency. Ultimately, NextFlow demonstrates superior performance in creating and refining complex visual content compared to existing diffusion and autoregressive frameworks.

More from Best AI papers explained

All 475 episodes
NextFlow: Unified Sequential Modeling Activates Multimodal Understanding and GenerationBest AI papers explained · 15 min
Listen in VO