In short
Race to production-grade diffusion LLMs—why diffusion language models can be faster/cheaper at inference than autoregressive LLMs, how discrete diffusion for text/code works, and what’s needed for serving and RL post-training.
Guest backgrounds
Stefano Ermon is an associate professor at Stanford University and CEO/cofounder of Inception. He helped pioneer diffusion models (early work in 2019) and has focused on extending diffusion beyond images to text/code/discrete objects. Inception launched Mercury 2, a commercial diffusion LLM with reasoning.
Key claims
Diffusion LLMs scale better at inference time, yielding lower price per token and higher tokens/GPU. They can trade quality for compute by changing the number of denoising steps, improving answers via in-place error correction rather than longer “thinking traces.” Serving requires custom infrastructure because existing LLM serving engines are optimized for autoregressive decoding.
Notable examples
GPT-2-scale diffusion vs autoregressive A/B test (same architecture/params/data; diffusion ~10x fewer neural evaluations). Token-masking noise objective. Copilot Arena ranking where diffusion completions score at the top; Mercury models embedded in IDEs (e.g., Cursor, Zed, Helocode). 128k context; not yet multimodal.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOIntroduction to Stefano Ermon
0:48 to 1:40
Host Sam Sharrington introduces guest Stefano Ermon and their discussion topics.
“the price per token or the what's needed per token becomes the key metric that you care about.”
Stefano's Journey in Generative Models
1:40 to 3:00
Stefano discusses his long-term work in generative models and their evolution.
“Before we get going, be sure to take a moment to hit that subscribe button wherever you're listening to today's show.”
The Development of Diffusion Models
3:00 to 5:00
Stefano explains the creation and advantages of diffusion models over GANs.
“And now the bar has shifted a little bit in terms of what these models can do.”
Challenges in Text and Code Generation
5:00 to 8:02
A deep dive into the challenges faced when adapting diffusion models for text and code.
“Like where did the inspiration come from?”
Embedding Spaces and Diffusion Models
8:02 to 10:00
Discussion on the use of embedding spaces to build diffusion models for language generation.
“in those chains of thought and being able to kind of like adjust the amount of compute at inference time.”
Performance Comparison of Model Types
10:00 to 12:50
Stefano compares autoregressive models with diffusion models regarding performance and efficiency.
“In the context of text and code, everything is very discrete and so it's not obvious how you get the mathematics that were developed for continuous spaces do not translate immediately to discrete spaces.”
Advancements with Inception and Mercury 2
12:50 to 14:00
Stefano shares insights on his company Inception and the launch of Mercury 2 model.
“Like you could generate the same quality of text in about 10x less, so 10 times less, sort of like number of neural network evaluations.”
Innovations in Diffusion Language Models
14:00 to 17:30
Learn about the advancements in diffusion models and their performance.
“How did you overcome the discrete challenge in training that model?”
Mathematics Behind Diffusion Training
17:30 to 22:00
Understand the mathematical principles that facilitate diffusion in text processing.
“And crucially, the network can output more than one token at a time.”
Comparative Analysis of Model Inference
22:00 to 26:30
Explore how diffusion models differ from autoregressive models in inference time scaling.
“because it's, you know, autoregressive is next token.”
Show all 30 chapters
Challenges and Strategies in Post-Training
26:30 to 28:00
Discover how reinforcement learning applies to diffusion language models and its implications.
“in post-training apply also to these diffusion models?”
Understanding RL Post-Training for Diffusion Models
28:00 to 29:28
Explore the challenges and strategies in using RL for diffusion language model training.
“Yeah, we're training our own models we have our own pipeline we have not disclosed a lot of a lot of detail in terms of like how we do it.”
Challenges in Transitioning from Autoregressive to Diffusion Models
29:28 to 31:28
Learn about the difficulties in adapting traditional models for diffusion training.
“to make the best possible use of the data and the flops that you can have access to.”
Serving Diffusion Models: Current Limitations
31:28 to 33:32
Understand the challenges in deploying diffusion language models in production.
“generate synthetic data that then you can use to train your word diffusion language model.”
Balancing Quality, Speed, and Cost in LLMs
33:32 to 35:38
Discuss the trade-offs between quality, speed, and cost in language models.
“but it's still not nearly as developed as for autoregressive models.”
Evaluating Diffusion Model Performance: Benchmarks and Insights
35:38 to 37:58
Discover how performance is measured and compared across different models.
“And are you giving the user all three or are they sacrificing?”
User Experiences with Diffusion Models in IDEs
37:58 to 40:04
Explore user feedback and performance comparisons in coding environments.
“qualities about the same as the speed optimized models from Frontier Labs.”
Limitations of Current Diffusion Models and Future Directions
40:04 to 42:00
Identify key limitations and future goals for the development of diffusion models.
“where there's basically an ELO score that they come up with for code generations.”
Challenges and Iteration in Model Scaling
42:00 to 43:35
Explore the iterative process and engineering challenges of scaling AI models.
“And what are the key impediments, steps, that kind of thing?”
Open Science Questions in Diffusion Models
43:36 to 45:28
Delve into the unresolved scientific questions surrounding diffusion language models.
“Can you talk about some of the open science questions?”
Generalization and Hallucination Issues in AI
45:29 to 47:20
Understand the limitations and generalization challenges faced by AI models.
“And so there is a lot of research to be done there as well.”
Controllability in Diffusion Models
47:21 to 50:38
Learn about the controllability advantages of diffusion models in AI applications.
“That's generalization and deep learning work.”
The Future of Diffusion Models in AI
50:39 to 54:42
Discuss the potential of diffusion models to challenge existing autoregressive architectures.
“from the very beginning, the full object, as opposed to generating it token by token where you can only check whether or not it satisfies the constraint at the end.”
Applications of Diffusion Models in Agents
54:43 to 56:00
Examine the current use of diffusion models in developing agentic applications.
“Yeah, when I think about latency sensitive, the things that come to mind most immediately are things like voice interactions.”
Using Fast Agentic Interactions with Models
56:00 to 56:39
Learn how quick interactions with models can enhance productivity.
“So yeah, you can use it if you can plug it in.”
Tracking Google's Innovations in Diffusion Models
56:40 to 58:19
Explore Google's advancements and comparisons in the diffusion model space.
“And so eventually you can get to the final result in less time, which is the thing that actually matters to developers.”
Emerging Research and Efforts in China
58:20 to 59:39
Discover significant research contributions from Chinese teams in the field.
“but it's a little bit harder through these things, this place, I think, in a big company, big lab.”
The Evolution of Diffusion Language Models
59:40 to 1:00:56
Understand the rapid development and excitement in diffusion models for language.
“But then generally, I think that the whole academic community, there is a bunch of interesting papers coming out.”
Cross-Pollination Between Image and Text Techniques
1:00:57 to 1:02:06
Learn how techniques from image processing are influencing text model development.
“There is a lot of cross-pollination, I would say.”
Advancements in Multimodal Diffusion Approaches
1:02:07 to 1:03:09
Explore credible approaches to multimodal diffusion models and their implementations.
“at Inception or elsewhere that points to an approach to kind of a credible multimodal approach based on diffusion?”
Transcript
Automatic transcript. May contain errors.0:00Sam Charrington:A big thanks to Blitzy for supporting the podcast and sponsoring this episode. Want to accelerate software development velocity by 5x? You need Blitzy, which brings autonomous software development to your enterprise codebase. Your engineers declare intent and Blitzy agents map your codebase and generate an agent action plan. Once approved, Blitzy gets to work, autonomously generating hundreds of thousands of lines of validated, end-to-end tested code. More than 80 % of the work completed in a single run. Blitzy is not just generating code, it's developing software at the speed of compute. Experience Blitzy firsthand at blitzy.com slash twiml.
0:42Sam Charrington:That's B-L-I-T-Z-Y dot com slash twiml. If you need to scale up these models and they are actually getting into production, the price per token or the what's needed per token becomes the key metric that you care about. And so what we're seeing with diffusion language models is that they scale better than autoregressive models at inference time. They're cheaper to serve, they're faster, you get more tokens per GPU, which means that the price is actually lower. And so that's why we felt like, yeah, this is the time to do it. And in fact, that's what we're seeing.
1:31Sam Charrington:All right, everyone, welcome to another episode of the Twimble AI podcast. I am your host, Sam Sharrington. Today, I'm joined by Stefano Ehrman. Stefano is associate professor at Stanford University and the CEO of Inception. Before we get going, be sure to take a moment to hit that subscribe button wherever you're listening to today's show. Stefano, welcome back to the podcast. It has been a while. Yeah, thank you for hosting me again. Yeah, it's been a very long time since we last chatted. Yeah, I think about eight years or so. Certainly lots has changed. And we'll get into some of that, in particular, what you've been doing with diffusion models.
2:11Sam Charrington:But to get us started, why don't you tell us a little bit about what you've been up to for the last eight years, maybe? Yeah, so I've been working still in the same space. So I've been working in generative models, I guess, my whole career, my whole life. Now what has changed is that the field really took off. I guess now it's called generative AI and everybody is paying attention to it. And, you know, it's become the thing that everybody is looking at and everybody is trying to, you know, get into. So, yeah, it's been exciting to see the growth of the field and the capabilities of these models.
2:49When I started back in 2014 or so, we were barely able to model any images and it was all very blurry and that was already a big result. And now the bar has shifted a little bit in terms of what these models can do. So yeah, it's been exciting. And, you know, more specifically, my lab at Stanford has always been kind of like at the forefront of what these models can do and kind of like always been innovating at the model level, at the architecture level, the MLCs level. So I did early work on diffusion models back in 2019 when everybody was using generative adversarial networks, if you still remember.
3:34yeah so we kind of like came up with this alternative approach that is what's now called a diffusion model which is now used in pretty much every generative solution for images, video, music and yeah since back then I've been trying to get these models to work on text and code and DNA like discrete objects and finally been able to get some really really good results with this approach and that's what I've been doing at Inception I'm currently the CEO and one of the founders of a startup called Inception, where we are developing a new kind of LLM that is based on diffusion. And these new LLMs are way faster, more efficient, higher quality.
4:18So we just launched our newest model, Mercury 2, a couple of days ago. and so you know if you want to play with a different kind of LLM something that is fundamentally different in the way generate text and code give it a try it's really really fast it's a great solution especially if you're thinking about latency sensitive applications of LLMs with a very tight latency budget these models are really really quick and they get really high quality answers so a lot of developers are already building a bunch of real time AI applications on top of them And so that's what I'm most excited today. That's what I've been spending a lot of my time, kind of like figuring out we'll get these models to work even better.
4:59Sam Charrington:Take us back to the creation of the fusion models. Like where did the inspiration come from? Yeah, so back then the field was dominated by GANs, Generative Adversarial Networks. And, you know, that's that old approach where there is two neural networks. There is one that generates images and there is one that is kind of like trying to discriminate and figure out if the images are real or fake. And then you train them one against each other. And it's a very kind of like unstable and challenging kind of optimization problem because there is this game theoretic kind of aspect to it where they need to out-compete each other, these two neural networks.
5:35And it was very, very unstable, very, very difficult to get it to work well. Lots of tricks were needed. And so we were trying in my lab to experiment with alternatives. And one alternative was the usual autoregressive approach where you kind of like generate the image, let's say one pixel at a time. and that's never worked particularly well for images and video and it still doesn't. It's just very slow and not very accurate. And so we came up with this alternative approach which is now called the diffusion model where essentially you generate an image by starting from noise and then iteratively refining it until you get a crisp kind of like nice image that is consistent with the prompt at the end.
6:20And the key benefit is that the training objective is very stable. The neural network is trained to just denoise an image. You take images, you add noise, and you train the neural network to remove noise. It's a fairly standard, relatively easy kind of like optimization problem that you can use to train neural networks over large data sets, and it works reasonably well. And essentially, you can then use these neural networks at inference time to generate images because the networks have been trained to remove noise, to improve the samples to correct mistakes. And so it turns out that you can just start with pure noise and then you apply this denoising network a bunch of times and at the end you get a really nice image.
7:03And the key benefit, I mean back then people were not thinking about test time inference and those kinds of things. But one of the reasons we really wanted to get this to work was that it had the flavor of a kind of neural network where you have a very deep kind of like inference path because you are chaining together many, many evaluations of this neural network at inference time. So you have a very, very deep kind of like computation graph that can do very, very powerful things, but it's still very scalable during training because you don't have to unroll all of this computation during training.
7:38During training, you just train the model to remove noise. So you just basically need a single neural network evaluation during training. So this idea of having something that is cheap to train, yet very powerful at inference time has always been something that was on the back of my mind and trying to think about ways to do this sort of computations efficiently, which is now kind of like showing up in a different form in the context of LLM where people are very excited in those chains of thought and being able to kind of like adjust the amount of compute at inference time. I feel like it's a similar idea, although implemented on top of a very different kind of generative model.
8:13Sam Charrington:Talk a little bit about the path to getting from diffusion models for images to diffusion models for text? Yeah, so it took a while. So immediately after, you know, getting the good results on image generation where, you know, initially we showed that, you know, these models were better than GANs and then very quickly like the field switched to diffusion models and stable diffusion came out and mid-journey and then quickly basically took over the whole field. And since basically back then, I started thinking about, okay, how do we get these kind of ideas to work for text and code, but we want the model to somehow generate discrete objects.
8:54Sam Charrington:You've mentioned discrete a couple of times as opposed to continuous. Can you talk about why that presents a challenge for diffusion models? Yeah, of course. So if you think about an image or just even a single pixel, you can think of it as a bunch of colors and the interesting thing is that if you change the colors a little bit, you know, the meaning doesn't change. So in particular, you can kind of think about two possible colors for a pixel and all the kind of things in between them still make sense and they don't change the meaning of the image in any dramatic way, right? But if you think about text and you take two words, then it's not clear what's in between the meaning of two different words, right?
9:41and so there is no real geometry to the space of possible tokens or possible words and so that makes the idea of denoising much more challenging because it's not clear what it means to perturb other noise to text it's not clear how you build like the whole geometry does not exist and so a lot of the concepts that were defined for that were invented to get diffusion models to work on images and video they were kind of like relying very heavily on the fact that there is some kind of continuum of possible images and you can kind of like interpolate between them and it makes sense to get the model to kind of like smoothly move from one image to another.
10:26In the context of text and code, everything is very discrete and so it's not obvious how you get the mathematics that were developed for continuous spaces do not translate immediately to discrete spaces.
10:41Sam Charrington:When you talk about the idea of words between points and, you know, words in a neighborhood, it calls to mind embedding spaces and the like, you know, to what degree I imagine that's been tried or, you know, and or maybe part of the ultimate solution to getting it to work for text. Yeah, so that can be, there are approaches that essentially try to build diffusion models for language generation, kind of like in the embedding space. So first you embed everything, then you build a diffusion model. And then the problem is that essentially you have to eventually decode back to text, right? Eventually you cannot give embedding to your users or your customers.
11:22and so that's always the problem that essentially at the end of the day the diffusion model will make some small mistakes and it might not end up exactly in a point that corresponds to one of the existing words in the dictionary and so it's actually pretty challenging to get these models to diffusion models to work well in latent spaces there's been a number of papers including from academia, industrial labs, but it's not been very successful, but it is one of the approaches that people have taken.
11:58Sam Charrington:So what has been demonstrated to work for text with diffusion? So the initial results were still sort of like in the academic setting. It was actually, again, from my lab where we had a paper a couple of years ago, essentially showing that for the first time, it was possible to train a transformer-based model. So you basically took a GPT-2-sized model and then you train it autoregressively the usual way. You train it to predict the next token the way everybody else is training LLMs. And you can train the same neural network as a diffusion model. And in that paper, we showed that for the first time we were able to match the quality.
12:40So in terms of perplexity, in terms of like the quality of the text that these two models are able to generate, it was about the same. but the diffusion model was significantly faster. Like you could generate the same quality of text in about 10x less, so 10 times less, sort of like number of neural network evaluations. So diffusion models are significantly more efficient at the GPT-2 scale.
13:06Sam Charrington:And so just so I understand the setup there, are you saying, you said train it autoregressively and train it via diffusion? Is that two different models that you're comparing or are you sequentially training it autoregressively and then with diffusion? So it's really just like almost like an A-B test where it's not very much a fair comparison in the sense that you take the same exact neural network architecture with the same number of parameters, you train it on the same amount of data, you just train it on the one hand as a typical autoregressive model where you just predict the next token and that's how you use it at an interest time.
13:47And then on the other hand, you can train it as a diffusion model. And so at that point, kind of like the difference in performance is entirely due to the different modeling paradigm, diffusion versus autoregressive model. How did you overcome the discrete challenge in training that model? Yeah, so that was kind of like the main kind of like idea in that paper. like there were some new mathematics, some new methods of basically figuring out what it means to do diffusion in the context of discrete text-like objects. And then it was demonstrated to actually work well in practice up to the GPT-2 scale.
14:30The next step was that with Inception, and I was very excited about those results. And so I started a company called Inception where we've been scaling that up. And so, you know, we've trained commercial scale diffusion language models. So much larger models, some more data. And now the results are extremely good. Like the latest model that we announced this week, Mercury 2, is actually matching in quality some of the best speed optimized models from Frontier Labs. So we'll think about the Haiku models, the Flash models, mini models from OpenAI. So it's at that quality level. But again, it's about 5 to 10x faster in terms of the time it takes you to get an answer using a diffusion model versus an autoregressive model.
15:20Sam Charrington:Are you able to give us an overview or a summary of some of the mathematics that kind of make this work? at an intuitive level, it's all somewhat similar in the sense that there is still a neural network that is trained to remove noise. It's just like the noise process is no longer kind of like adding small numbers to the pixel intensities. It's more like there is different kinds of noise processes that you can use. One that works pretty well is basically one where you mask out tokens. So you kind of like hide them. you take a sentence and then you remove some of the tokens you hide them from the neural network and then you ask the neural network can you predict what those tokens were and so it's similar in some sense to next token prediction except that things are done out of order and the network needs to be able to use information from you needs to use context to the left and to the right and combine it in some interesting ways to figure out how to predict all these missing tokens from the sentence.
16:28Sam Charrington:So in some ways, you're changing the definition of noise to one that makes sense in the context of text. Exactly, exactly. And so actually, that kind of training objective is very similar to the birth style models from, again, many years ago. But that was the thing that for a while, you know, was sort of like widely used in natural language processing. People were training these neural networks exactly on the same objective. this idea of, oh, let's train the network to predict some of the missing tokens. You know, if it, in order to do that, it really needs to understand the meaning of the other tokens, and that's a good way to get representations.
17:07In that ICML paper that I mentioned, basically we show that, well, once you can do that, you can also generate content from scratch because essentially you can start with a sentence where everything is masked, and then you can let the neural network figure out how to fill in pieces, but it does so out of order. So instead of generating left to right one token at a time, it does it in any order. And crucially, the network can output more than one token at a time. And that's why these models are so much faster, because in the autoregressive world, if you want to generate a thousand tokens, you need a thousand neural network evaluations.
17:45In the context of a diffusion language model, the neural network can output many tokens at every step. And so to the extent that you don't need too many steps, 20 denoising steps, then these models can be much, much more efficient.
17:59Sam Charrington:In the image world, I think we're familiar with these like progressively enhanced images where, you know, you see the image taking shape. Are you able to see the same thing with text? Like does text, you know, start out horrible and get better over time? Yeah, there is definitely something like that going on. In fact, if you go on our website, you can see some little animations so to kind of give you a sense of what's going on under the hood. It's not as interpretable, I would say, as what you see in the image space where you see really the details emerging as you go through the process. I think that has always been a little bit less interpretable to me.
18:41But you can definitely, especially in code, you can see the structure kind of like emerging. Sometimes at least you're kind of like able to see some interesting patterns in terms of like how the model is producing the answer. And for sure, like this idea of kind of like being able to control the quality of the answer as a function of the number of improvement steps, the number of denoising steps is actually very exciting because it gives you another direction to do test time scaling, test time inference, applying compute at inference time to control the quality of your answers. So if you have an autoregressive model, the only way you can actually kind of like control, you know, the quality of the answer is by basically producing a longer and longer thinking trace.
19:27You know, that's what these, you know, reasoning models are doing. They produce a thinking trace before actually providing the right answer. And the longer you let them think, the better usually the quality of the answer is and the more expensive it becomes and the slower, of course, it becomes. A diffusion language model has a different kind of axis to do something similar where you can kind of like control the number of denoising steps, the number of iterations. And the more iterations you do, the higher the quality becomes. But all the edits are essentially happening in place. So you don't necessarily have to make the trace longer.
20:02The model is actually able to do error correction. It's able to improve its own answer without having to make it longer and longer, which saves memory and it's significantly more efficient.
20:14Sam Charrington:When you talk about reasoning models and thinking in the context of diffusion, even beyond this idea that you can change the number of denoising steps, should we be thinking about thinking in the same way with diffusion models as autoregressive models? Like, do you still have thinking traces or is all thought, if we're even there yet, is all thought in diffusion models, you know, out of band? Yeah, it's a great question. And in fact, the space of like reasoning and diffusion language models, it's pretty new. Mercury is the first, Mercury 2, like the model we released this week, is the first commercial scale diffusion language model with reasoning capabilities.
21:06So it's still this capability, this technology, it's all brand new. In our case, we do still have reasoning traces. They're just produced in a different way, and the models have been trained to generate them through denoising, through a different kind of like training process. But the idea of a reasoning trace is still there. And in fact, we're able to provide summaries of the reasoning trace to our users if they want to. So it's actually pretty similar. And in fact, all the API, it's all open and compatible and you can still use some parameters to decide and to kind of control how quickly you want your answer, how much you want to trade off compute for quality at inference time.
21:46Sam Charrington:Another thing that I'm thinking about in comparing these two types of models is that with your traditional LLMs, autoregressive LLMs, to get more tests, you just continue generating tokens because it's, you know, autoregressive is next token. For diffusion models, you know, what does that mean? And what are the implications on things like context windows? Like, are you doing rolling windows of generation or how does that all translate? Yeah, that's another kind of like a kind of capability that it's not, you know, there's many different ways of handling like outputs of variable length. We have figured out a way to do it at inception.
22:36I'm sure there are other, there is also the idea of, yeah, doing rolling blocks, which also makes sense, has been published in the literature. so there's different ways of handling variable length it's possible to do it with a diffusion language model it does not affect sort of like the scaling with respect to context size that's more affected by the architecture which is kind of like completely orthogonal to the training objective so at inception we're still using transformers as the underlying neural network and so you know we're using attention and so we have the same kind of like benefits and downsides of attention in terms of like how it scales with respect to the to the sequence length uh but that's an orthogonal kind of like direction it's possible to train uh and in fact we have prototypes of diffusion language models where the backbone the network uh is not a transformer it's maybe state space based or mamba based so that you have better scaling with respect to context length, subquadratic scaling.
23:40But for our main models, we're still transformer-based.
23:44Sam Charrington:I guess the question that is coming to mind for me is why now? Like, is this model enabled by, you know, particular other things that are happening in the space? Or is it just the, you know, the time that it takes or took you to kind of get to this point? Yeah, it's kind of like a combination of both. One is the, you know, it was just like, you know, it took us a while to figure out how to do it. And, you know, it was the right timing to sort of scale things up because finally things were working at least at the sort of like on academic benchmarks at academic scale. The other one is that people are starting to realize that it's all about inference scaling, right?
24:26So for a while, the main axis that people care about and all the interest was around scaling laws in terms of like, okay, how do things scale at training time, at pre-training time, right? Now everything has shifted to inference time scaling and because of several reasons. One is just like, you know, that's where you're seeing the biggest benefits like RL post-training, test time inference, but also just like the economics. Like if you need to scale up these models and they are actually getting into production, you know, the price per token or the what's needed per token becomes the key metric that you care about.
25:08And so what we're seeing with diffusion language models is that they scale better than autoregressive models at inference time. They're cheaper to serve, they're faster, you get more tokens per GPU, which means that the price is actually lower. And so that's why we felt like, yeah, this is the time to do it, because if we are just able to match the capabilities in terms of intelligence of autoregressive models with a solution that scales better along the axis that actually matters, which is cost and speed, then we would have something that can actually be very, very valuable and that customers would jump to.
25:47and in fact that's what we're seeing that there is really really a lot of demand for speed for cost and we're seeing also other competitors are sort of like trying to get fast versions of their models partnering with the AI inference chip companies Cerebras, Grox and Bonova to kind of like try to get the fastest models out there except our solution is software based so it's much more scalable we're still running on GPUs so you can get as much capacity as you can get GPUs for, which is relatively easier compared to specialized data inference chips.
Read the full transcript
26:23Sam Charrington:To what degree do all of the techniques that we have learned about and apply regularly now in post-training apply also to these diffusion models? Some do, some we have to kind of like reinvent from scratch. So if you think about pre-training, mid-training, SFT, a lot of that is actually relatively simple in the sense that you can essentially use the same kind of data sets rough. The architectures don't have to change too much and the loss is, you just need to change the loss function essentially, right? From next token prediction to denoising. If you think about reinforcement learning, that's where things become more interesting.
27:07Because in the context of reinforcement learning, whether you do it from human preferences or you do it if you have some kind of verifiable or non-verifiable reward that you need to optimize for, the sampling process is quite different. And so the way you would propagate that information back into the network is different. And in fact, it's actually beneficial to diffusion language models because in the context of RL post-training, the real bottleneck is inference. Like if you're doing RL post-training of an LLM, you're going to spend most of your time doing rollouts. You get the model to provide a bunch of candidate solutions and then you score them using your reward function and then you somehow figure out how to teach the model to do better, to put more probability mass on the rollouts that were good and avoid rollouts that were not good as evaluated by the reward function.
28:03and because diffusion language models are so much faster at inference time then you can kind of really do different things in the RL post-training stage and that's where we're spending a lot of our time right now which is to figure out what is the right way to do RL post-training for a diffusion language model
28:21Sam Charrington:In terms of pre-training are these models pre-trained from scratch? Yeah, we're training our own models we have our own pipeline we have not disclosed a lot of a lot of detail in terms of like how we do it. But we have our own recipe. We have our own stack for training our models. And are the recipes substantially different from beyond the loss function? Or are they like, you know, if you squint, you can kind of see the echoes of the way we train autoregressive models. There are some similarities, but I would say, yeah, it's been non-trivial to figure out, you know, what is the right way to get these models to work.
29:04And yeah, in fact, I mean, there have been attempts over the years to get diffusion language models to work, including from Google, from other places. And, you know, it wasn't, you know, they were not successful for a long time, right? So it is non-trivial to figure out how to train them well, how to get them to scale in the, you know, in the best possible way, to make the best possible use of the data and the flops that you can have access to.
29:31Sam Charrington:Is it a foregone conclusion that there's no way to like, you know, does the math, for example, say that there's no way to start from a pre-trained autoregressive model and somehow, you know, through some magic transform that into a base that you can then diffusion train? It seems like that would be really interesting given how much, you know, energy and investment has been placed in training traditional models. Yeah, so there's been a number of papers in the academic literature, kind of like trying out various recipes for doing exactly that. And so, you know, to some extent, the embeddings can be still reused and to some extent, you know, networks.
30:24You know, the real challenge is that sort of like the attention mask that you use in a traditional autoregressive model is causal. So the model only knows how to use context to the left as it figures out what to do next. And in a diffusion language model, you really want to be able to have access to the context to the left and to the right as you decide what to change. It's like one of the key properties that make these models potentially much higher quality than compared to autoregressive models. And so that's sort of like the challenging bit. And people have explored ways of kind of like annealing the attention mask to make it go from causal to something non-causal, slowly kind of like making the model drift away from the initial autoregressive thing.
31:07there are more mathematically sophisticated ways of kind of like converting the likelihood from an ultra-aggressive model to a score function, which is what you need in the context of a diffusion model. More broadly, I think one thing that always sort of works is that you can get samples from the ultra-aggressive model. That's usually a good way to, at the very least, generate synthetic data that then you can use to train your word diffusion language model. That's almost like a black box that you can always kind of like use to combine and try to get knowledge out of an existing model. And as we know, I mean, with the distill gate, I think that has been on the mind of a lot of labs and a lot of researchers.
31:50And it seems to be, you know, something that is really going on on a pretty massive scale in other places, but that's always possible.
31:59Sam Charrington:How does the serving setup change for diffusion models? Yeah, that's a great question. And it's another pretty challenging kind of aspect. And I think one of the reasons why there are still no other providers that are able to serve diffusion language models in production today, you cannot run a diffusion language model on existing serving engines. So if you think about BLLM, SGLang, TensorRT, these frameworks that exist and are not even open source, and they are really, really good, at serving autoregressive LLMs very efficiently. So they will handle things like continuous batching for you. Like when there is a stream of requests coming in, how do you batch them together to serve them efficiently?
32:48And there is all kinds of optimizations that you need to do once you have access to multiple GPUs and many requests. And there is a lot of existing kind of frameworks and great work that has been done for autoregressive models. The space for diffusion language models, is much, much less developed. So we had to build our own serving engine. Over the last maybe month or two, there's been some support for diffusion language models in SGLang for the open source models that have been open source diffusion language models that have been developed by the community. So there is starting to be a little bit of ecosystem, a little bit of tooling, a little bit of community support for diffusion language models in the open source community.
33:32but it's still not nearly as developed as for autoregressive models.
33:38Sam Charrington:You talked earlier about the ability to change the number of refinement steps and kind of how powerful that is in diffusion zone type of inference time scaling. Is that something that is currently, you know, I'm thinking like on the static to dynamic spectrum. is it fully dynamic? Is it fully static? Is it per request static? Like how do you think about that, that the knobs there? Yeah, so it's a design choice. I think, you know, like to some extent, it's a choice that we, you know, as developers we made to sort of like figure out how to expose this kind of functionality to the user. So right now our Mercury models allow you to select different kinds of efforts, essentially.
34:32So we tried to basically kept it compatible with the existing autoregressive OpenAI kind of like frameworks so that it's very easy for people to essentially plug in our diffusion language models into their existing apps or IDEs and they can just be used seamlessly. You just need to change the API key, everything works out. So we still basically use the reasoning effort parameters to control, you know, how much compute is used under the hood. But potentially, yeah, you could think about alternative ways of exposing the knob. It's just like as, you know, there is already a very well-developed market right now.
35:13And so we've tried to be, you know, for us, it's very important to be backwards comparable so that our customers can very quickly kind of like switch out whatever they were using before. To a diffusion language model, it's very easy for people to try our models and see how fast they are. And so that was kind of like a design choice that we made because it makes it easier for us to go to market with diffusion language models.
35:37Sam Charrington:beyond speed and cost, which are, you know, these metrics that we've talked about, you know, there's also quality. And are you giving the user all three or are they sacrificing? Where are they sacrificing? How do you know or how do they know what the sacrifices are? And, you know, what's the strongest evidence you have that the speed gains, you know, survive under like real production load at an acceptable quality? Yeah, that's a great question. And I think it boils down to, you know there are three things that matter when you think about llms it's quality speed and cost right and it's always a trade-off between those things and there is you know you can actually plot this uh uh you know the the where existing llms stand in terms of these three things and you know that's what you find if you go on artificial analysis or you know you look at the kind of like you know providers that are kind of like benchmarking llms in terms of like the capabilities that they have in terms of like cost, price, cost, speed, and quality.
36:37And so we've benchmarked our models using this existing methodology. And so, of course, measuring speed is easy. Measuring cost is also easy. Quality is always the hard one, right? Like, what does it mean that a model is better than another one? It's all very tricky to actually measure quality in a good way, right? But the way it's usually done is, you know, there is a number of benchmarks that have been established and that people, you know, that kind of like try to measure things that people care about, like coding ability, question answering, instruction following, stuff like that tool use.
37:17And so what we do is we basically just compare the quality of our models on these existing benchmarks. and we've actually given our models, for example, to artificial analysis. Artificial analysis did their independent evaluation. They tried a model on a bunch of benchmarks that they use to come up with their own intelligence score, which is exactly like a quality metric. It's basically trying to see how different models compare in terms of their capabilities on these benchmarks, which reflect kind of a real-world use cases. And again, the result is that our latest diffusion language model, Mercury 2, is comparable in qualities about the same as the speed optimized models from Frontier Labs.
38:02So Haiku's mini flash models, but significantly faster, 5, 10x faster, depending on which one you compare against. So the big limitation is that it's not the highest possible quality. So if you have a workload where you want to have the most intelligent model, the latest Opus model or the latest Pro model from Google Gemini Pro or something like that, we are not at that quality level. So we have not yet trained a diffusion language model that matches the quality of the best models from Frontier Labs. That's kind of like the key limitation at the moment. So we've been able to show that we can shift the Pareto Frontier of quality versus speed at the level of the speed-optimized models from Frontier Labs.
38:50But we need to do more work to basically keep increasing the quality of our models, train bigger diffusion language models, use more data, figure out better training techniques to close a gap. And at that point, we would have something really, really, really valuable.
39:04Sam Charrington:Do you find when you're comparing these models empirically, do you find any qualitative differences between the types of generations that you see? We heard anecdotally from our users and customers that, yeah, it does feel different, but it's hard to quantify again. I mean, you can use the benchmarks as a good way to measure how well these models do. There are some that I think are, you know, basically editing-like tasks. If you think about autocomplete or edit suggestions, that's the kind of task where intuitively, like you can kind of see that you really want to be able to use context to the left and to the right if you're doing autocomplete in an editor.
39:56And indeed, we're seeing that diffusion models do really, really well. So, you know, there is this thing called Copilot Arena. it's kind of like the LM arena for code generator models where there's basically an ELO score that they come up with for code generations. So it's literally an IDE and developers get to see autocomplete suggestions from two models. They don't know what the models are and then they rank them, which one is better. And we are at the top of that ranking in terms of like the quality of the completions that you get from a diffusion language model. It's really, really fast. So we're seeing a lot of usage right now.
40:34Our models are already embedded in a number of IDEs, Continuous, Zed, a bunch of others, Helocode. And yeah, developers are actually loving the experience, the quality of the generations that we get and the speed at which we can provide.
40:54Sam Charrington:Are there areas where diffusion struggles relative to traditional models and not... You know, granted at a consistent like, you know, tier, like if you're comparing the smaller, faster models to Mercury 2, you know, like long horizon coherence or, you know, needle in a big haystack or like. Yeah, so the context that we, our Mercury 2 model has 128k context. So that would be, you know, that's, you know, if you have a task where you maybe need more than that, that's again, probably not the best use case. Again, I don't think it's a fundamental limitation of diffusion language models. It's just like we haven't trained models with longer context.
41:42We're not multimodal yet. so that's another limitation at least right now so if you have a task where you need vision inputs or you're thinking about audio or outputting images and video or something multimodal we do not yet support those kind of functionalities I mean there's no technical fundamental reason we cannot do it's just like we didn't have the time to train the multimodal models yet
42:13Sam Charrington:In terms of getting to a larger scale, what does that look like for you? And what are the key impediments, steps, that kind of thing? Yeah, it's a process. It's a new technology. And so a lot of the time, we still have to reinvent new things. And it doesn't make sense to do all the R &D at the largest possible scale. so it's you know we can iterate much more quickly if we try out our ideas try out our methods at you know medium scale small to medium scale kind of like uh models size um just because yeah iteration is faster and so there is still a lot of rnd to be done uh before we can kind of just okay let's just scale up right and but fundamentally yeah it's it's uh there are some science questions that still need to be solved.
43:11Then there is engineering. Of course, every 10x in data parameters comes with a lot of engineering challenges, and it often means that you have to change a lot of the infrastructure because there is a bunch of new problems that didn't show up at that previous scale that now become important at the next scale. And so as we go through this process, we're learning a lot about scaling up to a much larger number of GPUs and bigger data sets and kind of like various kinds of engineering problems that infrastructure problems that you know they are not there's not a lot of technical risk it just takes time to figure out how to you know come up with a solution in turn.
43:54Can you talk about some of the open science questions? Yeah it's still pretty open in terms of like you know what is the best way to train one of these models right what is the right noise process there's many choices there we have some things that work but there could be better ones if you think about the inference the interesting thing about diffusion language model is that training and inference are decoupled so in an autoregressive model you train to predict the next token and then at inference time the only thing you can do is to basically reuse exactly the same process over and over in a diffusion language model you're essentially solving a differential equation to generate samples.
44:34And at least for image and video generation, there is a lot of methods that you can use to accelerate sampling. A lot of techniques from the, a lot of numerical methods techniques like fancy ODE, ordinary differential equation solvers or stochastic differential equation solvers. A lot of those techniques have been ported over to machine learning and they've led to really, really fast and high-quality sampling algorithms for traditional continuous diffusion models, the space of discrete language, diffusion language models, it's still the Wild West. Nobody knows what's the best way to do things.
45:19Architecture-wise, I think there is still a lot that can be changed. If you think about RL, what is the right way to do RL using a diffusion language model, even in the context of just traditional image and video models there is still a lot of research that's still kind of like wide open what is the best way to incorporate the information through the diffusion process what's the most efficient way of doing it I'm still involved through my lab at Stanford and some research projects there some collaborations with NVIDIA where we're training big video models, Cosmos we've been working on trying to figure out what is the right recipe for RL on these more standard diffusion models that have been around for six or seven years, the space for language, discrete, it's still much less mature.
46:06And so there is a lot of research to be done there as well.
46:10Sam Charrington:And presumably because you're ultimately based on transformer models, all of the limitations of traditionally trained LLMs are similar in diffusion models. Is that the case? Hallucination is one that comes to mind, for example. Yeah, so hallucinations, yes. I think it's not necessarily an issue with, or at least the way I think of it, it's not necessarily a problem with the architecture. I think that's just like a fundamental issue whenever you fit a statistical model, right? You know, there is data, you're fitting a statistical model, there is a regime where you're going to be interpolating and, you know, maybe the answers that you get are going to be reliable.
46:54that there's always going to be a regime where you're going to be extrapolating at that point. You know, there's going to be mistakes. And I think that's to some extent unavoidable, no matter whether it's a diffusion, autoregressive, no matter what is the architecture. We're learning from limited data. We need this model to generalize. And yeah, generalization is very, very, very, you know, nobody really understands it, basically. That's generalization and deep learning work. I mean, even for classification, like there's been, you know, people in ML theory, very, very smart people have spent a lot of time trying to understand how does generalization work in deep learning.
47:35The progress has been very, very limited. Like to this day, you cannot, there is no predictive theory that can tell you will this neural network generalize. It's all very empirical. You have to try and then you know, you know, did it work or not? But there is no theory that is, or that is, at any reasonable scale that people would care about that is predictive and that will tell you how well a neural network will work in practice, even for classification. For generation, generating models, it's even worse. It's a problem that fundamentally should be impossible to solve, right? There is a curse of dimensionality.
48:13There are some pretty good arguments for why what these models are doing should not be possible, yet they work. So I feel like there is something fundamentally missing there from a scientific point of view in terms of like understanding how these models work, why they work, under what conditions they will work. It's still very, very open.
48:33Sam Charrington:And how about things like explainability or the model's ability to estimate its uncertainty? Are there any, I guess I'm looking for, are there any fundamental differences either to the benefit of diffusion models or to transformer-based models, in these kind of core dimensions? Yeah, so we've not explored much interpretability. I would not expect particular differences in the sense that it's yet again one of the spaces where if the moment you start using deep networks, I think I'm personally pretty skeptical about the whole interpretability research direction. And so that's not something that we've invested in at the moment.
49:19one interesting direction that I think is actually exciting and it's also practically relevant is controllability that's a space where people do care about being able to control the outputs of these models and usually that's done through a prompt, maybe some guardrails at the end there is a certain stack and a certain set of things you can and cannot do with an autoregressive model, a diffusion model at least for images, diffusion models are known to be much more suitable for controllable generation. And the reason is that because the object, let's say the image that you're generating, is sort of like available to the model from the very beginning, it's very easy for the model to check whether or not this object that it's generating is consistent with, say, some constraints or some kind of control signal that you want to use to make sure that the output is consistent with whatever you want the model to generate.
50:24And not only you can check whether it matches your conditions, but you can also steer the generation process in a direction that makes it consistent with these external constraints. And that's only possible because you have the object from the very beginning, the full object, as opposed to generating it token by token where you can only check whether or not it satisfies the constraint at the end. And so that's why diffusion models have been used a lot as priors for solving inverse problems in medical imaging. Like there is a lot of applications where this ability of controlling the output through some external signal has been really, really important.
51:06So I was on some papers where we're doing medical imaging And the idea is that, you know, when you do a CT scan, you're basically taking some projections of your body cross-section. And then, you know, you're trying to reconstruct what your body looks like from some measurements that you get from the machine. And the more measurements you get, the higher the quality of the reconstruction. But it also means more radiations for the patient, right? But if you had a good prior model of what the body looks like, which can be given by a diffusion model, then you can kind of force the model to say, okay, produce something that is likely to be the, you know, to correspond to an actual human body, but it's also consistent with these measurements that I'm getting, you know, for this particular patient.
51:49And that can significantly reduce the number of measurements that you need to take for the same quality level, which means less radiations for the patients. And there's a number of problems that kind of like have that flavor where diffusion models have been really, really good. and so now how to do that for text there's some work again there but that would be pretty exciting because people care about being able to stay on brand or be of course safety constraints there is a bunch of settings where you do want to be able to control the output of the model and so I think that's an exciting capability that is pretty unique to diffusion language models
52:28Sam Charrington:Looking forward what's your a kind of mental timeline for, maybe I should even ask this, maybe even more in a more open-ended way. Do you ultimately see diffusion, challenging, autoregressive models at frontier scale? Yeah, yeah. I think that's our bet. I think there is no reason it shouldn't. I don't know how long it's going to take us to get there. And I think that, I guess the challenge a bit is that the frontier keeps moving. like if you tell me this is the frontier how long do you need to get there I think I could probably come up with a reasonable estimate for that the problem is that it keeps shifting and so the models keep getting better and the speed keeps accelerating so it's hard to predict how long it's going to take and again there is still a lot of R &D unfortunately which has a lot of risks but also a lot of upside like it's entirely possible that we come up with a new algorithm that is way better than what we have.
53:30And so that could accelerate progress by a lot, especially because the diffusion language space is very, very unexplored. I think there's still a lot of low-hanging fruits, a lot of room for improvement, a lot of room for, you know, wildly better solutions to what we're currently doing. So it's hard to predict how quickly it's going to take. It's hard to predict, you know, how well it's going to work if we were to scale up to those sizes. And what's exciting is that, you know, it's unlikely that, you know, one architecture is going to dominate the other one. So maybe the best case scenario is, sure, you know, diffusion models are better.
54:11Everybody will switch. That will become the architecture for LLMs in the future. Even if that doesn't happen, there's got to be some use cases, latency sensitive on device. Like there's going to be some use cases where an alternative architecture is just going to be better. and it's going to be such a big market that even the worst case scenario is actually pretty good for us, right? Because there's going to be just so many use cases in these other labs. But as long as we can win on a reasonable subset of them, that's still going to be extremely valuable.
54:45Sam Charrington:Yeah, when I think about latency sensitive, the things that come to mind most immediately are things like voice interactions. But then I think about all the activity around agents and how they're running a loop And anytime you have looping, if you can compress the, you know, the time for one run through that loop, then, you know, that compounds. Are you doing a lot or seeing a lot with regards to these diffusion models and agentic applications? Are they powerful enough to be used in agents now? Absolutely, yes, yes. So we're already seeing a lot of usage. I mean, you nailed the two main ones that we're seeing.
55:22Voice, a lot of voice, customer support, the educational kind of like agents. People love the speed of the diffusion language models. They always had this issue that they would want to be able to use a thinking model, like a reasoning model. But usually the latency is just not enough. And so maybe they use, unless they use specialized AI inference chips, but that's too expensive and they cannot scale to large volumes. So we had a bunch of customers that are building voice agents on top of diffusion language models. And agents, that's another one, just general agents. MercuryT works actually pretty well in OpenClaw, for example.
56:00So yeah, you can use it if you can plug it in. It's already, yeah, you can use it also for coding, you can decline a kilocode. So it's all, you know, it can use tools, it can reason, and it's really quick. So, you know, especially kind of like, it's not the best model if you're thinking of, okay, I'm going to let it run for 24 hours, I'm going to come back and see whether it solves my problem. Then it's probably not a good use case for that. But if you think about fast agentic interactions and loops where you're actually there and you want to be able to get an answer quickly and there is a human in the loop, then it's a really good model because as you said, it's significantly faster.
56:39You can iterate more quickly. And so eventually you can get to the final result in less time, which is the thing that actually matters to developers.
56:47Sam Charrington:I think last year, I think it was at Google I.O. Google announced and kind of previewed their play in the space. I don't know that I've seen much of it since then. Have you tracked what they've been up to? And can you give us a summary? Yeah. So, I mean, I don't have any inside information in terms of like what they're doing. But yeah, as you said, they also announced the Fusion language model, Gemini Diffusion, a few months after we announced our first Mercury model. I'd like to think that maybe that had a little bit of an influence and a little bit of impact up there, pushing them to actually show they have something too.
57:31I think what they published back then, like those numbers were very comparable to our initial Mercury 1 model. so I don't know whether you've been able to improve what I know is that it's not yet in production it's not yet available to customers so I'm guessing they've maybe not figured out or they're still working on figuring out how to serve it efficiently and what are the best use cases and you know my sense is that there is a big switching cost they're very focused on Gemini and their main model And so, you know, it could be that's kind of like the issue with these big labs is that, you know, they're all in in one direction and then it's hard for them to really focus on an alternative direction.
58:17As a startup, we're in much better positions to do that because we, you know, we're laser focused on one thing and we can really deliver and build everything that's needed to get that technology to succeed. but it's a little bit harder through these things, this place, I think, in a big company, big lab. I think they already kind of like have a direction set and there's a big opportunity cost if you want to switch.
58:39Sam Charrington:What other labs or teams, academic or industry, do you kind of keep an eye on for doing interesting things with the Fusion? Yeah, so there is a lot of good work coming out from China, like the LADA models. These are like several Chinese universities collaborating with Alibaba. So I think they get all the computer funding from industry, and they're doing good work in terms of thinking about models, architectures, how to train them. There's still a huge gap between these ladder models and what we have internally, but they've been doing good work in terms of doing research and pushing the field forward.
59:19ByteDance has a pretty serious effort internally. They've also at least published a few papers with some internally built diffusion language models by that seed, which is kind of like the fundamental research group within ByDance. A lot of smart people, a lot of good researchers. I think they've been doing good work in the space too. But then generally, I think that the whole academic community, there is a bunch of interesting papers coming out. I was at NeurIPS in December and yeah, it was crazy to see how many papers are there on diffusion language models. and if you were to plot it, you'll see that there's been an explosion since that original paper from my group in 2004.
1:00:01Now everyone is kind of like looking at this new paradigm. I mean, of course, it's exciting, right? Because there is this approach that works really well for image, video, and music, and then this design approach that works well for text and code, and then what's going to be the winning solution is there a way to, you know, is there a way to unify everything and have a single kind of generative model that works well across all modalities, of course everyone is excited about LLMs but it's surprising how similar all the different models from Frontier Labs are they're all kind of like clones of each other there's very very little differences and so now there is an alternative approach an alternative path so of course that is generating a lot of excitement in the research community because that's an opportunity to do something new something at the frontier something to have really impact on the conceptual foundations for this approach.
1:00:56Sam Charrington:Do you see image and text as kind of these two divergent paths or are there techniques, you know, with image being further ahead or the techniques that, you know, are created on the image side that you, you know, can pull over or have pulled over to facilitate your work on the text side? Yeah, yeah. There is a lot of cross-pollination, I would say. I myself started out working on images. A lot of the researchers in our team, because there was not really a set of researchers that were working on diffusion for language, or not that many, a lot of the people on our team actually started out as just, you know, pure old kind of diffusion for images or diffusion for video kind of researchers.
1:01:41And so a lot of the know-how did indeed transfer reasonably well. And so, yeah, we also do pay attention, close attention to what's happening in that community. in terms of like distillation, techniques to accelerate the models. As I mentioned before, inference tricks to make diffusion models go even faster. So all those advances have been pretty exciting.
1:02:06Sam Charrington:Is there any work happening either, you know, at Inception or elsewhere that points to an approach to kind of a credible multimodal approach based on diffusion? So yeah, at Inception, we've not been prioritizing our multimodal yet. But in the academic community, yeah, there's been a number of papers that have come out over the last year or so, including from one of my co-founders, Aditya, who was a former PhD student with his lab. He's done some really, really good work in terms of like showing how to build diffusion models that are truly multimodal. And so there's been some really, really good results.
1:02:48in the academic space on getting that unifying model based on diffusion that can handle different models.
1:02:56Sam Charrington:Well, Stefano, it's been great catching up with you and getting a complete download on text diffusion. I feel caught up now. Thanks so much for jumping on and sharing a bit about what you've been working on. Yeah, thank you so much for hosting me. Yeah, it was really fun chatting. Same, thank you. Thank you.
From the publisher
Today, we're joined by Stefano Ermon, associate professor at Stanford University and CEO of Inception Labs to discuss diffusion language models. We dig into how diffusion approaches—traditionally used for images—are being adapted for text and code generation, the technical challenges of applying continuous methods to discrete token spaces, and how diffusion models compare to traditional autoregressive LLMs. Stefano introduces Mercury 2, a commercial-scale diffusion LLM that can generate multiple tokens simultaneously and achieve inference speeds 5-10x faster than small frontier models, paving the way for latency-sensitive applications like voice interactions and fast agentic loops. We also cover the open research challenges in diffusion LLM training, serving infrastructure requirements, and post-training for diffusion-based systems. Finally, Stefano shares his perspective on whether diffusion models can rival or surpass autoregressive LLMs at scale, the advantages for highly controllable generation, and what the future of multimodal diffusion models might look like.
The complete show notes for this episode can be found at https://twimlai.com/go/764.




