Diffusion for Text: Why Mercury Could Make LLMs 10x Faster

24 Feb 2026 · 49 min · 20 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Notes: The Neuron - Diffusion for Text: Why Mercury Could Make LLMs 10x Faster

Overview Podcast Title: The Neuron Episode Title: Diffusion for Text: Why Mercury Could Make LLMs 10x Faster Host: Grant Harvey and Corey Noles Guest: Stefano Ermon, Stanford Computer Science Professor and Founder of Inception Labs Description: This episode explores how diffusion models, originally developed for image and video generation, can be applied to text generation, potentially enhancing the performance and efficiency of large language models (LLMs).

Key Topics Discussed Introduction to Diffusion Models

  • Definition: Diffusion models are generative AI models that generate a complete draft and then refine it by correcting errors, contrasting with autoregressive models that produce output token-by-token.
  • Advantages:
  • Parallel Generation: Allows for simultaneous modifications of multiple tokens, which increases speed and efficiency.
  • Reduced Latency: Especially advantageous in real-time applications where speed is critical.

Mercury's Approach

  • Inception Labs' Mercury Models:
  • Trained as diffusion models rather than autoregressive.
  • Capable of faster token generation, which lowers the cost per token.
  • Larger models trained on extensive datasets for enhanced reasoning and planning capabilities.

Technical Aspects

  • Memory Constraints of LLMs:
  • Current autoregressive models often run into memory bottlenecks, making their inference workloads inefficient.
  • Diffusion models shift towards a compute-bound regime, optimizing GPU utilization.
  • Inference Mechanism:
  • Instead of generating and evaluating one token at a time, diffusion models can infer multiple aspects of a draft simultaneously, leading to faster outputs.

Potential Applications

  • Industries and Use Cases:
  • Real-time Applications: IDEs, voice agents, customer support, and educational technology (EdTech) where human users cannot wait for responses.
  • Future Applications: Further developments could expand to areas such as robotics or multimodal generative tasks (image, video, and audio processing).

Model Evaluation

  • Benchmarking Methodology:
  • Use of offline evaluations combined with A/B testing against relevant business metrics to assess model performance in real-world applications.
  • Quality vs. Speed Trade-off:
  • Users can control the balance between output quality and response time through adjustable parameters in the diffusion model.

Future Directions

  • Upcoming Enhancements:
  • New models with improved planning and reasoning capabilities are anticipated, which will enhance the effectiveness of agentic tasks.

Conclusion Stefano Ermon emphasizes the transformative potential of diffusion models for LLMs, particularly in reducing costs and latency while improving the quality of outputs. The conversation highlights the advantages of rethinking model architectures to leverage parallel processing capabilities, which is crucial for real-time AI applications.

Additional Resources

  • Website: [Inception Labs](https://inceptionlabs.ai)
  • Newsletter Subscription: [The Neuron Newsletter](https://theneurondaily.com/subscribe)

---

Key Takeaways

  • Diffusion models offer a faster and more efficient alternative to autoregressive models for text generation.
  • Real-time applications benefit significantly from lower latency and improved token generation speed.
  • The future of AI applications may encompass not just text but also multimodal processing, further expanding the capabilities of models.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Introduction to Mercury Models

0:00 to 1:08

Learn about the advantages of Mercury models in terms of efficiency and capabilities.

“We were matching the perplexity, but we were able to be like 10 times fast.”

Understanding Image Diffusion Models

1:35 to 2:16

Gain insights into how image diffusion models differ from traditional models.

“computer science professor and the founder of Inception Labs that created the mercury diffusion large language models.”

Exploring Diffusion Models

3:21 to 5:48

Stefano Armand explains diffusion models and their unique training process.

“So I guess to start, would you mind kind of explaining Diffusion in a fairly simple way for viewers who maybe aren't familiar?”

Commercial Viability of Diffusion for Text

5:48 to 9:10

Discussion on the commercial potential and advancements in diffusion models.

“It's very fast because the key thing is in an autoregressive model, you get one token at a time.”

Performance Advantages of Diffusion Models

9:10 to 14:08

Stefano discusses how diffusion models optimize efficiency compared to autoregressive models.

“There is a lot of engineering work that went into post-training the models and making sure that they would be useful for tasks that people care about, like commercial kind of use cases of LLMs.”

Challenges in Transitioning Labs to Diffusion Models

14:08 to 15:00

Explore why labs are hesitant to switch to more efficient diffusion models.

“So why wouldn't all the big labs like immediately switch to this?”

Optimizing Production Workloads with Diffusion Models

15:00 to 17:18

Learn about the current state of serving ultra-aggressive and diffusion models in production.

“If you think about actually serving production workloads, there is a decent amount of open source and closed source, of course, solutions for ultra-aggressive models.”

Context Length and Efficiency in Language Models

17:18 to 19:31

Understand how context length affects performance in autoregressive and diffusion models.

“Yeah, so that really depends more on the architecture than whether it's a diffusion model or an autoregressive model.”

Trade-offs in Quality and Speed of Diffusion Models

19:31 to 21:51

Delve into the balance between quality and processing speed in diffusion language models.

“Would you agree with that, if that makes sense?”

Agentic Flows and Interaction Speed

21:51 to 24:44

Discuss the significance of fast interactions in agentic flows and model performance.

“So, you know, even in the context of image and video generation, you can usually control the number of denoising steps.”
Show all 20 chapters

Open Sourcing and Competitive Landscape in AI

24:44 to 27:33

Examine the balance between sharing research and maintaining competitive advantage in AI.

“It's something that we've thought a lot about before.”

The Impact of Publication Policies on AI Research

28:00 to 29:00

Learn about the challenges researchers face with publication restrictions in industry versus academia.

“And you've contributed a lot of research over the years as well on the topic, though, publicly.”

Mercury's Target Applications and Latency Sensitivity

29:00 to 30:20

Discover how Mercury is targeting latency-sensitive applications for AI in IDEs and customer support.

“I'm curious, are you, with Mercury, targeting any specific industries where you feel like Mercury and what it has to offer could really make a bigger impact than maybe LLMs could?”

The Potential of Diffusion Models for Speech Generation

30:20 to 32:50

Explore the possibilities of using diffusion models for real-time speech and music generation.

“And that's where we're seeing a lot of the initial traction.”

Multimodal Applications of AI and Diffusion Models

32:50 to 36:00

Understand the future of AI with multimodal models that can process various types of data inputs.

“and have a real world model that understands everything and puts together all the learnings and the signals from all the different modalities.”

Challenges in Generalization and Hallucination in AI Models

36:00 to 40:00

Investigate the issues of model generalization and hallucination behavior in AI systems.

“and we really need to do this at some point.”

Evaluating Diffusion Model Performance

40:00 to 42:00

Learn about the evaluation metrics used for diffusion models and their significance in AI.

“more than 10 years ago and I thought it was going to keep me busy for my whole career because it's such a hard problem.”

Evaluating Model Performance Metrics

42:00 to 44:08

Learn about the key metrics for evaluating AI models, including quality, speed, and cost.

“So if you think about the quality metrics, they are often very similar.”

Implementing AI Models in Business Workflows

44:08 to 46:00

Discover strategies for integrating AI models into business processes and assessing their effectiveness.

“I actually have a question related to this, which is a lot of our readers and viewers are trying to figure out how to implement these systems, like, like in their business, in their workflows, right.”

Future Developments of Mercury Models

46:00 to 47:20

Explore upcoming advancements in Mercury models and their potential impact on AI capabilities.

“Well, I know we're getting tight on time, but I have one last question before I let you go.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00We were matching the perplexity, but we were able to be like 10 times fast.

0:03Stefano Ermon:That was super exciting to me, and I really wanted to see what happens if you train something bigger than a GPT-2 model, possible to build something commercially viable. And that's why I started the company to scale things up. The arithmetic intensity of inference workloads that we have today with an autoregressive model is very bad. The utilization is very low, and that's why people are building massive data centers or or even building custom chips, AI inference chips that are better suited for that kind of work. Basically, if you can generate more tokens per second, what this means is that for the same amount of hardware for the same number of GPUs, you can produce more tokens.

0:38Stefano Ermon:And so the cost per token is going to go down. And that's why we're able to serve our models much more cheaply than what you could get because we make better use of the existing hardware. So now the Mercury models that we have in production are significantly larger. They've been trained on more data. that's going to enable Mercury models to be even smarter. It's going to have much better planning and kind of like reasoning capabilities. And so that's going to enable a lot of agentic use cases that people really care about. They're going to make them really, really fast. Welcome, humans, to the Neuron AI podcast.

1:10I'm your host, Corey Knowles, and I'm joined as always by the man who can turn a GPU benchmark into a bedtime story, Grant Harvey. How's it going today, man? It's going great. don't put me on the spot to do that right this moment, though. I'd have to think of some mechanics there. Oh, well, here in just a few, we're going to be joined by Stefano Armand, Stanford University computer science professor and the founder of Inception Labs that created the mercury diffusion large language models. But first, Grant's going to share us a little context before we get in there. Yeah, so image diffusion models work in an entirely different way than the next token predicting GPT models.

1:53So we've invited Stefano today because he's taken that same technology and applied it to LLMs. And it has the potential to transform how AI is used in all types of settings from agents to complex enterprise workflows. Excellent. Well, before we bring him on, I want to take just a quick second to show you Mercury in action, because I think seeing it really matters and will keep you really interested. you'll understand why you need to be watching this video. So what you see here on my screen, this is the Inception Lab site. And if you go up top, you can go to Mercury Chat. And down here, I'm just going to grab one of their suggested prompts.

2:27And I love this one. Simulate a roundtable discussion between Einstein, Ada Lovelace, and Alan Turing. Now, you want to make sure you click this diffusion button because the diffusion button gives you the visual of how cool this is. So watch this. All right.

2:47watch how this works whoa and if you go there's so much cool into the typewriter effect of the ai is that not insane right yeah yeah yeah that's awesome it's really cool too when you do this like when you build a game in html5 like how quickly it can make you something like pong or or you know any kind of like 2d game it's it's amazing yeah it is it is well uh before we get to this interview, please take a quick second to like and subscribe to the channel so we can keep bringing you the most interesting people in tech and AI. And with that, welcome to the Neuron, Stefano. It's great to have you.

3:24Stefano Ermon:Thank you. Pleasure to be here. Good to see you again. Excellent. Well, we're so excited to have you on. As I mentioned before we started, we chatted in Vegas, did a short interview, and I've really been looking forward to this because I had so many questions when I walked away still that I was like, oh, we've got to get him on. We've got to get him on. So I guess to start, would you mind kind of explaining Diffusion in a fairly simple way for viewers who maybe aren't familiar? So Diffusion is a type of generative AI model. It's the kind of model that is commonly used to generate images, video, music.

4:03And you're probably familiar with the, you know, chat GPTs or Gemini's or Claude, where

4:08Stefano Ermon:kind of like see the models generate text kind of like left to right one token at a time. A diffusion model works very differently in the sense that it generates the full object from the beginning and then it refines it by kind of like fixing mistakes making it sharper making it look better and better and it's a very different kind of like solution that it's more parallel in the sense that the neural network is able to modify many components of the image or the text at the same time and that's why diffusion models tend to be a lot faster than traditional autoregressive models that kind of like work left to right one token at a time okay right right and how how is it that they're they're actually like reasoning over the original version that they they create like how like how do they know that the first version isn't good yeah so it's a great question yeah that's a great question and it really it really stems from the way the models are trained.

5:04Stefano Ermon:A traditional autoregressive model like a GPT model is trained to, there is a neural network and it's trained to predict the next token, the next word. And that's how you use it at inference time. You give it a question and then it will try to predict the answer left to right, one token at a time. The fusion language model, it's trained to remove mistakes, fix mistakes. So you kind of like start with clean text or clean code. You artificially add mistakes and then you train the model to fix those mistakes. And that's how the model is also used at inference time. You start with kind of like a full answer and then you refine it.

5:41Stefano Ermon:And so it's a very different way of training the models. It's a very different way of using the models at the inference time. And it happens at the speed of lightning. It's very fast. It's very fast because the key thing is in an autoregressive model, you get one token at a time. You have to process a massive neural network with hundreds of billions or trillions of parameters. And at the end, you only get a single token. Yeah, very inefficient if you think of it that way. You still need big neural networks, but each forward pass, you need to still evaluate the whole thing. But then at the end, you are able to modify more than one thing.

6:20Stefano Ermon:And so as long as you don't need too many denoising steps, too many diffusion steps, this can be really, really fast. Wow. Stefano, what did you... I know you've done a lot of research on this topic along the way. What did you see that convinced you diffusion for text was commercially viable now? Yeah, so it started... I mean, I've always been passionate about diffusion models. The kind of like original idea came out from my lab at Stanford back in 2019. back then pretty much all the image generative models were based on GANs generative adversarial networks which were very difficult to train, very unstable there is this kind of game between a generator and a discriminator it's a pretty complex kind of model to train and there is all kinds of issues in scaling that approach up and we came up with this alternative approach of training the model to remove noise and and then generating in a course-defined way.

7:25Stefano Ermon:And we showed that it was working much better. And eventually that took off and everybody kind of switched to diffusion models for image, video generation, mid-journey, SORA, stable diffusion. They were all based on those original ideas from my lab at Stanford. And since back then, I kind of tried to see, can we get this to work also on text and code generation? and it took a few years to figure out how to do it properly. But then in 2024, we kind of had a breakthrough. We figured out how to adapt the math and the underlying sort of like algorithms from continuous spaces to discrete spaces like text and code.

8:07Stefano Ermon:And we had some really promising results at the GPT-2 scale. So at Stanford, in university labs, you don't have access to a lot of GPU, a lot of compute. And so the largest model we were able to train was a GPT-2 sized model. But basically what we did is we took a GPT-2 sized model and we train it as a diffusion model and as an autoregressive model on the same data. And so what we found was that the quality was the same, so you were able to actually, if you think about perplexity, which is the way people usually use to figure out how good the model fits the data, what's the quality of the generations you get.

8:46Stefano Ermon:we were matching the perplexity, but we were able to be like 10 times faster. And so that was super exciting to me. And I really wanted to see what happens if you train something bigger than a GPT-2 model. Is it possible to build something commercially viable? And that's why I started the company to scale things up. And you've been scaling up from GPT-2 caliber since then, right? Since then, yes. Now the Mercury models that we have in production are significantly larger. They've been trained on more data. There is a lot of engineering work that went into post-training the models and making sure that they would be useful for tasks that people care about, like commercial kind of use cases of LLMs.

9:27Are they kind of similar in size? Should we be thinking about them the same way? Oh, like parameters and such? How do we compare them to a traditional language model?

9:40Stefano Ermon:Yeah, so the models are still fairly large. in terms of the number of parameters. We're still using actually similar architectures. So under the hood, it's still a transformer. Okay. So what we found is that that kind of architecture actually works pretty well. Also for, you know, as a backbone for a diffusion language model. And so, you know, we've kind of used what we knew worked well and is well supported by existing, you know, frameworks and open source code. And so, yeah, the neural networks are not that different. It's just like they're trained in a different way and they're used in a different way at the inference time.

10:23That's so interesting. I didn't even realize that, that essentially you're still looking at transformer technology, just not wrapped in an autoregressive approach, right?

10:33Stefano Ermon:That's right, yes. And I think perhaps that's actually suboptimal. I mean, people have kind of converged on transformers as being a really good architecture for autoregressive models. I mean, these days people also use them for diffusion models, like people use diffusion transformers. So it's kind of like an architecture that it's widely used across different modalities, across different kinds of generative models. but it's possible that there might be better architectures that shine even better once you change the generative model, it's no longer autoregressive so I think the design space is different, I think there is a lot more room for doing R &D and coming up with further improvements just by kind of matching the neural network architecture to the training objective and to the inference kind of computations that we do right that makes sense very cool when we talk about parallel generation what is parallelized are we talking about tokens spans edits is something i maybe don't know about yeah basically what's parallelized is that the the network is able to essentially modify multiple tokens at the same time wow and so that's kind of like what you were seeing i think if you tried our website and you kind of like see the animation of how the diffusion model works, you're going to see that it constantly changes the answer and it's not one token at a time.

12:06Stefano Ermon:Many things get changed at the same time and that's what makes it more parallel. It makes it much more suitable to GPUs. GPUs are built to process many things in parallel. We're going to like apply the same computation across across different data points effectively. And it's the kind of computation that we do when you sample from an autoregressive model does not map well at all to a GPU. It's a very memory bound kind of computation where you're going to spend most of your time moving around weights from slow memory to fast memory where you can actually do the computation. So the arithmetic intensity of the inference workloads that we have today with an autoregressive model is very bad.

12:54Stefano Ermon:The utilization is very low. And that's why people are building massive data centers because it does not, or even building custom chips, AI inference chips that are better suited for that kind of workload. We just saw a model drop a couple of days ago where in order to get it two and a half times faster, the cost is 6x. Right. Because it's a sequential computation. You cannot generate the third token until you've generated the first and the second. And so it's just a structural bottleneck. There is no way to parallelize it because there is sequential dependencies across the computation. And so you can't process something into the future until you've generated everything before it.

13:38Stefano Ermon:And so there's just no way to parallelize that. It's kind of funny. Basically, what you're talking about is like instead of scaling the amount of GPUs widely, it's like let's scale the amount of tokens that we're actually dealing with at a certain time. And then you can do less on a GPU. That's awesome. Exactly. Exactly. Like it's shifting from a memory bound regime to a compute bound regime where you basically are bounded by the number of flops that you have available on the GPU, which is a much easier quantity to scale. Like if you were a chief manufacturer, it's a lot easier to add flops that do increase memory bandwidth.

14:13So why wouldn't all the big labs like immediately switch to this? Like if it's that much more efficient, because I guess they have to train it. They don't have the same skills that you do. What's your thought?

14:22Stefano Ermon:Yeah, some of the labs are very entrenched in a certain stack. And so there's a big cost if they were to switch to something different. There is quite a bit of secret sauce involved in terms of like, what is the right way to train these models? what is the right way to even just sample from them. It's not as obvious as, okay, generate one token, append it, generate the next one, append it. There is not much you can do on the inference side if you have a traditional autoregressive model. But on a diffusion model, the design space for inference algorithms is much broader. And then there is also the issue of kind of like the MLCs level.

15:04Stefano Ermon:If you think about actually serving production workloads, there is a decent amount of open source and closed source, of course, solutions for ultra-aggressive models. Things like VLLM, SGLang, TensorRT, there are pretty mature serving stacks for ultra-aggressive models. For diffusion models, it's much earlier. We have our own stack, but it takes a significant amount of work if you want to figure out how to actually make things efficient in practice on real-world GPUs. And there's all kinds of optimizations that you can do at the system as well. Well, I've got to say, at a dollar per million output tokens, you seem to be doing okay with that.

15:53Stefano Ermon:For sure, for sure. That's the thing. I keep thinking when we talk like GPUs and what it takes to do the autoregressive approach, you know, this kind of, in a lot of ways, could be a smart way to sort of side skirt things like the current memory supply issue, the need to go acquire brink trucks of money and back them up at Jensen Wong's patio door. you know I really think that this is an interesting approach at a prime time for that yeah basically if you can generate more tokens per second what this means is that you know for the same amount of hardware for the same number of GPUs you can produce more tokens and so the cost per token is going to go down and that's why we're able to serve our models much more cheaply than what you would get if you were to use traditional ultra-aggressive models because we make better use of the existing hardware and so the costs are actually significantly lower.

16:58That makes sense. That makes sense. So how does this behave with long context? Are we looking at it getting specifically any more expensive or is parallelism playing more of a role as they grow?

17:18Stefano Ermon:Yeah, so that really depends more on the architecture than whether it's a diffusion model or an autoregressive model. Right now, as I mentioned, we're still using self-attention, which unfortunately scales pretty poorly with the context length. So I would say there is no difference. It's not better, it's not worse than an autoregressive model as you think about longer context. Our models are supporting roughly 100k tokens of context length. we could potentially scale it up more. Again, it's not something that is very different. If you think about an autoregressive model versus a diffusion model, it's more a function of the underlying architecture.

18:02Stefano Ermon:And in fact, we can actually use alternative architectures that scale better with respect to the context line, like state-space models or other attention variants that are more efficient. We have some preliminary results, So everything is compatible with different kind of backbones, but not in the production process at the moment. That's cool. Yeah, I was wondering, like, I've always wanted to ask a researcher this. Have you seen anything over the past year or like six months even that has lit you up in terms of like, oh, this could be a good alternative for that memory context problem? Or are you like, we still haven't seen anything that's even close?

18:45Nothing particularly.

18:47Stefano Ermon:I think it's just like a fundamental problem for which it's going to be hard to get a real breakthrough. There is just inherent trade-offs. I think of them in terms of sufficient statistics. What do you store about your past and how do you keep track of... You want to remember the things that are useful. You want to discard the things that are not useful. And that's just fundamentally a hard problem. Like there is no, there's always, you know, there is some kind of no free lunch involved where ahead of time, you don't know what you should remember and what you should discard. And some things are going to be useful for something and they're going to be not useful for something else.

19:25Stefano Ermon:And so I think it's a fundamentally very difficult problem where you have to make trade-offs. Although I will say with that equation, being able to work in parallel, I think makes you maybe perhaps make those trade-offs or you can make some of those calculations more efficiently. Would you agree with that, if that makes sense? Yeah, I think it changes a little bit the design space in terms of how many flops you have access to and how memory bound you are, but not fundamentally. You still need to process and you still need to be able to look at all the context, all the past information to be able to generate good quality answers.

20:12Stefano Ermon:whether you do them one token at a time or you do them in parallel you kind of have to look at the past and so there is something pretty fundamental there you know i would say probably outside of most you know maybe enterprise and software applications but you're dealing with what say the average worker uses 100k context is is plenty for most things you can really do a lot in that range. Yeah, I was kind of wondering how coherence plays together by not going left to right. And I guess that's me thinking of it through how my human brain works. Yeah. So essentially, there is an element of error correction.

20:55Stefano Ermon:And so the models are trained to fix mistakes. And then they constantly revise the answer. And so initially, the answers are not coherent and then they get increasingly better as you throw essentially more compute they also think of it as a as another dimension over which you can scale compute at test time so test time inference test time compute to trade off quality for speed or cost and so that's kind of like the fundamental trade-off that is exposed by a diffusion language model it just provides you an axis to control quality versus speed. And there is a fundamental trade-off. Like if you want to be, you know, if you don't want to do too many denoising steps, too many passes over the output, the quality is not going to be as good as what you would get if you refine it many, many, many times.

21:50Same as with if I'm running, you know, stable diffusion and CompUI or something and you're doing your denoising runs.

22:00Stefano Ermon:That's exactly the same. Exactly, exactly. So, you know, even in the context of image and video generation, you can usually control the number of denoising steps. The more denoising steps you take, the higher the quality. But of course, the more expensive it becomes, the more time. How does this work in a Gentic context then? Like, how would it be different? Like, let's say, like in something like Cloud Code, like where, you know, it's going out, it's doing a bunch of tool calls, it's reasoning. How does it play out in that scenario if different? So it's related in the sense that, you know, it needs to output something.

22:38Stefano Ermon:And then, you know, if there is a tool call involved, then it kind of you need to essentially wait until the result of the tool call comes back. But it's not too different. Like essentially, we're able to serve the models in an open AI compatible way. We support tool calls. And so people are already using our models in, you know, a variety of authentic frameworks, including some of the open source ones. I think people have figured out ways to also use it in cloud code. I think there are some wrappers that allow you to use other models. I've been eyeballing how to connect it to my OpenClaw. That works.

23:19Oh my gosh. Have you done that actually?

23:22Stefano Ermon:Have you played with OpenClaw? I haven't, but some of my team members have and they've used Mercury. So I know for sure it works. Excellent. That's cool. Excellent. I haven't tried to connect it yet. I've still been stumbling my way through, but it's very much on my list because I keep thinking of like, I have such a variety of tasks. I'm having this thing do for me now. And I think some of these, that super speed would be so huge on. And so much more affordable, frankly, than running plenty of the other models I have access to as well. so it's definitely on the radar and I feel like in many agentic flows there's a lot of value for this approach yeah completely agree I think at that point speed of interaction with the environment becomes the key bottleneck you really want to be able to use the tools that you have access to and collect the information needed to solve the task and the faster you interact the faster you can take actions, collect feedback reason, decide what to do next the better the experience the better the model is going to work and so speed I think we're seeing it also from other labs people are pushing more and more okay let's make the models faster, faster, faster because that really improves the user experience Yeah because you kind of have to scale that at the same time you're scaling intelligence otherwise I feel like what you're going to wind up with is a 48 minute wait for everything you ask for like you kind of have to lift them together Yes, yes um you mentioned earlier so like one of the the um unique things and benefits of being the person who basically made this um uh is that you know all of the tricks of how to you know run it and deploy it um and you know there's a lot of frameworks that support regular llms but maybe not for as many frameworks or you know you have an in-house framework for dealing with diffusion would you ever make like an open version of that like do you want everyone to go through you like What's your business plan there, I guess?

Read the full transcript

25:25Stefano Ermon:It's a good question. It's something that we've thought a lot about before. I mean, a lot of the team is very academic. A lot of us have been researchers, have been publishing everything we do in our labs, and we see a lot of value in sharing our ideas and having a community of researchers that work together to improve and make progress. I think the constraint is that, as you know, it's an extremely competitive kind of landscape at the moment. IP, it's still a big moat. It's still an important kind of part of the company. And, you know, we feel like it would be hard for other labs to reproduce what we have.

26:10Stefano Ermon:and so unfortunately open sourcing anything would reveal probably quite a bit about how the you gotta keep your advantage and so that it's expensive it's expensive to do this work i'm sure it's gotta cost a fortune yeah no that's fair no it's totally fair i was just thinking like you know like in one of nvidia's strengths right is is cuda which kind of keeps everyone locked in so i was wondering if there was a similar kind of play there eventually where you know you make it easy for everyone to do this and but then you're still the gatekeeper in some way no and then that would also be like i think a lot of value just like uh you know if we could open source something just get uh contributions from the community get feedback i think there would be a lot of value if we could do it so that's why we know we thought a lot it we thought about it um a lot and we'd love to do it at some point maybe when the team becomes bigger and maybe there is a way to release a smaller model or some more like a research type model that people can play with and make improvements.

27:13Stefano Ermon:I'm sure there is a lot to be invented and it would be good if the whole community works towards making the fusion language models become the default for the next generation of other people. Yeah, that's right. Yeah, and I get it. It's a tough sword too because the fact is, you've got to continue to be able to push forward. I've wondered about the open source value proposition for companies that have done it for quite a while. And I get the idea and I love that they're available, but at the same time, I'm always left wondering, but how do you continue to innovate if you can't make any money? The fact is it takes money to do this.

27:53These researchers aren't coming cheap these days and neither are GPUs, you know? All of it's really, really expensive. And you've contributed a lot of research over the years as well on the topic, though, publicly.

28:10Stefano Ermon:Yeah, that's the benefit, I think, of being in academia, that everything is open and you're allowed to publish all of your work. And, you know, that's the whole point for advancing the field together as a community. I love that aspect. And I think, I mean, a lot of the researchers do and, you know, I could sense a lot of unhappiness from colleagues and other researchers at, you know, in industry working in the big labs, you know, as people, as the publication policies, you know, started to tighten and people were not allowed to publish anymore. I think there was a lot of people were not happy.

28:53Stefano Ermon:And so... When it's businesses doing the research instead of universities, essentially. Yeah. Yeah, that's definitely a struggle. I'm curious, are you, with Mercury, targeting any specific industries where you feel like Mercury and what it has to offer could really make a bigger impact than maybe LLMs could? Yeah, at the moment, we're going after what we think of like instant AI, sort of kind of like applications of LLMs where latency is critical, which typically means there's a human in the loop. and the human cannot wait and that human could be a developer. So we are seeing a lot of usage of Mercury models in IDEs where you're essentially providing suggestions or edits to the code, for example, directly to a developer.

29:43Stefano Ermon:And there you maybe have a few hundred milliseconds of latency budget and you want to be able to provide the best possible suggestion within the latency budget. But it could also be customer support, voice agents, ad attack. I want to ask about that. Yeah, any other situation where you have, you know, to give an answer, you have to interact with a human in real time, then latency becomes critical and the game becomes, again, sort of like what's the best quality result that you can provide within the latency budget for a reasonable cost. And that's where we dominate existing autoregressive solutions.

30:21Stefano Ermon:And that's where we're seeing a lot of the initial traction. I think eventually, as the intelligence of the models keeps improving, as we do more R &D, as we catch up with Frontier quality models, I think there's going to be more and more applications that we can go after. But right now we're going after latency-sensitive applications of other labs. That makes sense. There's a lot of stuff in that world that would fall in there. Grant, you said you had a question about the medical tech end of it. No, I was thinking, um so the immediate place that i went to was voice agents and i was wondering could you make a diffusion speech model and then i was thinking that would sound kind of wild if like all of a sudden it comes out and it's like and then you have the and then like this you hear the sound wave generate um but then also i was thinking i'd love your take on that if it's even possible um but then also i was thinking well it makes sense to try and you know perhaps match it with a speech-to-text model because it makes sense to that's where you need the most real-time interactions right as if you're having a voice-to-voice conversation so that makes obvious sense to me i'm curious how how that's going yeah and and i mean you're absolutely right that diffusion actually does work and it works really really well for speech and music generation like some i know that some of the open source models and some of the state-of-the-art actually closed source models are based on diffusion for text.

31:54Stefano Ermon:I didn't even know that. Yeah. Yeah. And so, you know, it does make a lot of sense. One of the challenges is that, you know, if you wanted to go straight from voice to voice, is that often these kind of interactions involve, still involve tool calls. So if you're doing a customer support, you might still need to be able to, you know, query the database or check a calendar for availability or look up the menu to get the prices. And so there still needs to be some text, I think, some code involved, which makes it a little bit more tricky to develop. But we are very excited about eventually getting to something that is actually multimodal.

32:35Stefano Ermon:The existing Mercury models are just text only or code only, but we know the future models work really well for image, video, music. And so if we put everything together, we could get to something like truly phenomenal handling different kinds of modalities and have a real world model that understands everything and puts together all the learnings and the signals from all the different modalities. But it's definitely something we want to do at some point. Yeah, that would be so awesome. Would that be something that would be useful for like a robotic kind of situation or would that be more for like simulations that you can use to train robots in your opinion?

33:14It could be a mix.

33:16Stefano Ermon:It could be decision making, like if you're using a video or other kinds of sensors as input and then use the model to make decisions or kind of like analyze what's going on in the surroundings. It's a very useful kind of like application of this technology. In fact, we've already heard it from some early adopters that they would love for our models to have image inputs because they're building computer agents. And so that's another space where you really need to be quick. You need to be able to interact fast with whatever software, whatever application the agent interacts with. But it's important not to just look at the text and the HTML code, let's say, of a web page, but actually seeing what's happening.

34:12Stefano Ermon:And so that would open up a lot of other applications, I think. That's blowing my mind. How would a computer use, and this might be giving secret sauce away, but how would a computer use Diffusion LLM work? Is it generating the whole screen space? How does that work? For computer use, no. It would be more like controlling the actions. So, you know, the agent can type, the agent can click, the agent can check out an item for you or can book a flight for you. So it needs to take a bunch of actions, let's say on a website or maybe on your apps or on some apps on your phone, like depending on what's the environment.

34:47Stefano Ermon:But how is it? I guess what I was wondering is how is it seeing it? How is it like see the... And that's, I guess, that's maybe more so on the vision model side of things. Is the visual of how diffusion looks? what brings you to this gap? Yeah, yeah, exactly. So existing models, yeah, they take a mix of sort of structured information about what kind of menus are available, where the buttons are, and images. And then they use that and they map it to an action. So you can imagine something similar where instead you have a diffusion model that processes the same inputs, but then produces the answer, not one token at a time, but through this refinement process, which makes a lot of sense because people are already using diffusion policies and flow-based policies.

35:37Stefano Ermon:Like if you look at RL and robotics right now, the way one of the best approaches for controlling robots and kind of like implementing the policy that decides what actions the robots take is based on diffusion. It's based on flow-based models, which is more or less the same thing. And so that's another kind of like data point that really excites me and gives me even more confidence that we are on the right track and we really need to do this at some point. Just to clarify, is that a vision action model or is that a different type of model? Yeah, vision action. Okay, cool. Well, I have a question that's kind of a different direction that I've been really curious about, and that's does diffusion change hallucination behavior or controllability of a model, or is it similar in nature?

36:28Stefano Ermon:So those trade-offs change in the sense that, you know, you're at the end of the day, hallucinations happen because, you know, you're building a statistical model. And so, you know, whenever you fit a statistical model to data, there is a certain regime where maybe you're interpolating, but then you might need to extrapolate and then mistakes happen unless you have a perfect model, but perfect model is never actually possible. because you're fitting a different model even if you use the same data you're going to get a different kind of behavior so it's going to interpolate it's going to extrapolate in different ways what we're seeing is that they still make mistakes if you try our Mercury model it's not perfect it does hallucinate but it does so in different ways and it's hard to quantify how there are benchmarks so we use benchmarks to see how well does it know you know does it general knowledge and instruction following doesn't make up things and we're seeing that it does really well but it's it's hard to actually uh precisely quantify or even qualitatively figure out okay there are certain things that it doesn't do well or there are certain things where it does better um that that's actually a pretty pretty hard problem like even in uh from a from an academic theoretical perspective kind of like understanding how generalization work, how these models are actually able to combine all the knowledge that they see in the training data in ways that make sense and ways that don't make sense.

38:00Stefano Ermon:It's widely open. Nobody really understands how these models work. And so unfortunately, that also means that it's very hard to compare two kinds. Yeah. Well, OpenAI had this paper not too long ago, right, that was saying that basically like hallucinations are a training problem, or at least that was their thesis, that, you You know, the models are effectively being rewarded for guessing as opposed to saying when they don't know. Do you buy into that? Do you subscribe to that idea? And if so, is that something that you could possibly, you know, work on to help reduce that even more? Yeah, I think fundamentally it is a training problem in the sense that it's, you know, you're fitting a statistical model.

38:42And fortunately, this is a very, you know, it's a very high dimensional space, what we would say.

38:48Stefano Ermon:There is an extremely large number of possible combinations that you could come up with. If you think about all the different sentences that you can generate, it's just like a combinatorially large space. And no matter how big your training set is, it's only ever going to be like a tiny little fraction of all the possible things that you could, all the possible sentences that you could generate. And so what this means is that the model has to essentially, you know, the training data itself will not tell you everything and you have to interpolate and extrapolate. You have to generalize and nobody actually knows how this model is generalized.

39:27Stefano Ermon:Even in simpler settings, like even if you take, you know, supervised learning, just training a neural network to classify images, nobody understands how these models generalize. here we're talking about something even more complicated where you're not just outputting a binary label or a thousand classes in ImageNet, but you have an extremely large combinatorially large space of outputs and nobody understands how generalization works in that space. In a sense, is it kind of a miracle that any of this works in the first place? Absolutely. I mean, I started working on this more than 10 years ago and I thought it was going to keep me busy for my whole career because it's such a hard problem.

40:10Stefano Ermon:That and, you know, it almost feels impossible for this to work. There's just so many combinations. Even if you think about fitting, let's say, a model over images, there is so many different kinds of images and different kinds of combinations of features. And so let's say you train a model over a data set where there is red cars and blue buses and red buses. then should the model generate a blue car? Right. And it's now clear, right? Yeah. Like, but fundamentally, that's what these models do. And it's, of course, much more complicated because there's many more different things and many different shapes and colors and some combinations make sense, some don't.

40:53And they're able to do it the right way.

40:56Stefano Ermon:But it's just a very, very bad problem. That's interpretability has become such a fascinating field in terms of just looking at, you know, yeah, we know these are the parts. We know this is how we build it. We know that these principles work on this end and these principles work on this end, but somewhere in the middle is that, like, leap of faith that we don't quite get. And it's really, it's fascinating. Grant, I'm glad you asked that question so much. You just got to ask, you know? That's right. I don't know. So I've got to ask the people who know. So do you feel like, does model evaluation, is it approached any differently with diffusion?

41:46Or are we essentially just looking at the answers in there? I just, I wasn't sure. I wrote this question down and I wasn't sure it wasn't really stupid. But I decided I was going to ask it anyways. is if not being sequential effects, how those metrics work and evals.

42:05Stefano Ermon:Not too much. So if you think about the quality metrics, they are often very similar. So we can use the same benchmarks as WeBand, CB or IFE. A lot of the existing benchmarks that people use to test how good are the models doing various things that we care about, following instructions, writing code for us, software engineering. like so we can still essentially test end to end you know we just feed the same prompt to our model we look at the answer and then we see how well it does at the task and that's a very useful way of measuring you know how useful the models are uh there are other metrics that are specific to to diffusion models like how many kind of like denoising steps do you need to do so essentially how fast they are that's another thing that we need to track and we and we need to optimize because there's always kind of like a trade-off between quality and speed.

43:00Stefano Ermon:And so everything becomes a little bit more complicated because there is an extra knob that you can kind of play with. And so there are a few other things that are maybe diffusion-specific. But broadly, I think it's still always going to be a matter of speed, quality, and cost. You always boil down to those three things. And speed and cost are relatively easy to measure and there's not a ton of wiggle room in terms of like how you do things, quality is really the harder one. That's why, you know, it's hard to say which model is better. There is still a lot of qualitative evaluations and a lot of vibes involved.

43:43Stefano Ermon:Yeah. Like seeing this model better than the other model. But on the other hand, it's so important, right? You cannot do good engineering. You cannot do proper R &D here and I was tracking the things that matter. And so. You have to track them, but also respect that benchmarks need a bit of a grain of salt with them as well. But, you know, in the end it's, it's down to your tasks and what you're doing, where, where you find that, I guess. Exactly. I actually have a question related to this, which is a lot of our readers and viewers are trying to figure out how to implement these systems, like, like in their business, in their workflows, right.

44:20As someone who builds and has worked with these types of things, of models for years how do you recommend that they i guess close the loop in terms of implementing this and and actually assess the quality when you're actually in production scale like do you have any like tips or tricks there of like how basically like what's your recommendation for how

44:42Stefano Ermon:to assess whether they're working or not i mean ultimately what we see with our customers the the gold standard is some kind of a b test on a business on the business relevant metric i think Ultimately, that is the thing that decides whether people are going to buy, are going to switch to Mercury or not. There is some business metric they care about. And then you do an A-B test and then you see Mercury is better. And if it is and the cost is right and the reliability is there and all the other things that matter are there, then they switch. And so I think that's kind of like the gold standard.

45:19Stefano Ermon:Unfortunately, it takes infrastructure. Not every client has the maturity to be able to run AB tests properly. And it's also expensive. So, of course, ideally leading up to the AB test, you want to be able to have some kind of offline evaluation to kind of like guide the model selection and even just make any sense of is it even worth doing an AB test. So coming up with some good kind of like offline evals, which could be based on rubrics, other lamps as a judge, or some kind of proxy for the thing you think will matter in production, I think that's always a good step. It's going to save you time.

45:58Stefano Ermon:It's going to save you money. And it's always important to do that as well. Awesome. Well, I know we're getting tight on time, but I have one last question before I let you go. And I want to make sure we talk a little bit more about Mercury because it's been around a minute now, and it's surprisingly good for something that some people don't even know exists at this point and I'm wondering like what's what's on the roadmap what's what's coming up in the near future maybe the longer term future is there anything you want to share about that I'm glad you asked yeah so as I mentioned we're working hard on on R &D and improving you know the frontier of what's possible with diffusion language models.

46:44Stefano Ermon:We're going to be releasing new models constantly. There is one that we're very excited about that we hope to release publicly very soon that's going to enable Mercury models to be even smarter. It's going to have much better planning and reasoning capabilities. And so that's going to enable a lot of agentic use cases that people really care about. They're going to make them really, really fast so we're excited about the results that we're seeing on these new models and so hopefully we're going to be able to share them with you and your audience pretty soon. Amazing, amazing. Awesome. Stefano, thank you so much, man.

47:27This has been absolutely fascinating. We love hearing from you.

47:31Stefano Ermon:Thank you so much for hosting me. Yeah, this was really fun. Where should listeners go to give Mercury a try and learn more about Diffusion Language Models? So they can come to our website inceptionlabs.ai we have a number of resources blog posts on our models there's a chat playground there's documentation to use our API if you want to build an app using our models that's probably the best place to start excellent, excellent well if you haven't yet please take just a second to like, subscribe go hit the newsletter and make sure you keep up with Inception and Mercury because Stefano's doing some neat stuff and we're really excited we can bring him to you today but that's it for today so farewell for now humans we'll see you next time

From the publisher

Diffusion models changed how we generate images and video—now they’re coming for text.


In this episode, we sit down with Stefano Ermon, Stanford computer science professor and founder of Inception Labs, to unpack how diffusion works for language, why it can generate in parallel (instead of token-by-token), and what that means for latency, cost, and real-time AI products.


We talk through:

  • The simplest mental model for diffusion: generate a full draft, then refine it by “fixing mistakes”

  • Why today’s autoregressive LLM inference is often memory-bound—and why diffusion can shift it toward a more GPU-friendly compute profile

  • Where Mercury wins today (IDEs, voice/real-time agents, customer support, EdTech—anywhere humans can’t wait)

  • What changes (and what doesn’t) for long context and architecture choices

  • The real-world way to evaluate models in production: offline evals + the gold-standard A/B test

Stefano also shares what’s next on Mercury’s roadmap—especially around stronger planning and reasoning for agentic use cases.


Try Mercury + learn more: inceptionlabs.ai


For more practical, grounded conversations on AI systems that actually work, subscribe to The Neuron newsletter at https://theneuron.ai.

More from The Neuron: AI Explained

All 106 episodes
Diffusion for Text: Why Mercury Could Make LLMs 10x FasterThe Neuron: AI Explained · 49 min
Listen in VO