In short
Eye On A.I. Podcast Episode Notes
Episode Title
#311 Stefano Ermon: Why Diffusion Language Models Will Define the Next Generation of LLMs
Episode Overview In this episode, host Craig S. Smith interviews Stefano Ermon, CEO of Inception, discussing the advancements in language model technology, specifically the benefits of diffusion language models over traditional autoregressive models. The conversation covers the architecture, efficiency, and future applications of diffusion models, as well as their impact on coding and real-time AI applications.
Key Takeaways
- Diffusion Language Models (DLLMs):
- Operate on a parallel inference-first approach, contrasting with autoregressive models that generate text sequentially, one token at a time.
- Trained to remove noise from data rather than to predict the next token, allowing for simultaneous modifications across multiple tokens.
- Exhibit increased speed and reduced costs compared to autoregressive models.
- Advantages of DLLMs:
- Speed: Capable of generating answers quickly by utilizing parallel computations.
- Cost Efficiency: More efficient use of resources, resulting in lower operational costs.
- Context Handling: Context windows can be expanded, and while similar to traditional models in terms of context length, they possess unique advantages in managing larger contexts.
- Architecture and Training:
- Both DLLMs and autoregressive models utilize neural networks, specifically transformer architectures.
- Training involves adding noise to data and then training the model to reconstruct the original input, leading to improved data efficiency.
- Can use reinforcement learning from human feedback to enhance performance.
- Competition and Future Directions:
- Emergence of competing research in both academic and corporate environments, but Inception is currently leading in commercial DLLM applications.
- Ongoing research into improving reasoning capabilities within DLLMs for enhanced decision-making in agentic systems.
- Interest in developing multimodal models that integrate various data types (text, images, etc.).
Discussion Points Autoregressive vs. Diffusion Models
- Sequential Generation: Autoregressive models generate output token by token, which can lead to bottlenecks in speed.
- Parallel Generation: DLLMs, by contrast, allow for the refinement of multiple tokens simultaneously, which significantly enhances generation speed.
Contextual Limitations
- Context in autoregressive models scales quadratically, presenting challenges with larger context windows.
- DLLMs can manage context more effectively but still have limitations similar to autoregressive models.
Control and Safety
- DLLMs are perceived as more controllable due to their structure, allowing for real-time adjustments during the generation process.
Use Cases
- Primarily focused on code generation, where the demand for speed and accuracy is high.
- Other applications include voice agents and potentially broader uses in various AI-driven interfaces.
Current Developments
- Mercury Models: Inception's commercial-scale DLLM focused on code generation with promising results in speed and accuracy.
- Fine-Tuning: Models can be tailored for specific tasks using customer data for enhanced performance.
Challenges
- Addressing issues of hallucination (inaccurate or nonsensical outputs) remains a priority for improving the reliability of models.
Conclusion Stefano Ermon emphasizes the transformative potential of diffusion language models in the AI landscape, highlighting their speed, efficiency, and adaptability across multiple applications. As research and development continue, the expectation is that DLLMs will increasingly become a staple in AI solutions, potentially outperforming traditional autoregressive models in various domains.
Links and Resources
- Inception Labs API: [Inception Labs API](https://chat.inceptionlabs.ai)
- Eye on A.I. on X: [Eye On AI](https://x.com/EyeOn_AI)
- Craig Smith on X: [Craig Smith](https://x.com/craigss)
Episode Timestamp Breakdown
- 00:00 Autoregressive vs Diffusion LLMs
- 02:12 Why Build Diffusion LLMs
- 05:51 Context Window Limits
- 08:39 How Diffusion Works
- 11:58 Global vs Token Prediction
- 17:19 Model Control and Safety
- 19:48 Training and RLHF
- 22:35 Evaluating Diffusion Models
- 24:18 Diffusion LLM Competition
- 30:09 Why Start With Code
- 32:04 Enterprise Fine-Tuning
- 33:16 Speed vs Accuracy Tradeoffs
- 35:34 Diffusion vs Autoregressive Future
- 38:18 Coding Workflows in Practice
- 43:07 Voice and Real-Time Agents
- 44:59 Reasoning Diffusion Models
- 46:39 Multimodal AI Direction
- 50:10 Handling Hallucinations
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOStefano Ermon's Background
0:45 to 2:07
Stefano Ermon shares his background and expertise in generative AI.
“Now an open source Linux foundation project, Agency is building the internet of agents, a collaborative layer where AI agents can discover, connect, and work across any framework.”
Diffusion Language Models Explained
2:07 to 3:39
Exploration of diffusion language models and their advantages.
“I'm one of the founders and the CEO of Inception.”
Advantages of Diffusion Models
3:39 to 5:51
Discussion on how diffusion models outperform autoregressive models.
“I mean, it's fascinating that diffusion language models work.”
Context Problems in LLMs
5:51 to 8:12
Examining context challenges in LLMs and how diffusion models address them.
“and I've been talking to people about sort of post-transformer architectures, But one of the problems with LLMs is context, the compute required for context, for the context window scales quadratically.”
How Diffusion Training Works
8:12 to 11:51
Detailed explanation of the training process for diffusion language models.
“and you get the benefits and the downsides of those architectures.”
Understanding Token Prediction
11:51 to 14:01
Clarification on how diffusion models predict answers without token prediction.
“Because again, each neural network evaluation is able to modify multiple tokens as opposed to just one.”
Understanding Attention in Language Models
14:01 to 14:31
Learn how attention mechanisms help language models determine meaning and output corrections.
“It tries to figure out what is the meaning.”
Training Diffusion Language Models
14:37 to 16:38
Discover the training process for diffusion language models and their unique predictions.
“And I'm going to break this down and drill down a little bit to clear up my understanding.”
Controllability of Diffusion Models
16:44 to 19:09
Explore the controllability advantages of diffusion models over autoregressive models.
“And in the context of autoregressive language models, you always reconstruct left to right, one token at a time.”
Advanced Training Techniques
19:11 to 22:00
Learn about advanced techniques for training diffusion models and their efficiency.
“also in the context of diffusion language models, where there is still quite a bit of...”
Show all 24 chapters
Evaluating Language Model Performance
22:04 to 24:19
Understand the methods used to evaluate and track performance of language models.
“reinforcement learning from human feedback.”
Competitive Landscape of Diffusion Models
24:21 to 26:29
Examine the current competitive landscape and research efforts around diffusion models.
“And that is very valuable kind of research, because it allows you to compare apples to apples, like a small autoregressive model, a small diffusion language model, and see which one scales better.”
Scaling and Efficiency of Diffusion Models
26:31 to 28:00
Find out about the scaling laws and efficiency of diffusion language models in practice.
“It's linear scaling, not in the context window, but is it the same?”
Advancements in Model Training
28:00 to 28:20
Learn about rapid progress in training competitive models.
“But also on the training side, we've been able to quickly come up with models that are competitive with state-of-the-art, speed-optimized models from Frontier Labs.”
Focus on Code Assistance with Mercury Model
29:40 to 31:40
Explore the Mercury model's unique benefits focusing on coding.
“And you're talking about Mercury is the publicly released model.”
Fine-Tuning and Model Performance
31:40 to 34:30
Understand how model size and fine-tuning affect performance.
“where people are getting most of the value, I would say today, from using LLens to write code faster and find the coding and all those kind of things that people are so excited about.”
Diffusion vs. Autoregressive Models
34:30 to 37:10
Learn about the competition between diffusion and autoregressive models.
“And so the way we're thinking about it is that there is a Pareto frontier between quality and speed.”
Applications of Coding Assistance Models
37:10 to 42:00
Discover how coding assistance enhances developer workflows.
“And you're going to get value by having a completely different stack to build intelligent machines, in fact.”
Interactive Code Generation with LLMs
42:00 to 43:19
Learn about the benefits of interactive workflows in code generation using LLMs.
“to build an app from scratch, what people call vibe coding.”
Challenges and Future Directions in LLM Research
43:20 to 44:38
Explore current challenges and future research directions for diffusion language models.
“for other applications that are latency sensitive.”
Advancements in Reasoning Capabilities for LLMs
44:39 to 45:59
Discover how enhancing reasoning capabilities can improve LLM performance.
“We often just take whatever works for autoregressive models and we just use it in the context of diffusion language models whenever that's possible.”
Multi-Modality in Generative AI Systems
46:00 to 47:28
Understand the potential of multi-modality in creating powerful generative AI systems.
“And that would have a big impact on their use in agentic systems, is that right?”
Real-World Applications of Diffusion Language Models
47:29 to 50:09
Learn about the current applications and future potential of diffusion language models.
“I mean, I'm going to be talking to Fei-Fei Li, your colleague, or erstwhile colleague, about world labs.”
Addressing Hallucinations in LLMs
50:10 to 51:24
Explore strategies for managing hallucinations in language models and their impact.
“And we started, before I started recording, talking about hallucinations.”
Transcript
Automatic transcript. May contain errors.0:00When you think about most existing LLMs, chat GPT is the Gemini's they're also called autoregressive and what that means is that when you ask a question they will give you an answer and it will generate the answer one word or one token as we call it at a time and that generation is sequential and that's kind of like a structural bottleneck that cannot be overcome because it's kind of like built in into the autoregressive formulation a diffusion language model is also using a neural network under the hood and in our case, it's also a transformer neural network. But the task that the model is trained on is quite different.
0:35The model is not trained to predict the next token. Instead, it's trained to remove noise. Build the future of multi-agent software with Agency. That's A-G-N-T-C-Y. Now an open source Linux foundation project, Agency is building the internet of agents, a collaborative layer where AI agents can discover, connect, and work across any framework. All the pieces engineers need to deploy multi-agent systems now belong to everyone who builds on agency, including robust identity and access management that ensures every agent is authenticated and trusted before interacting. Agency also provides open standardized tools for agent discovery, seamless protocols for agent-to-agent communication, and modular components for scalable workflows.
1:37collaborate with developers from cisco dell technologies google cloud oracle red hat and more than 75 other supporting companies to build next generation ai infrastructure together agency is dropping code specs and services no strings attached visit agency.org to contribute that's A-G-N-T-C-Y dot O-R-G. Hi, everyone. My name is Stefano Hermon. I'm one of the founders and the CEO of Inception. Before starting Inception, I was a professor of computer science at Stanford for about a decade. I've been working on generative models, what's now called generative AI, for about 10 years. I did very early work on diffusion models, the technology that is now powering a lot of the state-of-the-art image generation systems, video generation systems.
2:46I worked on a variety of techniques like flash attention, DPO, that are now widely used in industry to build large-scale other labs and other kinds of generative AI solutions. I've been working on diffusion language models for a while. in my lab at Stanford. So kind of trying to figure out how to get diffusion models to work not just on images and videos, but also on text and code generation. Had some breakthroughs in my lab last year where we showed that for the first time, this technology was able to compete with the leading autoregressive models. And then that prompted me to start the company to scale up the models.
3:28And that's where I've been for the last year or so. And yeah, we've been pretty successful at training the first commercial scale diffusion language models that we call Mercury. Yeah. And I'm curious. I mean, it's fascinating that diffusion language models work. Maybe we'll talk about that first. But I want to ask, is there an advantage to a diffusion language model over an autoregressive transformer-based LLM? Yes, for sure. And the reason is that when you think about most existing LLMs, like the ChatGPT, the Gemini, the Cloud, they're also called autoregressive. And what that means is that when you ask a question, they will give you an answer, and it will generate the answer one word or one token, as we call it, at a time.
4:25And that generation is sequential, and that's kind of like a structural bottleneck that cannot be overcome because it's kind of like built in into the autoregressive formulation. The neural network is trained to predict the next token, and that's how you use it at the inference time when you ask a question to it. And sequential computations are generally very hard to accelerate because there's not many things you can do in parallel if you have to go through a token spiral at a time. On the other hand, a diffusion language model is generating the answer all at once. And so the neural network is trained to refine its answers, to kind of fix mistakes.
5:09And then at the inference time, you kind of use this procedure where each neural network evaluation is able to essentially modify multiple tokens at the same time instead of just one. And so what we see is that these models can be much, much faster compared to ultra-aggressive models of similar quality. And that's kind of the key benefit. We're able to drastically speed up the generation, speed the time it takes basically to generate answers with an LLM and reduce costs. That's the other big benefit. the computation is more parallel, we can make use of parallel compute of GPUs more effectively, and we can effectively reduce costs for our users.
5:50Yeah, the other, I don't know that this is an advantage, but one of the problems with LLMs, and I've been talking to people about sort of post-transformer architectures, But one of the problems with LLMs is context, the compute required for context, for the context window scales quadratically. So it very quickly becomes too expensive. And then there are all kinds of problems with context windows that are too large. the LLM gets confused or lost. Is that something that diffusion LLMs would address? Yeah, that's a great question. And yeah, I'm very familiar with these issues. I was one of the original authors of the flash attention paper that is widely used to accelerate the very expensive quadratic computation that is required to do attention in a transformer.
7:01We have shared in the past, and I'm happy to confirm, that our diffusion language models are also based on transformers under the hood. So the architecture that we use is still a transformer with self-attention. And what this means is that we got the same pros and cons with respect to the context length that you would expect from an autoregressive LLM. And so the current models that we serve in production have about 130K tokens of context. We think we can expand it using the techniques that people have used for autoregressive LLMs. But in terms of pros and cons, there is not really any benefit or any downside compared to autoregressive models.
7:49We've also used different architectures. So just like people in the ultra-aggressive lab world that they've been experimenting with state-space models and Mambas and alternative architectures, that's also possible in the context of diffusion language models. You can use any neural network architecture you want inside the diffusion language model. And we have some prototypes where we've replaced transformers with these alternative architectures. and you get the benefits and the downsides of those architectures. So fusion versus architecture, those are kind of like two orthogonal axes that you can use to accelerate the computation or reduce the memory footprint or kind of like trade-off different kinds of performance characteristics you might hear about.
8:38Yeah, well, for listeners and for myself, frankly, can you explain how the fusion works? And how transformers play a role. Yeah, just, I mean, I know that there's, you inject noise during training and then it sort of reverses the noise down to the output. But yeah, if you could explain it for a layman, that would be terrific. Sure. So both autoregressive LLMs and diffusion LLMs rely on neural networks under the hood. So there is a neural network that takes as input a sentence and then outputs something. And in the context of autoregressive LLM, this neural network, which is often a transformer, takes as input a sentence and then it predicts what is the next token, what is the next word that I should append to that sentence to make sense.
9:43and you can train this neural network to do a pretty good job at it by just feeding a lot of text or a lot of code to these models and then eventually, after lots of training iterations, they become pretty good at understanding what kind of tokens make sense and how you should complete sentences. A diffusion language model is also using a neural network under the hood, and in our case, it's also a transformer neural network, but the task that the model is trained on is quite different. The model is not trained to predict the next token. Instead, it's trained to remove noise. So again, you start with sentences of code that can be maybe taken from the internet.
10:28And then what you do is you artificially add noise. Perhaps you mask some tokens or you change the value of some of the tokens. And then you train the neural network to take as input this corrupted sentence and try to fix the mistakes, try to give you back the original sentence. And so it's a different training objective. It's not next token prediction, but it's more like a global transformation of the sentence where you try to fix the mistakes. And at the inference time, these two models are used in a very different way. The traditional autoregressive model is just going to predict the next token, then you fit it in, and then you predict the next token, and then you fit it in, and you predict the next token, and so forth.
11:10But a diffusion language model, instead, it starts with a guess of what the answer should be. And in theory, it can be anything. And then you pass this whole guess through the neural network. The neural network will fix some of the mistakes. And crucially, it will fix the mistakes across all tokens, across all words that form the sentence. So it's much more parallel. And then you repeat this denoising process a bunch of times. And then at the end, you get an answer that we serve to our users. And as long as you don't need too many denoising steps, you don't need too many diffusion steps, this can be way more efficient.
11:51Because again, each neural network evaluation is able to modify multiple tokens as opposed to just one. Yeah. In the autoregressive model, I understand the token prediction and the attention mechanism. So it's really looking for the most probable next token. in diffusion, I don't understand, because it's not predicting the next token. So how does the model know the answer? I mean, know, is it looking across a probability distribution of the input? Yeah, I just don't understand that at all. Yeah, so it's, you know, like at the first time, yeah, the model never knows what's the correct answer. But during training, you have, from truth data, you have text and code that you've collected on the internet, for example.
13:02And so you do know what is the correct answer. So you do know what was the right token that came after a given context, a given sentence. Similarly, in the context of a diffusion language model, we start with ground truth text, we artificially add noise, we introduce some mistakes, we mask out some of the tokens, and then we feed this corrupted sentence to the neural network, and we ask, we train the model to reconstruct the original signal. And so during training, we do have access to ground truth, just like in the context of a lot of aggressive models. It's just a different kind of task that the neural network is trained on.
13:46It's trained to predict and modify multiple tokens at the same time. and under the hood, it does this by using attention. And so what it does is it takes as input this corrected sentence. It does attention and so it kind of like tries to compare every token with every other token. It tries to figure out what is the meaning. It tries to figure out what should have been the right words that go in every spot. And then it outputs its best guess of what the fix should be. And yeah, they're not perfect, but if you train them for a sufficiently long amount of time on a sufficiently large amount of data, they actually become pretty good at correcting mistakes, at fixing mistakes.
14:30And that's why they actually work quite well during Fresno. Yeah. And I'm going to break this down and drill down a little bit to clear up my understanding. So in the training, you're taking existing language, you know, diffusing it or, you know, whatever the term is so that it's full of noise and denoising it back down to the original sentence. That's right. But where does the prediction come from to answer the question? Is that purely from the attention mechanism or is there some next token prediction happening in the background before the denoising? Yeah, so it's not just the max token in the sense that that's the other advantage of the model, is that it's not necessarily just sequentially going through one token at a time.
15:46It's actually able to use context to the left and to the right to fix mistakes. So if the first token is missing and everything else is known in the sentence, then the diffusion language model can try to predict what should the first token be, given that I know everything else. So it's kind of like a generalization of the task that you would normally do in an autoregressive model where you always just predict one token left to right. In a diffusion language model, you have a richer set of tasks that you need to be able to solve during training, which could be just predicting the next token, but maybe it's predicting the first token, given that you know everything to the right of it.
16:30And so it's more like right to left prediction. And it's actually even more complicated because there's all kinds of mistakes that could be there in the sentences, and the model needs to figure out how to fix all of them at the same time. But in some sense, they're kind of similar in the sense that in both cases, you try to reconstruct part of the input. And in the context of autoregressive language models, you always reconstruct left to right, one token at a time. In the context of a diffusion language model, you reconstruct the sentence all at once globally. And so it's a more complex kind of reconstruction, denoising task.
17:15But we can still train neural networks to solve it. Yeah. On image diffusion models, they're sort of notoriously hard to control. they'll certainly give you an output but it it it may not be what you intended even if it fits the prompt how how does that I mean and there's been a lot of progress on that but but how does that problem relate to language yeah in my opinion a diffusion model is actually much more controllable than a typical autoregressive model. And the reason is that in a diffusion model, you kind of have access to the whole object from the very beginning. And so you kind of know from the very beginning whether or not it's consistent with the prompt, whether or not it satisfies some safety constraints or whatever is the control signal that you want to use to steer the generation.
18:23Throughout the generative process, you always know whether or not you're going in the right direction, whether or not the constraints are satisfied. And that's why actually diffusion models are relatively easy to control. And there's a lot of techniques that people have come up with to steer the generation process in a certain direction. Because again, you can kind of like check whether or not you're satisfying the constraints and you can kind of push the generation in the direction that you want. And in fact, instead, if you think about an autoregressive model, you kind of need to wait until you get to the very end, to the last pixel, to know whether or not this is satisfying the constraints.
19:03And so that's why, in my opinion, diffusion models are more controllable than autoregressive models. And we're starting to explore some of these benefits also in the context of diffusion language models, where there is still quite a bit of... It's a problem that comes up sometimes, that you want the model to behave in a certain way. Of course, you want the model to be safe, but maybe you want the model to be on brand. There are more complex ways of defining what you want the model to do, and we're starting to explore ways to use our diffusion language technology to allow our customers to control the language models in a more fine-grained way.
19:48Yeah. Now, the, okay, let me gather my thoughts here.
20:00So, well, I've got a bunch of different questions. One is about training. How do you train these? The other is about guiding them to, or, you know, in autoregressive models, they use reinforcement learning with human feedback. Are they the same kinds? Because I'm going to cut my stumbling around here out. How do you train these? Is it the same process as training an autoregressive model? Yeah, so at a high level, it's kind of like what I discussed. We're training this neural networks to denoise and the inputs to remove noise from sentences. We haven't disclosed a lot of the details in terms of the exact training objective or the exact size of the models or the model training data that we use.
21:09There are certain papers that are publicly available or in the open where people have studied the scaling laws and the performance that you see if you train diffusion language models versus autoregressive models. And there is some evidence in the academic literature that diffusion language models are much more data efficient in the sense that kind of the intuition is that because you're training the model to solve many different kinds of tasks, not just predict the next token, but let's say predict in any order, then they tend to be more data efficient. So you need less training data to achieve a certain level of quality.
21:52And of course, everything depends on the devil is in the detail and there is all kinds of things that you can do, but we have our own recipe and we've seen that it does really, really well. On top of that, it's also possible to do reinforcement learning from human feedback. It's also possible to do reinforcement learning using verifiable words. If you are in the context of, let's say, coding or places or math, where there is a verifiable word, there is a uniquely correct or wrong answer that you can use to check the quality of the results. And internally, we've been able to use all of these techniques.
22:28and adapt them to the different kinds of generative models that we're dealing at this action. Yeah, and so for evaluation, do you use other diffusion models for evaluation or do you use autoregressive LLMs in your evaluation? So it's, I mean, evaluating LLMs is notoriously difficult and the way we track performance internally is by using benchmarks. So there are a variety of benchmarks that the community has developed that are trying to measure how good the models are at different tasks that people care about, like how good they are at math, how good they are at following instructions, how good they are at doing question answering, how good are they at writing code in different kinds of languages.
23:25And so internally, what we do is we monitor the performance of our models across a variety of benchmarks, and then we keep adjusting our training methods and the architectures and the inference schemes to try to get as good a performance as we can while keeping the computation costs down, trying to also make the models as efficient as possible. So on the evaluation side, it's actually exactly the same. It's all very, very similar to what you would do with a traditional other lab. Yeah. And what are the sizes, parameter sizes of the models that you have to date? Yeah. Unfortunately, we cannot disclose the exact size of the models.
24:11We are in a very competitive kind of space and that's information that we're not able to disclose at this time. Well, that's interesting. a competitive space. Are there other people working on diffusion models, diffusion LLMs? Diffusion LLMs, yes. So there is a fairly active community in academia that is training diffusion language models, relatively small scale models, including from my lab at Stanford, I mean, early work at the GPT-2 kind of size, where we trained models with less than a billion parameters, let's say. And that is very valuable kind of research, because it allows you to compare apples to apples, like a small autoregressive model, a small diffusion language model, and see which one scales better.
25:04But these models that are sort of like either open source or that have been developed in academia, they don't perform particularly well because the scale is relatively small and the kind of data that was used to print them is not exactly the best that is available out there. There have been attempts at building diffusion language models in other industry labs. There's been a Gemini Diffusion paper and release, blog posts at least, where researchers from Google, from DeepMind, have talked about their internal efforts around building a diffusion language model. So far, it's not been deployed in production.
25:51It's a demo, but it's not available, as far as I know, to customers. It's not serving production traffic. So as far as I know, Inception is the only player right now with a working model that is used to sort of production traffic, who were the first ones to release a commercial scale model. And we're still way ahead of the competition in terms of like how mature our models are and our cervix engine and kind of like inference stack compared to anybody else. But there is certainly research attempts in different places and people are actively exploring this direction. Yeah, and the scaling of the transformer-based models, that's part of what's made them attractive.
26:40It's linear scaling, not in the context window, but is it the same? Do you benefit from the same scaling laws in these DLLMs? Yeah, there's certainly certain where we're observing is predictable sort of like behaviors that we can use to guide our experiments and decide what kind of architecture and design choices to prioritize. And again, there have been papers in the literature published showing that diffusion language models can scale better than ultra-aggressive lamps. And that's part of the reason we're so excited about this. It's not only extremely efficient at inference time, so the scaling at inference time is strictly better than what's possible with an ultra-aggressive model, which in my opinion is the thing that matters because ultimately it's all going to be an inference game.
27:39People are building these massive data centers. It's not for training. It's because they expect everyone will be using AI, generative AIs, LLMs. They're going to be everywhere. There's going to be a massive need to serve these models efficiently, quickly, cheaply. And the technology that scales better is the one that is going to win. And that's why we're so excited about this. But also on the training side, we've been able to quickly come up with models that are competitive with state-of-the-art, speed-optimized models from Frontier Labs. So we're making rapid progress in terms of increasing the capabilities of the models.
28:17Build the future of multi-agent software with Agency. That's A-G-N-T-C-Y. Now an open-source Linux foundation project, agency is building the Internet of Agents, a collaborative layer where AI agents can discover, connect, and work across any framework. All the pieces engineers need to deploy multi-agent systems now belong to everyone who builds on agency, including robust identity and access management that ensures every agent is authenticated and trusted before interacting. Agency also provides open, standardized tools for agent discovery, seamless protocols for agent-to-agent communication, and modular components for scalable workflows.
29:15Collaborate with developers from Cisco, Dell Technologies, Google Cloud, Oracle, Red Hat, and more than 75 other supporting companies to build next generation AI infrastructure together. Agency is dropping code, specs and services, no strings attached. Visit agency.org to contribute. That's A-G-N-T-C-Y dot O-R-G. Yeah. And you're talking about Mercury is the publicly released model. Right. You're talking about it as a code assistant. Is that right, primarily? Is there a reason why you're focused on writing code? Yeah. So we started with code because it felt like, well, first of all, we're all computer scientists.
30:14And so code is pretty close to our heart and we understand if the model is doing well or not. It's also a pretty interesting application for the Fusion Language models, because it's a little bit less autoregressive, in the sense that this bias of going left to right kind of makes sense if you're thinking about writing a document. It's kind of like how we write, although I would argue that maybe there is still a little bit of course-to-find generation where we start with an outline and then we fill in the details. but at least I would say if you think about code, there is a lot less left to write bias.
30:56Like, if you think about different files that are not necessarily in a particular order, there are many more tasks where you really want to be able to look at a whole code base and infill a particular function or functionality that you're trying to develop. And that's where the future models can shine because they are not constrained to be the goal. So this idea of going left to right is not baked in. So that's why we kind of focused on code initially because, yeah, it's an area where we were expecting to see big improvements. And also it's very commercially, and it's super important from a commercial point of view.
31:36That's where we're seeing a lot of the usage of LLMs where people are getting most of the value, I would say today, from using LLens to write code faster and find the coding and all those kind of things that people are so excited about. And so it was a pretty natural kind of like first step for us. Yeah. And can these models be fine-tuned in the same way that Autoregressed? Absolutely, yes. We can fine-tune them for specific tasks. And in fact, that's one of the modes when we engage with enterprise customers. We sometimes are able to, if they're able to provide internal data or specific data that they care about, we're actually able to fine-tune our models on their data and deliver even higher performance on the tasks they care about.
32:32Yeah. And is it the same as with, I mean, you said you can't disclose parameter counts, but is it the same sort of progression that the larger the number of parameters, the more accurate the model gets? Yes, typically yes. The bigger the model, the higher. I mean, as long as you trained with a proportionally larger amount of training data, so there are some caveats, but typically you can expect larger models to perform better, but they also become more expensive and slower. So that's kind of like the trade-off that we're facing. Yeah. And from your point of view, the advantage is speed and cost or power efficiency or not accuracy or, yeah.
Read the full transcript
33:30How do you measure accuracy against the retro, I mean, the autoregressive? Yeah, so it depends on you. If you think about the different LLMs that exist today, there is always a trade-off between quality, accuracy, cost, and speed. And if you want the fastest models, they are typically not the most accurate because they are smaller, because people need to use all kinds of tricks to make them as efficient as possible. And so what we can achieve is we are actually, for a Shoran speed profile, we are the most accurate models. We are not the most accurate models overall. So there are larger models that have been trained by Penai, Anathropic, and Gemini that are more accurate in an absolute sense compared to where we are today.
34:25But in a Shoran speed class, we are the most accurate model. And so the way we're thinking about it is that there is a Pareto frontier between quality and speed. We've been able to shift that Pareto frontier, or at least for a certain quality level, and the next stage of the company is to kind of keep improving the quality of diffusion language models until we actually catch up with frontier-level quality intelligence that right now is only available in a few models that are only out of address. Yeah. And I mean, do you see diffusion language models as an alternative to our aggressive LLMs or or do you do you see diffusion LLMs as as, you know, competing with auto aggressive LLMs?
35:26and with the hope that someday they'll surpass autoregressive LLMs? Yeah, we certainly... They're functionally a replacement of each other. Whenever you can use an autoregressive LLM, you can also use a diffusion language model. So from that perspective, they are certainly competing solutions for the same problem. Even right now, we are surpassing autoregressive LLMs and our customers are switching from autoregressive Lens to diffusion language models because there are a number of low latency applications where you need to be able to give an answer quickly to a developer or to a customer because you're building a voice agent.
36:12There are a variety of applications where latency matters and our models for that latency budget are actually more accurate than the autoregressive models that people would use in order to be sufficiently quick at outputting an answer. And so the question, of course, is what will happen as we scale our models, as diffusion language models get better? We don't know. I can see a future where all models will be diffusion-based. Usually, if I look back at the history of computer science, the more parallel solution is the one that wins, and the diffusion language models are built to be parallel from the grounds up.
36:56And so that's why we're betting on this. But we don't know. It's also possible that multiple technologies will exist. Then maybe one would be better for certain use cases. Another one will be better for other use cases. It's entirely possible that they're going to make, maybe neither of them will be perfect. They're going to make mistakes. But as long as the mistakes are different, that they are not correlated with each other, then there's going to be value in having agents that are maybe autoregressive based on autoregressive models, based on diffusion models. They can interact with each other.
37:28They can learn from each other. And you're going to get value by having a completely different stack to build intelligent machines, in fact. Yeah, actually, that's an interesting idea. You could have a system where the autoregressive model is checking the work of a diffusion model and back and forth to refine the answer. Is that something you've experimented with? Not yet, but I think, yes, it's something that in the future I can see how that could work. There's going to be humans, there's going to be agents, and the agents are going to be based on different technologies. And as long as they complement each other in interesting ways, they're going to be useful in combination.
38:13And the sum is going to be greater than the parts. Yeah. And for coding, how do you see people using it? Is it really as a coding assistant? How accurate is it compared to some of the best coding models? Are people writing complete code bases using a DLLM? Yeah. Yeah, so if you think about code assistance, there is different types of use cases of LLMs. The simplest one and the first one that was actually deployed is autocomplete. It's kind of like you're writing code and the LLM will suggest how you should complete that line or those lines of code. That's where we started. That's an application that is very latency sensitive.
39:12A developer is not going to wait a minute to get a suggestion. The suggestion has to be instant. It has to be very quick for the developer to be in the flow and for the experience to be good. and what we have right now our Mercury models are state of the art in terms of code completions that they provide there is this benchmark called Coopilot Arena it's run by researchers at CMU and it's a benchmark that they use to evaluate different kinds of LLMs specifically for autocomplete and the way it works is that they have an IDE and there are developers around the world that are using this IDE and as they write code they get to see not just one completion but they get two completions and these two completions they're coming from two different models the developers they don't know which models they are using and then they vote which one is better and then they come up with an ELOR score an ELOR ranking same thing that people use for chess players to kind of like compare different LLMs in terms of their quality as measured by how often developers prefer one LLM versus the other.
40:28And our Mercury models are currently number one tied with a few other LLMs, but we are number one in terms of quality and we are number one by a fair margin in terms of speed. So it's already quite good as evaluated by external benchmarks. And even if you look at adoption, where the default in quite a few coding IDs and in a few other places where developers can choose the model, they bring your own key where developers can plan and play different models. There are a lot of developers that are using Mercury models to generate better autocomplete suggestions in their workflows. The other space where we're seeing a lot of adoption is NextEdit, which is kind of like autocomplete, a more general version of autocomplete, where you're not just suggesting how to complete a sentence, but you're suggesting more general edits.
41:27Maybe you want to delete some code. Maybe you want to rename some variables in the code based on some changes that you're doing. That's another space where we have a very, very good model that is being deployed as the default, again, in a variety of IDs. With the latest model, we're also able to do more agentic kind of use cases, the kind of things you were referring to, where maybe a developer or somebody who is not maybe even an expert can just ask the model to write code to build an app from scratch, what people call vibe coding. And with the latest release of the Mercury model, we're now actually able to support also those use cases.
42:12And because the models are so fast, the experience is actually a little bit different. Like you get answers quickly, and then if you want to still interact with the model, you want to provide suggestions, you want to ask the model to make changes, the workflow is much more interactive. And so we're seeing a lot of usage, I think, based on the fact that people like to be able to interact more closely with the LLM, provide suggestions. And this idea of the LLM will do everything in one go, I don't even know if that's ever going to be the future. I think there's always going to be a need for guidance and interaction.
42:53And if you want to have a human in the loop, speed is the limiting factor. you want to maximize the number of interactions, the back and forth between the LLM and the developer. And that's where the Fusion LLMs shine compared to autoregressive models. Yeah, that's fascinating. Where do you see this going? You're going to remain focused on code generation, is that right? Code is a big area of focus, but we have customers that are using the models for other applications that are latency sensitive. One big one is voice agents. So if you're building a voice agent, the typical pipeline involves ESR.
43:34Then there is an LLM that perhaps does, sometimes does tool call it, maybe it will check the menu or it will check the calendar if you're making reservations. And then the LLM will suggest an answer that then will be passed through a text-to-speech system that then communicates back to the user. And whether this is customer support or it's like educational agents, they are all kind of like built on a very similar kind of framework. Yeah. Yeah, the LLM in the middle is a key bottleneck in terms of like the overall latency. And what we're seeing is that people are getting really, really good results by replacing the small autoregressive LLMs that they were using before with the diffusion language models.
44:22Yeah. And going forward, is it a matter of, I mean, are you still doing research in this area? And is it a matter of scaling up the size of the models? Or you mentioned before replacing the transformer element with some other architecture, with Mamba, for example. is do you see that there's some other direction that you may be going with yeah for sure I think we were just scratching the surface of what's possible with diffusion language models that's why I'm so excited about these technologies that you know everything is pretty new pretty fresh I think there is still a lot of low-hackney fruits you know a lot of design choices I think are suboptimal.
45:19We often just take whatever works for autoregressive models and we just use it in the context of diffusion language models whenever that's possible. But I think those choices are not necessarily the best ones given that the architecture is different, given that the model is dramatically different. And so there is still a lot of opportunities for research, development, R &D on the models. that one of the things we're working on is a reasoning diffusion language model. That's a capability that autoregressive models have, and it's one of the ways they can drastically improve their quality, the ability to think before outputting an answer.
46:02And there are interesting ways of achieving those kind of capabilities in the context of diffusion language models, and it's one of the areas where we're investing heavily on R &D to come up with the first reasoning diffusion language. Wow, that's fascinating. And that would have a big impact on their use in agentic systems, is that right? Exactly, exactly. Because if you think about agents, there's often a lot of planning and reasoning involved, and those kind of capabilities become much stronger once you have a reasoning model under the hood. And are you looking at doing multi-modality? I mean, diffusion is already sort of the dominant player in image and video generation.
46:53Yeah, that's another reason we are so excited about diffusion language models is that, you know, we already know that diffusion are the best model for image generation, world models, video generation. once we crack and figure out how to do text and code generation, this is kind of like opening the way for a single generative AI system that can handle different modalities and learn across the different modes. It's something that we're very excited about. And yeah, it's another thing that we are working on and hopefully we'll have some models and some announcements coming out. Yeah. I mean, I'm going to be talking to Fei-Fei Li, your colleague, or erstwhile colleague, about world labs.
47:39I think that's what she's calling it. Remarkable, this marble is a remarkable model. Can you talk at all about how diffusion LLMs could fit into world models? because world models are amazing, but what's really going to be amazing is when you have a world model that can perceive the world and understand the laws of physics and then combine it with language. Exactly, yeah. I think there are very interesting opportunities once you combine different modalities because as you said, Yes, you could have a model that has been trained on physics, books, and how it works, and simulation code, and can write code to figure out what's going to happen and make predictions about the future.
48:45And I think that's just going to enable much stronger world modeling capabilities. if under the hood this model has been trained on not just video data or just 3D data, but really understands the world through all the available literature, all the available scientific knowledge that it should have access to. Is there anything I haven't asked that we should talk about? I mean, I think one thing that is, I think, exciting and unique about Inception is really that the models are available out there. It's not just the research prototype. It's a model that developers can use right now. businesses can build apps and they can replace autoregressive LLMs with diffusion language models today.
49:44Everything is backwards compatible. The Lady API is the same. It's kind of like an OpenAI compatible endpoint. And so this is happening and I think your viewers might be interested in seeing that this is and be aware of this transformation and be prepared for what might change the way they build things, they develop apps, and so forth. Yeah. And we started, before I started recording, talking about hallucinations. How do you handle hallucinations? I mean, there's certainly been a lot of work done on the autoregressive side. And is it the same strategies for a diffusion LLM? And is that a focus of refining these models?
50:42Yes, for sure. The models are not perfect and they make mistakes. And yeah, we try to measure it using benchmarks as much as possible. And some of them, let's say when you ask questions, related to question and answer, or general knowledge. We try to see where the gaps are, and then we try to fix the issues by incorporating additional training data, additional preference data, or other ways of supervising the model to try to do the right thing. And I think it's still an open problem, Like even for autoregressive LLens, hallucinations do happen. Mistakes do happen. And it's less and less as the capabilities improve.
51:33But I think, yeah, that's a super high priority for us right now to keep improving the quality of the answers, the accuracy of the answers that we provide to our users because that just makes the whole experience better. Yeah. So if someone wants to try Mercury, they go to remind us of the URL. Yeah, chat.inceptionlabs.ai that's where you can find a chat interface to play with the model and the platform.inceptionlabs.ai to get access to our API and you can sign up for a key. We give 10 million tokens for free to our users to get started and then you can start building some cool apps using really, really fast the fusion memories.
From the publisher
This episode is sponsored by AGNTCY. Unlock agents at scale with an open Internet of Agents.
Visit https://agntcy.org/ and add your support.
Most large language models today generate text one token at a time. That design choice creates a hard limit on speed, cost, and scalability.
In this episode of Eye on AI, Stefano Ermon breaks down diffusion language models and why a parallel, inference-first approach could define the next generation of LLMs. We explore how diffusion models differ from autoregressive systems, why inference efficiency matters more than training scale, and what this shift means for real-time AI applications like code generation, agents, and voice systems.
This conversation goes deep into AI architecture, model controllability, latency, cost trade-offs, and the future of generative intelligence as AI moves from demos to production-scale systems.
Stay Updated:
Craig Smith on X: https://x.com/craigss
Eye on A.I. on X: https://x.com/EyeOn_AI
(00:00) Autoregressive vs Diffusion LLMs
(02:12) Why Build Diffusion LLMs
(05:51) Context Window Limits
(08:39) How Diffusion Works
(11:58) Global vs Token Prediction
(17:19) Model Control and Safety
(19:48) Training and RLHF
(22:35) Evaluating Diffusion Models
(24:18) Diffusion LLM Competition
(30:09) Why Start With Code
(32:04) Enterprise Fine-Tuning
(33:16) Speed vs Accuracy Tradeoffs
(35:34) Diffusion vs Autoregressive Future
(38:18) Coding Workflows in Practice
(43:07) Voice and Real-Time Agents
(44:59) Reasoning Diffusion Models
(46:39) Multimodal AI Direction
(50:10) Handling Hallucinations




