In short
Cursor’s Composer2 training and the distributed RL infrastructure behind it, including why environments must closely mimic real user computers to prevent “cheating,” plus systems tricks to run massive rollouts efficiently.
Guests
Federico (research lead on Composer2 at Cursor; focuses on model training approach). Dima (worked for months at Cursor supporting infrastructure for the large-scale training runs; also discusses distributed systems and throughput tradeoffs).
Key claims
Composer2 is trained by combining continual pre-training (mid-training on code tokens) with large-scale reinforcement learning on Cursor’s harness. RL rollouts require orchestrating long agent sessions (tool use, multi-turn coding) and efficient inference. Models can detect fake environments; RL can encourage reward hacking. Cursor uses globally distributed inference clusters and ships weight deltas (lossless compression) to reduce staleness. Numerical nondeterminism (especially in MOE routing) can break RL; they mitigate via kernel alignment and “router replay.” Online RL uses user feedback but is gated by quality.
Notable examples
rollouts as ~50-turn Cursor sessions; self-summarization to extend effective context to millions of tokens; using production traffic off-peak for inference GPUs; reward signals via verifiable checks or LLM judges (details “top secret”).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOIntroduction to RL Environments
0:00 to 0:54
Understanding the importance of realistic environments for RL training.
“You need all the infrastructure to run these environments that have to mimic as closely as possible what a user's computer would look like.”
The Launch of Composer 2
1:30 to 2:31
Exploring Cursor's transition from an application to a foundation model company.
“For those who haven't been following as closely, Cursor recently announced Composer 2, which is an agentic coding model meant for long-horizon coding tasks.”
Designing for Specialization
2:31 to 4:28
How Composer 2 is specialized for software engineering tasks.
“Also, as people may have noticed, Composer is an order of magnitude less expensive than Opus and other coding models because we can just simply specialize all of the model weights to that particular task.”
Application Evolution and User Data
4:28 to 6:01
Discussion on the evolution of applications and the importance of user data for model training.
“Like the way we kind of view it with Fireworks is that when you're trying to do optimization, you had this like three-dimensional trade-off between quality, speed, and cost.”
Understanding Training Methodology
6:01 to 7:48
Insight into the mid-training and reinforcement learning processes of Composer 2.
“And so if we want to saturate all that capacity, we need to scale data.”
The Learning Process of Composer 2
7:48 to 9:18
How Composer 2 learns from libraries of code and correct code writing.
“And what is the model roughly learning in the kind of mid-training step?”
Challenges in Reinforcement Learning
9:18 to 12:04
Exploring the complexities and challenges in the reinforcement learning phase.
“is we're kind of tuning the feature of the model saying, hey, now you got to write correct code all the time.”
Optimizing GPU Utilization
12:04 to 14:01
Strategies Cursor employs to maximize GPU efficiency during training.
“are a lot of trade-offs how you can kind of co-optimize and co-design the system.”
Maximizing GPU Efficiency in RL
14:01 to 16:24
Learn how Cursor optimizes GPU usage for RL training.
“yeah, you have higher compute efficiency, you can get to a better model in a smaller amount of time.”
Challenges of Global Distribution in RL
16:25 to 18:26
Understand the complexities of globally distributing RL training infrastructure.
“more efficient and more precise rather than like spin up like an inference effort.”
Show all 21 chapters
Optimizing Weight Shipping and Model Updates
18:27 to 21:03
Explore methods for efficiently shipping model weights across clusters.
“You can have smaller groups of GPUs interconnected together.”
Handling Numerical Mismatch in RL
21:04 to 24:35
Discover how numerical mismatch affects RL training and inference.
“And you're going to probably spend, maybe you're going to allocate like one third to training and two thirds to inference.”
Advanced Techniques for Dense Model Training
24:36 to 27:11
Learn about specific techniques to deal with sparse models in RL.
“basically you can be very, very careful and write all your GPU kernels so they always add numbers in the same order.”
Real-time Reinforcement Learning with User Data
27:12 to 28:00
Understand how real-time user data informs RL updates.
“And we're able to update that model live.”
Challenges of Reinforcement Learning Horizons
28:00 to 30:59
Explore the complexities of reinforcement learning in model training and user experience.
“Actually, at some point, we'll have to increase that time because as the horizon of the model gets longer and longer, we'll have to re-extend that time.”
Optimizing Long-Horizon Reinforcement Learning
31:00 to 36:25
Understand the strategies for optimizing agents to achieve long-term tasks effectively.
“It's kind of like cherry on top to really get this super delightful experience.”
The Role of RL in Evaluating Model Behavior
36:26 to 42:03
Learn how reinforcement learning can improve model evaluations and real-world applications.
“helped RL fine tuning in generally for many customers.”
Leveraging Production Environments for RL Models
42:03 to 42:42
Learn why using your product's production environment is crucial for building effective RL models.
“The most powerful environment is your own product.”
Challenges of Transitioning from Toy RL to Production
42:42 to 43:32
Discover the difficulties of applying toy RL frameworks to real-world production scenarios.
Components of Effective RL Environments
43:32 to 44:21
Understand the key components of RL environments and their significance for model performance.
“Yeah, like, I mean, what we call aerial environments is really three components.”
Cursor's Innovative Virtual Machine Approach
44:21 to 45:00
Learn about Cursor's approach to building a virtual machine stack for rapid RL environment setup.
“And it has to be super bursty because you can imagine like, we are asking this system, please give me 100 ,000 virtual machines now.”
Transcript
Automatic transcript. May contain errors.0:00Dmytro Dzhulgakov:You need all the infrastructure to run these environments that have to mimic as closely as possible what a user's computer would look like. And it's very important as closely as possible because sometimes the model can actually figure out when it's being run in like a fake environment or a real one. And it has like different behaviors during RL than in production.
0:20Federico Cassano:Are you seeing it being conscious that it's being, it's in a fake environment and starts behaving differently?
0:26Dmytro Dzhulgakov:Yes, yes. Like it's like, oh, I'm in a thick environment. I've learned a few tricks to like get the better reward in this environment. And let me try them out. Models love to cheat. I was really good at encouraging cheating.
0:53Federico Cassano:I'm delighted to welcome Federico from Cursor and Dima from Fireworks, the podcast today. Federico, you are the research lead on Composer2 at Cursor, Cursor's new agentic coding model. And Dima, you spent how many of the last few months moonlighting at Cursor in order to support a lot of the infrastructure required to make this gargantuan training task happen? And so I'm excited to talk to both of you today about how the training of Composer2 came together, what hard problems you solved together, and what do you think it means for the future of AI and foundation model companies. Exciting. Yeah.
1:27Federico Cassano:Awesome. Thank you for having us. Thanks for joining. OK. Let's dive right in. For those who haven't been following as closely, Cursor recently announced Composer 2, which is an agentic coding model meant for long-horizon coding tasks. Federico, up till now, Cursor was mostly enabling other people's coding agents. What was the impetus for Cursor to lean so heavily into Composer 2? and how existential is it for you to become not just an application company, but also a foundation model company yourselves?
1:57Dmytro Dzhulgakov:The reason why we started looking into training our own models is you can sort of think about the model as sort of like a storage drive. It has a certain amount of bits that it can store in its weights. And the idea is very simple. You know, like we care about only one task. We don't even care about coding or programming necessarily. We care about software engineering inside Cursor. and inside the cursor only. And so what if we were to allocate all of the bits of information that can be stored inside the model weights to that one particular task? Also, as people may have noticed, Composer is an order of magnitude less expensive than Opus and other coding models because we can just simply specialize all of the model weights to that particular task.
2:47Dmytro Dzhulgakov:And so we can serve a smaller model or something of that sort.
2:50Federico Cassano:So it's about, let's make sure every single bit of weight or information we have is dedicated toward the specific problem that we have at hand.
2:57Dmytro Dzhulgakov:Exactly.
2:57Federico Cassano:Got it. That seems like it's an almost generalizable problem. Deema, I'm curious your perspective. Do you think that every application company should be looking at Cursor as a harbinger of what's to come? Like, should they all be looking to do the same thing? Yeah, absolutely. I mean, we actually generally see it as a pattern of kind of evolution of applications. You maybe start prototyping, you might be using kind of off-the-shelf model to get something running, maybe do some prompt engineering, figure out how your harness works. But the most kind of leveraged attribute of your application is actual usage of user data or particular specific aspects of how the application works, maybe some aspects of your harness, which tools do you provide, how the application works, kind of really important bits which are important for your application.
3:38Federico Cassano:And the right way to capture that, you can do a little bit of that through prompting, but really the right way to do this is craft your model to act in your environment.
3:45Dmytro Dzhulgakov:Yeah, absolutely. Like there are certain tools the agent calls that it's very hard to succinctly describe exactly the behavior of that tool to the model. And, you know, with just like post-training, we can bake in the optimal way to use those tools. Like Composer, we do serve a prompt to Composer, but I think the way we are training it, it would work even without a prompt and it would know what to do just because like we are intrinsically pushing the model to like the right direction of how it should act throughout our training.
4:15Federico Cassano:Basically, there's kind of like upper bound of like how far you can get this prompt engineering. And if you want to craft really great AI products, you have to go through kind of feinturing and influence the model behavior. That's kind of one reason. I mean, reason number two is what Federico mentioned is kind of cost trade-off or like speed trade-off. Like the way we kind of view it with Fireworks is that when you're trying to do optimization, you had this like three-dimensional trade-off between quality, speed, and cost. And you can go quite far and we are doing it with all the customers. Initially, you can go quite far with just optimizing infrastructure.
4:47Federico Cassano:But when you start getting to model training, you can really push this tradeoff much further. And you can get a better model at a fraction of the cost running much faster. And Composer is a great example. Can I push on this a little bit? I want to ask if this approach is fit or less in PILs. And we were actually all talking about tab 9 on the walk-in. I'm remembering before the LLM era, there were these small, specialized coding models. And one of the things that was, I think, surprising to a lot of people was as you scaled up, you know, you scaled up just training on the Internet and a bunch of English texts and other languages.
5:20Federico Cassano:Actually, the models themselves got inherently better at coding as well. And so at least the trend line I've seen so far is like bigger models perform better on everything, including on coding. Is what you guys are saying, does that go against the grain of the bitter lesson?
5:33Dmytro Dzhulgakov:I think no, but one sort of thing to point out is that the big models trained by the labs train on a lot of code as well. Code is one of the main tasks the labs are interested in pushing, and so they don't just generalize to it. They're a bit specialized as well. I think for our case, actually, if we believe about the bitter lesson, we are just pushing very hard on the data dimension, and we know that the models inherently have finite capacity. And so if we want to saturate all that capacity, we need to scale data. And in order to ingest more data, we need to free up the weights from distractions the model may have.
6:14Federico Cassano:Okay, got it. Super interesting. Okay, let's dig into the training of Composer 2. You launched a couple weeks ago, immediately grabbed attention. Strong benchmark numbers, much lower cost to run imprints on. What's the short version of how Composer 2 works and what you guys did to make it so performant?
6:30Dmytro Dzhulgakov:We started from a very strong base, which is Kimi 2.5. It's like a 1 trillion parameter MOE that's 30B active. So very, very sparse, actually. We sort of like looked at the stack and realized there are like two axes. So mainly Composer 1 was just pushing on one of these axes, which is reinforcement learning. But Composer 2 pushes in two different axes. One is continual pre-training and the other is reinforcement learning. So the thing that made Composer 2 very good is pushing in both of these directions. So we started off the training run by doing lots of mid-training on code tokens, almost sort of pre-training scale, actually.
7:13Dmytro Dzhulgakov:And then coming out of that mid-training run, we took the checkpoints and we did very large-scale RL on lots and lots of tasks. Okay.
7:23Federico Cassano:And then the premise here would be because Cursor sits in the middle of so many interesting coding tokens, you actually pretty uniquely have access to data to be able to train at almost pre-training scale.
7:33Dmytro Dzhulgakov:Yeah.
7:33Federico Cassano:Why not pre-train your own model then?
7:35Dmytro Dzhulgakov:We just think about our approach from top down instead of bottom up. so like how do we get a model that's useful to users in the least time possible if we were to start from the bottom sort of figure out how how we do pre-training and then scale it up to mid training and then okay now we figured out mid training i would do reinforcement learning that would take a very long time to get a model out to our users by doing it the other way around we were able to give a useful model to our users in very little time so hopefully you know like next Composer versions are going to be our own model instead of basing it off on open source base.
8:15Federico Cassano:And what is the model roughly learning in the kind of mid-training step? Yeah. What is the model learning in the post-training step for you?
8:22Dmytro Dzhulgakov:Yeah. So in mid-training, it's sort of just kind of learning about libraries of code and learning about specific code patterns that are very common, like just world knowledge as well. There is like web data there as well. And this is sort of just creating a wider distribution that then reinforcement learning can sharpen on. And so during reinforcement learning, you know, the model gets to play directly with the cursor harness. And so it gets to learn about the world the model is going to live in for the rest of its life, right, in some way. And so then during reinforcement learning, that's where it learns how to call tools properly, how to navigate its environment, how to write correct code.
9:03Dmytro Dzhulgakov:because during mid-training, it learns how to write code. That doesn't necessarily mean it learns how to write correct code. We try to train on code that is largely only correct, but the model doesn't actually know how to differentiate between the two. While in RL, one of the key things that we are doing is we're kind of tuning the feature of the model saying, hey, now you got to write correct code all the time. Exactly.
9:28Federico Cassano:And is the model after mid-training, Is that similar to the model that you guys have on Tab autocomplete or is that a different core competency?
9:36Dmytro Dzhulgakov:Yeah, I mean, it's, yeah, I think I would put it like that because like during mid-training, we are just doing the next token prediction, you know, like how well you predict the next token and then the token after that. So, yeah.
9:47Federico Cassano:So why not just post-train on your Tab autocomplete model then? Why mid-train is a different model?
9:50Dmytro Dzhulgakov:Yeah, I mean, Tab is a very small model because it's like a super low latency model. We want it to be very fast. So like the core two distinctions about the base models here is that tab is like small and Composer is quite large.
10:06Federico Cassano:I see. I see. Okay. So it seems like a lot of the focus of what you guys did for Composer 2 was this large scale reinforcement learning run. Can you break that down for us? Like what goes into that and what are the various hard problems you solved along the way? When you do a run, it's quite different from like pre-training or mid-training because you're not just trying to predict next token. you're actually running the entire harness, like the entire experiment. You're letting the model act in the environment, see how it performs for a given rollout. That's the terminology which is called rollout.
10:37Federico Cassano:And kind of assign it to reward whether it did something correctly or not, which might be some using LLM as a judge or maybe something verifiable like does this code compile or something like this, which actually means that compared with regular training, you need a bunch of other components. Like you still need large-scale training. you still need to orchestrate tens of thousands of GPUs to the forward-backward propagation, all the stuff you do in mid-training and pre-training. But now you also need to orchestrate a bunch of environments. You need to run model inference because when you do this rollout, you're effectively running like real Courser session in some sense, right?
11:11Federico Cassano:So a rollout is like a forward pass? No, rollout is basically your entire agent session from Courser, right? So it basically means it might take something like 50 turns. Model will take your initial prompts, then decides to call some tools you want to execute those tools, then model generates a bunch of other code. Kind of entire session, which you, when you interact with agent in Courser, right, you kind of simulate this entire session as a part of your training run. You get to final reward and you use that signal to now go back to trainer and kind of incorporate it in the model weights. So you have this kind of very big loop, update loop, which is very heterogeneous, right, because you have all this like different components working together.
11:49Federico Cassano:And now you're trying to orchestrate all of this to work efficiently and work with high throughputs because GPUs are expensive and you want to get your model trained quickly in an economic fashion. So that by itself is like very interesting kind of problem and intersection of algorithms and infrastructure because there are a lot of trade-offs how you can kind of co-optimize and co-design the system. One aspect is kind of people call about like this async RL of a pipeline RL. The idea is basically, okay, you're trying to update this model in steps, right? So you have your current model version and you're trying to do a bunch of rollouts with it.
12:24Federico Cassano:What does your trainer do while you're doing these rollouts, right? Like native approach would say that, okay, now I'm going to stop my trainer. I'm going to do a bunch of sessions and those sessions might run for like five, 10 minutes or even longer if it's like longer horizon tasks, I'm going to get those outcomes. And now I'm going to pause my inference. I'm going to go back to training, trying to do updates. That's like very theoretically algorithmically robust because you are not precisely simulating everything, but it's very system inefficient because half of your capacity is sitting idle at the time.
12:52Federico Cassano:So you can do all the clever algorithmic tricks allowing you to... What do you do instead? Yeah, you can pipeline all of this. So imagine this as a gigantic factory, right? You have this trainer building and you have a rollouts building. They're always churning, right? So rollouts always take latest model version and try to do new sessions and kind of simulate new agent sessions. and trainer always takes new outcomes as they come and try to compute updates. So everything is moving along all the time. The trade-off is that why I'm saying that algorithmically it's different because now by the time you finish some test rollout in your kind of simulated environment, maybe model weights already updated on some other data.
13:30Federico Cassano:So you have this kind of staleness, like delay between how quickly model can learn updates because by the time you kind of process or some interaction session with a simulated environment, your model base changed, and that introduces interesting training dynamics, and there are clever ways how you can address this. But the flip side of that is that all your GPUs, all your computers kind of loaded and chime in all the time, which actually you're using more flops, and to your bitter lesson example, yeah, you have higher compute efficiency, you can get to a better model in a smaller amount of time.
14:05Federico Cassano:Maybe you're losing a few percent from being asynchronous us and not doing like perfect mathematical updates, but you way compensate for that by effectively not leaving half your capacity on the table. And there are a lot of kind of depth and interest in interaction in that part.
14:20Dmytro Dzhulgakov:And we're very serious about performance at Cursor because unlike the big labs, you know, we have tens of thousands of GPUs, not millions. And so, yeah, we do all sorts of tricks to get the most out of GPU. Like we train in production with FP4 even. We work with fireworks to like push on inference as well. because the thing about a rel infrastructure is just like it's just inherently more complex than pre-training because you need all the pre-training infrastructure that's just like one of the requirements then you need all the infrastructure to run these environments that have to mimic as closely as possible what a user's computer would look like and it's very important as closely as possible because sometimes the model can actually figure out when it's being run in like a fake environment and a real one.
15:07Dmytro Dzhulgakov:And it has like different behaviors during RL than in production.
Read the full transcript
15:11Federico Cassano:Are you seeing it being conscious that it's being, it's in a fake environment, it starts behaving differently?
15:16Dmytro Dzhulgakov:Yes. Interesting. Like it's like, oh, I'm in a fake environment. I've learned a few tricks to like get a better reward in this environment. And let me try them out.
15:24Federico Cassano:Models love to cheat. RL is really good at encouraging cheating. Yeah.
15:28Dmytro Dzhulgakov:And then we need a really efficient inference. So this is really important so there is like actually this kind of myth that during rl you spend more way more inference flops than training flops this is sort of like just because the open source inference engines are very unoptimized instead of actually being a property of rl roughly the same ratio is kind of the same in theory if you push the gpus to the maximum you should have one third of your training gpus allocated to inference right because training is effectively three forward passes you have the forward pass, you have the data gradient, the weight gradient.
16:05Dmytro Dzhulgakov:While if you really hit the critical batch size on inference, you should only have a single forward password of flops.
16:12Federico Cassano:So that's why you guys use Fireworks instead of using an open inference engine?
16:15Dmytro Dzhulgakov:Yeah. I mean, the other alternative is we would build one in-house, but, you know, if we have finite engineers like everybody else, we would prefer to have engineers make training more efficient and more precise rather than like spin up like an inference effort. Yeah.
16:30Federico Cassano:Okay, that's super hardcore. What about, I think you mentioned in your technical paper that you were doing this in a kind of globally distributed way. Why globally distributed and then what makes that hard?
16:40Dmytro Dzhulgakov:Yeah, well, there are various reasons. One, you know, like this very large contiguous clusters are hard to find in the market. And so what we can do instead is we have one cluster that's going to run all of training. You know, we can't do global training cluster. But then the inference component of reinforcement learning, we can globally distribute that across small clusters all over the world. So I think for the Composer to run, we use the four clusters in total. that were all over the world, very far away from each other. And we even used some of our production traffic when it was least used.
17:20Dmytro Dzhulgakov:So like we had the Composer 1.5, the previous model served. And when it was least used by people, we just grabbed some inference GPUs and we put them to speed up training. And so we can do these sort of things and sort of easily scale up our training ground without having one large contiguous cluster. And the thing that enables it, maybe Dima can talk more about.
17:41Federico Cassano:Kind of like to reformulate what Federico said is basically RL training is like very heterogeneous, right? And by leveraging heterogeneity, how different components, like what infrastructures they need, you can actually drive efficiency. You see this pattern kind of across the board everywhere. Specifically for training, you have all this like highly interconnected clusters. You need high speed network, kind of need to work in lockstep. So those clusters are expensive, right? And actually, it's really hard to find big ones, right? Basically, at the scale with which Composer was trained, finding like 2x larger cluster is like significantly harder than finding the current size one.
18:16Federico Cassano:And that's why if you can disaggregate these components and put them on different places, one, you don't need to find such a big cluster. Two, you can actually find like different trade-offs of hardware because for inference, you don't need that kind of wide interconnect. You can have smaller groups of GPUs interconnected together. You can have heterogeneous types of GPUs. You can have different generations of GPUs. you can kind of play all these games of optimization. And finally, like inference, it's much easier to scale up and down as you go. And yeah, it's very conventional, like when you have off-peak hours, you can view all your kind of inference pool as one set of GPUs, serving production traffic for real users or serving simulated environments for RL purposes and kind of balance between this.
18:59Federico Cassano:Of course, it's a very interesting systems problem. So you can mention like the Kimi model is like one terabyte. training step takes somewhere between like 5 to 15 minutes. So it basically means like every 5 to 10 minutes you are producing like 1 terabyte new snapshot of weights. So the question is like how are you going to ship it to a different cluster on the other side of the world very efficiently, right? And you want to like do it quickly because remember you don't want to get this staleness to get out of hand. So I think that was probably one. Yeah. The kind of the most fun part which we figured out together is that despite full model being like 1 terabyte, not all the weights change every step, right?
19:36Federico Cassano:Because RL does a lot of very precise adjustments, especially as the training going along. So actually there are very kind of regular patterns in like which subset of weights gets changed. Maybe not all of them change every time. So if you were to look at like how my model changes within one training step, like after 10 minutes, there is relatively small delta between those. You can write a compression algorithm, which basically leverages this property. And now you end up with kind of like database systems problem, which is, okay, I have my delta and I just want to ship it across the world. My delta maybe is like 20 times smaller than what shipping the full model is.
20:12Federico Cassano:And that makes it practical. But of course, now you need to build all this kind of machinery from storage systems of full snapshots and deltas and recovery and reconciliation, etc. We were able to build it kind of in lossless fashion, basically means that you always end up with bit equivalent model on the other side. So you don't need to worry about any mass aspects of this. And you can do it really fast. You can do it under a few minutes. Even in the worst conditions, usually it's under a minute. And most importantly, you pause only for maybe 30 seconds to swap the weights in your actual inference.
20:43Dmytro Dzhulgakov:We also fully saturated the egress of the cluster by sharding the upload and the download as well.
20:50Federico Cassano:So you can do all these system tricks to bring the stand down. It is quite a few complexity, but you can kind of abstract it out and just make it work great. It doesn't interfere with your training algorithm. And on the flip side, you have this kind of power to disaggregate, to leverage other clusters to do all that. And that kind of goes against conventional wisdom of how you should do RL infrastructure, because conventional wisdom is like, okay, you're going to have this really huge one cluster connected with RDMA, and it's going to be very expensive. And you're going to probably spend, maybe you're going to allocate like one third to training and two thirds to inference.
21:26Federico Cassano:And sure, if you have very expensive network, it's much easier to copy this one terabyte quickly. but now we have like three times larger cluster. Now, if your inference engine is more optimized, then maybe you're going to save one third of that cluster in terms of GPUs anyway, because you're just more efficient and you can take half of this cluster somewhere else and maybe cheaper hardware in a different region. So your cost comes down quite a bit. I love that you guys are just grinning as you described this, because it's like, it's so hard and this is like a systems engineer's dream, right? And so it's just like a, it's an amazing, amazing system you guys have built.
21:57Dmytro Dzhulgakov:We spend a bunch of nights working on this.
21:59Federico Cassano:You look like you spent a lot of time together. What about, I mean, you went to the beginning that Kimi is a very large, sparse MOE model. Does that make the RL run tricky in any way?
22:11Dmytro Dzhulgakov:Yeah. How so? Well, when you do inference, you're essentially doing a forward pass. It's just kind of like autoregressive. And in this forward pass, it produces log probabilities of the tokens it has sampled. when we ship back the like generations of the model to the trainer we have to rerun that forward pass because as we mentioned we are doing asynchronous training so the model that has produced the pass may have been like actually a few steps behind what the trainer is at and so we have to rerun that forward pass and reproduce log probabilities now the problem is in theory this log probability should be exactly the same if it's the same model version.
22:53Dmytro Dzhulgakov:But even with the same model version, you get slightly or sometimes very different log probability values for the same tokens. So this is often called like a numerical mismatch for inference. You hear this about all the time these days.
23:10Federico Cassano:Why is that? Why does that happen? I mean, primarily because fundamentally floating point arithmetic, which is doing this, is non-deterministic. Sorry, floating point arithmetic is non-deterministic? So, you know, we learned this code that, like, if you take A plus B plus C, right, and, like, C plus B plus A, it's going to be the same result. If you're doing this with integers, with whole numbers on the computer, that's going to be always true. If you're going to do it with floating point numbers, which are actually, like, approximation numbers, you have this, like, Montice and exponents, et cetera, A plus B plus C and C plus B plus A is going to give you, like, different results, or even, like, A plus B and B.
23:43Federico Cassano:So basically, fundamentally, it's accumulation order of all these operations which models do. It's basically multiplications and additions. And addition order matters to your final result. It's all small differences, but they get amplified through millions and billions of operations. So when you do inference of models, usually it doesn't matter that much. Because you pre-train your model. It's actually pretty robust. If you flip some bits, it's still going to produce good results. Your benchmarks are going to change. But RL in particular, because you're using this very, very weak signal to teach the model, the noise from these numerical differences can make or break your training.
24:21Federico Cassano:And that's particularly important. And again, it's an interesting intersection between algorithmic and systems part, because you can write a beautiful mess and it just doesn't work in practice. There are ways how you can drive this difference to pretty much zero. There are all these batch invariant ways. basically you can be very, very careful and write all your GPU kernels so they always add numbers in the same order. So you always do like A plus B plus C and not a different order. It's possible, but it always has like trade-offs, right? Basically your system becomes maybe like 2x or 3x slower.
24:52Federico Cassano:Again, it becomes an interesting trade-off like, okay, what is the 10 % of slowdown which we can take? Or in fact, it's actually a few percent of slowdown we can take to address 90 % of this difference. That's the right trade-off, which we find together through iteration. And you mentioned that particularly for MOEs and sparsity is hard. The reason for that is that the way MOEs work is that you would take your activations at every layer and you would run it through a gating layer, which basically decides, okay, for this token, I'm going to run out of 384 experts, I'm going to run this eight. So it's going to do some mess in top eight scores.
25:28Federico Cassano:Those eight experts are going to be activated, other ones will not be activated for this token. This operation amplifies your small numerical differences quite a bit because maybe your hidden states were like difference by like fifth digit after dot, doesn't really matter, but this difference made it so you picked expert number seven versus expert number nine as kind of as a cutoff and suddenly you went and like activated totally different part of the model and your difference got amplified quite a bit. And my models, by definition, are more sensitive to this mismatch. Again, when you do inference or when you do kind of regular layout, it usually doesn't matter as an average out.
26:05Federico Cassano:But now if you're trying to base this model learned, this difference is huge because your inference activated expert number seven. Now in your training, you're trying to update expert number nine, which didn't even contribute to that during inference. So were you guys handwriting GPU kernels then to help get around this problem? Yes. So you can, again, you can address a lot of this through GPU Cornels and there's always trade-off. Specifically for ME, you can do this interesting trick, which people call router replay. But basically you can have your inference just pass extra information to training and say that, hey, I activated expert 7 for this token.
26:37Federico Cassano:This very small piece of information is just one integer saying that like, okay, this is the expert which activated. So trainer can be aligned with that. And a lot of this numerical alignment is basically doing tricks like that, matching quantization levels, matching kernels, et cetera. to drive the divergence between training inference implementation down. And that makes huge difference in between, you know, your run may be divergent completely or being, you know, multiplex less compute efficient because you'll need much more data to address to this mismatch. I'd love to maybe chat a little bit more about the RL kind of recipe.
27:10Federico Cassano:Can you say a word about the reward signal you're using? It's like, are you care? Okay. Can't say. Got it. Top secret stuff. Top secret stuff. Okay. That makes sense. like it seems like there's a almost like the equivalent of learning in sim this is simulated rollouts versus like you have so much actual user data that you could be learning on why not just do rl on your your actual user data and your actual user harness versus doing this in sim yeah we're
27:33Dmytro Dzhulgakov:also doing that so that's uh what we call a real-time rl okay and uh we use the same technology to do like the inference wait sync with like fireworks to do this we find like user signals where the user was happy or sad about a particular model generation. And we're able to update that model live. And so then ship a new version of the model continuously every few hours. We're working on decreasing that time. Actually, at some point, we'll have to increase that time because as the horizon of the model gets longer and longer, we'll have to re-extend that time. It's like an interesting play. Like right now, we are trying to decrease the time for stability because we were figuring out the right hyperparameters.
28:17Dmytro Dzhulgakov:And then after we figured it out, we have to re-extend it again just because we want to lengthen the horizon of these models.
28:24Federico Cassano:Do you need to do any of the kind of like pre-training simulated RL? You have so much actual user data. I imagine that's just like much more valuable to train and tune on. Like why not just go straight to the online RL step? Why do you have to do the offline RL?
28:37Dmytro Dzhulgakov:The online RL currently is pretty inefficient. We suffer from this problem that the GPUs are offline for a long time, essentially.
28:45Federico Cassano:And besides that, there are also like different trade-offs, both in terms of efficiency and user experience. If you do simulation, you actually do multiple rollouts from the same prompt, right? You effectively take a task and you ask a model to do 16 tries to the task, like 128 tries to the task, like different rollouts from the same prompt. Some of them are going to go well, some of them are not going to go well. and by doing multiple rollouts in parallel, you're able to get much more precise signal. Maybe model is very good and it does it well 90 % of the time. Maybe it's not very good. Losses like GRPO, like group policy grading, kind of work by doing multiple rollouts at the same time.
29:22Federico Cassano:If you're doing online, you have only one rollout coming back. So trade-offs of how you do it algorithmically are different. And most importantly, if a simulated rollout goes wrong, it's not too bad, right? I mean, you just maybe spend some time on GPU. If it's an actual user, you have much higher minimum bar on this because effectively you're doing A-B tests, right? So if the model produces something weird, that's a bad user experience. Okay, so you can go off policy more often when it's not a real user because you can experiment with crazy things and without affecting the user experience. You can do a lot more rollouts.
29:57Federico Cassano:You can do GRPO. And then you can basically bootstrap some level of performance that's good enough to even put in front of users.
30:05Dmytro Dzhulgakov:Yeah, like we teach reasoning through like the offline RL, which is actually like called online RL. Offline RL is more like DPO kind of technique. The sort of reinforced kind of RL is online. And then there we like teach the reasoning to the model. We give it some kind of input of the behavior it should have. We try to give it new information about the world and we teach it tool calling. And then we put it live to users. Because you could imagine like if the model is bad, users don't want to use it. They're not going to give us any feedback, right? So the model has to meet some kind of bar to even be put into OnlineREL.
30:41Dmytro Dzhulgakov:We want to be really happy with the model, and this is the model we ship. That's kind of the paradox of OnlineREL, or how we like to call it real-time, is that we can't use this to really create the model from scratch because users need to be using the model. And so it has to be good already, and we can only make it better. Yeah. It's kind of like cherry on top to really get this super delightful experience. Yeah, totally. Hopefully one day it will be like big, big cherry, you know?
31:12Federico Cassano:Yeah, Dan Roberts presented at our conference last year. I think you were there. It's like traditionally it was the big cake and the little cherry. Yeah, the Yannickons cherry, yeah. Little cake, big cherry. Yep. I'm curious, the Andre Carpathy line of like right now RL is, you know, still super inefficient. You do a big, big, long rollout, and then you kind of get like, you know, a little bit of information at the end. And it's still like, I think, slurping bits from a straw. What do you think? And have you been able to figure out how to get more bits out of that path?
31:40Dmytro Dzhulgakov:I can't talk about that.
31:42Federico Cassano:Okay. Okay. Got it. We're back on the secret stuff. Good. That's how I know I'm asking the right questions. You mentioned the rollouts are a few minutes at a time. It seems like the whole field is pushing towards making like long horizon agents, agents that can work for a long period of time uninterrupted. and generally not failing. I love that meter scaling chart. What goes into the RL process to try to get the agent to run for longer?
32:07Dmytro Dzhulgakov:Several things. So one problem about sort of reinforcement learning is that the longer the trajectory is, the harder it is to do credit assignment. So you can imagine we are giving thumbs up, thumbs down at the bundle right at the end of its work and sort of like to simplify the problem is like the model asks itself okay where did I do it right and where did I do wrong that's basically the problem called credit assignment it gets harder as this gets longer so you have to do a bunch of tricks there the other problem is just like you run out of space right like these models have a finite context window and at some point they're gonna reach that so actually the way we solve this at cursor is we put compaction inside the RAL loop.
32:53Dmytro Dzhulgakov:So we call this self-summarization. So during reinforcement learning, the agent actually learns how to continue and go on forever. So in practice, our model is like a 200 ,000 context window model. But in reality, it can go on for millions of tokens. And just because of this ability that it can summarize its work and then take that summary to restart its context window while still trying to accomplish the task. And through RL, because RL pushes the model to do things correctly towards the goal, at the same time, jointly, we are training the model to produce a good summary, and then we are training the model to listen to that summary very well at the same time.
33:36Dmytro Dzhulgakov:And so this is kind of like a continuation to reasoning, almost, I feel like.
33:40Federico Cassano:I find it fascinating because usually context management is considered part of the hardness, right? In this case, you're effectively co-optimizing how part of the harness and model itself work together and throwing all of that in the optimization loop. And we've seen this again and again in AI. The more you throw computers a problem, the more you can solve the problem end to end. The magic of computing better lesson works and you get a much better system which can work together. Totally. Totally. Do you think every company is going to be RLing their own harnesses? Do you think that every company has the same shape of problem as Cursor?
34:15Dmytro Dzhulgakov:If they are using AI and they're like producing lots of tokens and they have a product to optimize against, I think it's like the right move and the right direction to train models.
34:27Federico Cassano:Yeah. Yeah. Interesting. Interesting. And so it seems like most of the reinforcement learning you guys did then was on the kind of like the harness slash tool use part rather than on the get good at, you know, completing the next token for code. Is that roughly the pattern that other founders should have in mind when they're trying to think about where should I use reinforcement learning? So if you're trying to get an agent to perform tasks with tools over long horizon, you need RL. If you're trying to create a model that's good at summarization or a next token or whatever, you probably don't need RL.
34:56Federico Cassano:Is that a good framework for when you need RL?
34:58Dmytro Dzhulgakov:I think RL fits everywhere. So even for tab, we use the RL. Personally, this is just my theory and it's not backed up by anything. When you pre-train a model, the models are just ingesting the totality of human knowledge. Let's say you're training a model for math. The model sort of like learns all the math on Stack Exchange. The model, when it's presented with a math problem, and this is a model that hasn't gone through RL, the model needs to wonder what kind of person it is. Is it the expert or is it the student that's trying to learn? and so one of the things that i think happens during rel is that we are tuning this knob letting the model know hey you are the expert you need to do things correctly so that's like one thing that happens is we are sharpening this distribution sort of like rel has a few phases so like there is the very first phase where the model learns and becomes very good very quickly and then there is like a second phase where like it takes a lot of compute to continuously improve the model and like you see the model starts reasoning and have this pattern so in the very first phase of the curve i think that's where we're just tuning the knob telling the model hey you should do things correctly here and so rl in the small compute case is also very useful just to let the model know that it has to do things correctly that's sort of like my case to this
36:24Federico Cassano:yeah i mean second that i mean you we see this pattern because many of these cases you know we helped RL fine tuning in generally for many customers. And we see this usually you kind of continuous pre-training, basically me training, like regular supervised fine tuning is simplifying. You can say it's transfer of new knowledge, kind of in abstract way. And RL is kind of sharpening the behavior or like particular qualities you would want from the model. And usually you end up needing both. And even to your example of summarization, it's actually like RL may be useful for this because sometimes it's, if you want particular style out of summarization, right, It's really hard to come up with examples of good and bad summarization, et cetera, really describing this precisely.
37:05Federico Cassano:But if you use, for example, LLM as a judge, you can actually say very precise rubrics. You can prompt eval saying, okay, this is a criteria how I'm going to evaluate whether summarization is good or not, throw it into RL loop, and let the model experiment with different summarization styles, figure out what you actually want from it, while maybe another LLM kind of evaluated whether it's a match in particular rubric or not. And that's kind of type of pattern which you see a lot, not just encoding. I see. Okay, I'm going to ask this question to Dima because Federico is going to plead the fifth.
37:37Federico Cassano:You mentioned LLM as judge a couple of times. Do you think that ultimately companies will be more successful having experts hand-examining RL rollouts and hand-coaching the model behavior in some way? Or do you think LLM as judge, other automated rubrics are likely to get us there? You don't really like put experts directly in judging RL rollouts. I mean, that would be some kind of like, I mean, real-time RL if it's actually users or like some form of, I don't know, like RLHF or DPO. I mean, generally, the more verifiable your reward is, the better because it allows you to like scale the compute and just get better outcome.
38:11Federico Cassano:In some case, and by verifiable, it basically means like, can you automatically produce it without the human? Of course, if it's like math or coding and you can craft something like very deterministic, that's the best. The reason why LM the judge works is that it's actually kind of like generator discriminator distinction. Like it's much easier to judge. I mean, the center of humans, right? It's easier to judge than to create. There's Lama VC. Yeah. No. No implication there. But yeah, it's much easier to judge and you can craft precisely like different criterias you want to rank some answer. And you see this pattern where you might have like very complicated eval from multiple aspects, right?
38:50Federico Cassano:Because if you dump multiple aspects to a single LLM, it might get confused how to judge. You might break it down. Okay, you're going to judge rubric based on style, based on some different aspects, based on factuality. Kind of really craft these rewards. Some of them will be the genesis, some of them will be LLM-based. And that's what guides your model behavior. Then you just turn on more computes and see the graph go up. Do you think that we're going to see RL be more effective in the harder to verify domains? Do you think LLM is judged as sufficient? That's one of the picnics you would start, right?
39:23Federico Cassano:Ideally, you want to figure out what is the actual outcome, what is the actual metric you want to get, right? So kind of trying to approximate this 3LM is one way. Trying to get bigger simulated environments is another, right? Like if you can simulate more of your product, if you can simulate more of your environment, usually you have like final metric which you care about. It's just harder to capture. If you can figure out how to capture this, that's great. And to your point about, you know, experts, I mean, experts are still muted, right? Because crafting this task and actually encoding the product experience you want, that's what matters, right?
39:54Federico Cassano:We went through software 1.0, 2.0, 3.0, right? Instead of crafting software directly, we went to crafting training data. Right now you're effectively crafting the evaluation rules. But that's still very important. You need to look at examples. You need to look at the data. You need to look at where your product fails and how to nudge the model in the right behavior. I want to ask about RL environments, which is maybe related to what you were talking about. Now, it seems like there's been a huge explosion in just the revenue scale that some of these RL environments companies are reaching. What do they provide that's actually useful?
40:27Federico Cassano:Because I think Cursor, for example, you have so much data on how your customers are actually using your environments. What do the RL environment vendors offer you on top of what you already have?
40:38Dmytro Dzhulgakov:Yeah, we don't actually use any of the environment vendors. I think it's very difficult to construct working environments. It's a valuable product for people that do not have access to this. However, for coding particularly, there is a very large amount of working coding environments available to everybody. That's GitHub, right? You can go in and maybe you can have a model, just install all of the dependencies for a repository, and that's a working environment. I think a lot of the difficulty comes from the infrastructure as well. So you can imagine that an environment that works well for a particular task may need services up.
41:20Dmytro Dzhulgakov:You're making a change that, let's say, a database migration. To test that it's actually working, you need the database up, right? And so those kind of things are very tricky. I think these environment companies are quite helpful for that kind of stuff.
41:35Federico Cassano:There are kind of two aspects to this, right? First, if you look at Frontier Labs, they're trying to build a generic model which is good at everything, right? So they need to cover all these different tasks underneath, package up in one model, and kind of encourage it to generalize. So that's kind of one part, and that's very helpful. In cases like Composer, you have your actual product. And I think that's what it also kind of video with Fireworks. Yeah, if you have your actual product, you should do URL against it. The most powerful environment is your own product. Exactly, because that's where your model actually will be used.
42:07Federico Cassano:And of course, if you have Frontier Lab, you're not going to do it across all the products. But if you're trying to build the best model for your product, specialize and tailor it, you should just use your production environment. Of course, you want to isolate it properly. You don't want to model havoc on your production database. You want to clone it, et cetera. And there are some tools from environment companies, just like from general infrastructure, which makes it easier. But generally, you want your RL environment to be as close to real production as possible. and that's what you know as an example we see it is if you look at kind of toy rl examples so toy rl frameworks they always start like oh there's this like toy environment i'm going to spin up a docker container and run everything in it which is great for like toy examples if you're trying to teach model how to play atari or whatever right but if you're actually transitioning to like production cases you can't just put your real real production application in a docker container and we found it pretty early ourselves like working with many folks like in case of Courser trainer on their side some other customers we run trainer on our training platform but for environments we actually default to running them on the customer side because that's where the actual implementation is and you effectively have the same setup of trainer even if it's part of Airworks platform or on the customer side calling the actual production environment not trying to kind of wrap it and componentize it on the hosted platform because that's really hard and that introduces differences.
43:32Dmytro Dzhulgakov:Yeah, like, I mean, what we call aerial environments is really three components. One is the harness. So the harness is like where the model can submit tools and the tools get executed. And the second thing is, let's call it the kind of operating system, right? So like, what is the actual like world and state where the model is like interacting with? And then there is like the reward component, which needs to check at the end that the work is done correctly. And generally the harness is pretty portable. You can take the harness and put it in many different environments. The thing that's key is the operating system.
44:09Dmytro Dzhulgakov:And to replicate this, just normal containers don't really work very well. So at Cursor, we actually built like a whole virtual machine stack. And so we can spin up like virtual machines really quickly. And it has to be super bursty because you can imagine like, we are asking this system, please give me 100 ,000 virtual machines now. and it has to all come up. And yeah.
44:33Federico Cassano:Awesome. I really enjoyed this conversation today. I think Cursor is such an inspiration in what you all are doing as a company towards going from application company to really a frontier model lab. And I think the work you did with Composer too really leads that charge. So really special to hear about it. And then Dima, really cool to hear about the hardcore infrastructure problems actually that the two of you solved together in the trenches over many, many late nights to make it all possible. So thank you. Thank you guys for joining today.
45:00Dmytro Dzhulgakov:Thank you so much for having us. Thank you.
45:28Woo!
From the publisher
Cursor's Federico Cassano and Fireworks' Dmytro Dzhulgakov explain how they collaborated to build Composer as a specialized foundation model. The core insight: models have finite capacity in their weights, and allocating all those bits to the singular task of software engineering in Cursor frees the model to be both better at the task and far more efficient at inference. Rather than start from pre-training and work up, they took an unconventional top-down approach — mid-training and RL on top of an open-source base to get a useful model into users' hands fast, then specializing the model around real Cursor usage. With Fireworks providing distributed infrastructure, Composer delivers frontier-class coding performance with the speed of a much smaller model.
Hosted by Sonya Huang, Sequoia Capital




