In short
Core Automation’s founders argue that replacing Transformers requires (1) architectures that learn at test time/continually adapt, and (2) end-to-end integration of pre-training + reinforcement learning, plus major efficiency work (especially GPU kernels) to make new architectures practical.
Guests
Jerry Tworek (cofounder of Core Automation; previously VP at OpenAI, led Strawberry and Reasoning teams; long-time RL scaling researcher). Rohan Anil (cofounder; previously one of four Gemini pre-training leads; earlier Google Brain optimization/efficiency work; “fix-it guy” at Google and later at Anthropic).
Key claims
Transformers are economically powerful but capped as static models; they don’t adapt well when real-world distributions change. RL is not the only “learning from experience” method, but current training/evals don’t match messy deployment. Inference is inefficient (token-by-token), and “computational depth” is poor; better depth/learning could come from new mechanisms beyond autoregressive generation. Labs struggle to unify pre-training and RL due to different optimization regimes and release-cycle incentives.
Notable examples
football as RL-like adjustment; math as different internal learning; Codex usage as a prompt for why tasks aren’t fully automated; QR-kernel competition on B200 nodes where human+search can yield ~60x speedup vs models failing to solve the kernel efficiently.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding Learning Mechanisms
0:00 to 1:00
Explore different learning mechanisms through examples like football and mathematics.
“If I play football, for example, it looks very, very closely to reinforcement learning.”
The Future Beyond Transformers
1:45 to 3:40
Discussing the limitations of Transformers and the need for new architectures.
“The first step to replacing Transformers is appreciating deeply how far they were able to carry us.”
Challenges in Reinforcement Learning
3:40 to 7:10
Examining the challenges faced in scaling reinforcement learning for AGI.
“automation and of the systems of the products that we have today that we essentially have built over those six years of scaling.”
Real-world Application Gaps
7:10 to 9:00
Identifying the gap between model training and real-world application.
“Our training data didn't really replicate the real world use cases and despite us basically maximizing all the tasks, if you see ask anyone training models, hey, what is one of your main issues?”
Learning at Test Time
9:00 to 11:00
The need for models that can learn and adapt during real-world usage.
“If those were easy to solve, some would already solve it.”
Economic Factors in AI Development
11:00 to 12:50
Discussing the economic implications of scaling AI architectures.
“It's not guaranteed by itself, but for LSTMs, it probably wouldn't be that way, which made it happen.”
Diverging from the Transformer Path
12:50 to 14:00
Reasons for starting a new company to explore alternatives to Transformers.
“And I think in many ways, it's likely a timing thing, timing issue.”
The Evolution of Transformers
14:00 to 17:43
Explore the history and current state of transformer architectures in AI.
“what the most successful labs are doing.”
Challenges in Computational Depth
17:43 to 19:24
Discuss the limitations of existing transformer models and their computational depth.
“It is something when it comes into practice.”
The Future of Transformers
19:24 to 22:27
Examine the potential future applications and limitations of transformer models.
“Depth is like, it's called deep learning because you wanted deeper representations.”
Show all 20 chapters
Learning from Experience
22:27 to 26:41
Analyze different learning methods and their effectiveness in AI development.
“There's some ability to adapt, but it's not very big and not very flexible.”
Optimizing Learning Algorithms
26:41 to 28:00
Understand the importance of measurement and optimization in AI training processes.
“Rohan, I'm curious, since a lot of your work has been around optimization and efficiency, how do we get to orders of magnitude more efficient, I guess more compute efficient and more data efficient learning algorithm?”
The Evolution of Optimization Techniques
28:00 to 29:50
Learn about the journey of optimization algorithms from Google to neural networks.
“and that's where one order of magnitude improvement would come from, and that's like a training procedure.”
The Interaction of Architecture and Optimization
29:50 to 31:30
Discover how architecture and optimization methods influence performance in neural networks.
“And in some sense, it's like two sides of this coin and optimization on architecture go together.”
Hardware and Biological Learning Efficiency
31:30 to 32:40
Explore whether machines can ever match or exceed biological learning efficiency.
“And I also see Aral as spending a lot of compute, not as efficiently.”
Automating the Research Process
32:40 to 35:00
Understand the vision for creating an automated research lab and its implications.
“maybe more analog, figure out how to deal with analog circuits and figure out how to do with error correction, figure out how to get information through, it will be much harder.”
Defining AGI and Its Challenges
35:00 to 39:20
Delve into the definition of AGI and the challenges of removing humans from the loop.
“And to start with, I think the automation, the version of automation by core automation is about giving each human maximum level of agency in some way.”
Core Automation's Roadmap and Technical Challenges
39:20 to 42:00
Get insights into the roadmap for Core Automation and the technical hurdles they face.
“on this company or some other to try to unlock how do we make our models learn and adapt at a test time on a deeper level than we've been so far.”
Exploring Kernel Efficiency and Architectural Innovations
42:00 to 46:04
Discover the challenges and innovations in kernel optimizations and transformer architecture.
“because that's our inner loop to having more efficient architectures.”
Defining Success in Automated Learning
46:04 to 48:06
Learn how to identify when you've found a successful architecture for automated learning.
“architectural ideas is exploring that space.”
Transcript
Automatic transcript. May contain errors.0:00Jerry Tworek:If I play football, for example, it looks very, very closely to reinforcement learning. I kick a ball a lot of times and every time I adjust it a little bit and I see if it roughly matches what I wanted and there's some self-reinforcement happening. When I learn mathematics, it's a very different type of thing. It's like reading about hard concepts and thinking about them very deeply inside my head until things click and until I have them connected. And both of those in some way are learning from experience. they are just very different. We probably are spending the most compute than ever on learning from experience but reinforcement learning is not the end of learning from experience and there will be better approaches that researchers will be coming up in the coming years on how to use that data.
1:00Rohan Anil:Jerry, Rohan, thank you so much for joining us today. The two of you are the founders of Core Automation, one of the hottest neolabs in San Francisco right now. And before starting Core Automation, you led some of the most important research projects of the AI era. Jerry, you were VP of OpenAI, where you worked, amongst other things, on running the Strawberry and Reasoning teams. And Rohan, you were one of two of the four pre-training leads at Gemini. and before that led a lot of the fundamental AI research at Google Brain and were the fix-it guy across Google and then at Anthropic. And so between the two of you, you've seen more than your fair share of what the world looks like in terms of doing frontier research.
1:42Rohan Anil:And so I'm very, very excited to dig in. Let's start with you, Jerry. You tweeted a very spicy take recently. The first step to replacing Transformers is appreciating deeply how far they were able to carry us. Is that a eulogy for the Transformer? What does that mean?
1:58Jerry Tworek:Thank you very much for inviting us here, Sonia. I feel like a lot of my interviews these days is explaining my tweets and what did I mean. But appreciating Transformer means understanding what it does well. So you're not solving the problems that it is solving well. You have to focus on its weaknesses. You have to understand good parts and bad parts. And it's very easy in a lot of the work, what people are doing in architectures is trying to make transformers cheaper and trying to make transformer more efficient. I very rarely see people thinking about how do we make transformers more powerful, trying to do more expressive.
2:36Jerry Tworek:But seeing someone weak parts and seeing someone strong parts are almost the same thing. It's just understanding the shape of transformer a little bit more. Well, I think right now we are in this stage. We got really, really good at training really, really big models. We mastered two algorithms. We mastered pre-training at a large scale, and we mastered reinforcement learning at a large scale. And I'm asking myself a lot, what is next in machine learning? And I think at this moment, what the bottleneck is to better models and to smarter systems is the architecture itself. It is this moment to revisit the train we've been riding for the last six years of trying to add more and more parameters to essentially two of the same operations, which is MOE and attention.
3:25Jerry Tworek:And when I'm thinking about it, like where we are today and what we are doing, I am thinking a lot about what codecs and what cloud code are doing for us. And I'm really, really appreciative of those systems and of the coding and of the workflow automation and of the systems of the products that we have today that we essentially have built over those six years of scaling. And I think this is the first step of thinking, what is the, if we want to work on replacement, we need to see where we are, what problems we have solved, to start seeing what the next stage is, what problems we haven't solved yet.
4:02Jerry Tworek:What kind of are we missing? And this is kind of whenever I use codecs and I am successful at a task, I also start thinking why didn't I try to push that thing harder whenever whenever I come to work there are a lot of things I do with codex but I still come to work I still ask it to do certain things for me and I'm always asking myself why am I even needed there why is core automation is name and its concept is we want to we want to be automating tasks and why why are those things not yet automated why why is not codex doing everything for me and this is the question of like the research where we want to go and with that research i'm trying to think what kind of models what kind of what kind of systems do we need what kind of qualities do we need what that that's we don't we don't have today and that's that's what i'm thinking a lot these days and you have this starting premise of the architecture is the issue which i think is a contrarian point of view so what what led you to that point of view what did you see that made you think the architecture was the issue it's fundamentally what is the issue it's It comes back from the previous implication.
5:11Jerry Tworek:What I think is the issue is that the models are being trained in the lab and are being deployed in the real world. That is the fundamental tension that is there. And a bit of my disappointment comes from my personal story. Whenever we were starting the research and progress on scaling up reinforcement learning at OpenAI, I basically believe that scaling up reinforcement learning is a necessary stepping stone on a path to AGI since I started working at OpenAI. And I was always reinforcement learning maximalist. I was always believed this is what we need to focus on. This is what we need to do.
5:56Jerry Tworek:I've seen LLMs being scaled up to higher and higher levels through GPT-3 to GPT-4. And we're still doing very little RL. And I had this internal belief that the moment we start scaling up RL, we'll solve everything. We'll be able to solve all the problems. And we eventually started scaling up RL. I was just in there. I was in the center of it. I was thinking, here we are. If you ask Jerry in 2024, when do we get AGI? I would say 2025 will be that year. This is where we solve everything. And I saw us training model after model. This model was getting better and better. All the benchmark scores were going up.
6:39Jerry Tworek:And did we also solve all the real world tasks at that moment? Unfortunately, unfortunately not. We still have work. And I realized there was this bit of distinction as all the benchmarks that we are evaluating our models, they were essentially the same thing as we were training the models. on. Like all the evals and training tasks are the same sides of the coin. But the real world distribution and real world task is much messier, much murkier, much more different. Our training data didn't really replicate the real world use cases and despite us basically maximizing all the tasks, if you see ask anyone training models, hey, what is one of your main issues?
7:23Jerry Tworek:I don't have hard enough tasks. I don't have what to train our model on. Yet we are still not covering the entirety of the real world distribution. From that, my conclusion is we need to have models that learn at test time. We need to have models that learn with users on their data, on their real world task, on the real world distribution. And there, when you are asked, why don't we have that today? Why are transformers not learning anywhere? And there are essentially two types of learning that we could be doing at test time. We could be doing and context learning essentially of transformers, which is, it doesn't have fundamental problems of catastrophic forgetting.
8:07Jerry Tworek:It doesn't have that issue. It is pretty data efficient. So that is great, but it's not very scalable. We only can have so much of it. It is limited and has some more, even more of mechanical limitations of what actually are you doing when you, when you build context, but maybe, maybe we can, we can come back to it later. but we have in-context learning, which is very limited and very small amount of data. Whenever I'm using codecs, roughly around 20 minutes of usage, I need to compact it and move it afterwards, which is not that much data. If all we can learn for 20 minutes, it's not that much.
8:44Jerry Tworek:And the second thing is fine-tuning. We could try to continuously fine-tune our models, but then those have the issues of catastrophic forgetting. We have issues of very low data efficiency and neither those are very solvable. Neither of those are very easy to find ways people have been trying. If those were easy to solve, some would already solve it. So my personal belief is we need to find an algorithm that we can meta-learn, that we can express on the architectural layer that can represent how How does learning look like? How does learning look like that can work on much, much longer horizons?
9:24Rohan Anil:Do you expect the architecture will look Transformer-like? Because my understanding from the chief seats is that, you know, OpenAI had been trying to scale up reinforcement learning for a long time. And it wasn't until the Transformer came about that it seemed like there was an even kind of scalable prior on the world upon which to even scale RL. And so how do you even go about trying to think about scaling up this new regime?
9:50Jerry Tworek:Yeah, that's a great question. I think those two things happened at the same time. But if anything, that happened there was mostly about economics, because technically you can scale up LSTMs. Just no one really, really dared to go in that direction. And they did scale much, much more poorly. they are they're they're scaling in a scaling loss paper there is uh presented a comparison of lstms and transformers and fundamentally the scaling love of transformers was better there is a world where we never invented transformers and we would be scaling lstms and we'll be having some models but because they would be much more expensive to train and much less impressive as a product we would have would have just a worse experience and maybe no one would be able to convince people to spend as many dollars training those gigantic LSTMs because we wouldn't get a market return.
10:43Jerry Tworek:The majestic thing about Transformer, which goes back to like, why do we have to appreciate Transformers so deeply, is that Transformers are economically valuable. Training them, the cost of training them is lower than the revenue that they generate, which is just magic of machine learning. It's not guaranteed by itself, but for LSTMs, it probably wouldn't be that way, which made it happen. But in many ways, you can scale most of the architectures. I think a lot of reasons why people didn't scale things before was because researchers before OpenAI had a lot of reluctance to scaling. It was often seen as unscientific.
11:23Jerry Tworek:And research in algorithms was providing how do we become more and more efficient? How do we, for the same compute budget, get better results? And it was a bit of a contrarian bet by OpenAI. at that moment to try to say, hey, we don't care about better and better algorithms. We care about more and more scalable algorithms. And how do we pour more and more compute and get better results, which OpenAI was criticized repeatedly by many people in the community for a long time. But thanks to that, we have the models that we have today. And I think there are tons of architectures that can be scaled up.
11:57Jerry Tworek:And I am part of the core automation's mission and our belief is that a lot of architectural research happened at too small scale for too long time. A lot of people are trying to say, hey, let's try to first try our architecture on a small data set in a small compute regime. And then see where it scales only after you prove itself. But for example, when you do work on reinforcement learning, you know that to get to any interesting results, you need a certain level of compute to even see the capabilities in the model. Reinforcement learning needs a baseline of ability to only start working. So where I am coming from, probably there are many architectures that need a baseline of compute to even start doing anything, anything interesting, anything useful.
12:47Jerry Tworek:Can I ask you then maybe a touchy question? Please do. If you need a baseline of compute, that sounds like a job that would be well served inside of a big research lab. Why start a company to go do this? It's a great question. And I think in many ways, it's likely a timing thing, timing issue. Market is right now in a very specific place where the biggest and the most successful labs, by coincidence or by fate, are probably in the most competitive market fight ever right now, which makes them not very keen on trying different paths, trying alternatives. if transformer is profitable and if you can spend more efforts and more resources scaling transformer to win in the next quarter it's very hard to put at least a lot of attention and a lot of energy to work on something that will that will maybe better or maybe or maybe will will redefine the field in a year or two so so i think the biggest labs and i talked to basically all of them don't don't have that much interest in trying the alternatives to Transformer.
13:57Jerry Tworek:And the labs that are not the biggest are doing whatever they can to do what the most successful labs are doing. And everyone is trying to train the same coding agent. If you look at the last week's releases, everyone is trying to release a coding agent right now. And I think we need different paths and different approaches here. So that's initially the ecosystem we are trying to fill in.
14:24Rohan Anil:And Rohan, you were at Brain when the transformer was invented. Do you agree with Jerry's eulogy for the transformer? Yes. In some sense, like, first when the transformers, Ashish, Noam, and others came up with it, I had, like, worked on my work on online distillation around the same time. We presented it at the same internal research conference. It wasn't a big deal internally. There was only a few people who actually got it. A lot of people were like, oh, it's like, yeah, it's another work. And people were finding ways to. And it was also very focused on, at least the original work was very focused on a real problem, which is translation.
15:02Rohan Anil:So they saw like the Beatles TM on translation. And it took opening. I mean, internally at Google, there was definitely like Noam and a few others were definitely interested in scaling language models. I think it is until GPT-2 and GPT-3 that we saw the benefit of transformers working quite well. At least the way I think about architecture is how do we spend computation? And Transformer is one very efficient way to spend computation. But now that I look at the industry, a lot of our computation is inference time and spending it on tokens. Let me ask this question. If I want to optimize for a better architecture, I want to look at both pre-training and RL together.
15:39Rohan Anil:And I would like to find architectures that spend computation much better than current chain of thought token generation. in at a like to give a much better overview it's i think of like pre-training has built the transformer with certain context length and rl comes in and it's like well that's not sufficient i need more computation let me do it via adding one token at a time this is quite inefficient from like inference perspective you're doing one token at a time so most of the solutions have been finding to do better ways of speculative decoding so it's like a band-aid to a problem that we've picked something that can only generate one token at a time, so autoregressive recording.
16:21Rohan Anil:There is problems with the transformer in terms of how do we spend the computation for the longest time. I think most of the world was training very large, dense models, and it took the industry roughly two to three years to get to refine the architecture to what we now take for granted. It was not obvious to a lot of people, sparsity and mixtures of experts and getting good training efficiencies with them, right? So then you can ask like, what's wrong with the transformer? Well, it's if the computational depth is poor, how do we increase computational depth? And just posing that question opens up like 20 new directions on how we can modify the mechanism to incorporate it.
17:03Rohan Anil:So I see like to do work like this, it takes time. and usually like fundamental research in the past have taken like five, six years to land into industry and it's largely from organizational, knowing that it is important. This is the bet. Like just like Jerry had the inner belief that RL is needed. Absolutely did not have that belief at Google. I was a pre-training maximalist, pre-trainer, biggest model. That's why you guys are a good fit. Right, like things are. And then, so that inner belief. And second is you need your architecture to run efficiently on hardware. A theoretically optimal architecture is not useful to anyone.
17:43Rohan Anil:It is something when it comes into practice. So you need the research inception to getting it productionized and getting kernels and everything written, the end-to-end loop. And there's only a few places right now which have integrated teams doing that. And I think we have built a team in a way that puts the experts together, not in different silos that we are accelerating on having everybody look at the problem holistically from end to end. So that's where I'm quite bullish. That's why I'm here. The current mechanisms are quite poor. And if you leave it to the world, I am afraid that it will take us a much longer time horizon before we replace the transformer.
18:26Rohan Anil:and I think a lot of folks are already complaining a lot on token costs. And that seems like as someone... We're not complaining. I mean, in terms of like, yeah, exactly. I come from the Google mindset where we had to like serve billions of people. So like finding more efficient architectures that have like a deadline on latency and the number of tokens that you can serve. So like when I look at that, like the amount of like the world that can use like frontier tech is very little. and someone or some group has to accelerate and make this better. And we are taking that shot at doing that. The current technology just scaled up is still only relevant to a subset of humans.
19:09Rohan Anil:And this is the bet we're making.
19:11Jerry Tworek:So one of the things I heard you say was the problem with transformers is the computational depth is poor. If that's the crux of the issue, tell us what does that mean? Why is that the case? How do you fix it?
19:23Rohan Anil:I can give you like one insight, like most transformers that we train are quite shallow. That's at most like 100 layers deep. Depth is like, it's called deep learning because you wanted deeper representations. There has been experiments on going into depth, but no one has actually shown us learning extremely deep representations. Chain of thought reasoning and RL to do chain of thought by model itself is one way to increase computational depth because every token you add, you add like one more pathway. And there has been, so then you can get out of like this bottleneck that the pre-trained architecture has set you up on.
20:03Rohan Anil:You can only do a number of layers times sequence length. Now you can increase the sequence length and you get much stronger results. You can do inference time scaling. Now the issue with inference time scaling is that models now have to produce more tokens to get better results and that's very one token at a time. And from this, you can see like you can directly address many of these things. And this is like a subset of work that we are looking at, right? Making this much more efficient.
20:32Jerry Tworek:What's your forecast for the transformer-based architecture? If it's not the end state, how far can it get us? When do we start to see it topping up? I think it all comes back to what we are training transformers for and what we can do with them. we're doing pre-training which is very good at distilling all the knowledge from the internet into transformers and then we can aral them through which is we basically can bake all the workflows that we want into a transformer so what what what transformer is like capped out is we have all the knowledge of humanity in the model together the relationships and how do they how do they work together how they can be combined and basically any task that we have training data for we can put into the model and this could be a gigantic model trained with a lot of compute on all the data in the world.
21:26Jerry Tworek:And then if we ever stop training that model, what would happen? The question worth asking often I'm thinking about transformer, what would happen if OpenAI and Anthropic stopped training new models and we got the transformer we have today and say this is it. This is the best model we have. Months pass, year your spas and the model is getting less and less useful. Maybe the lab really like recorded of every human on earth what they were doing and what their tasks were and their environments and put them in the model, put them in a learning environment. But then what happens if anything of that changes?
22:06Jerry Tworek:If there are new events in the world, if those new events have new relationships between them, if there are new types of tasks, if there are new code bases, new tools to use, transformers are getting a lot of their usefulness and value through the things that are valuable have to be present in training. And when they are not, they suffer. There's some ability to adapt, but it's not very big and not very flexible. So in my mind, this is kind of the level where the transformers top, which in many ways what I think is a tool to use for us. If we kind of know, if there's a human who knows the limitations of a transformer, they can schedule that model, they can write a prompt of what is the task that you want and by doing the training we are doing, you can get very successful in that and any task the model fails, you can add it to the training data and the model can succeed but that loop has to go through the lab training the model for you.
23:10Jerry Tworek:And if the model that fundamentally needs to be trained in the lab, how much do you think of it that this is the goal or you would want to be able to update the model somehow, not having to go back there?
23:22Rohan Anil:Have you read the Rich Sutton and David Silver have this paper, The Age of Experience? Have you read it? I'm curious how much you agree or if you have any different opinions, whether your opinions diverge.
23:32Jerry Tworek:Reinforcement learning is not a particularly new approach, particularly new thing to do. So in some way, age of experience, I think always has been there. And people have been criticizing a bit pre-training because pre-training very clearly is the other way of looking at the models, which is like we have static data, that data is mostly generated by others. Although I have this personal view that pre-training today is largely distilling other models, all their malls into the new mall because most of the tokens in the internet are coming from AI. But there's clearly pre-training, which is behavioral cloning, which is mimicry, which is compression of internet data.
24:17Jerry Tworek:But reinforcement learning is not something that people haven't been thinking and people haven't been doing. Reinforcement learning was used to solve by Gammon back in the day, that it used to solve Go, StarCraft, Dota, to solving programming right now and every time it comes down to model writing its own experience and learning from that experience but it is very clear and what I think is interesting and what I think is still perplexing to people that reinforcement learning is not really the only way to learn from experience and there will be more and there will be a little bit more of I think you can call it algorithmic but essentially innovation of how we learn from experience just because just because reinforcement learning is only one way to do it a mathematical formulation and especially right now how we are using it it really likes those parallel rollouts for variance reduction and for uh for comparing how the model does in parallel versions of the world which is not not how we do not how we learn from experience we learn from our experience much more efficiently and much more um we use we use those in many in many, many ways.
25:30Jerry Tworek:At some moment, I've been trying to explain to people that what brain does, how we learn. It's not that there's one learning algorithm in a brain. I think there are multiple, actually, and they work together. But if I play football, for example, it looks very, very closely to reinforcement learning. I kick a ball a lot of times, and every time I adjust it a little bit, and I see if it roughly matches what I wanted, and there's some self-reinforcement happening. When I learn mathematics, it's a very different type of thinking. It's like reading about hard concepts and thinking about them very deeply inside my head until things click and until I have them connected.
26:11Jerry Tworek:And both of those in some way are learning from experience. They are just very different. So summarizing my thinking of the learning from experience is that we've been doing it for a while. We probably are spending the most compute than ever on learning from experience. but reinforcement learning is not the end of learning from experience. And there will be better approaches that researchers will be coming up in the coming years on how to use that data in every nature settings.
Read the full transcript
26:39Rohan Anil:Interesting. Rohan, I'm curious, since a lot of your work has been around optimization and efficiency, how do we get to orders of magnitude more efficient, I guess more compute efficient and more data efficient learning algorithm? I'd start with measurement. I think pre-training as we define it right now is about compression. We look at perplexity and then measure how do we decrease the perplexity. And then we find that scaling and increasing parameter count and putting more compute is the way. And every time we increase compute in log scale, we get this epsilon more improvement in this matrix.
27:19Rohan Anil:This is, I think, this is fine for building the prior. but I think this is the wrong way to look at the problem. We should be looking at the end-to-end. What are we training these models for? Look at the outcome. Like, for example, I trained this model and give it to Jerry. Jerry will do RL and destroy all the perplexity metrics that I have created, right? So then it's sort of like it was the best way we had so far to attempt to solve the problem. And I think the labs and everyone else have done a great job and producing intelligence that's super valuable and makes my work so much fun. But it was the bootstrap process to get there.
27:57Rohan Anil:We have to combine pre-training and RL together, and that's where one order of magnitude improvement would come from, and that's like a training procedure. You can say it's a learning algorithm. In terms of optimization, my story is I started optimization at Google for logistic regression back in 2016, got nerd sniped by it, worked on some solvers for what we used to call SIBL, which was the large-scale linear solver that was used at Google before neural network took off and then replaced this. Then I asked myself, like, what do I want to work on with neural network? And it was quite clear, like, I want to understand the training algorithm and make it better.
28:38Rohan Anil:And then someone, Benit Gupta, just showed up one day at my desk. It's like, I heard you're really good at writing optimization methods, we have this idea that, you know, like we worked out on a whiteboard and what turned out to be the shampoo algorithm, can you help us implement scale, make sure it works at large scale for training? So when I was working on this, I thanked my manager, Yonghui Wu, who supported it throughout that until the end of my tenure in 2024. But largely, the community and most of the other people were not as excited by this idea. and for me this is the most exciting thing because i was like i'm putting in computation and making training better this is the thing i have to figure out i will spend as much time i would take to do it and then people were making this assumption oh why like what's the upper bond you could still use adam that's fine like why are we you could spend all the time on everything else not optimization but in some sense optimization is like like you have a model you're optimizing it you want to optimize it better.
29:41Rohan Anil:Now connected to some of the stuff that we talked about architecture, what has happened is that a lot of the work that we've done in architecture is to make these networks train. And in some sense, it's like two sides of this coin and optimization on architecture go together. You could have a stronger optimizer train a much more harder to optimize model and get better performance or we can use a weaker optimizer on easier to optimize models and get decent performance so there are these trade-offs that appear all over and for me it's i spend a lot of time working on it i think we used it for gemini 1.5 flash and then like the community started getting like more interested in it there was the soap paper published we have like an entire literature of like shampoo soap all like bath time the things that you would use and then it It was quite clear.
30:32Rohan Anil:So that was like maybe a 2x improvement over what was happening. But even then, if you look at Shampoo, it's quite weak in what it's doing. It's not using all the information that's available to you when you train. And as you use more and more information as part of training, you can get better improvement. And in some sense, like your optimization algorithm defines what architectures you will discover. Like I have colleagues, it's not very popular in the literature. it's only like maybe like four people in the world care about it kind of ideas that are extremely interesting like residual connections have been extremely useful for training neural networks there are folks who've now like gotten rid of them and learned deeper representations but they needed a better optimization method so like for me optimization methods and the question you asked is just like how do we get there it's combined with architecture and thinking about the problem end-to-end is where a lot of the computational efficiency is in.
31:30Rohan Anil:And I also see Aral as spending a lot of compute, not as efficiently. And so if you could spend it, because you don't get much feedback and you're spending a lot more compute because you have to decode all this long chain of thought to get this one bit of information into the network, seems quite an efficient and easy target to get orders of magnitude on Tompa. I can go on talking about optimization all day. I love it. Do you think we'll ever approach or surpass biological learning efficiency? I do not think so because I think we would need to change. Maybe like that was a strong statement. At least with the hardware we have, it seems pretty unlikely.
32:12Rohan Anil:Our biological, like we have something, as Jeff Hinden says, model computation. So we built our own circuit as we grow up and we built our own learning algorithm with the hardware. and then we die and then we're gone. Neural networks are quite different. The hardware stays, the neural network stays, but it's learning very inefficiently and you need a lot more of them and a lot of parallelism to get small amounts of information through. So until I think we design hardware to be much more like how humans operate, maybe more analog, figure out how to deal with analog circuits and figure out how to do with error correction, figure out how to get information through, it will be much harder.
32:53Rohan Anil:I think we are safe. Safe. That's an interesting way to put it. The idea that pre-training and RL should be optimized end-to-end seems like such a clear, maybe obvious statement. Do you think the labs realize this? And is it just hard for them to get rid of org charts and process to be able to make that come together? Or what stops the labs from being able to unify the two? I don't think it's as obvious because it's a completely, again, different optimization problem. You have a prior, you're doing roll-outs, you have higher variance, and then pre-training is much larger batch and like more parallelism or compute for the unit of time that you can spend.
33:40Rohan Anil:So it is not an obvious thing for folks to combine these two training procedures until you think a bit more like, why is it that the naive combination doesn't work? So that's one. The second one is, if I poll some of the best researchers in these labs, they would say, oh, this makes sense. We should probably explore it. But it would be probably not in the top bucket because they have to train a model for the next cycle before. It's like, as Jerry said, there are companies now competing for release cycles because tokens are not sticky. So it's much harder to do long, slight, like even long-term research of six months in many of these labs in the environment they added.
34:24Rohan Anil:So it seems like one of the core premises for core automation is, you know, you're starting at the lab at a time when, I know Sam's been talking about the AI scientist. I think Dario's been talking about the AI scientist. It seems like your job as researchers has actually fundamentally changed. And you get to start the company native to that era. And as a result, you're maybe able to run a lot more experiments than otherwise might be possible. How automatable do you think the research job is? And how are you guys approaching building your lab to be as, I believe your mission, one of your missions is to be the most autonomous lab there is?
34:59Jerry Tworek:Most automated lab. Automated. Yeah, lab in the world. And to start with, I think the automation, the version of automation by core automation is about giving each human maximum level of agency in some way. It is, we are not trying to really get humans out of the loop, which is like one version to automate, but it is about give humans ability to do the most with their amount of time. Whenever you are walking, you can get some distance. Whenever you get a bike, you can go walk a larger distance. whenever you are a car you can can go much much much much larger when whenever whenever uh humans started farming they had to farm by hand and work on a small plot of land when you have a machine you work on a much much larger plot of land uh personally i am i'm both really great fan of the of the current coding agents and and very happy it's in some way it is what i've been working for many years both doing coding research and working on various versions of of ai scientists inside inside of open ai and in the end i realized starting a company to realize that vision as is one of the best one of the best ways to realize it because the way you can do research today is very very different because like single researcher can do much more in the end the speed of iteration the speed of research, the speed of how quickly you can move through ideas and how quickly you can get data on your ideas is something very, very different.
36:29Jerry Tworek:And you can try to move the old structures around it, the teams, workflows, how data is gathered. or we can try to build, like you said, we can try to build natively for it, for processes that maximally empower each researcher and allow them to just iterate on their idea much quicker. We are here and we are trying to rebuild the full deep learning stack and try to think how we can do almost each operation differently, what are various options. And if we can execute at least even one of those experiments a day, that's already a pretty good iteration speed versus anything that was that was done before and there isn't really like any any fundamental like laws of physics reason why not and maybe maybe one day we get to 10 of those a day maybe one day we got we got 200 those a day and fundamentally for that like search process optimization process we should be able to just just find things that work in a better deep learning setting and like what we are trying to do is like we've been like almost all of us we are we are a team that is very agent-filled and automation-filled we are we're trying to do an experiment like how far how far we can we can push those things and how far how much an organization that tries to do as much as we can with a small team how far how far we can get with that when will we know that we've reached AGI at some moment I used to say it's very much in everyone's heart whatever whatever they consider AGI.
38:01Jerry Tworek:OpenAI is the system that can outperform all humans in economically valuable work, but it goes to my previous statement, what if OpenAI stops training models? Would that still keep working? Would that still keep the automation level the same or would that drive? And for me, AGI is a model that can improve itself without human in the loop in in any way. That's, I think, the moment where we can meaningfully talk about AGA because that is, in some way, it is a sub-definition of the previous one because improving AI model is actually a job that humans can do. It is economically valuable work and it definitely is the case.
38:45Jerry Tworek:But removing humans from loops with models has been actually notoriously, notoriously difficult so far. We haven't come anywhere close to it. It's very, very hard for me to find any task where all of us we're able to get humans out in the loop. We are, the human LLM hybrid is really, really successful right now, but LLMs without humans, not so much, not at all. And what I have seen there in 2024 and in already 2025 is that the current path doesn't get us there. And I think we need some pretty serious research on this company or some other to try to unlock how do we make our models learn and adapt at a test time on a deeper level than we've been so far.
39:32Rohan Anil:I feel like we've alluded to this throughout the conversation, but it's kind of been one of these five blindfolded people trying to find the elephant things. What is the grand master plan for Core Auto that you're willing to share? I can share our six-month roadmap. In some sense, building architectures, as I said, it's not just how good the architectures does it run well and can we get users, including ourselves, as part of the lab to use it, right? So then that's directly like we can do many things now, but the thing that is going to be difficult and we want to automate away is kernel generation.
40:09Rohan Anil:So we have a set of hardware, GPUs, blackwells that we have to train and run inference on. We will build the best model that we can to basically reduce the time from having a very cool idea that can make these bottlenecks go away in the architecture to having them run at the highest TFLOPs on GPUs. And in some sense, like current coding agents plus humans can go a long way. But like an example of this is our QR kernel competition that we hosted with GPU mode. It's for running this very old linear algebra operation QR. It's used for optimization, like Shampoo Line of Work uses it, many other places uses it.
40:56Rohan Anil:And you want to run this efficiently on B200 node. If you use CoSolver for the shapes that we care about, you get some efficiency. And then a human plus some search loop can get you something like 7x. But it requires the high-taste human, like there's maybe three people in the world, and spend about$100 ,000 on these coding agents over a span of four weeks to get to a solution that's 60x faster. So these models today are nowhere close to getting that 60x faster kernel. And there is a real bottleneck. Now, that was a single problem. It has like perhaps three different operators work on this panel, do this matrix multiply, fold it back in and do this repeatedly.
41:43Rohan Anil:That's what the skewer factorization of a matrix would look like. And if you give this problem to entropic models, OpenAI models, Gemini, it just wouldn't solve it. It just, it's not, our models are not even close to solving this problem. So for us, it's like something that we've talked about, something that we're getting close to as sort of like getting to that point because that's our inner loop to having more efficient architectures. Why kernels? Is it just because like maximize intelligence per flop of compute? You need to do kernels? In some sense, I've had like three projects two of them have kind of landed in the industry so first is secondary methods, kernels was a bottleneck because you have to run it well if you were at a place like Google you cannot spend 10x amount of compute and get a 2x win so I could only spend maybe a budget of 20 % and get the 2x win everyone's happy like great so like I think that's the market right you spend less than you get.
42:43Rohan Anil:So kernels ended up being a bottleneck there because most of the operations were novel that we haven't gotten a lot of people to look at it. There's only like two humans at Google who could write it, Rasmus and Peter Hawkins, because it was deep XLA LLO code that you have to write to make this work. And then that took them two years to do. The other idea I had with one of my workers at that time was like replacing some of the parameters in a transformer with extra memory and we called it ngrammer ngram memory we worked on it in 2020 it's a we had versions of it internally deployed not the big version the smaller version but there i needed like something that can accelerate sparse gathers and scatters as part of training yeah it required hardware change and hardware making use of the hardware it never arrived i had conferences set up with the tpu team us and a bunch of others we were talking about it and during covid like oh we're gonna have it happen and it never arrived i also was using tpus at anthropic while i was leaving just barely started the surface of being able to do it but at the same time six months before that DeepSeq wrote their Ngram, which is an improved version of adding more memory, showed scaling loss that, yeah, you don't need MOEs.
44:04Rohan Anil:You could actually replace it with these Ngram embeddings. For me, that was like, oh. Yeah. Very interesting. It was like a five-year thing. And I was very happy for them. So much of the state space to explore isn't even possible if you're not writing kernels. Kernels. And you need to be assisted in writing kernels or solve that kernel to have the highest performance. and the roof line is pretty high so it's like the QR. If I use QSolver's QR, I get some performance. If I use the competition winner's QR, you get 60x faster. And that is a completely different playing field. Now it opens up an entire new set of algorithms you can apply in terms of training transformers, training optimizers.
44:46Rohan Anil:QR is so fundamental in analyzing the two for eigen decomposition and many other things. So it is a thing that I think also if you think about it only few people have the skill set too and they're very much not at the same place it's like one person here one person there and it would be ideal if models
45:07Jerry Tworek:had those abilities yeah maybe i'll summarize a little bit and talk from the high level of what we want core automation is a lab created to build models that continuously learn and learn from deployment we believe as as i mentioned that transformers are incapable of continual learning there's no way how to put continual learning on transformers so we know we have to find a different different architecture some of anyway our quest is to find that that new architecture find that transformer replacement and we want to build the most automated lab to do it we want to be able to build experiments at scale the quickest we can iterate on them try a lot of new architectural ideas have strong priors of what we want to do to search the space of architectures efficiently to find to go to that go to that place fastest than than anyone else that's what we what we want to do and all the work we are doing on kernels on large scale training on trying new architectural ideas is exploring that space.
46:11Jerry Tworek:So you're going to experiment your way into finding a superior architecture. How will you know when you've found it? What are you looking for to say, aha, this is the one? That's a great question. There are always two angles. In my mind, in my experience, every successful research had a plot that shows something that other plots don't show. there's one line that is a little bit bending in a different different way and you're saying this this is what you want but at least it is my experience with research always has been that plot is already quite late in a journey where most of the time you already know what you want and already know what you are up to i i am a bit a bit joking but it's actually true that all the best plots in my life i have them in a dream before before actually they were they were real I kind of knew what I was looking for.
47:03Jerry Tworek:Just the question is like when it actually clicks, if you know what I mean. Because most of the time, you kind of know what you are looking for, but you are not finding it. You try one thing, it doesn't work. Try second thing, it doesn't work. But eventually, all the right pieces fall into it. And most of the deep learning systems are very intricate. So usually, you have to get five things right in a row for the mink to start working. And then eventually, you get the plot that looks like you want. and then you know. So I think what we are looking for is systems that learn at test time and if we see meaningful long-term adaptability of our systems and like we're joking, but it's a real, we want to be evaluating our systems of our everyday work.
47:52Jerry Tworek:They get better at doing the work of core automation scientists each day.
47:57Rohan Anil:Yeah, we like go on a vacation as a team and see if the lab produces something better for the week. Then what do you do when you get back? We'll see.
48:07Jerry Tworek:Extend the vacation two times instead of four times until we are on permanent vacation.
48:14Rohan Anil:That is a beautiful note to end on. Rohan, Jerry, thank you so much for joining us. You've both worked on really, really transformative work for where we are today. and I'm so excited to see you starting a lab on this new journey and very excited to see what you're able to come up with. Thanks for joining us.
48:36Jerry Tworek:It's been great to be here and chat with you. Thank you.
48:57Thank you.
From the publisher
Jerry Tworek led reasoning at OpenAI, convinced that scaling reinforcement learning was the path to AGI. Rohan Anil co-led Gemini pre-training and built the Shampoo optimizer. Now they've teamed up at Core Automation on a contrarian premise: the transformer has carried us as far as it can, and the bottleneck to smarter systems is no longer scale — it's the architecture itself. The missing capability is continual learning, models that adapt at test time, which transformers can't do. In-context learning taps out fast (Codex needs compacting after ~20 minutes) and fine-tuning invites catastrophic forgetting. Rohan argues pre-training and RL should be optimized end-to-end, and that transformers spend computation inefficiently. They lay out why the largest labs won't chase alternatives while locked in the coding-agent race, and why building the world's most automated lab starts with automating kernel generation—the one place frontier models still lose to a high-taste human.
Hosted by Sonya Huang and Pat Grady, Sequoia Capital




