In short
How modern LLMs are trained and adapted, focusing on pre-training vs post-training (especially RLHF/RL), and how AdaptiveML applies reinforcement-learning pipelines to tailor smaller models for enterprise use.
Guest backgrounds
Julien Launay is co-founder and CEO of AdaptiveML. He has prior experience creating LLMs at scale, including at Hugging Face and LightOn. He recently moved to New York from France. He also wrote a French Minecraft guide in high school and learned programming via Minecraft plugins/mods.
Key claims
- Pre-training mainly teaches next-token prediction on massive web text; by itself models can be unchatty and produce “related questions” rather than direct answers.
- Post-training “sharpens” behavior; RLHF uses thumbs up/down feedback, while newer approaches scale with verifiable execution rewards and AI-as-judge (synthetic feedback).
- Reinforcement learning is characterized by online, trial-and-error learning from model-generated samples, unlike supervised fine-tuning which often shifts the data distribution abruptly.
- AdaptiveML’s “Adaptive Engine” simplifies RL operations (RLOps) by orchestrating judges, environments, and distributed pipelines.
Notable examples
- Thumbs up/down in ChatGPT as RLHF feedback.
- Verifiable rewards via math/code tests (e.g., GitHub tests).
- GPT-4 used to replace expert human review for a specialized task when outputs were indistinguishable.
- Synthetic data scaling: ~70 human critiques/re-writes generating ~80,000 self-play conversations, then exploring 5–10 candidate answers per prompt.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOJulien Launay's Journey and Background
0:39 to 3:50
Julien shares his journey from writing a Minecraft guide to co-founding AdaptiveML.
“Yeah, it's great to have you in person in New York.”
Creating Large Language Models: Pre-Training Overview
3:50 to 7:25
An overview of the pre-training phase in developing large language models.
“adapted to a business using smaller cost-efficient models.”
Post-Training: Refining LLMs for Real-World Use
7:25 to 9:20
Discussion on post-training techniques and the shift towards reinforcement learning.
“I put high quality in quotes because the definition of quality is a more other subject that we could spend hours on.”
Advanced Reinforcement Learning Techniques
9:20 to 14:01
Exploration of various reinforcement learning strategies used in AI training.
“I think that shift has been happening behind the shadows for a while and now is getting fully executed.”
The Role of Synthetic Data in AI Training
14:01 to 17:32
Learn how synthetic data can enhance AI training efficiency and quality.
“on data that is produced by other models, also is seeing a very big growth and very big success because it works so well.”
Reinforcement Learning Techniques Explained
17:51 to 27:32
Understand various reinforcement learning algorithms and their applications.
“but our technical episodes are some of our most popular ones so we're going to dig into it a little bit here.”
Adaptive Engine for Enterprise AI
27:33 to 28:00
Explore how the Adaptive Engine utilizes reinforcement learning for enterprise applications.
“So I mentioned right at the top of the episode something called Adaptive Engine, the flywheel for enterprise AI, which is continuously evaluating, tuning, and serving LLMs uniquely adapted to a business.”
Challenges in Reinforcement Learning and LLMs
28:00 to 30:27
Explore the engineering challenges in implementing LLMs and their differences from traditional reinforcement learning.
“you had that experience of going, you were like, oh, these are amazing methods.”
The Power of Pre-Training in LLMs
30:27 to 31:42
Learn why pre-training is simpler and scales effectively compared to reinforcement learning.
“You compare the top one to what was in the text, and that's it.”
The Power of Pre-Training in LLMs
32:33 to 33:01
Learn why pre-training is simpler and scales effectively compared to reinforcement learning.
“I'm excited to announce that I've launched my own AI consultancy, a firm called Y-Carrot.”
Show all 27 chapters
The Power of Pre-Training in LLMs
33:05 to 33:16
Learn why pre-training is simpler and scales effectively compared to reinforcement learning.
“Again, that's ycarat, Y-C-A-R-R-O-T.com.”
Adaptive's Reinforcement Learning Tools
33:16 to 35:48
Understand how Adaptive's tools simplify the implementation of reinforcement learning for enterprises.
“the current adaptive platform is designed for people like software developers, ML engineers, who want to have reinforcement learning be easier.”
Expanding the Accessibility of Reinforcement Learning
35:48 to 37:15
Discuss the future of Adaptive's offerings and the potential for broader access to reinforcement learning.
“because we think it is at a point where that can be the case and where everyone should be able to run this sort of stuff, even for OB projects, actually.”
Synthetic Data and Model Performance
37:15 to 41:11
Learn about the role of synthetic data in enhancing model performance and reducing reliance on human-annotated data.
“these synthetic data pipelines, anyone can run them, anyone can define them.”
Layering on Foundation Models
41:11 to 42:00
Explore how Adaptive builds on existing foundation models to create tailored solutions for specific tasks.
“of foundation models to tailor them to final use cases.”
Understanding Pre-Training and Post-Training
42:00 to 44:14
Learn how pre-training sets a strong foundation for models and the role of post-training in enhancing abilities.
“We start from them and we tune them to perform better.”
Reinforcement Learning in Practice
44:14 to 46:32
Discover how reinforcement learning methodologies can empower data scientists and enhance model performance.
“Yeah, I think for us, you know, like a lot of what we do now is bringing, you know, this expertise in reinforcement learning to be something that anyone can do.”
Defining Success in Reinforcement Learning
48:25 to 54:38
Understand the importance of defining reward functions and measuring success in reinforcement learning systems.
“You mentioned there's something that I want to highlight a little bit that these amazing teams at AT &T are doing is they're getting that definition of the reward function, right?”
Challenges with Data and Future Solutions
54:38 to 56:00
Examine the limitations of data availability and the potential of post-training methods to overcome these challenges.
“And we are seeing this now, you know, for a while people were seeing reinforcement learning more as like, oh, this only specialization layer or only, you know, like post-training layer.”
Challenges in Computing Paradigms
56:00 to 58:35
Explore the difficulties and complexities surrounding current computing technologies and their evolution.
“What do you think about this hardware problem?”
The Future of AI and Experimentation
58:35 to 1:01:12
Discuss the potential of AI systems to run scientific experiments and the challenges that lie ahead.
“Is there room for improvement, for more specialized hardware, for that sort of stuff?”
The Nature of Superintelligence
1:01:12 to 1:05:38
Delve into the implications and varying levels of superintelligence in different domains.
“Like at some point there's obviously a bottleneck in the real world.”
Optimism for AI's Impact on Humanity
1:05:38 to 1:07:48
Consider the potential positive outcomes of superintelligent AI systems for society.
“And it's very hard to predict, you know, which direction is going to work so well, which is not.”
Post-Training in LLMs
1:10:02 to 1:10:42
Explore the concept of post-training in large language models and its relation to alignment.
“allows you to transform those aliens into helpful, grounded assistants.”
The Shoggoth Analogy
1:10:42 to 1:11:51
Discuss the Shoggoth analogy in the context of AI and alignment challenges.
“And yeah, that's an idea that comes from here.”
Julien Launay's Insights
1:11:51 to 1:12:48
Julien shares insights on reinforcement learning and how it shapes LLM development.
“S-L-I-P-P-Y-L-O-L-O, which is a very long backstory.”
Future of Intelligence
1:12:48 to 1:13:17
Julien predicts the timeline for achieving superintelligence and discusses computing technologies.
Transcript
Automatic transcript. May contain errors.0:00Jon Krohn:Welcome to episode number 913. My guest in today's episode is Julien Launay. He is unbelievably knowledgeable about training LLMs, the pre-training part, the post-training part. We spend tons of time talking about that so you can get a full understanding of how cutting-edge AI models are made and how his startup, AdaptiveML, allows enterprises to have fine-tuned models for their particular use case available much more easily than ever before. This episode of Super Data Science is made possible by Dell, NVIDIA, and AWS. Julian, welcome to the Super Data Science podcast. Thank you very much. Happy to be here today.
0:43Jon Krohn:Yeah, it's great to have you in person in New York. Totally, yeah. And so it doesn't sound like you have a New York accent, though. I don't, I don't. I come from France. I'm spotted within the first few minutes. I come from France. I actually moved to New York just a few months ago. So it's very recent for me. I think just over two months today. Welcome. How are you finding it? It's really good. I think, you know, people ask me this a lot. And I think it's an interesting question because I think it would be really hard not to enjoy New York. Like I feel every time I feel very boring saying, oh, it's really good.
1:10It's really good because I don't really know honestly what negative things. It's an amazing city.
1:15Jon Krohn:I think it's basically infinite because of its size. You know, the restaurant turnover, new galleries opening. There's always new things to be doing. but I think especially in your first few months like this, it's so exciting because you're like, wow, a neighborhood like this, I had no idea it existed. Yeah, yeah, it's really amazing. Very diverse, lots of stuff to do. Like it's kind of endless. Like there is always, always activity. Like it's super nice. Really, really nice place. Nice. And so speaking of exploration, you're actually, you're our guest on the show because you wrote a bestselling book.
1:42Jon Krohn:Absolutely, maybe. 10 years ago called Aventure sur Vie et Création, Le Guide, Minecraft. So it's a Minecraft guide that you wrote in high school, is that right? Exactly. In high school, I used to play way too much Minecraft, I guess, like many people of my generation. Maybe it was a new generation, apparently it's making a comeback. And I used to write for a website called Minecraft.fr, so French domain. And one day Pearson's editor contacted us and was like, oh, we could do a guidebook. It's having a lot of success and ended up being part of this project. And surprisingly, it ended up becoming, I think, the year it came out, it ended up being like one of the bestsellers in French which is really funny because I can say that I wrote a bestseller in French well it's a video game book so you know your knowledge may vary but yeah.
2:30Jon Krohn:And this actually this book didn't actually show up in our research of you but you mentioned it before we started recording however it is kind of interesting because this kind of got, did this get you interested in programming kind of in the first place? Yeah yeah so not so much a book per se but definitely Minecraft definitely got me into programming like you know plugins and mods, all of that sort of stuff. I used to run a few servers, a few very large servers with a fund. And this was, although I think very interestingly, this was a time where Minecraft kind of professionalized, where it was starting to be like a very large server with tens of thousands of players that started making a lot of money actually on the side, kind of like the ecosystem started picking up, which in itself, by the way, is an entirely other story.
3:12I think like the world of Minecraft is actually fascinating even from like a business perspective and how it grew and all of that. But yes, that was very much the beginning of this and ended up doing some modding, some all of that, learning Java, like doing this and spending, once again, too much time on this. Maybe, you know, school performance dropped a bit because of that. But in the end, it all worked out. Yeah, it seems to be paying off.
3:35Jon Krohn:You are co-founder now and CEO of a firm called AdaptiveML, who are makers of something called the Adaptive Engine, a flywheel for enterprise AI which continuously evaluates, tunes, and serves large language models, LLMs. So they're uniquely adapted to a business using smaller cost-efficient models. Before we get too much into Adaptive, your company, I'd love for you to talk about, based on your rich experience at Hugging Face, also at a company called LightOn that we'll talk a lot about more later in the episode, through that experience, you have tons of experience in creating LLMs that are useful.
4:14Jon Krohn:for real life and at the biggest scale that LLMs come. So I'd love for you to start off by providing us with an overview of the steps involved in creating an LLM like pre-training and reinforcement learning. Yeah, yeah. It's a very timely question as well, given that I think these steps are blending a bit these days. So take everything that I say with a crane of soul. There's always nuance in this. But very broadly speaking, the way that historically large language models have been kind of approached. First is through a pre-training phase, which is the bulk, you know, historically of where the computer has been spent.
4:48Pre-training, you know, is during pre-training we essentially collect data from all over the web, pretty much every book, every paper, pretty much nearly at the scale of modern pre-training, nearly every text in existence. I think it sounds very grandiose, but it's not far from being true. And even nowadays images, images, videos, and all of this. And essentially the model is trained to very roughly predict the next world, predict the next token. This is a step that is built to be scalable, to run at scale that are essentially everything we have ever produced on tens or hundreds of thousands of GPUs these days.
5:24But pre-training is only a first step because immediately after pre-training, models are actually a bit unwieldy. If you take really pure, pure pre-training and you try your model immediately after, it's not going to be very interactive with you. It's not going to be chatty. It's not going to answer your questions necessarily in the way that you expect. I think a failure mode that we used to see a lot immediately after pre-training is, let's say, I ask a model a question. And instead of answering the question, the model will come up with 10 more questions that are similar. And the reason why is because in its pre-training data, this is equally likely to have like a list of questions asked to have, you know, the answer following the question.
6:02And this led to the development of second phase in model training, which is called post-training. And the idea of post-training is to kind of like own in, sharpen the model to really fit how it's going to be used, which typically means making it, you know, a good chat assistant or something like that. And the methods that you use during post-training typically differ. I mean, strictly speaking, you could do post-training in the same way you do pre-training, but with just data that is specialized, you know, maybe like just only transcripts of chats and continue doing pre-training on transcripts or chats only, and you would de facto be doing a post-training towards a chat model.
6:38But very often people, like the big success of post-training has been the use of reinforcement learning, so essentially enabling models to learn not from an explicit demonstration of what they should be doing, which is what supervised fine-gening and what pre-training are, but instead from a feedback about how are they doing. So the model generates an answer, and then from a human, from another model or from many different possibilities, the model gets a feedback of like, this is good, this is bad. And just based on this positive or negative signal, the model learns to improve.
7:08Jon Krohn:So this is like the experience that a lot of us will have had in ChatGPT where there's like a thumbs up or a thumbs down that you can click after you get a response and that can then be used as a training data for this post-training phase. And that'd be reinforcement learning from human feedback, RLHF. Yeah, from a very, like, yeah, from a very, very high level point of view, this is an example of the sort of data you could be leveraging to power this phase of post training i think what's really interesting is you know right now i'm giving a description where pre-training and post-training are very separate things so reality is much less so these days first because now pre-training is very dynamic where you shift the data distribution so you know you might start with like the lower quality data as a more like bulk data and as you advance through you know through steps of pre-training, you will focus more on higher quality data, maybe more code, more mathematics, more, it could be, you know, many like more chat data, more, you know, like more of the higher stuff that you consider high quality.
8:05I put high quality in quotes because the definition of quality is a more other subject that we could spend hours on. And post-training itself even now people are starting to do reinforcement learning during, you know, the pre-training step or starting at some point, you know, where they start to incorporate mixed blend the two. It used to be that post training was a much smaller spend than pre-training. You know, most of the money used to go to pre-training and to like, you know, the millions, tens of millions, hundreds of millions of dollars used to go there. But now if you look at recent papers, you know, like Kimi or even Grok 4, not really a paper, but more something that they mentioned, which is that they spent nearly as much on post training as pre-training.
8:43So there's massive scaling up of this post training phase. Yeah. So.
8:47Jon Krohn:Yeah, exactly. That seems to have allowed Grok 4, for example, to be able to get the highest score yet on humanities last exam, at least at the time of you and me recording this. Which might change, you know, like in a week. With another model coming out, it's always moving. But yes, definitely, I think, you know, one of the reasons Grok 4 has been so impressive on many of the benchmarks is a larger part of post-training that goes into it, a larger focus on reinforcement learning. But I would say, you know, obviously props to the Grok team for being some of the first to put out this sort of artifact.
9:18But I think there is a lot more coming. I think that shift has been happening behind the shadows for a while and now is getting fully executed. And I think most of the modern models are going to be going through much more extensive post-training than pre-training. Part of the reason why is also because post-training data is, if you think about it, I don't want to say more plentiful because it's a very complicated subject, but you can generate new data and new problems that the model is going to solve. When I was mentioning the feedback before, the thumbs up, thumbs down, a big trend that people are probably aware of with DeepSeq, with DeepSeq R1, was verifiable rewards where essentially the model solves a mathematics problem or submits an answer to a mathematics problem and then that answers get evaluated and if it's right, that's a positive signal.
10:12If it's wrong, that's a negative signal. And obviously these sort of things are very scalable, you know, mathematics problem, code, like code problems, same thing like tests, you know, for codes, you could use that as a signal. So it's very easy to imagine, for instance, mining all of the GitHub repositories that are available, pulling all of the tests from them, having the model, you know, write, you know, code that needs to pass this test and using that as a signal at a very large scale. And this is something, you know, that frontier labs do, and it is extremely effective. So there's a plurality of signals that you can use that is massive.
10:42And I think now people are very focused on scaling these massive environments in which to run the models to get these signals.
10:49Jon Krohn:Very cool. And so we've talked about reinforcement learning now in this post-training step. And we've talked about RLHF, where you have human feedback, like the thumbs up, thumbs down. What other kinds of reinforcement learning approaches are out there? Yeah, so totally. So there is historically, you know, the big one, the big first one, the big acronym, you know, that caught up a lot was RLHF, which is reinforcement learning from human feedback, where you are using your typically annotators, a company like Scale, you know, recently semi-acquired, I guess, by Meta, companies like Surge as well, which has been the news a lot.
11:22Essentially, having annotators give this thumbs up, thumbs down, or different forms of feedback. Obviously, human data is only so much scalable, you know, like at some point having armies of people. annotating data is not an infinite source or something that really is desirable on getting the model to be more competent. So people have started looking for years into ways to get better signals. So one of them was what we just mentioned, verifiable rewards. So some people call this RLVF or you see a lot in the literature RLEF, so from execution feedback, because you are executing what the model is producing, testing the result in an environment, looking at that result and being like, okay, based on that, I'm giving a reward or not.
12:04And by the way, this execution feedback, if you think about it, is if you go back to the roots of reinforcement learning, you know, when people used to do AlphaGo or, you know, or even before the Atari games, this is essentially execution feedback. You know, the model plays a game. If it gets, you know, if it succeeds at the game, you know, then it gets a reward. So it's a much more classical setting, actually, if you think about it in some way. That's the second category, so all of this verifiable reward or execution feedback. But obviously, not everything is verifiable. Actually, a lot of tasks that we do with the models are not necessarily, strictly speaking, verifiable.
12:40If you think... I think one of the reasons why Atari was such a great place to start was because of how verifiable it was. You had point scores that you're trying to optimize.
12:47Jon Krohn:That's a very clear reward function. A game is very obvious. A game like Go, it's very obvious at the end. If you win, you lose, or if it's a tie. But there are many tasks that are not like this. maybe like it could be writing a report or it could be like pretty much any natural a lot of natural language tasks and this is where i think in terms of scalability and access that has been really really successful as well is rlif where you use ai feedback so feedback from another model and it's kind of this this observation which which in insight i think is it's very funny because when in insight it's very obvious like you know all of these things when you look at them in insight you're like oh it's obvious but you know when when they were getting started it's like wow it's magic that it works at all, which is to use another model to review the output.
13:29So for instance, saying, oh, you know, like, let's say that you are doing summarization, very basic task, but is a summary that's been generated factual? You know, does it stick to the fact of the original text? Does it, you know, is it formatted in the right way? Is it like all of this, you know, kind of plurality of things you would want to see out of your summary? And what's really interesting is that this is obviously very scalable because this comes from another model and you can run models, you know, infinitely as many times as you want. And so right now, this sort of AI feedback, which some people also bundle up into the idea of synthetic data, you know, of training based on data that is produced by other models, also is seeing a very big growth and very big success because it works so well.
14:14And it's obviously very like, it's a great way to scale beyond just having the thumbs up, thumbs down from expert annotators to potentially reserving the human for the much more expert stuff, and then kind of having a baseline from other models, other models which might be specialized, and also having the verifiable rewards for our task that can be verified. Big mix of everything.
14:39Jon Krohn:About a year ago, I did experiments internally at a company that I worked at where we had had a very specialized task and it was enormously painful for humans experts to review that. It took them so long and they just, they expressed real disdain for having to do this task because it was so challenging. And so we thought, well, what if we could use at that time GPT-4 instead of the humans? And so we needed the humans to do enough that we had a kind of a sample that we could compare and GPT-4, you know, they were comparable. It was the same quality results, indistinguishable. And so we were like, perfect.
15:17Jon Krohn:This means we can now scale up to as many samples as we want. Yeah, yeah, totally. And I think this is actually very interesting that you mentioned this species comparison with human, is that a lot of people have pushed back on synthetic, well, I would say synthetic data as a whole, but on data that is model-generated because they are like, oh, this is going to be degenerate data. This is going to fall down, collapse into something that's bad. And that's possible. You can do this sort of data the wrong way and you can completely mess it up as always. But in general, it actually works really well.
15:48And I think, you know, people have this idea of human data as being very perfect. But actually, if you look at the data that comes out of the typical annotation contract, and I won't cite any, but it's actually not necessarily the quality that you think it is. It takes a lot of review to get it right. There's a lot of issues and there are plenty of studies on what people call inter-rater agreement rate, which is how much, you know, if you submit to two different annotators, how much they agree, you know, in their rating, you know, maybe if it's a rating on like a Leikert scale from one to seven, or if it's just a thumbs up, thumbs down.
16:19And the numbers obviously are very task dependent, but when you see them, they are actually crazy. Like actually, it's a lot of noise. There is a massive amount of noise. And when you measure actually the same sort of agreement rate with models or between models and humans, you see actually numbers that line up where essentially the quality that comes out of model is as good as what comes out of senators obviously not true of every task it's our task if the like if the judge model is completely incapable uh judge model is obviously not gonna not gonna be not gonna be good at this um but there is also like there is another side to this coin which is that verification is much easier than generation so it's much easier you know for a model a posteriori to come and to check a result than it is to produce it and that's something that's very powerful and that's probably one of the foundation of why this works so well.
17:33Jon Krohn:integrated dell and nvidia capabilities accelerate your ai-powered use cases integrate your data and workflows and enable you to design your own ai journey for repeatable scalable outcomes visit www.dell.com super data science to learn more that's dell.com super data science nice and now this next question is going to get relatively technical but our technical episodes are some of our most popular ones so we're going to dig into it a little bit here. And then after that, we'll get back to kind of more applications. We'll talk about your company. We'll talk about adaptive. But really quickly, I want to get into something really technical here.
18:09Jon Krohn:So, you know, these different kinds of approaches you talked about, RLHF, where we have the human giving a thumbs up, thumbs down. RLEF, where there's, you know, this execution feedback, where there's something, you know, kind of innate about what we're evaluating, like an Atari top score that we're trying to reach for. Or RLAIF, which we talked about most recently, where you're using AI models to kind of give you a thumbs up, thumbs down with an AI system. Regardless of which kind of those approaches we choose, there's also differences in what reinforcement learning algorithm we select, right?
18:40Jon Krohn:So there's things like PPO, A2C. Do you want to tell us about the big ones there? Yeah, yeah, totally. So obviously, RL, you know, HF, IEF, EF is essentially changing the data on which you are training, but you could also change a method. You know, we keep talking about this RL, but what is, you know, reinforcement learning. On a very fundamental basis, what makes reinforcement learning so different, it's not really a spectrum like everything. From SFT to reinforcement learning, you can build step by step and there's really a spectrum of things. The moment at which it exactly becomes reinforcement learning might be - SFT, supervised fine-tuning.
19:13Supervised fine-tuning, yes, totally. It's like the moment at which the transition might be a bit of a question of where everyone puts it, but But the very, very general idea, I think the key components to reinforcement learning, and then we can go into different methods. I think a big first thing is that reinforcement learning typically will be online. There is a difference in literature between online or offline RL. This is actually one of the very big, let's say, theoretical, and I put this in quotes, for as much theory as it can be in machine learning, especially concerning LLMs. but like historically, you know, people have argued a lot about what is offline online.
19:52And what does it mean to be offline or online? Well, essentially online means that you are learning based on the sample you just produced. So let's say I have a set of weights of my model. I make an inference, you know, I get an answer to the question. I evaluate that answer, like saying, oh, this is good or bad in whatever of the three ways we mentioned before. And then I use that in the training process to say, okay, so now I update my weights based, you know, on that. thumbs up, thumbs down. But I do it with fresh data, with data that has just come off the press. And then I repeat that process.
20:24I update the weight of the model, and then I get a new sample. Now, what if I accumulate sample, and then I train the model a few times, and then I start collecting samples again? As soon as I do my first step of training, I'm online. But as soon as I do the second one, I'm not online anymore, because the data I've generated doesn't come from the same set of weight, it comes from a set of weight that existed before and that hasn't actually produced the final output. Obviously, this is a bit of a ship of CZ kind of thing where one step might be okay, like might still be more or less the same thing, but two step, is it really the same model?
20:58Three step, five step, ten steps. And then you veer into offline ARL. But one of the big success of reinforcement learning is that it is mostly online, but this mostly, like proximally online, where essentially you have samples that are relatively fresh, you evaluate the samples, and you learn from that. And this means something, right now when I say this, it sounds very abstract, but actually there is a very, I think, easy analogy to say to this, is that the samples come from the model, and the model gives, you know, like a suggestion of what it can do, and you tell it if it's good or not. If you think, we make a parallel with human learning, and let's say I'm teaching you course about general relativity and I'm teaching you something about spinning black holes, kerometrics, that sort of stuff.
21:42If I show you an exercise to do and I could show you the solution of the exercise, have you memorize it, just memorize it again and again and again, and then when I present you the exercise, you can run through it exactly the same again. This is essentially what pre-training or supervised fine-tuning do, where you are presenting to the model maybe once, maybe twice, maybe thrice, the same samples, and the model eventually learns from it. Obviously, pre-training still generalizes because pre-training is very diverse. In pre-training, I don't just show you one problem, I show you all of the problems that can exist and that expectation.
22:18But in post-training, if I just show you one exercise and now I tweak something in the exercise, now I say, oh, now actually the black hole is carrying a charge. and so now you are in a completely different setting, you have no idea what to do. You are going to reproduce your answer, it's going to be bad. If we were doing reinforcement learning, the way that it would work is that you would try to do the exercise and then as a teacher, I would correct it and I would tell you, okay, this is good, this is not good, kind of that iterative process. And this is really fundamentally, I think a good mental framework for the difference between supervised training and reinforcement learning, which is, and that comes to this online-ness, which is that in reinforcement learning, the samples come from the model itself.
Read the full transcript
23:00So it's always in distribution for the model. It comes within what it's capable to do. And then slowly you are shaping that distribution away towards what you want it to be. Whereas in supervised fine tuning, you are kind of plopping down the new distribution, which might be very out of distribution. And you are hoping that, you know, as you show more and more samples that are diverse in us, you are going to, you know, widen the distribution and hopefully, you know, connect it back to the original knowledge so that there is no gap in between. Because if you ask a question that's in between what the model used to know and what you have taught the model, well, there is no guarantee, you know, that you fall, that in between it's covered, that's something it has learned.
23:39So it's kind of like the difference between the two. And I think this is a very big aspect of reinforcement learning is that, like, that onlineness, that learning from trial and error, that part actually touches to the second point, which is something you can somewhat simulate with supervised fine tuning by filtering the data, but it's also that reinforcement learning learns from a much wider range of signals. So instead of learning from an explicit, oh, you need to imitate that, it's about, oh, this was good, this was bad, this was maybe okay-ish, you know, like there's kind of a subtlety to this.
24:09You can bring this as well a bit to SFT in some ways, but I think it's a big difference that learning from a reward essentially from positive and negative things.
24:18Jon Krohn:I love this example that you just gave talking about, you know, teaching me general relativity and how the supervised fine-tuning is kind of like memorizing a solution and the reinforcement learning is this online way of learning where as I'm producing my output, you're providing me feedback and nudging me in the right direction. Totally. And I think it's a very good analogy because I actually think it's quite true to what happens. There is a caveat, like, you know, obviously an adversarial argument to this, and I touched a little bit on it, but I think it's worth double clicking. It's like, oh, but pre-training works.
24:49You know, very clearly pre-training works and teaching the model many things. So why it is, why it's a question of scale is that in pre-training, you are showing not just one exercise, but all exercises that are possible. In post-training, often, I mean, you can somewhat afford to do this. We discussed this before. Post-training is becoming wider and wider. But when you are specializing a model, you want to be as effective as possible, you know, with this to learn as much as possible from every sample that you have, and you might not have the luxury, you know, of every case that is possible that you have in pre-training.
25:21So, there is kind of like a slightly different regime. I put a bit of a star on this because now people run post-training at a much larger scale and it works as well, it does its benefits, so this is a bit of a caveat here. But the fundamental idea of much more generalization from reinforcement learning because of this onlineness, because of this trial and error and all of that, I think is very fundamental and actually, you know, when thinking about reinforcement learning research, I think it's one of the big things to think about. And to go back to your general question on the different algorithms, so we hear a lot like, HCC is an older one, but we hear a lot these days, for instance, about PPO, GRPO, GPO, all of this sort of stuff.
25:58I think one of them, they are actually quite similar. The answer is, especially PPO and GRPO, there was a big debate in the community. PPO and GRPO in terms of these components on onlineness and everything share very similar characteristics. What they do differently is more in the question of how then do you attribute, you know, you have a reward. So I tell you, you know, I tell you, oh, you passed the exercise or you failed the exercise. And there's a question of how do you attribute this to individual steps in the exercise or to like individual parts of the messages. And so PPO does this through something that's called like a model that calculates an advantage, you know so this was a value model uh that is gonna try to go back from a reward which is sparse you know in the level of the tokens like you have some tokens are rewarded but not others very often might be just a final token but sometimes it might be a bit more dense uh to a reward that is to to a value that's like at each token this contributed that much or that so in ppo you are training a model to do this literally you are training a large language model uh to do this task, which is an interesting view.
27:03In GRPO, it's a bit different. You essentially do an average over multiple rollouts, but fundamentally, this is just a different way to attribute the blame. You know, the fundamental of the methods are still very, very, very, very similar and share a lot of similar ideas.
27:20Jon Krohn:Nice. Very cool. So with that kind of context, that kind of foundation in mind now, including getting into the detail a bit of algorithms like PPO, GRPO, and A2C. Let's talk about Adaptive. So I mentioned right at the top of the episode something called Adaptive Engine, the flywheel for enterprise AI, which is continuously evaluating, tuning, and serving LLMs uniquely adapted to a business. Tell us more about that. Yeah, so I think at Adaptive, all motivation has been that reinforcement learning is amazing. All of these methods are really amazing. And this is even more obvious nowadays, I would say, in the past few months.
27:59But when we got started a year and a half, two years ago, I think it was still true if you had that experience of going, you were like, oh, these are amazing methods. They can do amazing things. And there's clearly a lot of potential in them. But there is a bit of a problem, which is that typically they are quite difficult to put in place.
28:16You might remember from reinforcement learning days, AlphaGo or Atari that we mentioned before. And you might also remember that this was very challenging to get right. There was a lot of research on it. Very often a lot if you have studied ML in a bit more formal setting at school, RL often seems a bit opaque. And also when you have experimented with it, very often it depends on the seed, like on the initialization and all of that. It's not as bad with LLMs, but with LLMs, the difficulty is more on the engineering. Because firstly, you are going to be blending inference and training. Because as we mentioned before, we are teaching the model based on something that it produced.
28:51So we are going to have some time to do rollouts, so to do predictions, and then rate this prediction, and then use that as training. So there is not as much as before, you know, this dichotomy between, oh, I serve my LLM to millions of users and I train my LLM on this cluster. Now there is a bit more of a combined, you know, of the two, which poses an engineering challenge, obviously. The other aspect of this is that these are complex pipelines. So typically, you know, we mentioned HF, AIF, VF or EF, depending on how you want to call it. So this means that during training, the model is going to have to interact, maybe not with humans, because maybe this is something that you will put offline and train a reward model, but it will have to interact with other models, maybe two, three, five, ten of them.
29:35Like today we run pipelines which have like five to ten AI judges in them and it works perfectly fine. But also environments. So maybe you know, So you are going to teach your model to do text to SQL. If you do this, well, you have to run the queries on the SQL database to get the answer to be able to do execution feedback. Maybe you teach a model to do Rust. And so you need to have like a Rust compiler. And maybe you need, you are teaching the model tool use. And so you need access to these tools. Or maybe you are teaching the model computer use, in which case you need a VM box, you know, you need like a virtual machine.
30:05You need a box where the OS is running and where the model can go through, you know, oh, I double click on PowerPoint. I open this. So you get all of these things. and now suddenly engineering becomes a nightmare. One of the reasons pre-training scales so fast is because pre-training is very simple. It's very straightforward. You have this huge batch of tokens. You just predict just the logits. You don't even actually sample. So you just predict the logits. You compare them. You compare the top one to what was in the text, and that's it. So it's very, very easy to make it run at impossibly large scale because fundamentally, yes, there are engineering challenges to distributed computing.
30:42obviously, but fundamentally the algorithm is very simple, very limited in interaction with the external world. I put this in quotes. Whereas with reinforcement learning, now you have all of these environments, all of these other models that you need to interact with. So motivation at Adaptive is actually to make all of this easy. We think that reinforcement learning is the way to get the best performance out of a given model for a specific task. If you think about it from a Pareto frontier point of view, reinforcement learning will always get you the best cost to performance compromise. Always.
31:14It's like a new part of frontier. So, obviously this is very attractive for enterprise adopting AI because either they want a cheaper model, you know, same level of performance, but they want something that runs as efficiently as possible, or maybe they want something that's not possible now and so they want more performance. Reinforcement learning in both cases is the answer to get there. But the question is, doing this reinforcement learning? And so this is what we do at Adaptive. We provide essentially data science teams with what we call the RLOps tooling, you know, kind of for them. Reinforcement learning ops, RLOps.
31:47Exactly, exactly. So tooling that they need to make this super easy and so that they don't have to worry about all of this distribution, this interaction, but they can just focus on the logic. So they can just focus on like, oh, I want this judge that does X, Y, Z, you know, maybe there is three, four, five, 10, 15 of them. On top of that, I also want an environment in which something gets checked. I want this, I want X. put all of this together, some synthetic data generation as well, a lot of lack of these things. And then you don't have to worry about any of the actual implementation. All tooling essentially does kind of, I don't want to say compile because that's exactly compilation, but essentially interprets your instruction, your Python recipe, and then runs it on the cluster in a distributed way without you having to think about it.
32:29That's fundamentally what we do.
32:32Jon Krohn:Hey, hey, this is your host, John Crone. I'm excited to announce that I've launched my own AI consultancy, a firm called Y-Carrot. Yes, the letter Y and the deliciously crunchy veggie. At Y-Carrot, we combine decades of experience in machine learning and software development with internationally recognized expertise in all the cutting edge approaches, including Gen AI, multi-agent systems, and RAG. From problem scoping and proof of concept through to high volume production deployments, we can do it all. To learn more, head to ycarat.com. From there, you can click partner with us to tell us exactly how we can help.
33:09Jon Krohn:Again, that's ycarat, Y-C-A-R-R-O-T.com. So I guess the target, correct me if I'm wrong, but it sounds like on what you're saying so far, the current adaptive platform is designed for people like software developers, ML engineers, who want to have reinforcement learning be easier. So your target audience is probably a lot like my listeners in general, where they're people who are writing, say, Python code. Yeah, totally, absolutely. And, you know, most of our users are data scientists who write Python codes to interface with the system. That's totally the target. So as you know, if we get into the details of the business, it's always obviously a bit more nuanced than this, you know, on like we also sometimes work, you know, with companies that have much less technical expertise, where they don't have a data science team, they still want to achieve something.
34:00So maybe either we will do some of the work or we have partners like Deloitte that can do some of that work. So there is always a bit more complexity business-wise, but yeah, fundamentally our core idea is to build better tooling for reinforcement learning so that it's easier to get to value and obviously it's something great for data scientists.
34:19Jon Krohn:Nice. And so if we have a listener today that wants to get started with Adaptive, what is that journey like? Yeah, yeah. So we are still very enterprise focused. I think when we started our company, which is a bit over a year and a half ago now, one of our theses was to be very focused on enterprise, on like larger enterprise, because these companies, when they put Genera into production, they have a very unique scale, you know, maybe millions, tens, hundreds of millions of users. And that comes, you know, with obviously costs, you know, that are much larger, so bigger impetus to potentially optimize costs to pack into smaller models.
34:52but although that means that you have much more interactions with the model and we say more interactions means more data points for post training so part of our original thesis was to focus on this sort of company which means that right now you know a lot of our deployments are kind of deploying in our customer infrastructure and we don't really have a cloud available yet but you can come in with your mom's credit card and just get started well not your mom's credit card or your company credit card but but this is something that is coming very soon where we want to make this more available. Like we think also now our tooling is a lot more mature and could be put into more.
35:23And so the short answer is that at the moment, we don't have an immediate general availability. I got you.
35:29Jon Krohn:So if somebody wants to take advantage of being able to do reinforcement learning more easily, they reach out to the sales team from the adaptive website. Right now we are very enterprise focused, which means essentially reach out to a human. But we are also excited to change that actually in the future and to make our technology more broadly available because we think it is at a point where that can be the case and where everyone should be able to run this sort of stuff, even for OB projects, actually. I think there is, like, even for this, you know, like, try something. Maybe it's really good.
36:00Maybe you're able to build, you know, a model that, like, ends up being much better than what currently exists, and maybe that can be the next big startup that you create, you know, that you get started. It's part of the tooling, you know, for this sort of stuff. I think one of the reasons we are expanding now to this is, And, you know, I mentioned something about original thesis was, oh, you need all of these human data points. Because when we got started, it wasn't as obvious that IEF and synthetic data would work so well. I think it was definitely something we had in our roadmap. And if you actually go back to our fundraising pitch, we had it like as an end of first year kind of thing.
36:35But what has really positively surprised us is how well it works. like how much leverage does synthetic data give you to go from almost nothing to a lot. And this kind of removes, there's still value in having all of these production data points, like they have tremendous value, you can do a lot with them, but they are not strictly necessary. You can get started, you can get bootstrapped from much less, and you can take a very small model, 8 billion parameter model, to be to the frontier performance on a task, mostly entirely with synthetic data, which is quite incredible and very easy to do. So this is why now we are also thinking of widening availability in a way, because actually these synthetic data pipelines, anyone can run them, anyone can define them.
37:19You don't need all of these users already to take the benefits of that.
37:23Jon Krohn:Yeah, we had the same example that I was giving earlier, where we evaluated, we tested the inter-rater reliability, like you mentioned earlier, between the AI model evaluation and in the laborious, tedious human evaluation. That was actually for the purpose of what you just said, which was fine tuning an 8 billion parameter model, one that can fit on a single relatively inexpensive GPU and get a frontier model performance on just a small set of tasks. Yeah, using these kinds of approaches. Yeah, totally. It works amazingly well. I think the boom in synthetic data, like the success of synthetic data.
38:05And when I say synthetic data, by the way, it's a very broad world, which means many things. And it means, to me, it means like, first, problem generation sometimes. So it means like creating new sample, creating new scenarios, you know, maybe new scenarios of conversation, self-play, you know. So maybe simulating a user, you know, like having a model stand, kind of standing as a user to kind of drive a conversation for self-play where you have like, that's, you know, first category. But it also means all of the AIF components of giving feedback, of reviewing some of these things. It's quite broad, but I think synthetic data has been widely successful.
38:39We have a company we worked with, which I can't name, but essentially they were building a chatbot of quite a general use case. And the chatbot, they wanted it to have certain traits, certain psychological traits and that sort of stuff. And to do this, we ended up mostly using synthetic data And we start from something like, I think it's about 70, 80, maybe human annotated samples. And when I say annotated in this case, actually, I don't mean thumbs down. This is something that you find a lot in this reinforcement learning pipeline. We use critics and rewrites, where essentially a human comes in, looks at something that the model has produced, and writes a critique and a new version of it.
39:26So literally in natural language. This is really cool, by the way, because I think when you ask someone, especially someone skilled, like think of a lawyer, think of a psychologist, you know, someone, when you ask them to give thumbs up, thumbs down on like thousands of samples, I think it's very, they don't like it. I think they're, you know, they feel a bit like in the meat factory. They feel like their work is being, is being tailorized. Like I think they don't, you know, they don't like it at all. But when you ask them to like give feedback about something or write something, actually they really enjoy it.
39:58It's really funny. They feel more, I think, engaged in the process and they feel more in control. So anyway, we collected 70 of these annotations in a bunch of different contexts. And from these 70, we are able to generate something like 80 ,000 synthetic conversations through self-playing. About 80 ,000 synthetic conversations. Each of these conversations is made of about 10 terms. So you are looking at nearly a million message. And during the reinforcement learning process itself, we explore multiple possibilities. So for each of these messages, we might explore five, ten possible answers which gets rewarded by judges.
40:40So at the end, you are looking from less than 100 human samples at something like nearly 10 million data points from which you can learn. So obviously, a massive multiplying effect from these synthetic data pipelines.
40:52Jon Krohn:Okay, so that's been a fascinating journey that you've had us on talking about lots of the reasons why Adaptive makes things easier for us and how we can have more powerful models, have smaller models be able to do things that Frontier models might otherwise only be capable of. You've described previously Adaptive as a layer on top of foundation models to tailor them to final use cases. Tell us about this being a layer on top. Yeah, so the one thing I want to be very clear that we don't do is starting from scratch. Like, I think, you know, specialized model in the sense of starting from something from scratch, I think, you know, there are use cases where it might make sense, but I think for a vast majority, it doesn't because the reality is that the foundation models, you know, are this amazing engine, this treasure trove, you know, of knowledge and of understanding, which you can sharpen into exactly, you know, what you want.
41:48And I think, you know, for us, what's really important is that why we say we are layer on top is because we start from open source models. So this might be Lama, Kwen, Kimi, whichever one is your favorite flavor and whichever one you are allowed to use at work. We start from them and we tune them to perform better. But this is only possible because the base model is already amazing, actually. And if the base model is not good, you don't really get anywhere. And this is, you know, back to the example, you know, much earlier we chatted about reinforcement learning used to be even harder, you know, with like initial issue at initialization with seed and that sort of stuff.
42:21And part of the reason is because this was reinforcement learning from scratch. And when you are doing reinforcement learning from scratch, behavior initially is fundamentally unstable because you are asking a random policy, like a random model to take decisions. But obviously the decisions are random, which is a disaster. was when you start with a large language model, you are starting on easy mode because the model is already incredibly smart. Like one thing to note about pre-training is, we said, oh, the pre-training model is not chatty. And by the way, something I would invite people to do, it's harder these days because as I mentioned, the lines between pre-training and post-training are blurred, but is to look at some of these older pre-trained only models.
42:59I think some of the early LAMA might still be available this way, but you can also look at models like GPT-J, like that were some of the very big, like some of the first, you know, very big success of open source models that were just pre-trained and you can try to interact with them. Obviously, this might come from another generation, much less compute spent, but still you will see it's very different. But anyway, despite that, these models are still amazing. You know, compared to a random starting point, they're still amazing. They still contain insane knowledge. Like if you think about it, in pre-training, they've seen nearly everything.
43:30Like the knowledge that's contained in this model is insane. So it's mostly a matter of disentangling that knowledge to an extent, adding some of it as well. I think there's always a debate like is post-training actually adding knowledge to the model or not at all. I think it does as well in certain conditions. But essentially of disentangling the knowledge that's in the model, maybe pruning the part that you don't need as much, surfacing the part that you need the most and using that as scaffolding to learn even more, to acquire even more capabilities. But this is possible because we start from an amazing open model.
44:05And that's kind of the sense that we are layer on top, you know, is that we start from this base open source model and we take them to even better performance.
44:14Jon Krohn:Nice. Great example there. It makes it crystal clear. And it's interesting how, you know, you have all this background in developing frontier models, pre-training, post-training, And, you know, now you've found this niche allowing other people to leverage that kind of background that you already have and be able to accelerate their own use case development, particularly at the RL stage. Yeah, yeah, yeah, yeah, totally. Yeah, yeah. Yeah, I think for us, you know, like a lot of what we do now is bringing, you know, this expertise in reinforcement learning to be something that anyone can do. Like, you know, our view is that reinforcement learning to an extent is still a bit of a frontier subject.
44:51it's still a little bit of something that you know only a more maybe more experienced audience gets to experience but i think it shouldn't be the case like the reality is that these are exceptionally powerful method that should be in the hands you know of every data scientist of everyone must like you know prompting you know is in the hands of everyone i think you being able to beat these pipelines to leverage them should be in the hands of everyone to to build something really cool so that's you know ultimately that's really our goal is to spread you know this And we see it, you know, one thing that I find personally very exciting, you know, when we work with customers is that, you know, obviously very often on the first use case, we work very closely with them because we teach their teams how to use the tool, how to think differently as well.
45:30Because even teams that have experienced with supervised fine tuning, I think reinforcement learning asks you to think differently. You know, you don't think so much about the data that's going to be the explicit demonstration, but you think more about measurement of success. So you think more about what defines success and how do I measure it? So that might be something that's verifiable, that might be with an AI judge, and then you use that to tune the model. So it's a bit of a different way, I think, to think. But yeah, typically we start, we work very closely with them on subject, and then their teams kind of take it on.
46:00And we have a customer that can name publicly because they are happy with that. We work a lot with AT &T in the US. And at AT &T, their teams now are using the tool autonomously, And it's always amazing when in the sessions that we hold with them, they come and they're like, oh, I've seen this and you see the scores and it's really good. And you're like, oh, it's an amazing new tool. I think it's really fun for data scientists to be able to have these new capabilities to do more.
46:24Jon Krohn:Nice that the teams there at AT &T are all grown up now on your training. They are really good, actually. And I don't just say this because I can talk about them publicly. I think one of the things that's been really cool in working with AT &T is the maturity in terms of bringing Gen.AI to use cases. And whenever we have meeting with them, I'm always surprised by the penetration of Generative AI inside the organization and everywhere. In every aspect of the business, they are pushing models to do really amazing things that really create value for the company. And so I think it's really cool to see that because often there's always that discussion of oh is ai a bubble is blah blah blah you know grumpy people and sometimes you know you might be like is it like whatever and i think it's definitely not like i think you know yes some business are slower in adoption the real world is always slower in adoption um but there is like in companies that are that are moving forward there's insane value being created on this podcast i'm always going on about how claude code is mind-blowing but now claude Cloud Cowork is making my jaw drop as well.
47:27Jon Krohn:For example, I recently wanted to quantify how healthy my sales pipeline is for my AI consulting business. I simply asked Claude to estimate my sales for the coming quarter, and it brought info from relevant Google Sheets and my Gmail to create a professional spreadsheet of clients with estimated revenue for each one. Whoa, this might have taken me a day. Instead, it was done flawlessly with Claude Cowork in minutes. Claude is the AI for minds that don't stop at good enough. It's the collaborator that actually understands your entire workflow and thinks with you. Whether you're debugging code at midnight or strategizing your next business move, Claude extends your thinking to tackle the problems that matter.
48:02Jon Krohn:Ah, and you'll appreciate that I can ask Cowork to show me data, such as my sales spreadsheet, and it provides an interactive chart right in the conversation. For problems worth solving, get started with Claude at Claude.ai slash superdata. That's Claude.ai slash superdata. And check out Claude Pro, which includes access to all of the features mentioned in today's episode. cloud.ai slash superdata. Nice. You mentioned there's something that I want to highlight a little bit that these amazing teams at AT &T are doing is they're getting that definition of the reward function, right? That sounds like it's one of the hardest parts, right?
48:35Jon Krohn:So you're getting people's mindsets on the reinforcement learning cycle. And if you don't define that reward function right, your model isn't going to end up doing in production what you hoped it would. Yeah, totally. So I think this is the part where it gets in a different way to think about this problem is that in reinforcement learning, fundamentally what you are thinking about is what defines success. Like what, you know, how do I define a successful outcome or a bad outcome for the model? And how do I provide the model a signal about this? And that's really what it becomes all about. We, like, there's this quote that I really like for reinforcement to describe reinforcement learning, which is that if you can measure it, you can optimize it.
49:13And this is literally true, actually. This sounds very cheesy, but it's actually literally true in the case of reinforcement learning, which is that as soon as you can measure something, you can use it as a reward to optimize it. And because these methods are so powerful, because these models that we are using are so smart, even if the signal is noisy, even if the signal is kind of removed, you know, like quite complex and all that, the models are going to find a way, like the system is going to find a way to optimize for it. And that's uniquely powerful. So then it becomes entirely a game of like, how do I define it?
49:43And this is a part where reinforcement learning becomes more of a like, I like to describe it as a pipeline or as like as a system. Because it's not very often success is multifaceted. You know, there is not like just one criteria. So it might be, you know, like somebody needs to behave in this way, needs to follow these policies, it needs to achieve that X, Y. And then it's about finding the signals or even, you know, in multi-agentic system, finding like the signals for individual agents of like, okay, this defines success as this step. I can check it. I can evaluate it maybe with another model, that's a reward, that's good.
50:16Then I move to the next step and kind of building these things and also being able to do it end-to-end as well where ultimately there might be an overall success. Yeah.
50:24Jon Krohn:Yeah, so it's clear that you have a ton of experience. You and the Adaptive team have a ton of experience with getting real-world use cases spun up, particularly leveraging the RLOps that you guys specialize in at Adaptive. I now have a long question. There's a lot of context here. Okay, okay. So I hope you have a big context window, as well as our listeners. With one million, though, Nia Yoba. And because I'm going to dig into a bit about your past prior to what you're doing at Adaptive, but then I'm going to use that to talk about how we can be preparing for the future. So you previously worked as an extreme-scale team lead at the AI community Hugging Face, and prior to that at the Gen.AI platform LightOn.
51:07Jon Krohn:We'll talk about, I have another question about LightOn coming up soon. And there's a big question mark around scale these days where, you know, the idea of bigger compute, bigger networks, bigger data driving more model capability. One of those things, you know, more compute, okay, we can just have, you talked about hundreds of thousands of GPUs. You can have a million theoretically. You know, it's an engineering problem. Same thing with bigger networks. You know, we can have more model weights or we can have more clever mixture of experts models. But bigger data can be tricky because you already talked about earlier in this episode how, you know, the pre-training can involve all the literature that's ever existed, all of the internet.
51:47Jon Krohn:So that can theoretically run into short supply. And so different people have different opinions. Ilya Sutskiver said that if Gen.AI's fossil fuel is human data on the open internet, we've exhausted our supply. However, other people like Sam Altman, Dario Amadei, Satya Nadella from OpenAI, Anthropic, and Microsoft, respectively, they don't seem to think it's a problem, that scaling has no end in sight. Synthetic data seems to be part of the solution there. Yeah, do you think that, you know, kind of engineering tricks you've mentioned in past interviews, how things like cleaning up data to remove duplicates had a big impact?
52:24Jon Krohn:So do you think that this kind of massaging the data that we have can continue to give us great results going forward, regardless of whether we have more of it? I think, you know, I think there is a bit of truth in every one of these statements where definitely in terms of like readily available data, we are starting to hit a limit where, you know, there was a golden age where we are just starting to do it. Oh, starting to crawl the web and starting to improve your crawler. And then you add archive paper. And then like there was kind of like, you know, a golden age where data seem unlimited. And obviously, no one of this is the case.
52:54You can massage this data to improve its quality, to get better results out of what you get. You can order it differently during pre-training, maybe put the lower quality data first so that you get more impact from the later high quality data. But ultimately, at some point, we have only produced, as humans, we have only produced so many worlds. So there might be a question of this, of like, do we run out? I think, actually, I can't answer the question, have we run out now or when? I think it's quite complex. But I think I can answer the question of what's next and what's already actually the case, which is we go back to post-training, where what's very interesting about post-training is that post-training enables model, and I'm going to use the word of Richard Sutton because he put it in a very elegant way.
53:40It enables model to learn from experience. So the models actually do something, as we discussed before, get the feedback on that doing, and receive that. And this is infinitely scalable because this is essentially the experience of the model in the real world. Yes, currently we do this in simulators. We formulate these artificial problems, but it's already possible for models to conduct an experiment in the real world and use that as a reward signal for reinforcement learning. And this, the bits of data we can get from this, I mean, there is no limit. Like it's practically, you know, models can conduct as many trials and many experiments in the real world as resources alone.
54:20So I think this is a part, you know, if you look from a more like very large scaling perspective, like this is where reinforcement learning is very exciting for this, is that actually it's a gateway to a much wider range of signals, a much wider ability to learn. And we are seeing this now, you know, for a while people were seeing reinforcement learning more as like, oh, this only specialization layer or only, you know, like post-training layer. but actually it can be the bulk of the resources are going to be spent in the future on post-training, on reinforcement learning because the models are going to learn from trying again and again across billions of virtual environments and eventually also in the real world against trying their experiments, their own ideas and this is obviously a very big frontier right now and people are pushing really hard on it and it's really, really exciting.
55:13Jon Krohn:Nice, so that kind of covers the data problem. It sounds like we're good on that front. Let's talk a little bit about compute as well. And this gets into your experience of Lidon, which I think is interesting. So land use and power requirements of data centers are getting more and more ambitious. So Mark Zuckerberg recently announced several multi-gigawatt clusters, including a five gigawatt data center, which would cover more than three quarters of the area of Manhattan, where we're recording today. In a presentation a few years ago in 2021, you said that by this year, by 2025, hardware would become the bottleneck.
55:46Jon Krohn:And so you discussed how things like LightOn's photonic chips, you know, how these kinds of hardware, alternative hardware approaches can be the solution. So maybe neuromorphic chips or photonics. What do you think about this hardware problem? Yeah, so for context, I used to work, when I started, when I did my PhD, in France you can do an industrial PhD where you work with a company, a very good system, very surprising that it wasn't invented in the U.S. of all places, but this company Lighton, which now mostly does Gen AI, but used to develop a chip which worked with photons instead of electrons, so with light, essentially to do computation, to do certain type of computation.
56:24And this comes with a bunch of advantages, you know, such on power consumption, on parallelism, on things you can do. So it's alternative means of computation, sometimes related to neuromorphic, that sort of stuff. And yes, So today the bottleneck is compute. Definitely, I think this is very obvious given the money that people are spending towards trying to get more, given the insane valuation of Nvidia, which is gonna continue to increase. So the bottleneck is compute. Obviously we might ask, do we need a new compute padding? This is actually a subject on which I'm very bearish, personally, on which I've...
57:01And maybe this is because I got burnt once. And so I think it's really difficult to bring a new hardware paradigm to life. The current hardware paradigm has its issues. It uses a shitload of energy, blah, blah, blah. It's very rigid, but it also has tremendous advantages in that it works really well. Like you can implement this algorithm very effectively. It works really well. And there is a lot of money that goes into it. If you think about the latest chip from NVIDIA, if you think about the GB200 or the full rack, It's an entire rack. But if you think even about just the chip, it's probably the most complex object that has ever been built by humankind.
57:42Like when you hold, if you hold one in your hands, even an H100 or B200, you are probably holding the sum total of all of human achievement. Like all of human achievement has peaked to this thing, which is like absolutely insane in terms of engineering to get there. You know, at NVIDIA, it took a decade to build a generation of chip, you know, from ideation to implementing some of the R &D that is coming out of TSMC, ASML or others, to actually, you know, building the compilers for this chip, the first tape outs and all of this. So like over a decade of building it, this doesn't even account all of the R &D that goes behind, you know, into extreme UV lightography and all of this.
58:17Insane chain of technology. For a competitor to come, for, you know, an alternative mean of computing to come, well, you have to reproduce all of that. And I think this is going to take a while. Like, I think the reality is that this is really hard to get at. I think, you know, our current silicon-based paradigm or current paradigm of computing is really good. Is there room for improvement, for more specialized hardware, for that sort of stuff? Yes, and, you know, to an extent, the GPUs are already extensively specialized. They're not really GPUs anymore. You know, they're already extensively specialized to machine learning.
58:50And even some people say, oh, it seems to be specialized to transformer. But this is already happening. You know, if you look into the instruction sets that you have on these GPUs. There are operations, you know, that are increasingly becoming specialized to these. They are thinking about, oh, adding, you know, like some specific unit, for instance, you know, in the attention you have the softmax, which has an exponentiation phase. So it's our instruction set for this. Like they are thinking about, oh, can we increase a bit on the chip the part that is dedicated to this so that we get a bit more, you know, a bit more throughput with this.
59:19It's better. So there's already, all of this already goes into thinking at NVIDIA. So I think there's already that specialization motion is moving. I will bring another point of view to this, which is something I've been thinking about increasingly recently when chatting with friends of like, oh, you know, what do you think of like, should we change tokenization? Do we need, you know, photonic computing, do we need quantum computing for that? I think about it in terms of like, do we need it to get to AGI slash ASI? Like, is it something we are going to discover by ourself that we need to figure out by ourselves to get there?
59:50Or is it something that later we are going to figure out with the support of general intelligent or super intelligent systems? Because I think general intelligent systems probably already exist. Or is this something that we are going to figure it out with these systems? And my thinking on the subjects of like photonic chips or even quantum computing is that we as humans don't really need to worry about this right now. I think that we already have the capability, like what we have, the technologies that we that we have are already in us to take us to the level where we will build systems that will help us build this.
1:00:22I'm sure, you know, in a century, you know, I'm sure we will use photonic chip, I'm sure we will use quantum chips and all of that, but we will have built them with the help of what we are creating currently.
1:00:33Jon Krohn:Yeah, so basically to kind of summarize your big idea there, we can use these chips that were originally designed for graphics processing, and we can leverage those at huge scale to create an AI system so powerful that it helps us to crack all these other kinds of things. Exactly, that will help us do scientific research and everything. And, you know, a lot of the way that I think about these questions these days is like, what are the innovations that we still need to do to bootstrap this system that will then help us to get even more? And there is stuff left to do. You know, this is not a negative point of view.
1:01:04There is nothing left to do. No, there is stuff left to do. You know, I think, you know, we were thinking of reinforcement learning just before. there's plenty of stuff to do in that direction of how do we scale this, how do we enable models to experiment in the real world. Like at some point there's obviously a bottleneck in the real world. If you are a material scientist or a biologist, you conduct experiments in the real world. You don't just sit at your laptop all day. Models currently cannot do this. There is no way currently for a model to run biology culture in a scalable way. There is no way for a model to run...
1:01:35Jon Krohn:There's a tiny bit of prototyping in that space where you have, on relatively small scales, an AI system that can control a wet lab. Yeah, exactly. It's starting, yeah, exactly, with wet labs or even for material science. And it's starting, but we need, like now, people, I think, want to scale this because this is one of the next bottlenecks, is how do we enable models to run experiments in a scalable way in the real world? And it's super exciting. I find the first early experiments in this that you mentioned to be really key. I think, like, how do we scale this? How do we make a weight lab, a material science lab, or whatever else, something that's addressable to a model that can be easily reset, that can be easily experimented with in a safe way.
1:02:17I think these are really big challenges. So these are subjects that I think, for instance, we need, you know, still a lot of innovation and will be key.
1:02:26Jon Krohn:So it seems like you spent a lot of time thinking about, you know, these powerful AI systems. We might call them artificial general intelligence, if it's kind of at our level or above us, artificial superintelligence and helping us with these kinds of problems, handling material sciences problems, biological problems. And so this is something in my mind, I'd love to hear what you think about this. In my mind, it's always seemed to me like having an AI system, having this AGI kind of system, it isn't the singularity that we can't really see beyond that that unleashes, yes, things will be very different, but some things will still take a lot of time.
1:03:06Jon Krohn:It's not like instantly overnight cancer is solved because you have to run experiments on probably humans and tissues and other animals, and that could take decades. So you can have hunches the AI might be able to have insights by taking papers from all different kinds of fields and having insights that humans might not have had. But then we still need to run the experiments, and those could take a long time. Yeah, totally, yeah. I think, you know, one of the aspects is that I think the superintelligence that we'll create will at first be very spiky. You know, there will be domains where there will be disproportionately superintelligence compared to others.
1:03:45For instance, I think mathematics is a really good example of this, where we can build, like, formal verification system. We can build all of this. So, like, getting to mathematical superintelligence can happen in a box. you know, like literally can happen in a completely closed box. You could build a super, like a super intelligence in terms of mathematics, building a biology super intelligence. Some people have a different view of this. Some people think, you know, that like, you know, computer-driven biology, simulation and everything will be in us. But some others, you know, the state of the science at the moment is that you need experiments.
1:04:16And so we might build a system that is like beyond genius level at mathematics that can describe mathematics that's way beyond, you know, our ability to understand, but at the same time, if you ask it, you know, to do even the simplest, you know, of like medicine development might not be that amazing. Or the bottleneck might and might be really good at making up the plan at updating, you know, you want the plan once you give it the result and it's going to be like, okay, so now we should try this, blah, blah, blah, but still be bottleneck, blah, blah, blah. So totally, I think it's a very realistic future.
1:04:44I think something that's, that that's very clear. I think now is that the closer we get, like, I think it will definitely happen, you know, like, I think we'll definitely build within probably the next five years, that's my personal bet, but maybe even 10 years if you're a bit more bearish, we'll build super-inteligent systems. But this super-inteligent system will not be super-inteligent in everything out of the box. I think this will be very messy, actually. I think it will be a very messy time because in some domains, we will do more progress probably in the space of like a year than we have done in the space of all of the existence of our civilization, which will be astonishing, like discovery that we can barely imagine.
1:05:18And in some others, we will barely move. In some others, it will be like, oh, a new flu come around? Well, still have to do the work to come up with a vaccine for this year because the system, you know, doesn't do that yet. So I think it will be very messy. Like, I think it's one of the nuances that I would bring to the stories that you often read about, like fast takeoff and everything, is the messiness of it. And it's very hard to predict, you know, which direction is going to work so well, which is not. Maybe some of them, yes, indeed, simulation will work very well for some things, and maybe, you know, for some things, we'll be able to do tremendous progress just in a box, without going to experiments.
1:05:52And maybe in some other fields, we will desperately need the experiments to be able to make forward progress. I think it would be very, very unequal, very messy in many ways.
1:06:01Jon Krohn:Fascinating. And so it sounds like with me, this super intelligent system that we are careening towards in five to 10 years, in your view, it is largely a positive thing for humankind. This is a very complex topic. I'm personally, maybe I'm a more optimistic person. I personally think, yes, it's very positive. I think, you know, I view scientific progress, I view progress in general as, you know, as one of the main driver of what we do. Like, of, you know, I think it's a view that maybe not shared by everyone, but personally, I think it's some of the most, like, you know, beautiful achievement of humankind is progress and understanding of our universe.
1:06:38So I think not only, like, being able to create intelligence, to understand intelligence, to create it. Understanding might come after creating it, which is kind of funny, but, like, so being able to do this. is really beautiful. I think it's something amazing that we are doing. And I personally think it will be positive. I think there will be challenges. Will it create big societal change that might create unrest? Yes, that's very likely. But I think on a longer time scale of, I think the five, 10 years where it happens are going to be very messy for sure. But I think the time after that is going to be a time of probably the best time ever.
1:07:18I mean, it's always the case. I think the next year is always better than the previous one in human history, more or less, give or take a few accidents. But I think overall, the trend is always positive because I think our progress gives us more freedom and enables us to do more, to give us more freedom to have more independence. So I think that's going to be very positive. But obviously, this is not to say that there might not be some complexities along the way that there are problems that we need to solve. Obviously, there's a lot of stuff to figure out.
1:07:47Jon Krohn:articulately said. I couldn't agree more with everything that you said. We're on exactly the same page. You're preaching to the choir, as it were, at least with me and probably with a lot of our audience as well. Before I let you go, Julian, this has been a fascinating conversation. I know that you read a lot of sci-fi books. I think your book recommendation might be in that vein, which, you know, so it kind of gives us, we've just been talking kind of sci-fi a little bit in real life, like real life sci-fi. In a way, in a way, it's sometimes, you know, Actually, I had a reflection recently when reading sci-fi books, reading depictions of artificial intelligence, that they are actually, they fall short of the reality of what's happening.
1:08:24I think there are very few books that actually, where you read them and you are now faced with LLMs, with what we have. I'm like, oh, actually, what we have is better. Life surpassed fiction. Yeah, on the book recommendation thing, it's a book of people that have heard me before, will know, will say that I'm obsessed. It's a book I recommend a lot. I heard a lot of sci-fi and I personally have a fascination with alien contact. I think one of the most interesting subjects in sci-fi is the idea of alien contact, of contact with an intelligence that is different than ours. And I think based on our previous conversation, you might understand why.
1:09:02I think the idea of different forms of intelligence and how we might interface with them is a very fascinating topic. There's a really good book called Blind Sight from Peter Watts, which is essentially an alien contact story. And I won't spoil it, but humans actually in the book are very different. It's in the future. So humans themselves or intelligence have, and I say intelligences, plural, because they've evolved in different ways. But also the one of the alien is extraordinarily alien and kind of raise the question of, okay, how do you interface with that? How do you interact with that? And I find this to be a very fascinating topic.
1:09:37So, yeah, it's a fun book that I recommend and I will recommend it again. I think it's an amazing book.
1:09:42Jon Krohn:Yeah, actually, aliens came up in our research. So as usual, our researcher Serge Massis brought up way more topics than I could possibly cover, but it does help me kind of interview you, even the questions that we don't get to. But you've actually talked about aliens before in the context of pre-training, creating something like aliens of extraordinary intelligence yet little understanding. And then the reinforcement learning, the post-training that comes later, allows you to transform those aliens into helpful, grounded assistants. Yes. I think this is an idea that people use to frame under.
1:10:13It's a bit less popular this day, but you know the meme with the Shogot? The Shogos or the Shogot? I don't know how do you pronounce it.
1:10:19Jon Krohn:I don't know what word you're saying. It's like, I think it's from Lovecraft. So it's this weird creator. Oh, Cthulhu? No, no, no, no, no. It's a specific creator from the Shogos or Shogot. I don't remember how it's pronounced. But anyway, there's this idea. Like, I think people were comparing for a while models, you know, large language model with it. and there is a few memes on this of like alignment is just putting kind of like just a mask on a terrible creature on a very frightening and terrible creature and this kind of comes from that that after immediately after pre-training the models are very strange they know a lot about us obviously because we train them on everything we have ever done so how could they be so different from us but they interface with us in a very weird way and some of post-training is about auto-lining this.
1:11:09And yeah, that's an idea that comes from here.
1:11:12Jon Krohn:Yeah, yeah, yeah. And you had the word exactly right there. It was one I wasn't familiar with. It seems like it's related in the kind of Lovecraft universe, HP Lovecraft universe to Cthulhu in some way, but Shoggoth, S-H-O-G-G-O-T-H, I'll have links to images of them in the show notes. Yeah, yeah, definitely. I have plenty of very fun memes in machine learning about them. Nice. Fantastic. Julien, this has been amazing. For people who want to hear more of your brilliant thoughts after this episode, how can they follow you? Yeah, I'm in a very boring corporate way on LinkedIn, Julien Lundi, but otherwise on Twitter, I'm at Sleepy Lolo, which I think we can put a link instead of…
1:11:48Sleepy Hollow? Sleepy Lolo. Sleepy Lolo. S-L-I-P-P-Y-L-O-L-O, which is a very long backstory.
1:11:57Jon Krohn:Okay, yeah, we'll have that in the show notes. Exactly. I'm sure we'll arrange that. Thank you so much, Julien. It has been a treat to have you here. I learned so much. Thank you. Thank you very much.
1:12:09Jon Krohn:What an exceptional conversation with the brilliant Julian Lanet. In today's episode, Julian covered the evolution from pre-training, that's predicting next tokens on web-scale data, to post-training, that's reinforcement learning, as the dominant phase of LLM development. He talked about how AdaptiveML's platform makes reinforcement learning accessible to data scientists enabling companies like AT &T to autonomously tune smaller models to frontier performance. He went in detail on the three types of reinforcement learning feedback. That's RLHF from human feedback, like human thumbs up, thumbs down.
1:12:43Jon Krohn:RLAIF, where AI models evaluate performance. And RLEF, where we have verifiable rewards from code execution or game scores. Julian gave his prediction that we'll achieve superintelligence within 5 to 10 years, but that it will be messy and spiky, revolutionary in domains like mathematics, while still requiring real-world experiments for things like biology and medicine. And he talked about why current silicon-based computing is likely sufficient to bootstrap AGI, which will then help us to scale new computing paradigms like photonic and quantum computing technologies. As always, you can get all the show notes, including the transcript for this episode, the video recording, any materials mentioned on the show, the URLs for Julian's social media profiles, as well as my own at superdatascience.com slash 913.
1:13:36Jon Krohn:All right. Thanks to everyone on the Super Data Science podcast team, our podcast manager, Sonja Breivich, media editor, Mario Pombo, our partnerships team, which is Nathan Daly and Natalie Jaiske, our researcher, Serge Massis, writer, Dr. Zara Karche, and our founder, Kirill Aramanco. Thanks to all of them for producing another excellent episode for us today. For enabling that super team to create this free podcast for you, we are so grateful to our sponsors. You, listener, can support this show by checking out our sponsors' links, which are in the show notes. And if you're ever interested in sponsoring an episode yourself, you can find out how to do that at johnkrone.com slash podcast.
1:14:17Jon Krohn:Otherwise, you can support us by sharing the show with people who would enjoy the episode, reviewing the episode on your favorite podcasting platform, subscribing, obviously, if you're not already a subscriber, but most importantly, I just hope you'll listen to us. You'll keep on tuning in. I'm so grateful to have you listening and I hope I can continue to make episodes you love for years and years to come. Until next time, keep on rocking it out there. And I'm looking forward to enjoying another round of the Super Data Science Podcast with you very soon.
1:14:49Thank you.
From the publisher
Julien Launay launched Adaptive to give data science teams in business enterprises their “RLOps tooling” to make reinforcement learning easier. Talking to Jon Krohn, Julien says, “Most of our users are data scientists who write Python codes to interface with the system”. Adaptive is also able to work with companies without data science teams, collaborating with partners like Deloitte to add the necessary personnel. Julien is currently working on making his platform more widely available.
Additional materials: www.superdatascience.com/913
Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information.




