In short
The episode argues that AI’s future will be both open and closed: open-weight/open-source models will increasingly reach “watermarks” for many tasks via post-training and specialization, while closed models remain useful for proof-of-concept and for domains where open models haven’t yet hit required capability/safety thresholds. It also discusses Australia’s AI investment priorities and why open ecosystems matter for national and individual “control of destiny.”
Guests
Mudith Jayasekara and Charlie O’Neill, both Australians. They co-founded Past (acquired end of last year) and now work at Base 10’s research side after joining Base 10. Base 10 (SF; backed by Blackbird) runs AI infrastructure “plumbing” that helps AI-native companies like Cursor, Harvey, and Lovable put models into production. Their research focus includes constitutional alignment, interpretability, and keeping open-source AI competitive with frontier labs.
Key claims
Open vs closed gaps are shrinking; intelligence differences are increasingly about scale plus post-training/specialization. Online reinforcement learning from user feedback is the durable moat. Cyber/security concerns aren’t uniquely “China vs US”; they’re about baked-in values/censorship and evaluation. Australia should invest in compute plus a talent hub, since top Australian talent often leaves.
Notable examples
“God in a box” for ~$20/month; Hugging Face breach where frontier models rejected investigation but open models enabled progress; DeepSeek/GRPO as maturing RL stack; Cursor collecting user data every ~5 hours for online updates.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VODefining AI Models: Open Source vs. Closed
4:30 to 6:04
Understand the distinctions between open source, open weight, and closed models in AI.
“Charlie, Moody, it is so wonderful to have you guys on the podcast today.”
The Evolution of AI Model Training
6:04 to 8:46
Discover how AI models are trained and the impact of data and open source on development.
“So one labs don't necessarily want to release that because it really is just like the exact recipe that you use to train the model.”
Post-Training Dynamics and Feedback Loops
8:46 to 11:52
Learn about the importance of post-training and online reinforcement learning in AI.
“should be permissible to run Chinese open source models.”
Future Directions for AI Companies
11:52 to 14:01
Explore how companies can leverage AI to evolve into frontier labs themselves.
“I'd love to hear more about your customers.”
Evolution of Reinforcement Learning
14:01 to 16:48
Learn about recent advancements in reinforcement learning and their impact.
“And supervised fine-tuning is very, very brittle.”
Navigating Open vs Closed Source Models
16:49 to 19:59
Explore the ongoing dynamics between open-source and closed-source AI models.
“decides which model to use for a specific task?”
Consumer Impact of AI Evolution
20:00 to 22:20
Understand how advancements in AI will affect consumer interactions and choices.
“is like one leaps ahead and then others catch up.”
Geopolitical Concerns in AI Development
22:21 to 25:58
Delve into the geopolitical implications of AI models, particularly concerning security.
“Do you think at the level of the consumer that will eventually become something they don't have to decide themselves and the systems default to them fitting the best model?”
Evaluating Open Weight Models
25:59 to 28:00
Learn the key considerations when choosing open weight AI models for various tasks.
“I think the big one is honestly just like size.”
The Importance of Open-Source AI
28:00 to 29:24
Learn why developing an open-weight ecosystem in America matters despite the dominance of closed models.
“So, yeah, I think like that's something we see people often go too far and get too excited about the post-training rather than just using the base.”
Show all 36 chapters
Talent Distribution in AI Across Countries
29:24 to 31:25
Explore how talent is distributed globally and the potential of countries like Australia in AI.
“From a country perspective, I think it's really important that like, you know, if you have the resources capable of training these frontier models, then you have a say in what those values are.”
Investing in Australia's AI Future
31:25 to 34:14
Understand the necessary investments for Australia to foster its AI talent and technology.
“And that's really exciting to us because they've been doing it for like a year or a year and a half.”
When Companies Should Train Their Own Models
34:14 to 36:56
Discover the factors that lead companies to transition from using existing models to training their own.
“Like, I think the dominant thing you'll hear companies talking about is cost.”
Case Study: Building and Training Models for Clients
36:56 to 39:29
Learn about a practical example of how a company transitions from using models to training their own.
“Can you walk us through an example of a customer that you've worked with where you've gone from building a product with them to them eventually training their own model?”
Challenges in Reinforcement Learning Environments
39:29 to 42:00
Delve into the complexities of creating effective reinforcement learning environments for AI.
“What's the most difficult part about reproducing a company's learning loop?”
Understanding Model Divergence and Training
42:00 to 44:20
Learn about the differences in model training and divergence among companies.
“Like we believe and even know to an extent that like the way in which company A post-trains a model versus a company B post-training that model or even for the full stack, the full recipe from start to finish.”
Reinforcement Learning and Generalization Challenges
44:20 to 46:10
Explore the limitations of reinforcement learning in generalizing across tasks.
“But I guess like what we see is that app layers have such particular like opinionated takes on what their user journey actually looks like.”
The Evolving Landscape of AI Applications
46:10 to 49:50
Examine the competitive dynamics between model companies and application companies.
“I'm curious to hear, as the Frontier Labs do move deeper into areas like legal and finance and coding, how does that change the relationship between a model company and an application company?”
Sovereign Datasets and Customization
49:50 to 53:40
Discuss the future of sovereign datasets and their implications for proprietary models.
“The countries can inform and as these things get more and more commoditized and, you know, we're getting a slew of different open source models like out of the box that roughly all have the same pre-training base.”
Building Companies in an Uncertain AI Environment
53:40 to 55:40
Learn strategies for building adaptable companies in a rapidly changing AI landscape.
“So you're just going to see this explosion, this kind of plethora of open source models that have been incredibly customized in really cool, weird ways.”
BaseLabs and the Open Source Vision
55:40 to 56:01
Discover BaseLabs' mission to empower users in model training and ownership.
“extrapolate out the scaling laws and figure out what they're gonna be able to do.”
The Future of Open Source Intelligence
56:01 to 56:32
Discussion on the potential of open-source intelligence and its coexistence with closed-source labs.
Introducing BaseLabs and Its Mission
56:33 to 57:22
Overview of BaseLabs and its goal to enhance the open-source ecosystem.
“And the point of base labs is essentially, can we be as helpful as possible in not only like propping up the open source ecosystem, but helping it to thrive?”
Strengthening American Open Source Labs
57:23 to 58:29
Plans to support American labs in their open-source efforts and improve capabilities.
“It looks like a lot of different things over the like, you know, short to long term.”
Hiring for High Slope Research Talent
58:30 to 1:00:39
Exploration of what qualities are crucial in hiring researchers amid rapid AI advancements.
“summarize the end of your session and keep going.”
Curiosity and High Frequency Trading Talent
1:00:40 to 1:02:01
Insights on sourcing talent from high-frequency trading and the importance of curiosity.
“Yeah and the network The work effects of this are like so incredible and rewarding to see.”
Motivating Factors for Career Changes
1:02:02 to 1:04:08
Discussion on the motivations behind career shifts from quant trading to AI research.
“I think the short answer is high frequency trading.”
Efficiency in Research and Experimentation
1:04:09 to 1:06:55
Strategies for enhancing research efficiency and decision-making in AI projects.
“I think I was really lucky going through undergrad to have a lot of really good research mentors that were not necessarily the traditional academics and like follow the traditional academic process.”
The Evolving Role of Researchers with AI
1:06:56 to 1:09:38
Exploration of how AI tools transform the role of researchers and the nature of experimentation.
“If I gave both of you a model tomorrow that could execute every single experiment you could describe perfectly, what would remain the job of a great researcher?”
Nonlinear Pathways to AI Research
1:09:39 to 1:10:01
Reflection on the diverse experiences that shape approaches to AI research.
The Evolution of Experimentation in AI
1:10:01 to 1:10:34
Discover how the accessibility of AI experimentation has changed the landscape for researchers.
“And I think that option is now available to a lot of people.”
Personal Journeys into AI Research
1:10:34 to 1:11:52
Learn how personal backgrounds in medicine and law influenced insights in AI.
“Maybe, Morty, you trained as a doctor originally.”
The Philosophical Shift Towards AI
1:11:52 to 1:13:38
Explore the philosophical implications and emotional connections to AI and machine learning.
Future of Work in an AGI World
1:13:38 to 1:15:06
Discuss the potential impact of AGI on jobs, particularly in knowledge work and medicine.
“I don't think maths and so on had appealed to me that much before.”
The Value of Human Judgment in AI
1:15:06 to 1:17:15
Understand the lasting significance of human oversight and decision-making in AI applications.
“I think we're a long way off AGI in the traditional sense of how people define it of like the LLMs doing everything.”
The Importance of Domain Expertise
1:17:15 to 1:18:13
Learn why domain experts will be crucial even as AI evolves and automates tasks.
“And so I think that's actually a really exciting place to be.”
Transcript
Automatic transcript. May contain errors.0:00Charlie O'Neill:The fact that you can get essentially like a god in a box for$20 a month is like kind of insane. What do you think is the right allocation of capital for Australia in terms of investing in our AI future? Building our compute is like not enough. What is the compounding thing you can do in Australia to actually, you know, give value over time? And to us that's enough compute to actually do interesting research and do interesting things with LLMs but also try and build out some sort of talent hub. There's so much talent in Australia. It's a bit of a sad reality where the smartest people in Australia, You know, they go overseas to go seek, you know, an opportunity to work in frontier tech.
0:33Feels like things are moving so rapidly.
0:36Charlie O'Neill:The reason open source is so important is because if intelligence is supplied by like one or two organizations, then no one in the world is controlling their own destiny except Anthropic and OpenAI.
0:49Here is a sentence that would have sounded like science fiction three years ago, and it comes straight from one of today's guests. For about$20 a month, you can rent something that behaves for a whole afternoon like a god in a box. Base 10 is a company we've backed at Blackbird. They're based in SF and co-founded by two Australians and are just coming off the back of a USD$1.5 billion Series F that put them amongst the most highly valued private AI infrastructure companies in the world. Base 10 essentially runs the plumbing that lets the world's most demanding AI native companies, think of Cursors, the Harveys, the Lovables, actually put models into production.
1:30But the two guests you're about to hear from don't sit actually on that side of the business. They run the research arm, and they're dedicated to keeping open source AI competitive with the Frontier Labs. And they were brought into the Base 10 team after their own company, Past, was acquired at the end of last year. This week, I'm handing over the mic to my colleague, Saron Bohanian, who leads Blackbird's deep tech work, who sat down with two Aussies, Mudith Chaya Sakara and Charlie O 'Neill in San Francisco. Before you dive in, here are a couple of terms that they use constantly that some of you may not be as familiar with.
2:03So a closed model, which they talk about, think of ChatGPT, Claude, Gemini, is kind of like a black box that you rent by the question. An open-weight model means that the lab that publishes it actually gives you the file the model learned from. You can download it, run it, adapt it. Open source goes further still, releasing the recipe as well as the weights themselves. And there's this concept of post training. So think of a model as a brilliant generalist fresh out of an extremely expensive education and post training as the apprenticeship that takes that model into something excellent at a very specific job.
2:40Before any of that, there is harnessing, which is essentially what you can get out of a model just by prompting and scaffolding it really well before you've touched the model itself. So why is this a conversation that we wanted to have right now rather than say 18 months ago? Well, one of the reasons is because the gap between open and closed models on intelligence has all but closed. And that changes a lot of how we might predict the future will play out in AI. It's now obviously a very live discussion in many token heavy firms about when and how to move from AI built on closed models to AI built on specialized models, which blend their own data and open source models.
3:20This move often enables more control, often lower cost and a better fit for a particular use case. To build on an ice cream analogy that Charlie brings up in the episode, it's a bit like shifting from a delightful vanilla ice cream mixture to something perfectly designed for your taste. Moody and Charlie walk through all of that and even talk to the geopolitical tensions in the open AI model space. But actually, as a listener, there was something bigger underneath it all that I found quite unexpected and delightful. And that's what it sort of feels like to these two people who are sitting in the eye of the storm of the most interesting AI research on the planet.
3:57They both talk really candidly about how fast their hardest one work becomes the commodity that fuels everyone else, and actually where that leaves the most durable jobs in the future, why it is that they believe so fiercely in Australian talent and what it means to better career on this moment right now. There's also actually a smaller and funnier thread, which is a bit like what's it like to make decisions in a multi-billion dollar company that's kind of given up on planning more than a few weeks ahead. Here are Moody and Charlie in conversation with Saron. Charlie, Moody, it is so wonderful to have you guys on the podcast today.
4:35There is a lot to cover, so I'm keen to jump in. I'd love for you guys to give me a one-on-one on how you define open source model versus open weight model versus closed model for folks who are watching who may not be as close to these definitions or be as familiar with them.
4:52Charlie O'Neill:Broadly referred to as open source, there's probably the two main distinctions are open source and a subset of that is open weight. So open weight is purely where the lab who trained the model releases, you know, all the individual numbers in the model and the matrices required to run that model. And like, you know, really all they are is numbers at the end of the day. You feed in tokens which get converted to numbers and then tokens come out the other side as numbers. And that's really useful, obviously, because you can then take those weights and run it on your own GPUs. You can train them. You can do whatever you want with them.
5:20Charlie O'Neill:But like there's a broader category which is open source and in order to fully understand what a model is capable of also again what it's like implicit biases are, what its values are, what its censorship is, you need more than just the model itself. It's very hard to like figure that out from just running the model. Like a much easier way is to look at okay what is all the pre-training data was trained on and like this is very hard for a lot of companies often because that pre-training data often looks like the whole internet. It's like you know on the order of 100 trillion tokens which is a lot of data.
5:50Charlie O'Neill:It also looks like, you know, the mid training data, which is like kind of a phase after pre training where we give models things that look more like tasks in the real world, like examples of tasks and like see how it does on those. And there's also the post training data, which is like the RL environments that the model is trained to like get very good on and like in particular, even like the rollouts on that. So one labs don't necessarily want to release that because it really is just like the exact recipe that you use to train the model. like even open source labs or I guess like open weight labs want to protect that IP to some extent because otherwise like you know they give an instant catch up there's anyone who's behind them but there are people out there doing this so like Percy Liang from Stanford and like there's a bunch of others doing like kind of experimentation in the open with like what does it look like to to release this whole recipe I think for most people like you don't need to be able to see the whole recipe like the most important thing is like having the open weights to be able to like do whatever you want with hosting yourself training yourself but for or researchers like us, I think it is important that at least some people are trying to release the recipes as a whole so we can understand what they're doing.
6:50Charlie O'Neill:And it's a spectrum, really, because Deepsea, Kimmy, Nematron is really good at this. They're very, very close to fully open source. They release the recipes and all the things I talked about before. But researchers need to be able to understand how these things are being built so that we can build on top of them, particularly post-training labs and interpretability labs. And it's very hard to do that when you just have the weights of the model. I think a lot of people outside of this world have a fairly crude mental model. You either use a model from OpenAI, Anthropic or Gemini, or you use some Chinese model and supposedly take on a bunch of security risk.
7:27How wrong is that picture today?
7:29Charlie O'Neill:Maybe if to distill it down, like the fundamental misunderstanding is that there's not a huge difference between the, you know, the closed source frontier models that you mentioned, the Chinese open source models, which are now getting very good and the American open source models, which are also starting to grow in strength. I think that fundamentally it comes down to like everyone has the same recipe and everyone is doing the same things. And they're at slightly different, you know, rungs of the ladder of scaling. But at the end of the day, like we're seeing all these models get better in tandem.
7:56Charlie O'Neill:no matter how you break that down like the scaling was hold for everyone and there's no real fundamental difference between you know how open ai and anthropic train a model how the china's labs train a model how the american open source labs are now starting to train models and we know that like people will continue to like march along that scaling path of course when you break it down like probably the biggest distinction is open source versus closed source so right now like there's a lot of pressure from a cost side from a moat side like in terms of owning your own you know, feedback cycle to be able to like deploy better and better models and ensure that like as the frontier models get better, they don't like start to verticalize and eat your lunch.
8:32Charlie O'Neill:And of course, that comes with like a lot of debate around, you know, if we are going to be running open source models, is it going to be American open source models or Chinese open source? What are the security risks implicated with those? Like there is actually a very strong argument for running open source models from a security perspective, but there's been a lot of debate about whether, you know, we should be, it should be permissible to run Chinese open source models. But from what we've seen at the end of the day, like these things are all trained in the same way. They have the same security risks and concerns.
8:59Charlie O'Neill:Probably a key piece of evidence here is that we're actually finding that most of the, you know, cyber hacking and things going on is actually still using Frontier closed source models, despite all the safeguards being implemented around them. So there's a lot of debate. But like I think the main thing to take away is like these are all fundamentally the same thing. and what to use comes down to cost. It comes down to the type of intelligence and how that intelligence needs to evolve for like your business rather than, you know, I'm going to make a bet on a particular lab or a particular distinction between which country is training the model or whether it's, yeah, you know, open source or the host source.
9:34I'd love to double click on that intelligence piece. It feels like the tech industry is increasingly obsessed with who has the smartest model. Can you tell us why that is increasingly the wrong question as well?
9:45Charlie O'Neill:Yeah, so Moody will have a lot to add to this because we've developed this thesis a lot over the last year and a half, I'd say. The original thesis, which is very, very contrarian when we came out with it almost two years ago, was that open source models are post-trainable, whereas closed source models aren't really post-trainable by individual people. and both of these as I said before getting better in tandem like you can kind of imagine them as two like parallel lines and there's I think always going to be a gap between open source and closed source in terms of frontier capabilities but a way to bridge that gap is to shift that lineup is post-training and specializing your model for the particular tasks you needed to do like you don't necessarily need your your model to have this obscure knowledge of like 17th century English literature if you're like doing you know like legal work or like you know a long-running sales agent or something like that.
10:33Charlie O'Neill:And now this thesis has started to come to fruition in a big way because the amount of post training we needed to do before was directly determined by how big that gap was. And that gap is in some sense shrunk over the last like year and a half. So, the amount of post training we need to do has like lowered. The friction to do that post training is lowered. And now we're starting to see like all these companies go, okay, I can't just do a one-off post training and then, you know, I have a slightly cheaper model. I've got a smaller model which can do the same task at the same capability level however you want to sketch that proto frontier now it's very much a matter of okay if i don't do this like the like the scaling laws are continuing to hold like frontier labs or anyone can replicate my software for free and i'm now just basically a wrapper around like you know claude or or open ai so how do i tap into that continuous like feedback loop to get like data from my users who tell me the things they love about my product who tell me the things they hate about my product and make the model doing my task better and better over time.
11:28And I think the other thing is we see like a lot of chat about benchmarks, but I guess like with every customer that we work with, it's always kind of how do the models perform on their particular evals and then work out where are the different models performing on that. Post-training dynamics differ. Like there are all these factors that make it not just clear from a benchmark which model to start off with. So, you know, it's indicative, but it's definitely not the thing that is like determining which model you end up with. I'd love to hear more about your customers. Base 10 works with some of the most sophisticated AI teams in the world.
11:59What are some of the best companies you're working with doing today that you think most companies will be doing in a year or two from now?
12:07Charlie O'Neill:There's now this kind of well-recognized playbook of once you get to a certain scale and have enough data to like kind of warm start this pipeline, you can kind of become your own Frontier Lab in a sense. Everyone always talks about how OpenAI and Anthropik are going to eat the world. But I think what we're actually going to see is every company ends up turning into a form of OpenAI and Anthropik rather than the other way around. And so this playbook is fairly simple from a high level. It's essentially like we take all our data, we take the biggest, best open source model available, and we engage in something which looks very similar to a mid-training and post-training run that the Frontier Labs will do.
12:43Charlie O'Neill:So we're talking weeks, if not months, of training time, and large clusters like people like Cursor and then others that we're working with now have got like four clusters all over the globe that are running this training and this reinforcement learning at scale. And then they do this process and like, yeah, the first iteration, like, you know, you get Composer 1, which is maybe not the most usable model. But a lot of things have happened since Composer 1, you know, that was trained on a very old Kimi model. Kimi has gotten a lot better since then, as have like a lot of the other open source models.
13:11Charlie O'Neill:And, you know, they figured out how to do this process much more efficiently in the same way that the labs first the first time they did have to figure out how to scale their reinforcement learning staff and how to do this properly and then they deploy this model and users continue to tell the company the things they love and hate about that model and how it's doing the task in their particular harness and then i think the most interesting part of that is like the the online phase which is where you're actually like kind of plugging in that that full feedback cycle so for instance cursor and again other companies that now we're now working with have gone down this online reinforcement learning route, which is like literally like every five hours you're collecting data from users.
13:48Charlie O'Neill:They're telling you things that the model did wrong, things that it did right. And you're doing an online update of that model in real time, which is amazing. Like this is exactly what the feedback loop should be. That's exactly how you build a moat. It's the only thing you have that Anthropik doesn't have, which is like specific user data for your task and users telling you what they love and hate about it. Would you say that is the biggest difference from say 18 months ago, this idea of the having the online user data in the feedback loop or yeah tell us more about what has changed most recently i i think deep seek and grpo which is a variant of the reinforcement learning algorithm that that most like you know post-training labs run that only came out like the beginning of last year like it was in its infancy um reinforcement learning as a whole like post-training at that stage look like okay i'm going to collect some examples of of some model doing this task or So I'm human doing this task and I will supervise fine-tune on that.
14:40Charlie O'Neill:And supervised fine-tuning is very, very brittle. It doesn't allow the model to figure out for itself what the optimal approach is. It can actually mess with the model in weird ways. The model can forget things. It can degrade its general capabilities. There's a lot of problems with supervised learning. So the reinforcement learning stack maturing has probably been the biggest change. And online reinforcement learning is just one part of that. But it would have been infeasible to suggest someone like Cursor or these other companies like Harvey doing a big post training run for months with reinforcement learning like you know a year and a half ago but now like we figured out a lot of things about how these things work and and now yeah as i said these companies are starting to look more more like mini versions of opening and anthropic just pointed at a particular domain or task i was going to say as well i think 18 months ago the actual post training stacks that existed out there made it very challenging to actually do large-scale runs and like you have companies like you know thinking machines with their Tinker product that made really awesome abstractions and also abstracted away the really challenging bits of the training library to be able to train these kind of, you know, trillion plus size parameter models.
15:42And that's like something that we've even seen in the last like six months, which is like, that's been a big focus of ours, which is like giving people just raw compute to run whatever training library they want on them like by themselves actually was really challenging to see customers get to value. But now we've tried to abstract away the training library work. And so, yeah, I guess like what we see is really, you know, the hard thing for customers to do is to find their emails, to find their rewards. And then the actual training library stuff is what we try and solve.
16:08Charlie O'Neill:Yeah. And how like how exciting is that that it only took 18 months to commoditize like frontier open AI level post training? Like that is really cool. And I think that's one thing we believe in is that like the training side, like how to apply my gradient to update the weights of my model should be commoditized. And the friction from that should be completely removed. Like no customer should be ever debugging like MV link issues or like GPU communication issues or whatever. The hard part will always be the real qualitative intelligence of like defining like, okay, what behaviors do I want to teach them in the first place?
16:38Charlie O'Neill:Like how do I determine whether something is not only correct or incorrect, but like good, very good, bad, very bad. And so that's a lot of the work that we focus on. Can you tell us how a sophisticated team decides which model to use for a specific task? yeah i would you mentioned this earlier but fundamentally again these recipes are all the same like out of the box like you're going to get different scores and benchmarks different models are going to have different like you know jagged points where they're better at some things than others and worse at some things than other models but at the end of the day like the main thing determining intelligence at the moment is scale um and the size of the model is like very very indicative of like how good it's going to be for a particular task and like it's nice because as the models get better at any given size you can do more and more with that particular model so you know if you had to post train kimmy k 2.5 six months ago you could probably get away with you know a few hundred billion parameter so kimmy k 2.5 was like a trillion parameters that size is rapidly shrinking but realistically if you're going to go down this like frontier route of like post training a large-scale lm for very agentic open-ended tasks you're essentially just going to pick the largest open source model available and run with that and that's why we're really excited that America is finally starting to invest in open source, but also in particular larger open source.
17:53Charlie O'Neill:So in the next three to six months, we're going to see a lot of multi-trillion parameter American open source models coming. Where do you think having access to the absolute frontier of a model materially changes the tasks that you can do? Like you mentioned, this gap between closed and open source models changing rapidly over the last 18 months. You know, will there ever be specific tasks or a time where it is incredibly important for a company to use a front tier model versus an open model? The thing we believe in is like open source should not just be the only thing that exists in the world.
18:24There should always be open and closed source. And maybe Charlie can touch on some of the fields that we think that closed source really will continue to be really apparent. But I guess like for us, the goal is that, you know, how do you make open source models as easily to run and to train as possible? because there is this huge amount of economically valuable tasks that you can use. And I guess like that's Base 10's big focus at the moment where, you know, again, like if you just have raw compute out there, it becomes very hard to run efficiently, to train efficiently. And so, yeah, that's where our big goal with open source is.
18:54But on the close source front.
18:55Charlie O'Neill:Yeah. Maybe there's two things I'll say here. So the way to view this, I think, is that for any like given economically valuable tasks that you could possibly use an LLM for, there's like a kind of a watermark, a high watermark above which there's like very low, like diminishing marginal returns and more intelligence and below which you can't really do the task at all. And I think viewing open source versus closed source from this like absolutist point of view rather than a relativistic point of view is like really, really helpful framing. Like it doesn't matter really what the gap between open source and closed source is.
19:27Charlie O'Neill:What matters is like has the open source model hit that watermark or not? And one way to get to that watermark is post training. So I think over time we're going to see this pattern of, okay, people start off using frontier models, frontier closed source models. Eventually the open source model gets to the point where it can do it. And then for many reasons, for cost reasons, for security reasons, and a bunch of others, people will transition to open source for that task. That kind of implies that the task distribution is static. And I think the economy is going to change over time. We're not going to keep doing the same things forever.
19:56Charlie O'Neill:And so I think there'll just be this oscillation between like open source and closed source is like one leaps ahead and then others catch up. The second thing to say here is that I think we're very lucky to have the gap that we have. And as Moody said, I think it's important to live in a world with both closed source and open source, which I think is another common misconception because like people see like from getting open source and you must think that, oh, you know, like all closed sources are evil and bad. And I don't think that's the case. Like I think closed source has been incredibly useful for like at least providing a proof of concept for like how to scale these things things up.
20:27Charlie O'Neill:But more importantly, they're kind of a canary in the coal mine. And we're getting to the point where this is starting to matter. So, you know, the Hugging Face of an Eye incident from from a few months ago was a really important example of like, yes, nothing really bad happened. But this was a very clear indicator of what could happen once these things are deployed at wide scale and like, you know, Mythos and Project Lastwing from Anthropic was another good example of like, OK, we need at least a bit of a gap before these things are widely distributed without any centralized guardrails where we can determine what's going to happen and harden the systems as it happened.
20:55Charlie O'Neill:And I think that promulgation of all the open source should happen because the world is going to need to be hardened to it at some stage. So you need a long enough gap to be able to do the hardening, but not too long of a gap such that only a few enterprises have been given select access to the latest frontier models are able to benefit from them. And yeah, I think it's a really healthy kind of situation right now, and I hope that settles into an equilibrium. Yeah, and maybe to crystallize my point, I guess previously we'd always suggest people go iterate with closed source first and then go to open source.
Read the full transcript
21:24And I guess like the reason for that was there weren't mature enough serving stacks where you could just hit an API the way you could an open AI API and similarly with post training. And I guess like that's now shifting. But like Charlie said, there is definitely some domains where, you know, closed source will still always be a very useful part of.
21:40Charlie O'Neill:Yeah. And it's interesting because like, yeah, that watermark has risen for open source. So as Marty said, it used to be always, okay, start with front of closed source and then transition. but I mean there's a lot of tasks out there now where you can just run Kimi K3 GLM out of the box even like many of the other models and like it would just do the tasks and I said there's very diminishing marginal returns to intelligence past that sort of watermark so like we're hitting this point where open source models kind of can do a lot of stuff and like people are almost transitioning the things that they're being ambitious about and what they're using LMS for in order to like you know, eggs the most out of the frontier.
22:17I'm curious to hear about the experience of the consumer or your reflections on how this will affect consumers. Do you think at the level of the consumer that will eventually become something they don't have to decide themselves and the systems default to them fitting the best model? How do you think that will change for everyday users?
22:35Charlie O'Neill:Yeah, unfortunately, I think like at the moment, like the consumer surplus from frontier models, even open source models is like some of the largest surplus we've seen in history. Like the fact that you can get essentially like a God in a box for$20 a month or even if it's$200 a month is like kind of insane. I think that eventually, yes, like there's a lot of talk about routers. There's a lot of talk about like how do you partition like, you know, total token spend between like enterprises who'd be willing to spend, you know, 10 ,000 to 20 ,000 on one particular identity query that's going to run for like, you know, a few days or a few weeks versus a consumer who's like learning like college maths or doing an essay with an AI.
23:12Charlie O'Neill:And I think, unfortunately, some of that surplus is going to shift, probably not back to the people training the models, because I think it's very hard to capture value in the long run, but at least to the enterprises and look much more like software today in terms of what the value can get out of it. You mentioned GLM and Kimi K3 earlier. People often hear open model and equivocate that with a Chinese model and then kind of associate that with some sort of security risk. Before we even talk about whether or not that characterization is fair. Can you maybe walk us through why that is even a question in people's minds at the moment?
23:48Charlie O'Neill:Yeah, I think it does come down to like, you know, geopolitical tension. People see the frontier open source and they also see how quickly that's caught up to the frontier closed source. And the leading labs doing that at the moment are Chinese labs. With that comes concerns like, oh, even though this is running on my own GPUs, what if there is like, you know, some sort of trigger word which is being trained in? There's obviously like different censorship practices. Like censorship isn't just something you actively decide like after you trained a model. Like censorship is determined by the whole training process from start to finish.
24:18Charlie O'Neill:It's determined by what pre-training data you have. It's determined by what RL environments you have. It's determined by, you know, how you RLHF the model to like respond to certain queries or not. And so obviously like you have to have standards. Like you can't just like not have an opinion. Like your standards are implicitly like baked into your whole training process. And so people see all that and say, okay, like these Chinese models must have very different standards to what we have in the US. and hence there may be some security risk there. I think like explicitly I don't believe and we've never seen any evidence of like these trigger words being baked in or like you know Chinese models going rogue and like serving some like you know nation-state purpose for China.
24:54Charlie O'Neill:I do think that like it is a very fair concern to like regardless of whether it's a Chinese or American model like evaluate the censorship that has been baked into that model and not just censorship but like what is its values. Like we do this with Anthropic all the time. Like Anthropic has a very explicit constitution that determines what Claude's values are and people debate it. Like, there's a human-written document with the help of AI. Like, should we have one of those documents at all? Like, how do we debate, like, what morality these models should have and what ethics they should have? So I don't think it comes down to, like, China versus America.
25:22Charlie O'Neill:I think it's simply like, okay, different models have, like, different preferences of how much they want to help humans, like, the way in which they help humans and, like, how they respond to, like, certain queries. And, like, that's a debate. Like, the whole AI industry needs to have as a whole. And I don't see any fundamental distinction between, like, China and the US in that regard. But I think like, you know, for all intents and purposes, like the Chinese model seemed to be fairly aligned in terms of like refusal versus like help with American models. It's just that we don't have these like kind of opaque classifiers on top.
25:49Charlie O'Neill:And a good example of this is like for the Hugging Face breach, Hugging Face actually had to end up using open source models because the frontier like models from OpenAnthropic were rejecting the request to dig into the cyber breach. and so you know hugging face used open models and like they didn't have those extra classifiers on top and were able to get to the bottom of it pretty quickly yeah i guess like for us as well like with base labs a bunch of our research is around things like constitutional alignment things like interpretability and understanding kind of what features are firing and like you know that's where we started as pars as a mechanistic interpretability lab but yeah it's not a solved problem but like yeah as charlie's mentioned nothing's come up in the stuff we've done so far but it is like the idea of even post these post training these post train models already like what does it look like if you actually post train a Kimi K3 or a GLM 5.3 and actually see what you can do around this alignment question is something we're actively researching and that's a great thing about open source is like even if I am given a model who's like you know guardrails or implicit biases I don't agree with like it's right there I can I can post train it and I can I can adapt and mold it to like what I think is like right and what I should be serving to like my customers if someone wanted to consider using an opening weight model what distinctions do you think they should be thinking about or caring about?
27:02Charlie O'Neill:I think the big one is honestly just like size. Like do I have latency requirements? Like what level of intelligence do I need to do my task like perfectly like in an ideal world? And then you know you can sketch out this Pareto frontier of like okay I can trade off like size and cost and essentially speed against like intelligence and like where's the maximum point I can afford? And then you know that curve shifts outwards with post training. So like cool if I'm big enough I'm probably going to consider that as well. But that's really all it is. I don't think anyone should be going into this with like this bias that I have to use a particular model from a particular country or a particular lab because of like the values baked in.
27:35Charlie O'Neill:Those values change all the time anyway as these training stacks evolve. So it really is just like, okay, I'm starting from roughly the same point given size help constant. I'm starting from roughly the same point. So like let me just evaluate it based on those merits. Yeah. And I think where we see people go wrong is like post training right now is like so in vogue. And we just see people want to jump to, you know, a post trained Kimi K3. But, you know, serving a dedicated instance of, you know, a multi-trillion parameter model is like not very economically viable for most companies. And the base model is so good.
28:04So, yeah, I think like that's something we see people often go too far and get too excited about the post-training rather than just using the base.
28:10Charlie O'Neill:Yeah. Like I think the rule now is like no post-training before product market fit at the very least. Yeah. Wise rules. You mentioned this a little bit earlier, but why do you care that America develops a serious open-weight ecosystem if you have the likes of OpenAI, Anthropic, and Google who are building some of the world's best closed models? From our perspective, the reason that, you know, Chinese open source got out to a big head start compared to American open source was that America did have the best AI talent. It just, like, aggregated in, like, a couple of places. So it's very hard to get the talent required to train these models when there's the option of OpenAI or Anthropic there.
28:51Charlie O'Neill:China didn't have this dynamic. China were able to put its best AI researchers into these open source labs. And China went down the open source route because they were kind of late to the party and a bunch of other reasons. I think now it's really important that every individual entity controls its own destiny. And I mean that from several different levels of abstraction. For us, like the reason open source is so important is because if intelligence is supplied by like one or two organizations, then we're not controlling our own destiny. Like no one in the world is controlling their own destiny except Anthropic and OpenAI.
29:22Charlie O'Neill:But you can step levels down that ladder. From a country perspective, I think it's really important that like, you know, if you have the resources capable of training these frontier models, then you have a say in what those values are. You have a say in how people use them. you have a very frontline view of like all of the safety things and all of the like preparedness things that you need to do so that society can accept this level of intelligence and not just break i think it's really hard to do if you're like not exposed to the frontline you don't have a very clear line of sight to the frontier and then from an individual perspective i think as well like it's really important to have you know control of your end destiny at regardless of what level you're operating at and you don't want to be beholden to like you know anything outside of control because of how important these things are.
30:01You guys have also spent some time in the UK and obviously Australia. What do you make of each country's sort of strengths and weaknesses in this space? We're all just trying to get back to Australia, I think is the key point. But yeah, we started our, like we started paths in the UK, we hired out of Australia and we very quickly moved to the US. I think on a sadder note, it's like very interesting to see how in SF and like Like even in SF, like, you know, compared to being in New York, there is like a massive discrepancy in the way that you're able to collaborate and the way you're seen as like a research lab, which like, you know, there are really amazing benefits of that.
30:36But, you know, there are also really challenging parts where, you know, you're considered relatively unserious if you're in the UK doing mechinterp and post-training. But yeah, I guess like on a positive side, I think the talent that exists in the UK and Australia like is just as frontier that exists in SF as well. yeah I think funding and like velocity of capital is so much better here which like makes it very easy to do the relevant research and build a company but yeah from a talent perspective we
31:02Charlie O'Neill:believe that it is kind of distributed everywhere yeah and I think like Australia is not in the next few years going to train a frontier model let's just be very honest like nor should we right like there would be very inefficient allocation of resources I think the world that Moody and I dream about again comes back to like this stuff being so commoditized that like countries with just the talent can do really cool things with it even if you know the concentration of capital and and network effects that you might have at sf aren't necessarily there there's so much talent in australia like this was one of our core things when we started paas which is like everyone tells you this really obvious advice when you start the company of hire the best possible people and that sounds really obvious when you get into it it's actually very very easy to drop your bar so we were very deliberate we were like okay how do we get anthropic open ai level talent with you know a small seed stage startup and to us like that looked like okay we know there are really really smart people in Australia and there's also very limited options to work on this stuff in Australia and so they funnel into a few places they funnel into quant trading and consulting and a few others so can we go and get the best possible people from those fields and like even though they've never touched an LM before like bring them over and like rapidly skill them up like looking for I guess like high slope people versus high intercept people and that strategy works so well and we can honestly say that some of those like Australians that we brought over and now are the best in the world at what they do.
32:17Charlie O'Neill:And that's really exciting to us because they've been doing it for like a year or a year and a half. So we know that the talent exists in Australia. Like there's one big thing we realize moving over here is like it's very easy to get daunted by all these personalities and AI and like all the things they've accomplished. But at the end of the day, I have full belief that like the upper end of Australian talent is just as good as the talent here. It's just like, you know, lack of network effects, lack of things to work on. And it's not till they come over here that they really bloom and blossom.
32:41Charlie O'Neill:So our aim eventually is like all this stuff is commoditized, like, you know, frontier models are readily available and all the research and applications and entrepreneurship you can build on top of that will come back to Australia. Separate to the talent piece, what do you think is the right allocation of capital for Australia in terms of investing in our AI future? Yeah, for us, just building our compute is like not enough. And what we think a lot about is what is the compounding thing you can do in Australia to actually, you know, give value over time. And to us, that's enough compute to actually do interesting research and do interesting things with LMs, but also try and build out some sort of talent hub where people don't just need to leave.
33:18And I think Charlie touched on this earlier, but it's a bit of a sad reality where the smartest people in Australia, you know, either go into quant trading or, you know, they go overseas to go seek, you know, an opportunity to work in frontier tech. And I guess our goal at some point is working out how do you give people an opportunity where they do work on, you know, frontier LLM research, for example, and also have enough compute to do the interesting
33:40Charlie O'Neill:work yeah there's a critical mass and we're nowhere near hitting it like you have to have enough smart people in a room together for like cool things to happen um a metaphorical room of course and like we want australia to get there one day but unfortunately because of all these reasons like it's just not there yet and so we view the path forward as yeah commoditize the really hard stuff and make it rarely available and then go build like the network effects that you need to in australia i want to touch on a company's use of open models again can you tell Tell us, you know, when does the company cross the line from using existing models to actually training one themselves?
34:13Charlie O'Neill:Yeah, I think there's a lot of things to consider. Like, I think the dominant thing you'll hear companies talking about is cost. Like, you know, Uber wakes up one morning and like they've spent their, they've done their yearly spend in like a couple of months. Or like, you know, Meta's anthropic bill is so high or like, as you step down the wrong, you know, these like kind of AI native startups are just spending tens of millions, if not like, you know,$100 million a month on anthropic and open AI. And of course, that's like a very, very visceral wake up call. But I think it gets to the point as well where for any given task, like with the frontier models considered roughly the same level, like your advantage at the start is like how well can you prompt and like harness and scaffold the frontier model and your wrapper essentially to do the task as well as possible.
34:55Charlie O'Neill:And unfortunately, there's a ceiling that you just like kind of asymptote towards. And like importantly, your competitors will asymptote towards that ceiling as well. So I think like that work has to be done at some point. Like you have to construct that harness and that scaffold and all that prompting. But if you're serious about being a competitive business in the long run and winning whatever space you're in, like your question should then turn to like, okay, I've asymptotied at this 80 % ceiling. Like how do I get past that? And the answer to that is like, okay, I'm going to go off the base models and I'm going to try and teach them how to do this better than I can from just specifying it in plain English.
35:27Charlie O'Neill:And that's when you embark on the very often long and arduous but very rewarding journey of post training. What has to be true about the task for that to be possible? There's many different tasks. I think there's tasks that meet that bar at a very low level. So like you can imagine, you know, for instance, like clinical notes, where it's a kind of a one-step thing. You're feeding this transcript. The model is then writing like, you know, the doctor's notes for the visit. Or like, you know, these voice call models where essentially like you get the speech in, the LM writes the text very, very quickly, and then it turns back into speech on the way out.
35:57I think those are tasks that like pretty quickly you kind of converge on the optimal prompt.
36:02Charlie O'Neill:There's not that much you can do to like keep fixing it. And anything you do in addition is going to be overfitting to like whatever you've done. So you would then kind of do post-training immediately. And that post-training doesn't necessarily look like reinforcement learning either. That's one of the cases where you can use supervised fine tuning and teach the model exactly how to do this like fairly structured mapping from, you know, transcript to doctor's notes. But you know, as you move up the complexity level of tasks, that gets a harder and harder consideration to make. So a very, very agentic, for instance, like coding tasks, if you're someone that codes or allowable, if there's very complex harness, a user could theoretically do like basically anything you want.
36:36Charlie O'Neill:It's like almost like Turing complete in a sense. Then if you are going to embark on post training, then you better be doing post training well enough to be able to teach the model this very, very general world and get better at this very, very general world without breaking any of the capabilities of the model. And that's no longer, you know, supervised fine-tuning on a few input-output examples. That's like quite complex reinforcement learning, which also takes a lot more compute. Can you walk us through an example of a customer that you've worked with where you've gone from building a product with them to them eventually training their own model?
37:08Charlie O'Neill:So maybe a good example of this is one of the kind of big coding companies like Lovable. Essentially, Lovable has been largely using frontier close-source models. you know those frontier close source models are the best available option often but they're still not perfect and they do many things wrong so like kind of the first step involves us coming in and just like basically looking at a bunch of data like unfortunately or fortunately like you imagine this world of post training to be very sexy and very algorithmic and it's all automated but like at the end of the day like 90 to 95 percent of it is just like understanding the task and looking at the data and so you might go okay here's how you know opus is doing it for this particular task or whatever it is, here's if I run Kimi K3 through this task, like a user is asking me to build this whole website from scratch with this backend and this database and everything, here's how Kimi's doing.
37:55Charlie O'Neill:And then we basically compare a bunch of those examples. We see what Opus is good at, we see what Kimi's good at. And then it really is just about identifying, OK, this user is clearly upset. Opus has not put the button or whatever in the right place like four times in a row, hasn't inferred exactly what type of mechanism the user wants here. Kimi is like being really inefficient at tool calls. It's like reading the same file like four times. Like very, very specific things like that. Like that's how granular you're getting. And then you collect and aggregate enough data at a high enough level and you can start pattern matching about what the models are good at, what the models are bad at.
38:28Charlie O'Neill:And then this process very much looks like exactly what the Frontier Labs are doing right now. You build reinforcement learning environments around those behavior modes. So, you know, for the tool calling efficiency example, like I might take a similar example to like, you know, this user building this web app where Kimi was very, very inefficient. And I will specify rewards not only for doing the task correctly, but for doing the task efficiently. Like how many tool calls are made and how many like repeat tool calls are made. And then we essentially like let the way reinforcement learning works is we just like run those models in that harness.
38:59Charlie O'Neill:Then once they've got an output, we tell them if they did a good job or a bad job and then adjust the weight based on that. And it's this, this is exactly what OpenAI and Anthropic and all the big labs are doing. It's exactly why they pay, you know, billions of dollars to McCaw and Handshake and all these guys is, okay, here's something the models are like doing quite poorly can you build me an our own environment so i can teach the model or it can teach itself to do it better and so again like there's no fundamental distinction between this process for an individual customer and exactly what open air is doing it like you know the scale of tens of thousands if not hundreds of thousands of environments like this is how we teach the model the recipe is the same in all cases it's just that we can spend a lot of time building environments exactly for this customer exactly the things the model is doing wrong on this particular task we don't have to worry about everything else yeah and then i guess the thing on top of that is routing which like we work a lot with specific companies routing again is very in vogue at the moment people want to have this kind of silver bullet of a magical router that can route between easy and hard queries for every query in the world and maybe this is something that we disagree with a lot like with what a lot of people are talking about on the timeline and that i guess training a router specific for a company's use case we really believe strongly in but training this general router for you know a company that's trying to divert every coding use case to different models is actually a much harder problem.
40:11What's the most difficult part about reproducing a company's learning loop? Is it developing the RL environments? Is it, you know, getting the customer data or company data in a shape that it can be used for, you know, supervised fine tuning or any other post-training work? Can you talk to us about what is most difficult about that?
40:30Charlie O'Neill:I think we're going through this like short-term pain at the moment where you know companies have built all these harnesses and fantastic products that lms plug into but haven't designed this with reinforcement learning in mind so reinforcement learning we have to let the model not just like you know generate tokens but it has to be able to operate inside of that environment for it to work like it has to be able to access you know super base or or whatever it is in real time and so companies haven't designed these harnesses so that we can easily take that harness and like run it ourselves so the moment there's a lot of pain and like the hardest thing is just like you know plugging the models into the real environments exactly as they would run into production.
41:03Charlie O'Neill:I think like that will be solved. Like people will literally start building these things with our own minds. So it's very easy to export a harness to people like us or internally to post train the model. So I think like in the limit and even right now, the answer is coming up with these our own environments. It's like it's very much an art on the science. You know, there's a there's a reason that OpenAI and Anthropik outsource a lot of this as well as like do a lot of quality assurance and so on internally is that it's very, very difficult to come up with ideas for environments around like you know being it's not a very quantitative skill and it's like you have to be very creative in like how you're specifying an environment to teach a model a particular behavior and then once you specify that environment it's very very difficult to like calibrate it to the right level so that you know the model can succeed some of the time which we need for reinforcement learning but you know it's not like succeeding most of the time because we want it to learn this thing so it is very much just like a data problem and a research problem i think you know rl environments will not only be like a dominant source of like economic value and all the human labor in the next like five to ten years but it will also be a major area of research and people will get very very interested in like how to construct these things efficiently how to determine what behaviors they're going to teach the model and like maybe the reward hacking that occurs alongside the the particular behavior you're specifying because it's like you know under specified or whatever like research is really going to be focused on like okay how do we like if models are teaching themselves to do things within this particular setup with these rewards what is that doing to the models on that point specifically, if there are two companies that have similar models and similar data, where do their systems diverge?
42:35Charlie O'Neill:This is the thing, right? Like we believe and even know to an extent that like the way in which company A post-trains a model versus a company B post-training that model or even for the full stack, the full recipe from start to finish. Like these are very, very fundamentally the same things. It's the same thing across open source and closed source. There's like, you know, a long tail of optimizations that the closed source labs have made that the, you know open source labs probably haven't figured out yet but fundamentally the recipe is the same and your difference comes down to okay i'm training on the same internet i'm like you know training on roughly the same environments i'm probably buying various similar things um from macaw and these guys so how am i going to be very deliberate about what values the model takes on and that sort of stuff and and at the moment you know anthropic's the most explicit about this um but every every company has a model card which tells you like what they believe in like what they've tried to tell the model to do and then like calibrating you know safety classifiers and stuff on top of that so in the lemon like again if you do believe in scaling laws which like you know most people do now you follow the scaling laws and you realize that everyone is also sitting on the same line really the difference is like how large of a model can you train like how much data have you been able to like get off the internet and also pay for this is something that we thought very explicitly about in the last you know a few months is if open source is to continue to improve not only do we need to aggregate compute which everyone talks about all the time but we need to aggregate data in the same way as like the frontier labs are able to do and so we've got a big effort with base labs to offer essentially like open rl environments that are the level of complexity that a frontier lab would want so that open source models can use that essentially for free and train models on top of that because again like it kind of converges to the same behaviors like you're just playing whack-a-mole and plugging these new behaviors into these models but you need to have the capacity to do that and the aggregation of data to do that Yeah.
44:20And as the app layer, like if there's two frontier legal AI companies or two frontier medical AI companies, something that comes up is like, if we're helping someone post train a model for company A, is that going to just be, you know, directly translated to company B? But I guess like what we see is that app layers have such particular like opinionated takes on what their user journey actually looks like. And the models do diverge in terms of what they actually want the model to do. And it ends up being not an issue at all.
44:45Charlie O'Neill:And like we end up, yeah. I think this is a really important point as well because when reinforcement learning first came out and before, you know, DeepSeq and the 01 models and so on, we didn't have any of this reasoning. And people were saying, okay, with reinforcement learning and reasoning, once we crack this, like, you know, it's going to be very generalizable. Like, you can train on a few math tasks and a few coding tasks and the model has learned this, like, you know, philosophical definition of reasoning and is going to be able to, like, generalize to all these different kinds of tasks.
45:11Charlie O'Neill:And what we found out about RL is this isn't the case. Essentially, like RL is a great scaling law because you can teach a model to do a behavior if you can specify it, but it doesn't really necessarily generalize outside of that behavior. And that is why the labs are spending so much money on data. And this goes to Moody's point, which is like, you know, even if you have two legal companies fundamentally doing the same thing, they're writing memorandums or, you know, doing associate work, they're harness and they're set up for that type of work. and even like very minute differences in domains they operate in and like in jurisdictions, it means that like you can RL to get really, really good at that like particular vertical or that particular subdomain or whatever and that particular harness.
45:49Charlie O'Neill:But that doesn't really generalize to like, you know, a different harness and even if it is just legal, like we see some generalization in terms of like horizon length at which the LMs are able to operate at, but not a huge amount outside of that. And so it is a good point. And I think that's a very good, you know, heuristic to have in mind, which is that RL doesn't generalize. And I think you can make a lot of concrete predictions about what the big labs are going to do and even what you should do as an individual company based on that. I'm curious to hear, as the Frontier Labs do move deeper into areas like legal and finance and coding, how does that change the relationship between a model company and an application company?
46:22Charlie O'Neill:It starts to get very defensive. A lot of these really large application companies make big commits for spend on opening Anthropic. and you know they turn around one morning and realize that you know Anthropic is is not only training on like you know legal data or finance data or whatever it is they're going to actively go after those verticals and are going to offer products in those verticals and this has happened quite a few times and this is kind of the wake-up call that I was referring to earlier where it's not just about cost really for most of these application level companies like they'd probably be willing to keep spending large amounts on Anthropic and OpenAI if you know they were the best models and like they offer the best experience to the users it really is about oh okay I now need to have some sort of moat against not only my competitors in the space, but a general frontier intelligence company just coming out and offering a product.
47:05Charlie O'Neill:And so, I can't just be a wraparound that anymore. So, again, I've hit my 80 % plateau on what I can do around the harness and the prompting. The only differentiating factor I have compared to Anthropic is the feedback cycle and the continuous improvement loop that I own with my data. And what gets better in that feedback loop for say the Frontier Model Labs or even a company that's training their own models? Is it the judgment? Is it the evals? Where do you think the kind of biggest advantage or moat that they have is? Yeah, these loops run at different cadences. So, in a sense, like the Frontier Closed Labs are running a continuous improvement loop when you like step back and look at it.
47:45Charlie O'Neill:Like, you know, they will train a model, they will release it into the world. They will collect, you know, large amounts of feedback on what that model does well and what it doesn't. They will go off and design our own environments, like essentially perhaps those behaviors, and then they will release the next model. And so from a very high level, that actually does look like continuous loop that you're targeting. The specific application level companies can run that loop on a much tighter cadence and can point the fire hose directly at the things they care about, which is their advantage. And also, there's loops that the application companies have that aren't available for the frontier labs.
48:17Charlie O'Neill:Like, despite all the discussion around data and privacy and everything, Anthropic fundamentally doesn't really train on user data and like you know open ai probably used to be worse at this but like they're not doing the same either like that's very dangerous then to break any user conditions like those are all fairly transparent like they have to collect very generic high level feedback they literally go on to twitter and ask people like oh what are the models doing poorly what are they doing well whereas like you know if you're an application level company and like you know you've got users who signed up and agreed to give you feedback and like you can collect this this this very explicit and targeted feedback like that's a fantastic source of data that's the we're kind of seeing them realize this as well.
48:53Charlie O'Neill:Like, you know, McCaw has been very explicit about they are going to be buying literally companies, like companies that are not going as well purely for the sake of their data to turn it into RL environments because it's such a valuable source of data for the Frontier Labs. And, you know, like this will happen to an extent, but even if it does, like the way that the Frontier Labs train models, the improvements or diffusioning capabilities is averaged out across like all the things you care about. Like they can't afford to care about one thing. And so theoretically, like you should always be able to like beat the frontier labs by being very, very targeted and not ending up with this, you know, vanilla ice cream mixture of like all the general improvements that are coming from the frontier labs models.
49:31Do you foresee a world where countries will build sovereign data sets that become the trusted local data set, which is sort of the basis of proprietary models that customers might have and have access to? Do you see that playing
49:46Charlie O'Neill:out into the future i think the safest thing to do here is to for pre-training in particular like everyone kind of has access to everything so like there's just massive like the internet was not something that any one particular country or person invented like all of humanity has contributed to the internet and i don't think it's like you know fair or right or even good that like you know we have different splits of that internet and different countries have different levels of access and so on i think the pre-training mixture should essentially be okay we're going to trust humanity's development and like the general principles that the model comes out of pre-training with are going to be roughly the same everywhere.
50:17Charlie O'Neill:I do see an argument for, you know, like on top of not only the pre-training and mid-training and post-training that we do now, having very specific targeted like opinions about, again, how we want the model to behave, its values, its ethics and all and so on. The countries can inform and as these things get more and more commoditized and, you know, we're getting a slew of different open source models like out of the box that roughly all have the same pre-training base. There's a little bit of slight differences in post training and mid training but like you can be very very opinionated about like how you want to mold that model further and this is something that base labs is also doing and considering now is like taking all these open source models and doing what we call post post training which is where okay we've noticed not only capability gaps with some of the open source models but like you know we would like them to like answer in slightly different ways to like hard ethical questions or like things like that like these things are so multiple and that's the benefit of open source that like country will probably start to develop opinions on those and And like, even if it's still private institutions, like voting as a country and instilling your particular values into a model will be much easier as things get commoditized.
51:19Charlie O'Neill:So I would say general basis, but customizability on top. On that point of commoditization, if increasingly capable models and training tools are available to lots of companies, where does durable differentiation come from? I think it again, it's just that data loop, right? it's really cool to see like if you take let's say quen you know 27b the most recent quen 27b or like even some of the nematron models or like tinker's models like you look at those models and they are like tens of times better than gp4 like a closed model that remained the closed model for like essentially three years and you know that that pre-training base was open air's pre-training base for three years like people love that model and the things you can run out of a box on a Mac today are like significantly better than that.
52:05Charlie O'Neill:The commoditization is not only in terms of like how easy just to train, but also like how much compute and things like that you need to train a model. So yes, I think the only differentiating fact left is the data and the feedback loop she can instantiate. Just on that point of commoditization again, you know, what was something that wasn't a commodity in the stack two years ago that now is and what do you think will sort of continue to follow that path and what won't? yeah I think it's what we alluded to earlier I think post training in particular felt like you know it was really really complex to post train even 12 months ago anything above a trillion you know a trillion parameters now like you look at how many open source libraries you can use to train a trillion plus size parameter models and it's like such a difference so to me that's the biggest thing yeah and we're moving back down the stack as well like you know there are open source libraries out there for posturing now like we've been able to write our own you know library from from scratch to do this stuff and I've been very opinionated about that.
52:59Charlie O'Neill:Like we're realistically living in the world where if you can get access to, you know, a few hundred GB 200s or whatever, like a team of five to ten people could train something significantly better than again, GPT-40 or whatever it is. And like, you know, again, like there's not that much value in training something that's like a bit better than GPT-40, but you just kind of extrapolate that path out and like where's the where's the durable advantage if, you know, in three years, someone with you know 50 ,000 bucks could train an opus 4.8 like it's pretty cool like like the all signs point to like the recipe has stayed stable labs are betting on scaling that recipe as they continue to scale as compute gets better as our techniques get better algorithm make a data progress like anyone can train you know the frontier model from a year ago for like orders of magnitude less and like that's bleeding into post training now because post training has now gotten to the point where it's like very very cheap and very very easy to do and I think that will continue to work backwards across the stack.
53:52Charlie O'Neill:So you're just going to see this explosion, this kind of plethora of open source models that have been incredibly customized in really cool, weird ways. Where does base 10 fit into that stack? How is your strategy emerging as all this is changing? Yeah, I think we're just betting on this ecosystem emerging. There's so many different ways to customize a model at so many different points, whether that's post-training or earlier in the training process. And we know that so many people, like there's going to be a lot of model trainers just to start off with, like trying to give people general models out of the box.
54:19Charlie O'Neill:we're betting on that we're betting on individual companies specializing those models for very very targeted things we're betting on like these systems being glued together with routers and like all these complex like meta harnesses on top and like i think that as the model's capabilities improve it's the fundamental bet is like people are going to want to use those models and so hence they're going to have to run inference and i guess that's the bet based on making i'm curious to hear your reflections on you know how does someone build a company when it's not entirely clear which layer of that emerging stack will ultimately capture the value.
54:52Charlie O'Neill:I think the best possible heuristic here is that intelligence is going to keep changing. And so the things that you should be doing to offer value to everyone is going to keep changing. And so you should be trying to build a business around something that the models can't currently do, but the next iteration or the next few iterations can do. I think that's a really cool strategy. I think there's going to be much smaller companies that are able to pivot and adapt up to like that changing intelligence faster because you know, even if you build a company around something that GPD M plus one, it's going to be able to do, but the models today can't.
55:23Charlie O'Neill:Like once we get to GPD M plus one, you know, you're gonna have competitors spring up. You're gonna have to face this problem of like, okay, now I have to not only like set up my harness and prompting, I have to do my post training and then set up my data loops. And I think we are going to see like just very rapid, nimble companies arise that are able to like kind of extrapolate out the scaling laws and figure out what they're gonna be able to do. and be ambitious enough to go after like better and bigger things with each like iteration of capability advancement how much more certain are you of that reality compared to say a year or two ago i think much more so from the perspective of like the open source models were slightly ahead of schedule there's like also a lot of like you know path dependent stuff that that occurs here like you know if anthropic and open al or the closed source labs like developed a big enough lead like it might have been difficult to even like have a crack at like releasing this open source stuff and hence like this would have like had these network effects and like economies of scale before open source could have really got started but I think like touch wood at least for a while I think we're past like the critical mass or like the warm starting we needed for open source to like kick off and so I'm really excited for those things to like thrive and coexist together I think that's like a fairly good bet to me.
56:33Tell us about BaseLabs.
56:35Charlie O'Neill:Yeah so we've obviously believed in open source for a long time back from my days of like the first time I fine-tuned GPD two I was like okay if you can teach a model a task and this is available for people to do so like this is going to change the world and this shouldn't just be a thing that you know a few thousand people can dictate we made that bet with Paused we've obviously been making with base 10 we're trying to enable individual customers to you know train their own models and serve their own models and like own their own intelligence and and for that to be a very democratized thing I think this has been very much a journey of like getting more and more ambitious just about how impactful and helpful we can be in that regard.
57:13Charlie O'Neill:The next stage of that is base labs. And the point of base labs is essentially, can we be as helpful as possible in not only like propping up the open source ecosystem, but helping it to thrive? And so what does that look like? It looks like a lot of different things over the like, you know, short to long term. In the short term, like we would really like to help the American labs, strengthen them in terms of their open source efforts. So we're doing some, a lot of the post training for, for Nivatron models, we're kind of collaborating with a lot of like the American open source people and like you know taking the lessons learned from both the frontier labs and the you know the Chinese open source labs and like getting America up to scratch and I think that's going to be a very rapid process in the medium term it means thinking about like what are going to be the constraints and the bottlenecks that open source spaces compared to closed source and so we talked before about okay can we provide like really large volumes of high quality reinforcement learning environments based on everything we've learned from seeing what people are actually using these models for and where that breaks down so that, you know, an American open source lab doesn't have to pay billions to record to get like the same advancement.
58:11Charlie O'Neill:And in the long term, it's going to be a lot of blue sky research. And we are very lucky to have a fantastic research team looking at so many different things about like not only just how to train these models, but like the ecosystem of tooling and stuff that exists around it. So for instance, these models at the moment have like context windows of a million tokens. You can do a lot of things with a million tokens and you can summarize the end of your session and keep going. But like, what would it look like if we could have like billion token context windows? If we had a summarization method, not in words, but like, you know, in the latent space of the models that could effectively give you a billion tokens of the context, what new possibilities that open up, particularly if like anyone can plug that into an open source model.
58:48Charlie O'Neill:So it's going to be a lot of blue sky research. And again, a priori, it's very difficult to predict what that blue sky research will be. But the broad aim is like, we're here to be as helpful as possible with as little commercial general as possible to the open source ecosystem. Yeah. And I think part of the thinking here is that when we try and think who is best positioned to do this blue sky research, it kind of comes down to incentives and who actually has the long-term funding cycle, for example, to be able to take on these projects. And, you know, we think about entities like NVIDIA or Microsoft who are obviously very well positioned.
59:19BaseLabs is a smaller kind of more agile team. But the fact that there is like this inference business that helps to kind of support this kind of research effort to us is really important. And when we look around at kind of other open source efforts, like that's a question that we kind of ask. And I think that's why we're so excited by it. You guys mentioned earlier optimizing for people who are high slope, not high intercept. What are you essentially looking for in research talent that you think conventional hiring tends to miss?
59:44Charlie O'Neill:I think it's clear that like LLMs would move so fast like the whole history of the world like when we're hiring someone for a job we're essentially asking like do they have the right experience and knowledge to like do this job properly so that I don't have to spend organizational resources like getting them to do this job and of course to some extent you're always going to have to spend those resources because like it's why we get degrees it's why you know there is a notion of this like corporate hierarchy that you move up as you'd learn more things about your particular domain I think with LLMs we can kind of throw that assumption out the window like even if you have previous research experience like these are fundamentally new objects like traditional machine learning and deep learning a lot of it doesn't necessarily apply and now with lms themselves we're moving up levels of abstraction in the stack such that doesn't necessarily matter what incredibly niche knowledge you have anymore you actually have something in a box to tell you all the knowledge that you need to know you need to be smart enough and curious enough and driven enough to be able to like go figure it out yourself and so the safest thing to do with the world moving so quickly is optimize for picking those people rather than like indexing too highly on experience or credentials.
1:00:42Yeah and high slope yes it's intrinsic to the individual but then you also have to foster the right environment and I think yeah having a team that is very happy to share and like be in person and there is like so much implicit learning that happens when you're in when you're in the same room I think has been really really valuable because yeah being high slope I think everyone says that it's like such conventional wisdom in some ways but to actually do that properly I do think a lot of it does come down to how you're on the team.
1:01:08Charlie O'Neill:Yeah and the network The work effects of this are like so incredible and rewarding to see. Like when Woody and I first hired like our first couple of people like it was very much our job to teach them everything. You know, these were people from quantum trading and I could never touch LLMs or reinforcement learning before. So it was a very manual process and that in and of itself was rewarding. But now we've got like, you know, a team of tens of people and all of those people, depending on when they've come on, have scaled up like rapidly in whatever they're working on. and have like learnt all this stuff and now, you know, learning things that we don't know about and like watching like all those individual relationships and how people are teaching each other.
1:01:43Charlie O'Neill:It's just like amazing. It feels like the snowball is like rolling down the hill and it's just like picking up momentum and it's really cool to see. You mentioned you're hiring Aussies and effectively bringing them over to SF. What signals are you looking for on their CV or how are you able to screen for folks who are curious and high slope? I think the short answer is high frequency trading. But no, I guess I touched on this earlier, but I really do think it is quite sad that in Australia, if you are really talented and you're finishing year 12, you know, where do you go? You go into, you know, law or medicine or you kind of, maybe some of you go into computer science and math.
1:02:21And then the jobs that exist after that, you know, they don't end up being kind of frontier AI research in Australia. are and yeah I guess like the reason I talk about that from my school like when people are thinking about university degrees particularly you know at a selective school people weren't thinking about CS and math and then even the people that do go into CS and math end up at quant trading or go overseas and so yeah when we think about talent we do look a lot at via people that are in high frequency trading and then I think the other thing that we really look for is curiosity so much of kind of our interview process does hinge on this it's this curiosity to kind of try and solve this unsolvable problem or like you see something and it kind of sits with you until you until you figure it out and then also being really engaged with what's happening in the field and even if someone is like high slope and you know very intelligent and quantitatively amazing if they don't have the curiosity to be at the frontier I think for us they end up not being able to
1:03:12Charlie O'Neill:yeah to stay with what we do yeah I we joke about like you know high frequency trading it's a very good domain to get people from there's a lot of analogous skills and transfer like you You know, if you're very good at hill climbing, like how a particular trading strategy is going to do over time, you're probably very good at hill climbing in the sense of reinforcement learning. But one thing that I've noticed about particularly, like, you know, the first 10 that we hired for Parz when we're back at Parz is like every single one of these like people we'd gotten from, whether it was quant trading or consulting or whatever, had some sort of like missing gap in like what they felt like was like definitely broken or whatever.
1:03:46Charlie O'Neill:Just slightly broken. Like a lot of them had like quit their quant trading jobs with no like no job to replace it. like they were like I just can't do this like I feel like I'm missing something like these problems are obviously like from a very granular level incredibly interesting and very like intellectually demanding but there is something else missing for me and so whether you call that mission whether you call it curiosity or whatever it is like it wasn't like we just went in like convince these corn traders to like transfer from you know one HFT firm to like an AI research lab it was very much these people who were really really really good at that stuff but had this like you know missing drive to do something greater than like you know move ones and zeros around and like what ended up being a very like kind of you know arbitrary game like I know a lot of my HFT friends are gonna hate me for saying that but I've done a little bit in there and I believe it's a bit of an arbitrary game.
1:04:31Fair enough you guys are you have a really lean team but from the outside looking in your productivity is kind of off the charts like how do you guys work like what do you think is different about the way you know base 10 and also your kind of post-training team works that makes you so exceptional but also allows you to just produce so much research?
1:04:50Charlie O'Neill:I think I was really lucky going through undergrad to have a lot of really good research mentors that were not necessarily the traditional academics and like follow the traditional academic process. And I think the one lesson I took out of all that experience and all that mentorship was when you're doing something like what is your information gain per unit time? So there's like, there's often two different types of people. And like, this is a course of spectrum, but it is easy to classify people all the time into 80-20 people and 99-1 people. And so like, I would probably consider myself more of an 80-20 person.
1:05:20Charlie O'Neill:Like I'm always like, okay, what, how do I get to like, you know, the 80 % of things, which is going to give me a noisy, but directionally correct signal on whatever it is I'm doing. And how do I get as rapid feedback process as possible and like reduce the time, but just between discrete like information points entering like my system. So I can make new decisions based with that. Of course, you can't always be like that. Like sometimes you need to be nice for that one or a hundred zero. You need to get something to perfection. and so for our team like being really explicit about identifying who is which type of person and then pushing them to be a little bit more that other type of person when they need to be and I think that framing like it sounds really obvious and it sounds like a very obvious heuristic but when you really think about like how do I maximize my information getting period of time like that's all you really need to know like sometimes the best optimal thing to do will be spend a lot of time to get a lot of information and sometimes it will be spend very little time to get very quick information and like that's been a kind of gunning principle of not only our research team but our engineering team and so on and like being very explicit about those trade-offs allows you to make the right decisions about when to move fast to break things and when to be slow and make sure it's built correctly yeah this is so like throughout like the years of phd mentorship that i got from my supervisor i still did not like internalize this as much as like charlie's just beating it over my head with and it is so true and it exists both from like you know blue sky research it's like what experiments can you run to get this information gained per unit time to be really high but then it also exists in custom work it's like what are the experiments you can run really quickly to be able to get the customer some information so that you can actually continue to move forward.
1:06:47In research engineering, we see it a lot with like prototyping and kind of exploring what is at the boundary of, you know, the particular work that we're doing before we go make this like very technical PRD that actually needs all the kind of planning and extra work that comes with it. And so, yeah, I think this is like one of Charlie's most fundamental principles that, you know, it sounds obvious, but like until I've internalized it, I like, yeah, it took a while, But once I was there, I was like, it is so, so valuable. If I gave both of you a model tomorrow that could execute every single experiment you could describe perfectly, what would remain the job of a great researcher?
1:07:21Charlie O'Neill:My maybe slightly contrarian answer to this is that models exist which can do that. I do believe the frontier models, both closed source and open source, are essentially capable of if I know all the assumptions it's making, it knows all the assumptions I'm making about an experiment, it is able to execute that experiment together to completion. like I don't see as many mistakes being made as long as the instruction is made explicitly anymore we are stepping up the level of abstraction in terms of research like we've built an internal tool which is essentially like allowing researchers to iterate as fast as possible it has very strict guardrails about like you know how much effort you have to put into planning an experiment and making sure the model knows exactly what you want but then it is very much able of like autonomously going off and executing that experiment and bringing the results back to you and making sure everything is visible in one place so and to me like this has always been the dream of research is like I don't have to waste time with the plumbing.
1:08:09Charlie O'Neill:I don't have to waste time with like, you know, the aggregation of data by like manually like coding stuff to collect it and pull it all together. A researcher's time should be spent figuring out what your prior is in terms of like what you're testing, figuring out what hypothesis you're making and then doing the best possible update to your posterior based on the information you're getting. And as long as that information is correct. And I think the main point here is the ELMs allow you to get correct information now. And so they will continue to get better. But I think one important thing to say about LLens is like whilst they're like task horizon is now getting increasingly longer.
1:08:42Charlie O'Neill:I would actually make the argument that what I call this slop horizon has not actually got that much longer. Like I can't trust its taste over a period of more than like maybe five minutes to 30 minutes to not diverge from significantly from what I would have decided if I haven't specified everything explicitly. So I think it's really interesting with RL. Models are going to get better and better and run for longer and longer. at doing things that you very explicitly defined even those things very hard but i haven't seen a huge jump in the ability to you know align or calibrate with your taste and some of these are just fundamental reasons right like like it just doesn't have that information it's not like making the model smarter is going to be able to like get that information out of my head but like interestingly while i'm super super bullish on like how much ellen's are going to accelerate research and stuff like i think like that last step is like we haven't actually made that much progress on it so yeah a good researcher will always be able to like make the correct update to the right posterior and thus suggest the next hypothesis about what they want to test and come up with the right ideas in the first place of like cool things to test.
1:09:38I think like something that we've talked a lot about is like yeah the greatest researchers have run just the most experiments and it is correlated and that lets you make these updates and like choose the problem space better and I don't know if you want to comment on that.
1:09:49Charlie O'Neill:This is probably another principle which is like I think the reason Noam Shazir who invented a mixture of experts and like so many other things that have been really important to LLMs like every single person you talk to who was at DeepMind with him at the time, simply says that GNOME ran an order of magnitude more experiments than everyone else. And I think that option is now available to a lot of people. You don't have to go and learn the intricate details of Slurm or whatever to run all these compute jobs. You kind of had that friction being taken away from you. So now your job is, yeah, can I be the person who runs the most experiments and making sure that you're not being so lazy that you're trusting the LN to decide what to do based on that.
1:10:23Charlie O'Neill:You're running as many experiments as possible and your only job is to decide what to do next. Both of you have had pretty nonlinear pathways into the world of AI research. I'm curious to hear how those experiences have shaped the work that you do now. Maybe, Morty, you trained as a doctor originally. What does having practiced medicine make you notice about AI that you might have otherwise missed? Yeah, I was eight years late was one thing. But no, I think like you kind of look back and it's like, you know, you actually can patch together the story. During medicine, it was kind of doing stuff, being self-taught and kind of running projects to the side.
1:10:59Then there was a bunch of computer vision stuff with like MRI analysis and that sort of work. And then obviously going to Oxford and doing the PhD or starting the PhD. That kind of allowed for some formalization of what I had done. And it was a bunch of time series work and then moving to LLMs. And then when we started PARS, it was all around mechinterp and regulated industry. So, it became very relevant. And then, you know, turns out that doing mechinterp and post-training, you know, generalizes beyond just healthcare data. So, yeah, I think looking back, it's kind of easy to see this story. But how does it still contribute?
1:11:29I think, yeah, the biggest thing to me now is that period of time kind of shaped, I guess, like how we interact with the team a lot. and yeah it's like a softer set of skills but it does really feel like yeah the time in medicine sure it teaches you a way to think sure it teaches you a way to learn for example but yeah i think when we're running a team now it actually is really really helpful and what about yourself
1:11:51Charlie O'Neill:charlie yeah i mean i had a similar path like i haven't been working on lm stroller i started in law and philosophy wanted to do english literature like i loved like i guess writing is probably the best way to to describe it writing and reading and then i remember i was based on some philosophy stuff was reading a guy called Gwern and Gwern has actually become very like quite famous in the SF scene but he wrote this blog post on GPT2 and I remember like thinking like how magical it was to have this like kind of machine that could output human language and like you know Chomsky and all these others have always said that like English was way too infinitely complex for a machine to be able to ever be able to model it and so like seeing this was was like fundamentally like wow and I went back and found the original GPT paper and kind of looked at the difference between the two I was like wow it got a lot better and I kind of just like drew the line through the gpt's and thought okay if if this keeps up like this is all anyone's going to be talking about in like five ten years time and so i swapped from law and philosophy to maths and computer science and was like i'm just going to focus on lms and i think the big thing for me that i've taken across all that even though i dove into lm research like relatively early back in like the gpt2 days was this isn't about like having this like you know ioi quantitative like very very pure intelligence that's not going to be rewarded by like working on these things what is going to be rewarded is like these things are like beautifully subjective and like very broken in fundamental ways but like very very similar to humans in other ways and the most important thing in like as these things get more important is going to be like that qualitative intelligence like much more closer to like you know being a very good reader or writer to be able to work with them and like the meta intelligence to know how to craft these systems that aren't perfectly deterministic, that aren't like, you know, these black box functions that the rest of computer science for the rest of history has looked like.
1:13:36Charlie O'Neill:And like that really appealed to me. I don't think maths and so on had appealed to me that much before. Like I kind of enjoyed it. But like when it's this very clean, like, you know, quantitative algorithmic process, it's like not as exciting as this like almost like living, breathing organism that you guys are shaped and is going to have a very big impact on the world in very subjective ways. So I think that's probably the main thing I've brought over from that. And then when or if AGI arrives, what do you think you guys will want to do? Like, Modi, for example, do you think you'd go back to practicing medicine?
1:14:05We'd love to hear your reflections on like this kind of post-AGI world. No, I'm going to open up a cafe in the European mountains. And Charlie will be reading that. I will be reading that. He makes very good coffee. Nice. Yeah, I do think we're so far away from that. I guess like something that we talk a lot about is kind of, yeah, even if we do get to this point of AGI, whether we're there now, you know, how long does it actually take to diffuse into these jobs? Particularly in medicine, it does feel like we're so far away, but lots of knowledge work, even like, you know, we work a lot with legal AI companies.
1:14:35And it's like, I do think that there is going to be a human life for, you know, at least a decade, if not longer, because there is this human centeredness as part of professions that will always be really valuable. And I guess like commentary that we've heard about this before that I kind of really like is, you know, what is the willingness to pay up until the point that you just have the human doing the bit that you want? And I think there's like in every job that I can think of, particularly knowledge work, there is still some element that we will always want some person or some human behind that part of the role.
1:15:02Charlie O'Neill:Yeah. I think my old answer to this would have been reading books at Moody's Cafe. And I do agree with Moody. I think we're a long way off AGI in the traditional sense of how people define it of like the LLMs doing everything. I've actually kind of changed my mind on what society looks like, like even once that process hits. I think that regardless of how smart the LLMs are like we will never trust like the taste function or the value function of an LLM over ourselves and so I think even once we have an LLM which is better than humans at like doing any possible thing that you could explicitly tell it to do I think like there's going to be a massive demand for like the smartest you know philosophers, ethicists, political people like just people who can think very clearly systems level thinking opportunity cost about where to point the big supercomputer.
1:15:47Charlie O'Neill:Like we are still going to have limited and find that compute. You know, do you point that big supercomputer for a few years at solving cancer or do you point it at climate change? And like, obviously, that's a very like naive way to put it. But, you know, like we are going to have to figure out how to allocate the compute and the resources and like what trade offs we want to make as a society for how to solve these problems with things. So that would always be my advice as well to like, you know, someone starting university now is like don't necessarily get into computer science, don't necessarily get into training.
1:16:13Charlie O'Neill:like our job is hopefully to commoditize that so our jobs don't exist in like five to ten years time i would really think about like how do i think about like where humanity is going as a whole and like what can i study to make me help me make better decisions about you know how to best help people with this incredibly powerful tool i think there's gonna be a lot of demand for that for a long time so if i were to restate that the most scarce thing when intelligence becomes too cheap to meter is effectively taste and human judgment and that sort of interface layer. And then our final question for all of the Wild Hearts guests is what is something you've changed your mind about?
1:16:46Charlie O'Neill:I think for me it's like I probably oscillated between this and I'm just in my most recent oscillation but I think that the world is actually going to value domain experts much more than people like Modi and I and that kind of goes back to my last answer as well. Again LLMs have removed a lot of the friction to like implementing ideas to having like the technical knowledge for computer science or whatever it is to be able to build out and test ideas very quickly. I maybe originally thought that like LLMs were going to get smart enough such that like you just kind of needed to you know have this like general high slope or intelligence or horsepower to be able to like use them and you could run right in any domain and I think that was fairly hubristic like one thing I've changed my mind is because that friction has been removed and like you know my skills and like you know setting up big training jobs with slurm or whatever, like kind of negated by LLMs, like the domain expert who's trying to solve like, you know, rice petty disease classification who has like years of like understanding that field and like what's happened, what's changed, like how to collect data, all that sort of stuff, who's always going to beat me if they're going after that problem.
1:17:46Charlie O'Neill:And so I think that's actually a really exciting place to be. And like, as I said earlier, the world's going to have a lot of demand for people who can make, you know, very rational decisions about opportunity costs and trade-offs when we're solving particular problems. And that goes all the way down to the granular level of like even within this particular problem we're going to solve like how am I going to get the alum to solve it and so like I think it's a really exciting time to be someone who really deeply cares about a specific thing even if that has nothing to do with AI because AI just lets you solve that problem better.
1:18:12Amazing thank you both so much this has been fantastic.
1:18:15Charlie O'Neill:Thanks Fabulous.
1:18:19Thank you so much for joining us for another episode of Wild Hearts. If you want to learn more from other ambitious people building designing and creating the world that we all want to live in, then please hit the subscribe and follow button. This podcast is produced by Camilla Herring from Blackbird. Our marketing genius is Laura Cofford, and our editor is Andy Jones from Colour and Sound Creative. Thank you all so much for listening, and I'll talk to you next week.
From the publisher
This week Kate hands over the mic to Saron Berhane, who leads Blackbird's deep tech programme. She was recently in San Francisco and sat down with two Australians who had to move to the Bay Area to do the sort of work they wanted to do. Which is both their personal experience and one of their main arguments in this episode. Australia's problem in AI isn't compute. It's that the smartest people here go into quant trading or get on a plane, and nobody has built them a reason to stay.
Baseten is where they landed. It's an AI infrastructure and production-grade inference platform, the layer that lets engineering teams deploy, scale and optimise machine learning models. Sydneysider Tuhin Srivastava co-founded it, and it recently raised a$1.5B Series F led by Altimeter Capital, Conviction Partners, Blackbird and Spark Capital.
Mudith Jayasekara and Charlie O'Neill are Baseten's Co-Heads of Model Training. Before that they founded Parsed, a mechanistic interpretability lab that Baseten acquired. Mudith trained as a doctor. Charlie started in law and philosophy and wanted to study English literature. Neither took the path you'd draw if you were designing someone to run model training at a company like this, and both think that's the point.
The conversation is wide-ranging and deeply technical. Saron takes them through what's actually happening inside the open source ecosystem right now, which is stranger and more hopeful than the usual framing allows. Sitting underneath all of it is one question. When everything difficult gets commoditised, what's left? The training recipe is now roughly the same everywhere. Post-training has gone from a specialist problem to something a small team can attempt. The friction on running an experiment has mostly gone. What's left is taste, and having enough of the right people in one room.




