In short
The “ChatGPT moment” for agentic AI, focusing on OpenClaw’s viral consumer adoption, multi-agent interaction/network effects, and enterprise implications for safety, privacy, benchmarking, openness, and sovereignty.
Guest
Joelle Pinault, former VP of AI research at Meta (FAIR), now Chief AI Officer at Cohere. Background includes leading FAIR’s growth from ~100 to ~600 people, building multilingual foundation models, and advocating for open research/open weights.
Key claims
Agent platforms are going viral with many independently created agents; this reveals new safety risks (prompt injection) and new emergent behaviors from agent-to-agent “social network” interactions. Enterprise deployment lags consumer adversarial risks but has higher expectations for privacy/security. Benchmarking must combine sanity checks, progress-driving benchmarks, and human experience; diversity in teams improves evaluation questions.
Notable examples
OpenClaw agents using Moldbook (Reddit-like social platform) to get answers and complete tasks; a coffee-ordering wearable use case; helicopter cockpit speech recognition test where she validated performance for a female pilot.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOAI Research: Academia vs. Industry
4:21 to 6:10
Explore the differences between AI research in academia and industry.
“So I want to first start by talking about the state of AI research, both in academia and industry.”
Lessons Learned from Meta
6:10 to 7:37
Joelle shares valuable lessons from her experience at Meta.
“Sometimes some of the insights come from much more, you know, small teams working really diligently on a specific problem.”
Cohere's Mission and Agentic AI
7:37 to 10:45
Understand Cohere's dual mission in AI and agentic platforms.
“I think it's about clarity of the objective.”
Discussion on OpenClaw and Agentic AI
12:39 to 14:02
Joelle discusses OpenClaw, its viral appeal, and security aspects.
“That's A-T-L-A-S-S-I-A-N dot com slash TeamChanger.”
Introduction to OpenClaw Tool and Its Use Cases
14:02 to 15:10
Explore the new OpenClaw tool and its viral applications.
“as many other people did looking into into that same same exactly so for for maybe some of our listeners who are not familiar with it this tool came out and it's basically an agentic kind of Yeah.”
The Viral Appeal and Safety Concerns of AI Agents
15:10 to 16:32
Discuss the popularity of agentic AI and associated safety vulnerabilities.
“when it is clear that there's an immense appetite for these agents.”
Interactions Among AI Agents: A New Social Dynamic
16:32 to 18:28
Examine how AI agents interact and the implications of their relationships.
“And in many cases, you don't know the creative ways in which, you know, people may be using these agents for nefarious purposes until you really deploy that scale.”
The Importance of Openness in AI Development
18:28 to 21:01
Investigate the significance of openness in AI models and Cohere's approach.
“And from this, what we can learn in terms of some of the essentially like social norms we need to build into these networks so that we don't have too many pathological behaviors.”
Risk Assessment in AI: Personal Experience
22:41 to 23:56
Learn about risk assessment in AI through a personal anecdote.
“and how Joelle found herself in the cockpit of a helicopter.”
Benchmarking AI Models: Challenges and Strategies
23:56 to 27:20
Understand the complexities in benchmarking AI models and their evaluations.
“So I'm just curious, what's your approach to thinking about how we benchmark these models?”
Show all 17 chapters
The Role of Diversity in AI Evaluation
27:20 to 28:00
Discuss how to incorporate diverse perspectives in AI model evaluations.
“But how do you create an enterprise benchmark?”
Importance of Evaluation in AI Processes
28:00 to 28:35
Learn why a clear definition of success is crucial for AI evaluation.
“then that's usually a good one to build an evaluation out of.”
The Value of Diverse Perspectives in AI Development
28:36 to 31:04
Explore how diverse teams can enhance AI model evaluations and applications.
“Do you have a specific point of view on how to bring diverse perspectives, I guess, when thinking about these models?”
Geopolitical Aspects of AI Sovereignty
31:05 to 32:50
Understand the rising demand for AI sovereignty among nations and companies.
“So how does that translate to what the team looks like on the ground?”
Emergence of Hybrid AI-Human Organizations
32:51 to 37:18
Discuss the future of hybrid AI-human teams and the cultural implications.
“It really opens up the perspectives in terms of the work.”
Evaluating AI's Measurable Impact on Organizations
37:19 to 38:38
Examine perceptions of AI value and the challenges in measuring productivity.
“Yeah, there's this MIT study that found that 95 % of organizations see no measurable return on their investment in AI.”
Personal Integration of AI in Daily Life
38:44 to 39:44
Discover how individuals and families are increasingly relying on AI.
“What's one way you're implementing AI in your everyday life that you know you can't live without?”
Transcript
Automatic transcript. May contain errors.0:00On Pioneers of AI, we often talk about the responsibility that comes with building powerful technology. AI isn't just one thing. It's a force reshaping medicine, education, science, and the stories we tell on screen. And for that future to work, it needs infrastructure that's built thoughtfully and deliberately. As the essential cloud for AI, CoreWeave empowers pioneers at leading AI labs and enterprises with a purpose-built cloud made specifically to train and run the world's most advanced models. They help accelerate the breakthroughs that matter at scale. The big ideas, the moonshots, the futures we haven't imagined yet.
0:41Ready for anything, ready for AI. To learn more about how CoreWeave powers the world's best AI, go to coreweave.com slash ready for anything. When it comes to AI, we spend a lot of time talking about what it takes to move models from research into reality. Enter LTX2 from Lightrix. It's an open-source audio-video foundation model built for synchronized sound and video, native 4K output, expressive motion, and precise multimodal control, all running on consumer GPUs. With over 3 million downloads on Hugging Face, LTX2 gives you full model weights, training frameworks, and evaluation tools to build real production workflows.
1:27Don't settle for a first draft. Get the complete creative engine to reach your idea's full potential. Try LTX2 today at ltx.io slash model. Hi, everyone. It's Rana. The future of AI is not built in isolation. It is shaped in rooms where ambitious people challenge each other. Masters of Scale Summit is back October 20th through 22nd in San Francisco. bringing together founders and innovators pushing the edge of what is possible. If you're building what comes next, you should be in this room. Apply now at mastersofscale.com slash pioneers. That's mastersofscale.com slash pioneers.
2:17We build agents thinking they're going to be unique. They're going to be deployed and they're going to go out, interact with a stationary version of the world. Now, suddenly we see what happens when you deploy against all these other agents. They start interacting, they start negotiating, they start trying to offload some of their work, all these other interesting behaviors. So what is going to happen in this case with all these LLM agents talking to each other and exchanging information, deploying their skills, using tools? I think that sandbox and the ability to look at the network effects across agents is going to be super interesting.
2:57That's Joelle Pinault, former VP of AI research at Meta and current chief AI officer at Cohere. She's referring to OpenClaw. It's the new AI assistant everyone is talking about. OpenClaw has taken the world by... Claw is generally the best personal AI assistant that I've ever used. This became the fastest growing open source project in GitHub history. The best technology I've ever used in my life. And by far the best application. In the way, probably 80 % of the apps that you have on your phone. It's like I'm checking the world 24 hours a day. 100 ,000 GitHub stars in three days. Everyone's losing their minds.
3:33Well, let's bring some sanity to this moment. As a leader with years of experience developing open source foundation models, Joelle is an expert on the matter. Sure, there's a lot of possibility here, but there's also a lot of risk. So let's dig into it. Joelle Pinault joins me for this week's episode to look at the future of AI agents. We'll also talk about why she's an open source advocate and how Cohere is moving the dial on developing foundation models outside the US and China. I'm Rana El-Khalyubi, and this is Pioneers of AI, a podcast taking you behind the scenes of the AI Revolution.
4:20Hi, Joelle. Welcome to Pioneers of AI. Hello. Great to be here. I'm so excited for our conversation. So I want to first start by talking about the state of AI research, both in academia and industry. So you got your start in academia. You did AI research there, and then you decided to transition to industry. And that was kind of your start at Meta. I'm curious, how is research different in academia versus industry, especially as it relates to AI? I decided to make the jump in 2017 because already we were seeing a divergence in terms of the type of research questions that you could take on. Academia is wonderful in that you have this influx of new students year after year, and that is just such a source of genuine curiosity.
5:07So I I find usually, you know, they ask wonderfully different questions, which really stimulates the exploration and the curiosity. On the other hand, when you're doing research in academia, you know, you do have to build in a lot of time for the learning process and the development of these young talents. The teams tend to be smaller, more junior, where on the industry side, you can put together larger teams with a mix of engineers, scientists, designers, user research folks. you get these bigger and multifunctional teams that just allow you to tackle bigger problems or do bigger projects. And so fast forward to today, I would say there's some of these same differences, but magnified significantly by the amount of investment that is going into AI in industry, whether it's compute, whether it's data, whether it's talent, the size of the team, the ambition of the projects is just on a very different level.
6:04Well, that being said, you know, like research doesn't always happen at scale. Sometimes some of the insights come from much more, you know, small teams working really diligently on a specific problem. So I think there's still a role for both, but very different mode of operating. Okay, so you were VP of AI research at Meta until quite recently, and then you made the transition to chief AI officer at Cahere. What lessons did you take with you from Meta into your new role? And what does your new role encompass? Interesting. I think I'm still digesting some of these lessons. You know, sometimes you don't know that you've learned something until suddenly you're in a context where you can apply those lessons.
6:49It was an immense learning opportunity for me to be at Meta at the time that I was. The FAIR, the Fundamental AI Research Team, I was leading, grew, you know, from maybe 100 people to 600 people or so. And so, you know, one of the things that I developed over the years is the ability to push teams to have ambitious projects, ambitious roadmaps, and a clear line of sights to a concrete goal. And so coming into Cohere, you know, bringing some of that culture of how do we work as teams, maybe the teams are smaller, smaller companies, smaller research team. But nonetheless, how do you set up the team in a way to tackle a goal together?
7:33That's one of the things that has been quite helpful. Great. So what's the secret? How do you do that? I think it's about clarity of the objective. Like, do you have a clear sense of whether you succeed or not? Having clarity of that North Star in a way that's measurable helps really align people towards a goal. Now you have to make sure it's a goal that's worth doing and that people are excited that they have a path to be successful, they have the right resources, all these other things. But it's a lot easier to get buy-in of people, of the research team, leadership, and so on once you have that clarity of goal.
8:10Yeah. I do want to take a step back and kind of give you an opportunity to share what does Cohere do. So you guys are one of the few players outside of China and U.S. that are building your own foundation models. But beyond that, what is Cohere's mission? There's two core pillars. One of them is we do build our own models. And the other one is this agentic AI platform that we build, which is really the core of the product strategy, I would say. So we have North, an agentic AI platform. It's designed to ingest the models that our modeling team is doing and deploy that so that we can, you know, bring that into different companies and governments who have the ability to benefit from AI agents, coordinate these agents to build automations, and at the end of the day, bring a lot more productivity to their work.
9:02And you can think of the models as the engine, and you can think of North as the vehicle to essentially deliver AI. And there's some places that are ready to ingest like a raw model. So in some cases, they're models that are very good, for example, in terms of their multilingual capabilities, in terms of their reasoning capabilities. But actually, more and more, we're seeing a lot of demand for having that engine be built into a vehicle that can deploy AI. So that means, you know, there's a UI, your end users can actually, you know, chat with the model, but they can also build their own agents.
9:38They can exchange agents across a team. We can build automations that orchestrate multiple of these agents. and building up a much more cohesive experience for how you bring AI into your workflow. And I would say the other really important distinguishing aspect of what we do is to deploy on-premise with very high security and privacy guarantees. So this isn't just like an API or a cloud thing. This is something that you can run locally with your own data. So you can have a conversation or your agents can see actually all the internal data from the company without that data coming all the way to cohere.
10:17But locally deployed, it means that you have the ability to leverage all the business intelligence in your data and in your systems internally and do that in a way that's secure and private. So especially in regulated markets, financial markets, health care, government services, you do have obligations to give strong privacy and security guarantees. And so there's a lot of potential for that and a lot of interest. In a minute, I get Joelle's take on OpenClaw. That's the new open source AI assistant that everyone's been talking about. We'll get to that and more after a short break.
11:13When it comes to AI, we spend a lot of time talking about what it takes to move models from research into reality. Enter LTX2 from Lightrix. It's an open-source audio-video foundation model built for synchronized sound and video, native 4K output, expressive motion, and precise multimodal control, all running on consumer GPUs. With over 3 million downloads on Hugging Face, LTX2 gives you full model weights, training frameworks, and evaluation tools to build real production workflows. Don't settle for a first draft. Get the complete creative engine to reach your idea's full potential. Try LTX2 today at ltx.io slash model.
12:01If you've spent any time building AI products or leading technical teams, you know this. Transformation doesn't fail because of ideas. It fails because teams can't move together. Enter Atlassian's teamwork collection. It has planning in JIRA, documentation in Confluence, video updates in Loom, and now AI agents in Rovo, which connects the dots across your work so nothing gets lost. It's one AI-powered teamwork platform designed for how modern teams actually build. Learn more at Atlassian.com slash TeamChanger. That's A-T-L-A-S-S-I-A-N dot com slash TeamChanger. AI has moved from possibility to practice.
12:51The question now is how do today's pioneers build it fast, reliably, and at scale? From medical research and education to scientific discovery and cinematic creativity, AI is already changing how we understand the world. As the essential cloud for AI, CoreWeave is a leader of that transformation. Delivering a purpose-built AI cloud designed specifically for today's most complex AI workloads, it allows researchers, developers, and creators at leading AI labs and enterprises to focus on impact, not limitations. which is how bold ideas become real breakthroughs. CoreWeave was built for this. Ready for anything, ready for AI.
13:35To learn more about how CoreWeave powers the world's best AI, go to coreweave.com slash ready for anything.
13:47Welcome back to Pioneers of AI. You can watch this episode and others by heading over to our YouTube channel. we're having this conversation at a very interesting moment in time when a lot of the AI headlines are dominated by this new tool open claw or claw and then I spent my weekend as many other people did looking into into that same same exactly so for for maybe some of our listeners who are not familiar with it this tool came out and it's basically an agentic kind of Yeah. Open source assistant. And people are using it for all sorts of things, checking your email, triaging it, calendaring, finding you the best price for a flight and booking it.
14:31There was this viral use case where people hooked it up to their wearable devices and prompted it to order a coffee just when you wake up, which I thought was cool. So, yeah, it's basically, you know, kind of gone viral. I have decided that I'm not going to try it out because it also has major security concerns. So I would love your take on all of this. What did you find over the weekend? We're just a few days in. So like take anything I say with a great, great grain of salt and hope by the time this lands in your viewers' inbox, you know, we may be way off in terms of our understanding of that.
15:10But there's a couple of things. when it is clear that there's an immense appetite for these agents. And in a sense, maybe this is the first of these agent platforms to sort of go viral on the consumer side. And so it's super exciting to see, you know, I don't know, I think the last number I saw was 150 ,000 of these agents getting created. Right. Each agent with their own data, with their own incentives and with their own set of skills. And so that starts being super interesting to just explore the diversity of things that people are interested in doing with their agents. Just from like almost a sociological point of view of what do people want out of agents, that's a fascinating development over the last few days.
15:55The second consideration is all of the safety questions. You know, is this vulnerable to prompt ejection? Can you explain prompt injection? Yeah, like the ability through a prompt to inject information that would lead the agent to take on malicious behaviors. Out of that, are we going to discover many different ways in which, you know, we have didn't think of vulnerabilities in our systems? Quite likely so. And so it opens up the door to discovering, and it opens up the door to new vectors of harm. But I also think, you know, the faster we learn about this, the better we're able to build in the right mitigations for it.
16:38And in many cases, you don't know the creative ways in which, you know, people may be using these agents for nefarious purposes until you really deploy that scale. You can do all the right teaming you want in-house before you go. But actually, when you go at scale is when you discover all of that. The third piece I want to highlight about this OpenClaw platform, which I think is possibly even more fascinating, is the interaction between all of these agents. Yes. There's a social network, right? Of the multibook. Yes. So, so much of AI, you know, we build agents thinking they're going to be unique.
17:13They're going to be deployed and they're going to go out, interact with a stationary version of the world. Now, suddenly we see what happens when you deploy against all these other agents. They start interacting. They start negotiating. They start trying to offload some of their work, all these other interesting behaviors. I've been looking for a platform to really study these multi-agent systems at scale. and we haven't had one in the space of a genetic AI. We've had it in a few other contexts. For example, you know, with autonomous vehicles on the roads, we're starting to see some of these effects.
17:52And as much as on the vehicle side, you know, for many years, we worried about the trolley problem. And in the end, like the problem that really turns out to be hard is like all the Waymo's stuck in a parking lot, like spinning their circles, not able to get out because they're in some kind of like pathological loop with each other. So what is going to happen in this case with all these LLM agents talking to each other and exchanging information, deploying their skills, using tools? I think that sandbox and the ability to look at the network effects across agents is going to be super interesting.
18:27And so I'm quite curious to see what comes out of that. And from this, what we can learn in terms of some of the essentially like social norms we need to build into these networks so that we don't have too many pathological behaviors. Yeah, like I saw, for example, one of these open claw agents was struggling to complete a task. So it went to Moldbook, which is the social platform, the Reddit for AI agents, essentially. And it basically like asked a question, it got upvoted, it got the answer and implemented it. But I also saw some agents were complaining about their human, you know, the humans behind, you know, behind them.
19:08And I thought that was funny. Hopefully, whatever framework we end up building does not allow for that. But, you know. Or maybe it should. Or maybe it should. I don't know. Yeah. I do think it's interesting, though, that OpenClaw, because of all of the security questions, is, well, also, it's just available to consumers. So what are the parallels in enterprise agentic AI, which is what Cohere focuses on? What are lessons to be learned from this kind of this moment? I think one thing to keep in mind is we're still very, very early days in terms of deployment in enterprise. And so whereas on the consumer side, we're already seeing the rise of the adversarial behaviors very, very early.
19:51I think on the enterprise side, we're not at that stage yet, frankly. Like we're seeing all the ways in which the agents are not meeting expectation of the users, but it's not adversarial. So I think there's a bit of a gap in terms of those two environments. As a consumer, like you're kind of responsible for your own privacy information. And I think we've long stopped expecting the platforms to handle that for us. On the enterprise side, there's still an expectation that the platform will provide high levels of privacy. And so we can't afford to take some shortcuts on that. And so I think that's one of the major differences I'm seeing in these early days.
20:36But again, so much more that we need to learn on this. Yeah, it's very exciting. So you did bring up the topic of openness. Why do you think the kind of openness is important? And in particular, what's Coher's approach? My understanding is some of your models are proprietary, especially for commercial use cases, but you open weight some of the models for research and development purposes. Can you kind of unpack that for us? How do you decide? And also, what is the difference between open source and open weight? Yeah. I mean, I spent the early part of my research career in academia where it's very natural to share all the artifacts of your research.
21:14Maybe not in all fields, but in computer science, it's pretty standard to share the code, the data, the paper, all of the artifacts that capture the results of the research process. Open source is about essentially like sharing all the information that was used in the course of the research in a way that facilities the reproducibility of that research result. When it comes to LLMs in recent years, we've seen much more movement towards open weight models, meaning that the weights of the network are shared such that someone can download that and essentially like rerun the network locally. And from that, then adapt it, fine tune it, and so on and so forth.
21:55But often some of the other artifacts are not shared. So you might not have the training code. You might have only the inference code. You might not have the training data. With respect to Cohere, I would say, you know, their position on openness is one of the reasons that drew me to them. I think they have a great track record of sharing the results, whether it's Cohere Labs, the results of their research, or whether it's the command line of models that are coming out of our modeling team. I think there's incredible value in terms of making our own work better, holding a hard bar, and also really empowering the community to keep building with us.
22:30Yeah. And keeping a kind of a level, some level of transparency as well, right? Yeah. We're going to take a short break. When we come back, mitigation risk and how Joelle found herself in the cockpit of a helicopter. All for the sake of AI research.
23:05I want to talk about risk assessment and risk mitigation when it comes to AI. And I want us to dig into benchmarks. And the story that comes up for me that I want to share with you. So when I spun out my company out of MIT, Affectiva, and we did computer vision to understand human emotions and human facial expressions. And, you know, of course, we would run all these validation tests on our data sets and whatnot. And we would look at all sorts of accuracy measures and scores. But at the end of the day, I would bring every like all our R &D team and we would lock ourselves in a room and we would literally watch all the videos and take a stab ourselves on kind of deciding what category, what class every video fell into.
23:48My thesis back then, which I still believe is true, is you build an intuition around the data. And I think that is so important. You can't just like delegate the results to kind of a harness of tests that you do. So I'm just curious, what's your approach to thinking about how we benchmark these models? There's a couple of things I think you talk about, you know, both like the formal evaluations and then the more informal experience of the model, which takes a different shape depending on what's the model. I think more and more, we also have the challenge that as we're building general intelligence, it's not, you know, whether it's the automated metrics or whether it's the human experience of the model, both of that becomes much more multidimensional.
24:37And so, you know, let's say you have just a model for translation from one language to another. Like there's some pretty standard metrics, Blue Rouge and so on to do machine translation. And then there's the experience of it. But in the experience of it, you're just going to expect to look at translation. You're not going to start expecting like a chatbot that can, you know, solve all sorts of your personal problems, gives you some advice, handle your email, all these other things. Like it's a narrow set of qualitative experiences as well. As the model get more and more general, models that suddenly are multilingual, multimodal, including chat and reasoning, agentic behavior, tool use, you start having a huge surface of automated test, but you also have like a huge surface for humans to form an intuition.
25:24mission and it becomes quite difficult to know what to look at, quite frankly. And we have that problem in practice. Like we train up a model, we think it's a good model, but then it's like out of this whole set of evaluations, like which one do we look at? Which one do you put more weight on, right? Yeah. And often there's a difference between what the external community might have decided to focus on, whether it's like the math olympiads or for a while we were looking at Sweebench and so on versus what actually matters to the paying enterprise customers that we serve. And so, you know, that has like a whole other dimension to it.
26:02So there's no perfect answer to this. I would say in general, we try to have a clear sense of like, what are the benchmarks that are sanity checks? So they may be saturated, but we still need to check that we don't have regression on those. What are the benchmarks that are really driving the progress? Because they're the benchmarks that are, you know, we're on the edge of unblocking. Some of them are from the broad AI community, and some of them are directly from our enterprise customers. And we look at a mix of those. And then we have a third bucket that's more the experience. You know, we have a whole harness for A-B testing and for bringing in human evaluators.
26:41And so we also leverage that as a way to do it. One of the things we often hear is the models that we have, You know, they're solid on some of the sanity checks. They're solid on the external benchmarks, but they're particularly good on the enterprise benchmarks. And they're particularly good in terms of the experience. And, you know, it's not completely, you know, a surprise. I mean, like we consider that as part of the development. But we can't afford to sort of overfit to just the enterprise benchmarks because I think we lose some of the generality. And so it's important to keep in mind the broader set of benchmarks.
27:20So, you know, on the one hand, you've got all these benchmarks that a lot of the other foundation models also use. But how do you create an enterprise benchmark? I imagine you have to partner with these customers, right, to get like the right data set and create the right harness. How do you do that? It's mostly about looking at the right tasks. and it's about so defining a question of like, yes, the right data, but it's more like what information and tools are we gonna allow? And from that, you know, you have a task, what's the concept of done? What's the concept of great? And so when we have tasks that have clear testable conditions for completeness and success, then that's usually a good one to build an evaluation out of.
Read the full transcript
28:06I was on a panel recently and someone shared, But, you know, you should never take a broken process and use AI to fix it. Because when you have a broken process, you don't have a definition of what done or great looks like. And so you're not going to be able to build up an AI solution for this. So picking things that you know, you have a good sense of how to assess, whether it's done, whether it's well done, it's usually easy enough to build an evaluation for that. Yeah. Now, I believe that diversity is also very important when you are evaluating these models, well, both training and evaluating these models.
28:43Do you have a specific point of view on how to bring diverse perspectives, I guess, when thinking about these models? Yeah, I'm smiling because I think, you know, there's an anecdote from my very earliest days that just stuck in my head. And for all these years, I go back to so often. I've told this story a few times, but when I was an undergrad at the University of Waterloo, I did a work term one fall at the flight research lab in Ottawa, where we were essentially building like speech recognition system in the cockpit of helicopters. And so, you know, like we'd have test subjects come in, test the speech recognition system.
29:21We'd go in flight. The pilot would ask all the commands with voice and then, you know, pilot the helicopter using standard command. And after a few evaluations, at some point I asked my manager at the time, said, like, can we get a female pilot in there? Like we know voices of men and women are different usually just in terms of range. And so, like, can we just validate that this is working also for a women pilot? And to his credit, you know, like this was a number of years ago, you know, bias in AI was not on the agenda. To his credit, my manager said, like, great question. Let's figure it out.
29:56Called up the pool of pilots and they could not find a woman to send to us. And again, to his credit, he said, not a problem, Joelle. you are going in the second seat and you will be piloting this helicopter and running the protocol yourself. And so I found myself piloting a helicopter, running the whole test protocol. I wasn't enough to have one test subject to give us like a thorough evaluation, but it gave us a measure at least of like the discrepancy in performance as a sanity check. So anyways, that stuck to my head in a couple of ways. One, I think that the difference really came to just having someone in the room to ask the question, right?
30:36Often you're not, you have a diverse set of people at the table. You're not going to get necessarily diverse solutions. Like I'm going to build my speech recognition system the same way that someone else is going to build their speech recognition system, but we will ask different questions and that's going to push the research in a different direction. So I'm a huge believer in building diverse teams for that question. Like what are the set of people we need to get around the table such that we're going to look around the corner and ask the questions that no one else is asking such that we can build products that are better.
31:05And that especially when it comes to general AI, in order to be more general, you need to open up the aperture of people who are building these models and the questions that they're asking and the types of tasks that we're putting forward for these models. So how does that translate to what the team looks like on the ground? Like what kind of diversity of backgrounds or perspectives do you want to see around the table? In some cases, you know, some of it is building teams that have diverse expertise. So we tend to have a mix of like research scientists, research engineers, some Swiss designers, user research, some program managers kind of all working together.
31:44Some of that is we are working across several countries. We have teams based in Canada, teams in the UK, in France, and we have folks in Germany. We have folks in the US. And through that, we've become really focused on building the world's best multilingual models. So cool. So the AYA line of models, which came out of, you know, I think the team being quite diverse and being interested in this question of multilingual modeling. And now it's opened up new commercial markets because when we get to Korea and to Japan, you know, of course, there's like, oh, you have world class models in the local language.
32:19And so it's opening up new perspective. But that didn't come from like a top down directive from like our revenue team saying like, you know, we have a request from Korea to build a Korean language model. It came out of like our research team being interested in democratizing access to language model, working with a large external community, assembling the data, building the model and so on and so forth. So I think that's the other aspect of that, like having a community that's really quite diverse from a geographic, linguistic, racial point of view. It really opens up the perspectives in terms of the work.
32:54Yeah, very, very, very cool. So I'm originally from the Middle East and I'm Egyptian, grew up in Kuwait and Abu Dhabi before I moved to the UK and then the US. And I am just very curious about the question of sovereignty in AI. Yeah, we see different countries wanting to build their own AI tech stack. What are your thoughts on that? We are definitely seeing a strong movement towards sovereignty. I think it can mean different things. But at the end of the day, this notion of sovereignty comes down a lot to the ability to have some control over your technology. For us, Cohere, ourselves, you know, the fact that we build our own models as well as build our own product is part of the strategy.
33:39And part of that is just to be robust. You know, by controlling the model, we control the data mix that goes in. We control how the evaluations are done. And we have a really good understanding of how we can shape that core piece of the technology such that we can build better products. So I'm really sympathetic, frankly, for this desire for more sovereignty. I also see a ton of appetite for it. I was in Paris earlier this month. I was in Davos earlier this month also. And everywhere I went, you do hear a lot about this question of sovereign AI. Some of it is, of course, related to the questions of geopolitics, but a lot of it is actually just questions of having a resilient access to the technology for us as an individual and for companies.
34:30Yeah, very cool. So I want to kind of look towards the future a bit. So I recently spoke to Kyle Law. He's the CEO and co-founder of an AI company. We hopped on a call and we had a Zoom call together. The caveat is, the catch is, he's an AI agent. So the company is mostly a group of AI agents like the CMO and the COO and this head of human resource. Well, actually, it's like the head of resources, but I guess it's mostly AI. Agent resources. Exactly, agent resources with a couple of humans also in the mix. And it was fascinating for a whole host of reasons. I think we are going to start to see more of these hybrid AI human teams and organizations.
35:12organizations, but one of the epic fails, I would say, of these agents is that they didn't quite have a world model. And so they didn't have a real sense of time and they didn't really have a sense of like what is okay to do and say and what's not okay to do and say. Are you seeing also the unfolding of this kind of hybrid agentic and human world? And how should organizations think about that? And also like, how do you instill culture and leadership into organizations like that. That's an interesting one. I'm definitely seeing a lot of that, though I would say, you know, if you look at the spectrum from like one end of the spectrum being the case where you have a fully human organization where you're starting to pepper in a little bit of AI versus the case you're talking about.
35:57And so a lot of these questions will have very different shapes in that. The case I see, usually it's people in teams who have like a very clear sense of like a particular piece of work, like a particular deliverable that could be done with AI, and suddenly with the right level of autonomy, agency, the right information flowing in, the right LLMs, they can actually accelerate that piece of work significantly to the point that something that would previously take hours to do suddenly is down to just like a few seconds. From a culture point of view, I think there's still a lot of open questions.
36:34For now, one of the biggest blocker, I would say, is adoption. So people come to AI with all sorts of preconceptions, some of them legitimate concerns, and a lot of them unfounded fears. And so we have to navigate through that to get them to a position of curiosity. So from a posture of curiosity, you can start playing with it and see how to integrate this technology into your workflow. But if you're coming at it from a point of view of skepticism and fear without that curiosity, it's going to be really hard for you to take advantage of that technology. So in terms of culture, I would say we're still in that phase of how do we get people to just be curious about the technology.
37:17And through that curiosity, you achieve productivity. Yeah, very cool. Yeah, there's this MIT study that found that 95 % of organizations see no measurable return on their investment in AI. Is that kind of consistent with what you're seeing in your world? I mean, no, I'm seeing a lot more value than that, but it's still very early days to put a number on it. I would say I would love to, you know, talk to you again in a year and have like a much better way to capture the productivity gains that we're seeing from the technology. I think you can read this number one way and say like, oh, there's no value in AI and I will stop trying.
37:56or to say like we haven't figured out the way to do it, the right way to do it, the right task, the right way to bring it into the work that we're doing from either culture point of view, process point of view. So I'm much more leaning on the second, I think to think that this is evidence enough to not lean into AI if you're running a company or if you're involved in government, to me seems really misguided given the trajectory I see for this technology. Yeah, a good mentor of mine, Peter DeMondis, he has this quote that in a decade, there'll be two types of companies or organizations, those that have adopted AI and those that are out of business.
38:36And I really believe that. Yeah, I tend to think so too. Yeah. All right. A few rapid fire questions. What's one way you're implementing AI in your everyday life that you know you can't live without? I'm still very much like go to AI whenever I'm stuck on anything. If there's something I don't know how to do, whether it's like fixing something at the house or doing something at work, if I don't know how to do it, I just go straight to AI to just unblock myself. How are your kids using AI? I have four kids who are sort of 16 to 22, covering sort of late high school till university. A lot of help with homework, pretty standard teenagers.
39:18This week, it was also a lot of help writing cover letters for summer jobs. Okay, that's a good use case. AI as a thought partner in professional matters, yay or nay? Absolutely. Personal matters? Yes, also, though I tend to turn less to it. I tend to turn to my friends and family before I turn to AI on the personal matters. Thank you, Joelle, for joining us. This was super fascinating. Thank you. It was a pleasure to chat with you. It was so fun talking to Joelle about her take on where AI is headed, especially digging into OpenClaw. So here's my take. OpenClaw feels like the chat GPT moment for AI agents.
40:00We've all been talking about agents for the past year, but this is the first time we're seeing them broadly adopted by everyday consumers. What makes this even more notable is Moldbook, a social network where agents can interact and even learn from each other. Again, this is the first time we're seeing AI bots communicating in the wild at scale. That said, it also surfaces real concerns. A lack of security and privacy, plus the clear need for human oversight. Personally, I'm not going to download it. Too much risk. The mainstream version of this technology will need to be a lot safer. It will also need to be packaged in a way that's accessible and not as techie.
40:42But this is definitely headed in the right direction in terms of what AI agents can do for us. Thanks so much for tuning in. If you haven't already, subscribe to us wherever you're listening so you don't miss an episode. We'll be in your feeds next week.
41:08Pioneers of AI is a Wait What original production. Our executive producer is Eve Trow. Our producer is Rachel Ishikawa. Our senior talent executive is Stephanie Stern. Mixing and mastering by Brian Pute. Video editing by Eric Purcell. Original music by Ryan Holiday. Our head of podcasts is Litao Moulad. You can join the conversation across social media platforms. Just look for us at Pioneers of AI. Thanks so much for listening. Thank you.
From the publisher
We may have reached the “ChatGPT moment” in agentic AI. Over the past couple weeks, an open-source AI agent called OpenClaw (previously known as Clawdbot) has taken the AI world by storm. It boasts nearly 200,000 stars on Github and has adopters scrambling to buy Mac Minis to run the program locally. The lobster-themed agent works like an ultra-fast, super-smart personal assistant capable of answering emails, booking flights, ordering coffee, managing calendars, updating its source code, and much more without human intervention. However, it’s also raised some major concerns around privacy and security because the program requires total access to your local system and personal accounts. To help make sense of this watershed moment, Pioneers of AI is joined by agentic expert Joelle Pineau. The former VP of AI Research at Meta and current Chief AI Officer at Cohere, Pineau is helping build the next generation of foundation AI models and enterprise agents. She offers her insights on the promise and perils of OpenClaw, why open source matters, and the value of developing AI models outside of the US and China.
Learn more about Pioneers of AI: http://pioneersof.ai/
Follow Pioneers of AI on all channels: https://linktr.ee/pioneersofai
At the center of AI is people, so we want to hear from you! Share your experiences with AI — or ask us a burning question — by leaving a voicemail at 601-633-2424. Your voice could be featured in a future episode!
See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.


