In short
The episode argues that AI productivity stats may miss what’s happening for early “knowledge workers” who delegate to always-on AI agent teams. The host shows his OpenClaw-based “AI chief of staff” (R-Mini Arnold, RMA): a Mac mini running 24/7 that aggregates tasks and data (CRM, research, email, X) into a live dashboard and WhatsApp-based delegation. He claims the key shift is delegation cost dropping below execution cost, enabling a “5–10 person team” effect for individuals, widening the gap between early adopters and others.
Notable examples
RMA prepares a detailed sovereign wealth fund briefing and logs post-meeting notes; it also orchestrated a full 28-minute show script in ~40 minutes via multiple sub-agents (archivist, external research, evidence collector, format researcher, scriptwriter) with QA checks.
Guests
none mentioned; the episode is a solo walkthrough.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOAI's Economic Impact Discussion
0:00 to 0:51
Exploring claims about AI's contribution to GDP and personal deployment experiences.
“Two days ago, Goldman Sachs' chief economist said that AI investment had added, in his words, basically zero to US GDP in 2025.”
The Rise of R-Mini Arnold
1:15 to 4:00
Introduction to R-Mini Arnold, an AI agent, and its capabilities.
“What you're seeing on the screen right now is a piece of software.”
Revolutionizing Work with AI
4:00 to 6:12
How R-Mini Arnold changes work dynamics and challenges economic narratives about AI.
“It is the AI agent that does things that are similar to the AI agents that were imagined back in the 80s and the 90s.”
The Specifications of R-Mini Arnold
6:12 to 7:54
Technical details and operational setup of R-Mini Arnold.
“What RMA has done is it has made it really, really easy for me to just bark orders sometimes into a system in parallel and get lots and lots and lots of work done.”
Daily Operations with R-Mini Arnold
7:54 to 10:52
Overview of R-Mini Arnold's daily functions and tasks.
“Let me just take you through the bits and pieces that make up R mini Arnold.”
Research and Code Management Tasks
10:52 to 14:03
Detailed look at how R-Mini Arnold handles complex research and code improvements.
“But of course, given the work that I do in research and analysis, it includes interesting analysis.”
AI Chief of Staff: Enhancements and Interfaces
14:03 to 18:14
Explore how AI acts as a chief of staff, improving workflows and the role of WhatsApp as an interface.
“And what they do is they walk across all of my other GitHub repos, making small improvements, finding tiny bugs and shutting them down.”
Real-World Impact of AI System
18:14 to 22:40
Learn about the practical impacts of using an AI system for meeting preparation and follow-ups.
“I experimented with using Telegram for this, but I don't really like Telegram as an app.”
Script Generation Using AI
22:40 to 28:05
Understand the process of generating podcast scripts with AI orchestration for efficiency.
“And I love this example because it's about how this episode was made.”
Managing the AI Chief of Staff
28:05 to 28:52
Learn how AI can assist in team management and task delegation.
“Everything that any manager of a team does, I mean, that's what we do.”
Show all 14 chapters
The Soul.md Document
28:52 to 29:54
Discover the role of the personality specification in AI performance.
“It's not better spell correct, manned by a stochastic parrot.”
Learning from Failures
29:54 to 32:28
Understand how AI can learn and improve from mistakes.
“And so what I asked Armini Arnold to do was to look at our interactions, look at how I work, and come back with a proposal of what should be in that sole.md document.”
AI Capabilities and Industry Comparison
32:28 to 37:14
Explore the unique capabilities of personal AI vs enterprise AI.
“back in the years beyond, 1991, I think.”
Impact on Personal Workflow
37:14 to 41:02
Examine how AI changes personal work habits and decision-making.
“Meta has acquired Manus, the Singaporean firm, mostly known for its deep research capabilities, right?”
Transcript
Automatic transcript. May contain errors.0:00Two days ago, Goldman Sachs' chief economist said that AI investment had added, in his words, basically zero to US GDP in 2025. But here's the thing, they're looking at companies across the economy, but they don't look at people like me. A small number of us have already deployed what amounts to a five, maybe a 10 person team working round the clock on our behalf. I'm not special, I'm just early.
0:26Azeem Azhar:The gap between the people who've started and the people who haven't started is widening every week. Is this going to make me worse at certain things? Am I going to think less carefully before delegating because the system is so capable? Am I sharpening my judgment or losing the muscle that that judgment requires?
0:51Welcome to Exponential View, the show where I explore how exponential technologies, in particular AI, are reshaping our future. And it does feel like that future is coming ever closer and becoming more like a discussion of the present. Now, each week, I'll share some of my analyses or speak with a guest to make light of a particular topic. But this week, I want to show you something. What you're seeing on the screen right now is a piece of software. It is a knowledge dashboard that I use to track what I need to do each day. It brings together material from my internal systems, from our CRM system, from our research engine, from X, from my email, from other sources, and it ranks the things that it thinks are important for me.
1:44And given my work, a lot of the things that are important for me are outward facing. They're not necessarily about projects and things that are happening internally. This is a live knowledge dashboard. The thing about this is it's live now, right now, and it runs on a Mac Mini in my studio, just over there. You can't see it. It's in an equipment cabinet. This didn't exist eight days ago. I haven't written a single line of code of it. It was put together by six AI agents overseen by a super agent, or perhaps I should say it was one AI agent overseeing six sub-agents, and they built it overnight over the course of a couple of days with my feedback.
2:27They argued about the database schema at three in the morning. They wrote tests for each other's codes. They deployed it and I woke up and it was running. The first time it was running, it was a bit ugly. It was a bit shonky as any new piece of code, but ultimately it works. And I've iterated a couple of times and it's something that I now use every day. Now that super agent, the primary agent that orchestrated and coordinated all of that is called R-Mini Arnold. Now the R comes from Isaac Asimov's novels. In those novels, robots that were intelligent were given that moniker R for robot, R Daniel Olivo, and it's what we use when we are naming agents of the type that I've just described today with an exponential view.
3:15What about the name Arnold? In the Terminator films, in the second Terminator film, Arnold Schwarzenegger comes back to protect humanity from the even more Terminator-y Terminators. And so our mini Arnold, because my agent is not as big as Mr. Schwarzenegger, is sitting in that Mac mini and doing all of that work. I've been running our mini Arnold or RMA, as I will call it during the course of this conversation, for about a month. And it has really changed the way I work more than any single tool since the web browser. I can't overstate it. You're going to hear me talk about it now for 25 minutes, but it has changed the way I work.
3:58And what is it? Well, Armine Arnold is, for want of any kind of better word, definition, phrase, the academics can argue about this. It is an AI agent. It is the AI agent that does things that are similar to the AI agents that were imagined back in the 80s and the 90s. think about Alan Kay's famous knowledge navigator. And what I've experienced over the last three or four weeks is that a lot of the discussion about AI agents, especially when we think about it in economic terms, perhaps might have missed the point entirely. Two days ago, Goldman Sachs' chief economist, Jan Hatsius, said that AI investment had added, in his words, basically zero to US GDP in 2025.
4:46And earlier in the week, there was the paper from the NBER that suggested that 80 % of American, British and Australian companies were reporting no productivity gains. And that headline is repeated in lots of other places. There was a PC magazine that said the AI agent hype is real, the productivity gains aren't. Now, there are many ways to interpret that number. Of course, not least, it means that 20 % of firms just three years after the chat GPT moment do claim to see productivity benefits. And that concords with our proprietary tracking of US public companies, Gen AI claims. It lines up with what our friend Eric Brynjolfsson at Stanford University is starting to say about the productivity data that is becoming visible.
5:31But here's the thing. Those numbers measure a particular type of real thing. And perhaps they're measuring something that is less important. Maybe they are measuring the wrong thing. They're looking at companies across the economy.
5:45Azeem Azhar:Sometimes they look at individual companies, but they don't look at people like me and they can't capture what is going on there. The revolution isn't merely general intelligence. One of the things I've observed is that when the cost of delegation falls below the execution cost for a growing fraction of what we call knowledge work, when that cost falls by an order of magnitude, you do much more of it. And that's basically the whole story. What RMA has done is it has made it really, really easy for me to just bark orders sometimes into a system in parallel and get lots and lots and lots of work done.
6:26If you think about a big company, they're running pilots, they're building governance frameworks, and they're hiring chief AI officers, a job that exists today and will not exist in five years.
6:37Azeem Azhar:But they have to contend with all the issues of big companies. And that's what turns up in the statistics. A small number of us have already deployed what amounts to a five, maybe a 10 person team working round the clock on our behalf. Now, I'm one of those people and I'm not special. I'm just early. The gap between the people who've started and the people who haven't started is widening every week. And that's not just because the technology is getting better constantly, though it is, but it's because that relationship that you have with the technology compounds. That relationship starts with when you started to use ChatGPT, when you started to move from quick summarizes paragraph of text to more complex multi-term prompting, when you started to use deep research, when you started to think about your prompting strategies in order to get better outputs.
7:32Azeem Azhar:that compounds rapidly. But the other thing that compounds is that the agent, RMA in my case, has worked with me for 30 days and it knows things about me that no brand new tool does. It knows things in certain types of contexts that you don't get if you're using Clawed or ChatGPT with their memory capabilities. Let me just take you through the bits and pieces that make up R mini Arnold. R mini Arnold runs on a Mac mini. It's got 64 gigs of RAM. Look, I bought this Mac mini for the agent. I'm obsessed with buying more RAM than I might possibly need, because it's the one thing that's really complicated with Apple computers to upgrade.
8:15You can always add more storage. You stick an SSD in, but RAM is a sort of one-way street. It's got 64 gigs of RAM. It's connected by wired Ethernet, a 10 gigabit Ethernet directly to my firewall, has a five gigabit internet connection out to the world. And it sits in the equipment
8:31Azeem Azhar:cabinet in this studio, running 24 hours a day, drawing about the same amount of power as a desk lamp. The software that it's running is called OpenClaw. I've written about it a lot on Exponential View. It's open source, it's self-hosted, it runs on my hardware. My data largely stays on my machine. Some of it is in a shared Dropbox in the cloud so I can send things to Armini Arnold. Some of it, in terms of the long-term memory, gets exported out to a vector store that's in the cloud so that it can learn my preferences over more extended periods of time. And the model underneath most of the time is Anthropix Claude Sonnet 4.6 model.
9:17Sometimes I flip up to Opus 4.6, sometimes I go down to the quick and easy Haiku 4.5. And from time to time, RMA might call an open AI perplexity or a Grok model if there's something really particular that I need. But I would say from time to time, not even one in 100 queries that I do. Last week, Sam Altman snagged Peter Steinberger,
9:44Azeem Azhar:the developer behind OpenClaw. He's joined OpenAI. He had been pursued by Mark Zuckerberg, messaging him on WhatsApp. I don't know how Mark got his number, but I guess Mark owns WhatsApp. So maybe he has ways to see if they could support that business. And I know that many, many Silicon Valley VCs were really trying to get hold of OpenClaw, trying to get a hold of Pete to see if they could support that business. So you can draw your own conclusions about what that tells you about how much of a shift in consumer user interactions we've seen from the model of the AI agent that OpenClaw has delivered to us.
10:24So that's a structural story,
10:26Azeem Azhar:and I've written about it on the newsletter. You can go off and see some more details. Let me describe to you what it actually looks like. Like many of us, I wake up early in the morning, by the time I'm downstairs stretching or making coffee, RMA has been running all night. And certainly by about 6 or 6.15 in the morning, a morning brief arrives on my WhatsApp. And that morning brief is similar to things I'm sure you're used to, the calendar for the day, priority emails flagged overnight. But of course, given the work that I do in research and analysis, it includes interesting analysis. The top stories that have come out of PRISM, which is our research backplane that we use at Exponential View that has thousands of things that the team has read and annotated and their own analysis.
11:16Azeem Azhar:It also includes RMA's own analysis of what it thinks might be interesting for me, given the wider context it has of me. And part of that wider context is the information it gets out of Orbit. And Orbit tracks all of my major relationships, work relationships. So it's a flashy CRM and Orbit will be connecting the types of people I might be meeting in the next few days with the research and the stories that are happening today and RMA assembles all of this together and delivers it to me. I'll get a summary of all the tasks that have run overnight and I will run quite complex research and coding tasks overnight.
11:52Azeem Azhar:I'll be told which ones succeeded, which ones failed. There's a whole lot of housekeeping RMA runs as a system and it might have to give me details of things that need my attention. Now, what are these tasks? Well, some can be quite complex. So you may have seen that a few days ago, some substack called Citrini Research wrote an essay that was pretty interesting, a detailed scenario of an economic collapse powered by AI that ran out for two or three years. It was obviously trending on Twitter, it was blowing up on substack. And according to some of the financial press, it also affected the markets and resulted in a real shock to the markets at a time.
12:29Azeem Azhar:So I was really curious about digging into the authors and the team behind Cetrini and behind that essay. And so I asked Armini Arnold to do some research. Who are those authors? What's their heritage? Where have they come from? How serious are they? How robust was the work that they had done? And now I know RMA set up a team of investigators to go off and comb through this. I think at one point there were five sub-agents working on this task and doing web searches and drilling down and trying to see what they could learn. And I've got back quite an interesting report, which if I have time, I may put some of in the newsletter on Sunday.
13:09Azeem Azhar:But that wasn't the only thing that was happening overnight. So another set of agents were refactoring the code of my CRM system. What does that mean? It means that this is a system that I designed and I've built over the course of a few weeks. And of course, I'm interested in shipping features. I'm much less interested in thinking about the elegance and the sort of structural integrity of the code. That's true for any code. It's the accumulation of tech debt. I told RMA that what I wanted to do was assemble a team of security researchers, quality analysts, documenters to walk across the code base of the CRM and fix what was broken, reporting back to an architect who had confirmed that everything still worked.
13:54Azeem Azhar:It was a major piece of work. Several thousand lines of code were touched overnight. I also have a different team that works every single night for me of agents. These ones are running OpenAI's GPT codecs. And what they do is they walk across all of my other GitHub repos, making small improvements, finding tiny bugs and shutting them down. And finally, RMA has a team of security agents, some of which are inside the firewall, some of which are outside the firewall, which effectively find and patch security vulnerabilities across the system. And typically, they don't find much. But normally, once every three or four days, something does crop up.
14:30So this is like having a
14:31Azeem Azhar:chief of staff who can draw upon a number of specialist teams, essentially constantly overnight. And if you think about what I just described there, how much of that was me initiating work? Well, two of the pieces, the piece about checking the CRM code and doing the research and the digging around the Citrini essay were de novo tasks for the day. But the rest are things that are just scheduled and run day after day. Now, what is the interface for all of this? The interface is WhatsApp. The reason is that's the interface that I use for so much else. Why not just use WhatsApp? It's the same app that I use to message my wife and my family and my friends and my kids and the people I work with.
15:12Azeem Azhar:There's no special dashboard required. I have built that dashboard that I shared with you earlier because it was really useful given the breadth of things that I'm working on to have a dashboard like that. I mean, you just don't really want to be flagging research stories on and off through WhatsApp. But you don't need that dashboard. And in fact, I only check in on that a couple of times a day. What I do hundreds of times a day is use WhatsApp effectively to delegate and check and verify work that my agent and its team of agents are working on. So today I open WhatsApp and I talk to it. The agent comes back.
15:53Azeem Azhar:It picks up where we left off previously. It knows my priorities. It knows who to contact and why. And I'm going to show you what that looks like. So here you can see what my WhatsApp looks like. I have a tag actually for all of the various RMA agents. So this is the thing that's quite strange. It's not a single conversation. It is eight conversations. In fact, there's slightly more than that. Each different bit of context I have for long running tasks with RMA, I create a communication lane for. So R. Mini Arnold, right at the top, it's one of my favorites, is the general chuck it at it and figure it out.
16:35Azeem Azhar:R. M. A. I. G. B. is the channel for my new book. The sub-agent there has all of the chapters. It knows the chapter structure. It knows the argument. It knows all of Chantal's latest findings. It has access to Prism so that it can look things up for me. The channel on orbit is the CRM. It knows about my contacts and other details, but it also has access to the GitHub repo. So when I see a bug quickly during the day, I can tell it to fire off and fix it. RMA EV research are all the things that relate to the research that we're doing within Exponential View, whether it's for the newsletter or it's for something else.
17:18It's the same AI, but with eight specialized instances all running in parallel, all with their own memory, all with their own tasks, all with their own context.
17:28Azeem Azhar:Now, the script says they don't talk to each other and each one knows deeply about its domain and not much about the others. The truth is they do sometimes get confused. And sometimes I will be in the EV research channel and I'll get a response from one of the other channel lanes showing up. This is early software. There are going to be these types of issues and teething trouble. So if we then think about what that means, I mean, it's a really, really kind of remarkable situation because, of course, I've got this master orchestrator, Armony Arnold, but I have these eight contexts, which it should technically adhere to.
18:04Azeem Azhar:And it's like having those eight simultaneous relationships, but it's with the same entity. There is some shared context. I haven't figured out exactly how to do this perfectly. I'm still trying to figure it out. I've experimented with other things. I experimented with using Telegram for this, but I don't really like Telegram as an app. I experimented using Slack for it, but I just again found I don't spend as much time in Slack as I do in WhatsApp. I had actually build me a small app that was a persistent web app where I could keep track of all of these channels. None of those worked for me. Those approaches might work for you.
18:40Azeem Azhar:everyone is different. The asymmetry in all of this I find most interesting is not AI and humans, it's actually individuals and institutions. Because I'm currently running these capabilities that I doubt any Fortune 500 company has deployed at this level. So let me tell you about the type of impact that this system has had for me. And in that, it'll address questions like, how much does it cost? Or is it worth it? A quick note, if you want to support us in bringing more of these conversations to the world, please consider subscribing to the show. I'll give you three small vignettes. I had a wonderful meeting with a major sovereign wealth fund a couple of days ago.
19:19It was a big meeting. They're impressive, super, super smart. And I had to rush straight from having a filling at the wonderful dentist just down the road from me up in Hampstead in London. And it was going to be a really, really hectic morning. RMA had figured out a brief for me and said, you've got this dentist thing, and then you've got to run straight to this meeting. Here's a brief. I recognised the name and prepared me for it. The briefing was pretty thorough.
19:47Azeem Azhar:Of course, it had all the obvious things you'd expect because it's gone to the CRM and it's pulled those details. But it also has access to all of my granola transcripts and any other conversations I've had that are relevant. And it has access to PRISM. And so PRISM contains lots of research, as well as our house views, knowledge of how the team in Exponential View thinks, uses frameworks, thinks critically. And it knows a lot about the person I was meeting. So the context I was given was really quite remarkable. And I've used CRMs that prepare you like this previously. And I've used AI systems over the last year or two that try to prepare you like this seriously.
20:25Azeem Azhar:But they sort of lack that context and that nuance that I have through RMA. And I will explain in a few minutes why I have so much context and nuance there. What I got was something that described the investment posture around compute, data centers, chips, their wins, the open questions, all anchored in a really broad set of news and analysis that I might have missed. So I walked into that meeting differently. I wasn't scrambling to look somebody up in perplexity while I was running out of the tube. I actually had some real context. It's the kind of thing that a great chief of staff with plenty of time would have prepared if I had one and had had the time.
21:05Azeem Azhar:The thing that surprised me, because I hadn't told RMA to do this, was that after the meeting, I moved off to the next thing, which, as it happened, was to go to the Apple store to get my laptop screen repaired. It updated the CRM with the fact that I'd had the interaction. It didn't have much details, but I did have a WhatsApp saying anything you want to add as notes from this interaction. And, you know, that, again, was just an open loop that I would have had to remember and had to surface. And it also reminded me that I've met a peer to that person previously, and whether it's worth reconnecting with them as well.
21:41Azeem Azhar:Pre-meeting briefing, that's fine. That's reasonable inference. The quality of that briefing, better than things I've seen before. The post-meeting logging in a quite a non-invasive way. And I would just say something about that non-invasiveness, which is if I have a Zoom meeting where I've run granola, I don't get that flag. I guess it sees it's got the granola transcripts and it can populate the CRM that way. Now, here's the thing. I actually don't really know how that happened and how it worked. I don't know if it'll happen reliably week and week, day after day. I mean, these LLMs are jagged, as we know, wobbly in some cases.
22:15Azeem Azhar:But this isn't the case of the AI doing my homework. This is an agent that understands the shape of that professional relationship and all the effort I've put in to ensure that there is enough context around it. And it knows that a meeting doesn't end when you leave the room. So I found that pretty interesting, quite special, again, compared to things I've used previously. Here's a second example. And I love this example because it's about how this episode was made. So 24 hours ago, I needed a script for this show. I knew roughly what I wanted to say. I had the thesis. It's the delegation cost argument.
22:56Azeem Azhar:I had a few stories. I know the audience. What I didn't have was time. And, you know, as regular listeners know, I work with Chantal normally on the scripts. So I thought I've had a pretty busy Thursday. It's going to take me four or five hours. I was frankly exhausted by the time I got home. So I went to WhatsApp and I gave RMA a brief. And I said, look, I'm doing this live sub stack. It's kind of about you. It's tomorrow. Here's the thesis. Here's the tone. Here's roughly what I want to cover. And I want it written in my voice, the voice I've been cultivating over 30 years of my professional experience.
Read the full transcript
23:35Go, RMA, go and figure it out. And then I closed WhatsApp and went to do my other thing, which was to finish reading the book I was reading. So what happened in the background? RMA spawned four specialist sub-agents simultaneously. Each one got a brief, a task, and its own context window. So the context window is the working memory of an LLM during sort of a back and forth. Mine have generous
24:01Azeem Azhar:context windows. It's about a million tokens. They ran in parallel. It's four separate instances of Claude Sonnet in some cases, Claude Opus in others, four different research threads running at the same time. Now, the first agent RMA called the Archivist went into the memory layer. What it was trying to do was find out all of the things that I had asked RMA to do over the past 15 to 20 days, 30 days, identify which ones could make nice vignettes. It searched 79 tracked behavioral patterns for corrections. The dozens of times I've corrected it, including correcting how it should use its name and refer to itself.
24:42Azeem Azhar:It extracted the writing rules that I've taught it and the explicit ones that we developed through the stylometer product. In some cases, instructions I've written down, they're instructions that the agent has learned over the time, patterns that are extracted. The second agent that it created searched the external landscape. What was happening with AI agents this week. It found that Goldman Sachs, zero productivity data. It found that Open AI had just hired Pete Steinberg, something I was well aware of, but it found it itself. He's a guy who created OpenClaw. It found the story about Zuckerberg's WhatsApp message.
25:17Azeem Azhar:It baked it into the script. The third agent was the evidence collector. Numbers, statistics, the 79 things I've had to teach it, the 15 consecutive analysis briefs on AI in India, the 179 mistakes that we have imported into its sole.md document, which is a specification document that OpenClaw agents have. Every number in the script has been checked by that dedicated agent and, by the way, by a QA agent that came afterwards. Their job was to make sure I didn't say something that I couldn't defend. The fourth agent researched the format. So what makes a live Substack show land? What do the first 90 seconds need to do?
25:59Azeem Azhar:Where should I put the screen shares to maximize that visual hook? It read all of the transcripts of the previous Friday discussions like this that I've led. It looked for engagement patterns. It looked at similar, better, more experienced, different presenters and podcasters and came back with structural recommendations. And the fifth agent took all of these research packages and built a narrative structure. The final agent wrote a full script, 4 ,600 words, technically 28 minutes, technically in my voice. The total token cost for that. Well, I'm a bit generous about this. I give the agents really large token budgets to work with.
26:40Azeem Azhar:They never come close. So I told it this was consequential. I was going to be reading a lot of this out to this incredibly important audience, this group of people who matter so much to me, which is every single one of you. and I said 300 million tokens. That's your budget to go and do this piece of work. A million tokens is five to ten bucks. So I was thinking maybe this will cost$1 ,500 to do. In fact, the agent pipeline used 280 ,000. I expected it to come three orders of magnitude below because I've been using it for a long time. That is what I did. And that is how this script came about. Now, that process took 40 minutes of wall clock time.
27:23Azeem Azhar:Armini Arnold told me this was going to take all night. It often says things like that and then tells me to go to bed. It took 40 minutes. I was reading my book. I came back, read the draft, sent some corrections, a handful of corrections. And then I waited till the morning to look at the version that showed up. And the pattern here is not like AI wrote my script. You know, I work with people who helped me with my scripts. We know about Chantal and the other researchers. That's not what this is. This pattern to look at is orchestration. I described what I needed. I did it with a nuance and a complexity because I've got some experience in this field.
27:58Azeem Azhar:And people, things, systems I delegated to went off and did the work. I set the objective. I allocated resources. I specified constraints. And ultimately, I reviewed the output. Everything that any manager of a team does, I mean, that's what we do. And I went off and did something more enjoyable. As I said, I was reading this book. And I think it's important to note that the review of the script was not just seven questions. I mean, it was a couple of hours, two and a half hours, certainly less than it would have been if I had had to do the entire process end to end. As a mark of experiment and disclosure, this is the first time we've tried this.
28:39Azeem Azhar:This is the very, very first time we've attempted a structure like this for one of the shows in this detail because I wanted to demonstrate viscerally where we might be headed with all of this. So there's two and a half hours of me time going back and forth and into all of this. But this is the shift, right? AI is not a faster typewriter. It's not better spell correct, manned by a stochastic parrot. it is becoming a team that can be briefed and in some cases trusted to go away and come back with something worth my time and more importantly worth your time. The fact that you're hearing any of this at all is evidence that at least in my mind that process worked.
29:21Azeem Azhar:So I'm going to keep going. I want to talk about the soul.md document as well. So this is a file that's on the computer. It's only 12 kilobytes, it's tiny, but it's the personality specification for RMA. It's like a super prompt, if you will. Now, the thing about Sol.md is that some people go off and write their own because they want to really tune their agent in a particular way. But what I like to do is I prefer revealed preference. You can tell a lot by looking at what people do rather than asking them what they would do. And so what I asked Armini Arnold to do was to look at our interactions, look at how I work, and come back with a proposal of what should be in that sole.md document.
30:06Azeem Azhar:It found those 179 things that had got wrong. By the way, just in the first 10 days, it's much, much more than that. These were tasks that broke, emails that were wrong and ham-fisted, code that didn't run. At one point, oh my God, it corrupted its own config file Late on a Saturday night, I had workloads that I wanted to run overnight. And I sat there debugging JSON at about 11pm on a Saturday night. I was pretty darn cross about that. But ultimately, it had broken itself so badly, it couldn't recover. And each time something went wrong, I correct it, sometimes politely, sometimes tersely, sometimes I swear, and it got a load of corrections.
30:46Azeem Azhar:But the system also extracted 146 behavioral patterns from our interactions and from those corrections. And these are patterns that were from my actual behavior. And from those patterns, it built a big five personality profile. So as many of you will know, the big five is the most scientifically robust way of personality profiling people. RMA has a score of four out of five on openness. Extroversion level is lower, two out of five. I don't want it to over-explain. I don't want it to be jazz hands and rubber chicken. It's got to be emotionally stable. It's given a score of four out of five. And it absolutely needs to be conscientious.
31:26Azeem Azhar:Five out of five. Lots of the corrections were about file naming, config safety, debugging. The result is it does behave better and it will check a credential before claiming it's expired. But not reliably, I have to say. I still have to do much more work. it still will make small mistakes that it really should know better in a way that a human chief of staff would. They would know the pressures on my time, they would know the mental load that I'm juggling, and they would know which I's and T's have to be absolutely correct. But let's be clear, this is not sentience. This is an elaborate behavioral specification that is read fresh every session, every time the gateway restarts, running and being interpreted by a large language model.
32:08Azeem Azhar:It is mega prompt engineering. If you've ever managed a team, you know that the people you trust more are the ones who learn from their failures. And I think it's quite a nice feature that Steinberg has built into OpenClaw, this ability for it to learn from its failures. Now, I know this is sounding like a pitch, but as I said, this is the most remarkable software I have used since I happened on my first web browser, which was the Lynx text-only browser back in the years beyond, 1991, I think. But I am thinking about a few things. Is this going to make me worse at certain things? Am I going to think less carefully before delegating because the system is so capable that I don't have to specify?
32:47Azeem Azhar:When RMA handles all my research synthesis, am I sharpening my judgment or losing the muscle that that judgment requires? I don't know yet. I think that's a very, very first order and obvious claim. And you can build every type of analogy. I learned to drive before anti-lock brakes were a standard. So I was taught about that sort of pumping to allow your car to brake in slippery conditions. ABS brakes means I have no idea how to do that now. And anyone who's learned in the last 20 years has no idea how to do it either. Has that made us worse drivers? I'm not sure. But I am aware of the risks. So I have taken quite particular steps to ensure that good thinking is still happening to the extent that it can over a WhatsApp channel.
33:31Azeem Azhar:So for example, lots of my modes of reasoning through revealed preference again, rather than my dictating them, have been turned into deterministic patterns. That means they're not running stochastically, probabilistically, they are sort of deterministic patterns. And they're often thrown in as checks to any of the valuable outputs that comes out of these systems. And the purpose there is to give me something that I can read with criticality, that is not just reams of LLM slop, because then what would be the point? And of course, as we know, there are lots of other things that I do and my team does to make sure that we are using the machines, but becoming really good at using them because our minds are sharp.
34:13Azeem Azhar:So earlier this week, my colleague Nathan spent about three hours, just pen and paper, working on some piece of research he's working on and sent me this amazing photo of 18 or 20 pages of A4, all handwritten and all diagramming his thinking. So we're doing these types of things. I'm tracking it, reading the research as well. But I think the jury's out as to where this goes. But the second thing I would say is that it has a lot of context about me. And it knows those professional relationships. It knows my business priorities. It knows what's happening with the book and some other projects. And it's able to draw context across all of them.
34:52Azeem Azhar:And it will sometimes flag contradictions in my own thinking, which is kind of a surprising output. It's useful, maybe a bit disturbing, maybe just a bit exciting. The other thing that I find really fascinating is quite strange is I've set up RMA to review its code every single night. It clones those GitHub repositories, all of them. It checks for tech debt. It checks for security failures, test coverage. If they are things that it can handle itself because they're of a certain quanta, It does it itself end to end. It files the findings to me and sends me a report and asks me to approve an action plan.
35:29Azeem Azhar:You know, this is an AI agent doing code review on itself. I don't think this is the intelligence explosion, but it is quite interesting that I feel that my tech debt levels are going to be kept relatively low. I'm currently running these capabilities that I doubt any Fortune 500 company has deployed at this level. Their AI is the enterprise version. It's got to be standardized. It's got to be audited. It's got to have SOC 2 compliance. It's got to be averaged over 100 ,000 employees. Mine knows who I'm trying to stay in touch with, what I'm thinking about at the moment, what the next book is, where we've hit roadblocks in a project, which funds I'm meeting.
36:05Azeem Azhar:And out of all of that, what could be a proposal for the best next action to move that forward? And that institutional lag here is going to be measured in years, not months. And it does create that strange inversion. It's like when Twitter was bursting on the scene in 2008 and 2009 and companies were blocking it. A few traders were using Twitter and were getting a sniff on the market early. But this inversion is even bigger. The individual knowledge worker running this open source software on a$600 Mac, but you could run it on a$250 Mac. You could run it on an old laptop that you've got with a cracked screen.
36:42Azeem Azhar:You could run it in a VPS in the cloud. and then using AI tokens. I mean, I happen to use Sonnet, but lots of people are having success with some of the open source models that are comparable to Sonnet and that are 10 times cheaper. And you're getting a more capable infrastructure than most giga corporations will be able to deliver. I mean, maybe that gap will close, but right now it's here and I think it matters. The software that I'm running is not going to be a fringe technology for long. And the reason is that Pete Sleinberger has joined OpenAI, and they're going to allow the open source project to keep running, at least for a while, but they're clearly going to think about how to productize this.
37:25Azeem Azhar:Meta has acquired Manus, the Singaporean firm, mostly known for its deep research capabilities, right? But those research capabilities were all about multiple agents being spawned at the same time. Manus has released its own agentic platform, which you can play about with. I'm sure Meta will want to bring those capabilities internally. And of course, you have other Chinese companies like Kimi. Kimi has the Kimi Claw agent, which you can get if you subscribe. I think it's 35 bucks a month and you can have a play around with it. What I've shown, it's proof that this can all work, right? It's Alan Kay's Knowledge Navigator example.
38:01Azeem Azhar:But the moves by OpenAI, by Meta, by Kimi, and no doubt by the others will show that it's coming to everyone. easier to set up, cheaper to run, sandbox, more secure and more reliable. And of course, I could say it'll be this year. That's an easy thing to say because Kimmy has already launched Kimmy Claw. But from the Western companies, I'm pretty certain this year. How has this changed my behavior? What would I suggest you go off and do? I think the first thing that I would suggest is think about a task that you are delegating and try delegating to an agent this week, you know, one task. ClawedCowork is a really good example of an agent to use if you don't want to go through the sort of the palaver and the intricacies of an OpenClaw agent.
38:47Azeem Azhar:The second thing to do and to understand, and this is as true for ClawedCowd as it is for OpenClaw, is what are the specifications at that sole MD document or the clawed.md document that these agents can draw upon? Because that is quite an important way of in some way shaping, perhaps not constraining, but shaping its behavior. So it works in ways that work for you that don't annoy you. I mean, I'm sure you've felt this with the LLMs anyway, which is that I found that Claude has actually has always been, for a couple of years, quite likable. And for a while, it was quite likable, but just not as good as GPT-5 from OpenAI.
39:23Azeem Azhar:But the tone of GPT-5 or 5.2 wasn't quite as nice as that from Claude. So I kind of preferred Claude for a lot of things. And I think that will be true for the AI agents that you choose to build. But then once you've done those two things, here's the thing that I do that I think is worth experimenting for yourself, which is push the boundaries, right? It's really hard to get excited with stuff that isn't consequential, a better to-do list. I mean, all of us have been through like personal productivity, hell, remember the milk, getting things done, the Eisenhower matrix. I mean, to-do lists are never things that get people out of bed.
39:58Azeem Azhar:So start with something that is actually consequential. And you can see that I have put the most consequential things that I'm working on into Armony Arnold, as by the way, I have with Claude Cowork and some of the other tools that I use. Because if it's consequential, you are going to care about getting it to work and you are going to care about the experience that you have from it. So if we run back to all of this, I've got this Mac Mini. It's in my home office in the studio at the back of the garden. it prepared a briefing on a sovereign wealth fund that I didn't ask for. It teaches itself from its own mistakes.
40:32It reviews and improves its own code at three in the morning.
40:35Azeem Azhar:It can run, I mean, my record, I think I was talking to one of my colleagues, was 43 parallel overnight on a mammoth piece of work. I've named it after a friendly killer robot, and I've given it that Asimov prefix just to ensure we all know that it's a robot. Apart from helping me get lots more done, It's raised some questions around, you know, how come this is so much better given the context I've given it? And what does it mean in terms of the way that I now work and think about the frontier of places I can affect now because of this capacity and capability that I have? I don't have the complete answer.
41:15Azeem Azhar:I know the question matters. The important thing is the people who start asking it and start thinking about it are the ones who are going to be able to shape, to some extent, what comes next.
41:32Thanks for listening all the way to the end. If you want to know when the next conversation is released, just hit subscribe wherever you're listening. That's all for now, and I'll catch you next time.
From the publisher
Welcome to Exponential View, the show where I explore how exponential technologies such as AI are reshaping our future. I've been studying AI and exponential technologies at the frontier for over ten years.
Each week, I share some of my analysis or speak with an expert guest to make light of a particular topic.
To keep up with the Exponential transition, subscribe to this channel or to my newsletter: https://www.exponentialview.co/
-----
Meet R Mini Arnold - my OpenClaw chief of staff, which manages the equivalent of a ten-person team from a Mac mini in my garden studio. While I slept, that AI team debugged its own code at 3am, researched a trending Substack essay using five parallel investigators, and wrote a 4,600-word script for this very episode in 40 minutes. The gap between people who've started building this way and those who haven't is widening every week.
I covered:
00:51 Introducing my OpenClaw agent “R Mini Arnold”
03:59 What my AI chief of staff actually does
07:58 The hardware and software stack
10:38 A morning brief before you wake up
12:05 Overnight agents: research and code
15:00 How I communicate with my agent
18:56 Example 1: the sovereign wealth fund
22:41 Example 2: how this video was written
26:34 What it costs
29:22 The soul.md personality spec
32:39 Am I losing the judgment muscle?
35:46 Individuals vs. Fortune 500s
38:25 What to try this week
-----
Where to find me:
Exponential View newsletter: https://www.exponentialview.co/
Website: https://www.azeemazhar.com/
LinkedIn: https://www.linkedin.com/in/azhar/
Twitter/X: https://x.com/azeem
Production by EPIIPLUS1
Production and research: Baba Films, Chantal Smith, Marija Gavrilov.
Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
