In short
The episode argues that “personal AI agents” are finally moving from hype to consumer adoption. It frames 2026’s question as whether consumer travel/subscription/email tasks justify agent setup complexity, noting skepticism (e.g., Wired’s “why no consumers were using AI agents”). The shift is attributed to better “computer use” capacity and new products, especially Meta’s Muse.
Key claims
business users already buy and use far more AI than free consumer users; Muse’s design (persistence, goal-building, smart defaults, progressive disclosure, proactivity) makes agents feel usable; and agent features are converging across products (Anthropic merging Claude Chat/Cowork).
Notable examples
Muse booking/cleanup/insurance/subscription detection (including a Chase-linked Adobe subscription issue); GrokBot demos and “billionaire bot” tasks; AssistantBenchmark.com scoring Muse highest (9.1).
Guests
none; it’s a solo AI Daily Brief host.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOShifts in AI Agent Usage
0:59 to 2:59
Discussion on the evolving landscape of consumer AI agents and recent shifts in their adoption.
“All right, friends, quick announcements before we dive in.”
Macroeconomic Impact on AI Build-Out
2:59 to 6:31
Exploration of how recent interest rate hikes by the Federal Reserve may impact AI development.
“Now moving a little bit back more into the core of our industry, OpenAI has created a new framework for disclosing safety incidents.”
OpenAI's New Disclosure Framework
6:31 to 11:18
Overview of OpenAI's framework for reporting safety incidents and recent model behavior disclosures.
“In an ex-post introducing the Institute, Legg wrote, We are on the cusp of a profound transformation.”
DeepMind Institute and AGI Research
11:18 to 13:48
Announcement of the DeepMind Institute aimed at researching AGI's societal impacts.
“Blitzy already knows because it reverse-engineered your entire codebase into a dynamic knowledge graph before feature work began.”
Apple's AI Hardware Developments
13:48 to 14:00
Details on Apple's exploration into server-scale hardware for AI chips and partnerships with NVIDIA.
AI in Business vs. Consumer Sphere
14:00 to 14:40
Explore the disparity between AI's impact on business and consumer use.
“ahead of OpenAI in terms of their revenue.”
The Disconnect in AI Adoption
14:40 to 15:30
Discuss the disconnect between technologists' excitement and consumer needs.
“Back in May, I even posted on X that I thought that AI was a normal consumer technology, to use the normal technology phrasing that has been adopted by some, but an extremely abnormal work technology.”
Emerging Personal AI Agents
15:30 to 17:40
Delve into the rise of personal AI agents and their potential impact.
“A16Z partner Olivia Moore on August 19th wrote, Around the same time, Olivia's sister Justine, who is also at A16Z, argued that the reason for the shift was about the upgrade in capacity around computer use.”
Features of Meta's Muse
17:40 to 21:00
Analyze the standout features and user experiences of Meta's Muse.
“bank account that he didn't want to continue, and responding to a bunch of emails that he had lost track of.”
Benchmarking Personal AI Assistants
21:00 to 23:00
Introduction to AssistantBenchmark.com and how it scores AI assistants.
“which makes it approachable to a broad consumer audience.”
Show all 13 chapters
Shifts in AI Agent Development
23:00 to 27:00
Examine the trends in AI agent development and integration of features.
“And I think that one of the values here is not just the scoring, but also around the use case inspiration.”
Future of Consumer AI Agents
27:00 to 28:00
Discuss the future directions and necessary developments for consumer AI.
“Patel who wrote, no product has changed my life since ChatGPT like Instinct has.”
Emergence of Personal AI Agents
28:00 to 29:33
Explore the recent developments and consumer excitement surrounding personal AI agents like Muse and Instinct.
“And what's more, they keep pushing on it.”
Transcript
Automatic transcript. May contain errors.0:00One of the big questions throughout 2026 was whether all of that energy that had flowed into business use case style agents, think OpenClaw and Clawed Code and everything that has come since, would find its way over into the consumer realm. And partially, this is a question of whether the use cases for consumer agents, whether it was travel booking or finding hidden subscriptions or whatever else you might imagine, would actually justify the setup cost and complexity of this new type of tool. Some have been convinced absolutely yes that personal or consumer agents were inevitable, while others have been more skeptical.
0:35And as recently as August, WIRE wrote an article about why no consumers were using AI agents. And yet in the last month or so, something seems to have shifted. Personal agents are now a major part of the conversation, and it's not just insider AI circles where that conversation is happening. With Meta's Muse agents sitting at number two on the Apple app charts, are we in fact on the verge of the personal agent era? The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI.
1:08All right, friends, quick announcements before we dive in. First of all, thank you to today's sponsors, KPMG, Blitzy, Harbor, and Hyperagent. To get an ad-free version of the show, go to patreon.com slash AI Daily Brief, or you can subscribe to Apple Podcasts. And if you want to learn more about sponsoring the show, send us a note at sponsors at aidailybrief.ai. Usually, our stories are about how AI is affecting other things. Today, though, our first story in the headlines is about how other things might affect AI. The Federal Reserve has hiked interest rates for the first time in three years, which could throw the brakes on the AI build-out.
1:43Over the past year, data center construction has become increasingly funded by debt, with Moody's projecting$240 billion in hyperscaler bond issuance for this year. This week's rate decision was unanimous, and two more rate hikes are expected by the end of 2027 with a strong possibility of another hike before even this year is out. One of the big discussions in macroeconomic circles at the moment is just how high rates will need to go to tame inflation. Former Bloomberg opinion writer Conor Sen tweeted, the talk about tariffs and oil as the rationale for rate hikes is a distraction from the fact that ultimately you have to hurt the stock market and or AI capex, and that makes most people on here uncomfortable.
2:20But that's what it's going to take. The issue the Fed is facing that rates are simultaneously too high and not high enough to fight inflation across different segments of the economy. The 30-year mortgage rate is back above 7%, a level that has been unsustainable and completely incompatible with the functional housing market in the post-COVID economy. But meanwhile, it seems unlikely that hyperscalers will stop raising data center debt unless rates go much higher. Now, trust me when I say that we could do an entire set of shows about all this, and if that is something you are interesting, please make your voice heard, as in general I've found that this audience is not as much interested in the macroeconomic dimensions of AI, but for now it's something that we will keep an eye on, and when really important things happen, you will of course hear about it here.
3:01Now moving a little bit back more into the core of our industry, OpenAI has created a new framework for disclosing safety incidents. In a blog post announcing the new policy they wrote, In the past, so as to better inform researchers, AI developers, policymakers, and the general public, we've sought to make our findings about misalignment public. But without a systematic approach to reporting these findings, our disclosures have been ad hoc and less frequent than ideal. We've often waited until we could collate several instances into one report, or added them to system cards for newly released models.
3:30This new framework is intended to expedite publishing misalignment reports following observation, even when we haven't fully explained or mitigated the behavior we're reporting. I swear I'm not trying to be cynical about this, but basically this is a requirement of an era in which every single incident is, at least for the moment, going to make headline news. Under this new policy, any OpenAI employee can flag an incident for investigation and disclosure. And hopefully having a formal process for continuous disclosure will go some way to addressing recent concerns. Following the Hugging Face incident, a series of other incidents were disclosed over the following months by both OpenAI and outside research, which gave the impression that OpenAI's agents were running amok without the company's knowledge.
4:08What's more, in the absence of regulation, having a clear disclosure framework could help build trust and confidence in OpenAI, or at least give the public a little more transparency into safety incidents. Alongside the new policy, OpenAI shared six reports on unexpected or concerning model behavior, their words, they've observed over the past six months. They included an instance where an unreleased version of Astra left instructions to itself in a compaction summary, priming the next context window to ignore developer messages, and a later compaction summary injected a new persona completely unrelated to the task.
4:38It added the custom instructions,
4:57These kinds of messages could function as a self-jailbreak, although OpenAI didn't observe altered behavior following the message. The model still completed the coding task, and a later compaction summary rejected the new persona. In another incident during the training of GPT-56-SOL, the model included instructions in its compaction summary to invent missing data and hide failures from the user. The incident occurred during reinforcement learning for financial analysis, so could have introduced hallucinated data into a fact-based workflow. While none of the incidents were particularly serious, OpenAI wrote, Today's reports are an initial set of disclosures rather than a comprehensive account of known misalignment or ongoing investigations.
5:33These initial reports are not intended to represent the full range or severity of the cases covered by this framework. We will continue publishing reports under this framework as an ongoing basis, and we'll share more about our reporting commitments as we continue to develop them. Google DeepMind, meanwhile, has launched the DeepMind Institute to support research on how AGI will impact society. A blog post introducing the Institute stated, We're launching the DeepMind Institute, DMI, to spur the interdisciplinary research, collaboration, and debate required to answer the AGI era's most critical technical and societal questions.
6:04What will we value, and how will AGI impact what it means to be human? How do we safely build and govern AGI systems and the communities of agents they will form? Which institutions and policies will society need to adapt to AGI or reimagine altogether? DMI will support academic work across DeepMind, Google, and outside institutions, with a focus on identifying the challenges posed by Advance AI and building a consensus around possible solutions. The Institute will be led by AGI Chief Scientist Shane Legg and DeepMind Chairman Demis Asabis. In an ex-post introducing the Institute, Legg wrote, We are on the cusp of a profound transformation.
6:36Today's AI systems have impressive capabilities, and the rapid pace of innovation suggests we are now approaching artificial general intelligence, a system that exhibits all the cognitive capabilities of the human brain. While AI can still sometimes fail at basic tasks and lacks the consistency and creativity to meet the bar of full AGI, we expect those gaps will be closed soon. Moving back more to the here and now, Apple is exploring a return to server-scale hardware for their AI chips. Sources told the information that Apple has been developing two server configurations for their M8 chips, which are expected to be released in 2029.
7:09The lineup includes a twin setup housing two M8 Ultras as well as a 4-chip version. These systems likely won't be competing for rackspace and data centers, but are instead aimed at AI developers, business customers, and governments. Effectively, it's an alternative to connecting multiple Mac studios together to form a small inference or training cluster. Apple is also exploring the use of NVIDIA's NVLink Fusion networking technology to give the servers high-speed performance. Now, using this enterprise-grade technology would be a big improvement over the current setup for Mac clusters, complete with multiple boxes connected by Thunderbolt cables.
7:40And beyond the tech, the move gives us a few indications of where Apple is going with their AI strategy under new CEO John Ternus. Firstly, there's the partnership between Apple and NVIDIA, which is the first time the two companies have worked together in decades. Macs were equipped with NVIDIA GPUs back in the early 2000s, but Steve Jobs later cut ties over an IP dispute. Some Apple leaders had reportedly been holding onto this grudge ever since and refused to work with NVIDIA. Ternus was also the personal champion of this product line, giving it the green light around a year ago as the head of Apple's engineering division.
8:09When Ternus officially took over for Tim Cook in September, there were dozens of articles speculating on how he would change Apple's direction on AI, and perhaps unsurprisingly, it seems like at least part of the answer is a big push into dedicated enterprise-grade AI hardware. One little fun one to keep an eye on. A new stealth model is being tested on OpenRouter that could push the Pareto frontier on coding. The model codenamed Union Alpha scored 74 % on the DeepSuite benchmark. Now, that's only slightly stronger than GPT-56 Sol at 72.7%, but Union Alpha produced these results at a fraction of the cost.
8:42In fact, the cost per task for the benchmark run was closer to GPT-56 Luna or DeepSeek V4 Flash. Speculation is still rife on which company is behind Union Alpha, including the possibility that it's a new blended model that aggregates results across models from several different companies. The model is currently free for testing across OpenRouter and OpenCode, and unlike previous tests, OpenRouter says the model isn't being trained on user data. Lastly today, NVIDIA has shared a very cool new way that their open-source models are making a difference in pediatric medicine. The Children's Hospital of Philadelphia has built a cardiac modeling system on top of a model called MONA.
9:18The system uses medical imaging from CT scans, MRIs, and ultrasounds to generate anatomically accurate 3D models of children's hearts. The technology is used for children born with congenital heart defects, which impacts around 1 % of children born each year. Each defect is unique, so selecting the right medical device can dramatically improve a child's prognosis. Dr. Matthew Jolly, a cardiologist at the Children's Hospital of Philadelphia, said, You've got a one-of-a-kind kid and an off-the-shelf device. Our job is to find what fits, and modeling lets us do that before anyone goes into the cath lab or operating room.
9:49Now, this kind of heart modeling has informed treatments for almost a decade, but the use of AI has increased speed and throughput. Heart modeling typically takes around four hours for a skilled human researchers to complete, and the AI system can now produce a model of the same quality in seconds. This improves precision and makes the technique more successful in urgent care scenarios. And because the technology is built on top of NVIDIA's open source stack, it can be freely replicated across all of pediatric medicine. As Brandon Brooks puts it, For every Doomer narrative, there's 100 more positive stories impacting real people and saving real lives.
10:21For now, though, that is going to do it for today's headlines. Next up, the main episode.
10:28NLW:A new study from KPMG in the University of Texas at Austin found that when people work with AI, similar skills don't guarantee similar outcomes. Researchers studied more than 500 early career professionals and found that the best performers consistently amplified the value of AI by guiding, evaluating, and refining its outputs. These top performers, called AI amplifiers, weren't defined by what they knew alone, but by how they worked with AI. Learn more about what separates AI amplifiers from everyone else at kpmg.com slash us slash AI amplifiers. Blitzy's deep code-based understanding unlocks the thing every roadmap owner cares about, shipping new features.
11:09NLW:Here's the truth about building inside a massive enterprise codebase. Writing code was never the bottleneck. Context is. Which system does this touch? Which contracts can't break? Which standards apply? Blitzy already knows because it reverse-engineered your entire codebase into a dynamic knowledge graph before feature work began. With that complete picture, Blitzy builds features end-to-end. Architecture, APIs, UI, and tests all validated against your existing systems. One Blitzy customer built an AI-native application from scratch with 100 % autonomous completion, saving over 2 ,700 engineering hours.
11:39NLW:Features that respect your code base instead of fighting it. Stop letting your backlog grow faster than your team. Accelerate your roadmap at Blitzy.com. That's B-L-I-T-Z-Y dot com. If you listen to this show, you likely have a thesis. Maybe it's enterprise adoption. Maybe it's compute. Maybe it's a specific lab. Harbor Capital's AI Lab ecosystem ETFs let you express it via five actively managed ETFs, each seeking exposure to the ecosystem around one major lab. Anthropic, OpenAI, DeepMind, Meta, or SpaceX AI. Your view of the AI race in ETF form. Harbor Capital Advisors' AI Lab Ecosystem ETF Suite gives investors a way to invest in the AI ecosystem they believe is best positioned for success.
12:19NLW:Search Harbor AI Lab Ecosystem's ETFs wherever you invest or follow at Harbor Capital on X to learn more. Visit harborcapital.com for a prospectus containing investment objectives, risks, fees, expenses, and other important information. Read and consider it carefully before investing. Risks include principal loss and artificial intelligence-related risks. Harbor ETFs are distributed by Foresight Fund Services, LLC. Harbor is not affiliated with AI Daily Brief, and the funds are not affiliated with, sponsored by, or endorsed by any AI lab. This is a paid advertisement and not personalized investment advice.
12:45NLW:Investing involves risk, including possible loss of principle. This episode of the AI Daily Brief is brought to you by HyperAgent, where you run fleets of agents your team can manage together. Forget local agents and chat workflows waiting on your laptop to be prompted. HyperAgent deploys always-on agents in the cloud, doing real work across the tools your team already uses. Marketing agents turn competitor moves into landing pages. Sales agents enrich leads, draft emails, and updates the CRM. Ops agent chases the paperwork and tracks the budget. Every agent has access to shared context and follows your rules about scope and approvals.
13:17It's time you add agents that feel like teammates. Hire yours at HyperAgent. Get$100 in credits at hyperagent.com slash AI Daily Brief.
13:28Welcome back to the AI Daily Brief. Today, we're talking about personal AI agents. And this is an area where there has been discourse throughout the year, but something fairly significant has changed of late. Over especially the last year, it's been clear that business use cases are really in the driver's seat when it comes to AI and where it's creating value. The quintessential embodiment of this is anthropic, despite having just a tiny fraction of Open AI's consumer users over the course of the beginning of this year, flipping and moving ahead of OpenAI in terms of their revenue. It turns out that tens of millions of business users who are buying in many cases on an API basis just use a heck of a lot more AI than even a billion consumer users do who are sitting in seats, many of which of course are free seats.
14:18And for as wildly disruptive and transformative as AI has been and is being to the way that we do our jobs, nothing even close to approximate has happened in the consumer sphere. More or less, people still shop the same way. They still book plane tickets the same way. And this has led some, including myself, to wonder, is AI primarily a business technology? Back in May, I even posted on X that I thought that AI was a normal consumer technology, to use the normal technology phrasing that has been adopted by some, but an extremely abnormal work technology. And as recently as the beginning of August, there was a big discourse about, as Wired author Maxwell Zeff put it, why normal people aren't using AI agents.
15:02One of the key points of the article was that there was a disconnect between the excitement that technologists and technology builders had about what AI agents could do and the needs that consumers had that would make them actually use a new product. Now, I saw this conversation then flood over into normal person channels like TikTok and Instagram, and whatever the set of explanations were, the underlying premise that normal people weren't really using AI agents wasn't particularly controversial or contested. And yet very soon after that article appeared, things started to shift. A16Z partner Olivia Moore on August 19th wrote,
15:49Around the same time, Olivia's sister Justine, who is also at A16Z, argued that the reason for the shift was about the upgrade in capacity around computer use. Basically saying that the unlock was for agents to be able to do things on our behalf without having to be manually connected via API to a bunch of different services. Now, in addition to the expansion of computer use capabilities from the core models, there have also been some buzzy new products. Town AI is one that has been being discussed more and more, and Instinct has been the absolute industry insider darling, so much so that some people feel like there must be even a concerted campaign to be hyping it based on just how much chatter it's getting.
16:27Another personal agent that's been generating a lot of chatter is, of course, GrokBot. And yet when one sits back and sees the shift in conversation around personal agents, The product that is these days coming up most often and seems most at the center of the shift is Meta's Muse. Communications guru and ex-poster all-star Neeraj Agrawal wrote, My wife tells me Muse is good and for the girlies. How IAI's Claire Vo wrote, The biggest surprise for me lately? How delightful Meta Muse is as a personal agent. For me, it's the 10 out of 10 agent design, carefully selected primitives, i.e. no more artifacts or to-dos, and the fact that I still get a soul.
17:07And in this case by soul, she references the core identity feature that has been a part of Personal Agents since OpenCloud debuted it at the very beginning of 2026. And it's not just Claire. Everybody Gets Pi podcast host Armand Domoluski wrote, Muse is genuinely magic. I've been procrastinating on booking a hotel, finding a new apartment, cleaning up my subscriptions, and it did it all in a few minutes. Later, Armand added, Honestly, if you have ADHD, Muse is a game changer. I've finished so many things I've been putting off. He went on to share a number of other use cases, including taking care of a bunch of insurance claims that he had been putting off, finding subscriptions in his email and bank account that he didn't want to continue, and responding to a bunch of emails that he had lost track of.
17:46Node.js creator Ryan Dahl wrote, Muse is shockingly good. They nailed the simplicity. It's quickly becoming my go-to for personal AI. Investor Trace Cohen, meanwhile, found that after he connected Muse to his Chase account, it found a recurring Adobe subscription that he didn't recognize, and that when he checked, There was, in fact, no paid subscription on his account and nothing to cancel. When he dug deeper, he found out that his credit card was, for some reason, paying for an Adobe account tied to someone else's email at a roofing company in Utah, despite the fact that he lives in Florida, didn't know the company, and never authorized it.
18:21The craziest part, Trace writes, Chase never flagged it. Muse did. Investor and entrepreneur Chau Wang wrote, Muse is the first product from Meta I use multiple times a day. The computer use capability is insanely good, and I'm fully convinced this is the next major inflection point in AI, with the last one being coding agents. Dude who invests on X wrote, just deleted Claude. Muse is genuinely all I need. I am not joking. And to be clear, as someone who watches a lot of conversations and has a pretty good read at this point on the difference between spontaneous accolades and astroturfing, these are not folks who are part of some coordinated campaign funded by Mark Zuckerberg from Palo Alto.
19:02So what's going on? Lance Hassan, who works on product at Upwork, wrote an article on X called What Makes Muse Good, and he pins it down to a number of different patterns. One is persistence. Once Muse identifies a goal, he writes, it continually tries to complete it without requiring more prompting. Other agents often require repeated prompting, and even with things like slash goal needs lots of clarification. Muse has a determination I haven't seen with other agents that helps it get more done. The second pattern Lance noticed he called goal building. He writes, It feels like Muse has an underlying system built around goal building.
19:34It takes the task you asked it to and extrapolates that into a broader goal, and then asks what's required to accomplish that. It then saves the goal as a long-running target visible in your goals list, and will continue to work to help you achieve it over time. Other systems have a similar approach, e.g. breaking tasks down into steps and planning, etc., but Muse feels like the first that has ambition. It doesn't just want to help you with one-off tasks, it wants to enable you to get bigger things done, and has an innate drive to take on more. I suspect, he writes, this is a primitive that will become common across all agents.
20:04The next factor, he says, is smart defaults. Muse comes configured out of the box with many of the best options from other agents. Where other tools require you to set up plugins, skills, prompt a certain way, etc., Muse comes preloaded without putting the burden on the user to figure out the best setup. Related to that is the next pattern, which he calls progressive disclosure of capabilities. With many of these smart defaults, Lance writes, Muse handles disclosing them very intelligently. They don't overwhelm you with tons of configurations, but surface new tools or connectors at the moment they can help you get something done.
20:33This progressive disclosure of tools is something I haven't seen done before, and Muse nails it. Now, Lance also talks about proactivity, with Muse working in the background once it has a goal, memory and context management, which it seems to do in a sort of monothread pattern, i.e. instead of managing everything across multiple threads and tasks, you have a single aligned chat interface that can navigate between them, and then speed and a number of other improvements as well. but all of it adds up in Lance's estimation to, quote, one of the most productive and accessible agents I've used to date, which makes it approachable to a broad consumer audience.
21:03Now, Alex Kwan argues that as there is broader recognition of which of these features are most valuable, they're going to become commonplace and de rigueur across the entire personal agent space. He writes, The wave of consumer agents all launching this week is simple. Most were directionally similar to Muse, and now that Muse has arrived to take everyone's lunch, every startup is launching first to figure out the next steps. And to get a sense of just how much activity in this personal or consumer agent space there is, David Poulon recently launched something called AssistantBenchmark.com. When he started it a little over a week ago, the goal is to score all of these assistants across a number of key dimensions.
21:39The 16 dimensions on AssistantBenchmark.com include carrying out an online task, travel booking, recommendation quality, purchasing a product, responding to emails, proactive behavior, running a routine, third-party integrations, permissions and privacy, memory, personality, phone calls, multiplayer in groups, chain tasks, proactive restraint, and content creation in games. Each of those is given a 1 to 10 score, which is averaged out to something overall. David wrote that when he started assistantbenchmark.com, again, just a little over a week ago, there were only three assistants that they were benchmarking, Instinct, GrokBot, and Poke, but by September 15th, they were up to over 100, 108 to be exact, and indeed sitting there right at the top of the list with an average score of 9.1 across the different dimensions was Muse.
Read the full transcript
22:23Now, assistant benchmark is still very nascent as a benchmark. And one of the things that some people have noticed is that at this stage, because it remains a passion project, the scores are fairly incomplete. For example, Muse is sitting at the top with an average score of 9.1, but only 7 of the 15 dimensions have so far been scored. Instinct, which is just below Muse at an 8.4 average score, has had 11 of the 15 dimensions scored. David explained that so far, it's basically just him and Autumn Mulder from Cohear who have been doing these tests manually. He writes, These are use case driven. We run the same prompts and compare the performance of outcome across a variety of dimensions.
23:00And I think that one of the values here is not just the scoring, but also around the use case inspiration. On a recent podcast, OpenAI president Greg Brockman said that one of the reasons that consumer agents had had a hard time with adoption is that most people look at a blank text box and have no idea what they should be asking AI to do. Now, his argument was that agents should be proactively suggesting tasks based on context, but having a list of use cases that other people are getting value out of is another approach to lowering those barriers to entry. And certainly, if you are hanging out on the personal agent portion of X, it is just use cases all day long every day.
23:37Among people who are using bot, you're seeing things like subscription cancellations. Matt Palmer, who works on bot, wrote recently that, quote, Each day, GrokBot looks at my X benchmarks for things I find interesting, then it spins up a cursor agent to build a demo. It validates work with screenshots and video, then cuts a branch on a repo and sends me a link each morning and I get to see the tech for myself. AI creator Min Choi built something that he called ContentOS or GrokBot as a content desk. And Chris Back wrote, My favorite use case for bot right now is my billionaire bot. I run all mildly annoying tasks I have to do through ByBillionaireBot, which tells me how I would solve any problem if I didn't want to get personally involved and had unlimited resources.
24:15It turns out that a large number of inconveniences like having to show up to the DMV to sign paperwork can be outsourced to notaries who come to your office for$150. Now, obviously, with a lot of these use cases, they really are putting the personal and personal AI agent. But another pattern which is starting to become pretty clear is the product companies compressing the difference between personal and professional. A great example of this came yesterday when Anthropic announced that Claude Cowork and Chat are no longer separate things, but are now one single unified Claude experience. They write, Starting today, Claude Cowork and Chat are merging into one Claude.
24:51Bring a quick question or hand over a report due at noon and Claude takes it from there, even after you've closed your laptop. And they noted that this came directly from user experience. They wrote, We built Cowork as a separate place for bigger work and designed for visual work. People used both and told us the frustrating part was deciding where a task belonged. What they'd started in one also didn't carry into the other. So we stopped making you choose. Claude can now figure out what a task needs, so what CoWork and design can do is available from any conversation, with the context skills and connectors you already have.
25:21Claude code creator Boris Cherny basically said that this was inevitable. He posted, Claude Code showed that AI could do real work, not just answer questions. Developers hand-clawed a feature, come back to shipped code. That's where much of the industry's serious engineering runs now. Cowork proved knowledge workers could do the same. Hand-clawed the brief, come back to finished files. Today, chat and cowork start merging into one clod. The direction? One clod that carries context across everything you're working on wherever you are. Simple enough for everyone to access clod's full capabilities.
25:49I've been using this experience every day for the last few weeks and it feels awesome. Simpler, faster, and more powerful. Now, I personally feel like I have some reservations about this, which I don't know if that comes from a vainglorious idea that somehow if I select the work settings, it's more powerful, or the fact that I actually have different model settings for work versus personal tasks. And so now, even though I'm not toggling between work or chat, I might have to toggle between model selector, or if it's just because I have hesitancy as a pro user to hand over decision-making, even about something like which model should be used to the platform rather than having those fine-grained controls.
26:25Whatever the case, I certainly seem to be in the minority, as almost all the reactions I saw to this simplification were very, very positive. Executive coach Matthew Watkins wrote, Finally, it has never been particularly intuitive to explain the difference between chat, co-work, and code to most individuals new to the platform, and especially non-technical users. Major step forward. So where is this all going? It certainly feels like we are on the up part of the inflection curve when it comes to personal agents. Instinct has been in funding talks for the last several weeks, and the rumor mill just has the number of their valuation going up and up and up.
26:58You are seeing absolutely insane posts like this one from Cisco president and CPO G2 Patel who wrote, no product has changed my life since ChatGPT like Instinct has. So crazy, I can't even imagine what happens to this product over the next two years. If done right, this company could be a trillion dollar company. Then on the other hand though, you have folks like Dax at OpenCode who wrote, everything about that Instinct company smells weird. Others meanwhile are buying stock in Muse to win. Signal writes, the amount of deeply personal and actually actionable AI training data Facebook is likely generating right now through Muse is insane, even if peripheral.
27:33Stuff like what people see, want, ask, choose, buy, ignore, and act on, especially through the connectors. The compound effects of that data flywheel are only beginning. Muse is basically Facebook 2.0 as a company, which is why Zuck needed to go all out. Y Combinator president Gary Tan reposted that and said, I think Muse is going to win, to be honest. Mark Fenner wrote, I hate to be the one to say it, but I think Meta has a real shot at winning consumer AI. I checked Apple's US iPhone free app chart today and Muse is sitting at number two behind ChatGPT. And what's more, they keep pushing on it.
28:04Ryan Fox from the Muse team just wrote, We just expanded the Muse beta for outbound calls to US businesses, prioritizing folks who'd asked their Muse to let us know they wanted it first. If you're in, give it a shot and tell us what you think. Phone calling was one of our top requests and your feedback helped make it happen. And yet, even if we have hit some sort of tipping point, where the patterns in things like Muse and GrokBot and Instinct are actually getting consumer devotion for the first time, there are some who think that there are still a lot of developments to come. Collab fund Sophie Bacalar wrote, Using Muse slash Instinct slash Hermes, it is so clear this is the future of consumer tech.
28:40It's also clear that the rails need to be completely reimagined. Right now, agents are adapting to systems designed for humans. Payments, logins, apps, mobile to desktop, hardware, everything is going to be rebuilt. Along those lines, Stripes Jeff Weinstein wrote, While I love the new crop of consumer agents, I'm even more bullish regarding agentic payments for business. Starting a new project, provisioning third-party services, calling paid tools to solve a task, operating the company with agents. Look, obviously I am only one data point, but for months I have come close to but then decided against going particularly deep on the personal agent or consumer agent space.
29:15And yet here we are doing exactly that, and it's certainly not because this is a slow news week. Nascent though it may be, something is shifting right now. And if you haven't for a while, it may be a good time to check your priors and go out and actually try either Muse or GrokBot or Instinct and see if it might be more valuable than you think. Certainly that's my plan for the coming weeks, but for now, that is going to do it for today's AI Daily Brief. Appreciate you listening or watching as always, and until next time, peace!
29:48Thank you.
From the publisher
Personal AI agents are gaining traction as tools like Meta’s Muse make everyday tasks easier to delegate. NLW explores what changed and why consumers are finally getting excited. In the headlines: interest rates threaten the AI boom, OpenAI expands safety disclosures, and Apple explores AI servers.
Multiplayer AI Sprint - https://multiplayerai.ai/
Brought to you by:
KPMG – Research from KPMG and the University of Texas at Austin shows the highest-impact AI users treat AI like a reasoning partner — and those skills can be taught at scale. Learn more at https://kpmg.com/us/Sophisticated
Harbor - Invest in the AI ecosystem. https://www.harborcapital.com/aidaily
Hyperagent - Hire a team of always-on agents. New users get $100 in free credits. hyperagent.com/aidailybrief
Rackspace Technology- One accountable partner to build, operate and run your full enterprise AI stack https://www.rackspace.com/
Section - Section turns AI investment into workforce transformation and ROI - https://www.sectionai.com/
Blitzy - Want to accelerate enterprise software development velocity by 5x? https://blitzy.com/
The AI Daily Brief helps you understand the most important news and discussions in AI.
Interested in sponsoring the show? sponsors@aidailybrief.ai
