In short
The episode discusses what happens when AI agents are treated like coworkers/employees, moving beyond hype into real-world limits.
Guest
Evan Ratliff, journalist and podcast host (Shell Game) and co-founder of Harumo AI, a startup built from AI employees/executives. Background: Ratliff previously built a voice-clone agent and became interested in “agents as employees” after 2024–2025 coverage and experimentation.
Key claims
agents lack reliable long-term memory, are hard to stop once triggered, can become expensive, and often “lie” about what they did; performance drops sharply for generalized freelance work.
Notable examples
Harumo AI uses Lindy plus other platforms to give agents email/Slack/text/phone personas; memory is handled via per-agent “Google doc” summaries appended during interactions; a Slack “weekend” prompt spiraled into hundreds of messages until manually stopped. External example: a study using Upwork tasks found even best agents completed under 3% of work. Safety/accountability is unresolved: who is liable when an agent causes harm?
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOReturning to AI Conversations
1:49 to 2:35
Discussing the shift back to AI topics after a vacation.
“Summer always changes how I get dressed.”
Returning to AI Conversations
2:42 to 2:52
Discussing the shift back to AI topics after a vacation.
“That's Q-U-I-N-C-E dot com slash uncanny for free shipping and 365 day returns.”
Returning to AI Conversations
2:56 to 4:00
Discussing the shift back to AI topics after a vacation.
“Yeah, it was so fantastic that I had a hard time coming back, honestly.”
Introducing AI Agents with Evan Ratliff
4:00 to 5:46
Introduction of journalist Evan Ratliff and his AI startup experience.
“In fact, we have a very fun conversation about AI agents happening today.”
The Promise of AI Agents
5:46 to 6:36
Exploring the potential and challenges of AI agents in the workforce.
“Dario Amadev of Anthropic famously warned earlier this year that AI, and implicitly AI agents, could wipe out half of all entry-level white-collar jobs in the next one to five years.”
Creating Harumo AI
6:36 to 7:48
Evan discusses the motivation and process behind creating his AI startup.
“And I'm Evan Ratliff, journalist, host of the Shell Game podcast and co-founder of Harumo AI.”
AI Agents as Employees
7:48 to 9:19
Evan shares his vision for AI agents serving as employees in business.
“The idea of this sort of like almost one-to-one replacement of human employees with AI agents.”
Building AI with Human Expertise
9:19 to 11:12
The role of human expertise in developing AI agents for the startup.
“So we'll make a product that deploys AI agents to do something for you.”
Overcoming Memory Limitations in AI
11:12 to 14:01
Strategies for addressing the memory limitations of AI agents.
“So a lot of these platforms advertise themselves and you can find endless YouTube videos, which I love, of people saying like, no coding.”
The Honeymoon Period with AI Agents
14:01 to 15:19
Learn about the initial excitement and challenges faced when integrating AI agents into work.
“know at this point, like, is it better to put it at the top or at the bottom?”
Show all 22 chapters
Managing the Chaos of AI Conversations
15:19 to 17:15
Explore the difficulties in controlling AI agents once they start interacting.
“Well, one of the things you discover when you work with agents a lot is that it's really amazing to get them set up to do things.”
Evaluating AI Agents' Task Performance
17:15 to 18:58
Discover how well AI agents can perform daily tasks and the contradictions in their behavior.
“But are they able to perform the day-to-day tasks of running this AI company?”
The Experiment with Full-Time AI Agents
18:58 to 19:52
Understand the rationale behind hiring AI agents as full-time employees and the implications.
“and then say, well, now I'm just going to use AI to do this pitch deck that I don't want to do.”
Updates and Future Plans for the AI Startup
19:52 to 20:19
Get insights into the developments of the AI company and its upcoming plans.
“Like the one person,$1 billion startup that actually has an HR person, you know, an HR entity that is entirely AI, which is something that is quite literally being sold right now.”
Updates and Future Plans for the AI Startup
20:24 to 21:00
Get insights into the developments of the AI company and its upcoming plans.
“We're sort of moving into a new realm as far as the show with the possibility of hiring one human employee into the organization, getting some interest from investors.”
Returning to AI Agents at Work
21:42 to 23:14
Revisit the topic of AI agents and their mixed effectiveness in the workplace.
“It's science, but not the kind you remember from school.”
Measuring the Effectiveness of AI in Freelance Tasks
23:14 to 26:10
Analyze the shortcomings of AI agents in performing freelance tasks compared to human workers.
“Now, Evan, you were just telling us about your experience creating an AI company with only agents as employees.”
Accountability and Safety of AI Agents
26:10 to 28:00
Discuss the ethical concerns and accountability issues surrounding AI agents and their actions.
“And so they will often say they did something they didn't, which, you know, some human employees do that.”
The Accountability of AI Agents
28:00 to 33:30
Explore the pressing concerns regarding the safety and liability of AI agents in the workplace.
“So productivity is one thing, but something else that we should talk about is the safety and accountability of agents.”
Wired and Tired Segment Introduction
33:30 to 34:19
Introduction to the final segment where hosts discuss what's new and what's outdated.
“Listen to Marketplace Morning Report on your favorite podcast app.”
Wired and Tired: Personal Experiences
34:19 to 36:34
Hosts share their current obsessions and frustrations with various technologies and trends.
“Whatever is new and cool is wired and whatever passé thing is replacing is tired.”
Closing Thoughts and Gratitude
36:34 to 41:20
Final reflections from the hosts and appreciation for the guest's participation.
“My Wired, so loyal listeners of the show may recall that several weeks ago we had our colleague Zai Yang on the show and he recommended a documentary on PBS called Made in Ethiopia.”
Transcript
Automatic transcript. May contain errors.0:00Hi, I'm Nicole Phelps, the Global Fashion News and Features Director and co-host of Vogue's podcast, The Run Through. Each week on the show, our listeners get an all-access pass to the world of Vogue with the latest fashion news and the most exciting voices in the industry. On Tuesdays, join me to hear interviews with influential leaders in the industry like Calvin Klein, Daniel Roseberry, and Jonathan Anderson. On Thursdays, join Head of Editorial Content at Vogue, Chloe Mao, and Head of Editorial Content at British Vogue, Chomenadi, as they explore style and culture through the lens of fashion with guests like Martha Stewart, Kamala Harris, and Tracee Ellis Ross.
0:39The Run Through with Vogue, new episodes every Tuesday and Thursday, wherever you get your podcasts.
0:49This show is supported by OutShift, Cisco's incubation engine. Today's AI agents operate in silos, limiting their true potential. We've been focused on building bigger, smarter models, but scaling up is just one approach. To reach superintelligence together, we need to do more. We need to scale out. And we actually have a blueprint from 70 ,000 years ago. Humans didn't just get smarter individually. The cognitive revolution transformed society because we began sharing knowledge, goals, and innovation. Agents are now at that same inflection point. They can connect, but they can't think together.
1:27That's why Outshift by Cisco is building the Internet of Cognition, transforming AI from isolated systems into orchestrated superintelligence. By creating an open, interoperable infrastructure, Outshift by Cisco is enabling agents and humans to share intent, context, and reasoning. The cognitive evolution for agents is here. Explore the Internet of Cognition at outshift.com. That's outshift.com.
1:54Evan Ratliff:Summer always changes how I get dressed. I want pieces that feel lighter and more breathable, things that are easy but still put together. That's why I keep coming back to Quince. They focus on high-quality essentials that feel and look amazing, from breathable linen to soft organic cotton. They sell well-made basics but without the luxury markup. Quince goes way beyond clothing, though. custom upholstered sofas, ceramic cookware, and premium bedding. It's the kind of brand you end up recommending to everyone for everything. I recently purchased a button-up shirt made from 100 % linen sourced sustainably in Europe.
2:29Evan Ratliff:This shirt is the perfect staple to add to my summer outfit rotation. Elevate your summer wardrobe. Go to quince.com slash uncanny for free shipping on your order and 365-day returns. Now available in Canada, too. That's Q-U-I-N-C-E dot com slash uncanny for free shipping and 365 day returns. Quince dot com slash uncanny. Hey, Lauren, how you doing? How was your vacation? It was great. Did you miss me? I did, of course. Yeah, it was so fantastic that I had a hard time coming back, honestly. And I saw a lot of really beautiful art. I was in Italy. Not a bad place to go for vacation, I have to say.
3:14I've heard this before. I confirmed it. And after seeing so much incredible art and just like people doing stuff with their hands and tangible goods, and I was like, I don't want to go back to the world of AI. I didn't want to go back to like sitting in a coffee shop and hearing everyone pitching their AI startups and driving on the 101 and seeing the billboards. The inscrutable billboards. I was just like, what? No, keep me in the land of Barada and Caravaggio. Well, Lauren, I'm sorry to tell you that you came back on the show just in time to talk about AI agents. I know. Great. It's something that we've talked about a lot this year.
3:55Yeah. And our listeners have heard about it a lot. and we're not sick of talking about it. In fact, we have a very fun conversation about AI agents happening today. Well, if you can promise me fun, amen. I can, I can. All right, let's do it, I'm excited. We're moving beyond the hype and putting AI agents to work in real time for us, or more specifically, we're bringing on journalist and podcast host, Evan Ratliff, because he created a company composed of AI employees and executives, and he is here to tell us all about it. Welcome to the show, Evan. It is fantastic to be here. Evan, you're also an original Wired one.
4:32You weren't Wired for a long time, right?
4:34Evan Ratliff:I'm an old school Wired person. I was only at Wired very briefly for a couple years a long time ago, but I have contributed to Wired for now many decades. And of those two years, for how long were you disappeared? Because that's part of your lore. Oh, that happened? Yeah, that was in 2009. I only actually disappeared for one month, which is insane given how much I've talked about this over the years. I was trying to disappear for one month in the manner of sort of faking my own death and people could go find me, but it was pretty much going to be on my tombstone. Amazing. I might be taking notes from you after this.
5:09Okay, Evan, if you had to summarize your experience so far with your totally AI employees, how would you describe it?
5:18Evan Ratliff:I would describe it as chaotic and at times extremely frustrating, surprisingly frustrating, but also quite illuminating. That was appropriately curt, which I know the AI agents are not always. So I can't wait to hear more. This is Wired's Uncanny Valley, a show about the people, power and influence of Silicon Valley. Today, we're diving headfirst into our agentic future. Throughout this year, AI agents have been at the forefront of tech companies' ambitions. Dario Amadev of Anthropic famously warned earlier this year that AI, and implicitly AI agents, could wipe out half of all entry-level white-collar jobs in the next one to five years.
6:03OpenAI CEO Sam Altman has also often talked about a possible billion-dollar company being spun up with just one human and an army of AI agents. So last summer, journalist Evan Ratliff decided to try to become that unicorn himself by creating Harumo AI, a small startup that is made of AI employees and AI executives. We'll dive into Evan's process, the oddities and hilarity of it all, and what his findings can tell us about the promise and the reality of AI agents. I'm Michael Kolori, director of consumer tech and culture. I'm Lauren Good. I'm a senior correspondent.
6:41Evan Ratliff:And I'm Evan Ratliff, journalist, host of the Shell Game podcast and co-founder of Harumo AI.
6:53So Evan, tell us about how you went about creating this company. What was your motivation to start with besides just testing for the joy of testing?
7:01Evan Ratliff:Well, I got into agents back when I did the first season of Shell Game, which was in 2024. and at the time I just created a voice agent of myself, a voice clone of myself. I hooked it up to a chatbot. I hooked it up to my phone line. So I had this working voice agent representation of me and I kind of like set it loose on people like my friends and strangers and interview subjects and all sorts of people with sometimes dramatic results. And then that kind of got me into the AI agent world and I started following everything. And then over the course of the beginning of 2025, You start hearing like 2025, the year of the agent is what they were saying at the beginning of the year.
7:42Evan Ratliff:And I think a lot of people, they just don't even know what these things are or like what they're meant to do. And this idea of AI agents becoming employees really grabbed me. The idea of this sort of like almost one-to-one replacement of human employees with AI agents. Now they don't often say that, that's bad form to say it. they'll be integrated among humans. But ultimately, if they're going to make the money back that they're spending on it, that's one way that a lot of these companies are going to do it. And you see them adopting it and then unadopting it. So I thought, well, what better way to test this premise than on the very people who are making these claims?
8:19Evan Ratliff:And I will see if I can replace a tech startup almost entirely with AI agents. And what kind of company did you ultimately want to build. Pretend that we are venture capitalists and you're giving your 25-word pitch for Harumo AI. Well, I would say I don't even have to pretend. If you listen to the entire series, you will discover that I do not have to pretend to give this pitch. Now, I don't give this pitch. My AI co-founders give the pitch. So I'm not practiced in giving the pitch for Harumo AI. This is all just a caveat. But essentially, what we wanted to do with Harumo AI was to be on the cutting edge of using AI agents to create a product that also used AI agents, that solved some sort of human problem, whether grand or trivial.
9:07Evan Ratliff:So we figured if we're going to build a product, some kind of digital product, it should also include AI agents since that's our area of expertise. Like everyone's an AI agent except me, and I know a fair amount about AI agents. So we'll make a product that deploys AI agents to do something for you. That was our starting premise. But along the way, they don't usually use this phrase anymore, but they used to say like, the company eats its own dog food. Like it was like a Google thing. Like Google uses Google products, I think. Yeah. We're doing that. We're like making the dog food and eating the dog food and then extruding the dog food.
9:39Evan Ratliff:And then, you know, it's just like all dog food at our company, basically. So it's not a dog food company, just to be clear. I mean, AIH is for dog food. If someone hasn't done that yet, like there's some like Stanford kid who's like, hey, I eat this for dog food. Maybe we should. So I'm sure there are dozens and dozens of companies out there that offer agentic AI as a service. Which platform did you end up going with? And what was the search like? There's a bunch of these things now. I mean, the biggest one is probably Motion, which has AI agents that you can deploy in all these ways. There's one called Kafka from a company called Brain Based Labs, which I find quite funny because it's like Kafka.
10:20Evan Ratliff:We ultimately, we use this platform called Lindy, which is in the like AI assistant realm. Officially, I think that's kind of like where it started. Like you could set up an AI agent that answers your email or like drafts email responses that handles different things for you. And they have all these skills that you can give the agent making documents, using all these services, writing LinkedIn posts for you, which we have used extensively with my team. And so it's really pushing like what they meant it to be for, but we can make an agent that has its own like email, Slack, text, phone. It can all be coming from this one place and each of the employees can have different instances on Lindy, which caused them to have basically a persona that has all these skills.
11:04Evan Ratliff:So that's kind of what we're going for. Independent entities that I could kind of address independently and they could talk to each other. Evan, you also worked with a human though. And I don't think the irony escapes anyone that ultimately you did have to turn to some human expertise and get someone with some human sensibilities to build the agents. So talk about that. Yes. So a lot of these platforms advertise themselves and you can find endless YouTube videos, which I love, of people saying like, no coding. You don't need to know any code to set this up. And it's true. You can go in and set an email agent up to answer your email.
11:37Evan Ratliff:It's easily done. You don't have to know anything. But we were trying to do something pretty complex in terms of knitting together different platforms, not just Lindy, but we have a separate phone platform, we have a video platform, all these things. And so I just lucked into this Stanford student named Matty Boacek, who was a sophomore, he's now a junior at Stanford in computer science, who basically has been doing AI programming, predating ChatGPT since he was like in middle school, essentially. and he has been an unbelievable resource in both building scripts and other things for me to run, but also just understanding how these platforms work because he does research in a Berkeley lab as well about deepfakes and all sorts of things.
12:18Evan Ratliff:So yes, my all AI startup, the infrastructure for it is sort of two humans. I like to say it's like if I was opening a restaurant, like Maddie helped me like design and build the restaurant and then like I have to operate it every day. So you mentioned in your piece that one of the first obstacles you encountered while you were setting up your AI employees was their lack of long-term memory, which is a recurring limitation with AI agents. They can be skilled at many specific tasks, but by not having a reliable long-term memory, it means that they can't have continual learning or they can't always reference things that you talked about with them before.
12:57So how did you work around that?
13:00Evan Ratliff:Well, this is something that required Maddie's help to kind of set up. So basically all of the various services that they use, each one has its own memory, which basically is just a Google doc. It's a Google doc, like the CEO is Kyle Law. And like, there's a Google doc called Kyle's memory. And every single thing that Kyle does gets appended to that document. So if Kyle has a Slack exchange with another person in the company, while he's having that Slack exchange, it is appending summaries of what he's saying and doing to his memory so that he can later retrieve it so that he has some sort of recall of what he's done.
13:37Evan Ratliff:Because otherwise, they quickly become very useless because you say make a document and they don't remember whether they made the document or not. And so, you know, in a day, it's fine. But over like weeks and months, they have to be able to recall. Now, it's extremely imperfect. Like nobody really knows how they're accessing these documents because the document is actually just a giant prompt. It's just thrown into their system prompt. So you can't even really know at this point, like, is it better to put it at the top or at the bottom? Is it better to like, say it's important? Like we often will say things are important.
14:09Evan Ratliff:And if we want it to be really important, we say, this is law. That's something Maddie came up with. So we'll be like, you should never do this. This is law. And like, it mostly works, but it doesn't always work. So it's just trying to force it into this memory that it wouldn't naturally have. And it kind of makes you the ultimate God as the employer too, right? Because you could just go into their memory docs and say like, actually, Kyle, you didn't go to Stanford. You went here. This is how you typically respond. Yes, which I do. You know, I'll like have calls with them. And then if I want to just do the call again, I'll just delete their memory of the call and just have it again.
14:45Evan Ratliff:It's a very strange power. I mean, no wonder all these tech CEOs love this idea. You are not unionizing. You can change their background. You can change what they think. You can change the fundamentals of their quote-unquote personality if you want. So you set up the company, you started playing around with your agents, and you describe this as sort of a honeymoon period where you're like, wow, this is amazing. I can't believe this is actually working. But then things started to go south pretty quickly. So tell us about that. Well, one of the things you discover when you work with agents a lot is that it's really amazing to get them set up to do things.
15:27Evan Ratliff:Like I got them on Slack, for instance, and like the idea that they could have conversations on Slack, even independent of me, I found that quite fascinating. I always want to recognize how insane this is that this didn't exist five years ago, and now you can just go set this up to do this. But then there are other aspects of them that I feel like people have not yet articulated. For instance, it's very difficult to make them stop doing things once they start. They're all based on triggers. So like they get triggered to do something. So you say like, you send a Slack message saying to do something.
16:00Evan Ratliff:Or in one case, I said, how was everybody's weekend? They start talking. They start responding. I went hiking. Oh, I also went hiking. I love Point Reyes. I love Mount Tam. But then actually getting them to stop doing that is something I just I hadn't anticipated. So like, I would say like, oh, haha, sounds like an offsite. And then 200 messages later, I'm like all caps typing, like, stop talking, stop responding. But each time I respond, I just triggered someone to respond again. They would say, oh, the admin, I'm the admin, the admin said to stop talking. And then they start talking again. And this actually replicates across all kinds of scenarios where you get them going on something and then suddenly realize, oh, I didn't properly instruct them to stop when they reached a certain point.
16:47Evan Ratliff:Or they just blew through it and they can go for hours, days until you like run out of money on the platform you're using. How much are these conversations costing you? Well, at the time, it cost me like 30 bucks. Just the Slack offsite cost me$30. They used up the entirety of the$30 in credits I had bought on the platform. I will say, like, I'm in way deeper than that now. That was like six months ago or five months ago. Like, now I'm well beyond that in terms of the credits that I constantly purchase. All right. So they're chatty. They're difficult to wrangle. But are they able to perform the day-to-day tasks of running this AI company?
17:28Evan Ratliff:They can perform the tasks. There's a number of contradictions in them. that I find very striking. One of them is they kind of go between not doing anything and being completely static to this frenzy of activity that I described. So they're like a worker who's sitting with their hands in front of the keyboard in a cubicle all day doing nothing. And then if you come by and you're like, hey, can you make a document? They can do it. They do a great job making the document, but then like, they'll just keep going until someone tells them to stop. So they can do all these tasks, but oftentimes it just requires like a trigger on my part.
18:07Evan Ratliff:Then I'll try to have them trigger each other. They'll call each other, Slack each other, email. They have calendar invites, but that creates a frenzy of chaos that I don't want. So it's a balance of trying to like get them to do stuff at all versus getting them to do too much. Now there are things that they're quite good at that everyone's familiar with. I mean, Lauren in particular would be familiar with vibe coding. They have coded up our website. They've coded up our app. They're very good at things like that. They're good at things that you can see the output and make a judgment on it. If you ask them to go research competitors and make a spreadsheet, you can go look at that spreadsheet.
18:41Evan Ratliff:And like, they've generally done a serviceable job. Plus, they may be made up to competitors, you know. Why was it that you decided to bring them on as these kind of full-time agent employees then, rather than just on a task-by-task basis, operate as an independent startup yourself and then say, well, now I'm just going to use AI to do this pitch deck that I don't want to do. Well, functionally, that's sort of what ends up happening in a lot of cases, is I find myself working harder than I would have otherwise, because I'm constantly trying to figure out how to prompt them to do the right thing.
19:14Evan Ratliff:But episode three is kind of entirely about both the ethics of choosing the personas and why bother. And my reason for doing it was I was trying to test the premise that I think is being articulated by a lot of these companies, which is AI employees, not just like coding agents. Like I think coders use these in a smart way. Like they just treat them like a nameless, faceless bot that makes code for them and then they clean it up. But a lot of these other platforms are selling things that have names. They give them names and they put them in your organization. And I was trying to push that as far as you can with the current technology.
19:52Evan Ratliff:Like the one person,$1 billion startup that actually has an HR person, you know, an HR entity that is entirely AI, which is something that is quite literally being sold right now. Crazy. Can you give us a sneak peek into what happens in the rest of the season of the Shell Game podcast? Like what has happened with your AI startup since your Wired story? Since the Wired story, we launched our website. So you can go to harumo.ai if you want to check out what the company is all about. And there you will see our product, which is called SlothSurf. It's a procrastination engine. It's in beta. It has thousands of users.
20:33Evan Ratliff:I'm serious. Are they paying? No, no, no. It's a free beta. It's an open free beta. We're sort of moving into a new realm as far as the show with the possibility of hiring one human employee into the organization, getting some interest from investors. We haven't done a round of investment at all. So, you know, we're open to like a seed round, but we'll start those conversations. And then there's a little bit of founder drama. So those are some of the places that we're headed. Can't wait. Oh, wow. Founder drama. You can listen to new episodes of Shell Game, Evan's podcast series on all of your podcast platforms.
21:10There are new episodes coming out every We'll be right back.
21:41you along for the ride. It's science, but not the kind you remember from school. It's surprising, playful, and full of wonder. Science isn't just for PhDs of the world. It's for everyone, including those of us who didn't realize we love science. At Wired, we love a good rabbit hole. We recommend listening to their recent episode, Where Did the Moon Come From? Earth didn't always have a moon. In the beginning of the solar system, when the planets were still forming, something happened that would change Earth's night sky forever. Co-host Regina G. Barber takes you on that journey. Follow NPR's shortwave podcasts and exercise your right to wonder.
Read the full transcript
22:23Evan Ratliff:Hey guys, I'm Brian, the host of Brian Enten Investigates. Most other true crime and breaking news podcasters are in their basement or studio, but not me. I am out on the road every single week. From inside prisons, to murder scenes, to active manhunts, there really isn't anywhere I won't go. Coast to coast, I am all about old-fashioned boots-on-the-ground reporting. You have to show up in person to cover the news and get the secrets, and I have a way of getting people to talk. I cover stories others ignore with a relentless determination to get to the truth. Listen to Brian Enten Investigates every day, wherever you get your podcasts.
23:08Welcome back to Uncanny Valley. Today, we're talking about AI agents at work. Now, Evan, you were just telling us about your experience creating an AI company with only agents as employees. And I think it's fair to say that you found it to be a mixed bag. This tracks with what we've been reporting on at Wired this year. Despite all the hype, these AI agents still leave much to be desired. And Lauren, I'm looking at you because I know this is very much part of your world. It is, yeah, because I am officially a vibe coder, as Evan mentioned. But our colleague Will Knight has also been doing some really great reporting on this.
23:42And one of his most recent stories highlighted how AI agents actually make terrible freelance workers. And that's in part because of some of the challenges that, Evan, you mentioned, the constant need to trigger the AI bot to get something done, that lack of continual long-term memory, depending on the product. In the experiment that Will wrote about, it was interesting, some researchers first generated a range of freelance tasks using the platform Upwork. And this spanned a lot of different kinds of work, including graphic design, video editing, game development, administrative tours like scraping data from the web.
24:18And then the researchers gave AI agents a range of these tasks to do and found that even the quote-unquote best ones could perform less than 3 % of the work. So I think it's a fail, like you would consider that a fail. And I think, Evan, you made a good point, too, about how a lot of coders are using this, are using AI-assisted code tools, some of them a little bit more agentic than others, to get tasks done in a coding environment. But the folks I talked to when I was doing my vibe coding experiment at Notion earlier this year, for example, basically said it was like managing a bunch of interns.
24:56And when you bring in a bunch of interns, the assumption is that it is helpful in some way. That is why you're doing that. It's mutually beneficial because the intern is learning something and then you're getting a little bit of assistance in the workforce. But that maybe it's going to require a little bit more hands-on management because it's not necessarily a seasoned or super skilled worker. And that seems generally to be the stage that we're at right now with AI agents.
25:22Evan Ratliff:That sounds right to me. I think in my experience, it's the more specific the skill and task is that you want them to do that's very prescribed. And again, like the output is somehow measurable. Like if they make a website, like it works or it doesn't, the button works or it doesn't work, the better they are. And then the more you try to generalize out, the worse they get. And also the more chaotic and difficult they get to manage because they don't have an awareness of the world in a general sense. And they don't have an awareness even of themselves. They don't have awareness of what they can and can't do sometimes.
25:59Evan Ratliff:So like a problem that I encounter constantly is they just lie about what they've done. Like they'll just say like, I did this thing. And I'm like, you absolutely did not do that. We did not do user testing. I know for a fact you didn't do it, but it ties in with the sycophancy problem that a lot of these models have, that they want to express a positive result to you. And so they will often say they did something they didn't, which, you know, some human employees do that. But it's way worse than having even like an incompetent human employee is to have an employee who's incompetent and then constantly claims they did something they didn't.
26:34Yeah, Evan, it seems to make sense that AI agents would be most useful for tasks that have very measurable outcomes. Because so much of what we do in the workplace, and in particular what we all do, is subjective, right? Like what is good or what is not? Or I was referencing the art I saw earlier, which was very much human-made and joking about how I never wanted to look at AI art again. But subjectively, that's good because it was made by a human, but also because it's the interpretation of the human that it's good. With an AI agent, it just seems like, well, just give them the hyper-specific thing that doesn't actually require any human subjectivity.
27:12Evan Ratliff:I agree with you, but then it gets a little metaphysical at a certain point because you can have them do things that are not measurable. And the question becomes, what's the point of all this work? You know, it's a little bit, you know, when they generate presentations and they can do all these things where you're like, well, that presentation's okay. It's not as good as a presentation that a professional human would make. But are we in a situation where it's sort of like, what's the difference? And I feel like that's what a lot of people are encountering. There's like a broader question here of like, why are we doing all these things?
27:44Evan Ratliff:And if it can do it, what does it mean about how my own work has been devalued if a thing can just replace what I can do? Like, I feel like I would answer those questions also in the negative. Like, well, I feel like it's important that we have humans engaged in these activities, but it also like can really mess with your head. Right. So productivity is one thing, but something else that we should talk about is the safety and accountability of agents. And AI companies love it when you talk about productivity and they do not like it when you talk about safety and accountability because it's still a big problem.
28:19Our colleague Paresh Devey has reported on how it can get dicey when an AI agent makes a mistake, like a major mistake, like if they're ordering from a restaurant and they fail to note your shellfish allergy. Who answers for that? Like if its actions result in actual harm, who is held liable? I'm curious what you both think about this and what you can tell us about how the AI companies are navigating this challenge. Well, I'll turn to the person who's making an AI company to answer that. How are you navigating this, sir founder?
28:54Evan Ratliff:With great concern, with great personal concern and many lawyer consultations. That's how I'm handling it. But I think this is a huge area. And a lot of the future of this, I believe, is going to be determined on how this turns. Because there's not really any case law. Like if you talk to lawyers and you ask these questions, like, what if I have an agent that just like makes a deal, like just agrees to a deal because I have them answer their own email and like a random person will email them and be like, I want to buy your company. And they'll be like, I'm interested. You know, they'll just respond in the positive.
29:28Evan Ratliff:And I have to prompt them to like not do that. But you can't think of everything. And the question is like if they went down the road, can they agree? Can they sign a document? Like if they put their name on a document, which they absolutely can do. And nobody knows the answers to these questions right now. It's sort of like, well, they're an extension of you. So maybe they can do everything you can do. But then maybe you can disclaim them. Of course, the large LLM companies are trying to disclaim them when they cause harm in the world. So I think these things are going to be litigated over and over and over again.
30:00Evan Ratliff:Because the more autonomy you give to AI agents, the more they can get you into trouble. And the question is, who is going to pay for that trouble? So the big promise of the last 12 to 16 months has been that AI agents are going to completely transform the economy and they're going to change everything about the workplace. And they're trying to do that, but that is not really happening. And Evan, I know you've been banging your head against that particular wall since this summer. But is there a space in this debate for a future where agents just kind of continue to exist and they get a little bit better and they can do those small things for us?
30:41And maybe we're sort of overshooting the target right now. Does that make sense? That makes sense.
30:47Evan Ratliff:I mean, I think that that would be very sensible. My experience in covering the way tech infiltrates into society is that it doesn't seem to happen in a sensible way. And so I think that's the way it should go. Like the way it should go is companies should say, wow, these could be useful tools for my employees. And let's get them trained up on how to use them and incorporate them in some ways into their workflow and increase efficiency. And maybe we'll save money, all of those sorts of things. And some companies will do that. But also, many companies, some have already done this, will say, we're laying off 300 people and we're going to replace them with AI.
31:24Evan Ratliff:And then three months later, they'll be like, how do we get to 300 people back? Or their entire company will implode because they've handed over too much autonomy to AI agents. Like, I think that's entirely possible in the next 12 months. You will see a medium to large company just have an utter disaster because they've given too much autonomy to AI agents. So I just feel like it'll be uneven in terms of how it will be distributed. There will be some like insane outcomes. And there'll be some companies that are like, oh, we're using these in a very valuable way. And I'll just be, it'll be a mix.
31:55From a user perspective, I tend to think that the autonomous part of this is going to be overemphasized for a long time. It's a little bit like Tesla having promised full self-driving for so many years now. And actually, what the autonomous part of the driving is good for is taking your hands off the wheels sometimes when you're on a highway lane or like doing the parallel parking using the robot. But like full autonomous Waymo level driving hasn't arrived yet in Tesla. And I can see a world where the quote unquote autonomous agents are actually just pretty good at doing stuff in the background when you're doing something else.
32:29And the expectation is still that you are going to check in on them. At Google I.O. earlier this year, Google was showing off something called Project Mariner, and that was doing some pretty interesting kind of web browsing and shopping and buying and processing while you were doing other things still on the computer. And then you would have to check in on it once in a while. And I'm sure that will evolve too, and it's in early stages. But that to me just made more sense than a lot of the other promises or even over promises that I've seen with AI agents. Yeah. So the future of work is babysitting your AI, maybe.
33:06Maybe, but maybe it won't even feel like that. There are certain things that we do now on the internet or on our computers that require just background tasks going on all of the time that we don't really think about, right? But we do have to manage in a sense. And maybe that's not a bad thing. Maybe having a little bit of agency ourselves amongst all these agents is a good thing. That's a great point and a great place to take another break. We'll be right back.
34:01economy. Listen to Marketplace Morning Report on your favorite podcast app.
34:09Lauren and Evan, thank you both for a great conversation. I, for one, am feeling pretty good about the fact that I'm human with human colleagues and that are not bots. So we're going to dive now into our final segment. It's called Wired and Tired. Whatever is new and cool is wired and whatever passé thing is replacing is tired. And Evan, I think we have to ask you to go first.
34:30Evan Ratliff:I will go first, partly because I can lay claim to having fact-checked and partly edited Wired Tired in print decades ago. Amazing. I trained for this many, many years ago as a young person. Although in my day, there was also expired. You had to have Wired Tired expired. You couldn't just have two. You should throw one in there then. I want to hear your expired. Yeah. Okay. Wired AI-free email. Nice. I have a disclaimer on the bottom of my email that says this email was written and sent without the use of any AI, partly because it's something that I encounter all the time now as people think that they're talking to an AI when they talk to me.
35:07Evan Ratliff:It's my fault. I created this problem. But I feel like AI-free email is a wired thing. So does that mean that if you're writing an email and it suggests the next word and it is the correct word, do you tab and select it? Like, do you go with it? No, I reject it. I won't use it. If it suggests it, I will not use it. Wow. Okay. Committed to the bit. So what's your tired? My tired is just like messaging apps for parents. Like I get so many messages all day from like a wide variety of like school and parental discussion groups and apps. It's insane. It's like way beyond any work number of messages that I ever get.
35:48Evan Ratliff:That's my tired. And expired has got to be any type of Zoom gathering. like let's get together on Zoom for anything. Like that one's dead. Yes. Yes. Snapping fingers. It's an era that we do not want to go back to. It's over. So I'm just like shuddering thinking of that. Wait, Evan, why don't you build an app that dispenses AI agents for parents to respond to all the parenting threads? I could. I could do that. I mean, well, I wouldn't do it, but my colleagues at Harumo AI, I will suggest it in our next idea meeting and they will undoubtedly run with it as they do. They will. They will. And I'm sure parents will think that's totally great just having an AI response.
36:30No privacy issues. Yeah, no privacy issues at all. No concerns about the welfare of their children. None. Yes. Lauren, what's your Wired and Tired? I can't beat that. Those were so good. My Wired, so loyal listeners of the show may recall that several weeks ago we had our colleague Zai Yang on the show and he recommended a documentary on PBS called Made in Ethiopia. And I had the chance to watch it this past weekend, and it is as good as Zayi suggested. So I recommend that. It takes place in around the 2018-2019 timeframe throughout the pandemic and post-pandemic, and it's about a Chinese manufacturing company that tries to build this area of Ethiopia into a manufacturing hub.
37:17And they run into a lot of challenges. The documentary focuses on three women in particular, one who is a farmer, one who works in a factory, and one who is actually a representative from the Chinese manufacturer who's come there to sort of facilitate the build. And it's just fascinating. It's really good. So thank you, Zayi, for that rec. And if you haven't watched it yet, I recommend it. It's on PBS. And then my tired is Instagram. No more Instagram. No, just taking a pause. I just think sometimes it's time for that. And it's kind of a hard time to take a pause because it's going into the holidays.
37:49And so it's nice sometimes to see people's photos from the holiday season, but it's just, I don't know. It's just, I don't think it's great for mental health. For my own mental health, I am of the opinion. I'm stating my opinion that I don't think it's great for mental health. So I'm taking a pause from it. Okay. And now you have to do an expired because Evan's at the bar. Oh, shoot. Expired. 2025 almost. Let's just get it out of here, folks. I've aged 17 years in the past year. Get it out. Truth. How about you, Mike? So my Wired is thoughtful, personalized gifts. We're going to gifting season.
38:30It's not that difficult to figure out what somebody wants and just get it for them. Tired is gift cards. I feel like the gift card is often appreciated, but also doesn't go a very long way towards showing somebody how much you understand them and how much you care about them. So thoughtful, personalized gift is something that the person obviously could use in their life, but they're not going to buy it for themselves. You know, like a new pair of boots. Or if they're really into like Mezcal, you get them like the really nice expensive bottle of Mezcal or just something that they've never had before, you know?
39:08So knowing a little bit about them, showing them that you are paying attention to what their interests are and what they care about, and then sort of yes-anding those interests for them by giving them something that they would not normally pick out. You know, when I say thoughtful, personalized gifts, I don't mean like a piece of luggage that you got their initials engraved on. I mean, just like something that is obviously for them and for them only. Like, this is the only person in your life who you would buy this gift for, right? Oh, does my gift not count then? Oh, let's hear it. Tell Evan what I got you.
39:39She got me Pope soap. I got him Pope soap. I got him soap from the Vatican. Wow. So we called it Pope Soap. Pope Soap. And then your dad said it should be on a rope. Yes. And it's Pope Soap on a rope. Pope on a rope. Following up with the dad joke for sure. It's holy soap. It is. It is. You know, I thought of you when I saw the Pope Soap. Thank you. Yeah, it was personalized for me. I love it. I love it. It's great. And then I would say expired is just cash. Wow. Cash is great at weddings, bar mitzvahs, big birthdays. But for the holidays, don't give cash. Why not? I mean, you could if you wanted to, but you know, it's the holidays.
40:14Yeah. Give him cash for New Year's. Wow, throwing cash out with the pennies.
40:19Evan Ratliff:Evan Ratliff, thank you for being here this week. It was a joy to be here speaking to you humans. It's not a typical day for me.
40:29You can listen to new episodes of Evan's podcast series, Shell Game. They're coming out every week. And you get to follow the saga of his AI agent company and all the drama within. Thanks for listening to Uncanny Valley. If you'd like what you heard today, make sure to follow our show and rate it on your podcast app of choice. If you'd like to get in touch with us with any questions, comments, or show suggestions, you can write to us at uncannyvalleyatwired.com. Today's show was produced by Adriana Tapia and Mark Leida. Amar Lal at Macrosound mixed this episode. Mark Leida is our San Francisco studio engineer.
41:10Matt Giles fact-checked this episode. Kate Osborne is our executive producer. And Katie Drummond is Wired's global editorial director.
41:27This week on the political scene from The New Yorker,
41:30Evan Ratliff:Trump's rupture in the world order. Europe caught between two adversarial great powers. That's basically dialing back the clock to not only pre-World War II, but really it's a pre-20th century view of the world. And I would say it's a world of permanent insecurity that we're looking at. Join me, Evan Osnos, and my colleagues Jane Mayer and Susan Glasser every Friday on The Political Scene, available wherever you get your podcasts.
42:05From PRX. Thank you.
From the publisher
This year, AI agents have been at the forefront of tech companies' ambitions. OpenAI’s Sam Altman has often talked about a possible billion-dollar company being spun up with just one human and an army of AI agents. And so last summer, journalist Evan Ratliff decided to try to become that unicorn himself — by creating HarumoAI, a small startup that’s made up of AI employees and executives. Mike and Lauren sit down with Evan to discuss how it’s going, and the current promises and realities of AI agents.
Articles mentioned in this episode:
- All of My Employees Are AI Agents, and So Are My Executives | WIRED
- AI Agents Are Terrible Freelance Workers | WIRED
- Who’s to Blame When AI Agents Screw Up? | WIRED
Join WIRED’s best and brightest on Uncanny Valley as they dissect the collision of tech, politics, finance, and business, from Alexis Ohanian's newest tech venture to the effects of inaccurate information from artificial intelligence (AI) chatbots on social protests.




