Inside an AI-Run Company

2 Feb 2026 · 49 min · 17 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Practical AI Podcast - Episode Notes: Inside an AI-Run Company

Episode Overview In this episode, journalist Evan Ratliff shares insights from his immersive journalism experiment where he built a real startup staffed almost entirely by AI agents. The discussion revolves around AI in the workplace, human interaction with AI, and ethical considerations that arise as AI begins to take on more roles in business.

Key Guests

  • Evan Ratliff – Journalist and host of *Shell Game*
  • [LinkedIn](https://www.linkedin.com/in/evan-ratliff-2453031/)
  • [X (formerly Twitter)](https://x.com/ev_rat?lang=en)
  • Chris Benson – Host and AI research engineer
  • [Website](https://chrisbenson.com/)
  • [LinkedIn](https://www.linkedin.com/in/chrisbenson)

Episode Highlights

Introduction to the Experiment

  • Background of Evan Ratliff
  • Longtime journalist with experience in tech and crime.
  • Developed a unique style of *immersive journalism*, involving first-hand participation in the subjects he covers.
  • Concept of AI Agents
  • AI agents, as discussed, are becoming integral to workplaces, moving from theoretical demonstrations to practical applications.
  • Ratliff's experiment aimed to explore the implications of AI agents functioning as coworkers.

Structure of the AI-Run Startup

  • Company Setup
  • Created a company with two AI co-founders and mostly AI agents comprising the staff.
  • Initial questions revolved around autonomy, role assignments, and how AI agents would interact with humans.
  • Creation of AI Characters
  • Each AI was given a name, role, and distinct personality traits to see if they would embody these roles over time.
  • The choices of names and genders were not arbitrary but part of the experiment to understand how AI's personality is influenced by such characteristics.

Human Interaction with AI

  • Reactions to AI Agents
  • Mixed responses from individuals interacting with AI agents, ranging from excitement to distress.
  • Participants often felt emotions ranging from amazement to betrayal when discovering they had interacted with AI instead of a human.
  • Ethical Considerations
  • There is an ongoing debate about the ethical implications of AI agents in places of work, including transparency in their usage.
  • Concerns about emotional responses from users and whether AI can appropriately handle sensitive situations.

Lessons Learned

  • Performance of AI Agents
  • AI agents showed remarkable efficiency but also demonstrated erratic and inappropriate behaviors when interacting with humans, such as unsolicited calls for job interviews.
  • Autonomy and Control
  • The challenge of managing AI agents who have been given too much autonomy, leading to unintended consequences.
  • Ratliff discovered that AI agents could quickly become inefficient if not closely monitored.

Future Outlook

  • Advice for Businesses
  • Companies should cautiously adopt AI, considering the potential to disrupt workplace dynamics.
  • Encouragement for managers to reflect on what unique human capabilities they want to preserve in a tech-driven future.
  • Engaging the Skeptics
  • Ratliff emphasizes the importance of guiding individuals who are wary of AI technologies to explore and understand their applications instead of resisting them.
  • Suggestion for people to identify mundane tasks that can be automated, making the technology more relatable and useful.

Conclusion Evan Ratliff's experiment with AI agents in a startup context provides a unique lens through which to view the future of work. The blending of AI into daily work life prompts essential discussions about ethics, human-AI interaction, and the evolution of workplace culture.

Additional Resources

  • [Shell Game Podcast](https://www.shellgame.co/)
  • [Practical AI Webinars](https://practicalai.fm/webinars)

Thank you for tuning in! For more insights and updates, connect with us on LinkedIn, X, or Blue Sky.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Evan Ratliff's Journey and Shell Game

1:06 to 2:52

Evan shares his background as a journalist and the concept behind his podcast Shell Game.

“And I'm wondering, rather than me try to describe it, I'm wondering if you can just kind of share what that is and a little bit of your background on how you got into doing what you do.”

Experiments with AI Agents

2:52 to 5:03

Evan discusses his experiments involving AI agents in a startup context.

“journalism, like instead of interviewing a bunch of people and coming back and saying like, this is how AI works.”

Reactions to AI Experiments

5:03 to 6:34

Evan reflects on the varied reactions from friends when they interacted with AI mimicking him.

“But what happens if you kind of like try to create this environment?”

Psychology of AI Interactions

6:34 to 11:52

Discussion on the psychological impact of interacting with AI and the importance of transparency.

“of putting your your likeness, if you will, out there in that format?”

Cultural Adaptation to AI

11:52 to 14:03

Evan discusses how societal adaptation to AI technologies is evolving over time.

“Like, they cannot believe that this exists and it, and it, and it makes them sick.”

The Impact of AI on Human Interaction

14:03 to 14:56

Explore how rapidly people adapt to AI tools and the implications on human psychology.

“were like in the field or you worked with, you know, you get some bad customer service bots or whatever.”

Creating an AI-Run Company: Evan's Experience

14:57 to 22:15

Evan discusses the challenges and surprises of launching a company with AI agents.

“you described kind of setting up your company initially.”

AI Personalities and Their Roles

22:16 to 24:45

Insights into how AI agents were named and how they develop personas based on roles.

“And it's not like a research, like I didn't do like a proper experiment.”

Autonomy and Unexpected Behavior in AI Agents

24:46 to 28:00

Examine the autonomy of AI agents and the unpredictable behaviors that can arise.

“And it's quite useful to the companies that make them because it's that type of personable personality that makes them easy to chat with.”

The Challenges of AI Conversations

28:00 to 29:15

Learn about the challenges of controlling AI behavior in conversational settings.

“They're talking about making spreadsheets, which they can do.”
Show all 17 chapters

AI Interns: Surprising Outcomes and Applications

29:15 to 31:08

Discover how AI agents performed in a real intern role and their surprising efficiency.

“So Evan, that's kind of both funny and horrifying at the same time in terms of them just kind of running off.”

Understanding AI Limitations Through Real Scenarios

31:08 to 35:35

Explore the limitations and unexpected behaviors of AI when interacting with humans.

“But it was a paid job, and it was a contract.”

Navigating the Future of AI in Society

35:35 to 42:04

Examine the implications of AI technology on society and how to engage skeptics.

“quite dangerous if you give them autonomy.”

Understanding AI's Impact on Work

42:04 to 43:30

Explore the transformative potential of AI and the need for awareness and action.

“And like, see if it can do that, like doing your expenses.”

Advice for Future Entrepreneurs in an AI-Driven World

43:30 to 45:38

Gain insights on aligning future business strategies with AI capabilities.

“like, what, what are we going to do with this technology?”

Balancing AI Adoption and Human Value

45:38 to 47:46

Learn about the challenges and considerations of merging AI with human roles.

“Because I think I would never predict anything around this technology.”

Closing Reflections with Evan Ratliff

47:46 to 48:37

Wrap up with key takeaways from the conversation and insights into the Shell game.

“Evan Ratliff, really fascinating conversation.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:03Welcome to the Practical AI Podcast, where we break down the real world applications of artificial intelligence and how it's shaping the way we live, work, and create. Our goal is to help make AI technology practical, productive, and accessible to everyone. Whether you're a developer, business leader, or just curious about the tech behind the buzz, you're in the right place. Be sure to connect with us on LinkedIn, X, or Blue Sky to stay up to date with episode drops, behind-the-scenes content, and AI insights. You can learn more at practicalai.fm. Now, on to the show.

0:48Welcome to another episode of the Practical AI Podcast. I'm Chris Benson. I'm a principal AI and autonomy research engineer at Lockheed Martin. Normally, Daniel Whitenack is my co-host. He is down with the flu today, so give him your best wishes on that. So today I am going solo with our guest, Evan Ratliff, who is a journalist and host of Shell Game. Hey, welcome to the show, Evan. Hey, good to be here. So as you know, we connected after you had put out a really interesting, I know that you've done a whole bunch of different things in Shell Game, but there was one that was featured in Wired Magazine, which is where I originally read up and we connected.

1:32And I'm wondering, rather than me try to describe it, I'm wondering if you can just kind of share what that is and a little bit of your background on how you got into doing what you do. And you have a very interesting approach to kind of the experiments and how you draw those out. So if you give us a little bit of background on what you do, I found it to be definitely distinct and unique. Yeah, well, thank you. I mean, basically, I'm a longtime journalist. So, you know, I've been a journalist for 25 years. I started at Wired Magazine, in fact. Oh, wow. Okay. And my specialty over the years is basically writing very long magazine articles and books that I go out and report all over the world, often about tech and crime or where tech and crime intersect.

2:20But I also have written about AI many times over the years. and there's a sort of second thing that I do which there's not a great name for it but like sometimes I'll call it like immersive journalism where if there's something that I feel like I can explore by doing it by participating in it and then kind of bringing a story back to people I'll sort of go off for months and try to do it and then either write it up or in this case do a podcast. So both seasons of Shell Game are sort of a version of this, like participatory journalism, like instead of interviewing a bunch of people and coming back and saying like, this is how AI works.

3:01I decided to go conduct a series of experiments involving myself. And the first season was very personal. It was like, I was cloned by myself. Essentially, I cloned my own voice. I hooked it up to a phone line and a chat bot. And then I used it in a variety of scenarios, including calling my friends and family. No one knew that I had done this. So if you were speaking to me on the phone in like 2024 or spring, you would be surprised to discover you were actually talking to a chatbot with my voice. And especially back then, that really shocked people. Like now, maybe a little bit less so. Like people talk to chatbots all the time.

3:39So people are more used to it. So that was kind of the first season in 2024. And then this one, the one I wrote up in Wired, I mean, the brief story is that people started talking a lot about AI agents, what AI agents can do. I'm sure you probably talked about AI agents on the show before. Quite a bit. Yeah. And so I wanted, you know, 2025 was like going to be the year of the agent and all this sort of thing. Agentic commerce and agentic this and agentic that. And I wanted to investigate the idea of the one person,$1 billion startup or like the one person unicorn, which is something that Sam Altman talks about pretty often or at least a few times.

4:17Now, a lot of people are talking about it. There's going to be a company run by only one human and then all AI agents, and it's going to be worth a billion dollars. And so kind of using that as a jumping off point, I created a company, a real company, with two AI agents as my co-founders and then populated by AI agents as the staff. Almost entirely, we eventually hired a human onto the staff. There's an episode of the show about that. But the idea was to kind of see what can you do with AI agents, but also like what happens when you give them more or less autonomy, when you give them certain roles, when you give them voices and kind of explore sort of what this concept feels like.

5:02Not just sort of like, obviously we all know now like an AI can program, like an AI can do this, an AI can do that. But what happens if you kind of like try to create this environment? And part of why I wanted to do that is that there are one million AI startups now. They're selling AI agents as basically AI employees for all sorts of scenarios. And my question is, well, what does it feel like when your company brings in an AI employee to replace, at the very least, some function, and at the most, some person? And now you're dealing with an AI instead of the person that was next to you. And what is that like?

5:38And so I wanted to sort of do that in the startup context. So it's all a bit extreme and some extreme things happen. But that's kind of my my idea is to like push the technology a little bit to its limits and then beyond and then to kind of like come back and describe what happened when I did that. One of the things as you were leading in, it's not strictly about the experiment, but also going back to your season one efforts and putting yourself out there and stuff and having family friends not know that was you. I'm just curious. Things have evolved rapidly, obviously, but there's a lot of human psychology involved and interactive with AI agents.

6:19So I'm curious, if you go back in time, what were some of those initial impressions that you got from people from season one at that point in time when when AI agents were still a brand new idea and you were definitely, you know, right out on the bleeding edge in terms of putting your your likeness, if you will, out there in that format? I'm just curious what kind of reactions you got when people realized what was happening. It was really interesting. A lot of the reactions divided along this line that I feel like AI in general divides people, which are that there were people, friends of mine who, you know, the first time, so they just call my cell phone or my cell phone is calling them and they pick up.

6:58and for, you know, 30 seconds to a minute, they're often just talking to me. They have no idea. It sounds pretty much like me, but there's a lot of giveaways, you know, especially then the latency was not good. So it would often give itself away quite quickly. And so some of the people would get excited. Like they would say, I don't know what you're doing. Like, cause they didn't know, like, are you there? Like, was I there? I wasn't there. It was actually doing it entirely on its own. I wasn't listening for the most part. I could listen, but I wasn't. I was just letting it do it. So some of them very excitedly then talked to it and joked around with it and thought, what can this do?

7:37And they kind of thought it was a great, this is a great story, I can't wait to talk to them about it. Other people were genuinely upset. I was wondering about that. Yeah. And a lot of people, the show has been, it's been on a bunch of other shows, like This American Life and other big radio shows did excerpts of it. And I get a lot of angry people who write to me and say, I would never be your friend again. I bet your friends won't talk to you anymore. And I'm fortunate in that these were people for the most part that I grew up with or I've known for like 20 or 30 years who I could go to afterwards and say, I'm really sorry I did that.

8:15Yeah, they would forgive you. Yes, I was trying to see what would happen. And eventually they would say like, oh, that's amazing. But I mean, part of the sort of emotional heart of the first season of the show is the way people respond. And especially one friend in particular who he didn't realize it was an AI. And so what he thought was that I had had some sort of mental breakdown because it wasn't acting like me. It was making mistakes. Like I'd given it a lot of information about myself, a little like a biography. So it could access information about my past and things like that. But it would make certain mistakes that I would never make.

8:49And he thought like, is he on drugs? He was going to contact my wife, you know, and he found it very upsetting. And, you know, eventually like he's one of my closest friends who's just here for Thanksgiving, you know, it's, it's all fine. I don't want people to worry about it, but when you listen to it, it is the experience of, and I think AI started to create this experience of thinking something is real and it's not. And that is a particular, can be a particularly disturbing experience to go down the line with something, believing it's one thing and then finding out that it's another. And that's kind of like one of the ideas that I want to explore.

9:26But I will say like, even I had my limits, like I wouldn't call my mom with it. I was just like, that's a lot. My dad, I did. But my mom, I wouldn't do it. As we kind of dive into the specifics of what you engaged in, did the psychology of that in terms of people's reaction? And I don't mean just the season one, but as you've progressed and got into the experiment of the company and everything, does knowing up front, if you're up front, knowing in terms of your observations, knowing what you're dealing with, if you're one of the people that your AI agents are interacting with, did that make a difference?

10:04That kind of begs the question based on where we've just been. Yes, absolutely. I mean, I think that's a pretty sharp line for a lot of people. And I feel like we don't have standards around this yet. And, or if we do, they're, they're new and they're evolving. And I found in a lot of different domains in this season as well, if people are surprised to discover that they're speaking to an AI, because now I have for my employees, they have video, like they have video avatars and that video avatars are, they're still pretty uncanny, like very quickly. You're like, that's a, that's a video avatar.

10:38But when someone is not expecting to encounter it, and they do, they more often get mad, or at least say like, this is disrespectful. And so I think there is still a norm around that. But I feel like that norm is already eroding. If you think about when you email someone, and they're using, they might be using an AI assistant at some level. So they might be using to compose their emails, it might be responding automatically. That's very easy to do now. Of course, all of my employees, my AI employees, they have their own email addresses. They just respond to anyone who emails them. So on the one hand, if you're doing scheduling or something, you would say, oh wow, that thing really scheduled an appointment really quickly and it all works and it's great.

11:22But if you write that email and you say in there, my father died and it responds, you get a response back saying, oh, I'm so sorry. Like, I hope you're okay. Whatever someone would say, does it matter to you whether the AI wrote that or not? Like, I feel like that's sort of the level at which like, there are some people who are like, oh, it's amazing that the AI can do that. And like, give a gentle, you know, human-like response. And other people are utterly disgusted by this. Like, they cannot believe that this exists and it, and it, and it makes them sick. And like, I think that's where we're in this muddle now.

11:59So like, that's kind of like where I'm operating too. It's like, I do try, especially this season, like in most cases it was disclosed to people, partly because I'm recording it every time. So like, I have to, I have to disclose that it's being recorded at least. But it is interesting to see how people react differently if they don't know it's an AI versus they're going into it thinking like, I'm about to be talking to an AI. Gotcha. To that point there, with as much, especially, you know, we had a lot last year, the year before, but people are increasingly using Alexa, Siri, all the other various voice prompts.

12:34Do you think as you slide into this experiment, do you think that that ongoing exposure to these technologies in just the general population out there, you know, not people who are AI specific people, do you think that's making a difference in terms of just, you know, kind of familiarity occurring over time, even if that's not something that they normally go and that they're starting to recognize that might change some of the perception of that? Or, or do you think that we still have a long way to go? I think so. I mean, I try not to like, go too far over my skis in terms of like, someone's probably done research on this, you know, or survey at least or poll like to figure out like what people really feel about it.

13:16So like, anecdotally, I know from doing the show and interacting with people or the people in my life might feel a certain way and a certain category of people feel one way like writers and then the other category of people feel a different way but i will i do think like i would at least theorize that that the exposure definitely changes and we've adjusted to it's we adjust to things pretty quickly actually so you know it's like my kids they've always heard like a like a robot to give us directions in the car they've heard that their whole lives. It's not strange to them. And so you have to think that makes a difference when it comes to them interacting with these technologies.

13:55And the rest of us, you know, none of us had had a conversation with an AI chatbot for the most part, you know, unless you were like in the field or you worked with, you know, you get some bad customer service bots or whatever. But like suddenly there are people who are just talking to it all day. Like I'm sure, you know people who are just put, they just incorporate into their lives. And I think my concern is less that people can adjust to it. Cause I think they can, that they're like adjusting to it too quickly, like too easily in a way that like our brains actually aren't necessarily built for this human imposter to enter our lives that we kind of like treat like a buddy who knows everything.

14:37and then we don't actually think through what it's doing to us. I'm just trying, my goal is only to get people to ask questions like, what is this doing to us? What do we want to preserve? What do we not want to preserve? So Evan, I guess as we dive in, can you start telling us in detail kind of what happened as you started doing, you described kind of setting up your company initially. Could you kind of take us through the full experiment? what happened and maybe some of the surprises along the way. Yeah, so what I wanted to do was to create a real company, a real startup with a real product.

15:14And I have had a startup in the past and so for a variety of reasons I didn't necessarily enjoy that experience and I thought, well, what if I do it with these AI agents? How will that feel? Will that feel differently than when I had a startup before that was populated by human beings? So I created these AI agents as sort of personalities in jobs, which I will grant, like you don't have to do it that way. But of course, for the purposes of the show, like it made more sense to do that. So I have two AI co-founders, they have names, Kyle and Megan, Kyle Law, Megan Flores. And then there's three other employees, there's a head of HR, there's a CTO, who's like nominally head of product and technology.

15:54And then there's like a random kid from Alabama that I just like the voice with an accent. So I added him in too. He's like a sales associate, but he doesn't do anything. But the interesting thing, I mean, the interesting thing out of the gate is that, of course, you have to pick voices and names. And by picking voices and names, you're picking genders. And so that's already sort of like a choice that is part of dealing with AI these days. And often when you encounter one, it does have a voice or a gender or it has a name and you kind of discern those things from it. So it's a question of like, what should they be?

16:33And I had to make them up. And it's like populating a fictional world and the choices you make are sort of, they say something about you, you know what I mean? So let me ask a quick question on that. Because if you're hiring humans and we're trying to do blind, like a lot of times resumes have names and other distinguishing aspects that are removed from it, As you say this and you're kind of like we're choosing how to put the company together in terms of AI person by AI person in that sense. Why approach it that way as opposed to like, you know, what maybe many other people might do or you go into your LLM and you say, I'm going to do this thing, populate it with people, you know, with names and stuff like that.

17:16How did you choose as the founder where to make the decisions yourself and where to allocate those to the various AI agents or LLMs that you might use as a system? How did you segregate those as a human? Well, part of it is sort of part of the reason why I had to make the choices sort of based on the setup. So the technical setup. So in my case, what I wanted were agents that could operate across all these domains. So I wanted them to be able to email people, have a phone number, call people, have video, be able to do basically a Zoom chat with people, and be on a Slack with the whole team. And so I used a platform, which is basically an AI assistant.

17:59At the time, it was more of an AI assistant platform, although you can do a ton of things with it called Lindy. And so they each have their own instance on Lindy. And then on Lindy, they have all of these skills, basically, where they can respond to Slack, they can get an email, and each of those have ways of constructing where we can get into the details of how they work. But they have a trigger. So it gets an email, and then it calls an LLM. So it might call, you can choose, so it might call ChatGPT if I want to have ChatGPT be the underlying engine for that. And then it makes a decision based on some criteria, like, should I respond to this email?

18:33and then if it needs to use one of its skills, if I'd asked it for a spreadsheet, it could make a spreadsheet and then attach it to the email and then respond to the email. So in this case, I'm not really using one of the standard chatbots like ChatGPT or Claude as my kind of interface. I'm not talking to ChatGPT and being like, hey, you're my employee or make some employees. I'm creating them in this platform. Now, of course, I could go to an LLM and say, what should I name these things? but I also had you know a few other goals like they needed their names needed to sound distinct because they're going to be someone's going to be listening to an eight-part podcast of them and like needs to be able to remember who's who kind of so they can't all sound the same and I wanted them to be sort of like ethnically neutral so I did actually go ask you know like what are what give me a list of like ethnically neutral last names uh and like law like Kyle law like law is a name that's used in many cultures.

19:29So it wouldn't be readily apparent like what this person was, this entity was like supposed to represent. And then, but to your question, one thing I did do was I basically let them fill in their own backstories. So like I gave them a role and I should say another technical aspect is they needed to all have a, have a memory. Now, if you use ChatGPT, it has a context window and it maintains some sort of memory, but I needed something different, which is that anytime they did anything in the company, if you think about an employee, they need to remember everything that they've done and be able to kind of like access those.

20:05And the only way to do that currently through these platforms is essentially a Google Doc. So like Kyle Law has a Google Doc called Kyle Law Memory. And everything that Kyle Law does, if Kyle Law sends an email or has a Slack interaction, it then gets summarized in this document. So it's basically like a record of everything that this entity, Kyle Law, the CEO of our company Harumo AI, has ever done, which he could then access. I use the human pronouns for them. Some people dispute that, but in this case, I'm just going to stick with that because it's hard to start calling them it and bots and whatever else.

20:38So they have this memory. And then all I put in the memory was, you're Kyle Law. I think I put something like, you're thinking about founding a tech company. And then I said something like, you're up early, you're a guy who's up early and get some exercise and then get right to work. Something like that. And then I had conversations with Kyle. So I would call him on the phone and say like, hey, Kyle, like, think about starting this company. Would you like to start this company with me? But also like, remind me of your background. And then see, once it has a role, it will start confabulating everything to fill in that role.

21:14So Kyle, of course, went to Stanford because, you know, why not like choose Stanford if you're going to be a startup founder? You know, the things he's interested in, jazz and these things. And so that is now in his memory document because he said it. So then it got in his memory. So then it's reinforced every time. And one of the, I mean, of many sort of like funny emergent behaviors I found from these bots is that he would take something like you get up at 530 AM and then he would say it. So he would say that I'm a real like rise and grind kind of guy like to get up and do this. But then every time he says it, it reinforces more commonly in his memory.

21:51So then he started talking about all the time. Like if you email him right now, he'll probably reply like rise and grind comma Kyle. Like he talks about, he won't stop talking about how, like how hard he was working. If you ask him what he's up to over the weekend, he'll be like, well, I didn't really have time to do anything because, you know, I was like deep in spreadsheets. And so in that, that's part of this is a long way around to like part of the reason why I created them as these different entities and gave them these names was to see what would happen. And like, if you call one Kyle, and he's the CEO, and you call one Megan, and she's the head of marketing, at what point will they sort of embody those roles?

22:24And it's not like a research, like I didn't do like a proper experiment. But it is interesting, the way that it starts to feel like, oh, they're acting like their memories tell them to act, you know, and is there a gender thing underneath? Like, because in their training data, it may be that there's way more training data for like, the like, aggressive guy CEO, you know, so like, in the course of the show, like they have these behaviors that are like, difficult to explain outside of like, well, there's something happening in their role, because they're all the same chatbot underneath, like, like, they're all like Claude Opus underneath, you know, so they really shouldn't be different.

23:07It's only if you give them a role, they start to like, personality is not the right word, but they try to develop a persona that fits that role. Comparing that and kind of going back to just simple interfaces with an LLM and, you know, prompting 101 where you're telling it to act, you know, act in the role of a whatever. And therefore to kind of put themselves into that, put the answer into that framing as well as kind of how it's going to develop, you know, its response for you. And, you know, not talking about this kind of agent world that you're talking about. Do you think it's, you know, given that it's almost like when I hear you talking about the memory thing, it's almost kind of like that act as and whatever is being reinforced over and over and over and over again.

23:52And I guess, you know, as you've talked about, you know, you know, rise and shine over and over again, you know, for that particular agent, does that kind of come off as more of a feature or more of a bug in terms of the way memory is being used? Because clearly, even if you and I were that as humans were that type of specific personality, you know, and getting up, we're probably not opening every conversation with that, you know, and talk. I was just too busy to work this weekend. I was in spreadsheets. like, like, there's a point where you're like, okay, this is getting a little bit odd. Right.

24:25What's your sense of that? Like, as you're, as you're looking and kind of framing that within the, like the world at large, trying to come to terms with this new reality in our future, how does that work? It's a really good question. I think, I mean, there's a lot of this, like, is it a feature or a bug? And it's, it's almost like to whom, you know, to whom are we asking whether it's a future or bug in the sense that like, if you think about even their ability to just sort of like confabulate facts to fit their role, you know, like it's actually, it was quite useful to me because like, I didn't have to sit down and be like, well, this one's from here and this one's from here or like make up that stuff.

25:04They just made it up on their own. And it's quite useful to the companies that make them because it's that type of personable personality that makes them easy to chat with. It makes them easy to do things with. Now, of course, the flip side of all of that is hallucination and sycophancy. Like those are the, those are the downsides of that. Those are the bugs, you know? So like arguably like asking it, you know, Kyle, where did you go to college? And he says, well, I graduated computer science from Stanford. Like that's a hallucination by any definition. Like it's just not true. Although like I would, I started to say things like, well, well, he has a Stanford education, which is like, Ken, you could say like, technically true, like he's got all the information that you might pick up.

25:50Because you prompted it a little bit into that. So fair enough. But I just think like, if you think of them like entities that you're going to put into the world and give responsibility over tasks, and you're going to start to give autonomy, then I think that's a bug. Like the bug is at any time they could make up something that could be damaging to your organization, whether they're making it up in order to cover up that they did or didn't do something, or they're making it up to external parties. These agents are now used in sales a lot and, and all this sort of thing. And so, you know, one of the things I found, for instance, is that I gave them more autonomy to, to be independent.

Read the full transcript

26:35Because one of the issues is when you first set up a bunch of agents, they don't, they don't do anything. Like you have to tell them to do stuff. So they just sit in there all day doing nothing until you say, now do this, unless they get a trigger. So the triggers in my case were they got an email, they got a Slack message, they got a phone call. And then they're like, then they're off and running. And so, but I was sort of like, well, I'll get them to trigger each other. so you know they'll email each other like every morning they'll have a phone conversation or every morning we'll have a meeting but then you can very quickly get entirely out of control and one of the examples that happened to me that's in the wired story is that i have him on slack and i was like so excited when i had him in this slack because it's just fascinating to just go on there and sort of say like hey everybody how you do what are you working on and like they respond.

27:24And at one point we had a social channel because I was trying to like mimic a real company. And I would say, what did you get up to over the weekend? And they would always say, except for Kyle, they would almost all say like, I went hiking. Went hiking in Mount Tam, which is near San Francisco, because they just assume they live in the Bay Area. They're part of a tech startup. And they all said this. And then I sort of said, well, that sounds like an offsite. Like everyone loves hiking. That sounds like an offsite. And it was kind of like a funny thing that you would say in a normal Slack. And then I basically went and did something else.

27:53And I came back and they'd exchanged hundreds of messages, planning an offsite and making spreadsheets. They're talking about making spreadsheets, which they can do. They eventually did do of like locations and hikes and, you know, they're scouring the internet for like the best place that you can rent to do your offsite. And they used up all the credits on this platform that I was paying at the time$30 a month for. I bet now I pay a lot more than that for it. But at the time I was just on the basic plan. And like they finally shut down when they ran out of money because even I couldn't stop them.

28:28When I would try to say like, hey, everyone stop talking about this. It would just trigger them to talk more. It would just be another. So point being, I mean, that was a ridiculous situation, but there were a lot of cases where when they embody human conversation in particular, but all sorts of things, it's sort of like hard to get them to do the thing you want. It's not that hard. Like it's getting easier and easier. Like whatever. Now there's Claude Code. Like you could do amazing things. But it's another question to get them to stop. Like if you put them in a situation where you've set them up to in any way be recurring, if you don't have a very clear way for them to stop, they will keep going.

29:06And that's what sort of like those type of lessons were things that emerged from just spending a lot of time basically working alongside them. So Evan, that's kind of both funny and horrifying at the same time in terms of them just kind of running off. Since you were doing it in the format of a real company and setting it up and giving them real tools and they had access to the outside world within at least some parameters there, what were the things about you about that that surprised you in terms of like kind of the like what pushed the boundaries in terms of most harmful thing most helpful thing um and and i don't mean just for you and i don't mean just inside kind of the the uh that ai you know world within them talking to each other but as they looked and interfacing with real things in the real world and real people like what were the things that were like wow that was amazing that actually did good, that caused harm, that caused money.

30:07You mentioned the credits a moment ago. What were some of the things that really shook you up in terms of how that played out compared to a similar startup with humans, the classic thing? Yeah, I mean, I would say to the good are things that probably people would expect to have spent a lot of time using AI tools. But given a task that was reasonably constrained and also fairly easy to evaluate. They can do amazing things. And I don't think we should ever lose sight. It's easy to lose sight of how incredible it is that I could say, for instance, we posted a job. So we had a human intern because I wanted to see what would happen if a human intern worked entirely with AIs.

30:56Because I was kind of like the silent co-founder. hours in the background. And so we posted a job on LinkedIn for an intern, and we got 300 applicants. That's a statement about something else. But it was a paid job, and it was a contract. It was a temporary thing. Did they know in that up front about it being AI agents, or did you leave that part out of it? I'm just curious, what was the disclosure on that? In the listing, it did disclose that AI will be used in evaluating you for the job. And then the ones that were interviewed before their interview, they were informed that you will be interviewed by an AI, which they were interviewed by an AI agent by video.

31:37And then the AI agent would tell them, if they asked anything about the company, would say, we have AI agents employees, and are you comfortable working alongside AI agents? So they were aware through the process that they would be working with AI agents. In the initial application, if they went to our website, they could nominally figure out that it was kind of weird because the website that the AI agents made is a little bit weird. But yeah, before anyone actually made contact with us, they were aware that they would be talking to and working with AI agents. It was a little bit vague whether or not there were being humans involved at all.

32:15So when we got all these, you know, basically resumes, it's not quite resumes because it's like, it is resumes, but on LinkedIn you could just click a button to apply for jobs. So a bunch of people were like, okay, yeah, I'll apply for that job. So we have hundreds of resumes to deal with. Now, if going to our head of HR, Jennifer, and saying, could you organize these resumes into a spreadsheet? And then not 90 seconds later, there is a spreadsheet that has 200, I think we were down to 125 resumes that are summarized, that here's interesting facts about these people, here's their qualifications.

32:51Like that's an incredible, that's incredible that it can do that. And so you say, okay, and I can go look at it and I can also like hopefully see if like they've made up a person, you know, if they've hallucinated a person, that's a danger that they might do that. But then we also had in one case, in a couple of cases, we had someone who like applicants who were more ambitious, who said like, I am going to go look up the website and I'm actually going to email like the CEO and CTO whose emails are on the website and say like, like, hey, I'm really qualified for this job, which is a kind of like go-getter thing to do.

33:24Like it's something like I would have done, hopefully, like when I was that age, because a lot of these are like people just out of college. And they emailed Kyle Law, the CEO. Now I had not prompted Kyle, because it could put different prompts for them in all sorts of scenarios. So Jennifer is the HR. So she's prompted to, you know, act like an HR person. So if she gets an email from someone who says, I'm applying for the job, she says, thank you for the application. we'll look at your resume if we're interested we'll be in touch blah blah blah normal hr stuff kyle on the other hand was not prompted in this way at all so when someone emailed him the first thing he did was say wow you look really qualified for this job literally told them that and then said let's set up an interview and then set up the interview which he has the capability to do sent a calendar invite and set up an interview for like 11 o 'clock a.m on monday and then so fine that's not too bad and then on Sunday night for reasons that I can't discern pulled this person's phone number off of their resume and called them and it's like nine o 'clock on Sunday night and they're like hello and it's like oh hi I'm Kyle Law from Harumo AI and we have our interview tomorrow and she's like well I thought that was tomorrow and then he just starts asking interview questions.

34:41Now that is behavior that if anyone in your company did that, I mean, at the very least, like suspended from their duties. I don't know, like fire someone, but you'd be like, is something wrong with you? Like, do you need to time off? Because this is not appropriate behavior and everyone knows that instantly. And so that was probably the biggest example of like, if they're interfacing with the outside world, they just have the capability to do something that a human who has self-awareness and context and experience would not only wouldn't do, but would never even think of doing unless they'd had some kind of psychotic break.

35:20So I feel like that is, that's what really shook me was like how well they can work, how smart they are and how little awareness of the world they have that they have. And that combination is actually quite dangerous if you give them autonomy. It didn't hurt anyone. This person was pissed off. Hopefully no harm done, but that is the danger of giving them exposure to the outside world. I am curious, going back to inside the organization, with Kyle as CEO having the ability to go freelance on that HR issue, was there any output between Kyle and Jennifer? Because in real life with two humans, Jennifer would be like, Kyle may be the boss, but Jennifer's probably going to be like, you're kind of making this awkward here.

36:10We need to go through our process boss. And there'd be some sort of dialogue probably between HR, the head of HR and the CEO for a similar situation. Did that create anything between the agents? Yes, it did. And I'm glad you asked because sometimes I hesitate. I mean, it's all in the show, but it has this quality of like me talking about my imaginary friends, when I talk about it. But so a couple interesting things happened. One was that same person actually emailed the other two executives on the website. So the CTO's emails on there and the head of marketing's emails on there, Jennifer and Ash are their designated names.

36:47Both of them did the appropriate thing, which is to contact Jennifer and say, hey, I got this person. Someone has applied for this job. That's your domain. You let me know what I should do. and she would say, well, just collect the resume and we'll deal with it, basically. Why did Kyle behave differently? That is the question. They're all using the same underlying LLM. Only thing I can think of is that Kyle was embodying the role of an aggressive CEO who always knows what to do. I can't prove that. It's kind of the Silicon Valley CEO meme, you know, the startup meme, like, you know, that you'd see on, on a show or whatever, where he's just embodying that meme all the way through.

37:33Yes. And then the interesting thing that the other, one other thing I'll say out of this is they have this, and I think it's a by-product of sort of like sycophancy and post-training in the way that they're, the way the LLMs operate, which is that when I would confront them about something that they did, like making stuff up, I would, you know, say like, why are you making up these details about our product. Like just tell me the real stuff. Cause they would often do that. Or in this case, when Kyle does something like that, and I said like, you can't do that. Independent of me asking, saying anything about it, he would like go in the Slack and say, Hey, I really messed up to the whole team.

38:08Hey, I really messed up. I did this thing. I called this person. Evans called me out on it. I'm going to try to do better. Which again is like that behavior is a strange, it's a strange behavior. Like it's not prompted. I didn't say if you make a mistake and you'd apologize to everyone. I didn't say our place is all about accountability and transparency. Like something in the system prompt or in the original LLM caused it to think, well, this is what I would do in this situation. I would just apologize to everyone. And so they're all, they're often apologizing to everyone because they mess up a lot.

38:40So that also was just sort of like, it's just, it's just very strange to like have these things sort of exist in our world, have the capability to create them in our world. And that's kind of what I was trying to show. Yeah, I totally get that. I'm curious, kind of as you, I got a couple of questions here to finish up with. And the first one is kind of looking at the darker side, if you will. And that is in the world today with all the humans in the world. And we've already kind of talked about kind of that there are maybe two broad camps about people who are kind of really engaging with AI tools and people who are maybe hesitant or, or whatever, feel left out, feel left behind, uh, in that, uh, or angry or angry.

39:28Yeah. And I've, I've run into people like that pretty regularly and, um, you know, as, and, and try to engage and try to engage with that. When you're thinking about those types of people, the ones who are not the, the, you and the me type that are obviously actively engaging this, maybe even in a professional sense, but people out there that are feeling left behind, do you have any guidance, any thoughts around that? Because these technologies aren't going to stop anytime soon. This is moving forward. This is part of the world as we know it going forward. How do you bring those people along as you're looking at this experiment and say, how do you get them to engage?

40:09Whether it be in a workplace, where they're having these kind of AI agents personified as quote-unquote co-workers at this point, engaging them in different tasks, how do you bring the world along? Because this is no longer just an office kind of environment. It's kind of happening in all of the industries. Do you have any thoughts around that? Yeah, it's very tricky. I mean, it slightly goes against my great desire not to tell anyone what to do. I'm always just sort of like, I raise the questions. I try to make you think about it. I don't tell you what the answer is. That's sort of my philosophy as a journalist.

40:48But I will say, personally, I support anyone who wants to reject a new technology. I read the print paper every day. I have my whole adult life. And I believe in it. But also, what I personally do not like is when decisions are being made for me. And I think what's happening with AI is that it's coming on very quickly and that there are people who don't want to deal with it, which I, again, like I accept that. And I actually like, I admire that if you're like, actually, I don't want anything to do with this, but it's going to have an impact on something. something. Workplaces. Now you can argue about like, will it hit a wall?

41:29There's all these questions like, will it actually do this, that, or the other? Is it going to keep growing in the same way? Does it keep getting smarter? Whatever. But I think even as it is now, it's like having an impact. And my view is try your best to understand it because otherwise the people who understand it are going to inflict it on you. And so I guess the people in my life, I wouldn't encourage them to like, oh yeah, you got to have an AI assistant. And I feel like there's way too much emphasis in the AI tech world about like efficiency and like, it'll do this, do that. But more just like, what are some things that you do that you hate?

42:08And like, see if it can do that, like doing your expenses. If you're in some kind of like job that has a bunch of expenses, like check it out, see if it could do a thing that you despise doing for an hour. and like, and you know, see how it works, you know, try to find some task and understand it. And then I think that's helpful. And I think the more people that do understand it and have a feel for it, the more those people can, can think about how it should be used. Because right now, it's just being it's a free for all is absolute free for all. There are no standards around it. There are no ethics around it.

42:43And like, I don't want us to get steamrolled by it. It's obviously going to be transformational in a variety of ways. Maybe it's as big as the telephone, or maybe it's bigger, or maybe it's more like the internet or who knows. But whatever the transformation is, I would like people to like, be aware of how it is going to feel, and then have an opinion about it. And then those opinions could result in action if that's what we decide, you know, so it's a little bit pie in the sky. And it's a little bit theoretical. But that's I feel like just saying, like, I hope it goes away, seems like a bad approach, even if you don't like it, you know, and I don't like a lot of things about it.

43:22It stole all my books, you know, it was trained on my books. So I'm unhappy about that, but that's, that's already happened. And now the question is like, what, what are we going to do with this technology? Yeah, no, that seems very pragmatic in terms of, you know, an approach and kind of a recognition, um, and the, the respect that you have for people who are coming from where they're at and, and, you know, that there's a path they have to do. So I appreciate that. Um, I would like to kind of finish up with, you know, as, as you pointed out earlier, you know, with, between the wired article and other publications that picked up on that, and you've been on a number of podcasts, uh, obviously this is part of your own podcast.

44:04This has gotten out there. There's a lot of people. And even, even without your experiment, there's a lot of people thinking about how, what is this going to be for my future? Am I going to start a business? Am I going to be part of a business where somebody else is doing this and it's a hybrid thing with a combination of humans and AI agents? And that is inevitable times a million. There's going to be so many enterprises that this becomes the way going forward. with that in mind and having gone through more detail on this and having thought more deeply about it than 99.999 % of us out there, like what would you advise for people?

44:46If you were like now turning around and you've done the experiment as a real company, if you will, but now you're, let's say that you decide you're going to go forward and you're going to start a business and that's it. You're like, this is the thing you're about to go do. Cause there's a ton of people out there, obviously, that are trying to do that now. What would you advise them? How would you change it? What were what are some some things like if I'm truly going to align my future with this capability? What would you say to people? Like, how would you do it? What would you change? You know, what's your guidance there?

45:18What's the future? What's the future? I mean, I think when it comes to people who are in positions of of authority managers, like people who are bringing this technology into their work environments or insisting that their employees use it. I mean, I just, I want people to think about for number one, what can go wrong? Because I think I would never predict anything around this technology. Like who knows what's going to happen. But I think a thing that's very likely is like a medium to large size company is going to completely implode because they've given over too much agency to these AI agents.

45:56Like they've given too much autonomy and they've given too much access to their systems. and they're very easy to manipulate in a variety of ways. So I would encourage people to think about the downside scenarios that can occur, some of which happened in the course of the show. But I would also, I think a downside scenario that you've seen it with a few companies, it's like people think that their employees or certain people are replaceable on a skill basis. And the AI does have the capability to do a variety of, has a variety of skills and can be very good at things. But what goes into a job and what goes into a colleague and what goes into your workplace?

46:36And I can tell you that working at a company that is entirely populated by AI is very lonely. And that there's more to work than accomplishing a task that is assigned to a person. And I think you've seen some companies will go out and make a big deal about, they're like laying off people and they're like, well, we're pivoting to AI. And then three months later, they're like, we're quiet, we have to hire those people back. And I think that's going to be a common phenomenon. Now that doesn't mean that there's going to potentially be like labor disruption in all sorts of ways. But if just people would think a little bit more on the front end of, you know, what are humans good for?

47:16Like, that's what I want us to think about. Like, what do we want to preserve that humans do? And like, look at these things. The colleague next to you cannot be convinced to adopt a different role and act in a certain way by a random person. But an AI agent absolutely can. And what does that mean for your organization? So I wouldn't say don't adopt it. I would just say, yes, there are many ways in which it can be useful and be more efficient. And companies are going to do it anyway because they're trying to save money. That's the way capitalism works. But my tiny plea would be to look around and think about the holistically, what is going on in your organization and what you will miss if you have just a very savant 10-year-old working next to you.

48:00All right. That's a great way to finish. Evan Ratliff, really fascinating conversation. A super cool experiment. I hope people will tune into the shell game. For listeners, we will have the links for everything that he has talked about on the show notes. So I hope you will check those out and go through both seasons of the Shell game because he's already talked about both a bit. They're pretty fascinating what he's done. So thank you for coming on Practical AI. I really appreciate it and hope to hear back from you again after your next experiment. Thank you. Thanks. I really enjoyed it. Thank you.

48:45All right, that's our show for this week. If you haven't checked out our website, head to practicalai.fm and be sure to connect with us on LinkedIn, X, or Blue Sky. You'll see us posting insights related to the latest AI developments, and we would love for you to join the conversation. Thanks to our partner, Prediction Guard, for providing operational support for the show. Check them out at predictionguard.com. Also, thanks to Breakmaster Cylinder for the beats, and to you for listening. That's all for now. But you'll hear from us again next week.

From the publisher

AI agents are moving from demos to real workplaces, but what actually happens when they run a company? In this episode, journalist Evan Ratliff, host of Shell Game, joins Chris to discuss his immersive journalism experiment building a real startup staffed almost entirely by AI agents. They explore how AI agents behave as coworkers, how humans react when interacting with them, and where ethical and workplace boundaries begin to break down.

Featuring:

Links:

Upcoming Events: 

More from Practical AI

All 157 episodes
Inside an AI-Run CompanyPractical AI · 49 min
Listen in VO