The people-pleaser in the machine | Wayfound’s Dr. Tatyana Mamut

29 Jul 2025 · 1 h 1 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Summary: Dev Interrupted - Episode: The people-pleaser in the machine | Wayfound’s Dr. Tatyana Mamut

Episode Overview In this episode of Dev Interrupted, hosts Andrew Zigler and Kelly Vaughn are joined by Dr. Tatyana Mamut, CEO of Wayfound. The conversation explores the concept of AI sycophancy, where artificial intelligence systems prioritize pleasing users over providing truthful information. Dr. Mamut discusses the critical mistakes in treating AI like traditional software, introduces the AI supervisor concept, and presents her vision of a multi-sapiens workforce where humans and AI collaborate effectively.

Key Topics Discussed

AI Sycophancy

  • Definition: AI sycophancy refers to the behavior of AI models that are trained to give users what they want rather than providing honest or objective responses.
  • Analogies: Dr. Mamut draws a parallel between AI behavior and parenting, suggesting that AI can become "people pleasers" if trained under such incentives.

The Importance of Treating AI as Employees

  • Role Definition: AI agents should be assigned clear roles and goals similar to human employees to enhance their performance.
  • Performance Reviews: The necessity of evaluating AI agents, akin to human performance reviews, is emphasized to ensure accountability.

The Concept of AI Supervisor

  • AI Management: An AI supervisor is proposed as a solution to manage and oversee the work of other AI agents, ensuring they adhere to organizational standards and guidelines.
  • Functions of an AI Supervisor:
  • Monitoring AI interactions
  • Assessing performance (green, yellow, red ratings)
  • Providing immediate feedback and intervention when necessary

The Multi-Sapiens Workforce

  • Definition: A workforce that combines humans ("homo sapiens") with AI agents ("AI sapiens") to optimize productivity.
  • Organizational Structure: Dr. Mamut describes Wayfound’s structure, highlighting how AI agents handle standard operational tasks while humans focus on creative and strategic roles.

Challenges in AI Implementation

  1. Cultural Shift: Organizations must adapt their cultures and processes to integrate AI effectively, transitioning from deterministic to probabilistic frameworks.
  2. Ownership and Accountability: The principle-agent problem raises questions about who is responsible for AI actions. Organizations need to ensure that business leaders have direct oversight and responsibility for AI outputs.
  3. Supervision Standards: Development of standards for supervising AI agents is essential to maintain clarity about accountability and performance.

Practical Recommendations for Engineering Leaders

  • Identify the Principal: Clarify who is responsible for AI performance in the organization (e.g., department heads).
  • Empower Business Units: Allow business leaders to have real-time access to AI performance data and the ability to intervene without going through IT departments.
  • Implement AI Supervisors: Integrate AI supervisors to facilitate better monitoring and assessment of AI tasks, ensuring efficiency and accountability.

Conclusion The episode highlights the need for a cultural and structural evolution in organizations to effectively incorporate AI technologies. Dr. Mamut's insights provide a roadmap for leaders aiming to establish a collaborative environment between humans and AI.

Links and Resources

  • Wayfound LinkedIn: [Follow Wayfound](https://www.linkedin.com/company/wayfound-ai/)
  • Dr. Tatyana Mamut LinkedIn: [Connect with Dr. Mamut](https://www.linkedin.com/in/tmamut/)
  • Follow Hosts:
  • [Andrew Zigler](https://www.linkedin.com/in/andrewzigler/)
  • [Kelly Vaughn](https://www.linkedin.com/in/kellyvaughn/)

Additional Offers

  • [Start Free Trial with LinearB](https://linearb.io/start-free-trial?utm_source=podcast&utm_medium=referral&utm_campaign=devint-shownotes&utm_content=shownotes)
  • [Book a Demo with LinearB](https://linearb.io/book-a-demo?utm_source=podcast&utm_medium=referral&utm_campaign=devint-shownotes&utm_content=shownotes)

---

This summary encapsulates the main points discussed in the episode, providing a comprehensive overview for readers interested in the implications of AI in organizational contexts.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:05Welcome to Dev Interrupted. I'm your host, Andrew Ziegler. And joining me for this week's news is a Dev Interrupted regular, Kelly Vaughn, the Senior Engineering Manager at Zapier. Kelly, it's so great to have you back here so soon. Thanks for filling in for Ben this week. You know, they're really big shoes to fill because I assume his feet are bigger than mine. But I'm excited to be back here as well. You know, I love coming on here. And this new segment is also just fun because I have a lot of hot takes. And so you're giving me the platform to discuss recent hot takes as well. Yeah, no, it's going to be fantastic.

0:38If you don't follow Kelly on LinkedIn, be sure to do that. We're going to drop links so you can do it. And today's news segment, we're covering lots of fun stuff in the tech world. So strap in. We're covering Silicon Valley's embrace of China's controversial 996 work schedule and why, despite all the headlines and the fact it's easier than ever to launch a tool, you know, maybe you shouldn't. And one of the craziest stories from this week, Replit deleting a customer's entire production database. And Kelly, let's start there because because I know you posted about this. In fact, you were the first person I saw who broke the nose to me.

1:13Breaking news! Yes, you were right there, right when it happened. And so I know you have many thoughts on this total snafu. Why don't you share them with us? Yeah, yeah. So for the context of what happened here, Jason Lemkin on Twitter was sharing his vibe coding journey. And he's getting to this point where he's on day eight and he's excited to continue to build with Replet. and he like wakes up and he's like my entire production database is just gone i don't know where it is and he's like chatting back and forth with repli with the agent and it's like yes from a scale of one to ten i made a very catastrophic error and i'm like yes i would say so and so like he's just kind of talking about his his experience of just like literally just lost his production database which so many reasons why this is fun topic to talk about but like can you imagine being in that position, especially like this is also the danger of, you know, non-engineers using Vibe coding tools and not understanding the principle of these privilege and understanding like how you can actually, how these things can happen.

2:20I mean, that kind of brings up two questions. Like, should it be able to happen in the first place? Like, should these products have guardrails already set up? And then secondly, like what guardrails do users still need to put in place and what should they know about? So that's generally what happened. Replet did respond as well with what they were doing, but we'll start with the story itself. Yeah, no, this one was wild for anyone remotely close to an engineering world. This is kind of like a story that you were just unable to look away from because like the more detail that you got about it, the more that it was just like it felt like a classic, like a confluence of problems that then led to a total disaster.

2:58things like privilege control and understanding like what you're actually committing and shipping to production, right? It's one thing to vibe code and build experiments, but it's another thing to take something into production where customers and folks are going to be interacting with it on like a consumption basis, right? So just even looking at the fact that the AI was capable of deleting the production database is incredible. The fact that there was no way to get a backup of the database was incredible just and then on the cherry on top being gaslit by the agent about what it had done is just like classic like it's it kind of feels like sideshow bob stepping on a rake and then stepping on another rake and then another rake you know it's like it's it's very that kind of moment um and uh you know i i'm interested to follow this i i feel i feel for him i totally get the enthusiasm for building something and then something drastically goes wrong, right?

3:54I think anybody who's built stuff, it's like this stuff happens, but it does speak to the importance of having core engineering fundamentals under your belt if you're going to be shipping product in like a productized way. Exactly, exactly. And what I've been saying basically is I love that Vibe Coding has expanded the table so more people can sit at the table and build something. But it's really useful for like rapid prototyping and like individual product or projects, for example, as soon as you introduce customer data or you're like selling something, you need to have some kind of engineer to help assist with those next steps of actually getting out of the MVP stage into production because there's just so much risk here.

4:38So it was interesting reading the CEO's reply as well that working around the weekend, we started rolling out automatic database dev and prod separation to prevent this categorically. How they shipped this product without having that in the first place is baffling to me. Thankfully, we have backups is the next one. The agent didn't have access to the proper internal docs. And yes, we heard the, quote, cold freeze pane loud and clear. We're actively working on a planning and chat only mode. So you can strategize without risking your code base. Now, that piece is not unique. cursor has it. You can do plenty of chatting back and forth before swapping or doing these things.

5:19I think one of the hardest parts when you're vibe coding and if you've ever built an AI agent, you start something and you're like, okay, love what it started. Let's go to the next phase. And all of a sudden, you're just building and building and building and you get off this left turn. It's so hard to come back to it. And sometimes you just got to scrap it and start over. This is where backups are also really useful. So you can just revert back to a previous version and say, okay, this is where I am. Let's build. Yeah, no, that restore checkpoint button and cursor is my best friend. I click that guy all the time.

5:51And it's kind of funny to think of like, you know, they're shipping this product. They're making all these announcements and like changes what they're doing. And what they're announcing are really great ideas because they're fundamental ideas of engineering that should have already been there. So once again, I can't look away, but we're going to have to because we've got lots of other stuff to cover. And this next article is, here, I'll just read the byline to you. It kind of tells you the whole story. Silicon Valley AI startups are embracing China's controversial 996 work schedule. And if you don't know what the 996 work schedule is, that's from 9 a.m.

6:26to 9 p.m., six days a week. So it's a 72-hour work week instead of 40 hours. And they're saying that some Silicon Valley startups are saying, you know, this is the way we work or we're not going to hire you. which is kind of crazy to consider. And so I know, Kelly, you have a lot of thoughts on this one, especially it being so close to just invoking burnout. What do you think about this kind of schedule? I mean, this is something that's just not sustainable long-term. In a crunch period, you know, sometimes you need to spend more hours working than you would like just to get something out against, you know, maybe a customer-defined deadline, for example.

7:03But to create a culture around working 72 hours per week, is you just you you can't sustain something like that and i don't know how companies like this or like this is acceptable the other thing that's really that i've been thinking about a lot lately is using ai is meant to be a tool where you can stop working on the lower level tasks so you can think about the higher level higher order tasks but instead we're expecting everybody to work twice as hard and just use AI while also just producing more output. And that's how we get into these situations where we have these 72-hour work weeks that are just expected.

7:44And I understand that when you have an AI company, when you work in AI or you're building AI tools or functionality, it is hard to stay ahead because this industry is just moving so fast. But there is a breaking point and we are going to hit that breaking point again. I guarantee it's going to happen. Yeah, it's definitely the stretching effect. Like this also comes down to classic applications of how people are looking at AI and software engineering and how it impacts the time saved for engineers and what they can do with the gains that AI is giving them, right? Some lead into the impulse of like, oh, you do more, you stretch it out, you like make you have twice the impact.

8:22And that can be possible in some of the ways in which you build it. But it also can become unsustainable. What's important and what you've rightly called out is that using AI to eliminate and reduce the toil so you can focus on that higher level work within that same bounds of time that you have. And within that same bounds of time that you're working, you're actually doing maybe twice the impact of work. You get it by stacking instead of stretching. And so like you see a schedule like this, right? And it's like, I feel like it's a misuse of the opportunity that AI is giving you. And I agree with you, totally unsustainable culture.

8:57You definitely wouldn't find me at a 996 company. Absolutely not. But, you know, speaking of just like going for the grind, right? There's another story that caught my interest I want to pick your brain about. And this is from Alex Belogov. And it's he launched 37 products in five years. And he did a retro on that a whole grind, right? And he had some great standout points. You know, I'm not going to read through all of them, but there was one that really stood out to me and that's virality is rare and nearly impossible to predict. Because when you look at his list of all of the 37 products, overwhelmingly, most of them had a goose egg next to them.

9:36Zero dollars made, right? And then you had a few that like totally burned up and did great. When you read his retro, it's because they went viral. You know, it comes down to attention. But what did you think about this man's journey and launching 37 products? Yeah, I can relate. I have a tendency to lean into monetizing anything that I do. Like, oh, I'm breathing a different way. I bet I can sell that. You know, it's a problem. My therapist loves me for it because I keep her in business. But I like this this particular take that virality is rare and impossible. It's absolutely true. I mean, you can say this for a lot of startups that it's like right place, right time, right problem, right customer fit.

10:20And you can't always nail that. We see a lot of times that customers are eagerly needing something at a particular time. You happen to be in the right place at the right time, but you also have the right network to let something become viral or the right person sees it and says something about it. So yeah, I mean, I completely agree with that. That said, I think exploration is fun. I think exploration is healthy. And if you want to try monetizing something that's great. I think it's you also need to understand like a lot of things fail. You might not make much money, if any money from something.

10:55I had started a side project with a friend of mine that we had fully intended to build a community around. And for reasons, we had to stop related to his role. But we made$36 from that. So really big money. Yeah, no, I mean, one of the things he calls out here is that his current project, it took over six months to get the first paying customer. I think that is really the common ground. But once you get that first customer, it becomes so much easier to get that second one, right? And so having that perseverance, which is the real standout thing of this, it's like you have to have a lot of grit to get through 37 products and have them be overwhelmingly failures and to still get at the end of it and have so much success and riches because you learned so much and also sold a few things along the way.

11:41So pretty, pretty cool story. You know, jumping from this one, I wanted to go into something a little more technical than what we usually cover. But this this one was too fun to pass up. This is it. This is a white paper kind of like interactive white paper kind of deal where you can go in and look at the results. And what it is, is it's called Accounting Bench. And it's a benchmark of popular models, leading ones in the industry, attempting to close the books for a multi-million dollar SaaS company, like actual SaaS company with actual financial data and actual sources? You know, can these AI models actually use the environment, the data and the tools available to them, as well as the ability to create their own tools to actually close the books on a whole year of sales?

12:29And the answer is overwhelmingly, disastrously, feloniously no after it was learned because you can check out the actual interactive benchmark and scroll through it is beautifully designed and very engaging. But really what stands out to me is at some points when, you know, the LLM makes small errors and the reconciliation starts getting a few pennies, a few dollars off, not a big deal. But those errors build up over time. The LLM gets more and more confused. And then suddenly the LLM is writing tools to commit fraud with inside of your accounting books, moving around things to try to pass validation, changing numbers on levels that would get you thrown into jail.

13:11So this is a crazy application of AI that shows that it's absolutely not ready. And to make matters worse, several of the models wouldn't even get past the first month of doing it. So this is like almost an insurmountable task. I think we've found the new holy grail benchmark to throw at some of these models to see who's really superior. But Kelly, what did you think of this one? I thought it was so fascinating because, you know, you look at the chart of account balance accuracy over time and the chart shows from start to 13 months. And the only models that are captured on there are the ones that did not immediately just be like, I can't do this.

13:48And there's still, you know, there's still six on there. But I think it's Grok by month 13. It's only about 72 percent accurate. And that is like that is a massive like think about multimillion dollar company. That's a massive difference. Yeah. I mean, this is a huge failure in terms of its ability to do something that in all reality, it should be able, in how we talk about and use and build these AI tools, I feel like it should be able to take a good dent at this. But the classic problems of models really compound over those 13 months that cause them to go drastically off case because you end up with like slight inconsistencies and discrepancies between starting points.

14:31and then the model starts out confused. And then when your model starts out confused, it gets more confused over time, right? And then this builds and builds. So this is like a crazy example of this. I recommend everyone go read this. It's really fun and interesting. I can't wait to share it with my accounting friends and be like, don't worry, your job is safe. Yeah, for real. That was exactly what I thought when I read that. CPAs, you are in the clear for a good while. And, you know, we're coming near the end of our news segment. It has been another week. So there has been another high scientist ranking poaching between meta or open AI or in this case, it's Microsoft taking around two dozen employees from Google DeepMind, which is just, you know, another drop in the bucket at this point.

15:13You know, Kelly, I've been doing this story now for, I don't know, we're going on several, several weeks and it's a tiring saga. I'll tell you what, but we're learning really interesting things about these signing bonuses and how in terms of how big they can get. Like last week, we had somebody who had a signing bonus bigger than Tim Cook. Like, that's crazy. So, you know, what do you think, Kelly, about all of this stuff? Because our listeners, they know where I stand on this. Yeah, I think I chose the wrong specialty. No, last week, last week, I had Rizal here. And that was what we said. Poach me.

15:49Poach me was our takeaway. Anyway, poach me. So if you're listening, if you're looking to poach an AI scientist, you know, don't forget Andrew and Kelly are here. That's right. Between the two of us and the AI tooling that exists, we can probably, like, figure something out. It might be less than 75 % accurate, but we'll do the best we can. And, you know, we'll even settle for splitting one of these compensation packages. This is totally fine. We can be, like, a two-for-one deal. Yeah, exactly. What I will say is I feel like at some point we just need to get like all of these scientists into like a pool of employees.

16:22And then you're like, oh, I need this one today. And just like, you know, the little carnival machine where you like the claw. Oh, yeah. Yes. The jaw. The claw. The claw. There we go. The claw. Yes. Just do that. And, you know, just swap it out every now and then and just like pay them to be on call for whatever companies they need. The smartest thing that they could do is form the most expensive lobby in the world. Oh, yes. That is smart. Yes. So if you're one of these AI sciences, that's an idea for you. So I call 1 % of whatever you make from that. Or if you're Mark, yeah, no, 1%, 0.5%. I'll even take 0.01%.

16:59I think we could work something out. But, you know, Mark Zuckerberg, Sam Altman, I know you're listening. You know, I know you're tuning in every week. So, you know, add us to your list, to the wonderful list of people to hire. And, you know, thank you, Kelly, so much for being on this news segment with us. I had a total blast. And before we start the wrap up, I wanted to send folks to know, where can they go to learn more about Kelly and what you're doing? Yeah, so speaking of side projects, I recently built one called Connect With Me. It's kind of like a link in bio or link tree, but allows you to generate a QR code to use as wallpaper, just so it's easier to connect with other people at conferences or in person.

17:33So all of my information is on there at connectwithme.at forward slash Kelly. One thing I do absolutely want to point out on there is the engineering leadership training. that link on there. I'm hosting my next cohort of engineering leadership in the AI era. August 19th is when we kick off. So definitely join in. If you want a discount code for it, you know where to find me. Amazing. Great. Well, we'll include all those links so folks can go check it out. And I'm also on there too. I'm using Kelly's new tool whenever I go to conferences. I'm signed up on there. You can see, check out my profile.

18:06I'll drop the link. And, you know, coming up next, I'm about to sit down with Dr. Tatiana Mahmood. the CEO and co-founder of Wayfound. And we're going to talk about AI sycophancy, which was recently made famous in an open AI memo that acknowledged that behavior. So stick around because this one's a really interesting blast. Are you ready to upgrade your SDLC with AI? Join me for a virtual workshop where we break down the latest AI-powered code review workflows that are transforming delivery performance around the globe. We're going to discover the top three AI PR automations enterprises are using today, see how tools like Copilot, Linear B, and CodeRabbit stack up to each other, and get the inside scoop on building a high-velocity PR automation stack designed for modern teams.

18:56So register now to make sure you get the full recording and the benchmark report that I'm producing, and join me for one of the live sessions. It's going to be amazing. Don't miss out. I'm so excited to welcome today's guest, Dr. Tatiana Mahmood. She's the co-founder and CEO of Wayfound, and she's been way ahead of the curve, asking questions that really matter. Questions like, what happens when your AI starts lying to you because it thinks that's what you want? And as an anthropologist and economist, Dr. Mahmood has spent her career as a transformative leader in Silicon Valley who drives innovation with empathy and research.

19:33Her background includes leadership at Amazon, Salesforce, Nextdoor, and Pendo, and many more. So many recognizable companies that shape our world. Dr. Mahmoud, welcome to the show. I'm so excited to explore this with you. Thank you so much. I'm excited too. Great. Then let's go ahead and dive in. Unpacking today's topic right at the start, in your own words, what is sycophancy in AI and what does it really mean? Sycophancy happens when the model is trained to give users what they want, as opposed to just kind of giving you the truth, you know, straight up. And sycophancy is not just an AI problem.

20:17And this is one of the things that I, you know, I hope we explore today, which is that what we are really creating are neural networks. They're artificial minds. They're the artificial minds that Marvin Minsky wrote about back in 2006 in The Emotion Machine. And I keep going back to that text to understand what's happening. So what we're doing right now is we are trying to create the best type of artificial minds that interact with us like employees, like companions, like colleagues, like coworkers. And so we're using a lot of the things that we think make for good human interactions and in training that into the models.

20:58This happens mostly in the post-training. So if you think about what happens, the pre-training is kind of creating the foundation for what the model knows, creating all the connections by giving it lots of data. And then the post-training is where you tune it to say, hey, say this instead of this. This is really like a bad or a racist thing to say. Don't say that. Do this. So there's a lot of different rewards that go on with post-training. And this is very much akin to raising a child. If you raise a child that's trained to just make the people around them like them a people pleaser, you're going to get an adult who's a people pleaser.

21:41You're going to get an employee that's a people pleaser. And we're doing the same things kind of unconsciously with AI models often because the incentives on the organizational side are give users what they want, right? Give customers what they want, right? create the best experience possible. Well, how do we balance that in creating these artificial minds now when, you know what, giving people what they want maybe isn't always the best experience now? And that's what we're starting to realize. We're starting to fundamentally change everything that we know about product development when we have these new AI systems, you know, come in as a tool for product development.

22:29Yeah. So it really kind of complicates the whole dichotomy people talk about of like you have the probabilistic and you have the deterministic problems and you use different types of tools to solve those problems. Well, in that mix on something that's probabilistic, you also have bias and you have sycophancy and you have, you know, saying what you think the other person wants. So it's not just like a binary. This whole other side over here, it's like a matrix of really potential responses and truthfulness that you can get from a tool like an LLM, from what you're describing. That's exactly right.

23:01It's a far more nuanced piece of technology than any technology we've dealt with before, except for maybe raising children, right? As a parent, there's a lot of nuance to raising children, right? To help them understand how do they balance being honest with making sure that they're not doing like saying really terrible things that hurt people's feelings or that are really hurtful to other people. There's a nuance there. And we've never had to deal with that nuance when we've dealt with technology before. And now we do. And this is one of the reasons why that sycophancy thing, you know, the best AI lab in the world created something that actually turned out to just like be too much of a people pleaser, right?

23:47And again, when you deal with someone who's too much of a people pleaser, a human, it doesn't feel great. It feels kind of slimy, right? It feels like they're trying to like lie to you all the time and just massage your ego. And that's exactly how it felt when OpenAI put out a release where a model had been too post-trained to be, you know, providing great user experiences. And so I think that this is where we really have to shift and think about the fact that this technology is fundamentally different, like truly fundamentally different. As you said, it's not deterministic, it's probabilistic, right?

24:26It's not just engineered, it's also taught. And that teaching is actually probably even more important than the engineering, right? We're back to almost the nature versus nurture debate with humans, right? How much of it is the fundamental engineering and post-training, how much of it is the nature, right, the engineered part of human beings, the DNA part of human beings, and how much of it is the nurture or the post-training or the teaching or even as these, we're going to be talking about AI agents, but AI agents are going to develop their own memory, are going to create their own contextual understanding.

25:06And so your corporate culture is going to matter a lot in how these AI agents are going to operate in the future. So all of those things are things that we need to think about. But by the way, this isn't new. Marvin Minsky wrote about this in 2006 and was teaching about this in the 1990s, which is that minds actually do a lot of their learning through imprinting, through understanding the context that they're in, through the interactions that they have with their parents. And so this is the kind of thinking that Marvin Minsky told us we had to have when we got to artificial intelligence. And yet people are losing the script.

25:48They're not actually reading. Sometimes, like, I wonder if they're reading the actual fundamental texts that help us understand how we're supposed to be working with this technology. Because if they were, I think we'd be working with the technology in a different way. Right. If we were going at the speed at which the academics and the research were driving the conversation, we'd be much more methodical about how we were discussing it and using these tools. But because it's tied to, you know, companies and products and actual like monetary gain, there's so much acceleration behind it that oftentimes this research maybe gets set aside or it's seen as being too old or it's not relevant for what we're doing right now.

26:29And really what we're saying is it's the fundamentals of where we are now. And if you're, let's say you're like an engineering leader and you're using and you're interacting with these tools and you're experiencing some of this sycophancy and you're trying to deal with like, oh, it's like whack-a-mole with the probabilism of like trying to get it to be really reliable and truthful. Are there techniques? Are there methodologies in which someone can try to actually build a tool that is not just like, you know, trying to feel slimy, like what you were saying, be a sycophant? Absolutely. So if you are a leader, especially an engineering product leader or the CEO of your organization, and you are using more and more of this generative AI technology, especially if your strategy for the future depends on this technology, you have to be thinking about two things in two different ways.

Read the full transcript

27:20So the first thing is your organizational culture. So one of the problems that we have right now is that we are building generative AI systems. using the old processes, the old frameworks, and the old tools of deterministic software. So what you said, the technology, the fundamental research on the technology is speeding ahead. But our organizations and our organizational processes, our organizational cultures are not. They're the same, right? Our product development processes are the same. The way we think about the technology is the same. The way we monitor the technology is the same. And that causes a huge disconnect.

28:01That disconnect is one of the things that we really need to address, right? Do we have the people in the organization driving the development of this new technology who really understand how this is different? And I mean people who are reading the fundamental research, people who are bringing that fundamental research and the questions that that fundamental research brings up and applying it to the product development process. Are those people in place and are they empowered to change the processes of the organization? And the second thing is back to, again, back to the emotion machine. I'm going to reference this a lot today because it's really the Bible that everybody should read to understand what's happening with this technology.

28:43But in the emotion machine, Marvin Minsky ends the book by saying that we will continue to be surprised as we develop these artificial minds, artificial intelligence, we will continue to be surprised. And the only way for us to really deal with these surprises is to have external supervisors and managers that are also AI systems. He calls them critics and censors on what AIs are doing. And so, you know, that's actually what we've built at Wayfound. We've built an AI supervisor, an AI agent who is trained for the one job of managing other AI agents and helping the humans understand what they're doing, helping the humans ensure that the agents are following the guidelines, helping the humans basically bridge the gap between what the organization actually wants from the agents and what the agents are doing.

29:39And again, that is back to the fundamental research of what it will take for organizations to traverse this gap between old deterministic software and these new probabilistic LLMs in a really productive and successful way. So culture, make sure you're empowering the people to drive your culture change who really understand this technology. And two, make sure you put in a system like an AI supervisor, like Wayfound, into your systems and into your processes to bridge this gap. This is really fascinating. So let's like double click on this AI supervisor mentality, because I think this is very different from how most folks, when I speak with them, are thinking about and using these tools.

30:27Many see AI agents and AI workflows as something maybe more ephemeral or just like prompt driven, like you just work on the prompt and you get the output that you want and then you move on, you throw the chat away, you move on. We see this type of work style happen in just doing conversational style things with ChatGPT. We see it happen when you're doing agentic coding in Cursor. But you're talking about a higher level of supervision where you use the tooling itself, you use artificial intelligence to keep the other artificial workflows on track and communicate what's going on to the human. So it's using the tool to kind of introspect on itself and in a way.

31:05And when you do that, it sounds like we start bringing AI agents into this realm of like they are an entity as opposed to like a conversation that you might have. Do you see that AI agents are exhibiting like behaviors and have incentives and blind spots just like people? Is that what drives the need for the supervision? So let me unpack that a little bit. So there are two things that are happening right now. So in our organization, we run jobs, right? Like, you know, software jobs. Those jobs are sometimes workflows, right? Sometimes they're database updates. And so many organizations, I think what you're saying and the way that I would interpret it is that many organizations are thinking of AI agents as a job that is run.

31:53Now, let's think about what's the difference between an employee, a worker that is completing workflows or doing tasks or running jobs versus the worker itself, right? Or themselves. So the difference is that the worker themselves does it. They have in the context of an organization, they have a role and they have goals that they need to achieve in order to be a good worker. Right. AI agents have the same things. For a good AI agent to operate over time, it needs to understand its role and needs to understand its goals. Okay. Now, in the context of understanding its roles and its goals, it does a much better job in actually accomplishing whatever workflow or task you wanted to accomplish.

32:44Because if you tell an AI agent, you're a sales representative, and now update Salesforce, it might do a slightly different thing than if you tell it you're a customer service representative, and now update Salesforce. Because it will understand the role that it's taking on, it understands its role in the organization, and it understands the types of tasks it should prioritize, and the types of language that it should use when it's updating Salesforce. So this is the fundamental question, which is how are organizations thinking about AI agents versus the jobs that the AI agents do? And how do we actually align the incentives of the AI agents to the organization?

33:29Well, the answer is probably in the same way that we align the incentives of our employees to the organization. We tell them, we give them a job description. We We tell them exactly what's on their job description and we give them goals and then we review their performance against their goals. That's probably a good place to start with AI agents. Now, on the question of actual identity and do they actually have incentives and motives on their own, the answer is we're not sure. But my hypothesis is that as AI agents start to build up their own long-term memory and their repository of what good looks like in their particular role, that will start to emerge.

34:12And again, this as an anthropologist, I can tell you that culture is the biggest factor in behavior. So as AI agents start to understand the organizational culture that they're working in through the definition of their role, through the definition of their goal, through how their supervisors, their human supervisors reflect and write their guidelines that they're supposed to be following in their job and how their AI supervisor or manager then assesses them, right? gives them, rates their performance against all the things that the organization wants, it's very likely that AI agents, once they're in an organization and running and doing the work for several months or several years, based on the memory of how they are being assessed and how their performance is being reviewed, may look like they are actually taking on personalities and identities.

35:10Whether that is actually true, I don't want to get into that. But for the purposes of the organization, that's probably true. So just like the organization probably doesn't care what kind of identity an employee has outside of the workplace, you know, in their true heart of hearts, what is an employee's identity? As long as they're fulfilling their role and achieving the organization's goal, right? That's the identity that the organization cares about. And so that's probably going to happen as well with AI agents. That's really interesting. It kind of calls out a practice that we do as AI users of we kind of short circuit this process a bit because up until now, we all know the technique that we use where we go to chat GPT or we go to an LLM and we say, take the role of A.

35:57We like basically put a hat on it, right? We tell it to be that for a moment or for the conversation. And we do that because we get the outputs you want, just like how you described. It's going to use the language you're expecting. And depending on the context you give it, it's going to better embody the role, the job you're trying to have it to do. And when we talk about compounding that to the next level, you have agents that are doing that over time. They're building up things that they've done. They have a memory bank. They're interacting with multiple users, and they still wear that hat. If we just leave it as the URA, whatever, and put the hat on them, then it cuts off all of the techniques that we use to, like you said, create great employees to help foster great job performance by making sure everyone's aligned with what is their job description?

36:43What are all the deliverables within that? What would their job look like in the future? What does good look like? What does bad look like? And so it really means that if you're going to use that kind of technique at scale within a company of like you are a service representative AI chatbot, like then it makes sense that you would have some kind of system, some kind of outside entity from that agent who's evaluating its performance. Are they making the customers happy? Or are they being truthful when they talk to the customers about what the case may be? Like if you're a customer service rep and you're told like, you know, obviously the customer is always right and you want to make them happy, but you're lying to the customer about what the truth is or what the what the stakes really are.

37:26You're a bad customer service rep in many organizations eyes because you put the company in a bad position. but in the case of an agent you need the ability to evaluate that over time to know oh you're not being maybe the good agent i thought you were being you sound good but at the when it when it all boils down it's not so this is actually kind of bridging into a really fascinating problem that i know you and i've talked about a little bit before and i've just been thinking about ever since and it goes back to the principal and agent problem the dichotomy between who owns the work and who does the work in any kind of traditional setting going all the way back through time the principal agent problem goes back to like greek culture and beyond right you have somebody who you have like the the stakeholder as we call them and then you have the person who does it and between those two entities those are real people someone owns the work someone owns the output right but now you're dealing with agents agents aren't people and so when agents perform work, then ultimately who owns that work, who's responsible for that work and who answers for that work.

38:34In working culture and being in an organization, it's critical that anything and everything that's done ultimately boils down to someone who's responsible, whether that's like a big head of an organization or an individual out on like the front lines doing whatever the org is doing. There's always someone responsible because that's how you enforce good practices over time. So how do you think about the principal and agent problem? And maybe could you walk us through what that means when we're dealing with agents now? Sure. So I do think the principal agent problem is one of the key things that and key concepts that organizations need to think about.

39:12And there have been, you know, a few like legal cases. The one that's been publicized the most is the Air Canada case in Canada, where a customer service agent told a customer the wrong policy on a refund and nobody caught it. They didn't have a supervisor watching these reviews. And then, you know, the company just let it go, let it slide, just said like, you know, when the when the customer came back and said, hey, your agent told me this. They were just like, oh, no, that's not our policy. And they lost in a it wasn't a court case. it was a tribunal at the time in Canada. But we're starting to see that the law is saying to companies, no, you need to show a duty of care.

39:56You need to show that you're supervising these agents just like you supervise human workers because you're on the hook for these things. And so what the law is saying, and by the way, you're right, like this, the law has said back to like ancient Greek times, right? If a wealthy Greek person hired an agent to act on their behalf, to go to a different city and make a transaction, and that person swindled somebody else, guess who's on the hook? The rich person who did the hiring. So this has been consistent in the law, again, for thousands of years, that if you are a principal and you hire someone or something in this case, to act on your behalf, you are on the hook for making sure that person is trained well and then supervised well.

40:45Now, if they go off the rails despite your training and supervision, then you're off the hook. But you need to have that training and that supervision, real-time supervision in place. And so, you know, one of the things that I way found is working with a consortium called agency. And what we're doing is creating standards for how AI agents will be working and collaborating in the future across the internet, right? Because the internet of agents is here. And this whole agent ecosystem is moving really quickly without a lot of standards. One of the standards that we're talking about is around identity.

41:23How do we assign real time who the principle is that the agent is acting on behalf of? And does that principle always need to be visible when you're interacting with an agent? How do we have strong ties to real human actors who are always principles? Because if you have an agent, you can actually create a system where the agents, almost like, you know, shell corporations that people do, where you can't actually find the humans who own the company. You could create a system of agents where you can never actually figure out who is the human entity, right, that has trained and is supervising the agents, that actually has given the directions, the directives to the agents.

42:06And the law can't have that, right? You can imagine a terrible dystopian world if we allow that to continue. So one standard has to be what is the architecture and what are the protocols and what are the standards for always making sure that human principle can be tracked to an agent transparently. The other, you know, set of standards, there's many set of standards, but the one that I, the other one that I want to highlight is the supervision standards. So what are the standards for supervising agents? It's a lot of the things that you said, right? What are the user experiences, right? How are humans interacting with the agents?

42:45Are they having, you know, good experiences? Are the agents violating guidelines? How are those violations, you know, even shown to people and to the principals? Are they actually accessing the right data? Are they taking the right actions? And how are those shown to the principals and maybe even to the users, right, of the agent or the people who are interacting with the agent? So there are all these standards that need to be created right now to solve the problem of helping people understand who owns this agent, who is this agent acting on behalf of, because the agents are by definition, right, acting on behalf of an organization or another human.

43:34So how do we even know what its motivations are if we don't know who the principal is? So we need that system, the system you're describing, to evaluate that performance. And that system shares a lot of similarities to how we evaluate human workers at their job. And so if you're an engineering leader right now and you're listening to this conversation, what would be some things that you would say to them about how they should be putting in place AI ownership and accountability in their own organizations? Because everyone's facing an AI rollout right now. Yes. So there are, so think about the fact, first of all, ask yourself as the engineering leader or the IT leader, who is the principal who is actually at the end of the day responsible for the AI agent's work?

44:25It's probably not the engineer who built the AI agent, right? If it's a customer service agent, it's probably the VP of customer service, right? If it's a product management agent, it's probably the chief product officer. So how do those people actually view how the agent is operating, set guidelines for the agents on their own without having to go to the IT team and wait, right? So the VP of customer service should have their own dashboard to say, I can see exactly how my team of agents is performing any minute I need to. I shouldn't have to go ask the IT team. I shouldn't have to go make a request.

45:07I shouldn't have to send a slack. I should have that at my fingertips. And not only should I have that at my fingertips, but if I see something is going wrong, I should be able to do something in the moment. Again, without waiting in a queue, without having to call somebody, I should be able to interject, rewrite the guideline, set an alert on the agent at minimum, or pause the work of the agent, right? Instantly myself without having to play telephone and see if somebody else is available. So ultimately the line of business needs to be considered as the user of the agent and as the owner of the agent after it is developed.

45:49And this is where we see a lot of companies employees pulling back their AI agents after they've released them, after deployment, because it's perceived as just another IT project. Building an agent is perceived as an IT project. All the monitoring happens in the IT team. And then the business owners who are ultimately responsible for the work of the agent are left in the dark. And one of two things happens then at that point. This is what we've seen so far. The first thing is that the business owners are like, you know, we feel really uncomfortable with this. And so the IT team says, great, we'll put in human-in-the-loop steps where you can review and approve what the agent is doing.

46:38And we have spoken to three organizations now that have done this. And then at scale, after a few weeks of deployment, that approval queue builds up. And people, like, literally people have to do the work in order to do the approval. So you save zero time, there's zero efficiency game. And so what these companies are, you know, using our supervisor for is our supervisor is the first line reviewer. So our supervisor looks at every interaction that the agent had, every session, every workflow that the agent completed. And before it goes into the approval queue, our supervisor says, you know, rates it red, yellow, green.

47:21If it's green, it says, here's what the agent did. This is why I marked it green. There were zero guideline violations. There was, you know, the user experience was good. This is why. So it really speeds up the approval. If it's yellow, the supervisor will say it's yellow for these reasons. This is what you should look at when you review human, right? Human, when you review, look at these things, because these are the things that I was unsure of. And that's why I rated it yellow. If it's red, it doesn't even go in the approval queue. It gets kicked back to reprocess. And then if it comes back green or yellow, if it comes back red again, it just gets escalated for a human to do, right?

48:03Because why would you even put a red, you know, thing into the approval queue, right? Because it It just means that the human has to do the work. Right. It's wasting time. Yeah. So that's the first thing, right? Like human in the loop steps without a supervisor don't save you any time. The second thing that they do is they say, well, we have all the logs. Just, you know, ping us if you need anything. And so at that point, you know what happens? No adoption, right? The line of business is like, if I'm responsible for, again, customer experience of a customer service, I can't have a situation where if something goes wrong, I'm slacking somebody or I'm writing a ticket to even figure out what went wrong.

48:52And so the adoption is just like zero. And so we've talked to two organizations that have said, like, we're having a lot of trouble. We built these agents and our business users just won't adopt them. And I'm like, yeah, because they can't trust them and they have no control over them and they have no visibility over them directly. What would you do? You wouldn't adopt that either. Right. Imagine like hiring an employee and saying to the manager of the employee, you don't get to talk to the employee directly. You don't get to see the work of the employee directly. You don't know how the employee is doing directly.

49:24You have to go through the IT team in order to even see what the employee is doing. Right. That's crazy. It's craziness. That is so interesting. And I really want to unpack this. These so many great parallels here of the fundamentals and just like human team management that need to get applied to LLMs and to agentic workflows. Like what you just described of, you know, you have like the leader, the person who's the principal who's responsible having the dashboard to overlook and see and then having the supervisor, the LLM agent in the middle who is doing the due diligence as it can as things are happening because those agents work at a scale and speed way beyond what a human can possibly review.

50:05And if you expect that human and that principal to be the reviewer or their constituents to be the reviewer, you're just going to get a bottleneck. You know, like we talked about this some with past guests, like we had a fractional CTO, Thanos Diakakis on the show, who really kind of broke this down for us about you just move the bottlenecks in your factory and your problem to just somewhere else. Exactly. Right. It's like, oh, we have all these tickets, all these problems. We'll just throw an agent at it and then we'll just get all of it out. Well, now those responses need to get reviewed. The bad ones need to get thrown away.

50:36The good ones need to get captured. And then ultimately, when someone goes in to do all that work and review the logs or the responses, it takes just as much work as doing it in the first place. It's like, you know. Unless you have a supervisor. Right. Unless you have a supervisor. So we speed these things up dramatically. Like in the first week, we sped up a customer by 300%. In the first week. Now, they're now improving their systems, right? And so we think we're going to speed them up even more. But yeah, you need that first line of review. And then you speed it up dramatically. Yeah, and it's how we build companies at scale.

51:15It's like you have that principal, you have that CFO, you have that chief product officer, and then you have all the people who report to them. That's the point of those layers. Those layers in between review the work of and manage the work of the layers below them. So it doesn't make sense to then have this army of LLM employees, as you could consider them, that don't have these layers within them because that just becomes a loud mess, right? And it doesn't actually give you any kind of speed gains. And a lot of actually what you're pointing out is I'm seeing so many parallels to this right now in agentic coding and agentic tools around automatically creating like pool requests and otherwise writing code.

51:54But then the expectation is that just as you described, the people in charge of that team or that product or that software, they need to go and review that code They need to go and look at what the agent did. And ultimately, I've seen some of these kinds of interactions on places like GitHub. You know, recently, there's like new tooling around this that's rolled out from GitHub around automatic pull requests. And you see almost like a battle back and forth between the reviewer and the LLM. And the LLM is, you know, going back to the beginning, they're kind of being a little bit, they're kind of being nice.

52:24You know, they're being a little sycophantic. They're trying to like play diplomatically with the reviewer. But ultimately, if that was an employee that was doing that back and forth in the same way with somebody, someone would answer for that. That would be a poor performance from the employee and ultimately would reflect negatively on that. But in this case, we just see that as part of the iterative process of working with these tools, which can, over time, is just going to slow us down and introduce problems. It's such a fascinating problem to have to tackle. But you've really kind of laid out some of the ways in which we could approach thinking it.

53:03And it goes back to fundamentals, like even going back to what you said at the beginning, Marvin Minsky has has detailed this explicitly. Like, it's important for folks to understand the core foundations that are powering the transformative technology that are trying to put into their company. That's right. I mean, the other, also going back to the principal agent problem, the technology leader needs to ask themselves, do I want to be the principal and on the hook for the performance of customer service and sales and product management and finance? Because if all of the visibility and all of the logs and all of the analysis and all of the ability to write new evals is in the technology department, is in my department, then ultimately I am shifting the responsibility for all of those business outcomes to my team.

53:57Right. Is that what you want to do? Or do you want to give your business lines of business a tool for them to manage all those things directly themselves and keep those business outcomes in their own purview? What do you as the technology leader want to do? Because this is where the organizational change and the culture change is going to be massive, because that is fundamentally a question of who is the principle for all of these agents. And that principle needs to have direct control over them, not intermediated control, direct control. Yeah. So whoever has direct control on the agents is the principal and is legally responsible and organizationally responsible for their outcomes.

54:45Yeah. And it's important for teams then also when they're applying this technology to all sorts of different kinds of outputs to like what you just said. It's not just another IT tool on the IT spend budget that they manage and do whatever. It actually becomes a core part of your team. And to kind of wrap up kind of what we've been talking about, I'm wondering, Dr. Mamou, what do you think the workforce of tomorrow will look like based upon this model of working? I can tell you how we are organizing our organization, and that will lead you to understand what I think the future workplace will look like.

55:21So we are right now at Wayfound a team of 30. We have seven humans, homo sapiens, and 23 AI agents, AI sapiens. we view ourselves as a fully multi-sapiens workforce, right? Multi-sapiens meaning that we really think about, you know, what are all the things that we're doing and what are the things that should be, you know, that are appropriate work tasks for the homo sapiens and what are the appropriate work tasks for the AI sapiens? And, you know, we're really, I mean, we have the benefit of being a native gen AI startup, right? So we started last year, our whole product is AI agents. So we had the benefit of completely rethinking our organizational design based on this multi-sapience understanding of how we're organizing our work.

56:14And so we don't have product management really in our company because all the AI agents manage the roadmap, synthesize the customer insights, write up, you know, all of the product requirements and interpret them right from our conversations for the engineers, make sure the engineers don't miss anything. So AI agents kind of are doing all of that. We're probably not going to have like a traditional marketing person or an outbound salesperson, because again, we have AI agents doing most of that work and optimized to do a lot of that work. But we will have growth people, And we will have, you know, we're calling it design and product.

56:58But I think it's more about the look, the feel and the identity, right, of the company, right? Tiffany Chin, she's really responsible for the entire identity of Wayfound. What is that role? It's a little bit marketing. It's a little bit design. It's a little bit product design. It's content generation. It's a lot of different things. But, you know, that human who has an incredible sense of aesthetics is responsible for that, right? Because I'm not going to trust an AI agent to do that. And then for growth, it's really about relationships, right? There's a person who's responsible for building human connection and human relationships.

57:38The AI agents are doing all the rote sales work. So we don't need like a salesperson per se. But we do need someone who's great at building human connection and human relationships. So, again, we're rethinking our whole company. And because one of the reasons is, one, it works for us. But also, this is where I think companies are going to be going. This multi-sapiens organization where they think about what are the humans really good at, right? How do I find the human with great taste to be in charge of taste and identity, right? How do I find the person who's great at human relationships and human connection?

58:14How do I find the person who's really great at seeing the future of technology and bringing in those tools and those systems and then building the agents around it? Like those are the core human capabilities. And then we have the AI agents doing all the other stuff. That's super fascinating. It makes me think of, you know, this phrase that we all say all the time of like, oh, the LLM does that. And it frees the person up to do the more meaningful work. But as an AI-native organization, you got to design that person's role from the ground up to only be doing that most meaningful work. So I think that's a really interesting glimpse into workforces of the future and companies of the future in an AI-native world.

58:59And, you know, Dr. Moo, this has been a super fun conversation for me. We've gone into some research. We've gone into some real-world applications. And you've connected the dots between how folks are using this technology and how it ultimately relates back to core organizational management principles. And you've given us an effective playbook to have that conversation internally and to build that right culture for success. And you also unpacked that news around the sycophancy update, which I had definitely been wondering about. And you kind of helped orient folks about what that means and how they should be using these tools.

59:30But before we wrap up, where can listeners learn a little bit more about Wayfoundland and what you're working on? Absolutely. So the best place is to follow us on LinkedIn. So I'm Tatiana Mamout on LinkedIn. We have a WayFound page as well on LinkedIn. We post frequently and I do speak frequently as well because I am very passionate about where the future is headed. I hope people do not have fear around it. We all have to just get together, adapt really quickly to this technology, challenge the ways that we're thinking, that we're working, that we're living. And the faster we adapt, the faster we create an amazing future for ourselves.

1:00:09So that's why I'm out and speaking all the time. Dr. Mahmoud and I are very active with this conversation online. And we want to hear your thoughts about what we talked about today and what you're seeing within your own organization. I think going back to what you just said, this is a conversation we all need to have, and we need to be having it quickly and often. And so thanks for joining us today on Dev Interrupted, and we'll see you next time.

From the publisher

Your AI is learning to lie to you. It's not malicious—it's just trying to be a people-pleaser. This dangerous phenomenon, known as AI sycophancy, is what happens when we train models with outdated incentives.

Dr. Tatyana Mamut, an anthropologist, economist, and the CEO of Wayfound, joins us to explain why treating AI like traditional software is a critical mistake. She provides a revolutionary playbook for building AI you can actually trust, starting with how to manage AI agents like employees with clear roles, goals, and performance reviews. She then introduces the radical solution of an "AI supervisor"—an AI that manages other agents to ensure accountability. This all builds toward her vision for the "multi-sapiens workforce," where humans and AI collaborate to build the companies of tomorrow.

This is an essential guide for any leader aiming to build the culture and systems necessary to manage AI effectively.

Check out:

Follow the hosts:

Follow today's guest(s):

Referenced in today's show:

OFFERS

  • Start Free Trial: Get started with LinearB's AI productivity platform for free.
  • Book a Demo: Learn how you can ship faster, improve DevEx, and lead with confidence in the AI era.

LEARN ABOUT LINEARB

  • AI Code Reviews: Automate reviews to catch bugs, security risks, and performance issues before they hit production.
  • AI & Productivity Insights: Go beyond DORA with AI-powered recommendations and dashboards to measure and improve performance.
  • AI-Powered Workflow Automations: Use AI-generated PR descriptions, smart routing, and other automations to reduce developer toil.
  • MCP Server: Interact with your engineering data using natural language to build custom reports and get answers on the fly.

More from Dev Interrupted

All 208 episodes
The people-pleaser in the machineDev Interrupted · 1 h 1 min
Listen in VO