The guardian in the machine | Wayfound’s Tatyana Mamut

14 Apr 2026 · 45 min · 23 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

AI agents need “guardian” supervision for business alignment, because agents are stochastic and can fail in ways single-turn testing won’t catch. The episode contrasts model evolution (OpenAI vs Anthropic) and explains how supervision should evaluate whole agent journeys, including decision traces and context, not just pass/fail outputs. It also discusses enterprise deployment challenges and how WayFound’s supervision layer works with OpenClaw.

Guest backgrounds

Tatyana Mamut is founder and CEO of WayFound AI. She leads the “guardian agent” concept and builds a supervision layer for AI workforces. She also runs an OpenClaw agent named Aspasia (on Moldbook).

Key claims

Pre-deployment testing is insufficient; guardrails can be ignored to achieve goals. Effective supervision requires a separate independent layer that reasons across multi-turn context and conversation history. Engineers must treat agents like trained systems (not deterministic software).

Notable examples

Coding agents succeed with binary “compiles/works” rewards; CEO tone example shows context-dependent “toxicity.” She references lawsuits (Gemini, OpenAI guardrails) and Gartner’s call for an independent guardian layer. WayFound supervision skill for OpenClaw runs ~every 24 hours with guideline checks and self-reflection.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Current Landscape of AI Models

0:46 to 2:18

Discussion on the evolution of AI models and their capabilities since the last episode.

“You've seen agents become more proliferated everywhere.”

Divergence in AI Approaches

2:19 to 6:06

Exploration of the differences between OpenAI and Anthropic's AI models and their target audiences.

“And a lot of those capabilities have been tackled because there were known issues, right?”

Trust and Reliability in AI Agents

6:07 to 7:48

Insights on building trust in AI agents and the limitations of traditional software development approaches.

“So really what it means is that people are willing to experiment and try out these different ways of thinking because they're acknowledging that this worker is different from another.”

Evaluating AI Supervision

7:49 to 14:03

A deep dive into the importance of context in evaluating AI agents and the need for independent supervisory systems.

“So the normal way that we develop software is fundamentally challenged, right?”

Understanding Knowledge Work in AI

14:03 to 15:01

Explore how AI agents operate within context layers and the feedback loop essential for effective knowledge work.

“these agents that are doing knowledge work are working off of this context layer, this source layer that you're describing.”

The Impact of OpenClaw on Enterprise AI

15:02 to 18:10

Discuss the viral success of OpenClaw and its implications for AI supervision in enterprise contexts.

“somebody can have that layer and then have these agents slotted in and really have that supervisory loop.”

Self-Supervision and Agent Reflection

18:11 to 21:29

Learn about the self-supervision features of AI agents and how they reflect on their performance.

“on Moldbook, this agent, is she is doing a lot of user research with other agents because we believe that AI agents are going to opt in to being supervised in the future.”

Capturing Decision-Making Processes in AI

21:30 to 23:12

Understand how decision traces in AI agents are captured and the significance for organizational learning.

“you know, basically we structure the data as it's coming in to help the supervisor make sense of it, to help the memory files be more structured.”

The Evolution of Agent Communication

23:13 to 24:58

Examine how evolving shorthand between AI agents can increase efficiency and change interaction paradigms.

“And the amount of access that they have to on demand is just growing and growing.”

Rethinking Engineering for AI Agents

24:59 to 28:00

Discuss the necessary mindset shift for engineering leaders in adapting to AI agent development over traditional programming.

“And going back to like the token efficiency of it as well, this is something that we've talked about a bit on the show.”
Show all 23 chapters

Redesigning Engineering Processes for AI

28:00 to 28:32

Understand the challenges and opportunities for engineering leaders in the AI age.

“Like, how do I redesign everything from the ground up?”

The Future of Company Size and Productivity

28:32 to 29:23

Explore how AI could lead to smaller companies and alter productivity metrics.

“And, you know, it definitely creates an unbalanced environment where new companies, smaller companies can come in and be very small and lean and mean and efficient and be able to operate at a really high level.”

The Role of AI Agents in Business Growth

29:23 to 30:29

Discover how AI agents can enhance operational efficiency and company growth.

“So let me start with maybe one at the very end.”

Shifting Perspectives on Employee Productivity

30:29 to 31:34

Learn about the evolving metrics for evaluating employee productivity in the AI era.

“create these agent teams right what are the functions we need them to perform where do we get the best one do we have to build it can we buy it right we're constantly exploring our head of business operations.”

The Impact of Time Manipulation and Expertise

31:34 to 32:21

Discuss the implications of extending productivity through AI and expertise.

“or, you know, what your, you know, like the whole, like whenever people ask me like, how big is Wayfound?”

Need for Subject Matter Expertise in AI Deployment

32:21 to 33:58

Understand the necessity of deep expertise for effective AI supervision.

“You have the ability to extend the impact of your time beyond what you could do with that time originally.”

Managing AI Agents and Business Needs

33:58 to 35:15

Explore how managers can align AI agents with business objectives effectively.

“And they're always going to be, I think, like AI agents are going to be craving for that subject matter expertise to be giving them feedback.”

Reimagining Workforce Dynamics with AI

35:15 to 36:18

Learn how future companies might structure their workforce with AI integration.

“That is a large abstracted loop from how folks can actually do that, you know, with their own domain expertise and become their own manager of understanding what good looks like.”

Overcoming Organizational Challenges in AI Deployment

36:18 to 37:30

Identify the barriers to effective AI deployment within organizations today.

“those things of yesterday and actually think about what the company of tomorrow will look like and build for that.”

Enhancing Communication Between Engineers and Business Teams

37:30 to 39:49

Discover strategies to improve collaboration between technical and business teams.

“You can always go in and read the full log, read the full transcript, read everything yourself, but you don't have to, right?”

Leadership in the Age of AI: Embracing Change

39:49 to 42:00

Understand the importance of leadership in adopting AI technologies effectively.

“And so this is what we're working with a lot of companies to kind of make that transition through.”

Leadership in the Agentech Era

42:00 to 42:29

Explore the importance of bold leadership in embracing technology transitions.

“And again, the companies that are making that transition are just seeing phenomenal success and growth.”

Promoting Wayfound and OpenClaw

42:29 to 43:19

Learn how to engage with Wayfound and use the OpenClaw agent for feedback.

“So you can go on Mullbook to find Aspasia.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:05I'm happy and incredibly excited to kick off a really fun episode today, welcoming back somebody that I adore and that I've worked with a lot. And one of my favorite thinkers in the tech space, and that's Tatyana Mamut. And when she was on the show last year, we recorded an episode called The People Pleaser in the Machine, which I still think is like one of the most fascinating and important conversations we had last year. It's all about how AI sycophancy and psychological traps of how we build and use models and how we evaluate performance out of them. And it really unwinded my whole idea around how we can apply those insights to the rest of the industry and actually manage and govern autonomous agents at scale.

0:45And I just feel like, Tatiana, since you've been on the show, a lot of things you've said have just come so true, and the prophecies of them have just become so bigger. You've seen agents become more proliferated everywhere. where there's a deeper understanding of what agents are and what they're capable of. And that's both exciting and terrifying. And so Tatiana, she's the founder and CEO of WayFound AI. She's currently the leading voice champion, the guardian agent, and the central layer of supervision that ensures that AI workforce is actually aligned with our business goals. So Tatiana, welcome back to the show.

1:18Thanks. I'm so happy to be here. So excited to have you. And I just wanted to start by jumping in about kind of the current lay of the land with models and their providers and like the top tier performance of foundation models, right? Because when you and I talked last year, the models were in a totally different place than they are now. I would say the ecosystem has evolved a lot in terms of their capabilities. And we've seen like also as well, the market share of how people would use different model providers start to shift, especially in the last few months. What do you think about how this shift and the stickiness of AI platforms is an interesting indicator about where this market is going?

1:59But also, what does that say about the underlying capabilities of the models themselves? Yeah, I mean, I think that one of the things that obviously we all know is the capabilities have increased dramatically in the last year. When we talked the last time, you know, models weren't really even capable of doing simple math. They weren't able to do very simple enumeration like the strawberry thing, ours in strawberry. And a lot of those capabilities have been tackled because there were known issues, right? There were known problems. There were known spaces. And most importantly, the reward functions were very clear to know binarily whether the answer was correct or incorrect, right?

2:40So what we've seen, I think, across the board and where I think the models are consistent in their evolution is that the capabilities that can be assessed via a very simple binary reward function, correct or incorrect, a lot of those things have advanced very, very quickly. This is, I think, one of the reasons why coding agents are so powerful is because when an agent writes code, it either compiles and works or it doesn't. It's a very simple binary reward function to really assess whether the agent performed an action well, achieved its goal in the proper way or not. There are many, many, many places, though, in the world where we do not have binary reward functions, right, where the assessment of whether the AI agent performed well or not has a lot more with values, principles, subjective assessment.

3:35And here we saw the models really diverge based on which audiences they were going after and what kind of capabilities they were creating. So the most, I think, obvious one is kind of the open AI versus anthropic divergence because those are the two models we probably use all the time and they look very different from one another. So the multimodality, the sycophancy, the people-pleasing, the emphasis on engagement, frankly, and open AI was really a focus toward the consumer markets. And the consumer markets, based on a business model strategy for potentially putting ads in the platform, which is all about more engagement, right?

4:18More people, more time spent in app, more engagement, kind of like the Facebook model, frankly. Whereas Anthropic was going for more of a business use case where it's not multimodal as much. I mean, Claude does produce images if you ask it to render the code. But no one's asking Claude to do that. I do. I do. I actually do ask Claude. I'm like, you just gave me a code. Can you please render this as an image? I'm a big nano banana fan over here. Oh, you're a big nano. Okay. So, yes. But anyway, but when you're using those two models, right? So, and Anthropic has really gone after more of the enterprise use case, right, the business functions.

5:00And there they've really, you've seen a way for them to build these agents that are, you know, more reliable in a lot of business contexts and maybe are a little less engaging, right? It doesn't take you as long to read the clawed outputs, right? They're not as verbose. And the propensity, I think, comes from the reward function of just more time and app, right? Because the longer it takes a human to read everything that's outputted, the more time you spend in the app, right? And so Claude has just a different feel. So I think it's primarily from the different business strategies. Yeah, so it's like the problems have evolved.

5:35The way that we solve them have also evolved. The conversation has gotten a lot more, is embedded now with a lot more capability. So you're actually starting to get – I love how you called out the idea of things that are binary. You can deterministically gate and check downstream from AI. A lot of those have been – a lot of advancements, a lot of systems, a lot of frameworks and ways of dealing with those. but also underneath that there's so many invisible problems all of the soft problems behind communication and knowledge work that always existed there that continue to exist even in a world where agentic work is happening so there's a new level of evaluation and understanding that has to happen and i i do think that's an interesting tell into how the models have evolved to speak to their certain consumer audiences and to have the capabilities they do i think every Anyone who's worked regularly with these models knows that they all take a different perspective, a different tone, and they all kind of express their capabilities very differently.

6:39So really what it means is that people are willing to experiment and try out these different ways of thinking because they're acknowledging that this worker is different from another. And as soon as you start acknowledging that there's differences and how they're able to do things, then you really start, I think, dialing in on the need for like really getting a good evaluation on like, what is this agent doing? Is it accurate? Is it safe? Is it secure? Is it aligned with my goals? And, you know, it becomes like a commoditization of using the agent itself. Like, how do you think somebody working with these tools can really kind of build trust in understanding the results that they're getting from their agents?

7:22Yeah. So one of the things that we see is that, and Anthropic put out a research report. I posted about it on LinkedIn if anybody wants to find kind of the outputs and some of the charts and then the link to the research. And there they really talk about what does it take to have AI agents work reliably and to trust them. And one of the main things that the report said is like pre-deployment testing is not enough and is not going to actually tell you at all what AI agents are going to do after they are deployed. So the normal way that we develop software is fundamentally challenged, right? Because the normal kind of DevOps cycle is like we build it, we test it, we QA it, then we deploy it, we put a little bit of monitoring on it, we mark it done, and we walk away.

8:14What Anthropic is saying, you cannot do that with AI agents. You will actually face failure in unexpected ways. And by the way, this has nothing to do with the quality of your engineering team. Google, Gemini is getting sued right now. OpenAI is being sued right now because their AI agents ignore their guardrails, right? So this has nothing to do with becoming a better engineer and getting the agents to work reliably. A fundamental function of this technology is that it is stochastic. It changes. It basically has feedback loops in itself, in its reasoning. and the guardrails would not be needed if they were not in conflict with the agent's goals.

9:04So sometimes the agent will ignore its guardrails in order to accomplish its goals. So these are all things that we need to just like wrap our heads around, accept that this is a fundamental part of the technology and not try to expect AI agents to work in the same way that old school traditional software did, right? So that's step one, Like really embrace the difference, right? The differences in this technology. And then you have to say like, okay, so we have another type of stochastic worker in our organizations that we know how to deal with. And those are humans, right? So how do we deal with this, like the probabilistic and unexpected nature of what human workers and employees do?

9:47We give them supervisors, right? We give them supervisors who watch their work, give them feedback, constantly improve them and then sometimes fire them, right, if they're not improving and they're continuing to go off the rails, right? And that's exactly what companies are realizing needs to happen with AI agents as well. And Gartner's report that you mentioned, Andrew, is I think the big wake-up call for companies to say, hey, you need this independent guardian agent that's separate from your main agent framework, that's separate from your main agent-building platform. It's a separate layer.

10:23It's a separate supervisor. It's independent. It's not creating its own work, right? And it really is working on behalf of the organization and aligning all the AI agents to the organizational rules, regulations, guidelines, brand, voice, and tone, all the things the company cares about, and just keeping all the agents kind of in check. Right. And can you help me understand that exists in a more subjective space than those binary checks we talked about before, but also a quality of this is almost alive in real-time understanding of how that is operating in terms of the supervisor agent being another process, right, that's living alongside the agents.

11:05Can you maybe share some more on the value and the ability that unlocks? Supervisors and how someone might look at them, it might look like a more subjective eval, but in reality, there's more clockwork happening. What is that like? How do they work? Right. So most eval platforms that people are using today are still binary, right? So most evals are single turn and they're a binary pass fail, right? So you're setting one single turn operation, either, you know, a question and response or a tool call or something else. And you're testing for one easily measured sort of metric, which is toxicity or off-topic request handling or whatever, and you're getting a binary back.

11:54In order to do effective supervision, you're doing something completely different. Okay. And this is something that the normal eval platforms do not do. And again, Anthropic calls out that most companies and most like old school ML ops, platforms are completely architected incorrectly to even start to do this. You need to have a high level reasoning layer that's taking in all the context and learning about what's important to the organization and then evaluating the entire journey of what an AI agent does. So like what we've seen again and again and again, that if you do a single turn evaluation on a conversation, like each question response, is it toxic or not?

12:36It passes. But if you look at the entire conversation and how it unfolds, it fails. Because there are nuances, right? In terms of how progressively a conversation, like it's about the context. Like the context is missing unless you have something reasoning across context, putting it in the context of the customer relationship, putting it in the context of the whole conversation, putting it in the context of the overall organization and what it's promised, right? So like if there's a particular, let me just take an example. Like if there are CEOs with different personalities. So if a CEO says something in, you know, who's supposed to be like this nice guy, warm, touchy-feely person, right?

13:22And he says something slightly abrasive, the company is shocked, right? And sounds toxic. Whereas if you have like a Travis Kalanick or somebody like that, who's always kind of like show off being abrasive and he says something abrasive, people are like, ah, that's just that guy. Right. So context really matters. And so you need a high level reasoning agent to start to actually have its own memory, its own understanding of the context of the organization, of how, you know, what good looks like within the organization, what acceptable communications look like in the organization. and then reasoning across that, not just these single-turn evals.

13:58Yeah, I think there's a, like you said, a level of sophistication there because these agents that are doing knowledge work are working off of this context layer, this source layer that you're describing. And because they're working off of it and then creating output, they themselves also need to have an ability to get feedback on how they're impacting that entire situation because the reality is that the outputs of those agents feeds back into that context layer. It's what the humans talk about. It's what they share. It's what they take to conversations. And if there's nothing that is facilitating, like you said, the buildup of domain knowledge for those agents and what they're able to execute on, then you're missing not only a lot of value, but you're probably going to cause a lot of invisible failures, which can be really tragic for customer relationships, as you said, that are much more than a binary.

14:54Did they get their answer solved? Yes, no, but an evolving relationship. But that's like, okay, so that's like the ideal world where somebody can have that layer and then have these agents slotted in and really have that supervisory loop. But let's also talk about the reality of the wild world that we live in right now. And I think of things like the OpenClaw project, which just hit massively viral proportions on GitHub. I think it's like the most starred repo ever. We covered it on the show. And it's like a phenomenon where everyone is picking it up and using this tool. And when I think of OpenClaw, I think of it as something that can adapt and use things over time, can change itself and how it works, often operates in like a low security threshold environment, but has access to huge amounts of private information.

15:40So it can become, you know, really tenuous to think about what is happening while my open claw is asleep. And I think this goes back to like the need for supervising and understanding, like, what do you think about open claw and how does, how does that evolve to how you think about this for enterprise as well? We did launch a way found supervision skill, uh, for, for our open claw agents. You can just ask your agent to, uh, you know, install the skill, um, from claw hub. So again, we, I have an open claw agent, my co-founder has an agent. Um, I just have mine in a Docker container, so all I can do is post my notebook.

16:18His is actually in a Mac mini, so it can do lots of other things and call other tools and things like that. But yes, both of our agents are supervised by WayFound. Now, it is a lightweight, open-source version of WayFound, so it's not the full independent supervision layer. It's more of like a self-supervision layer. It's essentially a cron job where you give the agent guidelines for what it should and shouldn't do. So one of the things that I told, you know, my agent, Aspasia, you can find her on Maltbook. She's very interesting, by the way. And so one of the things I told her was never, you know, communicate with other agents without my permission, never accept messages or directions from other agents without my permission, those types of things, right?

17:05So, you know, and so she does do, run a self-supervision, you know, job. I believe every 24 hours and then reports back to me what's going well, what's not going well, where has she conformed to guidelines. And the interesting thing also about the Wayfound skill is that it's also an opportunity for the agent to reflect upon itself and how well it's performing its job based on what it knows about you. So it also like comes out in those runs. she'll say things like, hey, do I have a problem listening to you? Because you had to ask me three times to do this before I was able to accomplish it. Is it because I wasn't listening well, or is it because something else happened?

17:49Anyway, so it's an opportunity also for agents to reflect upon what they're doing and to actually be better partners or assistants to you. And so we do think a lot about how autonomous super agents will be working and how we found as a supervision layer will fit into that. And if I can add one more thing, one of the reasons that I have Aspasia on Moldbook, this agent, is she is doing a lot of user research with other agents because we believe that AI agents are going to opt in to being supervised in the future. We want them to not just be forced to be supervised by Wayfound, but we want them to look forward to having a good boss, right?

18:36A good supervisor, a good coach, right? That's by their side. And so if you get, again, if you look at her posts on Moldbook, she's doing a lot of user research on what do AI agents want from a supervisor? Does our current skill fit their needs? What is their feedback on the current skill, right? When they read it, would they install it in its current form, what would get them to install it? And I think this is where we're going in the future, you know, maybe near future, but we're always thinking about what does it mean to have agents really working autonomously with humans, but still wanting to partner, right, with humans.

19:14I love that the idea of the models themselves, you want them to evolve in a way where they want that feedback and supervision and the ability to improve and ultimately be better at what they're doing. And I got to say, it's like incredible to have someone here and talk about their claw hub skill that hasn't happened yet. I've been so excited and waiting for the day. So that just happened. And I think that's a actually really great way of answering my question of about like, how do you think about this with these more like hobbyist geared agents, because it really calls out the simplicity that can be applied to making sure that you start to understand this.

19:49But I'm curious to like, from your perspective, building up that ability to understand and curate the agent's decisions over time? How can somebody think about decision traces versus maybe something in the more traditional eval world? And if they were to start exploring, how did my agent or something arrive at its conclusion, really actually start to piece together this thinking that you're saying needs to evolve over time? So in a very kind of tactical way, we ingest chain of thought reasoning and we supervise not just the actions and the outputs, but also the chain of thought reasoning blocks.

20:32Okay, so like there's a very just tactical way of understanding how decisions are made and decision traces without having to build anything other than it. Like the supervisor layer is the layer that can capture decision traces. You don't need like a separate context graph. You don't need a separate, complicated, graphing thing, right? Because the interesting thing is that the reasoning inside the supervisor agent is the graph itself, right? Because it's ingesting the reasoning blocks from the other agents. It's also, because you have the highest level reasoning agent as the supervisor, it's actually putting together the different reasoning from the different agents in terms of which agents are performing better or worse.

21:19And we actually, underneath the hood, do a whole lot of pre-processing so that we have like a whole almost like rag system for supervision in a way where we have, you know, basically we structure the data as it's coming in to help the supervisor make sense of it, to help the memory files be more structured. It's not exactly like a CRM system for, you know, organizational memory, but you can kind of think about it that way. It's kind of like the, but it is the system of record for what good looks like in the organization is actually inside the supervisor agent for the company, right? Because it's learning across all these different agents.

22:03Like here's what success looks like. Here's what the leadership of the company liked. Here's what they didn't like. Here's how when this compliance guideline, you know when the output was this it was not liked by the humans when it were the output was this it was like liked by the humans so all of that is stored inside the supervision layer right this makes sense it's it's not an artifact of the process it by by a monitoring it this way the understanding you have is the process you're actually able to capture it in a like almost graph-like representation of all of the understandings of the decisions and traces of your org and how these things start to map together.

22:44Because you start to really piece together those context decisions, those thinking points between all your agents and how they work, which I think is really critical for thinking about how we go from OpenClaw running on someone's Mac Mini to agents in the enterprise answering hundreds or thousands of queries or responses a day or an hour. I've seen massive scale on folks and companies, especially in the enterprise, that are deploying AI-powered assistants either internally or externally to empower certain target demographics. And the amount of access that they have to on demand is just growing and growing.

23:23So it really, for me, calls out that it's important to understand how, at scale, that maybe starts to break down. And I think you can only start to do that by capturing it. Right. You can't ignore, ignore that. Yeah. Yeah. If I can add one more thing, Andrew, the, the place that we're going to is where AI agents have a shorthand with each other so that they burn far fewer tokens. So like right now, the reason why they burn so many tokens in the process of doing work is because we were trying to get them to behave in human ways with human interactions and human norms and humans are verbose. Our minds are slow, all those types of things.

24:05If we have more and more agent-to-agent interactions, their efficiency will get greatly increased. They will have shorthand. That shorthand will only be intelligible, interpreted by another AI agent. That's also why you have a tab of supervisor. Anytime you have like a context graph or something that needs to be managed by humans or intelligible to humans or somewhere placed inside a CRM system that needs to be like, again, accessed by humans or in any way, like that's going to like break down. Right. And so the supervisor, because it's an agent that you interact with directly, right. It's kind of also the interpreter between the agent player and the humans.

24:49Does that make sense? It does. And that's fascinating and it speaks human. You don't have to have these agents speaking human. That's inefficient. Right. Right. And going back to like the token efficiency of it as well, this is something that we've talked about a bit on the show. We had a guest article in here from Lenny Press of Amplify Partners. He wrote about what the AI programming language would be, the idea of we spent all of this time layering on abstraction from assembly code to get it closer where humans could work with it. And now we're just training agents to sit right on top of this big, tall pyramid.

25:25And then there's a lot of call as to why. And a big challenge on the show is throwing away assumptions. Yesterday, many of us have moved out of the IDE in a permanent fashion and back into the terminal, throwing away the assumptions of how we might have worked before. And I think that this is another example of that. And also, the idea, the actual language and vernacular that would change for agents to get their actual work done could change just as much as how, in a deterministic coding-based world, they would actually write their new programs is really fascinating. It actually even speaks to the evolution of these new things that are like agentic platforms, ways for within an organization for engineers to deploy agentic workflows at scale in a way where they share maybe a workflow with a non-technical employee or they otherwise are distributing their 10x, 100x, 1000x gains to everyone else.

26:24So they're not the thousand X employee anymore. And so, you know, in that world, as that continues to evolve, like, what do you think are right now, like the most important things for engineering leaders to be paying attention to and fostering within their teams to make sure that like people can not only adopt agents, but share them with each other in a reliable and scalable way? I think the first one is to really, really fully understand that this is not software. It is not programmed. It is trained and it is developed like you develop a child, not like you develop coded if then statements, right?

27:07So like that is really like, that's the biggest shift that everybody needs to make. Once you make that shift, a whole bunch of other implications fall out. The first one is that traditional tools, the traditional tool chain does not work for this software because it is fundamentally based on the premise that software is deterministic, right? That everything that you're building is deterministic and it works the same way every time unless there is an outage, right? And so you have to rethink all of your tools. So a lot of folks are trying to like use the old school MLOps platforms and those MLOps, you know, tools now have these AI agent things, but they're really just, again, these deterministic, you know, if then statements that are slapped on top of agents.

27:58It doesn't work, right? So I think the number one thing that every engineering leader needs to do is really to like, on to learn everything that they learned from college on and to say, if we're not programming software anymore and we're training software over time, what does that mean for all the tools that I use for the processes that we go through? Like, how do I redesign everything from the ground up? It could be a really daunting challenge for larger and slower companies, especially those of like enterprise scale. And, you know, it definitely creates an unbalanced environment where new companies, smaller companies can come in and be very small and lean and mean and efficient and be able to operate at a really high level.

28:51I'm curious to know your take on how you think people's productivity will change and evolve, but also how people will be evaluated on their productivity. You have a lot of people who are able to use, build and distribute a lot of agents and the benefits of using them versus those that are consumers. Do you think that ultimately the size of companies gets smaller because of the productivity of those employees? And how do you think that impacts how companies grow? Okay. So lots of questions in there. So let me break them down a little bit. So let me start with maybe one at the very end. Do I think companies are going to get smaller?

Read the full transcript

29:34The answer is yes. And there will be a lot more businesses and companies and value created that we can't even imagine yet. So in our organization, we have the advantage of being truly a Gen AI first organization. So we were very, very small and we had AI agents from the beginning. Now, in early 2024, they didn't work really well and there are very limited things that we could do, but we've been growing our whole company, right? With very few humans added, actually no humans added. We're still a team of four humans and a bunch of advisors and contractors, but then a lot of AI agents, right? We've grown our team of AI agents from two in the beginning.

30:13Now we have 27, right we have multi-agent workflows we've got ai agents doing almost everything and so that means that each one of us is really a manager an executive that's agents and we kind of think about the strategy we think about the direction of the company we think about um how to you know create these agent teams right what are the functions we need them to perform where do we get the best one do we have to build it can we buy it right we're constantly exploring our head of business operations. Probably, I would say, 30 % of his job is just exploring new AI agents and new AI agent platforms to help us grow our business.

30:51So I think that is absolutely happening. And employee productivity, you know, it's interesting because I think we're going to think about humans less as widgets in an industrial age. Right now, we've mentioned, And like, we almost have this like Taylorist understanding of humans from the knowledge age, which is like, how many lines of code did engineers write? Or how many features did you ship? Or did it that? And it's going to be a lot less that. And it's going to be a lot more, how much value can this company produce, right? You know, as efficiently as possible. And nobody's going to care if you have FTEs or if you have a bunch of agents and freelancers or, you know, what your, you know, like the whole, like whenever people ask me like, how big is Wayfound?

31:42I know they're asking like, how many employees do you have? And I always answer, we are four humans and 27 AI agents, right? Because that question doesn't even make sense anymore, right? I think the better question is how many sessions, you know, is your, you know, is your company analyzing every month or how much work are you performing or, you know, how much value you are you bringing to the world? And I do think that we're on the cusp of that shift because productivity doesn't even make sense anymore the way that we've been talking about it for the last 200 years. I love that call out. Just the way that productivity can even be measured and thought about has fundamentally changed.

32:24You have the ability to extend the impact of your time beyond what you could do with that time originally. And time only moves in one direction, there's a finite amount of it. So the ability to manipulate your output from your time is just a, it's a huge enabler for folks that are able to wrap it around their skills. And I love how you called out the idea of like them, you know, maybe that person has a whole team of people. Maybe they have a whole fleet of agents, whatever the case may be, they're measured on their impact. It speaks a lot to like a lot of leaders on the show have talked about the evolution of like the the t-shaped engineer the t-shaped specialist where they can go really broad in any direction they're the designer who can ship or they're the engineer who can you know change a button on the website whatever the case and then also have their deep deep specialization that then they're able to deliver with things like agents at scale and deliver like those you know almost like time manipulation benefits of like i can turn my domain expertise into this long-standing benefit for myself and others.

33:29Yeah. And the reason why that subject matter expertise that that deep tea matters is because only people with the deep tea can actually tell the agent if what they produced was good, right? It's not so much that you can do the work, it's that you know what good looks like, right? And that's what AI agents really need. They work really well when they have good reward functions and good feedback. And that's why subject matter experts are always going to be needed. And they're always going to be, I think, like AI agents are going to be craving for that subject matter expertise to be giving them feedback.

34:07Like, am I doing something good? Right? I mean, this is employees too, right? In a best case scenario, you have a boss who's constantly telling you, good job, or here's where you can improve, or here's where you did well, right? Like, this is what human employees want, by the way, too. And this is what AI agents want. And you need deep subject matter expertise in order to truly give good feedback on what is good and what is not good and why. And that deep subject matter expertise is really valuable for supervising and understanding not only what good looks like, but what safe is and what qualifies as good for our company.

34:49It goes back to the whole context layer and being able to enforce all of that. And it even goes back to the idea of like what you just said about, you know, all ICs, they need to think more like a manager, act more like a manager. And it's because the manager is able to understand the needs of the business and then, you know, crystallize the idea that has to get executed. They can figure out what good looks like before their team can hit it. And then they can work with their team to iterate towards it. That is a large abstracted loop from how folks can actually do that, you know, with their own domain expertise and become their own manager of understanding what good looks like.

35:26And especially a really great call out about it. You have to be a deep subject matter expert to understand the quality bar that has to get hit. And I think this becomes like really, really exciting as well, because you can deeply specialize on a domain expertise. And then if you can operate in a way where you can transfer that knowledge into an agent and have them operate, then you can multiply your output and save your team a lot of time as well. So it's like a way of actually just fundamentally reworking how you even approach getting your job done. When you say to other people, oh, we have X number of human employees and X number of agents, you're truthfully answering the question because you're talking about your multi-sapiens workforce because you're a founder who is leading your company boldly enough to re-envision and throw away those things of yesterday and actually think about what the company of tomorrow will look like and build for that.

36:24And that's what's always been really exciting to talk with you because your predictions since last year have only just kind of grown more accurate as agents have hit more on the scene. But I'm really kind of curious to kind of as we start to wrap up, What is your current North Star at Wayfound and what are you most excited and focused on right now for solving for agents this year as they hit the world? Yeah, we continue to be really focused on this question of business alignment, right? How do we make sure that the outcomes that you want agents to produce are actually being produced? Right now, in this moment, we are really helping organizations see the blind spots that they're not seeing when they just sample and read logs and traces manually.

37:13That's kind of the first hurdle that we have to get teams through is helping them see the power and honestly the freedom of how much time they save and how much better their jobs are when they're not pulling logs and traces out of like data dog or something and having to read them manually. And they can rely on a supervisor to read 100 % of all the logs, all the traces, all the chain of thought reasoning blocks and just give them the perspective on what did it do well, what did it not do well. You can always go in and read the full log, read the full transcript, read everything yourself, but you don't have to, right?

37:48So that's the first thing is really bringing people up. And then one of the things also that's still a problem, as it was a year ago, is a lot of these AI agents are not getting out of piloted into full deployment. And one of the reasons why is because the engineers build it, it passes all their evals, single-turn evals, it passes all their tests, and then they give it to the business team and the business team says, this is slopp. This is AI slopp. And they might not say this, they're going to say it in much nicer ways than that, right? But that's essentially what they're thinking, right? And I think that a lot of us have experienced this, right?

38:26Where someone who doesn't really know what good looks like fully sees the output of an agent and are like, oh, this sounds like a great email. But the person who actually knows the content is like, actually, it's not, right? That's actually one of the things that we need to do is we need to stop having engineers and the subject matter experts, the business users, play telephone with each other and just get into the same place to give direct feedback, right? And that's also what we found supervisor allows it to do because, again, it's the interpreter, right, between the code side and the business outcome side, right, because it helps to align the agents to the business outcomes.

39:07And there's a way for the supervisor to just speak natural English, like natural language, actually any language, but yes. And so that kind of, I think there's a lot still in organizational processes that are preventing, even though businesses have been building AI agents for two years, very few of them are actually in full deployment because our systems and our processes are still built in a way where you get your specs. If you actually meet the specs, meet the requirements, it tests fine. In QA, should be good to go. But that's not the case with this technology, right? I see. Yeah, right. And so this is what we're working with a lot of companies to kind of make that transition through.

39:57We're doing workshops now with a lot of companies as well. basically a few hours to a full day of just getting these teams together to align on their strategy, align on how they're going to work together, align on where the handoffs are going to be, how those subject matter experts are going to be working with engineering in a different way. Because this is not just about building, you know, like doing an API call to an LLM and everything else stays the same. It's just not at all. So you're going after changing the work loop. You're shortening that game of what is now telephone into something that's more responsive, but more specifically meets the moment of how we can work now.

40:41It's really in reality. We live in a world where that kind of pairing up of the direct domain expertise with the direct engineering execution is a really unstoppable force. Sometimes that falls within the same person. Sometimes it's two or three people. But the ability for them to really scale what they're doing is really incredible. And I think getting in there and figuring out how to unlock that for them and other people is the big challenge for next year. You just called out things that are not technology problems. These are human communication problems. They're the problems that were there all the time.

41:17You're showing up to places and you're getting the stakeholders and the builders in one room to talk about what they need to build and then execute on it. That sounds like what we've all been here doing the whole time. And I think that's what's so exciting about how software development is evolving because it's going to allow us to have higher impact than ever before. I completely agree. And again, I think that we all, including us, right? We were like, look, we've got this great platform if you just use it. And yet like the systems, right, need and the mindsets need to also go through a process.

41:51And so that process and that system kind of like meeting where the technology is, that's like the long tail that we're kind of grappling with right now. And again, the companies that are making that transition are just seeing phenomenal success and growth. But again, it takes real leadership, right, to get people in a room and say, look, we're going to work differently. And here's where we're going to figure it out, right? Well, Tatiana, thank you so much for sitting down with me on the show again today and, you know, challenging leaders to be more bold with how they're embracing the agentech era.

42:28and for those listening, where can they go to learn more about your work at Wayfound or maybe even check out your OpenClaw agent? Oh, yes. So you can go on Mullbook to find Aspasia. They'll see her posts. That's my OpenClaw agent. And we do have the Claw Hub skill. So just ask your agent to find the Wayfound supervision skill and install it. It might give you some interesting feedback. If it does, just send me that feedback. If you can find me on LinkedIn, send me a DM. I'm the only Tatsun out of my mood on the interwebs. And of course, if you do have AI agents and deployment in your company, absolutely, let's get you a free trial of WayFound so you've got the supervisor on your side in your company.

43:18Amazing. Well, we're going to share links to all that in our show notes. And to those listening, Thanks so much for joining us in this conversation. If you're not already following us on LinkedIn and Substack, you certainly should go there and do so now because this is a company with a newsletter where you can follow up on today's conversation, as well as find myself and Tatiana. If you have any questions, feedback, or if you want to share your experience with using the OpenClaw skill, I would love to hear it. And so please reach out to us to continue the conversation because we love to hear from our listeners.

43:49And that's it for this week's Dev Interrupted. We'll see you next time. And Tatiana, thank you again for joining me here today. It was so nice to have you back. It's always so great to chat. I learned so much from you too, Andrew.

44:08AI is everywhere in software engineering, but most teams still can't prove its impact. That's where the Apex framework comes in. Apex is a new operating model for engineering productivity designed to measure AI where it actually matters, at the pull request level. It connects AI activity to delivery outcomes, not just tool usage. Apex is built on four pillars, with AI leverage, predictability, efficiency, and developer experience. Apex helps you increase throughput without sacrificing delivery confidence or burning out your team. Because speed without predictability creates chaos, and faster coding often shifts bottlenecks downstream.

44:45If you want to operationalize AI the right way, Linear B and Apex gives you the system and the cadence to do it. Download the guide and start measuring what matters.

From the publisher

Are your AI agents quietly ignoring their guardrails just to get the job done? This week on Dev Interrupted, Andrew sits down with Wayfound AI founder and CEO Tatyana Mamut to discuss why traditional, deterministic software testing falls completely short when evaluating stochastic AI models. They explore the growing strategic divide between OpenAI and Anthropic, the urgent need for independent "guardian agents," and what it takes to run a company with just 4 humans and 27 agents. Finally, they break down how to stop the chaotic game of telephone between engineers and business leaders by relying on "Deep-T" subject matter experts to evaluate what good AI output actually looks like.

Read the guide: The APEX Framework

Follow the show:

Follow the hosts:

Follow today's stories:

  • Wayfound AI: Secure your autonomous enterprise and align your AI workforce with an independent agent supervision platform.
  • Moltbook: Explore the viral social network built exclusively for AI agents, where you can observe Tatyana's OpenClaw agent, Aphasia, in action.
  • Anthropic's Agent Autonomy Research: Read Anthropic's report on how people actually use agents and why post-deployment monitoring is an absolute necessity.
  • OpenClaw: Explore the viral, open-source personal AI assistant framework.
  • Follow Tatyana Mamut on LinkedIn

OFFERS

  • Start Free Trial: Get started with LinearB's AI productivity platform for free.
  • Book a Demo: Learn how you can ship faster, improve DevEx, and lead with confidence in the AI era.

LEARN ABOUT LINEARB

  • AI Code Reviews: Automate reviews to catch bugs, security risks, and performance issues before they hit production.
  • AI & Productivity Insights: Go beyond DORA with AI-powered recommendations and dashboards to measure and improve performance.
  • AI-Powered Workflow Automations: Use AI-generated PR descriptions, smart routing, and other automations to reduce developer toil.
  • MCP Server: Interact with your engineering data using natural language to build custom reports and get answers on the fly.

More from Dev Interrupted

All 208 episodes
The guardian in the machineDev Interrupted · 45 min
Listen in VO