In short
Zapier’s internal “Code Red” after GPT-4, and how Zapier thinks about agents, deterministic vs agentic workflows, model routing/efficiency, evals, and shifting from individual to institutional AI.
Guests
Wade Foster, Zapier co-founder (2011) and long-time CEO; led Zapier from Y Combinator startup to major SaaS automation platform. Host: Matt Paige (Talking AI podcast).
Key claims
GPT-4’s capability jump and falling costs forced Zapier to reset roadmaps via a week-long shutdown and hackathon; daily AI usage rose from ~11% to 50% in one week. Differentiation comes from application-layer routing and deterministic/hybrid workflows. Automation Bench shows top models score only ~18.1% on end-to-end business tasks, so hybrid setups matter. Best evals are tough, human-easy but model-hard, and use private data.
Notable examples
Zapier’s “daily recap” workflow; daily brief vs recap; model optimization (Coinbase saving ~50% tokens); hybrid agent + deterministic loop; MCP as agent-friendly connectors.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe Rise of Zapier and AI
0:18 to 1:01
Discussion of Zapier's growth and the impact of AI on its operations.
“I'm your host, Matt Paige, and we're here to demystify AI for you so you can get some value from it.”
The Code Red Moment
1:01 to 2:15
Wade Foster shares the implications of issuing a Code Red in response to AI advancements.
“Zapier is one of those iconic names in the space, and you've run it for 15 years now.”
Hackathon for Innovation
2:15 to 3:30
How a company-wide hackathon transformed AI usage at Zapier.
“The capabilities had just advanced a lot.”
Navigating Mixed Reactions
3:30 to 4:32
Wade discusses employee reactions to the drastic changes and the need for education.
“You literally essentially stalled the company for a week.”
Creating Lasting Habit Formation
4:32 to 5:36
Strategies to ensure that AI usage becomes a habit among employees.
“And so it definitely was like a moment of just, whoa, what's going on?”
Differentiation in AI Automation
5:36 to 6:14
Insights on how Zapier differentiates itself in an AI-driven market.
“So you give people more space to come back in and pick up where they left off, see the new capabilities.”
Zapier’s Approach to AI and Token Management
6:56 to 14:01
Wade discusses token efficiency and the dual role of Zapier in AI usage.
“And in most of these areas, and you just keep the exposure up, keep giving people places to do it, keep sharing.”
The Evolution of AI Models and Token Management
14:01 to 16:41
Discover how AI models impact efficiency and cost in businesses, especially with token usage.
“these GPT-5-6 model, like write all the code because it's like a really good engineer on this stuff.”
Evaluating AI Models: A Deep Dive
16:41 to 18:59
Learn how to create effective evaluation methods for AI performance and model selection.
“But it comes down to, especially when you're doing things in a more structured way, the evals and evaluations.”
Defining Agents in AI: Workflows vs. Agency
18:59 to 22:25
Understand the distinction between deterministic workflows and agent-based systems in AI.
“I think colloquially agents have just become basically like things that automatically do work.”
Show all 20 chapters
Iterating and Optimizing AI Workflows
22:25 to 23:36
Explore techniques for improving AI workflows through iterative testing and evaluation.
“How do you think of iterating on a workflow and optimizing it over time?”
API Evolution: From Traditional to Modern Connectors
23:36 to 25:00
Learn about the transition from APIs to modern connectors in the context of AI applications.
“And then they'll add that to the eval suite and then they'll run more through.”
Characteristics of Successful AI-Driven Workers
25:00 to 27:58
Understand what traits make employees successful in an AI-enhanced work environment.
“The MCP sort of gives the agent the ability to kind of go figure it out on its own.”
Daily Recap and Work Architecture
28:05 to 29:15
Learn how daily reflections can improve work efficiency and prevent mistakes.
“And so how do I try and architect my day and architect my work to avoid some of those scenarios?”
Floor Raisers vs. Ceiling Raisers
29:15 to 31:09
Understand the difference between building AI fluency and scaling it organization-wide.
“And I would bucket it into two categories.”
Challenges of Institutional AI
31:09 to 32:55
Explore the complexities of scaling AI solutions within organizations.
“what is actually happening in companies to be honestly pretty different in a lot of cases.”
Optimizing Company Operations with AI
32:55 to 35:02
Discover ways to leverage AI for operational efficiency and value creation.
“And I would say there's very few companies who have really pushed the envelope on this.”
Agile Product Roadmaps in AI
35:02 to 36:57
Learn how to adapt product strategies in fast-evolving AI environments.
“organization, that does take out a lot of that operational overhead where you can start to run closed loop meetings, where you can start to have the AI acting on your behalf in certain places.”
Following AI Leaders and Innovations
36:57 to 38:50
Gain insights on influential figures and companies shaping the future of AI.
“It's like there's tent poles in your strategy.”
Following AI Leaders and Innovations
38:56 to 39:30
Gain insights on influential figures and companies shaping the future of AI.
“If you're listening to this podcast, chances are you've been thinking about how to actually use AI inside your business.”
Transcript
Automatic transcript. May contain errors.0:00Wade Foster:We looked at those three things and said, holy cow, if this continues with any sort of pace at all, this changes the entire industry. And so we felt like we needed to take a quick breather and say, OK, we got to reset roadmaps. We got to reset how we think about our vision. We got to rethink. We think about automation. Hence the code red.
0:17Matt Paige:Welcome to the Talking AI podcast, where we talk AI with both experts in the field and early adopters. I'm your host, Matt Paige, and we're here to demystify AI for you so you can get some value from it. Let's talk some AI.
0:32Matt Paige:Zapier has more AI agents than it has employees, and across its platform, nearly 600 million tasks have already been automated, with the number climbing every second. Literally, there's a running counter on their website. And Wade Foster co-founded Zapier in 2011, and he has led it ever since, turning a scrappy Y Combinator startup into the$5 billion plumbing of the SaaS era on barely a million dollars raised. Now comes what may be his most exciting chapter yet, deliberately disrupting his own company before AI does it for him. Welcome to the show, Wade.
1:02Wade Foster:Yeah, thanks for having me.
1:03Matt Paige:And I'm excited for this discussion. Zapier is one of those iconic names in the space, and you've run it for 15 years now. But I want to go back to March of 2023 when GPT-4 first dropped. And for the first time in company history, you issued a company-wide Code Red and effectively shut down the company for a week. What was going through your head at that moment in time? I think it was a seminal moment for everybody, but I think few leaders took it to that extreme and really saw what was coming.
1:36Wade Foster:Yeah. Well, the GPT-4 launch for us was pretty eye-opening. We had seen, obviously, ChatGPT launch. We played around with it. Product was fantastic. Really enjoyed using it. But it didn't really create this crazy sense of urgency inside the company, at least not yet. It was, hey, this would be cool. How can we operate better with this? How can we make better products with this? It caught our curiosity. The GPT-4 launch, which was about six months later, had a couple of characteristics. One, six months later, it was pretty quick. Two, the difference between 3.5 and 4 was pretty big. The capabilities had just advanced a lot.
2:20Wade Foster:And three, the cost curve was meaningfully coming down. And so we looked at those three things and said, holy cow, if this continues with any sort of pace at all, this changes the entire industry. And so we felt like we needed to take a quick breather and say, OK, we got to reset roadmaps. We got to reset how we think about our vision. We got to rethink how we think about automation. Hence the Code Red. and a big part of what we did around the Code Red was, you know, we did a lot of stuff, but probably the most impactful was the hackathon where we paused the company for a week and we said, hey, everybody, just go build.
3:03Wade Foster:If you're an engineer, play with the APIs. If you're not an engineer and go mess around with ChatGPT, that was really the main tool at the time and just get a sense of like, what is possible? What is coming with these things? And, you know, we saw our daily usage of AI from our employee base go from about 11 % of folks using it to over 50 % in one week. And so that really was kind of like that jumpstart that we needed to say, okay, something important is happening here.
3:30Matt Paige:How was that received? You literally essentially stalled the company for a week. Like I got to imagine some people were like, yeah, okay. Weren't taking it as seriously. How did people react to that?
3:43Wade Foster:But yeah, you got to, this is 2023, right? The AI frenzy is just starting. It's not in peak fervor as it is now. It was a mixed reaction, honestly. I think there were some folks who had already been using the technology who felt like I was behind. They were like, come on, we should go faster. We should go faster. We should go faster. There was definitely folks in the company that felt like this was unnecessary. It was creating chaos where there didn't need to be any. I was called sensational and things like that. And it felt like the way I think about all of that was just it was a moment where I had to be really like really help educate the company, really had to be clear about what I believed and what I felt was coming to help people understand why this was such a critical moment.
4:31Wade Foster:And it wasn't just sounding the alarm bells unnecessarily. And so it definitely was like a moment of just, whoa, what's going on?
4:39Matt Paige:So you got people excited. You got up to 50 % usage. But I feel like after a lot of those, you have the excitement and then human nature kicks in. You go back to how you did things. How did you actually get that to stick in terms of habit formation, which I think is one of the biggest undervalued things of this entire transformation is we're creatures of habit at the end of the day. We like the way things are.
5:02Wade Foster:Yeah. I think there's a couple moves that help it stick. One, you got to just keep doing show and tell. So in our old hands, we would have show and tell where we'd show off, what are you building with AI? And it's not just engineers. We just have everybody from across the company, myself included, that would just show off things that we're doing. And people learn by seeing. You'd see someone do a cool thing and you go, oh, I should try that out. Or, oh, wow, I have that exact same problem. So just the act of just seeing people and getting exposure to it, you start to feel the art of the possible.
5:33Wade Foster:You periodically step back and run more of those hackathons. So you give people more space to come back in and pick up where they left off, see the new capabilities. That's one of the fun and exciting things is that these models are constantly getting released new models that have new capabilities, new powers. The application layer is figuring out how to do more things with them. The harnesses are getting better. So there's this just pace of improvement that is pretty invigorating because the things you can't do today, you are very likely to be able to do not that far in the future. And so you just got to get people in the habit of keep trying.
6:07Wade Foster:It really isn't. The answer isn't, hey, no, it's not possible. It's not yet. It's not possible yet.
6:13Matt Paige:Quick break in the pod. I keep hearing the same pattern with companies I talk to. Claude's helping employees move faster, but in many companies, the business itself hasn't changed. the value is still trapped in isolated chats and experiments. And that execution gap is why forward deployed engineers have become one of AI's most talked about deployment models. They embed with your team instead of advising from the outside. It's also why the FDE model is now central to every client engagement we lead at Hatchworks AI. As an official Anthropic partner, we embed Anthropic certified FDEs to identify high value business problems, build and deploy the solution, and put governance and security around it, then transfer the capability back to your team.
6:45Matt Paige:If your cloud rollout is still mostly individual usage, check out how Hatchworks AI FDEs work at hatchworks.com slash claw dash FDE. You can also find it in the show notes. Now back to the show.
6:56Wade Foster:And in most of these areas, and you just keep the exposure up, keep giving people places to do it, keep sharing. And along the way, you just kind of, the momentum starts to build. It's like a snowball. Yeah.
7:06Matt Paige:And you mentioned the application layer. There's different layers of the stock stack. You got the frontier models, you have the application layer, obviously chips and things like that. How do you think of the level of disruption modes, differentiation, all of these things, because literally Zapier is a workflow automation tool, right? And that's one of the things that AI is best at and people can go into Quad or Codex and have it do things. How do you think about differentiation in this new era? You're obviously not going out and building a model, but you're leveraging models in a major way throughout your entire platform.
7:42Matt Paige:Sure. Yeah.
7:43Wade Foster:I think a big part of it is you have to think about what are the things that these companies, the frontier, the models, what are the things that they can't do or they want to? That's a really important question. And Sam and Dario and the leadership teams at those companies have been pretty forthcoming about the places that they intend to invest and the places that they sort of want the ecosystem to invest in. And so you'd like to pay attention to that. And then you really just try and understand what are the things that are really valuable to customers that those companies can't fulfill. So for example, most of these folks are pretty nervous about vendor lock-in and they want to have the best capabilities no matter what model company it is.
8:20Wade Foster:And so, you know, a lot of these application layers have model routing baked into them or have exposure to different tools. Zapier does as well, where you can pick and choose from the best. You can allow Zapier to sort of figure out what the right tool for the job is. You know, you're trying to figure out what are the capabilities that is not best served by AI. So right now there's a lot of token maxing going on and it's in the best interests of these labs to have people continue to do token maxing. But it turns out there's a lot of workflows that you benefit from having deterministic workflows that run on code that are backed in a reliable way.
8:50Wade Foster:Those are things that are not possible. Or you have this like hybrid agent workflow setup and you want something like Zapier to sort of manage those processes end to end. And so, you know, a lot of this is just trying to think through like where are the places that customers really value and you're just not going to get it from one of those labs. And how do you build, you know, a great product for those types of use cases that are going to be really important to a segment of the market.
9:12Matt Paige:Yeah, there's several rabbit holes I want to go down here, but you triggered one thing, the deterministic versus probabilistic nature. And Xavier's put out this benchmark called automation bench, which is effectively this leaderboard. I actually like you to talk through what it is and what it's scoring, but I think it hits at that point at which when a model is executing on a workflow, it's not perfect all the time. Even when you're using some of the most, the best frontier models out there, even I see a GPT 5.6, which the day of this recording just launched, what, about an hour ago or it's listed on there.
9:46Matt Paige:So what's the purpose of Automation Bench? How do you think about that? And what's it like leaving this benchmark? So you bet.
9:53Wade Foster:So Automation Bench, it's a benchmark that is used to test these various models on end-to-end workflow execution.
10:01Matt Paige:And we have a whole set of tests that we test against.
10:05Wade Foster:And it's across six business functions. So sales, marketing, operations, support, finance, and HR. These are the types of workflows that you or I, any white-collar worker will do every day working inside these tools. And it's meant to test how effective these models are at completing those tasks and at what price points. And as you can see, if you look at the leaderboard, the top of the leaderboard right now is GPT 5.6 Sol Max. And it's scoring 18.1%. So the best model right now is performing these tasks only 18.1 % of the time. And I think what that screams to me is that, hey, for a lot of things, these workflow related tasks, you probably don't yet want to delegate these to a model to go do them.
10:49Wade Foster:And furthermore, they're all expensive as well. So, you know, if you can, if there are parts of your workflow that you actually can run deterministically, you're going to have higher reliability and lower costs. but there are definitely workflows you want an agent to do it. Working with unstructured tasks, writing code, generating emails. Like there's a whole bunch of things that only a model can do. And so that's where these like hybrid agent workflow setups, I think are so powerful and underrated right now, where you can get the best of both worlds. And that's where Zapier just invests a ton of time, where you can, in natural language, describe what it is you want to go build.
11:21Wade Foster:And we're going to help optimize what gets built on the other side. So it's not just purely a model running amok, token maxing, and you spending a lot of money on these things and getting out the other side, you're getting something really purpose built for the task you want done. And we're trying to be smart about how that gets built so that it isn't just willy nilly spending tokens for you left and right. Yeah.
11:41Matt Paige:And I think that it's an interesting pattern because you're allowing AI agents, the models to leverage technology in a deterministic way. It's that marrying, it's very similar to how a human would go and leverage technology in a sense. Are there other bottlenecks or constraints as to why the top tier model is only scoring 18 %? Are there other culprits at play other than the fact that it's probabilistic versus determinist? Are there avenues to actually increase that capability?
12:12Wade Foster:Yeah. You know, look, I can sort of only speculate because I'm not building these models. But, you know, from what I understand is a lot of how these frontier labs are optimizing is they're optimizing for coding models. And one of the great things about building these models to focus on coding is that coding is easily verifiable. You know, you can have tests that run against code and you can see, did the code work or did it not? Did it perform the task or did it not? And so it's a really great place for AI to work because you can know that there is a factually correct way to do this or not. Now, there is some knowledge work that has that same characteristic.
12:52Wade Foster:You know, maybe there's accounting workflows. You know, you either balance the budget or you did not, right? Those types of things, I suspect the models will continue to get a lot better at. But there's also a lot of white collar work where it's a lot more subjective. Is this a great email or not? Maybe we might have different styles of communication and we might say, well, I really like this type of email. It was short and tight. You might say, well, I like this other email that had a lot of detail in it. So there's these other types of tasks that happen inside of knowledge work that have a little more variability to them.
13:24Wade Foster:And so I think this is where when we talk about like white collar work and AI being able to automate a lot of this stuff, I do think it will be able to automate it. But the quality aspect, as long as there's sort of a human evaluating and judging it, those tasks, I think, are going to be trickier for these AIs to objectively say, hey, it is absolutely better. I think you could have some humans that say, yes, it is better. And I think you could have some that say, I don't think so. I think I'd rather have it this way. Which is also why I think you also want to have multi-models. And you're seeing this now, Bearout, where these models have different capabilities.
13:57Wade Foster:And so you might say, well, I really like having, you know, the Opus model be the one that like talks to me because I really like how it talks, but I really want to have, you know, these GPT-5-6 model, like write all the code because it's like a really good engineer on this stuff. So like, you're starting to see, it's weird to call it a personality, but it is almost like a personality that sort of emerges from these different models.
14:15Matt Paige:It gets into what I think Anthropic just put out something around consciousness and the J space, which is a whole interesting separate topic. And then you have people, what is it when you, they're identifying patterns of AI in terms of writing and whatnot. And they're purposely trying to do things to look not like it's AI, which is just productive in some ways. But on the token maxing front, I think you're in an interesting position in Zapier. You're both using AI internally, obviously, right? So there's a cost component there. But Zapier, literally the product itself, people are consuming tokens using the product.
14:51Matt Paige:So you're getting hit from both sides. So how do you think about model efficiency? We had Coinbase CEO recently came out and optimizing which models they were using, saving 50 % roughly on tokens and tokens usage is still growing. How do you think about that in the nature of your business where it's on both sides of the equation? Yeah.
15:13Wade Foster:Well, I think this is a big opportunity for us to help our customers navigate a lot of this stuff. When you go into, I think many organizations, they don't really know what what is the best model to choose for a job? Like most people don't sit down and think about that kind of stuff. And so they just say, oh, you know, I just want to use the best model for the job. And, you know, they might be doing pretty mundane tasks where it's like, oh, you know, build my daily brief, but you're sticking Fable 5 on it. And it's like, whoa, that's just way overkill. You don't need that type of model for the job.
15:41Wade Foster:And, you know, furthermore, I think these models are starting to get where they are very smart. And the incremental advancements that we're seeing in them are now past what is required for many types of tasks inside of the workplace. And so that's why you can see, you know, somebody like Brian Armstrong at Coinbase say, hey, we're going to do a lot of optimization around our spend here because some of these open source models are plenty capable for a wide variety of tasks inside of Coinbase. And so we no longer need to be on the frontier for certain types of tasks. And that is where I think we're going to start to see a lot more optimization into the future.
16:18Wade Foster:And so it's going to be really interesting to see how this plays out. My guess is that we're going to live in a world where most of the tokens are consumed by open source models, but most of the spend still goes to the frontier.
Read the full transcript
16:31Matt Paige:Yeah, that's a great point. It is funny. There is this element of a model FOMO where it's, yeah, maybe the lower tier model would be okay, but maybe I may get that extra bit of something from use. I'm guilty of that all the time. But it comes down to, especially when you're doing things in a more structured way, the evals and evaluations. So how do you think about that? And is there a method for actually creating evals to determine, okay, this model is perfectly fine for XYZ task? Yeah.
17:04Wade Foster:What I've learned about what makes a good eval is a couple of things. I think, one, you really do want to have your eval be tough. it should be really hard for these models to achieve these tasks. And the best evals are typically evals that are actually kind of easy for humans to do, but really hard for these models to do. Because now you're testing something really interesting where you're like, huh, there is a true capability gap there. So I think that's one that's really important. The second thing I think that's really important is it's really valuable to have the test data entirely private.
17:39Wade Foster:this really this prevents basically the models ending up like benchmark maxing on these by accident i think one of the things we've often seen is that as these benchmarks get leaked all of that stuff gets leaked into the training data and then all of a sudden the models like hyper focus on fixing it and so the model gets good at the benchmark but it doesn't necessarily get better at just the things you might actually be exposed to in real life because ultimately a benchmark is still just a proxy for the type of work you do in real life. It's not actually real life. And I think this is what can happen a lot of the times when folks get frustrated with AI in the real world.
18:17Wade Foster:You'll see these benchmarks where it's 99 % or whatever. It's, ah, this is a crazy model. And then you go use it yourself on a task and you're like, dang, it kind of sucked at that. How is it so good at this? But it kind of stinks at this other thing. And it's because the benchmark is still just a small representation of all like the knowledge that sort of humans might ask the model to go achieve.
18:39Matt Paige:Yeah, it totally makes a lot of sense. I want to get into the topic of agents, right? You obviously are leveraging agents within the business within Zapier, but the term gets thrown around like crazy. How do you define what an agent is for our audience? Because I'm sure there's a lot of folks that have differing opinions or have no clue how to actually define it. Sure.
19:03Wade Foster:I think colloquially agents have just become basically like things that automatically do work. But I do think it's important for folks to understand the differences between a deterministic workflow and a purely agentic system. A deterministic workflow is something that's been around for ages, pre-AI. It's a program. It's Zapier in some ways is build agents before agents exist, where it's like, hey, a trigger happens over here and we're going to automate adding this customer to a CRM and then we're going to alert somebody inside of Slack and all that sort of stuff. These are just the workflows that sort of exist inside of a company.
19:39Wade Foster:And these things have only grown in popularity in the age of AI. Now, an agent is slightly different. The way an agent actually performs is you give it a goal. You say, hey, I want you to go complete this task. And then the agent gets to go decide how it wants to go execute on that thing. So based on its training data, based on how the model works, it will say, maybe I'll do this task and then I'll do this next. And then with this combination of data, I can go complete my goal. And every time it gets a similar task, it may not execute it the same way every time. It might choose to follow a different path and it still might get to the same outcome.
20:12Wade Foster:It might get to a slightly different outcome, but it behaves more like a human does where you give it a task and it just goes about its day. It might mostly do it the same way, but it doesn't necessarily follow strict reels. Now there's pros and cons to both of these models. A deterministic workflow, what you get is reliability, consistency, and cost advantages. It does the exact task the same way every single time. So that's the real benefit that you get from these deterministic workflows. The failure case is it's rigid. What happens if something goes wrong? What happens if something unexpected comes up?
20:45Wade Foster:What if you have to work with like messy sets of data, like big unstructured tasks or images or things like that? All of a sudden it gets really, it gets a lot harder to do the task inside of these deterministic workflows. So you run into those problems. Agents are almost the exact opposite because they have agency to go complete the task. Reliability isn't quite so good. Cost is a lot higher. As a result, it can be a little bit slower to complete these things. However, they can be a lot better at handling edge cases because an edge case pops up and it starts to reason through it. What should I do about this?
21:18Wade Foster:How could I go solve this one? And it might find its way through the problem. It can handle a lot more randomness inside the situation. And so that is like the blurring of these two systems. Now, oftentimes, the magic that I find is inside of a business setting, you often don't have just purely deterministic setups and purely agentic ones. So the magic is like when you can actually blend these worlds together where you might say the first couple steps of these always just work the same way every single time. And so we're going to just have that be purely deterministic. Right here, we have one of these ambiguous situations that right now a human is making all these judgment calls on because we can't really do anything with it.
21:55Wade Foster:Okay, let's put an agent right in the middle of that. But then once it spits out an output, we're going to feed it back into a deterministic system to complete the loop. And I think the most sophisticated customers we see, the folks on the frontier, are often getting really smart about how they do these things. They're doing the stuff like Brian and Coinbase is doing, where they're mixing models or thinking about what is a workflow versus what is an agent. They're really starting to optimize these setups. Most people aren't there yet, but I do think that is the world we will find ourselves in as more and more tokens get spent in these organizations.
22:24Yeah.
22:24Matt Paige:And that's the beautiful thing you mentioned is the marrying of both, because to your point, when it was purely deterministic, you had to figure out every edge case or you had failure patterns all over the place, but now with probabilistic, it's a whole new world. How do you think of iterating on a workflow and optimizing it over time? Because I feel like there is that element of being able to improve something because back to the deterministic nature is it did the thing or it didn't. But now that you have AI in there, there's this using the same term again, probabilistic nature of the workflow where it can improve over time.
23:00Matt Paige:Any thoughts on like people that are actually doing these things in their business? How do you like methodically think about starting small and improving a workflow that you're trying to automate over time? Yeah.
23:14Wade Foster:Yeah. I think what the best folks are doing is they have these, they're building their own mini evals for these setups. And as more and more scenarios come through, anytime the agent fails to complete it, they're analyzing why did it seem to go wrong? and based on how it went wrong, let me see if I can give it another test case or another scenario that it can compare against. And then they'll add that to the eval suite and then they'll run more through. And then each time that happens, they'll just keep adding more examples to it. Now look, not always, more examples doesn't always make it better.
23:46Wade Foster:So sometimes they are having to subtract from it, but by and large, like the best way I can describe it is these little test scenarios act as like a guide to the model where you're trying to steer it in the direction that you want more over time. And the more you do this, you can usually eke out higher and higher percentage of accuracy with these models.
24:05Matt Paige:And so, Zapier, you were effectively built on the API explosion as the basis of being able to provide this product. And MCP is the new thing. Connectors, being able to connect all your tools together. What similarities do you see between those two? And APIs are still a massively important thing in this new age, too. So how do you think of MCP versus AI and this ability to connect things together?
24:29Wade Foster:Well, they're really just the extension of the same thing. At the end of the day, it really is just about connecting to the tools that you use. And so APIs, and for most of our history, have just been running a call in a very specific way to sort of extract data from a tool or perform an action in another tool, so on and so forth. MCP just allows that for an agent to do it. And the nice thing about what MCP gives agents is the ability to kind of search and discover and kind of work in a sort of fuzzier way. And so you don't have to be as like specific about like the way to perform the task. The MCP sort of gives the agent the ability to kind of go figure it out on its own.
25:04Wade Foster:But under the hood, it's still effectively doing the same thing.
25:07Matt Paige:So in terms of the nature of jobs and work, there's a lot of debate on is AI taking jobs, is it creating new jobs? What is your take there? And I guess what are you seeing in terms of your own people? How would you describe the ideal type of worker right now? what type of qualities do they have? So you're hiring somebody new at Xavier. What are you looking for? Yeah.
25:32Wade Foster:The most general thing I think I can give is that the people who are very curious, who are doers, who have a high level of caring about their work, are just absolutely thriving right now. AI is giving them a jetpack to go figure out all sorts of things. And so whether you're just entering the workforce or you're on the cusp of retirement, these characteristics seem to accelerate people who really orient themselves that way.
26:01Matt Paige:Yeah, totally. And so let's shift to this. I'm curious, how are you using AI in your day-to-day life? What is the most unique, interesting use case of AI that may be not a common one that everybody's playing with today? Anything weird off the wall?
26:19Wade Foster:Weird off the wall stuff. I do a lot of -
26:22Matt Paige:We're just super productive. It's something that just - Yeah.
26:24Wade Foster:I mean, my favorite thing is, I mean, one of my favorite workflows is a little bit mundane. So a lot of folks like to talk about their daily brief. Candidly, I think the daily brief is like just okay. It's like fine for most people. But the flip of that is the daily recap. I think the daily recap is one a lot of people are sleeping on. So what does the daily recap do? It basically reviews everything that happened for you in that day. So you can have it go loop over all of the digital exhaust that you have from the day. So this might be emails, Slack, meeting recordings, everything, your calendar, and it can help summarize everything that's happened.
27:00It can help track down the key action items.
27:04Wade Foster:Even better than that, it can actually start to take action items. So if you were in a meeting and you said, hey, I'll make sure to make an introduction to this person or I'll make sure I'll follow up with this. Great. It'll already have those emails drafted for you. It already have like sample comms ready to go for you. And so you can do a lot of this stuff. The other nice thing I have it do is I often have it prompt me to say, how'd the day go? And so it's a very lightweight journal for me. And so usually I'll just do voice to text. Yeah, I'll just do voice to text real quick. And I'll say, these three things were awesome about my day.
27:36Wade Foster:These three things sucked about my day. And in the moment, that's not hugely valuable to me. But now I've been doing this for probably about six months now. And so I've got six months of data about what gets me pumped up during the day and what sort of makes me frustrated during the day. And so now I can do a whole bunch of interesting analysis on that and say there's some commonalities here that makes Zapier more successful, makes me more successful. And then there's some anti-patterns here where I find myself not being at my best or find the company not being at my best. And so how do I try and architect my day and architect my work to avoid some of those scenarios?
28:09Wade Foster:And so I actually find the daily recap to be just under discussed. And I think it's way more valuable than the daily prep.
28:17Matt Paige:No, I love that. And there's this memory component as well. Like when you're doing that reflection, it's learning over time. That's one thing for the, for lack of a better word, daily brief that I do is it's actually learning about the previous day. If I've slipped on something three times in a row, it knows, right? So there's this element of not repeating the same thing again, essentially over time. And so in terms of work, so we wrap it here, but the companies are obviously, they know the impact of AI. They know it's important. It's the top of everybody's strategy. But from somebody that literally called the code red at the beginning of GPT-4, what recommendations would you give to companies to actually drive this adoption internally?
29:00Matt Paige:Because it's not just about rolling out a bunch of licenses saying, hey, go do this thing. What advice would you give to companies other than start using Zapier to automate some things throughout your business? Of course, use Zapier for everything.
29:13Wade Foster:Yeah. No, I think there's, I would probably boil it down to a handful of things. And I would bucket it into two categories. I would say one I would call floor razors and the other category I'd call ceiling razors. Now the floor razors, this one's really important. This is about building widespread AI and fluency inside your whole organization. It's building a comfort level and excitement and energy for what is possible here. Now, the best floor raiser activities are generally things like hackathons, workshops, lunch and learns, any place where you're asking people to put their hands on the keyboard and actually learn and play around with this stuff.
29:52Wade Foster:What can they do that's practical in their day to day? We just talked about the daily recap and the daily brief. Great, go build one. Everyone should have something like that they can go feel, figure out. And whether you're an engineer or a marketer or a sales rep or an accountant or an HR, there's probably something in your day that you can use to build something that would just make your life better. And then you can start to build stuff that helps your team and all that sort of stuff. So these types of hackathons are just so powerful for doing this. Now, the other thing you get from this is you build a culture of an experimentation and you build a culture of excitement around what's possible with AI.
30:26Wade Foster:this becomes really important because you're going to have to go do some more interesting and harder things down the line. You might want to rethink like how the company works top to bottom. You might want to rethink job families and job structures. You might want to rethink just a lot of core concepts. And it's really important if you've built some of these, this floor raising culture internally, because now people are going to be a little more supportive of that because they're going to have a much more practical understanding of what's possible with these tools? What's not possible with these tools?
30:56Wade Foster:And so any other changes you're making in the organization will be less scary because it comes grounded in actually information and knowledge versus just people responding to the external headlines. I find the external headlines and what is actually happening in companies to be honestly pretty different in a lot of cases. So that to me is like the first category. Then the second category is the ceiling razors. Here, this one's important. It's how do you go from individual AI to institutional AI? And this one, I think a lot of companies are struggling with. Right now, they have figured out how to accelerate a lot of individuals inside their company.
31:33Wade Foster:We can all point to that engineer, that marketer, whoever, who's doing like crazy stuff with AI and is massively more productive. But it's a lot harder to go find companies where you can say, wow, this company is growing so much faster. They produce their cost basis. All this, the company is totally transformed with AI. Those are just a lot harder to find those examples. And I think a lot of this boils down to just because you accelerate one individual inside of a company, it doesn't mean that you've actually accelerated the whole company as a whole. You've just moved the bottleneck around. And so these institutional AI use cases are a lot more challenging.
32:07Wade Foster:You have to look around and try and break down what are the bottlenecks that prevent your company from growing? What is the thing that prevents you from getting another customer or making that customer happier or making that customer grow faster? And oftentimes those are things like, how do we actually truly accelerate time to market for a key product? How do we actually improve conversion rates in our sales and marketing funnel? How do we actually take waste and cost out of this business that's unnecessarily there? And so those, you really do have to rethink some of the core ways in which you go to market or you build your product or you operate internally.
32:41Wade Foster:And that requires a lot more tops down leadership and identification of the big levers inside of an organization. Much harder to do. much harder to do than the individual AI setup, but the rewards are much, much larger if you do accomplish this. And I would say there's very few companies who have really pushed the envelope on this. Even the companies you read about, I would say are still not as good as the things you've heard.
33:03Matt Paige:Yeah, it's so spot on. The individual productivity, but it's not spanning across the org. How do you think of the org of the future? You have Jack Dorsey talking about this hierarchy to intelligence. There's a lot of hypotheses of how the nature of an organization is going to evolve. Any thoughts or perspectives?
33:22Wade Foster:Yeah, I think Jack has got a lot of interesting ideas. The shift from individual to institutional AI is one of the things I'm really excited about. And I think the big bottleneck is you want to figure out how to increase the amount of time that your organization spends on those activities that are more likely to create a customer. It's how do you ship a product faster? It's how do you do better sales and marketing? And so you want to increase the time, energy, people that are spent on that. The challenge is as companies grow, you often have a lot of work that goes into stuff that's not that. You have a bunch of managers inside of a company.
33:58Wade Foster:You have a bunch of telephone games that sort of have to transport information from one team to another. If you've ever seen that movie Office Space, that's kind of what happens to companies as they grow. It's the TPS report that gets handed from one department to another. and a lot of that stuff doesn't really actually add value to the customer. It doesn't make them happier at the end of the day. It's just something that kind of has to get done because that was the best way to solve the problem inside that company today. So what I get really excited is how do you actually put AI at the center of how your company operates?
34:29Wade Foster:How do you build a brain so that the company actually has all of its knowledge institutionalized around it? How do you put automation and tools on top of that brain So now the AI can take action on those efforts. How do you put a governance layer around that so that it's really clear what humans can and cannot do in that system? And also it's clear what agents can and cannot do inside of that system. And then how do you put a harness around that so that there is an interface between the human layer and the AI layer inside the company to take action on all of that stuff? I think if you do those things really well, you can start to build a system inside your organization, that does take out a lot of that operational overhead where you can start to run closed loop meetings, where you can start to have the AI acting on your behalf in certain places.
35:18Wade Foster:And if you do that well, you can push more of your energy and effort to the things that humans are uniquely good at, which is this product development, which is the sales and marketing side of the house. Those things are stuff that distinctly bring value to customers. And that, I think, is when organizations get a lot more fun, but I think it's also when they get a lot more leverage from AI.
35:36Matt Paige:Yeah, it's so akin to system design, system architecture in a way. And last question for you, how do you think about your product roadmap? Things are evolving so quickly. And do you leverage AI on the strategy front, thinking through the product roadmap, what you're going to build next and how Zapier evolves? Has that changed at all for you over the years? Well, yeah.
35:58Wade Foster:You know, I think the thing that we have noticed is that the like six month roadmap is basically dead. What I see internally, what I hear from other companies is that you're having to be a lot more nimble on your toes about where you're heading and what is possible with these models. And so, you know, you're often thinking two, three weeks ahead, you know, at most. Now that said, I do think there are certain things that you're always trying to figure out, which are like, what are the foundational things that we do care about? And how do we sort of keep those at the center of where we're heading?
36:29Wade Foster:We know that this is the mission. We know that these capabilities are where we're heading. And so we're going to go hill climb against this and we'll climb against this for six months, 12 months, for however long as we think is important. But the way in which we get there, this sort of scoped out epics and stories and things like that, that stuff is kind of falling to the wayside in favor of these much more agile two-week roadmaps, or it's weird to even call them a roadmap. It's sort of just like, we just know we need to go achieve these things. Let's go figure it out.
36:58Matt Paige:Yeah. It's like there's tent poles in your strategy. To wrap us up, who are you following? Other CEOs, other leaders, other people in the space? Are there a few folks that you consistently look to for leaders in the world of either your world, AI in general?
37:14Wade Foster:Yeah, there's a lot of companies I think right now that are trying to be on the frontier. Obviously, you pay attention to the labs themselves. Those folks are really interesting to pay attention to. On the application layer, people who are maybe more closer to my seat, You look at the Brian at Coinbase and the Ramp crew, the folks over at DoorDash are doing really interesting things. So there's a whole bunch of folks that are trying to figure out, like, how do we operate differently with these things? What do we build that's different with these things that I think are really useful to pay attention to these companies and CEOs and whatnot?
37:45Matt Paige:Awesome. Wade, well, thank you for being on. Thank you for talking AI. Where can people go learn more about Zapier? And honestly, if you've used Zapier in the past and haven't used it in a while, like the product has evolved so much. I was just playing around with it recently. But where can people, you know, find you and learn more?
38:03Wade Foster:Yes. I mean, you can follow me on LinkedIn or Axe. I'm Wade Foster there. But, you know, if there is one thing I'd recommend that's different about Zapier is I would try and install it into your agent hardest of choice. So you might be using Claude or Chachukiti. You might be using Codex or Claude Code or whatnot. Now, to me, this is the best new interesting capabilities that we've added to Zapier. And so I think you'll have a really fun time playing around with all of the connections that Zapier has inside of these new tools. I think that is the one really exciting thing that you should check out with Zapier.
38:35Matt Paige:No, such a great call. It's like a new modality and way of working in a sense. Wade, thank you for being on.
38:40Wade Foster:Thanks for having me, Matt.
38:42Matt Paige:Thanks for listening to the Talking AI Podcast. If you enjoyed the show, give us a follow or subscribe on your favorite podcast podcast. And don't forget to leave us a review. We love those. For more info on Talking AI, visit TalkingAiPodcast.com. Quick break in the pod. If you're listening to this podcast, chances are you've been thinking about how to actually use AI inside your business. And that's exactly why we built the AI Opportunity Finder. It's a free tool that helps you uncover high-impact, tailored AI use cases based on your business, your goals, your pain points, and your industry. No fluff, no generic use cases, just real ideas that fit your business and they're ranked by ROI potential.
39:22Matt Paige:It takes about three minutes to run and it's like having your own personal AI strategist for free. If you want to try it for free, check out the link in the show notes or go to hatchworks.com backslash AI dash opportunity dash finder.
From the publisher
The best AI model in the world just scored 18.1%. On Zapier's own benchmark for real business work — the cross-app tasks any white-collar worker does every day — even the top frontier model completes them barely one time in five. That's the number Wade Foster keeps pointing at, and he runs an automation company that stands to gain from the hype. Instead, he makes the case for what actually works right now: not turning a model loose, but blending deterministic workflows with agents where each is strong.
In this episode of Talking AI, Matt Paige sits down with Wade Foster, co-founder and CEO of Zapier, who built a scrappy Y Combinator startup into the $5 billion plumbing of the SaaS era on barely a million dollars raised. Foster called a company-wide “code red” the week GPT-4 launched, and he's spent the years since rewiring how Zapier — and its customers — actually use AI.
The conversation covers why he shut the company down for a week in 2023, how AI habits actually stick, what Zapier's AutomationBench reveals about the gap between benchmark scores and real-world reliability, why coding models improve faster than knowledge-work models, how to tell a workflow from an agent, and the difference between individual AI and the institutional AI almost no company has cracked.
In this episode, you'll hear about:
- The three things about GPT-4 that triggered Zapier's first-ever code red
- How daily AI use jumped from 11% to over 50% in a single hackathon week
- The moves that make AI habits stick: show-and-tell, repeat hackathons, and “not yet”
- Why the best model on AutomationBench still scores only 18.1%
- Why coding is easy to verify — and subjective knowledge work isn't
- The power of hybrid setups that blend deterministic workflows with agents
- Wade's prediction: most tokens on open-source models, most spend on the frontier
- What actually makes a good eval — hard for models, easy for humans, private data
- A plain-English definition of an “agent” versus a deterministic workflow
- The daily recap workflow Wade thinks everyone is sleeping on
- Floor raisers vs. ceiling raisers — and why individual AI isn't enough
- Why the six-month product roadmap is dead
Key Moments
- 00:04:40 — Making AI habits stick: show-and-tell and repeat hackathons
- 00:06:38 — Differentiation when AI is best at the thing you sell
- 00:09:34 — AutomationBench: the best model scores just 18.1%
- 00:11:31 — Why the top model stalls: verifiable code vs. subjective work
- 00:14:19 — Getting squeezed on both sides: AI in the company and the product
- 00:15:20 — Model efficiency, Coinbase, and the token-maxing debate
- 00:17:18 — What makes a good eval
- 00:19:30 — What actually counts as an “agent”
- 00:23:12 — Iterating on workflows with your own mini-evals
- 00:26:15 — The kind of worker thriving right now
- 00:27:36 — Wade's favorite workflow: the daily recap
- 00:30:44 — Floor raisers vs. ceiling raisers for AI adoption
- 00:34:55 — From individual AI to institutional AI
- 00:37:58 — Why the six-month roadmap is dead
Key Links:
Mentioned in this episode:
AI Opportunity Finder
Feeling overwhelmed by all the AI noise out there? The AI Opportunity Finder from HatchWorks cuts through the hype and gives you a clear starting point. In less than 5 minutes, you’ll get tailored, high-impact AI use cases specific to your business—scored by ROI so you know exactly where to start. Whether you're looking to cut costs, automate tasks, or grow faster, this free tool gives you a personalized roadmap built for action. 👉 Try it now at https://hatchworks.com/ai-opportunity-finder/
